GitCode, a git-hosting website operated Chongqing Open-Source Co-Creation Technology Co Ltd and with technical support from CSDN and Huawei Cloud.
It is being reported that many users’ repository are being cloned and re-hosted on GitCode without explicit authorization.
There is also a thread on Ycombinator (archived link)
Solution: create a GitHub repo with Markdown articles outlining human rights abuses by the CCP and have a large number of GitHub users star and fork the repo.
create a GitHub repo with Markdown articles outlining human rights abuses by the CCP
Once you have logged “China killed 100 Zillion people! End CCP now!” in Chinese GitHub, everyone in China will realize that their lives are actually very bad and they need to do a Revolution immediately.
fun to think that my shitty program is now stored in an artic vault and stored in some Chinese servers
So many bugs I never fixed and yet here we are lol
The great thing about China is that it’s got lots of people eager to fix those bugs
You haven’t worked with a lot of Chinese engineers, have you? https://www.chinaexpatsociety.com/culture/the-chabuduo-mindset
I love how they reinvent a universal experience as uniquely Chinese
The vast majority of projects on GitHub is open-source and forkable, why would that need authorization?
It’s… suspicious that China’s doing it en masse, but there’s nothing wrong in cloning or forking a repo last i heard.
It’s not about authorization. They want to build a knowledge base for when the Great Firewall gets some more filters. Just like russias mirror of wikipedia which is heavily edited to discredit the west.
It’s a bit odd, but isn’t it equivalent to forking and putting up a fork elsewhere?
I guess I don’t see the problem.
It will be funny to see folks who spent the last ten years posting “It’s not stealing, it’s copying” memes suddenly find religion because Evil Foreign People got involved.
I’m quite scared of how AI apparently pushes people in favour of significantly stricter copyrights. This is not a good trend.
This isn’t people being influenced by AI. This is Microsoft’s Godzilla battling the RIAA/MPAA’s King Kong.
The trend, to date, has been consolidation of media properties under fewer and more hegemonic distributors. And now we’re seeing a couple of economic Titans battle over the position of “Last Legitimate Music Vendor”.
Some random Chinese company: does something jenky
Blogger: “The entire country of China is doing this jenky thing!”
With the obligatory “fuck everyone who disregards open source licenses”, I am still slightly amused at this raising eyebrows while nearly no one is complaining about MS using github to train their copilot LLM, which will help circumvent licenses & copyrights by the bazillion.
If I look at a few implementations of an algorithm and then implement my own using those as inspiration, am I breaking copyright law and circumventing licenses?
That depends on how similar your resulting algorithm is to the sources you were “inspired” by. You’re probably fine if you’re not copying verbatim and your code just ends up looking similar because that’s how solutions are generally structured, but there absolutely are limits there.
If you’re trying to rewrite something into another license, you’ll need to be a lot more careful.
What’s the limit? This needs to be absolutely explicit and easy to understand because this is what LLMs are doing. They take hundreds of thousands of similar algorithms and they create an amalgamation of it.
When is it copying and when it is “inspiration”? What’s the line between learning and copying?
I disagree that it needs to be explicit. The current law is the fair use doctrine, which generally has more to do with the intended use than specific amounts of the text/media. The point is that humans should know where that limit is and when they’ve crossed it, with motive being a huge part of it.
I think machines and algorithms should have to abide by a much narrower understanding of “fair use” because they don’t have motive or the ability to Intuit when they’ve crossed the line. So scraping copyrighted works to produce an LLM should probably generally be illegal, imo.
That said, our current copyright system is busted and desperately needs reform. We should be limiting copyright to 14 years (as in the original copyright act of 1790), with an option to explicitly extend for another 14 years. That way LLMs can scrape comment published >28 years ago with no concerns, and most content produced >14 years (esp. forums and social media where copyright extension is incredibly unlikely). That would be reasonable IMO and sidestep most of the issues people have with LLMs.
I love how every Chinese company is called “China”
Yeah… That’s because they are. it is required that every Chinese Company has to be owned by “tHe PeOpLe” (CCP)
Gotta disguise being a capitalist country somehow.
Capitalism is when the public owns the things. And Communism is when a handful of private individuals owns the things.
So what’s it called when the government is a handful of private individuals, as opposed to representing the public?
The Aristocrats