cross-posted from: https://kbin.earth/m/piracy@lemmy.dbzer0.com/t/3120649

As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.

  • melsaskca@lemmy.ca
    link
    fedilink
    English
    arrow-up
    31
    ·
    6 days ago

    I guess the next step is picking specific people to memorize specific books and pass that knowledge down to an assistant/acolyte who will keep the memories alive. Rinse and repeat. Fahrenheit 451 anybody?

    • matlag@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      15
      ·
      6 days ago

      It’s worse than that. I’m sure they’re getting wet dreams about eradicating as many books as possible so that they can finally sell AI slop books at a premium price, claiming that will be “the last way to get something in [author] style!”.

      • BlindPenguin@lemmy.world
        link
        fedilink
        English
        arrow-up
        14
        ·
        6 days ago

        I’m rather thinking that hey want to make their slop machines the single source of truth. Much easier to get control over information, when all information is located in your own house.

  • TrackinDaKraken@lemmy.world
    link
    fedilink
    English
    arrow-up
    65
    arrow-down
    1
    ·
    7 days ago

    Anna’s Archive says ‘time is running out’ as ‘knowledge is permanently monopolized on private servers’

    Not only that, of course, but easily edited. It’s a “mandala effect” generator. AKA modern gaslighting. They name the thing, point to legitimate examples, then abuse our new “understanding” to cover their crimes. 1984 indeed.

    Oceania was at war with Eurasia; therefore Oceania had always been at war with Eurasia

  • TrackinDaKraken@lemmy.world
    link
    fedilink
    English
    arrow-up
    17
    ·
    7 days ago

    If you read the blog post, they don’t provide help or guidance for regular bookworms on how to scan. No information about what they expect.

    Also, flatbed scanning a whole book takes many hours. It’s one page at a time, line up each page, press the book flat and scan, for hundreds of pages.

    You’d likely need to buy, or have access to a book scanner. Which is fine, there are cheap models, but who knows if that would work for them? They don’t say. It doesn’t seem like they’re expecting help from you or me.

    • luciferofastora@feddit.org
      link
      fedilink
      English
      arrow-up
      21
      ·
      7 days ago

      Eh, perfect is the enemy of good. I’ll take the wins we can get. It’s not like having to buy the books is gonna stop the biggest offenders from burning investor money (along with the books).

      The defeat of LLMs isn’t gonna come on that front. Let the cultists believe that more words will somehow become more than words. Their silicon gods run on power and water. That’s where they’re most vulnerable.

      • LavaPlanet@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        10
        ·
        7 days ago

        It’s not necessarily trying to stop them burning the books, just stop their ability to monopolise the knowledge, that’s their aim, lock the knowledge behind a subscription, Anna’s archive will always be available for all, so why pay for their subscription if its available for free.

        • luciferofastora@feddit.org
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 days ago

          I was aiming for the poetic parallel between burning money and destroying books, but also, isn’t the difference between destroying the primary source and preventing access to it mostly academic? Either way, as you say, the AI grifters are looking to control that information. Once they do, the original text and knowledge might as well be lost, since we couldn’t verify anything without the original to compare to. Whether through deliberate human manipulation or random word soup mutation, they can’t ever be trusted to deliver a genuine, faithful reproduction of the contents.

          And the conclusion is the same: Better to have it freely available for everyone (including the grifters) than the grifters only.

    • qyron@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      10
      ·
      7 days ago

      I am all up for starving machine learning of new knowledge but I view this as the least friction path. Otherwise, if the machines try to brute force their entry, what harm can kt cause?

        • quakki@discuss.tchncs.de
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          2
          ·
          5 days ago

          Okay I’ll give you a hint. The bad thing about the nazis burning books wasn’t that they were destroying physical copies of books. I’ll give you a few keywords intent, selection, political control

          • 5318008@lemmy.ml
            link
            fedilink
            English
            arrow-up
            1
            ·
            5 days ago

            Why be condescending now?

            You already fucked up the point you were making. By all means, insult people’s intelligence at the start of the conversation, before you fuck up. Doing it at the same time as fucking up looks bad… but insults after you fuck up looks even worse.

    • tb_@lemmy.world
      link
      fedilink
      English
      arrow-up
      20
      ·
      7 days ago

      Those who want to conduct large-scale scans and upload many titles can reach out to the archive for support, as the archive says that it “can help pay for the scanning fees and other rewards.”

      • vala@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        7
        ·
        7 days ago

        Hmm, it’s just a few books but I do think they’re pretty rare. Not something I’m willing to send away. I wonder if a library would be willing to help?

        • njordomir@lemmy.world
          link
          fedilink
          English
          arrow-up
          13
          ·
          7 days ago

          My local library has some large format scanners. I’ve considered using them to do my photo albums. I have the individual photos scanned and backed up, but the layout and margin notes are worth preserving also.

    • Catoblepas@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      2
      ·
      6 days ago

      If your local library doesn’t own a book scanner, any colleges near you may have one. You usually don’t have to be a student to walk into the library, and the book scanners I’ve used don’t require a login. YMMV on whether this is how it’s set up where you are.

      • vala@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        3
        ·
        6 days ago

        I might try this. Im concerned that some of the books may still be under copyright. Not all of my rare books are that old. I collect indie poetry books and comics as well as old religious texts/pamphlets.

        Also concerned that some might be literally some of the only remaining copies and I really don’t want to risk damaging them.

  • Tollana1234567@lemmy.today
    link
    fedilink
    English
    arrow-up
    4
    ·
    5 days ago

    the AI ran out of Reddit slop material to scan, because reddit has been purging alot of content/accounts even suspected of spamming, so they go for physical books now.

        • mojofrododojo@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          5 days ago

          that’s what I’m saying. I’ve never understood the hoopla of reddit selling it’s corpus, aside from being fucking gross (spot on for spez) it’s a shitton of garbage. there’s gold in there, but it’s absolutely buried in shite

          • 3abas@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            5 days ago

            For an AI to troubleshoot problems with your computer (something it’s really good at today), it needs content like reddit’s. Books don’t and won’t detail weird software issues and their workarounds. One the current training set becomes irrelevant to whatever software people are running, this is a specific area where Reddit’s data is way more valuable than books.

            Of course Reddit locking down access to that information leads to fewer people contributing. Reddit is dead, it’s just struggling to go down.

            • mojofrododojo@lemmy.world
              link
              fedilink
              English
              arrow-up
              2
              ·
              5 days ago

              there’s gold in there, but it’s absolutely buried in shite

              yep, there is a lot of nuanced depth, but measured against the overall sea of chuds lol… and yep, the moment they started putting up walls and controlling the community it was doomed. it’s gonna thrash for a while then, when the only ones left are advertisers and bots, it’ll shit itself and stop twitching.

          • Tollana1234567@lemmy.today
            link
            fedilink
            English
            arrow-up
            2
            ·
            5 days ago

            im guessing reddit has been banning so aggressively lately, that AI scraping isnt getting as much “useful data” as much anymore (chatgtp/google dropped "references for reddit in thier slop summaries). so they are trying to force logins now to keep up with the user generated content so they can squeeze every last penny they can from reddit, i dont think ads were ever a big presence in reddit? its the data. people try to post content, like posts, subreddits, those can instantly get removed by redit filters. i think lost alot of traffic recently to, which was likely AI scraping.

  • AlecSadler@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    10
    ·
    7 days ago

    Any reputable orgs I can donate to for assisting the cause? I just don’t feel I have the means otherwise to make a dent in the buying and preservation.

    • late_pessimistic@slrpnk.net
      link
      fedilink
      English
      arrow-up
      1
      ·
      7 days ago

      Perhaps there is a public library in your area that accepts donations? The one I live by is STARVING for support, they accept books, old computers, anything.

      They do good work for the community too, like free tech literacy classes, chess for kids, etc.

      But if you are dead set on an online NGO, then maybe archive.org? You have probably heard about their Wayback Machine, but they preserve books and other media too.

  • WakeUpSmashPots@lemmy.today
    link
    fedilink
    English
    arrow-up
    2
    ·
    5 days ago

    I hate this timeline… It’s things like this that make me hope we actually do live in a simulation, and the devs are about to release a major bug fix…

  • tigermountain@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    arrow-down
    2
    ·
    7 days ago

    So, my question is, of all these books being scanned and then destroyed, aren’t there many more copies of these books? And BTW, knowledge isn’t permanent. It can be lost.

      • tigermountain@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        3
        ·
        7 days ago

        When someone can’t answer a question a common response to change the subject. Do you have an answer to my question?

        • jnod4@lemmy.ca
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          7 days ago

          They just cut and bin the books after they’ve been scanned, the bins are taken to the local disposal, most municipalities have a garbage dump, some burn the trash.

          An ai trained on unique data is more valuable if you don’t share the data it was trained on to your competition, your LLM gains an edge.

          If I found a unique handwritten book from the 18th century, than I proceed to read it and memorise then burn it, I’ll be the only one to posses the knowledge of it.

          • tigermountain@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            2
            ·
            7 days ago

            There are very few, if any, rare, one of a kind books that would go through this process. If one of these books were scanned for an LLM it would be done delicately in a more time consuming manner and the book would be returned.

  • baconsunday@lemmy.zip
    link
    fedilink
    English
    arrow-up
    3
    ·
    7 days ago

    With modern phones having a document scan option with our phone cameras, is it possible we can scan and combine into a PDF file?

    I am going to test it myself, see how easy it is, and find a way to ensure anyone can follow with low cognitive load, but hopefully it’s as easy as I think to do.

    The sad part, is to do so would require a slow tedious task for books with hundreds of pages. It would be nice to see a large group tackle chunks at a time to the point we eventually have groups dedicated to genres and scouring to find them to preserve them.

    • Eheran@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      5 days ago

      Which app is good for this task? Last time I checked the result was garbage, as in not really better than simply the picture itself. But that was years ago…