• minus@lemmy.zip
    link
    fedilink
    English
    arrow-up
    114
    arrow-down
    4
    ·
    4 days ago

    The worst part is that they are destroying the books afterwards.

    • crusa187@lemmy.ml
      link
      fedilink
      English
      arrow-up
      60
      arrow-down
      2
      ·
      4 days ago

      We also have no guarantee these books will be available at a later date in these digitized formats, nor that they won’t be altered to fit certain narratives which is of course trivial to do with digital copies. At a minimum these need public verifiable checksums to be relied upon.

      I don’t like this and don’t trust it one bit.

      These AI companies already showed their hand in trying to corner the PC component market with an end goal of users no longer being able to own their own PCs. Now they’re potentially trying to do this with all written recorded information.

      How long until Altman’s Firemen show up at your door to burn your books?

      • morto@piefed.social
        link
        fedilink
        English
        arrow-up
        25
        arrow-down
        1
        ·
        4 days ago

        They even hve an economical incentive to destroy the books, because by doing so, they will get the data and prevent the competition from getting it for them

        • BlaestEgnen@feddit.dk
          link
          fedilink
          English
          arrow-up
          4
          ·
          4 days ago

          The incentive is literally that it’s cheaper to scan, that’s all.

          But given they’re only interested in books with ISBN’s, I recon they’re already scanned and available in digital forms. If they’re newer than the 80s, they’re most likely in a document file in a server related to the publisher.

          Honestly surprised they’d rather source physical books, than just request various publishers for a price for their combined works - Only reason I can think of, is the physical scan makes other scans like a mobile photo easier to recognise and respond correctly to

    • AbouBenAdhem@lemmy.world
      link
      fedilink
      English
      arrow-up
      13
      ·
      4 days ago

      Reminds me of Judge Holden in Blood Meridian, who sketched ancient artifacts in his personal notebook and then destroyed the originals.

      • Omgpwnies@lemmy.world
        link
        fedilink
        English
        arrow-up
        11
        arrow-down
        1
        ·
        4 days ago

        More that they are destroying the books in the process of digitizing them, they basically cut the pages out of the book and feed them into a scanner with a document feeder attached.

        • T156@lemmy.world
          link
          fedilink
          English
          arrow-up
          17
          ·
          4 days ago

          It’s worth pointing out that this is standard practice when digitising books in most places, if they’re not irreplaceable. Mass-produced books would fall under that category.

          Turning the page and photographing what’s there is really only done for some books, which you can’t afford to destroy.

          • Omgpwnies@lemmy.world
            link
            fedilink
            English
            arrow-up
            7
            arrow-down
            2
            ·
            4 days ago

            From what I’ve read previously, they’re destroying all the books, regardless if they’re historic and one-of-a-kind or a mass produced paperback.

            • T156@lemmy.world
              link
              fedilink
              English
              arrow-up
              6
              ·
              4 days ago

              At least in this case, yes. They’d just not get them, since they’d want something they can already easily process, and the source for the quote is a bookseller, who isn’t likely to sell something like that to begin with.

              An old book that might disintegrate and damage the scanner and hold things up would be undesirable, compared to either a new one, or a reprint of an old one.

              • cavitationfetishist2@quokk.au
                link
                fedilink
                English
                arrow-up
                1
                arrow-down
                2
                ·
                4 days ago

                That’s not how most of that works.

                Things would be sold in bulk. Mistakes happen. Literally the only part of what you said that makes any sense is that fucked up pages could clog the pipe.

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      15
      arrow-down
      35
      ·
      4 days ago

      Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.

      Nobody complained when Google did this over a decade ago 🤷

      When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.

      Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.

      • Hawke@lemmy.world
        link
        fedilink
        English
        arrow-up
        55
        arrow-down
        1
        ·
        4 days ago

        Digitized for private consumption.

        Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)

        • SkaveRat@discuss.tchncs.de
          link
          fedilink
          English
          arrow-up
          11
          ·
          4 days ago

          in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner

          • Hawke@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            4 days ago

            Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.

            • Zarobi@aussie.zone
              link
              fedilink
              English
              arrow-up
              4
              ·
              4 days ago

              Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.

              It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.

              I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).

              https://en.wikipedia.org/wiki/Book_scanning#Methods

              • smh@slrpnk.net
                link
                fedilink
                English
                arrow-up
                3
                ·
                4 days ago

                This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.

                Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.

                • Zarobi@aussie.zone
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  4 days ago

                  Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted

        • antonim@lemmy.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          4 days ago

          Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.

      • minus@lemmy.zip
        link
        fedilink
        English
        arrow-up
        23
        arrow-down
        1
        ·
        4 days ago

        When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.

        Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          6
          arrow-down
          1
          ·
          4 days ago

          If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.

          Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.

          It’s ok to throw trash away! Really!

        • XLE@piefed.social
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          4 days ago

          At least digital storage prices have only been going down in recent months, right? /s

          • xthexder@l.sw0.com
            link
            fedilink
            English
            arrow-up
            4
            arrow-down
            1
            ·
            4 days ago

            I guess you must live in that part of Alaska where it’s night time for half the year and don’t have daylight.

      • UnderpantsWeevil@lemmy.world
        link
        fedilink
        English
        arrow-up
        14
        arrow-down
        1
        ·
        4 days ago

        Almost all these books were either headed to the dump or the recycling center.

        My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.

        It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.

        Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.

      • Jomega@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 days ago

        Nobody complained when Google did this over a decade ago 🤷

        Yes they did?! That was a big controversy back in the day. You are engaging in historical revisionism right now.

      • zarkanian@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        2
        ·
        3 days ago

        Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.

        Where does it say that in this article?