Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.
Nobody complained when Google did this over a decade ago 🤷
When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.
Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.
Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)
Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.
It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.
I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).
This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.
Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.
Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted
Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.
If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
Almost all these books were either headed to the dump or the recycling center.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
I imagine they were talking about destruction in a practical sense. Disassembling the book and scanning it like that is faster and cheaper than purpose build book scanning machines.
Additionally, court documents indicate they generally just throw them away afterwards. So the knowledge is retained, but that book is destroyed.
Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.
Nobody complained when Google did this over a decade ago 🤷
When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.
Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.
Digitized for private consumption.
Nobody cares when Google did this or when archive.org does this, because they’re sharing the results with the world. (Idiotic shortsighted lawsuits from the authors guild notwithstanding)
in this case they often are actually destroying them. they take them apart, beause it’s easier to scan than to use a proper book scanner
Yeah that’s true of Google as well I believe. I think Archive is more careful, but I’m sure some books get damaged in the process there as well.
Having done a lot of scanning in the past, no matter how gentle you are, the book will be damaged; at least a little bit. Even if you do it manually, by hand, really slowly.
It’s the unfortunate reality, because books are not designed to be held open and pressed flat onto an unyielding surface. If you don’t press flat, you’ll get curved pages and bad quality (dark) images, which is sometimes ok and fixable by software, but often not. You can get really fancy expensive machines, but they still don’t fully solve the problem.
I did scanlation editing stuff for a while, and half my job was making the raw page scans look presentable, because the scanners didn’t want to destroy their manga (understandable).
https://en.wikipedia.org/wiki/Book_scanning#Methods
This matches my experience. My college archive digitizes a few yearbooks a year (for whatever classes are having as big reunion) and it’s a pain. For modern yearbooks (70s and later) we have a dozen or more copies, so we take one apart and feed it through the feed scanner, which works well. We then store the loose pages in a folder, in case we need to rescan anything.
Older yearbooks (where we don’t have so many copies) we use the flatbed scanner and the pages come out warped, so we hand fix each page in Photoshop.
Yepp… memories of using the light level curve tool and white point and black point and the perspective tool. Then going in with the white brush and deleting any specks of dust that remained, and comparing to the original to make sure no lines got deleted
Google (and/or the libraries that collaborate with Google Books) and Archive.org usually scan non-destructively. Some of Google’s book scans still show the fingers holding the corners of the pages, some of them were taken mistakenly mid-flip, etc. So they clearly used whole, normally bound books.
When they destroy the physical copy they remove them from the antique stores market which often rely on circulation.
Edit: Also physical copies don’t require electricity, a device and internet access plus they are something you can own.
If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
It’s ok to throw trash away! Really!
At least digital storage prices have only been going down in recent months, right? /s
Good luck reading by candle light.
I guess you must live in that part of Alaska where it’s night time for half the year and don’t have daylight.
More just the part of Alaska where I work most of the daylight hours and get my reading in before bed.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
I imagine they were talking about destruction in a practical sense. Disassembling the book and scanning it like that is faster and cheaper than purpose build book scanning machines.
Additionally, court documents indicate they generally just throw them away afterwards. So the knowledge is retained, but that book is destroyed.
https://tagteam.harvard.edu/hub_feeds/3415/feed_items/14341220
Yes they did?! That was a big controversy back in the day. You are engaging in historical revisionism right now.
Where does it say that in this article?
ship of theseus ah shit