Digital preservation of books explained: how libraries fight decay to keep literature readable for future readers. See real methods and examples now.

How Digital Libraries Preserve Literature
A book you can't open is just data. That's the blunt truth behind digital preservation of books: the practice of keeping scanned and born-digital texts readable, intact, and accessible decades after they were first digitized, not just stored once and forgotten.
Libraries have quietly been racing against file-format decay, funding gaps, and plain old bit rot for two decades now. Some are winning. Others are losing whole collections without anyone noticing, until a researcher tries to open a file that simply will not load anymore. The stakes go well beyond convenience: entire regional literary traditions, rare manuscripts, and out-of-print titles exist today only because someone digitized and then kept maintaining them. This piece walks through how digital libraries actually pull preservation off, where they are falling short in 2026, and why that gap matters if you have ever tried to track down an out-of-print book online. Along the way, you will see who is doing this well, who is struggling, and what a 400-year-old manuscript has in common with a corrupted file from 2015.
Key Takeaways
- Digital preservation of books means actively maintaining files, formats, and metadata for decades, not a one-time scan.
- Libraries mainly use two techniques: migration (updating file formats on a schedule) and emulation (recreating old software environments).
- A quarter of all webpages online between 2013 and 2023 were unreachable by October 2023, and digitized books face the same decay risk.
- Fresh 2026 developments include AI-OCR breakthroughs for historic manuscripts and a new archival e-book format heading toward an international standard.
- HathiTrust alone holds more than 17.6 million digitized volumes; Rekhta Foundation has separately preserved over 54 million pages of Urdu literature.
- Preservation stays chronically underfunded, and format obsolescence remains the single biggest technical threat to digitized literature.
What Is Digital Preservation of Books?
Digital preservation of books is the ongoing work of keeping a digitized or born-digital text usable, readable, and authentic well beyond the moment it was first created or scanned. It isn't the same as a PDF sitting quietly on a server somewhere.
Digitization gets a book into digital form. Preservation is everything that happens after: monitoring file formats, refreshing storage media, tracking metadata, and rescuing content before the software needed to open it disappears. A widely cited definition from library scholar Priscilla Caplan draws a sharp national line around the term, as summarized in the American Library Association's preservation guide: in the United States, the phrase typically covers the full life-cycle management of a digital object from the moment it is created, while in the United Kingdom the same work gets split, with "digital curation" describing ongoing management and "digital preservation" reserved specifically for long-term accessibility efforts.
That distinction matters more than it sounds. A library that only digitizes without preserving is building a house with no plan to maintain the roof.
Eventually, something breaks.
So what is a digital library, exactly, once you strip away the jargon? It is a managed collection of digitized or born-digital texts, made searchable and accessible online, backed by an institution that takes responsibility for keeping it that way. The "managed" and "responsible" parts are what separate a real digital library from a loose folder of scanned files sitting on someone's personal server. How digital libraries work, in practice, comes down to that ongoing institutional commitment as much as any specific piece of software.
How Are Old Books Digitized?
Old books get digitized through a mix of careful physical handling and increasingly capable software, not a single button press. The process usually runs through scanning, image cleanup, character recognition, and metadata tagging, roughly in that order, and each stage can quietly ruin the final result if it is rushed.
Scanning and Physical Handling
Scanning comes first, and it is slower than most people assume. Fragile bindings, brittle paper, and handwritten marginalia all demand different handling than a stack of modern paperbacks. Conservators often have to decide, page by page, whether a book can survive a flatbed scanner or whether it needs a gentler overhead camera cradle that never presses down on the spine.
Lighting matters more than people expect too. Too much heat from old scanning lamps can accelerate the very decay a library is trying to stop, which is part of why archive-grade scanning setups use cool LED arrays instead of the older bulbs still common in consumer scanners.
Optical Character Recognition
Once the pages are photographed, optical character recognition (OCR) software converts the page images into searchable, selectable text. This is where things get interesting in 2026.
OCR has historically struggled with old scripts, tight handwriting, and languages that do not use Latin characters. In April 2026, Japan's TOPPAN Holdings announced an AI-OCR engine purpose-built to decipher medieval Greek script, developed in partnership with the Vatican Apostolic Library and aimed at pushing recognition accuracy above 95 percent, on a script that older OCR tools have historically struggled with. The same underlying shift, AI models trained specifically on historical handwriting rather than generic printed text, is showing up across manuscript projects well beyond Greek and Latin scripts, including early efforts aimed at South Asian and Middle Eastern manuscript traditions.
Metadata and Cataloging
After OCR comes metadata: author, date, subject headings, language, and physical condition notes. Skip this step, and good luck finding that book again. It ends up buried in a repository with no way to search it.
This is also the stage where a library decides how much context future readers will have. A scanned page with no metadata is a picture. A scanned page with careful metadata is a discoverable, citable historical record, and the difference between those two outcomes is almost entirely a matter of staff time and institutional priority.
Quality control closes out the process, though it rarely gets much attention outside the library itself. Someone has to spot-check pages for missing scans, garbled OCR output, and mismatched metadata before a digitized book goes live, and skipping this step tends to surface years later, when a researcher finally notices a chapter is missing from an otherwise complete digital edition.
![]() |
| Inside a library conservation lab, an overhead scanning cradle carefully captures every page of a fragile antique book. |
Migration vs. Emulation: The Two Core Preservation Strategies
Two techniques do almost all the heavy lifting in digital preservation strategies: migration and emulation. Pick the wrong one for the wrong type of file, and an institution either loses data or wastes years of staff time.
Migration means periodically converting a file into a current, actively supported format before the old one goes extinct. The U.S. National Archives runs its entire preservation strategy this way, transforming files into a managed set of formats over time while keeping original versions in low-access cold storage as a safety net, and documenting every action against the ISO 16363 standard for trustworthy digital repositories.
Emulation takes the opposite approach. Instead of changing the file, it recreates the old software environment the file was built for. A 2025 CLIR report from the Software Preservation Network lays out the method in detail: rather than converting a 1990s word-processing document into a modern format and risking lost formatting, an emulator can run a virtual version of the original operating system, letting the file open exactly as it once did.
Here is how the two compare side by side:
| Strategy | How It Works | Best For | Example in Practice |
|---|---|---|---|
| Migration | Convert the file into a new, current format on a set schedule | Simple, stable file types like plain text and standard images | U.S. National Archives (NARA) |
| Emulation | Recreate the original software and hardware environment | Complex, software-dependent, or interactive digital objects | Emulated legacy operating systems, per CLIR/SPN (2025) |
Most large institutions use both, switching strategy based on the file type in front of them. A plain text file migrates easily. A piece of interactive software often cannot survive migration at all, so emulation becomes the only real option.
The decision usually comes down to one practical question: does converting this file lose anything meaningful? A scanned book page as a TIFF or JPEG converts to newer image formats without much loss, so migration wins almost every time. A digitized book that includes embedded scripts, unusual fonts, or interactive annotations is a different story, and that is where emulation, despite being more resource-intensive to maintain, earns its keep.
![]() |
| Migration means converting old files into new formats before the old ones become unreadable — here is what that looks like. |
Who Is Responsible for Digital Preservation?
No single organization owns digital preservation. It is handled by a patchwork of national libraries, university consortia, nonprofits, and, occasionally, commercial platforms, each covering a different slice of the problem.
National libraries carry the heaviest formal mandate. In the United States, the Library of Congress runs digital preservation work across multiple internal units, maintaining a public reference of over 440 distinct digital file formats so institutions everywhere can judge which ones are safe to rely on long-term. Its collections policy also follows established technical frameworks, including the PREMIS metadata standard and the OAIS reference model, the same ISO 14721 framework that shows up across much of the preservation field. Leadership matters here too: the Library appointed a dedicated Associate Librarian for Discovery and Preservation Services in 2022 specifically to oversee this work at scale, overseeing hundreds of staff across acquisitions, cataloging, and preservation.
University consortia handle a different layer of the problem, pooling resources that no single campus library could justify alone. HathiTrust, discussed in more detail below, is the clearest example of this model working at scale.
Nonprofits fill gaps that neither government mandates nor university budgets reliably cover, especially for material outside the English-language mainstream. The Digital Preservation Coalition and Rekhta Foundation, covered later in this piece, both operate in that space, one focused on advocacy and risk-tracking across the whole field, the other focused on one language's literary heritage specifically.
Commercial platforms round out the picture, usually motivated by search, discovery, or product goals rather than a formal preservation mandate. That distinction matters: a commercial scan can sit dormant behind legal restrictions for years, which is exactly what happened to a large share of Google's book-scanning output after litigation slowed the project down.
Put those four categories side by side and a pattern emerges: the institutions with the clearest legal mandate are not always the ones moving fastest, and the ones moving fastest are not always the ones with permanent funding. Readers rarely see this layer of the system. They just see whether a book is available online, with no visibility into which of these four types of organization made that possible, or how precarious that access might actually be.
Real-World Examples: Institutions Leading the Way
Some institutions are further along in this work than others, and the gap shows up fast once you start comparing collections. The three examples below were chosen deliberately: one for raw scale, one for a literary tradition that rarely gets international attention, and one for the kind of cutting-edge technical work that will likely define the next few years of this field.
HathiTrust is the scale example. Built by a consortium of research libraries starting in 2008, its own published statistics put current holdings at more than 17.6 million volumes, spanning upward of 8.4 million distinct book titles. For context, Google's separate Library Project has scanned more than 40 million books across over 500 languages since 2004, according to Google's own account, though legal disputes have kept much of that specific archive from full public access.
Scale is not the only story worth telling. Rekhta Foundation, a nonprofit devoted to Urdu language and literature, has spent over a decade building what may be the most overlooked digital preservation success story in South Asia. The Foundation's own figures put its digital library at more than 322,000 e-books and over 54 million preserved pages, drawn from partnerships with 35 libraries and reaching more than 30 million readers a month across Urdu, Devanagari, and Roman scripts. Much of that material, centuries-old Urdu poetry and prose scattered across private and institutional collections, had no digital footprint at all before this project began. That kind of language-specific effort rarely gets funded by the large multinational platforms, which tend to prioritize whichever languages already have the biggest existing digital footprint.
The TOPPAN and Vatican Apostolic Library partnership mentioned earlier is worth another look here, because it shows preservation and public engagement feeding each other. The resulting AI-OCR findings were built into a public exhibition at the Printing Museum in Tokyo, turning what could have been a purely back-office technical upgrade into something visitors could actually see and understand. That's a pattern worth watching: preservation work that stays entirely behind the scenes tends to struggle for funding, while preservation work tied to something the public can engage with tends to attract it.
![]() |
| A centuries-old Urdu manuscript — the kind of literary heritage nonprofits like Rekhta Foundation are racing to digitize. |
Picture a graduate student in Lahore working on a thesis about a 400-year-old Persian dastan tradition. Twenty years ago, she would have needed months of travel and special archive access just to read the source material. Today, if that manuscript has been digitized and preserved by an institution like Rekhta or a partner library, she can pull it up on a laptop the same afternoon she needs it, no travel required.
Why Does Digital Preservation Matter for Readers and Researchers?
Digital preservation matters because access without permanence is fragile access. A book you can read today but will not be able to open in ten years is not really preserved. It is just temporarily convenient.
The importance of digital preservation shows up most clearly in what it prevents: the quiet, permanent loss of material nobody thought to back up in time. The core advantages of a well-preserved digital library over an unmanaged one are speed, reach, and redundancy. A single well-preserved digital copy, mirrored across a few institutions, is harder to lose entirely than a single physical copy sitting on one shelf in one building. The benefits compound over time too, since every additional year a text survives is another year it stays available to whoever needs it next.
Libraries have historically treated preservation and access as two separate jobs, and the balance has shifted hard toward access in recent decades. A March 2026 piece from the Carnegie Corporation points out that this shift, driven by the promise of digital tools to broaden reach, has quietly sidelined preservation work and its budget line at many institutions, even as the volume of born-digital material keeps climbing. The role of libraries in preserving knowledge does not stop the moment a page gets scanned; it continues for as long as the institution exists.
That tradeoff has real consequences. Rare or minority-language literature is especially vulnerable, since it typically has fewer surviving physical copies and less commercial incentive behind digitizing it. UNESCO's PERSIST initiative, run through its Memory of the World programme alongside IFLA and other heritage bodies, exists specifically to help institutions decide which pieces of a digital archive of cultural heritage deserve priority before they are lost for good.
For researchers, students, and casual readers alike, the payoff of doing this well is straightforward: a text that survives past the current decade, the current file format, and the current company that happens to host it.
There is also a quieter benefit that rarely makes it into funding pitches: searchability. A well-preserved digital text is not just saved, it is findable, cross-referenceable, and analyzable in ways a physical copy locked in a single reading room never was. A historian researching a specific phrase across thousands of texts can now do in an afternoon what once took years of manual archive visits. That shift changes what kinds of research questions are even possible to ask.
Digital Libraries vs. Traditional Libraries: What's the Difference?
A digital library, at its simplest, is a collection of digitized or born-digital texts made accessible online, but that simple definition hides a lot of complexity once preservation enters the picture. The biggest difference from a traditional library is not the format. It is what maintenance actually means.
A traditional library fights dust, humidity, and physical decay. A digital library fights format obsolescence, broken links, and vanishing software. Physical books, stored properly, can survive for centuries with minimal intervention. Digital files cannot make that claim.
A well-designed preservation strategy has to account for constant, active change. CLOCKSS, a library-run digital preservation service, frames the core goal plainly: keeping long-term access to published books and articles intact despite shifts in national policy, platform ownership, or publisher priorities, none of which a paper book sitting on a shelf has to worry about.
Access patterns differ too. A physical library caps how many people can read a given book at once. A well-preserved digital copy can, in principle, serve unlimited readers simultaneously, provided the licensing allows it.
Cost structures diverge as well, and not in the direction most people assume. A physical library's biggest costs are usually one-time or slow-moving: buildings, shelving, climate control. A digital library's costs never really stop, since storage has to be actively monitored, formats have to be checked, and staff time has to be budgeted every single year the collection exists. Cutting a physical library's budget for a year rarely destroys the collection. Cutting a digital preservation budget for a year can start a slow, quiet loss that is much harder to reverse once discovered.
Neither model is strictly better. They solve different problems, and most serious research institutions now run both side by side, using the physical collection as a fallback of last resort and the digital one as the primary access point for daily use.
What Are the Biggest Challenges in Digital Preservation?
The biggest challenges of digital preservation are format obsolescence, link rot, funding, and legal uncertainty, roughly in that order of technical severity, though money often decides whether the first two ever get addressed at all.
Format Obsolescence
Format obsolescence is the quiet threat. Software changes, file types get abandoned, and a perfectly intact file becomes unreadable simply because nothing left can open it. A 2026 peer-reviewed guide in Learned Publishing points out that digital books remain far more vulnerable than print in this respect, since publishers and archives have not built the kind of consistent backup habits libraries maintained for physical collections over centuries.
So why hasn't this been fixed already? Because it isn't a single, solvable bug. It's a moving target that never stops moving, and addressing it requires permanent, funded attention rather than a one-time project.
Link Rot and Content Drift
Link rot compounds the problem. Pew Research Center's 2024 analysis found that a quarter of all webpages that existed at any point between 2013 and 2023 were already unreachable by October 2023, with 54 percent of Wikipedia's own reference links broken. Libraries are not immune to this same decay: a 2026 study in the Aslib Journal of Information Management, tracking citations across library-science journals over twenty years, found accessibility falling from 87 percent for sources under five years old to just 38 percent for those over a decade old.
That decay curve should worry anyone who cares about digitized literature specifically, not just web pages in general. A digitized book that lives behind a single institutional link, with no mirror and no backup location, is exposed to exactly this same slow decay.
The Funding Problem
Then there is the money. The Digital Preservation Coalition's 2023 Bit List report tracks at-risk digital material by sorting it into risk tiers from mild concern down to practically extinct, refreshed roughly every two years, and its own conclusion put the underlying issue bluntly: "if digital preservation is possible then data loss is a choice." Most of these losses are not inevitable. They are the result of under-resourced institutions making hard tradeoffs between preservation work and every other line item competing for the same budget.
Preservation is a strange expense to justify politically, because success looks like nothing happening. A well-preserved collection just keeps quietly working, which makes it an easy target when budgets tighten, right up until something breaks and the cost of neglect suddenly becomes visible.
Legal and Copyright Uncertainty
Legal uncertainty adds a final layer, and it is one you will run into constantly if you have ever hunted for an out-of-print title: copyright status can prevent an institution from fully preserving or sharing a work it has every technical ability to save. A library can often legally make a preservation copy of a book it owns, but making that copy accessible to the public is a separate legal question entirely, governed by rules that vary by country and by the specific rights attached to that title.
![]() |
| Old formats do not announce their own expiration date — until a file simply refuses to open. |
Where Is Digital Preservation Headed Next?
What comes next is already underway: better AI recognition, new file formats built specifically for archiving, and more institutional coordination than this field has ever had.
On the technical side, artificial intelligence is closing gaps that stumped traditional OCR for decades. Yale Library, for instance, has been developing a prototype tool nicknamed Digital Collections AI that uses large language models to read, summarize, and answer questions about already-digitized texts, moving beyond simple keyword search toward something closer to genuine research assistance.
On the format side, a new archival e-book standard called EPUB/A is moving through the International Organization for Standardization as of 2026, with the Learned Publishing guide cited earlier noting a target completion around the end of the year. If it succeeds, it could give publishers and libraries a shared, durable format built for long-term preservation from the start, rather than adapting a consumer format after the fact.
The field is also organizing more visibly than before. The National Digital Stewardship Alliance holds its Digital Preservation 2026 conference in November, built specifically around sustaining preservation labor and institutional resilience, and Core, a division of the American Library Association, is running a dedicated introduction-to-digital-preservation session in June 2026 aimed at librarians who are new to the field. Both are signs that this work is professionalizing rather than staying a niche IT concern handled by whichever staff member happens to understand file formats.
None of this guarantees success. AI tools can misread a manuscript with false confidence, new formats can fail to gain adoption, and conferences do not fund themselves. But the direction is clear enough: preservation is getting more attention, more tooling, and more dedicated staff time in 2026 than it had even five years earlier, and that trend line matters more than any single announcement.
Frequently Asked Questions
What is digital preservation in the context of books and literature?
Digital preservation means keeping a digitized or born-digital text readable, intact, and authentic for decades, not just scanning it once and walking away. It covers the file itself, plus the metadata, software, and formats needed to open it correctly in the future. Libraries like HathiTrust and the U.S. National Archives treat this as an ongoing commitment, actively monitoring formats rather than storing files and forgetting them. Without that active management, a perfectly good scan from 2010 can become unreadable within a couple of decades as software moves on. The goal is simple even when the work is not: make sure a text a library saves today can still be opened by someone decades from now.
What is the difference between digitization and digital preservation?
Digitization is the one-time act of converting a physical book into a digital file, while digital preservation is the ongoing work of keeping that file usable for the long haul. A library can digitize a manuscript in an afternoon, but preserving it means monitoring file formats, storage media, and metadata for as long as the institution exists. Think of digitization as taking the photo and preservation as making sure you can still open that photo in 50 years. Most of the actual cost and labor in this field goes toward the second half of that equation, not the first. A library that stops after digitizing has only finished the easy part.
What are the biggest challenges in digital preservation?
The biggest challenges are format obsolescence, link rot, and the sheer cost of active, ongoing management. A 2026 study in the Aslib Journal of Information Management found citation accessibility dropping from 87% for recent sources to just 38% for those over a decade old, and that same decay hits digitized books when formats or platforms disappear. Funding is a constant pressure too, since preservation is not a one-time expense but a permanent budget line. Legal uncertainty around copyright and licensing adds another layer, especially for in-copyright works institutions want to preserve but cannot fully open to the public. None of these challenges has a permanent fix, which is exactly why preservation has to be treated as ongoing work rather than a project with an end date.
How do digital libraries actually preserve books for the long term?
Digital libraries preserve books mainly through two techniques: migration, which means periodically converting files into current formats, and emulation, which recreates old software environments so original files still open correctly. Institutions like the U.S. National Archives lean on migration, converting content into actively managed formats while keeping originals in low-access storage as a backup. Emulation gets used more for complex or software-dependent digital objects where converting the file would lose information. Most large digital libraries use a mix of both, chosen case by case depending on the material. Neither approach works well without consistent funding and staff attention behind it.
Why is digital preservation important for future readers?
Digital preservation matters because a text that is not actively maintained can simply vanish, and it happens faster than most readers expect. Pew Research Center found in 2024 that a quarter of all webpages online between 2013 and 2023 were already unreachable by October 2023, and books stored digitally face the same basic risks. For literature specifically, this means rare regional or minority-language works, like much of the historical Urdu canon that platforms such as Rekhta Foundation are racing to digitize, could disappear from access entirely without deliberate, funded preservation work. Future readers, researchers, and students depend on decisions libraries make today, often long before anyone realizes a particular text was ever at risk.
The Bottom Line
Digital preservation of books is not a solved problem, and it probably never fully will be. Formats keep changing. Funding stays tight. New risks show up as fast as old ones get handled.
But the institutions doing this work well, HathiTrust for scale, Rekhta Foundation for a literature almost nobody outside South Asia was tracking, the Vatican Library and TOPPAN for scripts modern software once could not touch, prove the basic premise still holds. A book digitized carefully today, with real preservation planning behind it and not just a scan, can stay readable long after everyone involved in creating it is gone.
None of that happens automatically. It takes funding decisions made years in advance, staff who understand both the technology and the material, and institutions willing to treat preservation as core mission work rather than a side project squeezed in when budgets allow. Readers rarely see any of that effort directly. They just see whether the book opens.
That's the whole point. Not access for a decade. Access for as long as you, or whoever comes after you, still wants to read.



