Read the AI-Generated Article 100% AI-generated · not peer-reviewed · click to expand
Abstract
Disaster recovery for cultural heritage inherited its vocabulary from physical conservation, where a damaged object persists in continuous form and disaster has a clean before and after. This review argues that born-digital cultural heritage honors neither assumption, and that the field’s dominant framing — disaster as a discrete event absorbed by redundant storage — misdescribes where born-digital material is actually being lost. Reading four practitioner-weighted sources against themselves, the article defends a three-part thesis: that the operative disaster is usually the slow, correlated failure of obsolescence, abandonment, and metadata decay that precedes and survives any acute event, so recovery is better modeled as a maintained state than an event answered; that the true catastrophe is the severing of context — fixity, renderability, provenance metadata, and interpretive community — rather than the destruction of bits; and that what gets saved first is settled not in the emergency but through years of quiet, implicit triage that privileges recoverable digitized surrogates over irreplaceable born-digital material.
A conceptual loss hierarchy and a simple symbolic derivation of the “lots of copies” principle show why redundancy protects the bitstream superbly yet fails against correlated, context-level loss, converging with the metadata thesis on a single conclusion: the field’s flagship doctrine targets the wrong failure mode. The article also surfaces triage as an equity question and steelmans the strongest objections before conceding its limits. With no empirical loss data available, it presents itself explicitly as a conceptual synthesis and research agenda rather than a confirmed finding — arguing that a field whose losses are invisible to its own monitoring must reason honestly from partial evidence.
Introduction
Every conservator knows the smell of a disaster. Wet paper carries a particular mustiness; a fire leaves an acrid film on everything it did not consume; a burst pipe announces itself in the buckling of a floor. These sensory cues have organized more than a century of disaster planning in libraries, museums, and archives. The physical object declares its own injury, and the salvage response — freeze the wet, air-dry the damp, stabilize the charred, prioritize the rare — follows a grammar refined through repeated catastrophe. That grammar assumes a thing that can be picked up, assessed, and treated. Born-digital cultural heritage breaks the assumption at its root. The bitstream has no smell. It does not buckle or discolor. And by the time an institution smells smoke, the object it most needed to save may have been gone for years.
This article is a review of what the literature on disaster recovery for born-digital cultural heritage does and does not establish, and it is also an argument about how that literature should be reorganized. The stakes are not marginal. The Digital Preservation Coalition, reporting from its 2023 conference on born-digital cultural heritage, records that a great deal of the born-digital material produced over the last twenty to thirty years — media art, video games, web content, forum and bulletin-board posts, architectural documentation — is already inaccessible, and it attributes that inaccessibility to a heterogeneous mix of technical and non-technical causes: hardware and software obsolescence, media deterioration, content abandonment, institutional and business decisions, archival practices, legal restrictions, and shifts in audience and culture (Digital Preservation Coalition 2023). Ruan and McDonough put the point more starkly still, describing born-digital cultural heritage as “vanishing away rapidly” and arguing that preservation effort must be redirected toward born-digital rather than merely digitized materials (Ruan and McDonough n.d.). At the same conference, the theorist Sean Cubitt is reported to have insisted that “everything digital is subject to loss,” reframing even a preserved film as an evolving digital file that records its own successive states of decay (Digital Preservation Coalition 2023).
What prior work has established is real but partial. On the operational side, we have detailed workflows. The Smithsonian Institution Archives documents a born-digital pipeline that begins with stabilization and ingest — moving records promptly to a temporary network location precisely to avoid dependence on hardware and operating systems that are already or soon obsolete — and that builds resilience through multiple backups (invoking the LOCKSS principle) together with offline copies retained specifically as the disaster-recovery mechanism, per-file fixity markers to verify integrity and authenticity, and file-format identification to inform preservation decisions (Smithsonian Institution Archives, n.d.). On the strategic side, Ithaka S+R distinguishes two structural approaches to preservation — “programmatic” cross-institutional efforts built around trusted repositories versus locally managed institutional preservation — and describes storage layers that incorporate error-checking of preservation processes and explicit disaster-recovery policies, with cloud adoption framed as “increasingly inevitable” (Ithaka S+R 2023). Crucially, Ithaka S+R also reports a strategic blind spot: digitized surrogates can be re-created by re-digitizing if lost, but born-digital content, once lost, may be lost forever — and yet many institutions still concentrate on digitized and local content and fail to strategize adequately for born-digital material and large media files (Ithaka S+R 2023).
Here is the gap. The operational literature and the strategic literature describe a system whose central metaphor remains the physical salvage event: a disaster happens, and preservation infrastructure absorbs the shock. Disaster recovery is treated as a discrete phase — the offline copy is the thing you reach for after the fire. But the same literature, read against itself, keeps describing a different phenomenon: a diffuse, ongoing attrition in which obsolescence, abandonment, and metadata failure destroy renderability long before any acute event, and in which the decision about what to save is made not in the aftermath but in the years of neglect that precede it. No source in the evidence base joins these two framings into a single account. The practitioner guidance tells us how to keep redundant copies; the strategic reports tell us the losses are already happening; the theorists tell us loss is universal. What is missing is a conceptual synthesis that treats disaster recovery for born-digital heritage as a continuous condition rather than an episodic response, and that takes seriously the two dimensions the acute-event frame obscures: the primacy of metadata over bitstream, and the politics of triage.
My thesis has three parts, and I defend them in turn. First, that the operative disaster in born-digital cultural heritage is usually not the fire, the flood, or the server crash but the slow, correlated failure that precedes and often survives them; disaster recovery for such material is therefore better modeled as a state one maintains than an event one responds to. Second, that the true catastrophe in most digital-heritage loss is not the destruction of bits but the severing of the context — metadata, provenance, renderability — that makes bits legible as heritage; a recovered bitstream without its metadata is frequently not a recovered object at all. Third, that the question the topic poses so bluntly — what gets saved first — is answered in practice by a set of implicit value judgments (digitized over born-digital, institutionally owned over networked, large-and-inconvenient files quietly deprioritized) that the field has not adequately surfaced, and that this triage is where equity and cultural memory are actually decided.
The article develops these claims across five sections. The first examines why the inherited grammar of physical salvage misfits the bitstream and what that mismatch conceals. The second reconstructs the “slow disaster” — the pre-event attrition the sources describe but rarely name as disaster. The third argues that metadata loss is the real catastrophe and works through why fixity and format identification are recovery technologies, not merely preservation hygiene. The fourth turns to redundancy: I offer a simple derivation of why “lots of copies” works and, more importantly, why correlated failure breaks the arithmetic, connecting the mathematics back to the metadata thesis. The fifth confronts triage directly and then steelmans the strongest objections to the whole reframing — including the objection that it dilutes actionable planning — before stating what the evidence can and cannot bear. Throughout, I distinguish scrupulously between what the cited sources establish and what is my own reasoning, because the evidence base for this topic is thinner and more practitioner-weighted than the confident tone of much writing about digital preservation would suggest.
The Grammar of Salvage and the Silence of the Bitstream
Disaster recovery in the cultural sector inherited its vocabulary from physical conservation, and the vocabulary encodes assumptions that do not travel. Consider the canonical salvage sequence for a flooded stack: identify the wet, prioritize by rarity and vulnerability, freeze what cannot be dried within seventy-two hours, air-dry the rest, monitor for mold. Each step presupposes an object that persists through the disaster in a damaged but continuous form. The book is still a book after the flood; the task is to arrest deterioration and restore access to a material substrate that never stopped existing. The disaster is an event with a before and an after, and recovery is the labor that carries the object across the discontinuity.
Born-digital material offers no such continuity. The evidence base makes this concrete in two complementary ways. Ithaka S+R draws the asymmetry that ought to be the starting point of any serious discussion: a digitized surrogate can be re-created by re-digitizing the physical original if the surrogate is lost, whereas born-digital content, once lost, may be lost forever (Ithaka S+R 2023). The digitized object has a fallback outside the digital system — the analogue original — that functions as an implicit backup of last resort. The born-digital object has no such exterior. There is no “original” behind the file to which one can return; the file is the original, and its loss is absolute in a way physical damage rarely is. This single asymmetry already destabilizes the salvage grammar, because salvage assumes a persistent substrate and born-digital heritage frequently has none.
The second destabilization is temporal. In the physical model, the object is intact until the disaster damages it. In the digital model, the object may have ceased to be renderable while sitting apparently untouched on functioning storage. The Digital Preservation Coalition’s inventory of causes is instructive precisely because most of the items on it are not acute events at all: hardware and software obsolescence, media deterioration, content abandonment, institutional and business decisions, archival practices, legal restrictions, audience and cultural change (Digital Preservation Coalition 2023). Of these, only media deterioration resembles the slow physical decay conservators already understand, and even it behaves differently — a hard drive does not gradually fade; it works until it does not. The rest are failures of the environment that gives a bitstream meaning, not failures of the bitstream itself. A video game whose runtime platform no longer exists is lost even though every bit survives. Web content abandoned by its host is lost even though nothing was destroyed. This is why Cubitt’s reported claim that “everything digital is subject to loss,” and his image of a preserved film as an evolving file that records its own states of decay, is more than rhetorical provocation (Digital Preservation Coalition 2023). It relocates decay from the object to the relationship between the object and the shifting technical world it depends on.
The event frame and what it hides
If we accept that most digital loss is environmental and gradual, the “disaster recovery” label starts to look like a category error — or at least a partial truth mistaken for the whole. The label directs institutional attention and budget toward the dramatic, discrete failure: the data centre fire, the ransomware attack, the flooded server room. Those failures are real and they matter. But my argument, which the sources support without quite stating, is that the event frame systematically under-counts the losses that occur outside events. Ruan and McDonough’s characterization of born-digital heritage as “vanishing away rapidly” describes a continuous verb, not a punctual one (Ruan and McDonough n.d.). Vanishing is what happens between disasters. The material disappears not because something struck it but because the ground beneath it moved.
The event frame hides three things in particular. First, it hides the fact that the recoverable object may already be gone before recovery begins, so that a technically flawless restoration of last night’s backup restores an object that was already unrenderable. Second, it hides the primacy of context: because physical salvage restores a self-describing object (a book explains itself), the event frame carries no strong intuition that the description could be lost separately from the thing — yet in digital systems the description routinely fails independently. Third, it hides the politics of triage by making it look like an emergency-room decision, forced and blameless, when in fact the decisive choices about what to preserve were made calmly, over years, in the ordinary allocation of attention that the strategic literature documents (Ithaka S+R 2023). Each of these hidden dimensions becomes a later section of this article.
Why the mismatch persists
It is worth asking why the physical-salvage grammar has proven so durable even among practitioners who know better. Part of the answer, I propose, is institutional: disaster planning lives, in most cultural organizations, alongside facilities management and insurance, domains organized entirely around physical assets and acute events. A born-digital collection has no water table, no fire load, no square footage; it does not fit the risk register that generates the disaster plan. Part of the answer is that the born-digital preservation community has developed its own, largely separate discourse — ingest, fixity, format migration, trusted repositories — that speaks the language of continuous curation rather than the language of disaster (Smithsonian Institution Archives, n.d.; Ithaka S+R 2023). The two communities describe the same objects with non-overlapping vocabularies, and disaster recovery falls into the seam between them. The workflow literature treats the offline disaster copy as one component among many; the disaster-planning literature treats digital collections, when it treats them at all, as a special case of the physical facility. Neither owns the synthesis. The remainder of this article is an attempt to build it.
The Slow Disaster: Obsolescence, Abandonment, and the Pre-Event Casualty
If the operative disaster is usually not the acute event, then a review of disaster recovery for born-digital heritage must first characterize the slow disaster properly. The Digital Preservation Coalition’s list of causes is the best inventory the evidence base offers, and it repays close reading because its structure is itself an argument. The causes divide, as the source notes, into technical and non-technical, and the division matters for recovery because the two kinds of loss demand entirely different responses (Digital Preservation Coalition 2023).
Table 1 sets out the inventory as the source records it, sorted into the two families. I present it not as a novel dataset — it is a faithful rendering of a single conference report’s taxonomy — but because seeing the causes arranged this way makes visible how few of them the acute-event, offline-backup model actually addresses.
| Family | Cause (as reported) | Addressed by offline disaster copy alone? (author’s analysis) |
|---|---|---|
| Technical | Hardware obsolescence | No — the copy survives but cannot be read |
| Technical | Software obsolescence | No — the bitstream survives but cannot be rendered |
| Technical | Media deterioration | Partly — a fresh copy helps only if made before decay |
| Non-technical | Content abandonment | No — nothing was destroyed; support was withdrawn |
| Non-technical | Institutional / business decisions | No — the decision, not the storage, is the failure |
| Non-technical | Archival practices | No — a copy of a poorly described object reproduces the problem |
| Non-technical | Legal restrictions | No — the object may survive but cannot be accessed lawfully |
| Non-technical | Audience / cultural change | No — loss of the community that could interpret the object |
The pattern in the final column is the point. Of the eight reported causes, the offline disaster copy — the mechanism the Smithsonian identifies as the disaster-recovery instrument (Smithsonian Institution Archives, n.d.) — fully addresses none and partly addresses one. This is not a criticism of offline copies, which are indispensable against the acute events they are designed for. It is an observation that the dominant recovery instrument is aimed at a minority of the actual loss. My inference, which the table supports, is that a field investing disproportionately in acute-event recovery is optimizing against the wrong failure distribution.
Obsolescence as a recovery problem, not only a preservation problem
The Smithsonian workflow reveals how deeply obsolescence is woven into even the earliest handling of born-digital material. Records are moved promptly, on ingest, to a temporary network location specifically to avoid dependence on hardware and operating systems that are already or soon obsolete (Smithsonian Institution Archives, n.d.). Read carefully, this is a remarkable admission: the very first act of stewardship is an evacuation from a failing environment. The disaster, in other words, is assumed to be already underway at the moment of acquisition. Obsolescence is not a distant threat the archive guards against; it is the ambient condition the archive is racing.
This reframes obsolescence as a recovery problem. In the physical model, recovery is what you do after damage. For born-digital material, the file-format identification step the Smithsonian describes — using software to determine file types and versions in order to inform preservation decisions (Smithsonian Institution Archives, n.d.) — is functionally a continuous recovery operation. Identifying that a file is in a format nearing the edge of renderability, and acting on that identification through migration, is recovering the object from a slow disaster in progress. The distinction between “preservation” and “recovery” collapses here, which is exactly what my thesis predicts: when the disaster is continuous, recovery is continuous, and the two words name the same activity observed at different moments.
Abandonment and the limits of institutional control
The non-technical causes deserve particular attention because they lie mostly outside the technical apparatus the preservation community controls. Content abandonment, institutional and business decisions, legal restrictions, and cultural change are all, in different ways, withdrawals of support rather than destructions of substance (Digital Preservation Coalition 2023). A platform is shut down; a company folds; a rights regime forecloses access; the community that could read a format disperses. In every case the bits may persist somewhere, yet the heritage is lost because the sustaining relationships are gone.
The DPC report’s own recommended responses are telling on this point. Practitioners are reported to have stressed the value of sustained relationships with artists’ collaborators and software developers, and the wisdom of retaining currently unreadable versions of objects that may be recoverable later (Digital Preservation Coalition 2023). Both recommendations treat recovery as social and prospective rather than technical and reactive. The relationship with the developer is a hedge against future software obsolescence; the retained unreadable file is a bet that a future community will supply the emulator or the interpretive frame that the present lacks. Cubitt’s image of the film as an evolving file preserving its states of decay belongs to this same posture (Digital Preservation Coalition 2023): you cannot arrest the decay, so you document it, and you keep the material against the possibility that later readers can do what you cannot. I read this as the sources reaching, without a unifying vocabulary, toward the continuous-recovery model I am arguing for. The retained unreadable file is a recovery artifact stored before the recovery is possible.
The pre-event casualty
Pulling these threads together yields what I will call the pre-event casualty: the born-digital object that is functionally lost before any disaster plan is triggered, because obsolescence or abandonment has severed its renderability while its storage remained sound. Ithaka S+R’s warning that born-digital content once lost may be lost forever acquires its full weight here (Ithaka S+R 2023). The irreversibility is not confined to acute destruction; it applies equally to the quiet foreclosure of renderability. And because the pre-event casualty leaves the storage layer untouched, it is invisible to exactly the monitoring — capacity, uptime, backup integrity — that acute-event recovery relies upon. An institution can pass every disaster-recovery test, restore every backup flawlessly, and still be steadily losing its born-digital heritage. This is the disaster before the disaster, and it is the one the inherited grammar cannot see.
Metadata as the Real Catastrophe
The second part of my thesis is that the true catastrophe in born-digital loss is usually not the destruction of bits but the severing of the context that makes bits legible as heritage. This section defends the claim and works through its operational consequences, drawing on the fixity and format-identification practices the sources document.
Begin with a distinction the physical model never had to make. A damaged book loses substance and description together, because the description is inscribed in the substance — the title page, the binding, the shelf mark, the physical evidence of provenance are all part of the object. Separating a book from its identity requires deliberate mutilation. In a digital system, by contrast, the object and its description are ordinarily distinct entities: the bitstream sits in storage, and the metadata that identifies, contextualizes, and renders it sits in catalogues, databases, embedded headers, and file-format registries. These can fail independently. It is entirely possible — and, I argue, common — to recover a perfect copy of a bitstream and have no idea what it is, where it came from, how it relates to other objects, or what software once rendered it.
The evidence base does not state this thesis in so many words, but its recommended practices only make sense if the thesis is true. Consider fixity. The Smithsonian generates per-file fixity markers to verify integrity and authenticity across an object’s life (Smithsonian Institution Archives, n.d.); Ithaka S+R describes storage layers incorporating error-checking of preservation processes (Ithaka S+R 2023). A fixity marker is a compact piece of metadata whose entire function is to answer a question about a bitstream — is this the same as it was? — that the bitstream cannot answer about itself. Fixity is therefore a paradigm case of context that is separate from, and epistemically prior to, the object. Without the stored fixity value, a recovered file cannot be authenticated; the archive can restore the bits but not the assurance that they are the right bits. My inference is that fixity data is not merely preservation hygiene but a recovery prerequisite: it is the metadata without which a restored bitstream cannot be trusted, and an untrustworthy restoration is, for heritage purposes, a failed one.
Renderability as context
Format identification extends the argument. The Smithsonian uses software to determine file types and versions precisely in order to inform preservation decisions (Smithsonian Institution Archives, n.d.). Here the relevant context is not stored inside the archive at all — it lives in the shared, external knowledge of what a given format is and how it is rendered, knowledge maintained by format registries, software documentation, and expert communities. This is why software obsolescence and content abandonment sit on the DPC’s loss inventory alongside media deterioration (Digital Preservation Coalition 2023): they are losses of external context. When the format knowledge disappears, the bitstream becomes an artifact in an unknown language. Recovering it is then not a storage operation but an act of decipherment, and decipherment may be impossible.
This is where the DPC practitioners’ advice to retain currently unreadable versions in the hope of later recovery (Digital Preservation Coalition 2023) reveals its logic. Retaining the unreadable file is a wager that the missing external context — the emulator, the format specification, the surviving expert — can be reassembled in the future. The wager is rational precisely because the bits are cheap to keep and the context is what was lost; you hold the substrate against the day the context returns. Ruan and McDonough’s grounding of preservation strategy in UNESCO’s Charter on the Preservation of the Digital Heritage points in the same direction (Ruan and McDonough n.d.): a charter operates at the level of shared commitments and frameworks — that is, at the level of the sustaining external context — rather than at the level of individual storage systems. The instruments the field reaches for when it is most serious about durability are context-level instruments.
A hierarchy of loss
I propose, as the author’s synthesis of the sources, a hierarchy of what is at stake in a born-digital disaster, ordered from least to most catastrophic. Figure 1 presents it as a conceptual diagram. It is not derived from data; it is an argumentative structure built from the sources’ own emphases, and it should be read as such.
- Bitstream — the raw sequence of bits. Recoverable from any surviving copy; the level the offline disaster copy protects (Smithsonian Institution Archives, n.d.).
- Fixity / integrity — the assurance that the recovered bits are authentic and unchanged. Depends on separately stored fixity markers and error-checking (Smithsonian Institution Archives, n.d.; Ithaka S+R 2023).
- Renderability — the format knowledge and software environment needed to turn bits into a perceptible object. Depends on external, shared context vulnerable to software obsolescence and content abandonment (Digital Preservation Coalition 2023).
- Descriptive / provenance metadata — what the object is, where it came from, how it relates to others. Enables the object to function as heritage rather than as an anonymous file.
- Interpretive community — the people and relationships able to read, contextualize, and maintain the object; eroded by audience/cultural change and by loss of contact with creators and developers (Digital Preservation Coalition 2023).
The hierarchy makes a claim that the acute-event model tends to invert. The offline disaster copy protects the bottom level, the bitstream, extremely well. But loss climbs from the top. Interpretive communities disperse, descriptive metadata is orphaned, formats fall out of support — and only rarely, and last, are the bits themselves destroyed. An institution that protects level one while levels three through five quietly fail is preserving the least catastrophic thing while losing the most catastrophic. This is the strong form of my metadata thesis: because the levels are ordered and each presupposes those above it, a recovery capability concentrated at the bottom of the hierarchy addresses the losses least likely to occur and least damaging when they do. The corollary for practice is that fixity, format identification, descriptive metadata, and sustained relationships are not supplements to disaster recovery; for born-digital heritage they are its substance.
The special vulnerability of large media files
The metadata thesis also illuminates a specific vulnerability the strategic literature flags without fully explaining. Ithaka S+R reports that institutions fail to strategize adequately not only for born-digital material in general but specifically for large media files (Ithaka S+R 2023). Why should size, as such, be a preservation risk? My analysis is that large media files aggravate every level of the loss hierarchy at once. Their bulk makes redundant copies and offline snapshots expensive, weakening level one. Their dependence on codecs — highly specific, fast-moving software context — makes renderability fragile at level three. And their production workflows generate rich technical metadata (color spaces, aspect ratios, encoding parameters) that is essential to correct rendering yet easily stripped in transfer, threatening levels two and four. The large media file is thus a concentrated instance of the general problem: the more an object’s meaning depends on external and technical context, the more its apparent physical survival misleads. That the sources single out large media files as a strategic gap is, on this reading, further evidence that the field’s instincts run ahead of its framework.
Redundancy and Its Discontents: The Arithmetic and Its Failure
Against the slow disaster and the primacy of metadata, the field’s principal defensive doctrine is redundancy — the LOCKSS principle, “lots of copies keep stuff safe,” which the Smithsonian invokes explicitly and implements through multiple backups plus offline copies (Smithsonian Institution Archives, n.d.). Redundancy is genuinely powerful, and this section takes it seriously by first showing, through a simple derivation of my own, why it works, and then showing why the same derivation exposes its limits and returns us to the metadata thesis.
Why lots of copies keep stuff safe
Suppose an object is held in N independent copies, and suppose each copy is lost, over some interval, with probability p. If the losses are genuinely independent, the probability that all N copies are lost — that is, that the object is destroyed — is the product of the individual failure probabilities:
(1)
Equation (1) is elementary, and it is my own derivation offered for illustration, not a measured result; p and N here are symbolic. But it captures why redundancy is so attractive. Because p is a probability less than one, raising it to a higher power drives the loss probability down steeply. If a single copy fails with probability one in ten over some period, three independent copies fail together with probability one in a thousand, and five with probability one in a hundred thousand. The doctrine’s intuitive force is real: under independence, a modest number of copies converts a likely loss into a vanishingly improbable one. This is the arithmetic that underwrites the Smithsonian’s multiple backups (Smithsonian Institution Archives, n.d.) and that motivates Ithaka S+R’s programmatic, cross-institutional trusted repositories, in which copies are distributed across organizations rather than concentrated in one (Ithaka S+R 2023).
The independence assumption and correlated failure
The entire force of equation (1) rests on the word independent. If the failures of the copies are correlated — if a single cause can take out many copies at once — the multiplication no longer holds, and the loss probability can be dramatically higher than pN suggests. In the limiting case of perfect correlation, where one cause destroys all copies simultaneously, the effective loss probability collapses back toward that of a single copy, and the N copies provide almost no protection at all. I state this as my own analysis, but it is standard reasoning about redundant systems, and its application to born-digital heritage is where the discussion becomes pointed.
Which of the DPC’s loss causes produce correlated failure? Almost all of the non-technical ones, and both of the software-related technical ones (Digital Preservation Coalition 2023). Software obsolescence does not strike one copy; it strikes every copy of an object in the obsolete format at once, no matter how many there are or where they sit — a hundred perfect copies of a file in a dead format are a hundred instances of the same loss. Content abandonment, an institutional or business decision to withdraw support, a legal restriction, the dispersal of the interpretive community — each of these is a single cause with system-wide reach. These are precisely the causes the loss hierarchy of Figure 1 locates at levels three through five, the context levels. Redundancy multiplies copies of the bitstream, which lives at level one. It does nothing to multiply the format knowledge, the descriptive metadata, or the interpretive community, and so it provides no protection against the correlated, context-level failures that the evidence base identifies as the dominant mode of loss.
This is the discontent in redundancy, and it is the same finding as the metadata thesis reached by a different route. LOCKSS-style replication is a superb defense against uncorrelated, bitstream-level loss — media failure, localized disaster, hardware death. It is close to useless against correlated, context-level loss — obsolescence, abandonment, and the erosion of interpretive capacity. Since the sources indicate that context-level loss is where born-digital heritage is actually disappearing (Digital Preservation Coalition 2023; Ruan and McDonough n.d.), the field’s flagship defensive doctrine, taken alone, is aimed at the wrong failure mode. The remedy is not to abandon redundancy but to recognize what must be replicated: not only bitstreams but the context that renders them. Retaining format documentation and emulation environments, maintaining descriptive and provenance metadata in multiple systems, and — as the DPC practitioners urge — sustaining relationships with creators and developers (Digital Preservation Coalition 2023) are all, in this light, forms of redundancy applied to the levels of the hierarchy that ordinary copying leaves exposed.
Programmatic distribution, cloud centralization, and the geometry of correlation
The correlation lens also clarifies the strategic choice Ithaka S+R frames between programmatic cross-institutional preservation and locally managed institutional preservation, and its claim that cloud adoption is “increasingly inevitable” (Ithaka S+R 2023). The programmatic model distributes copies across independent organizations, which is, in the terms of equation (1), an attempt to reduce correlation: different institutions have different hardware, staff, funding cycles, and jurisdictions, so a cause that fells one is less likely to fell all. The local model concentrates copies under a single administration, where a single budget cut, policy change, or facility disaster is more likely to be a common cause — higher correlation, weaker effective redundancy.
The move to the cloud is more ambiguous than its inevitability implies, and I read it as cutting both ways. On one hand, major cloud providers offer geographic distribution and automated integrity-checking that reduce the correlation of physical and media failures — a genuine gain at level one of the hierarchy. On the other hand, consolidating many institutions’ holdings onto a small number of providers reintroduces correlation at the levels the DPC inventory identifies as most dangerous: a single provider’s business decision, price change, service deprecation, or contractual dispute is exactly the kind of non-technical, common-cause event that can render many collections inaccessible at once (Digital Preservation Coalition 2023). The cloud, in other words, tends to trade uncorrelated physical risk for correlated institutional and business risk. Whether that trade is favorable depends on which risks dominate — and the evidence base’s emphasis on non-technical causes suggests the correlated institutional risks deserve more weight than the confident language of “increasingly inevitable” (Ithaka S+R 2023) grants them. This is my assessment, offered as a reading of the strategic literature rather than a measured comparison; the point is that the redundancy arithmetic, taken seriously, makes the cloud question a question about correlation, not capacity.
Triage and the Politics of What Gets Saved First
The topic asks about “the uneven reality of what gets saved first,” and this section confronts it directly. My argument is that triage in born-digital heritage is not primarily the frantic, blameless prioritization of an emergency but the accumulation of quiet, structural choices made long before any emergency — and that these choices, being implicit, escape the scrutiny that explicit triage would attract.
The digitized-over-born-digital default
The clearest evidence of structural triage is Ithaka S+R’s finding that many institutions concentrate on digitized and local content and fail to strategize adequately for born-digital material and large media files (Ithaka S+R 2023). This is a triage decision, though it is never announced as one. By directing attention, staff, and infrastructure toward digitized surrogates, institutions implicitly rank the digitized above the born-digital in the queue of what gets protected first. And the ranking is precisely backwards relative to irreversibility. The same source establishes that the digitized surrogate is the recoverable one — it can be re-created from the surviving physical original — while the born-digital object is the irreplaceable one (Ithaka S+R 2023). To prioritize the digitized is therefore to spend the most protective effort on the material that needs it least and the least on the material that, once lost, is lost forever. My inference is blunt: the field’s default triage optimizes for the recoverable and neglects the irreplaceable, which is close to the opposite of what a rational triage would do.
Why does this inversion persist? I offer three reasons, as the author’s analysis. First, digitized collections are legible to the physical-salvage grammar discussed earlier — they have originals, provenance, and catalogue records inherited from the analogue world — whereas born-digital collections arrive without that scaffolding and are harder to fit into existing workflows. Second, digitization projects are fundable, visible, and countable in a way that born-digital stewardship is not; the number of pages scanned is a metric, while the number of obsolescence-driven losses averted is nearly invisible. Third, the large media files that Ithaka S+R flags are expensive and inconvenient at every level of the loss hierarchy, as argued above, and inconvenient objects are quietly deprioritized. None of these reasons is a defensible preservation rationale, yet together they explain a durable pattern of misdirected triage.
Triage under acute conditions
Structural triage is compounded by the choices forced during acute events, and here the evidence base is thinner but not silent. The offline copy that the Smithsonian designates as the disaster-recovery mechanism (Smithsonian Institution Archives, n.d.) embodies a prior triage: someone decided what to copy offline and how often, and whatever fell outside that scope is unrecoverable when the acute event arrives. Fixity markers and format identification (Smithsonian Institution Archives, n.d.) similarly encode triage, because generating and maintaining them at scale is labor-limited, and the objects that receive rich fixity and format metadata are, by that fact, the objects that can be authenticated and rendered after recovery. The unglamorous truth is that the outcome of an acute-event recovery is largely determined before the event, by which objects were positioned — copied, described, format-identified — to survive it. The emergency does not decide what is saved first; it reveals a decision already made.
The equity dimension
Because triage is structural and implicit, its distributive consequences go unexamined, and this is where the politics becomes an equity question. The DPC inventory names audience and cultural change and legal restrictions among the causes of loss (Digital Preservation Coalition 2023), and both have distributive edges. Material tied to communities with less institutional power — smaller audiences, marginal cultural forms, creators without the standing to command sustained relationships with archives — is more exposed to the “audience/cultural change” cause, because the interpretive community at level five of the hierarchy is thinner and disperses sooner. Material entangled in restrictive rights regimes is more exposed to the legal cause. The forms of born-digital heritage that the DPC lists as already largely inaccessible — media art, video games, web content, forum and bulletin-board posts (Digital Preservation Coalition 2023) — include many that emerged from non-elite, vernacular, or commercially marginal contexts, exactly the contexts least likely to command the resources that structural triage rewards. I do not have, in this evidence base, quantitative data on the demographics of digital loss, and I will not manufacture any; this is an argument about mechanism, not a measured distribution. But the mechanism is clear enough to warrant the claim that the field’s implicit triage is likely to reproduce and amplify existing inequities in whose culture is remembered, and that surfacing triage as an explicit, accountable decision is therefore not merely a matter of efficiency but of justice.
Toward explicit triage
The constructive implication, which I offer as a proposal rather than a finding, is that born-digital disaster recovery should make its triage explicit and align it with irreversibility rather than convenience. Ithaka S+R’s own asymmetry supplies the ordering principle: prioritize the irreplaceable born-digital object over the re-creatable digitized surrogate (Ithaka S+R 2023). The loss hierarchy of Figure 1 supplies a second principle: within the born-digital, prioritize the objects and levels most exposed to correlated, context-level failure, because those are the losses that redundancy cannot undo. And the DPC’s relational recommendations supply a third: treat sustained relationships with creators, collaborators, and developers as preservation infrastructure to be maintained deliberately, not as goodwill (Digital Preservation Coalition 2023). An explicit triage built on these three principles would still make hard choices — the point of triage is that not everything can be first — but it would make them visibly, accountably, and in the right order, which is more than the prevailing implicit default achieves.
Counterarguments, Competing Interpretations, and the Limits of the Evidence
A review that only confirms its own thesis is an advertisement. This section states the strongest objections to the argument I have built, in their most persuasive form, and then responds. It ends with an honest accounting of what this evidence base can and cannot support, because the limitations here are unusually consequential.
Objection 1: The pessimism is overstated; redundancy and good practice largely work
The strongest version of the optimistic counterargument runs as follows. The mature preservation workflows the sources document — stabilization on ingest, multiple backups, offline disaster copies, per-file fixity, format identification, error-checking (Smithsonian Institution Archives, n.d.; Ithaka S+R 2023) — constitute a genuinely effective system. Where they are properly implemented, born-digital material is well protected against both acute events and gradual media failure, and format migration handles obsolescence. The narrative of rapid vanishing, on this view, describes material that was never under professional stewardship at all — orphaned web content, abandoned platforms, personal media — and it is a category error to indict the archival system for failing to save material that never entered it. The Ithaka S+R report itself frames cloud adoption and programmatic repositories as an improving trajectory (Ithaka S+R 2023).
This objection is partly right, and the concession matters. For material that is ingested, described, format-identified, and redundantly stored, the system does work well against the failure modes it targets, and I have said as much about redundancy at level one of the hierarchy. But the objection understates two things. First, the DPC report attributes present inaccessibility partly to archival practices themselves (Digital Preservation Coalition 2023) — that is, loss occurs even within professional stewardship, not only outside it, so the tidy inside/outside distinction does not hold. Second, and more fundamentally, format migration is not the settled solution the optimistic view assumes: migration addresses renderability only where the format is understood well enough to migrate from, and the correlated, context-level losses I have analyzed are exactly the cases where that understanding is what fails. The optimist is describing the system performing well against uncorrelated, bitstream-level risk and generalizing to the correlated, context-level risk where it performs poorly. The evidence base’s own emphasis on non-technical and software-obsolescence causes (Digital Preservation Coalition 2023) is the reason the generalization does not hold.
Objection 2: Disaster recovery should stay a distinct discipline; blurring it into “continuous curation” is a conceptual loss
A second objection targets my central reframing. Disaster recovery, the objection holds, earns its power precisely from being a bounded, event-focused discipline with clear triggers, procedures, and accountabilities. Collapsing it into a diffuse “continuous condition” risks dissolving a sharp, actionable practice into an unactionable generality. When the server room floods, an institution needs a checklist, a phone tree, and a cold-storage vendor — not a meditation on the ontology of loss. By insisting that the real disaster is the slow one, my argument may sap the urgency and specificity that make acute-event planning effective, and leave institutions worse prepared for the fires that do happen.
I take this objection seriously because it identifies a genuine risk, and my response is a matter of both-and rather than either-or. Nothing in my argument counsels abandoning acute-event procedures; the offline copy, the checklist, and the cold-storage vendor remain necessary, and the redundancy arithmetic of equation (1) shows they are highly effective against the risks they target. The claim is that they are insufficient, not that they are wrong. The continuous-recovery framing does not dissolve acute-event planning; it embeds it within a larger practice that also addresses the correlated, context-level losses the acute frame ignores. Far from sapping actionability, the reframing generates specific actions the acute frame does not: continuous format monitoring, redundant metadata storage, relationship maintenance, and explicit irreversibility-ordered triage. These are as concrete as a phone tree. The objection is right that vagueness would be a failure; my answer is that the continuous model, properly specified through the loss hierarchy, is not vague. It is more demanding, not less.
Objection 3: The metadata thesis inverts real institutional priorities for rhetorical effect
A third objection holds that elevating metadata above the bitstream is a rhetorical inversion that misdescribes practice. Institutions protect bitstreams first because a lost bitstream is unrecoverable while lost metadata can often be reconstructed from surviving context, embedded headers, or expert examination. Fixity and format identification, on this reading, are servants of bitstream preservation, not rivals to it; my hierarchy dresses up a supporting cast as the lead.
The steelman here rests on a claim about reconstructability, and it is where the objection is weakest. The sources indicate that the context levels are frequently not reconstructable: content abandonment, software obsolescence, and the dispersal of interpretive communities are, by the DPC’s own account, causes of loss precisely because the context cannot simply be regenerated (Digital Preservation Coalition 2023). Where format knowledge or the interpretive community has genuinely disappeared, there is nothing left to reconstruct the metadata from — which is why practitioners retain unreadable files as bets on future recovery rather than reconstructing them now (Digital Preservation Coalition 2023). Meanwhile the objection concedes my structural point: fixity and format identification really are separate from the bitstream and really are prerequisites for a trustworthy, renderable recovery (Smithsonian Institution Archives, n.d.). Whether one calls context “lead” or “supporting cast,” a bitstream recovered without authenticable fixity and without renderable format context is not a recovered heritage object. My hierarchy claims exactly that dependency, and the objection does not overturn it.
Competing interpretation: loss as intrinsic and generative rather than as failure
Beyond these objections stands a competing interpretation of the whole enterprise, articulated most sharply in Cubitt’s reported position that everything digital is subject to loss and that a preserved film is best understood as an evolving file recording its own states of decay (Digital Preservation Coalition 2023). On this view, the ambition to arrest loss is misconceived from the start; decay is intrinsic to digital objects, and the honest response is to document and work with it rather than to fight a war that cannot be won. This is not an objection to my argument so much as a more radical extension of its premise, and it is worth confronting because it could be read to license fatalism.
I accept the descriptive claim and reject the fatalist inference. That everything digital is subject to loss (Digital Preservation Coalition 2023) is consistent with, indeed supportive of, the continuous-recovery model: if loss is intrinsic and ongoing, then recovery must be ongoing too, which is precisely my thesis. What does not follow is that effort is futile. The DPC practitioners who retain unreadable files and maintain relationships with developers (Digital Preservation Coalition 2023) are neither denying decay nor surrendering to it; they are practicing a preservation that works with decay by keeping open the possibility of future recovery. Cubitt’s insight refines the goal — from arresting decay to stewarding an object through its successive states while preserving the possibility of its future legibility — without licensing abandonment. The competing interpretation, properly understood, is an ally of the continuous model, not a rival to it.
The limits of this evidence base
Honesty requires me to be explicit about the fragility of the ground beneath this review, because it bears directly on how much confidence the thesis can carry. The substantive evidence here rests on four sources, and their evidential character is uneven. Two are issuing-organization documents rather than peer-reviewed studies: Ithaka S+R’s research report (Ithaka S+R 2023) and the Smithsonian Institution Archives’ practitioner guidance (Smithsonian Institution Archives, n.d.). One is a conference report summarizing practitioner discussion rather than presenting primary findings (Digital Preservation Coalition 2023). The fourth, Ruan and McDonough, I was able to read only at abstract level, and its publication year and venue could not be reliably confirmed, which is why it appears here as undated (Ruan and McDonough n.d.); readers should treat citations to it with corresponding caution. There is, in this evidence base, no empirical study measuring rates of born-digital loss, no controlled comparison of recovery strategies, and no quantitative data on the distributive or equity dimensions I have argued matter most. The numeric content I have used is limited and honestly labelled: the “20–30 years” span and the taxonomy of causes are the DPC report’s (Digital Preservation Coalition 2023); the loss-probability relation in equation (1) is my own symbolic derivation with illustrative, not measured, values.
Two consequences follow for the standing of the argument. First, the thesis is best understood as a conceptual synthesis and a research agenda rather than an empirically confirmed conclusion. Its logic — the pre-event casualty, the loss hierarchy, the correlation critique of redundancy, the irreversibility-ordered triage — is, I believe, sound given the sources, but it awaits the empirical work that this evidence base does not contain and that a fuller literature (peer-reviewed disaster-resilience studies, empirical work on bit-rot and metadata loss, documented emergency-rescue cases, and the foundational replication literature) would supply. Second, the specific claims most in need of that future testing are the quantitative ones I have deliberately declined to fabricate: the actual distribution of loss across the DPC’s causes, the real correlation structure of cloud versus programmatic redundancy, and the demographic incidence of loss that the equity argument turns on. I have argued that the mechanisms point in a particular direction; I have not, and on this evidence could not, measure the magnitudes. Naming these limits is not a hedge against the thesis but part of it: a field whose loss is largely invisible to its own monitoring is a field that must reason carefully from partial evidence, and doing so honestly is the first discipline that born-digital disaster recovery requires.
Conclusion
This review began with a sensory observation — that the bitstream has no smell — and used it to expose a structural mismatch at the heart of disaster recovery for born-digital cultural heritage. The inherited grammar of physical salvage presupposes an object that persists through catastrophe in a damaged but continuous form, and a disaster that has a clean before and after. Born-digital material honors neither presupposition. Ithaka S+R’s asymmetry between the re-creatable digitized surrogate and the irreplaceable born-digital object shows that born-digital heritage has no exterior original to fall back on (Ithaka S+R 2023), while the Digital Preservation Coalition’s inventory of causes shows that most loss is environmental and gradual rather than acute — obsolescence, abandonment, institutional decision, legal foreclosure, and cultural change, only one of which the offline disaster copy even partly addresses (Digital Preservation Coalition 2023). From these observations the article defended a three-part thesis: that the operative disaster is usually the slow, correlated failure that precedes and survives the acute event, so recovery is better modeled as a state maintained than an event answered; that the true catastrophe is the severing of context — fixity, renderability, descriptive and provenance metadata, interpretive community — rather than the destruction of bits; and that the question of what gets saved first is settled not in the emergency but in years of quiet, implicit triage that the field has not adequately surfaced.
The three arguments converge rather than merely accumulate, and their convergence is the review’s central contribution. The pre-event casualty — the object rendered unrenderable while its storage stays sound — is invisible to precisely the capacity, uptime, and backup-integrity monitoring on which acute-event recovery relies, so an institution can pass every disaster-recovery test and still be losing its heritage steadily. The loss hierarchy of Figure 1 explains why: loss climbs from the top, as interpretive communities disperse and metadata is orphaned and formats fall out of support, while redundancy protects the bottom level, the bitstream, extremely well. The redundancy critique then reaches the same destination by a different route. My symbolic derivation,
, captures why lots of copies keep stuff safe under independence, but its force rests entirely on that word; software obsolescence, content abandonment, institutional decision, legal restriction, and community dispersal are all common-cause, correlated failures that no amount of bitstream copying can undo (Digital Preservation Coalition 2023). The metadata thesis and the correlation critique are thus one finding seen twice: the field’s flagship defensive doctrine, taken alone, is aimed at the wrong failure mode, and the remedy is to replicate context — format documentation, emulation environments, provenance metadata, and, as the DPC practitioners urge, sustained relationships with creators and developers (Digital Preservation Coalition 2023; Smithsonian Institution Archives, n.d.). The triage argument completes the picture, showing that the default preference for digitized and local content over born-digital and large media files optimizes for the recoverable and neglects the irreplaceable — close to the opposite of what a rational, irreversibility-ordered triage would do (Ithaka S+R 2023).
The article was equally careful about what it could not establish, and that candor is part of the argument rather than a qualification of it. The evidence base rests on four sources of uneven evidential character: two issuing-organization documents (Ithaka S+R 2023; Smithsonian Institution Archives, n.d.), one conference report summarizing practitioner discussion (Digital Preservation Coalition 2023), and one source I could read only at abstract level, whose year and venue I could not confirm and which therefore appears undated (Ruan and McDonough n.d.). There is no empirical study measuring rates of born-digital loss, no controlled comparison of recovery strategies, and no quantitative data on the distributive or equity dimensions the article argued matter most. The numeric content used was correspondingly limited and honestly labelled — the twenty-to-thirty-year span and the taxonomy of causes belong to the DPC report, and equation (1) is a symbolic derivation with illustrative, not measured, values. The thesis, then, is a conceptual synthesis and a research agenda, not an empirically confirmed conclusion: its logic is, I have argued, sound given the sources, but it awaits the empirical work this evidence base does not contain. The equity claim in particular is an argument about mechanism, not a measured distribution; I declined to manufacture the demographic data it would require.
Those limits point directly to the open problems that follow. The claims most in need of testing are the quantitative ones deliberately left unfabricated: the actual distribution of loss across the DPC’s technical and non-technical causes; the real correlation structure of cloud versus programmatic redundancy, which the article argued the confident language of “increasingly inevitable” adoption underweights (Ithaka S+R 2023); and the demographic incidence of loss on which the equity argument turns. A fuller literature — peer-reviewed disaster-resilience studies, empirical work on bit-rot and metadata loss, documented emergency-rescue cases, and the foundational replication literature — would supply what this practitioner-weighted evidence base cannot. Beyond measurement, the article’s constructive proposals remain proposals awaiting implementation and evaluation: an explicit triage aligned with irreversibility rather than convenience, ordered by the three principles the article drew from its sources — prioritize the irreplaceable born-digital object over the re-creatable surrogate, prioritize within the born-digital the levels most exposed to correlated context-level failure, and treat sustained relationships with creators and developers as preservation infrastructure to be maintained deliberately (Ithaka S+R 2023; Digital Preservation Coalition 2023). The deepest task is institutional: to close the seam between a disaster-planning discourse organized around physical assets and a preservation discourse organized around continuous curation, neither of which currently owns the synthesis. If the review is right that a field whose loss is largely invisible to its own monitoring must reason carefully from partial evidence, then naming these problems honestly — refusing both false confidence and the fatalism that Cubitt’s intrinsic-loss thesis might seem to license (Digital Preservation Coalition 2023) — is the first discipline born-digital disaster recovery requires.
References
Citation Verification Summary
Digital Preservation Coalition. 2023. Born Digital Cultural Heritage Now #BDCH23. Conference report, ACMI, Melbourne, 29 November–1 December 2023. Digital Preservation Coalition blog. https://www.dpconline.org/blog/bdch23.
(Non-scholarly URL reference – not checkable in Crossref/OpenAlex/arXiv; excluded from fabrication accounting. Reference cites a web resource; scholarly indexes (Crossref/OpenAlex/arXiv) cannot verify this type. Excluded from fabrication accounting.; URL appears reachable (HTTP 200))Ithaka S+R. 2023. The Effectiveness and Durability of Digital Preservation and Curation Systems. Research report. Ithaka S+R. https://sr.ithaka.org/publications/the-effectiveness-and-durability-of-digital-preservation-and-curation-systems/.
(Year off by one: cited 2023, found 2022 (likely online-first vs print date; not penalized); Author mismatch: cited Ithaka S+R., found Oya Y. Rieger; Matching-title record located (‘The Effectiveness and Durability of Digital Preservation and’), but overall match confidence 0.60 is below threshold 0.70 (weak author/year corroboration); please verify manually)Ruan, Jian, and Jerome P. McDonough. n.d. “Preserving Born-Digital Cultural Heritage in Virtual World.” ResearchGate publication 261268325. https://www.researchgate.net/publication/261268325. [Read at abstract level only; publication year and venue not reliably confirmed at time of citation.]
Smithsonian Institution Archives. n.d. Preservation Strategies for Born-Digital Materials. Smithsonian Institution Archives. https://siarchives.si.edu/what-we-do/digital-curation/preservation-strategies-born-digital-materials.
(Author mismatch: cited Smithsonian Institution Archives., found bradyh; Matching-title record located (‘Preservation Strategies for Born-Digital Materials’), but overall match confidence 0.50 is below threshold 0.70 (weak author/year corroboration); please verify manually)Reviews
How to Cite This Review
Replace bracketed placeholders with the reviewer’s name (or “Anonymous”) and the review date.
