A pensive figure whose mind merges, on one side, with a sunset oil derrick and refinery and, on the other, with a luminous field of language, data, and knowledge
The Density Archives

The Fossil Logic of theIndustrial and AI Revolutions

How Civilization Learned to Mine Energy and Now Knowledge.

By
David C. Krakauer
Series
The Density Archives
Published
June 30, 2026
Read
18 minutes

AI Image

For most of the history of life, energy has moved through trophic, or feeding, networks. Producers capture sunlight, consumers eat producers, and predators eat consumers. At every link in the natural chain of "eat or be eaten," energy is lost as heat and the throughput of the network is bounded by the flux of solar radiation at its base.

The Oxford zoologist Charles Elton, who published Animal Ecology in 1927 at the age of twenty-six, was the first to give this cascade a precise geometric interpretation. Watching the unforgiving communities of the Arctic on summer expeditions to Spitsbergen, Elton noticed that the structure of any community was always a pyramid, in which many plants supported a smaller mass of herbivores, which supported a still smaller mass of carnivores, which supported a handful of large predators at the top. Each predator-prey interaction lost most of its input to heat. This is why the biggest and fiercest animals are rare not by accident but by thermodynamics.

Elton's contemporary at Yale, the Cambridge-trained polymath G. Evelyn Hutchinson, carried this logic outward into geochemistry. Hutchinson, who is sometimes called the father of modern ecology, treated lakes and forests as accounting systems for the elements, phosphorus, nitrogen, carbon. He proposed that an ecosystem was a network of feedback loops in which living and non-living processes obeyed a common mechanism, and he and his student Raymond Lindeman translated Elton's pyramid into a theory of energy flow at trophic levels. The total growth could not exceed the rate at which the sun delivered new fuel to the bottom of the trophic network. The evolutionary biologist Olivia Judson has placed this trophic regime in a longer arc, framing the history of life as a sequence of five energy expansions, each marked by the evolutionary appearance of organisms exploiting a new source, from geochemical energy, sunlight, oxygen, flesh, to fire.

Beneath the surface of darwinian competition that Alfred Lord Tennyson described as "red in tooth and claw," a second energy deposit had been accumulating for hundreds of millions of years. From roughly 360 to 300 million years ago, in the Carboniferous era, photosynthetic organisms died, descended into anoxic swamps, and were buried beneath sediments before they could be eaten. Heat and pressure transformed the buried matter into peat, then lignite, then bituminous coal, and finally into anthracite. Marine plankton, deposited under analogous regimes, became petroleum and natural gas. It is thought that climate, tectonics, and the slow uplift of Pangaean basins did most of the preserving work. Roughly four trillion tonnes of recoverable fossil carbon now lie under our feet that records several hundred million years of compressed sunlight. This is the primordial planetary archive that built Standard Oil, BP, ExxonMobil, Saudi Aramco, and Gazprom, and that has, since 1750, raised the carbon content of the atmosphere by half.

The steam engine unlocked this energy dense archive. Beginning with Thomas Newcomen's atmospheric engine of 1712, refined by James Watt, and perfected by Richard Trevithick, a new class of technology bypassed trophic networks entirely and drew energy directly from geology. The Industrial Revolution was a transition from network-mediated flows to archive-mediated extraction.

Watt was a Scottish instrument maker employed at the University of Glasgow, charged in 1764 with repairing a model Newcomen engine. Disturbed by how much steam the model wasted in alternately heating and cooling its cylinder, he proposed that the steam be condensed in a separate vessel, leaving the working cylinder hot. The separate condenser is the foundational patent of the age of steam. Trevithick, born in 1771 in Cornwall, built compact high-pressure engines that could be carried to a mine in a farm wagon. In 1802 he ran an experimental engine at Coalbrookdale at an unprecedented one hundred and forty-five pounds per square inch. The first running of his locomotive on the Penydarren tramway is reported in his wonderfully understated letter to Davies Giddy of 15 February 1804: "We put it on the tramway. It worked very well."

Newcomen's engine was built to pump water out of coal mines so that more coal could be drawn in order that more engines could be built. The technology that opened the thermodynamic archive ran on its own contents. There was no preexisting fuel for the first generation of engines other than coal, and no way to mine coal at scale without engines. This is the cybernetic logic of the industrial revolution, and of all technological revolutions, in which the outputs of machines are used to bootstrap better machines.

Part One

The Energy Archive

Deposition

Geological time is the hidden variable that sources the energy archive. The carboniferous forests that became the world's great coal seams flourished for sixty million years and were then buried with enough efficiency to preserve whole biomes. The plants did not decompose and were sealed under sediment to emerge hundreds of millions of years later as the working capital of England and Pennsylvania. Petroleum traces a parallel pathway from marine microorganisms in stratified Mesozoic seas. Roughly four trillion tonnes of recoverable fossil carbon, representing between three hundred and six hundred million years of net primary production, are now concentrated in geological strata to which we have technological access. The making of the modern world is the machine-assisted digestion of the Paleozoic.

Energy density is the property that makes this digestion lucrative and provides for the energy needs of the modern world. It describes the chemical energy released, per unit mass of fuel, when fuel is fully oxidized. Energy density varies across the spectrum of materials humans have learned to burn. A dry log of oak releases about sixteen megajoules per kilogram, bituminous coal around twenty-four, crude petroleum forty-two, and refined natural gas fifty-five. For calibration a cucumber sandwich delivers roughly eight to ten megajoules per kilogram and this doubles if it is a peanut butter sandwich.

A pound of TNT releases its energy in microseconds rather than the minutes or hours over which a coal fire burns. This difference is described as power density in contrast to energy density. TNT holds about a fifth the chemical energy per kilogram of coal. At the top of the energy-density spectrum is fissile uranium-235, with an energy density some two million times that of coal, and fissile plutonium-239 higher still. What differentiates a sandwich from a stick of TNT from a kilogram of plutonium is how tightly matter is bound but how rapidly the energy can be released.

The archive of fossil carbon is dense in a way that no living network can match. The standing biomass of every forest on Earth, harvested and burned in a year, would not power the global economy for half of one. Human metabolic output peaks near seventy-five to one hundred watts. A single industrial gas turbine, fed by methane drawn from a Mesozoic marine deposit, delivers six hundred megawatts and runs all year. The disparity between flow and stock is the basis on which the modern world has been built.

Extraction

Surface coal outcrops were burned locally for millennia. Burning an outcrop, however, is not the same as accessing the depths of the archive. Deep extraction required pumps to lift water, fans to move air, and lifts to move both men and materials. The energy cost of infrastructure often approached the energy yield of the coal itself, ensuring that the archive remained thermodynamically locked. The novelist Émile Zola descended into a working pit at Anzin in February of 1884 and came back with the material for Germinal. The mine he calls Le Voreux, the Voracious One, is described from the first page as a hunched, brick-built beast. In one passage of the novel his protagonist Étienne thinks of the mine as a "crouching and sated god, to whom ten thousand starving men gave their flesh without knowing it." On his descent Zola was shown a Percheron workhorse in the pit and asked how the animal was got in and out each day. The horse, he was told, came down once as a foal, grew up, went blind and died, and was buried in the darkness below ground.

Newcomen's atmospheric engine of 1712 was invented to pump water from coal mines. The earliest engine had as its primary purpose the extraction of the fuel that, a generation later, powered its descendants. Each technological generation extracted more useful work from the same mass of coal, and the surplus was plowed back into mining, drainage, transport, and metallurgy, made the next generation possible. The economic historian E. A. Wrigley has argued that the English Industrial Revolution is best understood as the crossing of a thermodynamic threshold past which extracted energy began to subsidize its own extraction.

The same cybernetic logic governs every general-purpose technology that has followed. The internal combustion engine made petroleum extraction economic. Petroleum extraction made internal combustion universal. The dynamo of the late nineteenth century turned coal into electricity, which turned electricity into the driver of the chemistry, the metallurgy, and silicon that made the dynamo obsolete. Geoffrey West has written about this kind of cascade in the context of cities and innovation, where the running-up of a positive feedback loop produces the characteristic super-linear scaling of technological output with population. Joseph Schumpeter described the same phenomenon as creative destruction.

Each generation of technology extracted more useful work from the same mass of fuel. The figure below traces this trajectory.

Thermal efficiency of fossil-fuel engines from 1712 to 2020, rising from ~0.5% (Newcomen) to ~62% (combined-cycle gas) along a sigmoid curve.
Thermal efficiency of fossil-fuel engines, 1712–2020. From Newcomen's atmospheric engine at ∼0.5% to the modern combined-cycle gas turbine at ∼62%: a two-order-of-magnitude gain over three centuries. The sigmoid fit suggests we are approaching the thermodynamic ceiling.

The data behind this curve are assembled from several primary sources, each with a particular history. The early Newcomen and Watt efficiencies are reconstructed from contemporary surveys made by Cornish mining engineers, who in the late eighteenth century compiled the so-called "duty" figures of working pumping engines: the foot-pounds of water lifted per bushel of coal burned, a unit invented for accounting and only later translated into thermodynamic efficiency. John Kanefsky and John Robey reconstructed these data systematically in the 1970s; their figures remain the standard. The mid-nineteenth-century points come from R. L. Hills's history of stationary steam, which draws on engineering reports of compound and triple-expansion engines installed in textile mills and on transatlantic steamers. The twentieth-century numbers come from publications of central station utilities, the U.S. Energy Information Administration, and the International Energy Agency. The whole curve is a history of accountancy as much as a history of engineering.

"Extracted energy began to subsidize its own extraction." The machine that opened the archive ran on its own contents.

Entropy

Combustion converts low-entropy fuel into high-entropy waste in the form of carbon dioxide, nitric oxides, ash, and heat. Every joule extracted from the archive is a joule permanently lost from a stock that took a hundred million years to assemble. The thermodynamic archive is read once and only once. It follows a combusting cryptographic logic worthy of an agent from Mission Impossible.

Humanity now burns roughly eight billion tonnes of coal, four billion tonnes of crude oil, and four trillion cubic metres of natural gas every year. The combined output, expressed as carbon dioxide, is approaching forty gigatonnes annually, or about 1,300 tonnes every second. A 2003 calculation by the ecologist Jeff Dukes, since refined, found that the fossil fuels burned in a single year of contemporary global activity correspond to roughly four hundred years' worth of the Earth's net primary production. This means that every year we metabolize four centuries of historical photosynthesis. A 2015 paper in the Proceedings of the National Academy of Sciences added a complementary figure. In the previous two thousand years, humanity has consumed roughly half of the planet's stored biomass, and ten percent of that consumption has occurred in the last hundred years.

4 trillion
Tonnes of recoverable fossil carbon
400 yrs
Of photosynthesis burned each year
55 MJ/kg
Natural gas energy density
2,000,000×
U-235 over coal
Part Two

The Knowledge Archive

Deposition

The knowledge archive could be said to begin with the temple administrators of late-fourth-millennium Sumer. The archaeologist Denise Schmandt-Besserat, working through Mesopotamian collections in the 1970s, found that proto-cuneiform tablets had evolved out of an older system of accounting tokens, small clay shapes used to record commodities of trade. The earliest tablets from Uruk, dated to roughly 3200 BCE, list quantities of barley, beer, and sheep. From this literary seedbed grew the laws of Hammurabi, the Egyptian Book of the Dead, the Vedas, the Iliad, the Confucian Analects, the Hebrew Bible, and the dialogues of Plato.

The greatest of the ancient repositories was the Library of Alexandria, founded under Ptolemy I in the early third century BCE, whose cataloguing project, the Pinakes, ran to one hundred and twenty scrolls. Estimates of the holdings at the library's peak vary from forty thousand scrolls to seven hundred thousand. Some part of this knowledge was lost in Caesar's siege fire of 48 BCE and the Christian destruction of the Serapeum in 391 CE. Roger Bagnall has argued that the more important loss was a slow one — papyrus has a working lifespan of one or two centuries in most climates, and any text not transferred to the codex during the third or fourth century CE was effectively lost regardless of fire.

The total corpus of human written output now numbers in the tens of trillions of words. Google's 2010 estimate of distinct book titles ever published, 130 million, is generally regarded as a lower bound, and ignores manuscripts, periodicals, and the digital sediment of the last fifteen years. Each surviving document is the distillation of an observation, an inference, or a computation, retaining the structure it had at the moment of its creation.

Extraction

For most of literate history, the knowledge archive was accessed through similar network topologies as pre-industrial energy, one reader at a time. The bandwidth of human reading is set by cognitive constraints. The eye fixates on a printed word for two hundred to two hundred and fifty milliseconds, during which a span of about seven to twelve characters is processed. Saccades between fixations occupy another forty milliseconds. The result is a sustainable reading rate of roughly two hundred and fifty words per minute. A student working an eight-hour day might process about a hundred and twenty thousand words. Assuming forty productive years of work, this adds up to a paltry one to two thousand books. The Argentinian archmage of poetry and the short story captured this sentiment more exactly than he had any right to when he complained, "Sometimes, looking at the many books I have at home, I feel I shall die before I come to the end of them, yet I cannot resist the temptation of buying new books."

The great medieval collections were archives in the geological sense. They were vast, largely inert, and tapped at the rate of individual attention. The Vivarium of Cassiodorus, founded in sixth-century Calabria, may have held a few hundred volumes; the Abbey of St Gall in the ninth century catalogued just over four hundred. By the late Middle Ages the great monastic libraries of northern Europe had grown to a few thousand. The Vatican Library, refounded by Sixtus IV in 1475, held some 3,500 volumes by his death and would take three more centuries to clear a hundred thousand. The British Museum library held about half a million volumes by 1856, when Sydney Smirke's new Round Reading Room was built with shelving for a million. And the Library of Congress passed a million books in 1901 and ten million in 1950, and now holds about one hundred and seventy million items, most of them never read by anyone but the cataloguer who barely scans their surface.

This was the knowledge system into which the social network of scholarship, including universities, learned societies, and journals, was inserted. It functioned like a trophic web through which knowledge flowed, bounded by the number and bandwidth of its links. A scholar in early-modern Leiden or eighteenth-century Edinburgh participated in a system that was not very different from the food web Elton would describe two centuries later. Each individual was a node consuming texts produced upstream and contributing texts to be consumed downstream, with most of the upstream output dissipating, like Elton's biomass, into the heat of forgotten readings.

Within decades, scholars were complaining that they were drowning in information. The historian Ann Blair, in Too Much to Know, has summarized many of these complaints. Erasmus, in the Adages, asked: "Is there anywhere on earth exempt from these swarms of new books?" Conrad Gesner, in his Bibliotheca universalis of 1545, complained of a "confusing and harmful abundance of books." Leibniz, in a letter of 1680, feared a return to barbarism, to which result, he wrote, "that horrible mass of books which keeps on growing" might contribute very much. The complaint is older than the printing press. Ecclesiastes 12:12 (KJV) records: "of making many books there is no end; and much study is a weariness of the flesh."

The transformer architecture, introduced in 2017, and the large language models built upon it, are the first extraction technology for the knowledge archive. Earlier technologies — indexing, library cataloguing, full-text search — were technologies of retrieval. They located documents and pointed users to them but they did not derive work from the documents. A search engine returns a document whereas a large language model computes using its content. Trained on billions of documents, transformer models compress the statistical structure of the entire corpus into trillions of parameters, distilling regularities, relationships, and inferential patterns that no human reader could process within the small number of books allotted to a readerly lifespan. This is what makes inference resemble combustion: in both cases, useful work is performed on a static deposit and what is delivered to the user is useful work.

Bar chart comparing deposition and extraction timescales for energy and knowledge archives, with dark slivers at right representing the brief extraction era.
The asymmetry between deposition and extraction. Both archives accumulated over timescales vastly exceeding the period of technological extraction. The dark slivers at right represent the extraction era, a vanishing fraction of the deposition era in each case. Ratios at right: deposition time to extraction time.

A search engine returns a document. A language model computes using its content.

Part Three

The Density Archives

The structural parallels between the energy and the knowledge archives are profound but there is one essential asymmetry between them that is likely to define the trajectory of the coming century, and perhaps distinguish progress in artificial intelligence from progress in mechanical engineering — the knowledge archive is not consumed by extraction.

When coal is burned it is depleted because the combustion is irreversible. When a language model is trained on a text corpus the text remains. The same document can in principle be processed, compressed, and extracted from an unlimited number of times. Newton's Principia is not degraded when it is read by a student or ingested by a physics game engine. Euclid's axioms are not reduced when one proves the infinitude of the primes from them. Knowledge, in this sense, is not a stock but a renewable, and even an inflating, capital.

This non-depletion property has consequences. In thermodynamics, efficiency improvements extend the life of a finite resource but cannot escape the constraint of exhaustion. In knowledge extraction, efficiency improvements compound without limit. If a model trained on ten trillion tokens achieves a certain performance level, and a future architecture matches that performance on one trillion tokens, the saving is pure gain. No conservation law is violated and no entropy debt should accumulate in the archive itself.

Unfortunately every extraction process produces entropy in the form of waste. Fossil fuel combustion produces physical entropy in the form of carbon dioxide, particulates, and heat. Knowledge extraction produces social entropy in the form of hallucinations, confabulation, and misattribution. Modern pollution is not only particulate but ludicrous.

The data centers in which language models are trained and run consume electricity at industrial-revolution scale. The International Energy Agency's 2025 Energy and AI report puts global data-center consumption at roughly 415 terawatt-hours in 2024, about 1.5 percent of global electricity use, growing at twelve percent a year, and projected to more than double to 945 terawatt-hours by 2030. Training a single frontier model can require tens of gigawatt-hours. AI inference, the operation of trained models in deployment, is rapidly becoming the larger share. The largest data-center buildouts of 2025 are sited, geographically, where the largest coal-fired power stations were sited a century ago.

A coal seam is not made worse by being half-burned and the unburned half is still good coal. A knowledge archive can be poisoned. If unreliable extractor outputs are re-deposited into the archive, and subsequent extractors are trained on the contaminated archive, the contamination propagates forward and amplifies. Shumailov and colleagues at Oxford and Cambridge showed in 2023 that successive generations of generative models trained on each other's outputs suffer "model collapse," a progressive loss of the tails of the original distribution and a concomitant collapse of lexical, syntactic, and semantic diversity. Subsequent theoretical work demonstrated that even small fractions of synthetic data, on the order of one percent of the corpus, can be sufficient to induce strong collapse if the training is iterated. The loss is not in the size of the archive but in its veracity. The energy archive degrades through depletion and the knowledge archive degrades through contamination.

The energy archive degrades through depletion. The knowledge archive degrades through contamination.

Part Four

The Minimality Loophole

A thought experiment. What is the smallest set of texts from which a sufficiently powerful extraction engine could reconstruct the core findings of science, law, engineering, and mathematics?

Knowledge, unlike energy, is generative. Reading Darwin's Origin of Species does not consume any of the prey species that he so lovingly describes. A knowledge archive is closer to a seed bank than to a coal seam, and what is extracted can be grown, recombined, and replanted. This is the special property of knowledge that can make it resistant to human avarice.

The question can be framed empirically. For a given model architecture and a fixed performance benchmark, define C*(t) as the minimal corpus size required at time t. Tracking C* over successive generations of extraction technology yields the knowledge-extraction analogue of the thermal-efficiency curve from Newcomen to the gas turbine. If C* falls sufficiently low, then a relatively small, curated library, perhaps a few hundred canonical texts, suffices as the seed corpus for an extraction engine of arbitrary capability. It would be a library built on the logic of the Svalbard Global Seed Vault which seeks to preserve the kernel from which all agriculture might be reconstituted.

The thought experiment is clearly not uniform across domains of knowledge, and one might expect formal disciplines like mathematics to differ from data-rich disciplines like biology in their dependence on a store of knowledge.

Zermelo-Fraenkel set theory and Peano's axioms for arithmetic can both be fit on a single page of paper. From these primitives much of working mathematics is, in principle, derivable. But this obscures the labor of building an edifice out of the smallest parts. Russell and Whitehead's Principia Mathematica, attempting to ground arithmetic in pure logic, reached a theorem for 1 + 1 = 2 only at proposition *54.43 on page 379 of volume one. There is the tone of a pyrrhic victory when Russell and Whitehead write "from this proposition it will follow, when arithmetical addition has been defined, that 1 + 1 = 2." And the complete derivation came hundreds of pages later in volume two. Fermat's Last Theorem took three and a half centuries to be proved and the classification of finite simple groups required tens of thousands of pages of collaborative work. This kind of derivation, from elementary building blocks, can be very expensive. There is also no guarantee that it will work. In 1931 Gödel demonstrated that any consistent formal system rich enough to express elementary arithmetic must contain true statements that cannot be derived from its axioms.

The Game of Life provides a somewhat more optimistic foundation. Conway's two-dimensional cellular automaton has four rules of birth and death and its rule set is shorter than Peano's. The long-term behavior of this dynamical system includes oscillators, gliders, spaceships, glider guns, and constructions that implement a universal Turing machine. The minimal corpus of the game spans four lines but the derivative corpus of phenomena includes the space of computable functions. And music shows a similar pattern. Twelve pitches, a relatively small number of conventional formal structures, and a handful of harmonic conventions generate the Western canon and an unbounded space of further compositions. Short rule sets can produce expressive spaces of arbitrary depth, even where Gödel's devastating result prevents any closed system from being complete.

The empirical sciences are altogether different. To rederive special relativity, an extraction engine would need the Michelson-Morley result, the structure of Maxwell's electrodynamics, and the principle of the constancy of the speed of light. Not to mention the intuition derived from a handful of thought experiments concerning trains, lightning, and clocks. To rederive quantum mechanics would require the blackbody curves, photoelectric data, atomic spectra, and the diffraction patterns of single particles. Rules are not enough. And in biology, the theory of evolution by natural selection required the Voyage of the Beagle, since any amount of harvesting from prior works of natural history would simply perpetuate erroneous orthodoxies.

Michael Hla, taking up a challenge proposed by Demis Hassabis, recently trained a small transformer on twenty-two billion tokens of pre-1900 text and prompted it to make predictions about later experimental observations. His Machina Mirabilis model fails on most tasks. The minimal corpus for theoretical physics is not encoded in shelves of canonical books but in books together with the encoded observational record, and intuitions derived from a disciplinary culture.

Coal when burned is gone and knowledge when processed remains. The principal of any given knowledge trust fund is not depleting, provided the stock is not poisoned by the re-deposition of confabulated extractor outputs. The energy archive imposed a hard ceiling on extraction set by exhaustion. The knowledge archive imposes no such ceiling and it imposes a different one, set by the fidelity of the extraction itself. What gets stored back into the archive determines what subsequent extractors can recover. This means that the future growth of artificial intelligence requires that human beings act as their filters, thereby avoiding what Ted Chiang has described as "valuable human-generated text being drowned out in a sea of AI-generated nonsense."

David C. Krakauer is the president and William H. Miller Professor of Complex Systems at the Santa Fe Institute.

Share this story