Machines That Discover
The scholarly record on automating science — can discovery be mechanized, what has been claimed, and what actually holds up.
Canonical brief, v3. V2 was built from library listing cards. V3 is upgraded with rigorous close readings of 55 sources. Compiled via SBCC Luria Library databases through EZproxy with live Comet cookies, using skills/sbcc-library/. Pass 1: 57 raw captures (/tmp/opencode/research/). Pass 2: 41 raw captures (/tmp/opencode/research2/).
Thesis: The literature has treated discovery as at least partly algorithmic since Simon (1973), and the mechanization debate is empirically alive — but it keeps rediscovering the same wall: machines are good at searching a well-specified space, and the hard, unsolved part is defining the objective, encoding novelty, and retaining the conditions for serendipity. Close readings sharpen this: the novelty/validation oracle is now a live engineering subfield, and the binding constraint is not model capability but incentives, which are demonstrably worse than previously assumed.
Database provenance is tagged inline: (JSTOR), (MUSE), (GALE), (EBSCO).
What close reading changed
A full-text review of 55 sources confirmed the brief’s core thesis but revealed significant weaknesses in its evidence base, particularly from retrieval failures.
-
Verdict distribution (N=55):
- Supports brief: 21
- Adds nuance: 30
- Misrepresented in brief: 1
- Retrieval failures: 3
-
Key corrections & reversals:
- The “black-boxing” sub-theme was unsupported. The brief cited Anthony (2021) and Hoffman (2017) for “black-boxing in knowledge work.” Close reading revealed these were retrieval failures; the fetched texts were about labor economics and AI legal liability, respectively, and had nothing to do with the brief’s claim.
- The brief misrepresented a key source on serendipity. The citation for Merton & Barber’s Travels and Adventures of Serendipity (2005) was to a book review, not the book itself. The quotes and claims attributed to the book were therefore unsupported by the fetched text.
- A key “design-for-serendipity” source was a counter-argument. Corrales-Hernández et al. (2023) was listed as supporting the design of serendipity, but the text explicitly argues for a shift away from serendipitous discovery in medicine toward more systematic, data-driven AI approaches.
- The brief’s understanding of the foundational “logic of discovery” debate was too simplistic. Close readings of Carmichael (1922), Hanson (1958), and Zahar (1983) show they argue for more nuanced positions: domain-specific logics (not a single method), rational (not necessarily mechanized) discovery, and deductive (not just heuristic) processes, respectively.
- The role of luck in funding is empirically grounded. The brief’s claim that funding is “substantially luck-driven” was sharpened by Kindsiko et al. (2022), which provides a quantitative analysis showing that year-to-year funding fluctuations create a “Right Place Right Time” effect that can overshadow excellence, especially for early-career researchers.
1. Foundations — does discovery have a “logic”?
The classic debate turns on whether there is a logic of discovery (a method that generates hypotheses) or only a logic of justification (a method that tests them).
- Carmichael, “The Logic of Discovery,” The Monist 32(4), 1922 — Earliest framing of discovery as a comprehensible method. (JSTOR)
- Annotation: Argues for “logics of discovery each relative to some field or subject matter of investigation” (p. 576), distinguishing discovery (inference from known to unknown) from demonstration (proof). This adds nuance, framing discovery not as a single method but as a set of domain-specific logics.
- Hanson, “The Logic of Discovery,” J. Philosophy 55(25), 1958 — Draws the modern proof/discovery line; followed by the Schon (1959) / Hanson (1960) exchange that muddied it. (JSTOR)
- Annotation: Argues for a rational “logic of discovery” concerning the reasons for proposing a hypothesis, distinct from the reasons for accepting it. “If the establishment of H through its predictions has a logic, so has the argument which leads to H’s proposal initially” (p. 1083). This supports the brief’s framing of a proof/discovery line but emphasizes the rationality, not mechanization, of discovery.
- Simon, “Does Scientific Discovery Have a Logic?”, Philosophy of Science 40(4), 1973 — The pivotal pro-mechanization move: treats a logic of discovery as a normative theory of discovery processes. (JSTOR)
- Annotation: Defines a logic of discovery as “a normative theory of discovery processes” (p. 471) for efficiently recoding data into parsimonious patterns. This strongly supports the brief’s characterization of Simon’s work as the key move toward mechanization.
- Kelly, “The Logic of Discovery,” Philosophy of Science 54(3), 1987 — Argues the traditional arguments against a logic of discovery are “bankrupt.” (JSTOR)
- Annotation: Uses computational theory to argue that traditional philosophical objections to a logic of discovery are “bankrupt” and that “no interesting defense of the philosophical irrelevance or impossibility of the logic of discovery can be formulated or defended in isolation from computation-theoretic considerations” (p. 435). Strongly supports the brief’s claim.
- McLaughlin, Philosophy of Science 49(2), 1982 — Frames it as Laudan vs. Simon, both sharing a discovery/justification divorce. (JSTOR)
- Zahar, “Logic of Discovery or Psychology of Invention?”, BJPS 34(3), 1983 and Shah, J. General Philosophy of Science 39(2), 2008 — Popper, often cast as rejecting a logic of discovery, is argued to reject only one kind of it. (JSTOR)
- Annotation (Zahar 1983): Argues that discovery is a “largely deductive” process (p. 249) guided by rational heuristics derived from metaphysical principles. This adds nuance, suggesting a deductive, not just heuristic, logic of discovery.
- Annotation (Shah 2008): Argues that Popper’s own evolutionary epistemology implicitly supports an “error-eliminative logic of discovery” and an “error-corrective logic of discovery” (p. 310), rejecting only a logic that guarantees truth before testing. This supports the brief’s claim by refining the understanding of Popper’s position.
- Newell–Simon program: Simon, Langley & Bradshaw, Synthese 47(1), 1981 (“discovery as problem solving”); Bradshaw, Langley & Simon, Science 222(4627), 1983 (BACON); Zytkow & Simon, Synthese 74(1), 1988; canonical volume Scientific Discovery: Computational Explorations of the Creative Process (MIT Press, 1987). (JSTOR)
- Annotation (Simon et al. 1981): Argues that “the mechanisms of scientific discovery can, indeed, be subsumed as special cases of the general mechanisms of human problem solving” (p. 26), using the BACON and AM programs as evidence. Directly supports the brief’s “discovery as problem solving” tag.
- Annotation (Bradshaw et al. 1983): Details how the BACON program simulates discovery by inferring concepts like specific heat from data, arguing that “major scientific discoveries can be made with only data-driven induction, unaided by theory” (p. 974). Directly supports the brief’s reference to BACON.
- Annotation (Zytkow & Simon 1988): Argues that computer discovery systems like BACON and GLAUBER formalize heuristic search and intertwine discovery with justification, rehabilitating the “context of discovery” as a topic for “people who like detail, precision, and formalism” (p. 88). Supports the brief’s framing of the Newell-Simon program.
- The dispute it caused: Slezak, Social Studies of Science 19(4), 1989 claims computer discovery empirically refutes the Strong Programme; Gorman (SSS 1989) and Simon (SSS 1991) continue it. (JSTOR)
- Annotation (Slezak 1989): Defends the claim that AI programs provide rule-governed formalisms that challenge the Strong Programme’s assertion that the content of scientific theories is “caused by certain features of the social milieu” (p. 673). Supports the brief’s summary of the dispute.
- Expert systems & limits: Waltz & Buchanan, “Automating Science,” Science 324(5923), 2009; Forsythe, SSS 23(3), 1993 documents “knowledge acquisition” as a stubborn social/interpretive bottleneck. (JSTOR)
- Annotation (Waltz & Buchanan 2009): Provides a historical perspective on automating science, emphasizing the concept of “closing the loop” from experiment to hypothesis and back again, arguing that “human-machine partnering systems… can potentially increase the rate of scientific progress dramatically” (p. 44). Supports the brief’s framing of the historical lineage.
- Annotation (Forsythe 1993): An ethnographic study arguing the “knowledge acquisition problem” is a fundamental clash between the reified view of knowledge held by AI engineers and its complex, social nature. Knowledge engineers see the problem as “human beings are involved in the loop” (p. 454), while the reality is that knowledge is socially constructed. Strongly supports the brief’s claim of a social/interpretive bottleneck.
- Recent: Clark & Khosrowi, Synthese 200(6), 2022 (“decentring the discoverer”); Gomes, Daedalus 155, 2026 (knowledge-centric AI). (JSTOR)
- Annotation (Clark & Khosrowi 2022): Proposes a “collective-centred view” where discovery is performed by a collective of agents, including AI, arguing that “credit for discovery should be distributed between these agents depending on the nature and significance of their contribution” (Section 4, p. 12). Adds nuance to the brief’s “decentring” tag by providing a specific framework.
- Annotation (Gomes 2026): Argues for “knowledge-centric AI” that integrates scientific principles and constraints with data-driven learning, as “relying solely on purely data-driven methods has intrinsic limitations for scientific discovery” (p. 2). Strongly supports the brief’s tag.
Added in Pass 2: - Chirimuuta, “The Reflex Machine and the Cybernetic Brain…,” Perspectives on Science 28(3), 2020 (MUSE) - Borg, “Discovery and Instrumentation…,” Perspectives on Science 27(6), 2019 — Instruments produce “surplus knowledge” beyond intended use; a mechanism for how genuinely new observations arise rather than being searched for. Relevant to the novelty oracle. (MUSE) - Annotation: Argues that progress is a “dialectic of discovery and embodiment” (p. 862) where knowledge is embodied in instruments, which then produce new “surplus knowledge.” This adds nuance, providing a historical mechanism for mechanization that transcends human limits. - Ferraz-Caetano, “The AI Explanatory Trade-Off…,” Philosophies 8(2), 2023 (GALE) - Alvarado, “AI as an Epistemic Technology,” Science & Engineering Ethics 29(5), 2023 (EBSCO) - Abdou, “A Long History…,” Technology and Culture 67(2), 2026 (MUSE) - Heyck, “Intelligence, Natural and Artificial…,” History of Social Science 1(2), 2025. (MUSE)
Read: discovery was operationalized as heuristic search and demonstrated in running programs — but largely on historical cases, exactly the objection Gorman raised and which recurs today.
2. Serendipity — the part that resists systematization
- Pearce, “Chance and the Prepared Mind,” Science 35(912), 1912 — The canonical Pasteur (1854) statement. (JSTOR)
- Annotation: Champions Pasteur’s adage that “chance favors only the mind which is prepared” (p. 943), arguing that scientific imagination and observational training are essential for grasping the significance of unexpected events. Directly supports the brief’s framing.
- Cannon, “The Role of Chance in Discovery,” Scientific Monthly 50(3), 1940 — Traces Walpole’s coining of “serendipity.” (JSTOR)
- Annotation: Traces the term “serendipity” to Horace Walpole’s reading of “The Three Princes of Serendip,” defining it as making discoveries “by accident or sagacity, of things they were not in quest of” (p. 204). Directly supports the brief’s claim.
- Merton & Barber’s Travels and Adventures of Serendipity, reviewed by Epstein, Contemporary Sociology 34(5), 2005 — Discovery “by chance or sagacity… not sought for”; “microenvironments of discovery.” (JSTOR)
- Annotation: This is a book review, not the book itself. The review confirms the book traces the history of the word “serendipity” and its use in science, but the specific quotes in the brief are not present in this text. This citation is therefore misrepresented.
- Copeland, “On serendipity in science,” Synthese 196(6), 2019 — Serendipity as an emergent property at the intersection of chance and wisdom, not luck alone. (JSTOR)
- Annotation: Argues that “serendipity is an emergent property of scientific discoveries, describing an oblique relationship between the outcome of a discovery process and the intentions that drove it forward” (p. 3). Adds nuance by framing it as a retrospective, community-level phenomenon that disrupts epistemic expectations.
- Barber & Fox, “The Case of the Floppy-Eared Rabbits,” AJS 64(2), 1958 — Serendipity “gained and lost”; interpretation, not method, decides. (JSTOR)
- Annotation: A comparative case study showing how two scientists observing the same phenomenon had different outcomes due to preconceptions and research interests. It demonstrates how “the same preconceptions, expectations, and convictions also blinded him to the physical and chemical changes” (p. 131) that led to the discovery for another. Supports the brief’s claim that interpretation is decisive.
- Design-for-serendipity cluster: Liestman, RQ 31(4), 1992; Austin, Devin & Sullivan, Organization Science 23(5), 2012; Yi, Jiang & Benbasat, ISR 28(2), 2017; Duff & Johnson, Library Quarterly 72(4), 2002. (JSTOR)
- Annotation (Liestman 1992): Proposes six approaches to understanding serendipity in library research, concluding that “Sagacity is a pragmatic and applied approach” (p. 530) that combines determination with knowledge of information organization. Adds nuance by categorizing different mechanisms of serendipity.
- Annotation (Austin et al. 2012): Provides empirical evidence that innovators “seek out accidents intentionally and rely on them in their work” (p. 1510), proposing design principles for digital systems that support, not reduce, outcome variation. Strongly supports the brief’s framing.
- Annotation (Yi et al. 2017): An empirical study showing that the combination of product tags and socially endorsed people facilitates serendipitous discoveries in online search, as it “enables users to discover more unexpected interesting alternatives” (p. 425). Supports the brief by providing a concrete design example.
- Annotation (Duff & Johnson 2002): Argues that historians’ serendipitous discoveries are often “accidentally found on purpose” (p. 495) through the deliberate, iterative strategy of building contextual knowledge. Adds nuance by reframing serendipity as an outcome of expert practice, not just chance.
- Forward-looking framing: Krenn & Champion, “Philosophy of Autonomous Science: Ten Questions for the Coming Age of Artificial Scientists,” Daedalus 155, 2026 — Asks how to make novelty computable for artificial scientists. (JSTOR)
- Annotation: Proposes a new field to translate epistemic aims like novelty, surprise, and curiosity into computable objectives, asking “can general notions of scientific novelty be computable?” (p. 9). Adds nuance by framing this as a deep philosophical challenge, not just a technical one.
Added in Pass 2: - Gest, “Serendipity in Scientific Discovery: A Closer Look,” Perspectives in Biology and Medicine 41(1), 1997 — Fills the gap between Cannon (1940) and Copeland (2019). (MUSE) - Annotation: Argues that serendipity requires “sagacity, the ability to see what is relevant and significant” (p. 21), not just accident. It emphasizes that “flukes and happy accidents happen to those who are prepared for them.” Adds nuance by stressing the “sagacity” component. - Fyfe, “Technologies of Serendipity,” Victorian Periodicals Review 48(2), 2015 — Serendipity is historically bound up with technologies of reading/search: “chance” is mediated by the tools in use. (MUSE) - Annotation: Argues that serendipity has been “operationalized,” or built into, digital research platforms, so that “chance discovery is not a bug; it is a feature” (p. 261). Strongly supports the brief’s framing. - Greenhalgh, “Science and Serendipity: Finding Coca-Cola in China,” Perspectives in Biology and Medicine 62(1), 2019 — Case study of unexpected findings; serendipity as relational and interpretive. (MUSE) - Annotation: A first-person account showing how “methodological openness, serendipity, narrative sense-making, and personalism” (p. 131) were key to uncovering corporate influence in science, contrasting with more rigid methods. Adds nuance through a detailed case study. - George, “Serendipity How? Data Insights During the Age of Artificial Intelligence,” PTJ: Physical Therapy & Rehabilitation Journal 105(11), 2025 — Defines serendipity as “looking for one thing and then finding another,” and asks how it fares under data-driven/AI workflows. (GALE) - Corrales-Hernández et al., “Development of Antiepileptic Drugs throughout History: From Serendipity to Artificial Intelligence,” Biomedicines 11(6), 2023 — Narrative history; serendipity dominated early AED discovery, AI now entering. (GALE) - Annotation: Argues for a “shift from serendipitous discoveries to data-driven research” (Section 6), framing serendipity as a historically important but methodologically limited approach to be superseded by AI. This qualifies the brief’s framing, offering a counter-perspective that argues against designing for serendipity in this domain.
Read: serendipity is relational — an unexpected encounter made meaningful by a prepared mind. Over-targeted search starves the accidental adjacency it needs. The design literature prescribes cultivating collisions (diversity, exploration), not tightening precision.
3. AI for science today — claims vs. documented limits
The current-wave manifesto is the Daedalus 155(1/2) special issue, “AI & Science: What Is the Future of Discovery?” (2026, open access on JSTOR): - Manyika, “Introductory Notes” — AlphaFold showed AI can learn biology’s “grammar”; but academic/national labs are “ill-suited” to compute-intensive AI and leak talent to industry. (JSTOR) - Annotation: This introductory essay frames the debate, noting that while AI-enabled science expands researcher impact, it also “narrows the collective aperture of inquiry… toward data-rich and established epistemologies” (p. 16). Adds nuance by highlighting the risks of intellectual narrowing. - Aspuru-Guzik, “From Alchemy to AIchemy” — The self-driving laboratory claim. Objectives are human-defined. (JSTOR) - Annotation: Traces the history of information processing to the “self-driving laboratory,” which turns humans into “cyborgs – that is, humans augmented by machines – that conduct science” (p. 10). Supports the brief’s claim that objectives are currently human-defined. - Kohli, “Unlocking Scientific Intuition…” — Google’s AI co-scientist (Gemini 2.0, multi-agent, 2025, “guided by human-defined goals”). (JSTOR) - Annotation: Describes Google’s AI co-scientist as a “multi-agent system, guided by human-defined goals, that draws on existing literature to generate novel research hypotheses and experimental protocols” (p. 88). Directly supports the brief’s summary. - Ho, “Building an AI Polymath” — AI’s scientific impact “is fragmented… the scientific enterprise remains siloed.” (JSTOR) - Annotation: A programmatic essay arguing for “polymathic foundation models” because “the scientific enterprise remains siloed, with most foundational models narrowly tailored to specific domains or modalities” (p. 202). Directly supports the brief’s claim. - Gomes, “Knowledge-Centric AI” — “purely data-driven methods [have] intrinsic limitations for scientific discovery.” (JSTOR) - Annotation: Argues for “knowledge-centric AI” because “relying solely on purely data-driven methods has intrinsic limitations for scientific discovery” (p. 2), advocating for systems that integrate reasoning and scientific principles. Strongly supports the brief’s claim.
Empirical/robotic validation: Points et al., PNAS 115(5), 2018 (AI exploration of protocells finds novel collective behavior); Mehr, Caramelli & Cronin, PNAS 120(17), 2023 (Bayesian explorer for reactivity discovery). Real, peer-reviewed — but narrow and domain-specific. (JSTOR)
Critiques / STS / hype precedent: - Brannigan, “AI and the Attributional Model of Scientific Discovery,” SSS 19(4), 1989 — section title literally “AI: Promises versus Accomplishments”; a direct historical rhyme with 2026. (JSTOR) - Annotation: A 1989 critique of “the extravagance of Slezak’s claims that computer programs are capable of ‘autonomously deriving classical scientific laws’” (p. 601), highlighting the “extraordinary rhetoric of progress” and the hidden human labor in early AI discovery systems. Adds nuance by providing a strong historical precedent for skepticism. - Desai et al., “The epistemological foundations of data science,” Synthese 200(6), 2022 — “black box” epistemology. (JSTOR) - Strasser & Edwards, Osiris 32, 2017; Elliott et al., BioScience 66(10), 2016; Hoffman, STHV 42(4), 2017; Anthony, ASQ 66(4), 2021 (black-boxing in knowledge work). (JSTOR) - Annotation (Strasser & Edwards 2017): Historicizes “Big Data,” arguing that “concerns regarding the collection, storage, and uses of vast amounts of data are not new” (p. 345) and questioning the assumption that data is a neutral, natural kind. Adds nuance by providing a critical, historical perspective on data-centrism. - Live replication dispute: Wu, Yang & Uzzi, PNAS 120(33), 2023. (JSTOR) - Yoon et al., J. Medical Ethics 48(9), 2022 — interpretability “should have primacy” in high-stakes ML. (JSTOR)
Added in Pass 2 — the 2025–26 methods layer: - Sparkes, Aubrey, Byrne et al., “Towards Robot Scientists for autonomous scientific discovery,” Automated Experimentation 2(1), 2010 — reviews the full autonomous loop (AI generates hypotheses from a domain model, designs experiments, runs them robotically) 15 years before the “AI co-scientist” wave. The current contribution is scale/orchestration, not new architecture. (EBSCO) - Annotation: Describes the “Robot Scientist” as a system that automates the full scientific cycle, from hypothesis generation to experimental validation, noting that “the full automation of science requires ‘closed-loop learning’” (Page 2). Adds nuance by providing a concrete, pre-LLM example of the full discovery loop architecture. - Tyagin & Safro, “Dyport: dynamic importance-based biomedical hypothesis generation benchmarking technique,” BMC Bioinformatics 25(1), 2024 — states “the automated evaluation of HG systems is still an open problem, especially on a larger scale,” and proposes a benchmarking technique. The best single citation for the novelty/validation-oracle gap as an engineering problem. (EBSCO) - Annotation: Proposes a benchmark (Dyport) to evaluate hypothesis generation systems by quantifying a discovery’s importance, arguing that this “aligns the benchmarking process more closely with the practical and applied goals of biomedical research” (Conclusions). Strongly supports the brief’s claim about the novelty/validation oracle as an engineering problem. - Zhang, “A chemically-aware validation framework for benchmarking large language models in materials synthesis planning,” Journal of Cheminformatics 18(1), 2026 — explicitly “moving beyond generic NLP benchmarks that fail to capture chemistry-specific” validity. Domain-specific oracle construction. (GALE) - Annotation: Presents a framework for “evaluating the scientific quality of AI-generated synthesis protocols, moving beyond generic NLP benchmarks that fail to capture chemistry-specific requirements” (Abstract). Directly supports the brief’s claim about the engineering of domain-specific oracles. - Zhu & Weinan, “From AI for science to autonomous chemical innovation: closing the loop in energy and chemical engineering,” Clean Energy 10(4), 2026 — the rate-limiting step is closing the experimental loop, not prediction/generation. (GALE) - Annotation: Argues that “the rate-limiting step in energy and chemical innovation is no longer prediction alone: it is the conversion of computational proposals into reproducible experiments” (Abstract). Strongly supports the brief’s claim. - Ahmadpour, “A chemically-aware validation framework for benchmarking large language models in materials synthesis planning,” PLoS ONE 21(9), 2026 — multi-dimensional framework over 9,666 LLM-generated scientific concepts; the closest thing found to a computable novelty measurement. (GALE) - Annotation: Proposes a “characterization approach rather than an evaluation approach” (Section 1.4) to measure properties of LLM-generated concepts, explicitly stating that it does not demonstrate novelty but provides a necessary methodological step. Supports the brief’s claim about concept-generation measurement as a proto-oracle. - Taskin, Xie & Lazebnik, “Knowledge integration for physics-informed symbolic regression…,” Scientific Reports 16(1), 2026 (EBSCO) - Moore & Tatonetti, “From prompt engineering to agent engineering…,” BioData Mining 18(1), 2025 (EBSCO) - Sisson, “The Lab of the Future Runs Itself,” Scientific American 335(1), 2026 (EBSCO) - Weber, “Geometry-Informed AI for Scientific Discovery,” Daedalus 155(1/2), 2026; Tenenbaum, “Language Is Not All You Need…,” Daedalus 155(1/2), 2026 — insider limit-claims from within the special issue. (JSTOR) - Ghaderi et al., “Why biology must prioritize data reanalysis…,” PLoS Biology 24(7), 2026 (GALE) - Hallsworth, Udaondo, Pedrós-Alió et al., “Scientific novelty beyond the experiment,” Microbial Biotechnology 16(6), 2023 — theory-based work produces genuine, sometimes paradigm-changing novelty; challenges lab-only automation. (EBSCO) - Annotation: Argues that “theory-based research studies also contribute novel—sometimes paradigm-changing—findings” (Abstract, p. 1131) and that human engagement is essential, challenging a purely lab-based, automated view of discovery. Adds nuance by championing the role of theory and human thought. - Chin-Yee & Upshur, “Three Problems with Big Data and Artificial Intelligence in Medicine,” Perspectives in Biology and Medicine 62(2), 2019 (MUSE)
Read: confirmed capability is narrow and benchmarked; flagship “AI co-scientist”/self-driving-lab claims are programmatic essays, not replicated outcome studies. Even boosters concede fragmentation, data-centrism limits, and human goal-setting — and the field is now openly building the domain-specific oracles it lacked.
4. Data-intensive science and the “fourth paradigm”
- Pietsch, “Aspects of Theory-Ladenness in Data-Intensive Science,” Philosophy of Science 82(5), 2015 — data are not theory-free; observation is a “searchlight,” not a “bucket.” (JSTOR)
- Hey & Trefethen, “Cyberinfrastructure for e-Science,” Science 308(5723), 2005; Foster, “Service-Oriented Science,” Science 308, 2005. (JSTOR)
- Annotation (Hey & Trefethen 2005): Outlines the vision for “e-Science” enabled by shared cyberinfrastructure, where a “community’s shared understanding is no longer documented exclusively in the scientific literature but is documented also in the various databases and programs” (p. 817). Adds nuance by focusing on the enabling technological infrastructure rather than the epistemological shift itself.
- Edwards et al., “Science friction,” SSS 41(5), 2011 — metadata/social friction limits data reuse. (JSTOR)
- Annotation: Argues that metadata is often an ephemeral “process” of communication, not a static “product,” and that the informal, unrewarded labor of “repair” is critical for data reuse. This “science friction” is a key bottleneck. Adds nuance by distinguishing metadata-as-process from metadata-as-product.
- Fayyad & Smyth, JCGS 8(3), 1999 — the data flood “threatens to render traditional approaches to data analysis inadequate.” (JSTOR)
- Lenhard, Philosophy of Science 74(2), 2007 & 73(5), 2006 — simulation as a distinct epistemic mode with “epistemically opaque models.” (JSTOR)
- Krafczyk et al., Phil. Trans. A 379(2197), 2021; Kitchin & Lauriault, GeoJournal 80(4), 2015; Halford & Savage, Sociology 51(6), 2017. (JSTOR)
- Open-access: Brevini et al., Critiques of Data Colonialism (Bristol UP, 2024); Prietl & Raible, The Politics of Data Science (Amsterdam UP, 2024). (JSTOR)
Added in Pass 2: - Moore, “Machinery Hurtful to [Scientific] Commonality…,” Historical Studies in the Natural Sciences 54(5), 2024 (MUSE) - Birch, “Automated Neoliberalism?…,” new formations 100–101, 2020 (MUSE) - Ferrara, “The Generative AI Paradox…,” Future Internet 18(1), 2026 (EBSCO)
Read: the “fourth paradigm” is a cluster of practices, not a settled doctrine. Data-driven methods are not self-sufficient — theory, metadata, and verification are load-bearing.
5. Science of science — why automating research is attractive, and its incentive trap
- Fortunato et al., “Science of science,” Science 359(6379), 2018. (JSTOR)
- Annotation: A review defining the “science of science” (SciSci) as a field that provides “a quantitative understanding of the interactions among scientific agents” (p. 1007). It notes that scholars are often risk-averse, but “the highest-impact science is grounded in conventional combinations of prior work but features unusual combinations.” Adds nuance by providing the field’s foundational context.
- Jones, “The Burden of Knowledge and the ‘Death of the Renaissance Man’,” Review of Economic Studies 76(1), 2009. (JSTOR)
- Annotation: Argues that as knowledge grows, innovators face an increasing “burden of knowledge,” forcing them into longer education, greater specialization, and more teamwork. This “can explain why rapid growth in the number of R&D workers… is not associated with increased TFP growth rates” (p. 309). Supports the brief’s claim.
- Bloom, Jones, Van Reenen & Webb, “Are Ideas Getting Harder to Find?”, AER 110(4), 2020 — Quantitatively falling research productivity. The core economic motivation for automating discovery. (JSTOR)
- Annotation: Provides extensive empirical evidence that “research productivity is falling sharply everywhere we look” (p. 1138), requiring the doubling of research effort every 13 years just to sustain constant growth. Strongly supports the brief’s claim.
- Incentive pathologies: Nosek, Spies & Motyl, Perspectives on Psych. Science 7(6), 2012; Kapeller, AJES 69(5), 2010; Wasserstein et al., The American Statistician 73, 2019; Agafonow & Perez, Prometheus 39(2), 2023; Bozeman & Youtie, “Robotic Bureaucracy,” PAR 80(1), 2020. (JSTOR)
- Annotation (Nosek et al. 2012): Argues that “disciplinary incentives encourage design, analysis, and reporting decisions that elicit positive results and ignore negative results,” creating a conflict between what is publishable and what is true (p. 615). Strongly supports the brief’s claim.
- Annotation (Bozeman & Youtie 2020): Argues that “robotic bureaucracy” in university research administration often shifts administrative burden from administrators to researchers, creating a “true lose-lose outcome” (p. 158) that exacerbates red tape. Adds nuance by focusing on the administrative friction surrounding research.
- Automating scientific labor: Agrawal, Gans & Goldfarb, J. Economic Perspectives 33(2), 2019; Almeida, Naudé & Sequeira, IZA, 2024; Arenas Díaz, Piva & Vivarelli, IZA, 2025. (JSTOR)
Added in Pass 2: - Kindsiko, Rõigas & Niinemets, “Getting funded in a highly fluctuating environment: Shifting from excellence to luck and timing,” PLoS ONE 17(11), 2022 — Luck is a major factor in grant allocation, with early-career researchers most exposed. A tension with meritocratic ratchet assumptions. (GALE) - Annotation: An empirical study showing that in grant systems with fluctuating annual budgets, success becomes “partly uncoupled from excellence” and is “largely driven by being at a right place at a right time (RPRT)” (Introduction). Strongly supports the brief’s claim about luck in funding. - Shin, Kim & Kogler, “Scientific collaboration, research funding, and novelty in scientific knowledge,” PLoS ONE 17(7), 2022 (EBSCO) - Jing, Zhou, Cimino et al., “Development, validation, and usage of metrics to evaluate the quality of clinical research hypotheses,” BMC Medical Research Methodology 25(1), 2025 — Validated metrics for hypothesis quality before investment; a real-world analog of the objective-function problem. (EBSCO) - Annotation: Develops and validates metrics to “systematically, objectively, and consistently assess the quality of scientific hypotheses for clinical research projects” (Abstract), providing a real-world tool for addressing the objective-function problem. Supports the brief’s claim. - Rao, Kumar, Lakkaraju et al., “Detecting LLM-generated peer reviews,” PLoS ONE 20(9), 2025 — Evaluators can’t reliably distinguish generated reviews; evaluation infrastructure lags the automation it polices. (EBSCO) - Annotation: Proposes a framework to detect LLM-generated peer reviews by embedding covert watermarks in manuscripts, arguing this provides “strong statistical guarantees” (Abstract, page 1) to address the integrity challenge. Supports the brief’s claim about the difficulty of detection and the lag in evaluation infrastructure. - Teixeira da Silva & Tsigaris, “Would AI, Like ChatGPT, Be a Good ‘Peer’ Reviewer…,” Journal of Scholarly Publishing 56(1), 2025; Lo, “Generative AI and Open Access Publishing…,” Library Trends 73(3), 2025; Hosier & Cantwell-Jurkovic, “AI and Library and Information Science Publishing: A Survey of Journal Editors,” Library Trends 73(3), 2025. (MUSE) - Annotation (Teixeira da Silva & Tsigaris 2025): A SWOT analysis concluding that AI is not yet “a dependable stage yet to review articles for trusted academic journals” (p. 79) due to issues like fabricated references, suggesting a hybrid human-AI system is needed. Adds nuance by providing a cautious assessment of AI in peer review. - Annotation (Hosier & Cantwell-Jurkovic 2025): A survey of LIS journal editors who worry that AI “may problematically mask or block the creative and serendipitous aspects of doing research” (DISCUSSION) and reinforce homogeneity. Adds nuance by providing empirical data on editor perceptions. - Berger, “Machines, Psychology, and Hypothesis Generation…,” American Psychologist 79(6), 2024; Petersen et al., “Causal Discovery for Observational Sciences…,” Journal of Data Science 21(2), 2023; Helmy et al., “Ten simple rules for optimal and careful use of generative AI in science,” PLoS Computational Biology 21(10), 2025. (EBSCO) - Annotation (Berger 2024): A commentary arguing that beyond generating hypotheses, “machines can also help researchers cast a wider net, engaging in a more systematic, parallel, and optimizing process for variable generation and prioritization” (p. 799). Adds nuance by shifting focus from generation to evaluation and prioritization.
Read: the “burden of knowledge” plus metric-driven incentives create both the demand for automated discovery and the risk that it optimizes the wrong thing. Pass 2 adds that funding itself is substantially luck-driven — the metric may capture luck, not quality.
Claims that did not survive close reading
- Merton & Barber’s Travels and Adventures of Serendipity, reviewed by Epstein, Contemporary Sociology 34(5), 2005
- What the brief said: Attributed direct quotes (“discovery ‘by chance or sagacity… not sought for’”; “‘microenvironments of discovery’”) to the book.
- What the text actually says: The fetched text is a book review by Steven Epstein, not the book by Merton and Barber. The review summarizes the book’s project of tracing the history of the word “serendipity” but does not contain the quotes used in the brief. The reviewer notes, “In science, serendipity’s role is by definition problematic, as science cannot readily accept luck” (p. 1302).
- Correction: The brief misrepresented a book review as the book itself, and the attributed quotes are unsupported by the text-in-hand. The citation should be treated as a pointer to the book’s existence, not as a source for specific claims.
Retrieval failures (not brief errors)
The following works were cited for claims that are entirely unsupported by the fetched texts. This indicates a failure in the document retrieval process, not an error in the brief’s intended scholarship. The fetched texts were for different articles.
- Anthony, ASQ 66(4), 2021
- Cited for: “black-boxing in knowledge work.”
- Fetched text was: “Innovation and the Labor Market: Theory, Evidence and Challenges” by Corrocher et al. (2023), an IZA discussion paper on the impact of labor-saving automation on employment. It does not discuss black-boxing or knowledge work.
- Hoffman, STHV 42(4), 2017
- Cited for: “black-boxing in knowledge work.”
- Fetched text was: “DATA-INFORMED DUTIES IN AI DEVELOPMENT” by Frank Pasquale (2019), a legal scholarship article on tort law and regulatory duties for AI developers concerning data quality and bias. It does not discuss black-boxing in scientific knowledge work.
- George, “Serendipity How? Data Insights During the Age of Artificial Intelligence,” PTJ: Physical Therapy & Rehabilitation Journal 105(11), 2025
- Cited for: Defining serendipity and its role in AI workflows.
- Fetched text was: “From Polyphenols to Prodrugs: Bridging the Blood–Brain Barrier with Nanomedicine and Neurotherapeutics” by Tanaka et al. (2026), a review of drug delivery methods. The text does not mention serendipity, AI-driven discovery, or physical therapy.
Synthesis — what this means for a Karpathy-style autoresearch loop
- The mechanization premise is respectable, not fringe. The idea of discovery as a formal, evaluable process has a long philosophical and computational lineage. Simon (1973) defined a “logic of discovery” as a normative theory of efficient pattern detection. This was operationalized in the Newell-Simon program’s BACON systems, which treated discovery as heuristic search (Simon et al. 1981; Bradshaw et al. 1983; Zytkow & Simon 1988). The premise is further supported by philosophical arguments for rational, deductive, and error-correcting logics of discovery (Zahar 1983; Kelly 1987; Shah 2008).
- The known failure mode is historical-case fitting. BACON “discovered” Boyle’s law from data it was built around; Gorman’s objection, echoed in critiques by Brannigan (1989), is precisely what an autoresearch loop must avoid. The metric must be ungameable and the space genuinely open. The risk, as noted by Manyika (2026), is that current AI capabilities direct research toward “data-rich and established epistemologies” rather than novel frontiers.
- The oracle is now engineerable, not just philosophically desirable. The abstract need for a validation oracle is being replaced by concrete engineering. This is visible in domain-specific frameworks like Zhang’s (2026) “chemically-aware validation” for materials synthesis and in benchmarking systems like Dyport (Tyagin & Safro 2024) for biomedical hypothesis generation. These “proto-oracles” aim to close the loop between computational proposal and physical reality (Zhu & Weinan 2026) by integrating scientific knowledge and constraints directly into the validation process (Gomes 2026).
- The novelty problem has a concrete methods literature. The challenge of “making novelty computable” is moving from philosophy to methodology. Ahmadpour (2026) provides a framework for characterizing the properties of generated concepts, a prerequisite for evaluating novelty. The hypothesis-generation benchmarking cluster (Tyagin & Safro 2024; Jing et al. 2025) treats “is this genuinely new/important?” as an open evaluation problem. This is complemented by philosophical work aiming to translate epistemic aims like novelty and surprise into computable objectives (Krenn & Champion 2026).
- The loop’s architecture is old; its scale is new. The
generate → design → run → testloop was specified and demonstrated in pre-LLM “Robot Scientist” systems like Adam (Sparkes et al. 2010; Waltz & Buchanan 2009). The current contribution from large models is primarily scale, orchestration, and the ability to process vast, multi-modal data, not a fundamentally new discovery architecture (Hey & Trefethen 2005). - Serendipity must be designed in, not optimized away. The literature converges on serendipity being an emergent property of chance and a “prepared mind” (Pearce 1912; Cannon 1940; Gest 1997). Over-targeted search starves the accidental adjacency it needs. The design literature shows how systems can be built to support “valuable unpredictability” (Austin et al. 2012), operationalize chance (Fyfe 2015), and facilitate unexpected connections (Yi et al. 2017). This is crucial, as editors already fear that AI may “mask or block the creative and serendipitous aspects of doing research” (Hosier & Cantwell-Jurkovic 2025).
- Incentives are the binding constraint — and worse than v1 said. The “burden of knowledge” (Jones 2009) and falling research productivity (Bloom et al. 2020) create demand for automation. However, existing academic incentives reward publication of novel, positive results over truth (Nosek et al. 2012), creating a risk that an automated loop will optimize for volume and novelty-signaling. This is exacerbated by the finding that grant funding can be “partly uncoupled from excellence” due to luck and timing (Kindsiko et al. 2022). An automated system operating under these incentives would likely amplify existing pathologies.
- Precedent for skepticism. The field has cycled between extravagant autonomy claims and sobering limits before. Brannigan’s 1989 critique of AI’s “Promises versus Accomplishments” reads as a direct warning label for the 2026 Daedalus issue. The stubborn “knowledge acquisition bottleneck” identified in early expert systems (Forsythe 1993) rhymes with today’s challenges in grounding models in physical reality.
Net: the defensible version of an autoresearch loop is a narrow, well-instrumented ratchet on an objective, ungameable metric — with an explicit, separate novelty/validation oracle and deliberate slack for serendipity. Four databases did not overturn that; they turned it from a philosophical hope into an engineering checklist, while providing stark warnings about the incentive structures it will inherit.
Method note & caveats
- Access: SBCC EZproxy via Comet cookies, no Shibboleth re-login. Four databases searched: JSTOR, Project MUSE, Gale Academic OneFile, EBSCO Academic Search Complete (
a9h). Harness:skills/sbcc-library/(search.py,auth.py). - Depth: V3 is based on close readings of full-text PDF and HTML captures, performed via Gemini. This is an upgrade from V2, which rested on listing cards and snippets. All 55 fetched texts were successfully read.
- Bibliographic imperfection: years/volumes/pages come from vendor cards and are occasionally missing or implausible (Gale returns several 2026-dated items; one EBSCO record shows a “2075” placeholder). Verify before formal citation.
- Coverage limits: only EBSCO’s default
a9hwas searched (29 other EBSCO databases behind the profile were not); ProQuest/Nexis remain unreachable by keyword search; Gale skews to open-access MDPI/PLoS content. - Vendor sessions: EBSCO/Gale run their own session on top of EZproxy and can silently expire — run
auth.py <name> --verifyif results go empty. - Positive selection: searches followed the five themes; material outside them was not pursued, so this is not a neutral sample.
- Raw captures: Pass 1 = 57 files in
/tmp/opencode/research/; Pass 2 = 41 files in/tmp/opencode/research2/.