Machines That Discover

2026-09-18

The scholarly record on automating science — can discovery be mechanized, what has been claimed, and what actually holds up.

Canonical brief, v3. V2 was built from library listing cards. V3 is upgraded with rigorous close readings of 55 sources. Compiled via SBCC Luria Library databases through EZproxy with live Comet cookies, using skills/sbcc-library/. Pass 1: 57 raw captures (/tmp/opencode/research/). Pass 2: 41 raw captures (/tmp/opencode/research2/).

Thesis: The literature has treated discovery as at least partly algorithmic since Simon (1973), and the mechanization debate is empirically alive — but it keeps rediscovering the same wall: machines are good at searching a well-specified space, and the hard, unsolved part is defining the objective, encoding novelty, and retaining the conditions for serendipity. Close readings sharpen this: the novelty/validation oracle is now a live engineering subfield, and the binding constraint is not model capability but incentives, which are demonstrably worse than previously assumed.

Database provenance is tagged inline: (JSTOR), (MUSE), (GALE), (EBSCO).


What close reading changed

A full-text review of 55 sources confirmed the brief’s core thesis but revealed significant weaknesses in its evidence base, particularly from retrieval failures.


1. Foundations — does discovery have a “logic”?

The classic debate turns on whether there is a logic of discovery (a method that generates hypotheses) or only a logic of justification (a method that tests them).

Added in Pass 2: - Chirimuuta, “The Reflex Machine and the Cybernetic Brain…,” Perspectives on Science 28(3), 2020 (MUSE) - Borg, “Discovery and Instrumentation…,” Perspectives on Science 27(6), 2019 — Instruments produce “surplus knowledge” beyond intended use; a mechanism for how genuinely new observations arise rather than being searched for. Relevant to the novelty oracle. (MUSE) - Annotation: Argues that progress is a “dialectic of discovery and embodiment” (p. 862) where knowledge is embodied in instruments, which then produce new “surplus knowledge.” This adds nuance, providing a historical mechanism for mechanization that transcends human limits. - Ferraz-Caetano, “The AI Explanatory Trade-Off…,” Philosophies 8(2), 2023 (GALE) - Alvarado, “AI as an Epistemic Technology,” Science & Engineering Ethics 29(5), 2023 (EBSCO) - Abdou, “A Long History…,” Technology and Culture 67(2), 2026 (MUSE) - Heyck, “Intelligence, Natural and Artificial…,” History of Social Science 1(2), 2025. (MUSE)

Read: discovery was operationalized as heuristic search and demonstrated in running programs — but largely on historical cases, exactly the objection Gorman raised and which recurs today.


2. Serendipity — the part that resists systematization

Added in Pass 2: - Gest, “Serendipity in Scientific Discovery: A Closer Look,” Perspectives in Biology and Medicine 41(1), 1997 — Fills the gap between Cannon (1940) and Copeland (2019). (MUSE) - Annotation: Argues that serendipity requires “sagacity, the ability to see what is relevant and significant” (p. 21), not just accident. It emphasizes that “flukes and happy accidents happen to those who are prepared for them.” Adds nuance by stressing the “sagacity” component. - Fyfe, “Technologies of Serendipity,” Victorian Periodicals Review 48(2), 2015 — Serendipity is historically bound up with technologies of reading/search: “chance” is mediated by the tools in use. (MUSE) - Annotation: Argues that serendipity has been “operationalized,” or built into, digital research platforms, so that “chance discovery is not a bug; it is a feature” (p. 261). Strongly supports the brief’s framing. - Greenhalgh, “Science and Serendipity: Finding Coca-Cola in China,” Perspectives in Biology and Medicine 62(1), 2019 — Case study of unexpected findings; serendipity as relational and interpretive. (MUSE) - Annotation: A first-person account showing how “methodological openness, serendipity, narrative sense-making, and personalism” (p. 131) were key to uncovering corporate influence in science, contrasting with more rigid methods. Adds nuance through a detailed case study. - George, “Serendipity How? Data Insights During the Age of Artificial Intelligence,” PTJ: Physical Therapy & Rehabilitation Journal 105(11), 2025 — Defines serendipity as “looking for one thing and then finding another,” and asks how it fares under data-driven/AI workflows. (GALE) - Corrales-Hernández et al., “Development of Antiepileptic Drugs throughout History: From Serendipity to Artificial Intelligence,” Biomedicines 11(6), 2023 — Narrative history; serendipity dominated early AED discovery, AI now entering. (GALE) - Annotation: Argues for a “shift from serendipitous discoveries to data-driven research” (Section 6), framing serendipity as a historically important but methodologically limited approach to be superseded by AI. This qualifies the brief’s framing, offering a counter-perspective that argues against designing for serendipity in this domain.

Read: serendipity is relational — an unexpected encounter made meaningful by a prepared mind. Over-targeted search starves the accidental adjacency it needs. The design literature prescribes cultivating collisions (diversity, exploration), not tightening precision.


3. AI for science today — claims vs. documented limits

The current-wave manifesto is the Daedalus 155(1/2) special issue, “AI & Science: What Is the Future of Discovery?” (2026, open access on JSTOR): - Manyika, “Introductory Notes” — AlphaFold showed AI can learn biology’s “grammar”; but academic/national labs are “ill-suited” to compute-intensive AI and leak talent to industry. (JSTOR) - Annotation: This introductory essay frames the debate, noting that while AI-enabled science expands researcher impact, it also “narrows the collective aperture of inquiry… toward data-rich and established epistemologies” (p. 16). Adds nuance by highlighting the risks of intellectual narrowing. - Aspuru-Guzik, “From Alchemy to AIchemy” — The self-driving laboratory claim. Objectives are human-defined. (JSTOR) - Annotation: Traces the history of information processing to the “self-driving laboratory,” which turns humans into “cyborgs – that is, humans augmented by machines – that conduct science” (p. 10). Supports the brief’s claim that objectives are currently human-defined. - Kohli, “Unlocking Scientific Intuition…” — Google’s AI co-scientist (Gemini 2.0, multi-agent, 2025, “guided by human-defined goals”). (JSTOR) - Annotation: Describes Google’s AI co-scientist as a “multi-agent system, guided by human-defined goals, that draws on existing literature to generate novel research hypotheses and experimental protocols” (p. 88). Directly supports the brief’s summary. - Ho, “Building an AI Polymath” — AI’s scientific impact “is fragmented… the scientific enterprise remains siloed.” (JSTOR) - Annotation: A programmatic essay arguing for “polymathic foundation models” because “the scientific enterprise remains siloed, with most foundational models narrowly tailored to specific domains or modalities” (p. 202). Directly supports the brief’s claim. - Gomes, “Knowledge-Centric AI” — “purely data-driven methods [have] intrinsic limitations for scientific discovery.” (JSTOR) - Annotation: Argues for “knowledge-centric AI” because “relying solely on purely data-driven methods has intrinsic limitations for scientific discovery” (p. 2), advocating for systems that integrate reasoning and scientific principles. Strongly supports the brief’s claim.

Empirical/robotic validation: Points et al., PNAS 115(5), 2018 (AI exploration of protocells finds novel collective behavior); Mehr, Caramelli & Cronin, PNAS 120(17), 2023 (Bayesian explorer for reactivity discovery). Real, peer-reviewed — but narrow and domain-specific. (JSTOR)

Critiques / STS / hype precedent: - Brannigan, “AI and the Attributional Model of Scientific Discovery,” SSS 19(4), 1989 — section title literally “AI: Promises versus Accomplishments”; a direct historical rhyme with 2026. (JSTOR) - Annotation: A 1989 critique of “the extravagance of Slezak’s claims that computer programs are capable of ‘autonomously deriving classical scientific laws’” (p. 601), highlighting the “extraordinary rhetoric of progress” and the hidden human labor in early AI discovery systems. Adds nuance by providing a strong historical precedent for skepticism. - Desai et al., “The epistemological foundations of data science,” Synthese 200(6), 2022 — “black box” epistemology. (JSTOR) - Strasser & Edwards, Osiris 32, 2017; Elliott et al., BioScience 66(10), 2016; Hoffman, STHV 42(4), 2017; Anthony, ASQ 66(4), 2021 (black-boxing in knowledge work). (JSTOR) - Annotation (Strasser & Edwards 2017): Historicizes “Big Data,” arguing that “concerns regarding the collection, storage, and uses of vast amounts of data are not new” (p. 345) and questioning the assumption that data is a neutral, natural kind. Adds nuance by providing a critical, historical perspective on data-centrism. - Live replication dispute: Wu, Yang & Uzzi, PNAS 120(33), 2023. (JSTOR) - Yoon et al., J. Medical Ethics 48(9), 2022 — interpretability “should have primacy” in high-stakes ML. (JSTOR)

Added in Pass 2 — the 2025–26 methods layer: - Sparkes, Aubrey, Byrne et al., “Towards Robot Scientists for autonomous scientific discovery,” Automated Experimentation 2(1), 2010 — reviews the full autonomous loop (AI generates hypotheses from a domain model, designs experiments, runs them robotically) 15 years before the “AI co-scientist” wave. The current contribution is scale/orchestration, not new architecture. (EBSCO) - Annotation: Describes the “Robot Scientist” as a system that automates the full scientific cycle, from hypothesis generation to experimental validation, noting that “the full automation of science requires ‘closed-loop learning’” (Page 2). Adds nuance by providing a concrete, pre-LLM example of the full discovery loop architecture. - Tyagin & Safro, “Dyport: dynamic importance-based biomedical hypothesis generation benchmarking technique,” BMC Bioinformatics 25(1), 2024 — states “the automated evaluation of HG systems is still an open problem, especially on a larger scale,” and proposes a benchmarking technique. The best single citation for the novelty/validation-oracle gap as an engineering problem. (EBSCO) - Annotation: Proposes a benchmark (Dyport) to evaluate hypothesis generation systems by quantifying a discovery’s importance, arguing that this “aligns the benchmarking process more closely with the practical and applied goals of biomedical research” (Conclusions). Strongly supports the brief’s claim about the novelty/validation oracle as an engineering problem. - Zhang, “A chemically-aware validation framework for benchmarking large language models in materials synthesis planning,” Journal of Cheminformatics 18(1), 2026 — explicitly “moving beyond generic NLP benchmarks that fail to capture chemistry-specific” validity. Domain-specific oracle construction. (GALE) - Annotation: Presents a framework for “evaluating the scientific quality of AI-generated synthesis protocols, moving beyond generic NLP benchmarks that fail to capture chemistry-specific requirements” (Abstract). Directly supports the brief’s claim about the engineering of domain-specific oracles. - Zhu & Weinan, “From AI for science to autonomous chemical innovation: closing the loop in energy and chemical engineering,” Clean Energy 10(4), 2026 — the rate-limiting step is closing the experimental loop, not prediction/generation. (GALE) - Annotation: Argues that “the rate-limiting step in energy and chemical innovation is no longer prediction alone: it is the conversion of computational proposals into reproducible experiments” (Abstract). Strongly supports the brief’s claim. - Ahmadpour, “A chemically-aware validation framework for benchmarking large language models in materials synthesis planning,” PLoS ONE 21(9), 2026 — multi-dimensional framework over 9,666 LLM-generated scientific concepts; the closest thing found to a computable novelty measurement. (GALE) - Annotation: Proposes a “characterization approach rather than an evaluation approach” (Section 1.4) to measure properties of LLM-generated concepts, explicitly stating that it does not demonstrate novelty but provides a necessary methodological step. Supports the brief’s claim about concept-generation measurement as a proto-oracle. - Taskin, Xie & Lazebnik, “Knowledge integration for physics-informed symbolic regression…,” Scientific Reports 16(1), 2026 (EBSCO) - Moore & Tatonetti, “From prompt engineering to agent engineering…,” BioData Mining 18(1), 2025 (EBSCO) - Sisson, “The Lab of the Future Runs Itself,” Scientific American 335(1), 2026 (EBSCO) - Weber, “Geometry-Informed AI for Scientific Discovery,” Daedalus 155(1/2), 2026; Tenenbaum, “Language Is Not All You Need…,” Daedalus 155(1/2), 2026 — insider limit-claims from within the special issue. (JSTOR) - Ghaderi et al., “Why biology must prioritize data reanalysis…,” PLoS Biology 24(7), 2026 (GALE) - Hallsworth, Udaondo, Pedrós-Alió et al., “Scientific novelty beyond the experiment,” Microbial Biotechnology 16(6), 2023 — theory-based work produces genuine, sometimes paradigm-changing novelty; challenges lab-only automation. (EBSCO) - Annotation: Argues that “theory-based research studies also contribute novel—sometimes paradigm-changing—findings” (Abstract, p. 1131) and that human engagement is essential, challenging a purely lab-based, automated view of discovery. Adds nuance by championing the role of theory and human thought. - Chin-Yee & Upshur, “Three Problems with Big Data and Artificial Intelligence in Medicine,” Perspectives in Biology and Medicine 62(2), 2019 (MUSE)

Read: confirmed capability is narrow and benchmarked; flagship “AI co-scientist”/self-driving-lab claims are programmatic essays, not replicated outcome studies. Even boosters concede fragmentation, data-centrism limits, and human goal-setting — and the field is now openly building the domain-specific oracles it lacked.


4. Data-intensive science and the “fourth paradigm”

Added in Pass 2: - Moore, “Machinery Hurtful to [Scientific] Commonality…,” Historical Studies in the Natural Sciences 54(5), 2024 (MUSE) - Birch, “Automated Neoliberalism?…,” new formations 100–101, 2020 (MUSE) - Ferrara, “The Generative AI Paradox…,” Future Internet 18(1), 2026 (EBSCO)

Read: the “fourth paradigm” is a cluster of practices, not a settled doctrine. Data-driven methods are not self-sufficient — theory, metadata, and verification are load-bearing.


5. Science of science — why automating research is attractive, and its incentive trap

Added in Pass 2: - Kindsiko, Rõigas & Niinemets, “Getting funded in a highly fluctuating environment: Shifting from excellence to luck and timing,” PLoS ONE 17(11), 2022Luck is a major factor in grant allocation, with early-career researchers most exposed. A tension with meritocratic ratchet assumptions. (GALE) - Annotation: An empirical study showing that in grant systems with fluctuating annual budgets, success becomes “partly uncoupled from excellence” and is “largely driven by being at a right place at a right time (RPRT)” (Introduction). Strongly supports the brief’s claim about luck in funding. - Shin, Kim & Kogler, “Scientific collaboration, research funding, and novelty in scientific knowledge,” PLoS ONE 17(7), 2022 (EBSCO) - Jing, Zhou, Cimino et al., “Development, validation, and usage of metrics to evaluate the quality of clinical research hypotheses,” BMC Medical Research Methodology 25(1), 2025 — Validated metrics for hypothesis quality before investment; a real-world analog of the objective-function problem. (EBSCO) - Annotation: Develops and validates metrics to “systematically, objectively, and consistently assess the quality of scientific hypotheses for clinical research projects” (Abstract), providing a real-world tool for addressing the objective-function problem. Supports the brief’s claim. - Rao, Kumar, Lakkaraju et al., “Detecting LLM-generated peer reviews,” PLoS ONE 20(9), 2025 — Evaluators can’t reliably distinguish generated reviews; evaluation infrastructure lags the automation it polices. (EBSCO) - Annotation: Proposes a framework to detect LLM-generated peer reviews by embedding covert watermarks in manuscripts, arguing this provides “strong statistical guarantees” (Abstract, page 1) to address the integrity challenge. Supports the brief’s claim about the difficulty of detection and the lag in evaluation infrastructure. - Teixeira da Silva & Tsigaris, “Would AI, Like ChatGPT, Be a Good ‘Peer’ Reviewer…,” Journal of Scholarly Publishing 56(1), 2025; Lo, “Generative AI and Open Access Publishing…,” Library Trends 73(3), 2025; Hosier & Cantwell-Jurkovic, “AI and Library and Information Science Publishing: A Survey of Journal Editors,” Library Trends 73(3), 2025. (MUSE) - Annotation (Teixeira da Silva & Tsigaris 2025): A SWOT analysis concluding that AI is not yet “a dependable stage yet to review articles for trusted academic journals” (p. 79) due to issues like fabricated references, suggesting a hybrid human-AI system is needed. Adds nuance by providing a cautious assessment of AI in peer review. - Annotation (Hosier & Cantwell-Jurkovic 2025): A survey of LIS journal editors who worry that AI “may problematically mask or block the creative and serendipitous aspects of doing research” (DISCUSSION) and reinforce homogeneity. Adds nuance by providing empirical data on editor perceptions. - Berger, “Machines, Psychology, and Hypothesis Generation…,” American Psychologist 79(6), 2024; Petersen et al., “Causal Discovery for Observational Sciences…,” Journal of Data Science 21(2), 2023; Helmy et al., “Ten simple rules for optimal and careful use of generative AI in science,” PLoS Computational Biology 21(10), 2025. (EBSCO) - Annotation (Berger 2024): A commentary arguing that beyond generating hypotheses, “machines can also help researchers cast a wider net, engaging in a more systematic, parallel, and optimizing process for variable generation and prioritization” (p. 799). Adds nuance by shifting focus from generation to evaluation and prioritization.

Read: the “burden of knowledge” plus metric-driven incentives create both the demand for automated discovery and the risk that it optimizes the wrong thing. Pass 2 adds that funding itself is substantially luck-driven — the metric may capture luck, not quality.


Claims that did not survive close reading

Retrieval failures (not brief errors)

The following works were cited for claims that are entirely unsupported by the fetched texts. This indicates a failure in the document retrieval process, not an error in the brief’s intended scholarship. The fetched texts were for different articles.


Synthesis — what this means for a Karpathy-style autoresearch loop

  1. The mechanization premise is respectable, not fringe. The idea of discovery as a formal, evaluable process has a long philosophical and computational lineage. Simon (1973) defined a “logic of discovery” as a normative theory of efficient pattern detection. This was operationalized in the Newell-Simon program’s BACON systems, which treated discovery as heuristic search (Simon et al. 1981; Bradshaw et al. 1983; Zytkow & Simon 1988). The premise is further supported by philosophical arguments for rational, deductive, and error-correcting logics of discovery (Zahar 1983; Kelly 1987; Shah 2008).
  2. The known failure mode is historical-case fitting. BACON “discovered” Boyle’s law from data it was built around; Gorman’s objection, echoed in critiques by Brannigan (1989), is precisely what an autoresearch loop must avoid. The metric must be ungameable and the space genuinely open. The risk, as noted by Manyika (2026), is that current AI capabilities direct research toward “data-rich and established epistemologies” rather than novel frontiers.
  3. The oracle is now engineerable, not just philosophically desirable. The abstract need for a validation oracle is being replaced by concrete engineering. This is visible in domain-specific frameworks like Zhang’s (2026) “chemically-aware validation” for materials synthesis and in benchmarking systems like Dyport (Tyagin & Safro 2024) for biomedical hypothesis generation. These “proto-oracles” aim to close the loop between computational proposal and physical reality (Zhu & Weinan 2026) by integrating scientific knowledge and constraints directly into the validation process (Gomes 2026).
  4. The novelty problem has a concrete methods literature. The challenge of “making novelty computable” is moving from philosophy to methodology. Ahmadpour (2026) provides a framework for characterizing the properties of generated concepts, a prerequisite for evaluating novelty. The hypothesis-generation benchmarking cluster (Tyagin & Safro 2024; Jing et al. 2025) treats “is this genuinely new/important?” as an open evaluation problem. This is complemented by philosophical work aiming to translate epistemic aims like novelty and surprise into computable objectives (Krenn & Champion 2026).
  5. The loop’s architecture is old; its scale is new. The generate → design → run → test loop was specified and demonstrated in pre-LLM “Robot Scientist” systems like Adam (Sparkes et al. 2010; Waltz & Buchanan 2009). The current contribution from large models is primarily scale, orchestration, and the ability to process vast, multi-modal data, not a fundamentally new discovery architecture (Hey & Trefethen 2005).
  6. Serendipity must be designed in, not optimized away. The literature converges on serendipity being an emergent property of chance and a “prepared mind” (Pearce 1912; Cannon 1940; Gest 1997). Over-targeted search starves the accidental adjacency it needs. The design literature shows how systems can be built to support “valuable unpredictability” (Austin et al. 2012), operationalize chance (Fyfe 2015), and facilitate unexpected connections (Yi et al. 2017). This is crucial, as editors already fear that AI may “mask or block the creative and serendipitous aspects of doing research” (Hosier & Cantwell-Jurkovic 2025).
  7. Incentives are the binding constraint — and worse than v1 said. The “burden of knowledge” (Jones 2009) and falling research productivity (Bloom et al. 2020) create demand for automation. However, existing academic incentives reward publication of novel, positive results over truth (Nosek et al. 2012), creating a risk that an automated loop will optimize for volume and novelty-signaling. This is exacerbated by the finding that grant funding can be “partly uncoupled from excellence” due to luck and timing (Kindsiko et al. 2022). An automated system operating under these incentives would likely amplify existing pathologies.
  8. Precedent for skepticism. The field has cycled between extravagant autonomy claims and sobering limits before. Brannigan’s 1989 critique of AI’s “Promises versus Accomplishments” reads as a direct warning label for the 2026 Daedalus issue. The stubborn “knowledge acquisition bottleneck” identified in early expert systems (Forsythe 1993) rhymes with today’s challenges in grounding models in physical reality.

Net: the defensible version of an autoresearch loop is a narrow, well-instrumented ratchet on an objective, ungameable metric — with an explicit, separate novelty/validation oracle and deliberate slack for serendipity. Four databases did not overturn that; they turned it from a philosophical hope into an engineering checklist, while providing stark warnings about the incentive structures it will inherit.


Method note & caveats