Interview Alexandre Gefen

AI, Explanation and the Changing Conditions of Scientific Knowledge

Interview with Alexandre Gefen

September 15, 2026

This interview was originally conducted on May 19, 2025.

How does the growing use of artificial intelligence alter what counts as scientific evidence and understanding? In this conversation, Alexandre Gefen, Research Director at the CNRS and a member of the THALIM research unit, reflects on the opacity of connectionist models and the risks that arise when synthetic results circulate as scientific data. He also discusses intellectual responsibility and the conditions under which artificial agents might eventually develop more genuinely creative capacities. Throughout the interview, Gefen considers generative AI a powerful accelerator of research whose use still requires theoretical knowledge and sustained critical control. The interview was conducted by Andreas Sudmann in Paris as part of the media-ethnographic field research of HiAICS. It has been lightly edited for clarity and readability.

Andreas Sudmann: AI systems often excel at prediction on the basis of correlations without providing clear causal explanations. In The Scientific Image, Bas van Fraassen advocates constructive empiricism, according to which the aim of science lies in empirical adequacy rather than necessarily in identifying true underlying causal mechanisms. To what extent might the predictive power of AI encourage scientific disciplines to move toward such an empiricist position? Could it eventually devalue the traditional pursuit of deeper causal understanding?

Alexandre Gefen: This is a major epistemological issue, closely connected to the difficulty of making AI explainable. How should we deal with black boxes and with stochastic systems that move directly toward an answer? Even a reasoning model that is supposed to explain its reasoning, or a model that incorporates bibliographical research, may proceed directly to a conclusion without relying on a clear ontology or causal basis.

In that respect, we are no longer operating in a straightforward Popperian framework. It becomes difficult to establish the conditions under which a result could be shown to be false. At best, one might ask another AI system to contradict the result. The criterion of falsifiability associated with Karl Popper therefore becomes much harder to apply.

Moving directly toward a result could be a serious scientific drawback because it appears to make theory unnecessary. This question emerged some years ago in the well-known article “The End of Theory.” What happens if we merely identify empirical rules or correlations? I do not think we can conduct every form of science without a causal model of the world and without logical laws. This marks a limit of empiricism. At some point, we may therefore witness the return of certain forms of symbolic AI. Yann LeCun, for instance, has argued that the next major step will require more than another reasoning model. It may require ontology and logic. An alliance between connectionist and symbolic AI could become important. For the time being, we need methods that combine causal and theoretical models with the assistance of AI in discovering and verifying scientific laws.

Andreas Sudmann: The persistent opacity of complex AI models creates a related problem. Even when such systems produce accurate results, their black-box character challenges scientific norms that emphasize transparent justification. Alvin Goldman’s reliabilism proposes that justification may depend on the reliability of a belief-forming process, even when the agent has no internal access to its operation. Does the increasing scientific reliance on opaque yet potentially reliable AI systems require a move toward this kind of externalist justification? Would that change what counts as adequately justified scientific evidence?

Alexandre Gefen: Yes, it will change what we consider scientific evidence. The challenge may eventually become comparable to the epistemological revolution associated with quantum theory. Quantum theory confronted us with results that could appear unexpected and unpredictable from one point of view while remaining explicable statistically. There may be a parallel between a quantum world organized through probabilities and a macroscopic world increasingly explained through connectionist models. I do not know how far this parallel can be taken, but we should expect an important epistemological shift toward the use of results that cannot be fully explained.

There is, of course, considerable work underway to make AI more explainable. Even so, biases can make a result completely false. We have to accept a certain imperfection in stochastic results because we can no longer be certain that A will cause B and B will then cause C. Something in the data can make the model diverge from what we would logically expect.

Another issue appears when AI is used across scientific fields. A growing part of our results may be produced by AI, after which we use another AI system to formulate new assumptions on the basis of those results. Consider astronomy, where deep analysis of images is already common. A model might classify galaxies, and another model might then draw predictions about their development from that classification. We would thus be using AI on AI-produced results.

At some point, it may become difficult to distinguish between observational results and synthetic ones, including results that are only partially observational. A related problem arises when one AI system is trained on synthetic data produced by another. This can contribute to what is called model collapse, or at least to a significant model shift. We must therefore continue returning to observational data from the real world. Otherwise, the possibility of model collapse may blind us to what is happening.

Andreas Sudmann: This circulation of AI-generated outputs as inputs for further analysis makes it increasingly difficult to distinguish between human-generated scientific data and synthetic data. Scientific breakthroughs, however, often require researchers to challenge established assumptions about what is relevant. Can systems trained largely on existing data and present paradigms help us identify radically new areas of relevance? Or does their strength primarily lie in optimizing knowledge within established frameworks?

Alexandre Gefen: Statistical methods can use very large quantities of data to identify new correlations. That is certainly possible. In my own field of literary studies, one could discover previously unnoticed relationships between authors or texts by identifying a citation or a form of intertextuality. If a model is given hundreds of thousands of texts, one can ask it to trace the history of a concept or a particular mode of description. It may produce new classifications and reveal previously unseen paths between ideas.

This use differs somewhat from generative AI. Generative systems produce results in relation to a model. When researchers use machine learning to discover correlations, the operation is not necessarily the same as generating language through a large language model. A generative model already involves a form of simplification through the compression of information in a latent space. Because it works with a compressed representation of language or the world, its capacity to discover something genuinely new may be more limited.

The model tends to align its outputs with an expected standard or an average position. Its results are organized around what is most probable according to that model. In this sense, discovering something genuinely unexpected becomes less likely because every answer is mediated through a stochastic model that is often strongly aligned.

Andreas Sudmann: As AI becomes embedded in scientific workflows, it may analyse data, suggest hypotheses, or contribute to the writing of papers. How should we attribute intellectual responsibility and agency under these conditions? Can a non-conscious system share any accountability for a false scientific claim?

Alexandre Gefen: I confronted this question recently when I had to prepare a conference paper in a hurry on a subject I had not previously studied in depth. I gave Deep Research several books that I knew I could trust, together with the general idea of the paper. I corrected the resulting text and added some ideas, while retaining much of its structure. I included a note stating that the paper had been drafted with the assistance of OpenAI’s Deep Research.

One possible response would be to develop a declaration by the researcher. I have seen proposals for three labels indicating that no AI was used, that AI was used, or that a paper was generated by AI. Yet these categories are insufficient. Someone might use AI only to translate an article, while another person might rely on it for most of the text. It is difficult to quantify the contribution of AI because such systems have become pervasive, including in ordinary word-processing software.

The three labels therefore cannot explain precisely which parts of a research process involved AI. We could describe the particular operations for which AI was used, although this might become extremely tedious. When writing an article, we do not normally explain which idea was inspired by every book we consulted or by every conversation we had. We will need appropriate ethical rules, but I am not convinced that the distinction between “with AI,” “made by AI,” and “without AI” can capture the complexity of actual research practices.

Andreas Sudmann: The issue also depends on whether researchers are fully transparent about their use of these systems. Even when they are, can they reliably distinguish between a passage that was merely translated, an argument substantially generated by AI, and an idea developed through a continuous interaction with a model?

Alexandre Gefen: Intellectual influence has always been complicated. When designing an experiment or interpreting its results, we are constantly influenced by the work of colleagues. It is often difficult to explain exactly where influence begins. AI does not make this process simpler. On a single screen, the ideas produced by ChatGPT or Anthropic may appear beside one’s own notes and research. They become mixed at every stage. Researchers must therefore remain conscious of the system’s limits.

In my own research, I often use AI to see what has already been done or to identify the most common ideas about a subject. It helps me position my research in relation to what ChatGPT produces as a synthesis of the field. A large part of my use of Deep Research therefore consists in identifying what does not appear in its results and using that absence to develop something new.

Andreas Sudmann: This already addresses part of my next question. Could large language models function as a kind of co-philosopher, actively shaping philosophical inquiry? Do you see an emerging methodological divide between researchers who use LLM assistance and those who deliberately avoid it?

Alexandre Gefen: I think both positions can become excessive. Embracing AI does not mean giving it the authority to write in one’s place. For me, it may mean using AI to review literature while remaining aware of its limits. These systems do not necessarily reach deeply into the history of an idea or concept. They often skim the surface, and even when they go deeper, they can work only with the published literature to which they have access.

I need to know what kinds of sources the system can access and what it can reasonably do when producing a broad synthesis of an unfamiliar field. I also need to understand its ethical limitations. Above all, I want to remain in control. The system will generally give me common knowledge rather than genuinely new knowledge.

Using AI therefore requires a critical view of its results and a good understanding of the particular system involved. Embracing AI means checking the context and remaining aware of what the system can and cannot know.

Andreas Sudmann: If we consider Thomas Kuhn’s concept of paradigms and Imre Lakatos’s methodology of research programmes, what epistemological role should we assign to novel AI-generated models or visualizations when they conflict with established theory? Could they become anomalies capable of challenging the protective belt, or even the hard core, of a research programme? Which standards of validation would they have to meet before becoming serious theoretical contenders?

Alexandre Gefen: I am not sure that AI systems are serious theoretical contenders at present. They are highly efficient tools for particular processes. AlphaFold’s work on protein folding is one example, and a literature review is another. Yet we remain far from systems that can move inductively from data to a new theory. They are far from providing genuinely new theoretical frameworks.

Whether the AI revolution should lead us to adjust our view of the history of science is a very interesting question. For me, AI is primarily an important accelerator of research and theoretical work. It does not yet amount to a theory or a vision of the world in its own right because it synthesizes existing ideas with increasing speed. Reviewing literature and identifying opposing arguments have always belonged to scientific work. In the way I use AI to survey a field, develop an idea, shorten a long reading process, or obtain statistics more quickly, the underlying operation is familiar. What changes is the speed.

In that sense, AI does not constitute a new paradigm in my own practice. In sciences that depend heavily on data, however, it could contribute to a shift toward a more inductive and empirical mode of inquiry.

Andreas Sudmann: Beyond these gains in efficiency, what philosophical risks might arise if scientific discovery becomes increasingly driven by AI and organized around data? Could this development marginalize theoretical work or serendipitous findings? Might it also change how knowledge moves across disciplines and thereby alter the character of scientific progress?

Alexandre Gefen: I do not have a comprehensive view of how AI is changing every scientific discipline. I collaborate with researchers from many fields through the CNRS AI centre, but it is still too early to identify a clear paradigm shift. It is equally difficult to know whether AI will produce more interdisciplinary research. We do not yet have enough distance from these technologies to see what they will do to science.

They could encourage intellectual laziness and reduce inventiveness. They could also produce serious mistakes when researchers rely on synthetic material instead of data from the natural world. At the same time, they may lead to major discoveries or technological applications. It is still too early to draw a stable conclusion.

Andreas Sudmann: AI is increasingly performing tasks that were once assumed to require human scientific creativity. Margaret Boden distinguishes between psychological creativity, which is new for the person who produces an idea, and historical creativity, which introduces something new at the level of culture. In light of AI’s growing ability to explore or transform information, do traditional concepts such as intuition, genius, and subjective insight need to be revised? Are they still adequate for describing scientific work augmented by AI?

Alexandre Gefen: From an evolutionary perspective, a system must interact with the world and integrate different forms of sensory information before we can attribute creativity to it in a stronger sense. It must, in some sense, function as an organism. If it remains a reactive loop responding to human input, it cannot possess this form of creativity. Randomness is another necessary component. Without some form of randomness, creativity cannot emerge.

To determine when an artificial agent might create genuinely new knowledge, we will therefore have to wait until it possesses memory rather than merely a context window. It would also need to integrate different sensory modalities. Most importantly, it would need to act in the physical world and understand the consequences of those actions. These capacities would have to develop before we could speak of genuinely creative artificial minds.

Andreas Sudmann: You are pointing toward the possible future role of AI agents. The term itself suggests systems that can interact with an environment and extract information from it. What would these agents need in order to learn from such interactions?

Alexandre Gefen: They must be able to receive feedback from the real world and learn from it. As they are currently designed, foundation models cannot learn continuously. They are trained at a particular moment and subsequently return to the basis of that training rather than evolving through each interaction.

We would need a model that can learn and reorganize itself step by step. A large language model may have a context window that allows it to recover information from an earlier discussion. Yet it cannot integrate what it has learned from its interaction with you or with the real world into its own model of that world. This is a very important limitation.

Andreas Sudmann: We recently conducted an experiment for the HiAICS website in which one LLM interviewed another about how AI is changing science. The human role was largely limited to establishing the initial setting and transferring the questions and answers between the systems. What interests me is the emerging possibility of machine-to-machine communication in which the human is no longer continuously present in the loop. How do you assess this kind of experiment?

Alexandre Gefen: For a debate at the Centre Pompidou, I am preparing a related experiment with my former doctoral student Pierre Depaz. We will present an LLM debating with itself. Such a debate offers another way of using one system to assess the output of another.

Andreas Sudmann: Which questions will the systems debate at the Centre Pompidou?

Alexandre Gefen: The experiment will address the issues around which we are organizing the public debates. These include AI and work, as well as AI and creativity. We will also consider whether AI can be understood as a being and examine its role in education and scientific research. A further discussion concerns its political implications for democracy.

Large language models are very good at presenting a debate between different ideas because their latent or embedding spaces allow them to organize one position in relation to another. They are extremely good scholastic pupils in this respect.

Andreas Sudmann: It might be productive to extend such an experiment beyond a dialogue between two systems and organize a group discussion among several LLMs. We have already explored a related arrangement with our AI co-ethnographer, asking one model to assess outputs produced by humans and other models.

Alexandre Gefen: You should do it. I would be very interested in the experiment. We could organize a group conversation among LLMs and compare how they formulate the different problems. I would be happy to work on such an experiment with you.

Andreas Sudmann: If an AI system demonstrates high predictive accuracy or generates successful novel hypotheses within a scientific field, does this amount to genuine scientific understanding in a philosophical sense? Or does the system remain a sophisticated form of simulation? How important is this distinction for the aims and self-conception of science?

Alexandre Gefen: It is very difficult to move from a plausible and efficient analysis of data, which may help produce an important medical result, to a full understanding of the process that explains why the data behave as they do. There will always be a major additional step involved in producing such an explanation.

We can return to the earlier analogy with quantum physics. Even if quantum physics provides statistical predictions rather than certainty about a particular outcome, it still gives us an understanding of the phenomena involved in an experiment. Obtaining a prediction differs from understanding what is happening at the physical or biological level.

I think the human mind seeks more than efficient prediction when it engages in science. Science may continue to be defined through the attempt to develop a model of reality that explains what is happening behind the observable result. The distinction between prediction and understanding will therefore remain extremely important.

Andreas Sudmann: Popperian falsifiability remains a central criterion for scientific theories. How does the growing reliance on complex and potentially opaque AI systems affect the practical application of this principle? Does their opacity hinder our ability to test and refute hypotheses generated or supported by AI?

Alexandre Gefen: I addressed part of this question at the beginning. We can still try to preserve falsifiability by describing the model we use and the data on which it operates. We should also describe how the model was trained, reinforced, and aligned. In this way, we can document the experiment as precisely as possible.

This resembles the description of the chemical products used in a laboratory experiment. At best, we can provide a concrete account of our tools and data, together with the conditions under which the result was produced. Researchers already do this in many scientific fields. A publication often includes a precise description of the experiment so that others can identify where a problem may have occurred.

There are examples in biology and medical research of authors retracting an article after discovering that a chemical reagent was contaminated or unreliable. Comparable problems could be multiplied many times when the research process depends on a highly complex algorithm or foundation model. The need to document these systems and the conditions of their use will therefore become even more important.

Citation

MLA style

Sudmann, Andreas. „AI, Explanation, and the Changing Conditions of Scientific
Knowledge: An Interview with Alexandre Gefen, 19.05.2025.“ HiAICS, 15 September 2026, https://howisaichangingscience.eu/interview-alexandre-gefen/.

APA style

Sudmann, A. (2026, September 15). AI, Explanation, and the Changing Conditions of Scientific
Knowledge: An Interview with Alexandre Gefen, 19.05.2025
. HiAICS. https://howisaichangingscience.eu/interview-alexandre-gefen/.

Chicago style

Sudmann, Andreas. 2026. „AI, Explanation, and the Changing Conditions of Scientific
Knowledge: An Interview with Alexandre Gefen, 19.05.2025.“ HiAICS, September 15. https://howisaichangingscience.eu/interview-alexandre-gefen/.