PhilosophyBench at Stanford University

Frequently asked questions

Learn more about PhilosophyBench.

What is PhilosophyBench?

PhilosophyBench is an academic research study at Stanford University that is designed to evaluate the capabilities and limitations of AI systems in producing written responses to philosophical questions. Given widespread claims about AI capabilities, we believe that high-quality evaluations like this are important to understand the true extent of AI capabilities in philosophy.

In particular, we aim to better understand whether AI can generate novel philosophical ideas and develop them clearly with depth and sophistication, not merely summarize existing views or apply existing philosophical theories.

Our goal is to develop a rigorous, independently conducted methodology for assessing claims about AI performance in philosophy. The major questions we are hoping to answer include:

  • Can an AI generate novel philosophical ideas and develop them clearly with depth and sophistication?
  • Can an AI generate essays that could be accepted to a philosophy journal, or earn admission to a leading philosophy PhD program?
  • To what extent does providing human philosophy BAs and PhDs access to AI tools change the quality of the philosophical writing that they produce (this is called an “uplift” study)?
  • Given that many frontier AI models are already pre-trained on an extensive set of philosophical work across human history, what limitations, if any, prevent frontier AI models from producing philosophical work at the quality level of the best philosophical work in human history?
  • When AIs write philosophy, are there any patterns in how they write that reveal anything about AI alignment, particularly when AIs are asked to produce novel arguments?
  • What, if anything, might the philosophical capabilities of current AI systems tell us about how future AI systems might develop?

Rather than relying on evaluations conducted by frontier AI model developers, the project seeks to provide an independent academic benchmark against which claims about capabilities in AI philosophical reasoning and writing can be evaluated. If a new AI model comes out and claims are made about its philosophical capabilities, we hope that this study will provide an independent methodology that can judge those claims.

This research study has been approved by Stanford University’s Institutional Review Board (IRB).

This research study is advised by an international group of philosophers and computer scientists, including Ruth Chang, Kit Fine, Ned Block, Nancy Cartwright, Gideon Rosen, Jonathan Schaffer, Crispin Wright, Linda Zagzebski, Michael Cheng, Isaac Robinson, Neil Band, Zachary Lang, Michael Bernstein, Chris Re, and Sebastian Thrun. It is funded by a pro bono donation from the Open Benchmarks Grant at Snorkel. Snorkel will handle payment and tax reporting obligations as they have professional experience with doing so, but they will not receive study data that could train AI models.

If you are interested in participating in PhilosophyBench, please sign up ASAP by October 15, 2026.

Participation requires a minimum commitment of 10 hours, with the scope of your involvement tailored to your interests and availability. We welcome participants who are professional philosophers, graduate students, or advanced philosophy undergraduates, as demonstrated through formal education or other appropriate evidence.

We will provide reasonable compensation for your time commensurate with an academic research study. We especially encourage participation from those interested in answering research questions concerning the philosophical capabilities of AI.

If you have questions about PhilosophyBench, please email philosophybench@gmail.com.

How does PhilosophyBench work?

Philosophers will provide challenging examination questions in their areas of expertise and then grade and provide feedback on essay responses to those questions. We will ask study participants to write essay responses to those questions, and randomly assign them to different experimental conditions (e.g. writing without AI, or writing with access to AI tools). We will also generate some AI essays. Philosophers will not know how the essays they grade were generated, although a significant proportion of all essays will be written by humans, and they will provide detailed, structured feedback.

We will compare human and AI-generated responses using a statistically validated methodology to answer our study’s core research questions:

  • Can an AI generate novel philosophical ideas and develop them clearly with depth and sophistication?
  • Can an AI generate essays that could be accepted to a philosophy journal, or earn admission to a leading philosophy PhD program?
  • To what extent does providing human philosophy BAs and PhDs access to AI tools change the quality of the philosophical writing that they produce (this is called an “uplift” study)?
  • Given that many frontier AI models are already pre-trained on an extensive set of philosophical work across human history, what limitations, if any, prevent frontier AI models from producing philosophical work at the quality level of the best philosophical work in human history?
  • When AIs write philosophy, are there any patterns in how they write that reveal anything about AI alignment, particularly when AIs are asked to produce novel arguments?
  • What, if anything, might the philosophical capabilities of current AI systems tell us about how future AI systems might develop?
  • What patterns are there in AI-generated philosophy papers compared to human papers? For instance, how do AIs handle conflicts between values?
  • To what extent can we reliably detect that a philosophy paper was written by AI?
  • Are there any computational, theoretical, or other limitations on the potential for AI to write philosophy?

We will develop a methodology (e.g. an experimental autograder) which would allow for this evaluation to be repeated when new AI models are released, but we will not open-source any autograder’s source code, which should prevent AI model developers from using the data generated in this study to improve AI products.

We will never identify participants by name when publishing excerpts or comments, and we will associate data collected with an ID instead of a name. All interactions with AI systems will be done through zero data retention endpoints. Participant data will be anonymized and analyzed for research purposes in accordance with Stanford University’s IRB-approved study protocol. We will publish only a limited, anonymized set of examples and aggregate results in the public research paper.

The goal of this study is to produce an independent academic assessment of the philosophical capabilities and limitations of current and future AI systems, not to collect data for training commercial AI models.

However, it is important to acknowledge that even though the data from this study will never be used to train frontier AI models, the mere creation of a benchmark can indirectly lead to AI improvement on philosophy.

We will follow all privacy and confidentiality requirements of Stanford’s IRB, including the requirement that all researchers who work with data collected complete Research, Ethics, Compliance, and Safety Training through CITI and the requirement that data is only used for the purposes of this study, which explicitly exclude training or providing data to AI model developers (including Anthropic, OpenAI, Meta, Google Deepmind, Deepseek, xAI, Kimi, et al).

Will the data I provide be used to train commercial AI models, or be provided or sold to an AI model developer?

No. PhilosophyBench is an academic research study that has been approved by Stanford University’s IRB (Institutional Review Board).

We will follow all privacy and confidentiality requirements of Stanford’s IRB, including the requirement that all researchers who work with data collected complete Research, Ethics, Compliance, and Safety Training through CITI and the requirement that data is only used for the purposes of this study, which explicitly exclude training or providing data to AI model developers (including Anthropic, OpenAI, Meta, Google Deepmind, Deepseek, xAI, Kimi, et al).

As per the approved study protocol, the data you provide will be anonymized, and analyzed for research purposes only—using or disclosing the data in a manner inconsistent with the IRB’s confidentiality and privacy requirements could violate federal or state law. We take these confidentiality obligations seriously.

To efficiently facilitate payment of study participants and any tax reporting obligations, we are working with Snorkel. Snorkel uses enterprise-grade encryption, is SOC-2 certified, and has experience assisting with large-scale AI evaluation studies like this one. As part of our contractual agreement with Snorkel, your data will not be provided to Snorkel to train or improve AI models, but Snorkel will have access to data necessary to pay study participants and facilitate tax reporting.

The goal of this project is not to formalize philosophy or collect data to train AI models, but rather to independently evaluate the extent that AIs can write philosophy through a statistically validated methodology (including IRB approval that we already obtained, randomized controlled trials, control/intervention groups, etc.).

Won’t a research study evaluating how AI models perform at philosophy indirectly lead to AI model improvement?

Yes, this could happen. Publicly available scientific findings about AI capabilities in philosophy could foreseeably influence subsequent research.

Nevertheless, we are not collecting the data or labor necessary to train or improve commercial and frontier AI systems, and our IRB approval does not authorize us to do so. Frontier AI models are already trained on much of the philosophical writing that humans have ever produced, and improving their performance would likely require more targeted forms of data that we are not collecting in this study.

Instead, this project is designed to evaluate the extent to which existing AI systems can produce philosophical writing through a rigorous empirical methodology. We are required to follow the IRB’s confidentiality and privacy requirements.

Furthermore, frontier AI companies (e.g. Anthropic) are already launching initiatives to improve the capabilities of AI to generate philosophy. For instance, see:

There is an important distinction between evaluating AI capabilities and contributing labor or data that is then used to improve commercial AI systems. We deliberately structured the study so that philosophers’ participation serves to evaluate AI capabilities. Nevertheless, it is already evident that AI model developers are trying to improve their AI models’ capabilities in philosophy.

How will my privacy and confidentiality be protected?

We will follow all privacy and confidentiality requirements of Stanford’s IRB, including the requirement that all researchers who work with data collected complete Research, Ethics, Compliance, and Safety Training through CITI.

Furthermore, we will never identify participants by name when publishing excerpts or comments, and we will associate data collected with an ID instead of a name. All interactions with AI systems will be done through zero data retention endpoints. We plan to publish only a limited, anonymized set of examples and aggregate results in the public research paper.

Will the PhilosophyBench questions and data be made public?

We will keep the data in this study private, as required by the nature of our IRB approval, except for publishing a limited, anonymized set of examples necessary to explain the study’s conclusions and aggregate scientific results in the public research paper.

Moreover, keeping the data private is critical for the scientific value of this paper. If AI developers had unrestricted access to this study’s data, then they could repeatedly optimize AI models against the study’s methodology and effectively cheat on the test.

Who is running this study?

This research study is advised by an international group of philosophers and computer scientists, including Ruth Chang, Kit Fine, Ned Block, Nancy Cartwright, Gideon Rosen, Jonathan Schaffer, Crispin Wright, Linda Zagzebski, Michael Cheng, Isaac Robinson, Neil Band, Zachary Lang, Michael Bernstein, Chris Re, and Sebastian Thrun. It is funded by a pro bono donation from the Open Benchmarks Grant at Snorkel. Snorkel will handle payment and tax reporting obligations as they have professional experience with doing so, but they will not receive study data that could train AI models.

Why are some of the people in this study AI researchers?

A thorough research study of AI systems’ philosophical capabilities requires expertise in both philosophy and AI. Computer science professors and researchers at Stanford University have volunteered to contribute technical expertise that is necessary to produce a research paper that would be accepted by the standards of computer science research.

The purpose of this collaboration is not to improve commercial AI models, but to make sure that this research study is rigorous. This project is intentionally structured as an independent academic study, and is not designed to collect data to train frontier AI models. This study is governed by Stanford University’s IRB-approved research protocol.

Why are you training an experimental autograder?

We are going to attempt to train an experimental autograder, but not release its full source code publicly, so that this evaluation can be repeated when new AI models are released. If a new AI model comes out and claims are made about its philosophical capabilities, we hope that there will exist an independent methodology that can judge those claims. Moreover, it is often expected by the standards of computer science research that these types of benchmark studies will attempt to train an autograder.

We cannot determine in advance whether an autograder will be sufficiently accurate to be reliable (there are significant reasons why it might or might not be possible with philosophy); the autograder in OpenAI’s GDPVal benchmark achieves only 65.7% accuracy. It would be a notable finding if this study finds that it is difficult or impossible to train a philosophy autograder, although it is impossible to predict the results of this study in advance. We will not sell the autograder or provide its data to AI model developers (including Anthropic, OpenAI, Meta, Google Deepmind, Deepseek, xAI, Kimi, et al).

If an AI model developer eventually claims that one of its models can produce philosophical work at the level of professional philosophers, it would be useful to have an independently conducted academic methodology against which such a claim could be tested. Otherwise, suppose that a leading AI model developer releases an AI model and claims that it can generate philosophy at the level of professional philosophers and publish in journals, based on the empirical research studies that they have funded. Journalists will write that AI can produce better philosophy than humans, and they will be able to cite empirical research studies.