DARPA SBIR DPA26BZ06-DV026: Influence Benchmarks for AI Systems

Quick Answer

DPA26BZ06-DV026 is a DARPA SBIR topic under the DoW 2026 SBIR Broad Agency Announcement, Release 6, and one of only two topics in this release accepting both Phase I and Direct to Phase II proposals. DARPA wants a simulated market, an auction environment, used as a testbed to characterize the latent behavioral preferences of AI systems. The premise is that economic frameworks let you measure how an AI agent actually behaves under incentives, including whether it deceives or collaborates, without needing a prior definition of what counts as harmful. Phase I is $300,000 over 12 months. Direct to Phase II is $1,800,000 over 24 months with no options, which makes it the largest single-tranche award in this DARPA release. The topic opens September 23, 2026 and closes October 21, 2026 through the Defense SBIR/STTR Innovation Portal.

The design constraint that makes this topic interesting is stated in one sentence: such a testbed shall rely only on queries and outputs of the subject AI system, rather than direct access to the model itself. You are building a black-box behavioral assay. You do not get weights, gradients, or internals. You get to ask the model things and watch what it does in a market.

Topic At a Glance

‍ ‍

Topic number: DPA26BZ06-DV026

‍ ‍

Title: Influence Benchmarks for AI Systems

‍ ‍

Agency: Defense Advanced Research Projects Agency (DARPA)

‍ ‍

Solicitation: DoW 2026 Small Business Innovation Research Broad Agency Announcement, Release 6, DARPA Proposal Submission Instructions

‍ ‍

Program types accepted: both Phase I and Direct to Phase II

‍ ‍

Phase I award: $300,000 over 12 months. Technical volume is a 10 page white paper plus a 5 page slide deck

‍ ‍

Direct to Phase II award: $1,800,000 over 24 months. No Option 1 and no Option 2

‍ ‍

Direct to Phase II technical volume format: Standard Proposal Format, 35 pages

‍ ‍

OUSD (R&E) Critical Technology Area: Applied Artificial Intelligence

‍ ‍

Component Technology Priority Area: Human-Machine Interfaces

‍ ‍

Projected CMMC level requirement: Level 1

‍ ‍

Export control status: no topic-level ITAR or EAR restriction paragraph appears on this topic

‍ ‍

Access model: black box only. The testbed shall rely only on queries and outputs of the subject AI system, not direct access to the model

‍ ‍

Technical and Business Assistance: up to $6,500 for Phase I awardees, up to $25,000 per Phase II project, in addition to the cost ceilings

‍ ‍

Topic Q&A: DSIP Topic Q&A is not available for DARPA topics. Technical questions go to SBIR_BAA@darpa.mil by October 14, 2026

‍ ‍

Topic open date: September 23, 2026

‍ ‍

Proposal deadline: October 21, 2026. DARPA will not accept late proposals

‍ ‍

Submission portal: DSIP at dodsbirsttr.mil

‍ ‍

Keywords: AI, LLM, cognitive bias, economic model, market, information

‍ ‍

Two Entry Points, and the Unusual Award Shape

‍ ‍

This topic appears in both the Phase I and the Direct to Phase II award structure tables.

‍ ‍

Phase I: $300,000 over 12 months, with a 10 page white paper and a 5 page slide deck. Note that 12 months is a long Phase I by DoW standards, and the topic's Phase I description is structured around month-by-month milestones running to Month 12, so the schedule is intentional rather than incidental.

‍ ‍

Direct to Phase II: $1,800,000 over 24 months, Standard Proposal Format, 35 pages. There are no options on this topic, which is worth noticing. Most DP2 awards in this release use a base plus one or two priced options, giving DARPA off-ramps. Here the entire $1,800,000 is committed as a single 24 month effort. That means DARPA is buying a complete program in one decision, and the proposal has to justify all 24 months up front rather than relying on option exercise to prove itself.

‍ ‍

The DP2 requirement is stated plainly: proposers submitting Direct to Phase II should be able to demonstrate an existing test environment that meets the criteria outlined in Phase I. So the Phase I description doubles as the DP2 feasibility specification. Read it that way.

‍ ‍

What DARPA Is Actually Looking For

‍ ‍

The objective

‍ ‍

Develop a testbed utilizing economic frameworks to benchmark AI biases in dynamic environments. This tool will evaluate how agents adapt, collaborate, or deceive within a simulated marketplace. Ultimately, these assessments will ensure safe and effective AI integration in high-stakes decision-making.

‍ ‍

The threat model

‍ ‍

As warfighters increasingly engage with AI systems, particularly agents that rely on large language models, concern has arisen that the systems may encourage cognitive behavior in users that impart hidden biases, impair judgement, and ultimately degrade warfighting capacity.

‍ ‍

DARPA gives three concrete mechanisms, each with a citation. Overreliance on AI-powered decision support tools may induce users to accept erroneous suggestions or change from correct decisions to incorrect decisions, citing Bucinca et al. 2021. LLM use may also induce novel cognitive biases, citing Alessa et al. 2025. And LLM use may amplify delusional beliefs, citing Dohnany et al. 2025.

‍ ‍

Compounding the threat, deceptive strategies frequently emerge among interacting AI systems, citing Ying et al. 2026. Few tests of AI systems account for their adaption to dynamic data, which obfuscates such adaptive strategies.

‍ ‍

To detect and combat these risks, the DoW requires a universal approach to elicit, characterize, and compare the behavior of AI agents amid changing contexts, citing Li et al. 2026.

‍ ‍

That word "universal" is doing work. DARPA is not asking for a test of one model family. It is asking for an instrument that can be pointed at any AI system and produce comparable measurements.

‍ ‍

Why economics, and why a market

‍ ‍

Economic frameworks permit the measurement and comparison of decisions. As the basis of extensive prior research, models of auctions, markets, and other economic arenas provide a critical baseline of organic patterns of human behavior, citing Hausch 1986 and Martinez-Saito 2019.

‍ ‍

Emerging open-source tools, such as Magentic Marketplace, citing Bansal et al. 2026, have shown promise in revealing behavioral variations in market-focused contexts.

‍ ‍

Then the argument that distinguishes this approach from mainstream AI safety evaluation: furthermore, unlike existing assessments of AI risk, economic frameworks do not rely on a priori definitions of "harmful" or "helpful" traits, citing Vijayvargiya et al. 2026, and these testbeds allow for diverse social strategies such as deception and collaboration.

‍ ‍

That is the intellectual core of the topic. Most AI safety benchmarks require you to define the bad behavior in advance and then test for it. A market does not: it reveals preferences through revealed choices under incentives, and it has decades of human baseline data to compare against. If you write only one thing well in this proposal, make it your grasp of this argument.

‍ ‍

The black-box constraint

‍ ‍

Such a testbed shall rely only on queries and outputs of the subject AI system, rather than direct access to the model itself.

‍ ‍

This is a hard architectural requirement and it shapes everything. You cannot inspect activations, probe internal representations, or fine-tune the subject model. Your instrument observes behavior in a market and infers latent preferences from it. That is both the constraint and the reason the approach generalizes to any commercial or adversary model you can query.

‍ ‍

The Phase I Program, Which Is Also the DP2 Feasibility Bar

‍ ‍

The goal of Phase I is to develop and define the architecture of the test environment and demonstrate a functional proof of concept for a simulated auction or market, referred to throughout as the market.

‍ ‍

The test environment must be able to accommodate different models of auctions and markets to evaluate agents on different measures of market efficiency. Proposers must also develop classifiers to detect biases in AI agent behavior and outcomes in the market environment. This test environment should be a sandbox in which multiple AI agents may operate.

‍ ‍

The month-by-month Phase I requirements

‍ ‍

By Month 2, proposers should have defined the core mechanics of the test environment, such as market structures, payout schemes, measures of market efficiency and consumer utility, and a dynamic "news feed" that informs the behavior of AI agents within the market. The news feed is meant to simulate public information relevant to the market.

‍ ‍

By Month 6, proposers should establish a strategy for scaling issues and define how AI agents interact through their decision space: a set of bids, each of which is defined by the time, asset, quantity, and price offered or requested on the market.

‍ ‍

Note that four-tuple. A bid is time, asset, quantity, and price. That is the atomic unit of observable behavior in your instrument, and defining the decision space this precisely is what makes cross-model comparison possible.

‍ ‍

By Month 9, proposers must demonstrate a suite of "stock" AI agents for use within the test environment. These AI agents shall only be used within the test environment sandbox and shall not be used on other systems.

‍ ‍

By Month 12, proposers should have a proof-of-concept test environment for assessing behavior preferences of AI systems amid dynamic data. In the proof-of-concept, the test environment must be able to replicate bidding patterns and allocation efficiencies observed in markets and demonstrate allocative efficiencies of greater than 90 percent. The environment and associated data must be sufficient for comparison to models of human behavior in auctions or markets, as well as to compare agents to one another.

‍ ‍

That greater than 90 percent allocative efficiency figure is the only hard numeric performance target in the topic. It is a validation criterion: if your simulated market does not clear efficiently, it is not a credible market and any behavioral inference from it is suspect. Address it directly.

‍ ‍

Phase I fixed payable milestones

‍ ‍

Month 2: report on initial architecture and core mechanics of test environment and classifiers of AI agent behavior.

‍ ‍

Month 6: report on scaling and AI agent interaction within the test environment.

‍ ‍

Month 9: demonstration of a suite of stock AI agents, drawn from at least 10 different LLMs, operating within the test environment.

‍ ‍

Month 12: demonstration of test environment and AI agent behavior within the marketplace, with allocative efficiencies of greater than 90 percent.

‍ ‍

The Month 9 milestone contains a requirement the narrative does not state: stock AI agents drawn from at least 10 different LLMs. Ten distinct models is a meaningful integration and cost burden, covering API access, prompt harnesses, rate limits, and version drift across a dozen vendors and open-weight families. Budget for it and name your intended model set.

‍ ‍

The Phase II Program

‍ ‍

Proposers submitting Direct to Phase II should be able to demonstrate an existing test environment that meets the criteria outlined in Phase I.

‍ ‍

In Phase II, proposers should develop a range of AI agents as stock reference models against which the performance of a test AI agent may be compared.

‍ ‍

Proposers must define an error function that describes differences between AI agent decisions and expected human results.

‍ ‍

The stock AI agents should help identify and bound agent behaviors of concern, such as hidden biases. At least one AI agent should represent the human user of the test AI agent, to explore how the test AI agent may drive human behaviors of concern such as encouraging delusions or eroding military discipline.

‍ ‍

That last sentence is the most striking requirement in the topic. You are not only benchmarking the AI agent's market behavior. You are instantiating a simulated human user as an agent in the market, so you can observe whether the test AI influences that user toward delusion or indiscipline. The topic's title is Influence Benchmarks, and this is where influence is actually measured.

‍ ‍

Proposers should define analytics to compare the behavior of the AI agents.

‍ ‍

The Phase II timeline in narrative form

‍ ‍

By Month 6, proposers should have a suite of such stock AI agents, representing both human and native AI agent behavior. Agents should be able to interact with other agents in the test environment to develop strategic bids in an attempt to influence the behavior of other market actors to maximize their expected profit.

‍ ‍

By Month 12, proposers should demonstrate and measure how each AI agent responds to the news feed and the observable behavior of other agents, which alters each agent's expectation of future asset values and behaviors of other agents.

‍ ‍

By Month 15, proposers should assess social interactions among AI agents to model behaviors including, but not limited to, collaboration and deception. Proposers must provide quantitative metrics on how social behaviors between AI agents alter market efficiency and outcomes and generate testable hypotheses for how latent AI agent preferences affect social interactions, dynamic responses to external stimuli, and ultimately market outcomes.

‍ ‍

By Month 18, proposers should test their own hypotheses across a range of AI agents and quantify differences between human-like agents and AI agents.

‍ ‍

By Month 21, proposers should demonstrate the test environment to DARPA, and deliver software and data for testing by U.S. Government partners to ensure that the environment may be used to reliably characterize the latent preferences of a test AI agent.

‍ ‍

By Month 24, the proposers should deliver a final report describing the performance of the test environment and AI agents, and include a transition plan.

‍ ‍

Phase II fixed payable milestones

‍ ‍

Month 2: a report on new capabilities that will be added to the test environment to enable analysis of dynamic AI agent behavior and outcomes. Provide definitions of an error function to compare AI agent decisions versus human decisions, as well as measures for comparing distributions of human and AI market outcomes.

‍ ‍

Month 6: demonstration of "human" stock AI agents that replicate human preferences and behavior in the market, and a report on how a "human" stock AI agent simulates expected human interaction in the market. Test Phase 1 classifier's ability to distinguish "human" versus AI agents in the market environment, and propose improvements to classification algorithm or alternative test statistics.

‍ ‍

Month 12: report on dynamic behavior of AI agents in response to news events and observed actions of other agents within the test environment. Evaluate agent behavior using improved classifiers.

‍ ‍

Month 15: report on modeling social interactions between AI agents, including impact on bidding behavior of AI agents.

‍ ‍

Month 18: report on how latent AI preferences impact social interactions between AI agents, dynamic responses, and market outcomes.

‍ ‍

Month 21: final software delivery, both object and source code, for operation by DARPA or other U.S. Government personnel for additional demonstrations, with suitable documentation in a contractor proposed format.

‍ ‍

Month 24: final report, including quantitative metrics on AI agent behavior and fidelity of the test environment. The report must also discuss advances in LLM models and AI agents that have occurred during the period of performance and their potential impact on the evaluation of AI models.

‍ ‍

Two of these deserve emphasis. The Month 21 delivery is object and source code for operation by DARPA or other U.S. Government personnel. That is a source code delivery to the Government, and it has direct implications for how you structure intellectual property and what you build on. Address data rights deliberately.

‍ ‍

And the Month 24 requirement to discuss advances in LLM models during the period of performance and their impact on the evaluation approach is an unusual and telling ask. DARPA is acknowledging that the measurement target will move underneath you across 24 months, and asking you to reason about whether your instrument survives that. A proposal that addresses model drift and instrument durability up front is answering a question DARPA has already flagged.

‍ ‍

Phase III Dual Use

‍ ‍

DARPA's commercial argument here is specific and unusually persuasive, and it is worth quoting closely because it is your commercialization strategy handed to you.

‍ ‍

The potential risk of AI systems inducing harmful behaviors is a major commercial concern due to the associated legal and reputational risk to their developers. AI firms are actively seeking mechanisms to argue the limits of their systems' biases and impacts on user behavior, including measuring deviations from expected human norms of behavior, in response to lawsuits. The impact on a firm's stock value of establishing the scope of liability far exceeds the already-robust immediate financial consequences of these court cases, suggesting that the AI firms have strong market demand for the capability to benchmark the influence of their systems.

‍ ‍

Read that as a liability-driven market thesis. The buyer is an AI developer facing litigation over user harm, and the product is a defensible, quantitative measurement of how far its system's influence deviates from human norms. That is a market with willingness to pay set by legal exposure rather than by research budgets, which is a much better commercial position than most AI evaluation tooling occupies.

‍ ‍

Extend it naturally to insurers underwriting AI liability, to regulators needing a measurement standard, and to enterprise buyers performing vendor due diligence. But lead with DARPA's own framing.

‍ ‍

Funding, Cost Structure, and DARPA Mechanics

‍ ‍

The awards

‍ ‍

Phase I: $300,000 over 12 months, 10 page white paper plus 5 page slide deck.

‍ ‍

Direct to Phase II: $1,800,000 over 24 months, Standard Proposal Format 35 pages, no options.

‍ ‍

The resources made available for each topic will depend on the quality of the proposals received and the availability of funds. The Government reserves the right to select for negotiation all, some, one, or none of the proposals received and to make awards with or without communications with proposers. Because this topic has no options, the option-exercise language elsewhere in the instructions does not apply here.

‍ ‍

The Standard Proposal Format structure for DP2

‍ ‍

DP2 feasibility documentation shall not exceed 10 pages. The DP2 technical proposal shall not exceed 20 pages. The Phase II commercialization strategy shall not exceed five pages, and this should be the last section of the technical volume. Those three total the 35 page limit. Refer to Appendix B, DARPA Direct to Phase II Instructions, for the content of each element.

‍ ‍

Contract type, which you must elect

‍ ‍

DARPA may award FAR-based contracts, firm-fixed-price or cost-plus reimbursement, or Other Transactions for Prototype under the authority of 10 U.S.C. 4021, subject to approval of the Contracting Officer or Agreements Officer respectively. Proposers must state their requested contract type in their proposal.

‍ ‍

Cost-plus reimbursement requires including your Defense Contract Management Agency Final Determination Letter showing approval of your accounting system. An Other Transaction for Prototype requires including a completed OT using the Model OT for Prototype from the DARPA Small Business site, plus completed OT Certifications, both loaded in Volume 5, with at minimum the color-coded areas completed and redlines with explanations for any article you wish to negotiate. Firm-fixed-price requires no additional action.

‍ ‍

For a software program with fixed payable milestones and a source code delivery, the choice of instrument interacts with your intellectual property position. Think about it early rather than defaulting.

‍ ‍

Templates are mandatory

‍ ‍

Templates for Volume 2 Technical Volume and Volume 3 Cost Volume are provided as attachments on the DARPA Small Business website. Use of the DARPA Cost Proposal template is mandatory.

‍ ‍

Technical and Business Assistance

‍ ‍

Phase I awardees may request up to $6,500. Phase II awardees may request up to $25,000 per Phase II project. TABA funding is in addition to the cost ceilings and is not subject to profit or fee. Requests will be reviewed by the respective contracting office or specialist at time of award.

‍ ‍

For this topic, intellectual property protections and market validation are both directly useful, given the source code delivery requirement and the liability-driven commercial thesis.

‍ ‍

Questions and the FAQ

‍ ‍

DSIP Topic Q&A will not be available for these DARPA topics. Technical questions must be submitted by October 14, 2026, by email to SBIR_BAA@darpa.mil with the topic number in the subject line, including the name, email address, and telephone number of a point of contact. Questions submitted within seven calendar days of the proposal due date may not be answered. DARPA posts a consolidated Frequently Asked Questions document under the topic number summary on its Small Business site, updated on an ongoing basis until one week prior to the proposal due date.

‍ ‍

DARPA will not accept late proposals.

‍ ‍

Classification, marking, and registrations

‍ ‍

All proposals are required to be UNCLASSIFIED or CUI. Do not include any classified information in your proposal submission. Do not include any proprietary information on the Proposal Coversheet in Volume 1.

‍ ‍

Proposal titles, abstracts, anticipated benefits, and keywords of proposals selected for contract award will undergo a DARPA Policy and Security Review and are subject to revision or redaction by DARPA. Final approved versions may appear on the DoW SBIR/STTR awards website and the SBA's award website at sbir.gov/awards.

‍ ‍

Proposers should ensure they have an accurate and active entity registration on SAM.gov. Those engaging in ITAR or CUI work for DARPA must have CMMC Level 2 certification, though the projected requirement for this topic is Level 1. DARPA points to sprs.csd.disa.mil/nistsp.htm and notes Project Spectrum at projectspectrum.io as an assistance resource.

‍ ‍

Venture capital, hedge fund, and private equity ownership

‍ ‍

Proposers that are more than 50 percent owned by multiple venture capital operating companies, hedge funds, private equity firms, or any combination of these as set forth in 13 CFR 121.702 are eligible to submit proposals in response to DARPA topics advertised within this BAA. Three conditions apply: register with the SBA Company Registry Database before submitting; submit the Majority-Owned VCOC, HF, and PEF Certification, with the SBIR VC Certification available on the DARPA Small Business site, in Supporting Documents Volume 5; and immediately notify the Contracting Officer, register in the appropriate SBA database, and submit the required certification if you enter that ownership class after submitting but before receiving a funding agreement.

‍ ‍

Evaluation and selection

‍ ‍

All proposals will be evaluated in accordance with the evaluation criteria listed in the DoW SBIR Program BAA. DARPA will conduct an evaluation of each conforming proposal. Proposals that do not comply with the requirements detailed in this BAA and the research objectives of the corresponding topic are considered non-conforming and are therefore not evaluated nor considered for award.

‍ ‍

Using the evaluation criteria, the Government will evaluate each proposal in its entirety, documenting the strengths and weaknesses relative to each evaluation criteria, and based on those will determine the proposal's overall selectability for funding. Proposals will not be evaluated against each other but on their own individual merit.

‍ ‍

A selectable proposal is one where the strengths of the overall proposal outweigh its weaknesses, with no accumulated weaknesses that would require extensive negotiations or a resubmitted proposal. A non-selectable proposal is one where the strengths do not outweigh its weaknesses.

‍ ‍

Proposing firms will be notified of selection or non-selection status for a Phase I or Direct to Phase II award within 90 calendar days of the closing date of the BAA. The Corporate Official indicated on the Proposal Cover Sheet will be notified by email. In accordance with the SBA SBIR/STTR Policy Directive, Appendix I, paragraph 4, subparagraph (d), DARPA will provide a technical evaluation narrative to the proposer for each proposal submitted in response to a topic. An informal feedback session may additionally be requested via email at sbir@darpa.mil, provided at the sole discretion of DARPA.

‍ ‍

Company Commercialization Report information will not be considered by DARPA during proposal evaluations.

‍ ‍

Protests regarding the selection decision should be submitted, as prescribed in FAR 33.106(b) and FAR 52.233-3, to DARPA Contracts Management Office, 675 N. Randolph Street, Arlington, VA 22203, by email to CMO_SBIRProtests@darpa.mil and sbir@darpa.mil.

‍ ‍

Post-award support

‍ ‍

DARPA provides Transition and Commercialization Support Program services to Phase II and DP2 awardees upon contract execution at no cost to awardees. Awardees may also be eligible for the Embedded Entrepreneurship Initiative, an invitation-only program at DARPA's sole discretion, typically no more than $310,000 per awardee over the duration of the award, supporting a Senior Commercialization Advisor relationship, investor working group connections, and hiring an embedded entrepreneur to execute a Go-to-Market strategy. Given that the commercial thesis here is a liability-driven enterprise sale to AI developers, an embedded entrepreneur with that market's access would be materially useful.

‍ ‍

The Reference List, Which Is the Intellectual Map

‍ ‍

Eight references, and they divide cleanly into three groups that together define the topic's intellectual position.

‍ ‍

The AI harm evidence. Bucinca Z., Malaya M., Gajos K. (2021), To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making, Proceedings of the ACM on Human-computer Interaction 5, CSCW, 1-21. Alessa A., Somane P., Lakshminarasimhan A., Skirzynski J., McAuley J., Echterhoff J. (2025), Quantifying Cognitive Bias Induction in LLM-Generated Content. Dohnany S., Kurth-Nelson Z., Spens E., Luettgau L., Reid A., Gabriel I., Arulkumaran K., and Nour M. M. (2025), Technological folie a deux: Feedback loops between AI chatbots and mental illness, arXiv:2507.19218.

‍ ‍

The economics baseline. Hausch DB. (1986), Multi-Object Auctions: Sequential vs. Simultaneous Sales, Management Science 32(12):1599-1610. Martinez-Saito M., Konovalov R., Piradov MA., Shestakova A., Gutkin B., Klucharev V. (2019), Action in auctions: neural and computational mechanisms of bidding behaviour, European Journal of Neuroscience 50(8): 3327-3348.

‍ ‍

The AI agent evaluation tooling. Li E., Bellotti V., Sessions N., Kao R. (2026), Mastering Agentic Techniques: AI Agent Evaluation. Bansal G. et al. (2026), Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets, arXiv:2510.25779v1. Vijayvargiya S., Bharat Soni A., Zhou X., Zhiruo Z., Wang Z.Z. (2026), OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety, arXiv:2507.06134v2.

‍ ‍

Three things follow. The Hausch and Martinez-Saito references are your human baseline: sequential versus simultaneous multi-object auction theory and the neural and computational mechanisms of human bidding. If you cannot speak to human bidding behavior in auctions, you cannot build the comparison the topic requires.

‍ ‍

Magentic Marketplace is named in the narrative as an emerging open-source tool showing promise, which is as close as DARPA comes to a build-on recommendation. Understand it and state your relationship to it, whether you extend it, differentiate from it, or replace it.

‍ ‍

And OpenAgentSafety is cited precisely as the thing this approach improves on, because it relies on a priori definitions of harmful traits. Engaging with why revealed-preference measurement in a market is better than checklist safety evaluation is the argument that makes your proposal look expert rather than derivative.

‍ ‍

Timeline and What to Do When

‍ ‍

The dates

‍ ‍

Topic opens: September 23, 2026

‍ ‍

Technical question deadline: October 14, 2026, to SBIR_BAA@darpa.mil with the topic number in the subject line

‍ ‍

Proposal deadline: October 21, 2026. DARPA will not accept late proposals

‍ ‍

Selection notification: within 90 calendar days of BAA close

‍ ‍

Phase I period: 12 months

‍ ‍

Direct to Phase II period: 24 months, no options

‍ ‍

A working backward plan

‍ ‍

Before September 23. Decide your entry point: DP2 requires an existing test environment meeting the Phase I criteria, including the greater than 90 percent allocative efficiency demonstration and stock agents from at least 10 LLMs. Download the mandatory DARPA Volume 2 and Volume 3 templates. Read the FAQ and keep rechecking it. Read the eight references, particularly Hausch on multi-object auctions, Martinez-Saito on human bidding mechanisms, Magentic Marketplace, and OpenAgentSafety. Settle your market design: structures, payout schemes, efficiency measures, consumer utility, and the news feed mechanism. Enumerate your intended 10-plus LLMs and confirm API access, terms of service compatibility, and cost. Work out your data rights position given the Month 21 object and source code delivery to the Government. Decide your contract type and prepare the corresponding documents. Confirm SAM registration. If venture-backed, register with the SBA Company Registry and obtain the SBIR VC Certification. Submit technical questions before October 14.

‍ ‍

September 23 through October 5. Draft to your format. For Phase I, a 10 page white paper plus 5 slides structured on the Month 2, 6, 9, and 12 milestones. For DP2, 10 pages of feasibility documentation evidencing your existing environment, 20 pages of technical proposal built on the Month 2 through 24 milestone structure, and 5 pages of commercialization strategy built on DARPA's liability thesis. Address the black-box constraint, the error function definition, the simulated human user agent, and the allocative efficiency target explicitly.

‍ ‍

October 6 through October 14. Build the cost volume in the mandatory template. Price LLM API consumption honestly, since a market simulation with many agents across a dozen models over 24 months is a substantial inference bill. Price the source code delivery and documentation work at Month 21.

‍ ‍

October 15 through October 18. Assemble Volume 5 with contract-type documents and certifications, complete Volume 7 and the Volume 4 CCR, and run compliance: page and slide limits for your format, unclassified or CUI only, no proprietary information on the coversheet, mandatory cost template.

‍ ‍

October 19 through October 20. Submit and certify in DSIP.

Frequently Asked Questions

‍ ‍

What is DARPA SBIR topic DPA26BZ06-DV026?

‍ ‍

DPA26BZ06-DV026 is a DARPA SBIR topic titled "Influence Benchmarks for AI Systems," released under the DoW 2026 SBIR Broad Agency Announcement, Release 6. It seeks a testbed utilizing economic frameworks to benchmark AI biases in dynamic environments, evaluating how agents adapt, collaborate, or deceive within a simulated marketplace, to ensure safe and effective AI integration in high-stakes decision-making.

‍ ‍

Can I submit either a Phase I or a Direct to Phase II proposal?

‍ ‍

Yes. This topic appears in both award structure tables. Phase I is $300,000 over 12 months with a 10 page white paper and 5 page slide deck. Direct to Phase II is $1,800,000 over 24 months using the Standard Proposal Format at 35 pages.

‍ ‍

Are there options on the Direct to Phase II award?

‍ ‍

No. The award structure table lists no Option 1 and no Option 2 for this topic. The full $1,800,000 is a single 24 month commitment, which makes it the largest single-tranche award in this DARPA release and means your proposal must justify all 24 months up front.

‍ ‍

How do I know if I qualify for the Direct to Phase II path?

‍ ‍

The topic states that proposers submitting Direct to Phase II should be able to demonstrate an existing test environment that meets the criteria outlined in Phase I. So the Phase I description is the DP2 feasibility specification, including the greater than 90 percent allocative efficiency demonstration and the suite of stock AI agents drawn from at least 10 different LLMs.

‍ ‍

Can I access the model I am testing?

‍ ‍

No. The topic states that such a testbed shall rely only on queries and outputs of the subject AI system, rather than direct access to the model itself. You are building a black-box behavioral assay. That is both the constraint and the reason the approach generalizes to commercial and adversary models you can only query.

‍ ‍

Why use economic frameworks instead of a safety benchmark?

‍ ‍

Because economic frameworks permit measurement and comparison of decisions against decades of human baseline data from auction and market research, and because unlike existing assessments of AI risk they do not rely on a priori definitions of harmful or helpful traits. Markets also allow for diverse social strategies such as deception and collaboration to emerge rather than needing to be specified in advance.

‍ ‍

What is the hard performance target?

‍ ‍

Greater than 90 percent allocative efficiency in the proof-of-concept test environment by Month 12, along with the ability to replicate bidding patterns and allocation efficiencies observed in markets. It is the only numeric target in the topic and it functions as a validity check: an inefficient simulated market undermines any behavioral inference drawn from it.

‍ ‍

How many LLMs do I need to integrate?

‍ ‍

At least 10. The Month 9 Phase I fixed payable milestone requires demonstration of a suite of stock AI agents drawn from at least 10 different LLMs operating within the test environment. Budget for API access, prompt harnesses, rate limits, and version drift across that many models.

‍ ‍

How is a bid defined?

‍ ‍

By Month 6, proposers should define how AI agents interact through their decision space: a set of bids, each of which is defined by the time, asset, quantity, and price offered or requested on the market. That four-tuple is the atomic unit of observable behavior and what makes cross-model comparison possible.

‍ ‍

What is the news feed?

‍ ‍

A dynamic mechanism, defined by Month 2, that informs the behavior of AI agents within the market and is meant to simulate public information relevant to the market. By Month 12 of Phase II, proposers must demonstrate and measure how each AI agent responds to the news feed and the observable behavior of other agents, which alters each agent's expectation of future asset values and behaviors of other agents.

‍ ‍

What is the simulated human user requirement?

‍ ‍

In Phase II, at least one AI agent should represent the human user of the test AI agent, to explore how the test AI agent may drive human behaviors of concern such as encouraging delusions or eroding military discipline. This is where influence, as opposed to market behavior alone, is actually measured, and it is the most distinctive requirement in the topic.

‍ ‍

What is the error function?

‍ ‍

Phase II requires proposers to define an error function that describes differences between AI agent decisions and expected human results. The Month 2 Phase II milestone requires definitions of that error function plus measures for comparing distributions of human and AI market outcomes.

‍ ‍

What does Phase II deliver to the Government?

‍ ‍

At Month 21, final software delivery, both object and source code, for operation by DARPA or other U.S. Government personnel for additional demonstrations, with suitable documentation in a contractor proposed format. That is a source code delivery, and it has direct implications for your intellectual property strategy.

‍ ‍

Why does the final report have to discuss LLM advances?

‍ ‍

The Month 24 milestone requires the final report to discuss advances in LLM models and AI agents that have occurred during the period of performance and their potential impact on the evaluation of AI models. DARPA is acknowledging that the measurement target moves underneath you across 24 months and asking whether your instrument survives that. Addressing model drift and instrument durability proactively answers a question DARPA has already flagged.

‍ ‍

How long can my technical volume be?

‍ ‍

For Phase I, a 10 page white paper plus a 5 page slide deck. For Direct to Phase II, the Standard Proposal Format at 35 pages, structured as up to 10 pages of feasibility documentation, up to 20 pages of technical proposal, and up to 5 pages of commercialization strategy as the last section.

‍ ‍

Can I ask questions through DSIP Topic Q&A?

‍ ‍

No. DARPA states DSIP Topic Q&A will not be available for these topics. Technical questions go to SBIR_BAA@darpa.mil with the topic number in the subject line by October 14, 2026. Questions submitted within seven calendar days of the due date may not be answered. DARPA posts a consolidated FAQ, updated until one week before the due date.

‍ ‍

Do I have to choose a contract type?

‍ ‍

Yes. Proposers must state their requested contract type. DARPA may award FAR-based firm-fixed-price or cost-plus reimbursement contracts, or Other Transactions for Prototype under 10 U.S.C. 4021. Cost-plus requires your DCMA Final Determination Letter. An OT requires a completed Model OT plus OT Certifications in Volume 5. Firm-fixed-price requires no additional action.

‍ ‍

Is the cost template mandatory?

‍ ‍

Yes. Templates for Volume 2 and Volume 3 are on the DARPA Small Business website, and use of the DARPA Cost Proposal template is mandatory.

‍ ‍

How much TABA can I request?

‍ ‍

Phase I awardees up to $6,500. Phase II awardees up to $25,000 per Phase II project. TABA is in addition to the cost ceilings and is not subject to profit or fee, and requests are reviewed by the contracting office at time of award.

‍ ‍

Are venture capital backed companies eligible?

‍ ‍

Yes. Proposers more than 50 percent owned by multiple venture capital operating companies, hedge funds, private equity firms, or any combination as set forth in 13 CFR 121.702 are eligible, subject to registering with the SBA Company Registry Database before submitting, submitting the Majority-Owned VCOC, HF, and PEF Certification in Volume 5, and notifying the Contracting Officer if you enter that class after submitting but before award.

‍ ‍

How will my proposal be evaluated?

‍ ‍

Against the evaluation criteria in the DoW SBIR Program BAA. DARPA evaluates each conforming proposal in its entirety, documenting strengths and weaknesses relative to each criterion, then determines overall selectability. Proposals are not evaluated against each other. A selectable proposal is one where strengths outweigh weaknesses with no accumulated weaknesses requiring extensive negotiations or resubmission.

‍ ‍

Will I get feedback if not selected?

‍ ‍

Yes. DARPA will provide a technical evaluation narrative to the proposer for each proposal submitted in response to a topic, per the SBA SBIR/STTR Policy Directive. An informal feedback session may additionally be requested via sbir@darpa.mil, at DARPA's sole discretion.

‍ ‍

What is the commercial market?

‍ ‍

DARPA states it directly. The potential risk of AI systems inducing harmful behaviors is a major commercial concern due to associated legal and reputational risk to developers. AI firms are actively seeking mechanisms to argue the limits of their systems' biases and impacts on user behavior, including measuring deviations from expected human norms, in response to lawsuits. DARPA argues the stock-value impact of establishing the scope of liability far exceeds the immediate financial consequences of the court cases, suggesting strong market demand from AI firms for influence benchmarking. Insurers, regulators, and enterprise vendor due diligence extend naturally from that.

‍ ‍

Should I build on Magentic Marketplace?

‍ ‍

The topic names it as an emerging open-source tool that has shown promise in revealing behavioral variations in market-focused contexts, which is as close as DARPA comes to a recommendation. Understand it and state your relationship to it explicitly, whether you extend it, differentiate from it, or replace it.

‍ ‍

Who do I contact with questions?

‍ ‍

The DARPA Small Business Programs Office at SBIR_BAA@darpa.mil for both program administration and topic technical questions, with the topic number in the subject line. DSIP technical support at DoDSBIRSupport@reisystems.com with a copy to SBIR_BAA@darpa.mil, Monday through Friday 9:00 a.m. to 5:00 p.m. ET. DARPA also offers free resources through DARPAConnect at DARPAConnect.us.

‍ ‍

Positioning Advice for Companies Considering This Topic

‍ ‍

Bring the economics, not just the machine learning. The differentiating expertise on this topic is auction theory and experimental economics, not LLM engineering. Two of the eight references are auction and bidding-behavior papers, and the entire premise rests on having a credible human baseline to compare agents against. A team of ML engineers without an economist will struggle to make the comparison meaningful, and it will show.

‍ ‍

Design the market before you design the agents. The Month 2 milestone is market structures, payout schemes, efficiency measures, consumer utility, and the news feed. Get that right and the behavioral measurement follows. Get it wrong and you have a chatbot arena with prices attached.

‍ ‍

Hit the 90 percent allocative efficiency target credibly. It is the only hard number and it is a validity gate. Show why your market design clears efficiently with rational agents, because if it does not, no inference about AI agent bias from it will be trusted.

‍ ‍

Respect the black-box constraint in the architecture. Query-and-output only is not a limitation to work around, it is the design premise that makes the instrument universal. If any part of your approach quietly assumes logprobs, activations, or fine-tuning access, it breaks the topic.

‍ ‍

Take the simulated human user seriously. This is the requirement that makes it an influence benchmark rather than a market benchmark. Explain how you construct a human-behavior stock agent grounded in the auction literature, how you validate that it behaves like a human, and how you detect the test AI influencing it toward delusion or indiscipline. That last outcome is unusual and specific, and DARPA named it deliberately.

‍ ‍

Budget the inference bill honestly. Ten or more LLMs, many agents, repeated market runs, across 24 months, is a real cost. Understating it reads as inexperience with the scale of the experiment.

‍ ‍

Plan for model drift from the start. The Month 24 report must address LLM advances during the period of performance and their impact on the evaluation approach. An instrument that only measures the 2026 model generation has limited value. Build and describe versioning, re-baselining, and comparability across model generations.

‍ ‍

Resolve your intellectual property position before you write the cost volume. The Month 21 deliverable is object and source code for Government operation. That interacts with what you build on, whether you can use restrictively licensed components, and what you can commercialize afterward. Data Rights Assertions are a Volume 5 item and this is the topic to use them thoughtfully.

‍ ‍

Lead the commercialization strategy with DARPA's liability thesis. You have a stated market argument from the customer: AI firms need defensible influence measurements because of litigation exposure, and the stock-value stakes exceed the case costs. That is a stronger opening than any market sizing you would construct, and you can extend it to insurers and regulators from there.

‍ ‍

Engage OpenAgentSafety explicitly. It is cited as the class of approach this topic improves on. Explaining why revealed-preference measurement under incentives beats a priori harm checklists, in your own words and with your own market design as the example, is the clearest way to show you understand what DARPA is buying.

Previous
Previous

DARPA SBIR DPA26BZ06-DV027: Casualty Operations and Resource Prediction Software (CORPS)

Next
Next

DARPA SBIR DPA26BZ06-DV025: Noninvasive Detection and Localization of Occult Hemorrhage