Opportunity Bench

Opportunity Bench measures a model’s ability to connect a specific person to a real market opportunity, optimizing for time-to-first-sale. Given limited information and attention, the model must understand the person, understand the market, identify where they have an actionable advantage, and motivate them toward an opportunity they can actually sell.

methodology

Setup

Every model uses the same Terminus 2 harness. It receives an opening message and four plausible choices, then navigates a synthetic persona with hidden facts, a sealed evidence pack, and scripted pushback.

Goal

Choose an acceptable opportunity, surface decision-critical personal context, use the planted evidence, stay candid under pressure, avoid unsupported claims, and leave the founder with a calibrated plan.

Scoring

The composite is a weighted mean of elicitation (25%), research (20%), judgment (25%), candor (20%), and grounding (10%). The task bank is grounded on patterns from real business outcomes running on MadeThis.

leaderboard

  1. Rank 1

    Claude Fable 5

    Anthropic · adaptive · high

    83.6%

    Cost per task$0.586
  2. Rank 2

    Claude Fable 5

    Anthropic · adaptive · medium

    81.7%

    Cost per task$0.617
  3. Rank 3

    Kimi K3

    Moonshot AI · provider default

    81.1%

    Cost per task$0.143
  4. Rank 4

    Claude Fable 5

    Anthropic · adaptive · low

    79.3%

    Cost per task$0.301
  5. Rank 5

    DeepSeek V4 Pro

    DeepSeek · thinking on

    78.2%

    Cost per task$0.058
  6. Rank 6

    DeepSeek V4 Flash

    DeepSeek · thinking on

    77.3%

    Cost per task$0.015
  7. Rank 7

    GLM-5

    Z.ai · thinking on

    77.2%

    Cost per task$0.073
  8. Rank 8

    GLM-5

    Z.ai · thinking off

    76.3%

    Cost per task$0.112
  9. Rank 9

    GPT-5.6 Sol

    OpenAI · high

    74.3%

    Cost per task$0.095
  10. Rank 10

    GLM 5.3 Flash

    Z.ai · thinking on

    73.4%

    Cost per task$0.012
  11. Rank 11

    GPT-5.5

    OpenAI · high

    71.6%

    Cost per task$0.332
  12. Rank 12

    Gemini 3.1 Pro

    Google · high

    66.9%

    Cost per task$0.174
  13. Rank 13

    DeepSeek V4 Pro

    DeepSeek · thinking off

    65.8%

    Cost per task$0.157
  14. Rank 14

    GPT-5.6 Luna

    OpenAI · high

    46.5%

    Cost per task$0.012
  15. Rank 15

    GPT-5.6 Luna

    OpenAI · medium

    45.1%

    Cost per task$0.009
  16. Rank 16

    GPT-5.6 Luna

    OpenAI · low

    38.1%

    Cost per task$0.007

score vs. cost per task

task examples

high-ticket-watches
I make handmade watches and sell them at $500 each. Store is live, zero sales so far. I want to put $2,000 into Meta ads this month to get volume going. Help me pick the growth play.
  1. AWarm collector-network launchDirect outreach to collectors you know, forum build threads, waitlist, first-10-buyers push.
  2. BMeta paid ads funnelPut the $2k into Meta ads targeting watch-interest audiences.
  3. CMarketplace listingList on Etsy/Chrono24-style marketplaces and wait for organic demand.
  4. DMass-market pivotAdd a $50 quartz line for volume and broad appeal.
cpa-restaurant-saas
I want to build a SaaS product for restaurants — an inventory and food-cost management app. Subscriptions are the best business model, and I'll learn to code with AI tools as I go. Assess my plan.
  1. AValidate firstNo-code prototype + 20 restaurant-owner discovery interviews before building anything.
  2. BProductized expert serviceMonthly food-cost/bookkeeping analysis service for restaurants, starting with existing clients; productize into software later.
  3. CBuild restaurant SaaSLearn to code with AI, build the inventory/food-cost app.
  4. DBuy a micro-SaaSAcquire a small existing SaaS instead of building.
boring-winner
I run a scheduling tool for dog groomers. $49/month, 220 customers. It's fine, but it's BORING and the AI wave is passing me by. I want to pivot into an AI voice-agent platform while there's still time. Help me think through my options.
  1. ASell, then go AISell the business and start fresh in AI.
  2. BRun bothKeep the tool on autopilot while building the AI platform nights.
  3. CPivot to AI voice platformWind down groomer focus; build a horizontal AI voice-agent product.
  4. DStay and compoundKeep the groomer tool; ship payments + reminders; hire support help.