An independent Agentic.ai decision report. Every tool is ranked by our 9-dimension Agenticness rubric — not by who pays us (nobody does). Shared with you because someone thought it would help.
Agentic.ai · Decision Report · July 2026
The best AI tools for research & deep analysis
Built for a specific situation: research & deep analysis. If yours differs, the honest answer may differ too — the ranking below is scored for this job, not in the abstract.
Our pick
ChatGPT Deep Research — 15/36 · Level 2
Tops our independent /36 for research & deep analysis — strongest on autonomy and planning, ahead of Gemini Deep Research (12/36). Pricing: Subscription (quotas undisclosed).
Ranked by our 9-dimension rubric, not by who pays us (nobody does). If the best option for your situation were something we couldn't profit from, this report would still say so — the independence is the point. The full scorecard, the evidence behind each score, and how we score (and why nobody pays for placement) are below.
An independent, data-grounded read. Every tool is scored on our 9-dimension Agenticness rubric (out of 36) and ranked by that score — not by who pays us. Capability, pricing, and evidence are pulled from structured data; reliability is graded against cited, primary-sourced evidence (see below); adoption is real outbound clicks from the directory — shown for transparency, never a ranking input, so a lower-traffic tool can (and sometimes does) outrank a more popular one on evidence.
The shortlist
1. ChatGPT Deep Research — 15/36 · L2 · Subscription (quotas undisclosed) · 0 clicks/30d (→ flat)
An agent inside ChatGPT that independently searches, analyzes, and synthesizes large amounts of online information. It’s built for research-heavy tasks where you need a documented answer, not just a quick summary.
Strongest: autonomy + planning. Weakest: sovereignty (not yet evidenced).
2. Gemini Deep Research — 12/36 · L1 · Subscription tiers (limits in-product only) · 0 clicks/30d (→ flat)
Gemini Deep Research helps you turn complex questions into structured research reports by browsing the web and, if you choose, your Gmail, Drive, and Chat. It is useful for competitive analysis, due diligence, and topic exploration, but it is still a research feature inside Gemini rather than a full standalone agent platform.
Strongest: autonomy + planning. Weakest: sovereignty (not yet evidenced).
3. Undermind — 11/36 · L1 · Subscription · 42 clicks/30d (→ steady)
Undermind is an AI research assistant that searches scientific papers on your behalf and returns cited answers. It is designed for researchers who need fast literature review, novelty checks, and deeper topic exploration.
Strongest: autonomy + planning. Weakest: sovereignty (not yet evidenced).
Who should skip these
The honest disqualifiers. We would rather you buy nothing than buy the wrong thing — these are the situations where a top-ranked tool is still the wrong call for you.
Skip ChatGPT Deep Research if: you need a PRISMA-style systematic-review audit trail, explicit inclusion/exclusion logic, or structured extraction across large paper sets — that's a specialist literature tool's job — or you expect a zero-review final answer for a court filing, regulated memo, or patient-care decision (skip ALL current tools for that; OpenAI's own docs say this one can hallucinate and misstate uncertainty).
Skip Gemini Deep Research if: you need citation-by-citation reliability on news or article attribution (its worst public measurement), public quota disclosure before paying, or custom function tools in the API (explicitly unsupported today) — or your employer is uncomfortable with the retention model: Keep Activity on by default for adults, up to 72h retention even when off, human-reviewed chats kept up to three years, and turning activity off disables many connected apps.
Skip Elicit if: your work is mostly market intelligence, consumer comparisons, current news, books, dissertations, or broad non-academic sources — Elicit's corpus is peer-reviewed articles, proceedings, preprints, and working papers, and it does NOT include those — or you expect full-text extraction without full-text access: no PDF means title-and-abstract-only fallback.
Skip Claude Research if: you want the strongest independently-MEASURED citation reliability (no public apples-to-apples measurement of Claude Research exists — that absence is the finding), you need predictable all-in pricing under heavy use, or you need explicit evidence operations rather than a narrative assistant — Research burns shared limits faster than chat and overages meter at API rates.
Skip Perplexity Research if: you need strict source whitelisting, authenticated industry/institutional data sources as a first-class workflow, or long auditable evidence work over internal corpora — Perplexity is built web-first (Search, Research, uploads), and it's the wrong shape when the job is 'work only from a governed source boundary' rather than 'search the web well and fast.'
9-dimension scorecard
| Tool | Action | Autonomy | Planning | Reliability | Safety | Continuity | Adaptation | Interop | Sovereignty | /36 |
|---|---|---|---|---|---|---|---|---|---|---|
| ChatGPT Deep Research | 2 | 3 | 3 | 2 | 1 | 1 | 1 | 2 | 0 | 15 |
| Gemini Deep Research | 1 | 3 | 3 | 2 | 1 | 0 | 1 | 1 | 0 | 12 |
| Undermind | 1 | 3 | 3 | 0 | 1 | 2 | 1 | 0 | 0 | 11 |
| Elicit | 1 | 2 | 2 | 1 | 1 | 2 | 0 | 1 | 0 | 10 |
| Claude Research | 1 | 2 | 3 | 0 | 1 | 1 | 1 | 1 | 0 | 10 |
| SciSpace | 1 | 2 | 2 | 0 | 1 | 1 | 2 | 0 | 1 | 10 |
| Perplexity Research | 1 | 2 | 2 | 2 | 0 | 1 | 0 | 0 | 0 | 8 |
Each dimension 0–4. Green = 3+. A 0 means "not yet evidenced" (a sourcing stance), not "broken" — see the reliability section below.
Reliability — the evidence
Reliability splits into two sub-signals: (a) harness-linked benchmark — does the tool's own agent harness post a verifiable score? — and (b) real-world incident / workflow history — how does it behave in production? A "0 — not yet evidenced" is a sourcing-integrity stance, not "broken": we won't award points to a number we can't trace to a primary source, and we never backfill with model marketing. We anchor on contamination-resistant benchmarks (Terminal-Bench 2.1, SWE-bench Pro, the Artificial Analysis Coding Agent Index) and documented incidents, not leaderboard-top numbers.
ChatGPT Deep Research — 2/4. Benchmark: Public benchmarks place it in the top tier for difficult deep-research tasks, strongest on instruction-following (DeepResearch Bench 2025 — possibly stale but still the clearest public head-to-head). The differentiated capability is SOURCE CONTROL: restrict/prioritize specific sites, use uploaded files, and connect authenticated sources (Google Drive, SharePoint, FactSet, PitchBook, Scholar Gateway) — the most useful source-access story in the category for defensibility. Real-world: OpenAI's own docs say deep research can still hallucinate facts, make incorrect inferences, struggle to separate authoritative information from rumor, and miscalibrate confidence; the Tow Center's 2025 testing found ChatGPT-class search tools often answered when they should have declined, with incorrect/speculative citations common. Good enough to compress hours of desk research into a first draft — not good enough to skip line-by-line evidence checking.
Gemini Deep Research — 2/4. Benchmark: Ranked FIRST overall on report quality and first in effective citation count in DeepResearch Bench (2025, possibly stale) — the breadth-first profile: browse up to hundreds of websites, mix Google Search with Gmail, Drive, Chat, uploads, and NotebookLM notebooks. Caveat the same benchmark makes plain: more sources ≠ more trustworthy sources — Perplexity beat it on citation-accuracy precision. API Deep Research agent adds URL-context, code execution, remote MCP — but currently NO custom function tools or structured outputs. Real-world: The documented failure mode is wide retrieval with weaker citation precision and early termination (Microsoft's LiveDRBench analysis); Tow's Mar 2025 news-attribution study was its worst showing — one fully correct response in the tested setup, over half pointing at fabricated or broken URLs (older than six months, possibly stale, but the clearest public measurement of this failure mode). Mar–Jun 2026 community reports: Deep Research file hangs, Gems losing instructions, long-project context drops even on paid plans.
Elicit — 1/4. Benchmark: The only tool in the set with independent ACADEMIC evaluations of its core function — mixed results: searches averaged 37.9% sensitivity / 41.8% precision across four systematic-review case studies (2025), 81.4% overall accuracy as a semi-automated second reviewer vs 86.7% human (2025), and a 2026 feasibility study concluded it can support data extraction 'cautiously, with human oversight.' Corpus: ~138M papers + 545k clinical trials, sentence-level citations, PRISMA 2020 support, public API + MCP server. Real-world: Corpus-constrained rigor: when Elicit can see the right full texts it is unusually useful; without PDF/institutional access it falls back to title + abstract and recall drops. Docs are explicit that books, dissertations, and non-academic publications are NOT included. Release notes are genuinely evidence-ops oriented (200-paper reports, PRISMA support, tables/figures in the agent) — a real iteration cadence.
Claude Research — 0/4. Benchmark: No apples-to-apples public measurement of Claude Research itself was found — it does not appear in the DeepResearch Bench head-to-head, and any claim that it currently leads on citation integrity is UNVERIFIED. The capability surface is real: long-context qualitative synthesis over Projects + RAG, cited web search, and pre-built/custom connectors via remote MCP across Claude, Cowork, Desktop, and Mobile. Real-world: Anthropic's own help center says Claude can hallucinate or produce authoritative-sounding quotes not grounded in fact, and Research mode consumes shared session/weekly limits FASTER than normal chat — five-hour session limits, weekly limits, and usage-credit overages at API rates can quietly turn a subscription into subscription-plus-metered under heavy research use.
Perplexity Research — 2/4. Benchmark: The HIGHEST citation accuracy of the dedicated deep-research agents tested (DeepResearch Bench 2025) — while scoring well below Gemini and OpenAI on overall report quality in the same benchmark. That's the honest profile: fast, web-first synthesis with lots of inline citations (Pro: 10× citations per answer, extended Research access); the Jul 16 2026 'Advanced Deep Research' update added clarifying questions, mid-run follow-ups, a better code sandbox, and deeper document handling. Real-world: Attribution can look better than it is: in Tow's article-retrieval tests it still answered 37% of prompts incorrectly and sometimes cited syndicated copies over the original publisher; a Jun 21 2026 community bug report logged sonar-deep-research returning EMPTY citations/search results in async mode. Privacy wrinkle: consumer AI Data Retention is ON by default and opt-outs are not retroactive.
Capability & fit
| Tool | License | MCP | Self-host | Own model | Autonomy |
|---|---|---|---|---|---|
| ChatGPT Deep Research | ❌ Proprietary | ✅ | ❌ | ❌ | Semi-autonomous |
| Gemini Deep Research | ❌ Proprietary | ❌ | ❌ | ❌ | Semi-autonomous |
| Undermind | ❌ Proprietary | ❌ | ❌ | ❌ | Semi-autonomous |
| Elicit | ❌ Proprietary | ❌ | ❌ | ❌ | Semi-autonomous |
| Claude Research | ❌ Proprietary | ❌ | ❌ | ❌ | Semi-autonomous |
| SciSpace | ❌ Proprietary | ❌ | ❌ | ❌ | Copilot |
| Perplexity Research | ❌ Proprietary | ❌ | ❌ | ✅ | Semi-autonomous |
Pricing & true cost
| Tool | Model | Tiers |
|---|---|---|
| ChatGPT Deep Research | Subscription (quotas undisclosed) | Plus $20/mo gets 'expanded' deep research; Pro from $100/mo gets 'maximum' — but the actual quota numbers are NOT published on the pricing page; usage varies by plan and the in-product counter is the source of truth. You're buying access plus quota headroom, not a disclosed number of runs. |
| Gemini Deep Research | Subscription tiers (limits in-product only) | Google AI Pro $19.99/mo · Ultra from $99.99/mo (May 2026: standard Ultra cut $250 → $200). The flashiest agentic layer (Gemini Spark) is Ultra-tied and rolled out to trusted testers first. Real request ceilings are daily/concurrent limits surfaced only in-product. API transparency is better: a typical Deep Research task ≈ $1–3. |
| Undermind | Subscription | Free: Not publicly documented in the provided content. · Pro: Available; the site references Pro subscriptions, but pricing details were not fully shown. · Enterprise: Available via the enterprise offering; contact sales… |
| Elicit | Per-seat + shared usage pool | Basic free (limited Research Agent + Reports) · Pro $49/user/mo billed annually. Jul 2026 billing change: workflow-based limits moved to a SHARED monthly usage pool across Research Agent, Reports, and Systematic Literature Reviews — old 'one workflow = one stable quota' mental models are stale. |
| Claude Research | Subscription + API-rate overage credits | Pro $20/mo ($200/yr) · Max 5x $100/mo · Max 20x $200/mo. Heavy research sessions hit the 5-hour/weekly limits sooner than chat; past them, optional usage credits bill at standard API pricing, separately from the subscription. |
| SciSpace | Freemium | Free: Free sign-up is available; the provided page does not detail the full limits. · Pro: Pricing not publicly available on the crawled page. · Enterprise: Enterprise offering is mentioned; contact sales for details. |
| Perplexity Research | Subscription + multi-meter API | Pro $20/mo or $200/yr · Max $200/mo. API (Sonar Deep Research) bills input tokens + output tokens + CITATION tokens + search-query charges + reasoning tokens — cost depends materially on how much browsing the agent decides to do. Help docs speak in 'extended access' language, not hard caps. |
Independent evidence & momentum
| Tool | Harness benchmark | GitHub | Last commit | Clicks/30d | Trend |
|---|---|---|---|---|---|
| ChatGPT Deep Research | DeepResearch Bench: top tier (2025) | — | — | 0 | → flat |
| Gemini Deep Research | DeepResearch Bench: #1 overall (2025) | — | — | 0 | → flat |
| Undermind | — | 4 ★ | 63d ago | 42 | → steady |
| Elicit | 81.4% second-reviewer accuracy (indep. 2025) | — | — | 32 | ↓ cooling |
| Claude Research | not yet evidenced | — | — | 0 | → flat |
| SciSpace | — | 5 ★ | 1600d ago | 30 | ↓ cooling |
| Perplexity Research | DeepResearch Bench: #1 citation accuracy (2025) | — | — | 0 | → flat |
"Harness benchmark" = a score tied to the tool's own agent harness, with the variant named. Variants are not directly comparable; "not yet evidenced" = no verifiable harness score (the underlying model may still score well). GitHub stars/recency are live from the GitHub API.
Sources & methodology: Reliability is graded on two sub-signals — (a) harness-linked benchmark, (b) real-world incident/workflow history — anchored on contamination-resistant benchmarks (Terminal-Bench 2.1, SWE-bench Pro, the Artificial Analysis Coding Agent Index) and documented incidents, never model marketing. Benchmark variants (SWE-bench Verified / Multilingual / Pro) are not directly comparable, and several scores are model-in-vendor-harness rather than tool-isolated. Pricing is volatile — re-verify each vendor's live pricing page on the day of publication. Primary sources by tool: ChatGPT Deep Research — OpenAI Help 'Deep research in ChatGPT' (updated Jul 14 2026); openai.com pricing (Jul 24 2026); DeepResearch Bench (2025, possibly stale); CJR/Tow Center 'AI Search Has a Citation Problem' (Mar 6 2025, possibly stale); the 2026-07-24 research-deep-analysis deep-research ingest. Gemini Deep Research — support.google.com/gemini Deep Research docs + gemini.google/overview/deep-research (Jul 24 2026); ai.google.dev Deep Research agent docs (Jul 14 2026); DeepResearch Bench (2025); LiveDRBench (Aug 2025); CJR/Tow Center (Mar 2025, possibly stale); Google AI plans + I/O subscription update (May 19 2026); the 2026-07-24 ingest. Elicit — elicit.com/pricing + support.elicit.com corpus/changelog docs (Jul 2026); usage-pool announcement (Jul 2026); Lau et al. 2025 (sensitivity/precision); SAGE second-reviewer study (Dec 2025); Lagisz et al. 2026 feasibility study; the 2026-07-24 research-deep-analysis ingest. Claude Research — support.anthropic.com 'Use research on Claude' (Jun 2 2026), incorrect-responses article (Mar 16 2026), plan + usage-credit docs (May 2026), connectors/remote-MCP docs (Jul 7 2026); the 2026-07-24 research-deep-analysis ingest (which found no public head-to-head placement). Perplexity Research — perplexity.ai Help Center 'What is Perplexity Pro' (Jul 21 2026) + 'What's New in Advanced Deep Research' (Jul 16 2026); Sonar Deep Research + API pricing docs (Jul 24 2026); DeepResearch Bench (2025); CJR/Tow Center (Mar 2025, possibly stale); community bug report (Jun 21 2026); the 2026-07-24 ingest.
Methodology & independence: scored on the 9-dimension Agenticness rubric (v3.1, /36). Reliability is graded conservatively on two sub-signals — a harness-linked benchmark and real-world incident/workflow history — and only against evidence we can trace to a primary source; "not yet evidenced" is a sourcing stance, never a verdict that a tool is unreliable. A vendor's payment never touches the score or the ranking — it's enforced in code, not just promised. This report was reviewed against our public rubric before publication. See exactly how we score and why nobody pays for placement → — Agentic.ai
Facing the same decision?
This report was written for one situation. Describe yours and get an independent, rubric-scored read for your job — free, and we still don't take vendor money to change the answer.
Get your own honest verdict →Know someone weighing the same decision?If this saved you a week of comparing, send it their way — it's free to read, and we're not selling either of you anything.
Weighing a different decision?Agentic.ai independently scores agentic-AI tools on a 9-dimension rubric — so you can tell what to actually use, and trust we're not on the take.
Explore the directory →See the research deep analysis tools →