Introducing LongCat-DeepResearch

September 2026·Meituan LongCat Team

LongCat-DeepResearch combines a model built on LongCat-2.0 with improvements in deep-research capabilities and the LongCat-DeepResearch harness. These improvements are intended to be incorporated into the next general-purpose LongCat model release.

The system organizes open-ended research around an executable ResearchSpec. Multiple planners explore sources and refine a shared agenda; independent Researchers develop citation-bearing sections; direct assembly and coordinated editing turn those sections into a complete report without asking one final writer to regenerate everything.

LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. Among the four systems compared in the report, it has the highest overall score on all three public benchmarks.

Public-benchmark scores for LongCat-DeepResearch, Gemini-DeepResearch, ChatGPT-DeepResearch, and Claude-DeepResearch. LongCat-DeepResearch scores 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics.
Figure 1. Public-benchmark scores.

A Shared Agenda, Without a Shared Context Bottleneck

Deep research is not a linear retrieve-then-write task. New evidence reveals new questions, and each section can grow into a substantial investigation. Keeping every branch in one expanding history forces evidence, intermediate reasoning, and draft text to compete for context. Compressing those branches too early can discard details before their importance is clear.

LongCat-DeepResearch separates global coordination from local investigation. A compact ResearchSpec records what the report needs to establish and how the work is divided. The detailed evidence stays with the Researchers that use it to write each section. Completed sections, rather than recurrent summaries, become the artifacts passed into composition.

The LongCat-DeepResearch Harness

The harness moves from an inspectable research agenda to independent section research, then coordinates the resulting text through direct assembly and global-to-local editing.

Three stages: explore and plan, research and write sections independently, then assemble and edit the report.
Harness stage 1: explore and plan Harness stage 2: research and write sections Harness stage 3: assemble and edit
Figure 2. The LongCat-DeepResearch harness.

Explore before committing. Planning Writers briefly search and read before proposing candidate specifications. A Planning Judge integrates their perspectives, a Critic looks for missing questions and cases, and a Reviser updates the plan. The resulting ResearchSpec records scope, research questions, required entities or cases, provisional source leads, and section responsibilities.

Research and write in independent section contexts. Each Researcher receives the original query, the complete ResearchSpec, and one assignment. It continues gathering evidence and writes a complete cited section from its own research history. The ResearchSpec is shared; the detailed tool interactions are not forced into one context.

Assemble first, then coordinate. The harness mechanically assembles the sections without rewriting them. A Global Editor identifies conflicts, repeated material, and ownership; Local Editors revise assigned sections against those directives. Global judgment and local generation are separated, while the original section artifacts remain available for inspection.

Research Tasks and Trajectories

The same stage interfaces used at inference time provide concrete units for building and checking research-oriented training data.

Ground questions and rubrics in independent source material. A target profile and either a licensed review article or a frozen multi-source brief condition the research question and its hidden task-specific rubric. Factual criteria require supporting evidence; analytical criteria specify the comparisons or inferences the report should make.

Validate the task before collecting a trajectory. Bounded searchability checks look for alternative support after excluding the construction article. Additional checks cover duplicate criteria, answer leakage, temporal scope, and question-rubric alignment. KEEP, REVISE, and DROP decisions preserve the distinction between complete, partial, and missing support.

Evidence-grounded research-task construction, validation, trajectory collection, and structural filtering.
Stage 1: grounded task synthesis Stage 2: task validation Stage 3: trajectory collection and filtering
Figure 3. Evidence-grounded construction of research tasks and trajectories.

Collect the real stage dependencies. Accepted queries drive teacher executions of the harness, recording candidate and final ResearchSpecs, tool calls and observations, cited sections, assembled drafts, and editorial decisions. Filtering checks role identity, tool-call closure, and intermediate-artifact validity.

Research-related data are used alongside other data in the mid-training and post-training of LongCat's general-purpose models. A candidate trajectory does not, by itself, establish inclusion in a trained checkpoint.

Results Across Three Public Benchmarks

The comparison treats each Deep Research product as a complete system, including its native tools and budgets.

SystemDeepResearchBenchDeepResearchBench IIResearchRubrics
LongCat-DeepResearch55.2551.3579.83
ChatGPT-DeepResearch54.9547.1674.21
Claude-DeepResearch53.4348.1872.91
Gemini-DeepResearch50.2146.7264.92
+0.30DeepResearchBench
+3.17DeepResearchBench II
+5.62ResearchRubrics

The compared products are accessed through their official clients with Gemini 3.7 Flash, GPT-5.6 Sol (xhigh), and Claude Opus 5 (xhigh) selected for Gemini-DeepResearch, ChatGPT-DeepResearch, and Claude-DeepResearch, respectively. DeepResearchBench and DeepResearchBench II use GPT-5.5 (medium) as evaluator; ResearchRubrics uses Gemini 2.5 Pro.

LongCat-DeepResearch leads on comprehensiveness, insight, and instruction following in DeepResearchBench; on information recall and analysis in DeepResearchBench II; and on explicit requirements, implicit requirements, and synthesis in ResearchRubrics. The breakdown also identifies room to improve readability, presentation, citation quality, and communication.

On the separate in-house benchmark, LongCat-DeepResearch scores 76.04 overall, 0.55 points below ChatGPT-DeepResearch. It has the highest Content and Analysis scores in that comparison, while Presentation remains weaker.

What the Design Analyses Show

The report separates full-benchmark system comparisons from development-subset studies of planning, research context, and editing.

Planning

Integrating perspectives helps

On a fixed ResearchRubrics subset, integrating three Writers through the Judge raises the final total from 81.67 to 85.13 versus a fixed single-Writer plan.

Refinement

More planning is not automatically better

Critic/Reviser cycles raise planned coverage, but category and total report scores move differently. ResearchSpec is inspectable state, not a calibrated selector for the best report.

Research contexts

Section-level investigation matters

The full pipeline has the highest observed mean in the component study. Replacing parallel section Researchers with one whole-report Researcher lowers both displayed benchmark means.

Editing

Coordinated revision improves the aggregate trend

Three enhanced Editor rounds reach a 53.96 average automatic readability preference and the highest displayed average native score, although benchmark-level trends differ.

Model and harness are complementary

ModelHarnessDeepResearchBenchDeepResearchBench IIResearchRubrics
LongCat (previous release)ReAct / direct report47.6533.1360.33
LongCat (previous release)Current harness53.5848.6871.90
LongCat (current)Current harness55.2551.3579.83

All three configurations are evaluated on the full benchmarks. With the previous model fixed, the current harness improves every displayed score. Holding the current harness fixed, the current LongCat model improves them further. The result supports treating model capability and research orchestration as complementary parts of the system.

From a Research Question to a Traceable Edit

A recorded portfolio-management trajectory shows how planning changes propagate into section research and how editing resolves overlap without regenerating the report.

01Plan

The merged ResearchSpec assigns 19 subsections across the user's four themes.

02Refine

The Critic identifies Deep Portfolio Theory as a missing direction; the Reviser adds a concrete question to S2.2.

03Research

Independent Researchers expand assignments into cited prose, including CAPM evidence and the newly added topic.

04Coordinate

The Global Editor assigns detailed Fama-French coverage to one section; the Local Editor retains a concise finding and adds a cross-reference.

The artifacts make the change inspectable: the added research question appears in the revised ResearchSpec, the evidence appears in section drafts, and the ownership decision appears in the final edit. The same case also exposes limits: the final report retains an internal section identifier and still misses some benchmark requirements. Traceability helps diagnose gaps; it does not guarantee complete coverage.

Scope and Limitations

  1. System comparison. Public scores compare deployed Deep Research systems with their native tools and budgets. They do not isolate one model component or training stage.
  2. Automatic evaluation. Native benchmark scores, ResearchSpec coverage, and readability preference measure different properties. The readability study is model-judged, not a human rating.
  3. Development studies. Component and Editor analyses use previously inspected development subsets. Their trends should not be read as independent full-benchmark confirmation.
  4. ResearchSpec boundaries. The plan is fixed after dispatch in the current pipeline. Reopening the global agenda during section research remains future work.
  5. Data provenance. Validating a task or recording a trajectory does not establish that it was accepted into a training release.

Citation

If you find this project useful, please cite:

@techreport{longcatdeepresearch2026,
  title  = {LongCat-DeepResearch Technical Report},
  author = {
    Meituan LongCat Team and He Zhu and Yue Xu and
    Xunliang Cai and Yan Chen and Fan Yang and
    Lingchuan Liu and others
  },
  year   = {2026}
}