Back to research

Manual evaluation

These diagnostics do not score factual correctness. Each Run case button starts a full live research request using configured credentials and may incur API charges.

medical/scientific — What does research suggest about sleep and memory consolidation in healthy adults?

Expected topics: study design, population, limitations

Important sources: Systematic reviews and controlled studies

Review evidence quality manually; not medical advice.

technology — How do rooftop solar and small wind turbines compare in urban settings?

Expected topics: capacity factor, location, cost

Important sources: Field performance studies

Check geography and publication dates.

economics — What evidence links minimum wage increases to employment changes?

Expected topics: causality, effect sizes, regional differences

Important sources: Quasi-experimental studies and reviews

Look for conflicting findings and scope limitations.

general factual research — Why do seasons occur on Earth?

Expected topics: axial tilt, day length, hemispheres

Important sources: Authoritative astronomy sources

Academic search relevance may be low despite a straightforward question.

controversial/conflicting evidence — What are the benefits and limitations of intermittent fasting compared with continuous calorie restriction?

Expected topics: adherence, trial duration, confounding

Important sources: Randomized trials and systematic reviews

Check whether disagreements and uncertainty are disclosed.

little academic evidence — Does using a purple notebook improve productivity for freelance illustrators?

Expected topics: lack of direct evidence, assumptions

Important sources:

An explicit lack of evidence is appropriate; do not reward invented support.

Compare research runs

Latest 100 matching runs. Ratings are manual developer observations, never input to synthesis. Costs and latency include all attempts; models lists include historical attempts.

Question / runModels usedLatest judgeCost USDLatency msActive papersLatest manual ratings
#1 hello

COMPLETED

openai/gpt-6-astra, google/gemini-3.1-pro-preview, qwen/qwen3.8-max-0902x-ai/grok-4.200.03269410003539920Not rated

Run diagnostics

· Sept. 22, 2026, 6:21 p.m. · completed hello
total_duration_ms
35399
interrupted_execution_metrics_incomplete
False
successful_providers
4
failed_providers
0
successful_agents
3
consensus_success
True
synthesis_success
True
paper_count
20
total_estimated_openrouter_cost
0.0326941000
known_cost_subtotal
0.0326941000
attempts_missing_cost
0
valid_source_percent
100.0
removed_source_ids
0
agreements_count
2
disagreements_count
1
evidence_strength_distribution
{'limited': 2, 'moderate': 2}
judge_confidence
80.0
answer_length
42