LongShTA Benchmark for Omni-Modal Reasoning in Long Videos
Mohammed Irfan Kurpath*, Jaseel Muhammad Kaithakkodan*, Jinxing Zhou, Sahal Shaji Mullappilly, Mohammad Almansoori, Noor Ahsan, Beknur Kalmakhanbet, Sambal Shikhar, Rishabh Lalla, Jean Lahoud, Mariette Awad2, Fahad Shahbaz Khan3, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal
* Equal Contribution
Abstract
Bridging the gap in long-form video understanding
Q&A Pairs
Open-ended, intent-driven
Models Benchmarked
6 paradigms evaluated
Avg. Duration
Long-form video
Annotation Hours
Human validation effort
Human Verified
Manually reviewed
Long-form omni-modal video understanding requires models to integrate vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temporal scale, modality coverage, open-ended interaction, and interpretable scoring. We introduce LongShOTBench, a long video understanding benchmark designed around three coupled goals: holistic omni-modal integration, intent-driven open-ended interaction, and rubric-level diagnosis. Each item includes a reference answer and a weighted criterion-level rubric, enabling evaluation to identify which perceptual and reasoning steps are satisfied or missed.
We also introduce LongShOTAgent, a training-free omni-modal evidence-seeking agent that couples full-video preprocessing with targeted retrieval, query-adaptive segment refinement, and explicit claim verification. We perform comprehensive evaluation of 105 video-capable models spanning open-source omni-modal models, vision-language systems, audio LLMs, agentic pipelines, and closed-source APIs. The strongest closed-source API, Gemini 3.1 Pro, reaches 55.63%, the best open-source model, Qwen3-Omni 30B-T, reaches 64.05%, while our LongShOTAgent emerges as the strongest system at 66.64%.
Demo
See LongShOT in Action
A walkthrough of LongShOTAgent, from question to grounded answer.
Findings
What the benchmark reveals
Evaluating 105 video-capable models surfaces consistent, sometimes counterintuitive patterns about where long-form omni-modal reasoning breaks down today.
Small omni models beat much larger ones
A 30B omni model tops a 235B video model, and a 27B model beats a 241B one. The bottleneck is cross-modal alignment, not raw parameter count.
Models reason better than they perceive
Across methods, reasoning scores sit well above core perception. Grounding specific facts at specific moments is the harder axis, by up to 23 points.
Non-speech audio is the weakest modality
Within the audio channel, transcribable speech is handled well but ambient non-speech audio lags behind. Grounding door clicks and music shifts to the timeline remains an open problem.
Explicit thinking helps a lot
Within a family, thinking variants outscore their instruct siblings by 5 to 25 points, with the largest gains on the strongest omni models.
The Agent
Why the agent wins
LongShOTAgent is training-free. Its search-refine-verify loop lifts a base model to the top of the benchmark and outperforms prior vision-centric agents that overlook non-speech audio.
The agentic loop lifts the baseline model
Wrapping a base LLM in the search-refine-verify loop adds 12.95 to 38.52 points. The orchestrator, not just tool access, drives the gain.
Vision-centric agents don't transfer
Prior long-video agents all score below 15% on LongShOTBench. Treating non-speech audio as a first-class signal is what closes the gap.
Overall score (%), 3-verifier mean. LongShOTAgent leads the next agent by over 6x.
The advantage is comprehension, not multiple-choice shortcuts
On standard MCQ benchmarks like Video-MME and WorldSense, strong monolithic models sit a few points ahead. But strip the answer options and force open-ended responses, and they collapse while LongShOTAgent barely moves, because it generates answers from retrieved evidence rather than picking from a list.
Accuracy (%): change from standard 4-option MCQ to options-stripped open-ended. The solid bar is the open-ended score over the faint MCQ extent; the right column shows the open-ended score, its MCQ baseline, and the change. On both benchmarks LongShOTAgent has the smallest drop and leads open-ended.
Comparison
A comprehensive video benchmark
Comparing LongShOTBench against 18 existing benchmarks across 7 capability dimensions.
Our Benchmark
LongShOTBench
The only benchmark in our comparison to combine all three modalities with intent-driven Q&A, multi-turn dialogue, and custom rubrics for interpretable evaluation.
* Subtitle aided
Pipeline
How LongShOTBench is built
From raw long-form video to a diagnostic, human-verified benchmark: a five-stage pipeline with multimodal extraction, intent-driven Q&A generation, and graded evaluation rubrics.

Raw Video Content · Cooking Tutorial · 1h 20 min
Multimodal Signal Extraction
Cross-Modal Alignment & Fusion
Intent-Driven Question Design
Rubric & Evaluation Framework
Human Validation & Correction
Figure 1: The pipeline begins with raw video data, extracts multimodal signals, generates intent-driven Q&A, creates verifiable rubrics, and undergoes human validation.
Citation
Cite LongShOT
If you find LongShOT useful in your research, please cite our paper.
@misc{kurpath2026benchmarkomnimodalreasoninglong,
title={A Benchmark for Omni-Modal Reasoning in Long Videos},
author={Mohammed Irfan Kurpath and Jaseel Muhammad Kaithakkodan and Jinxing Zhou and Sahal Shaji Mullappilly and Mohammad Almansoori and Noor Ahsan and Beknur Kalmakhanbet and Sambal Shikhar and Rishabh Lalla and Jean Lahoud and Mariette Awad and Fahad Shahbaz Khan and Salman Khan and Rao Muhammad Anwer and Hisham Cholakkal},
year={2026},
eprint={2512.16978},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.16978}
}