AGIBOT's WITA-Omni Preview Tops Daily-Omni Audio-Visual Reasoning Benchmark

2026-07-28 20:37
AGIBOT's WITA-Omni Preview Tops Daily-Omni Audio-Visual Reasoning Benchmark
AGIBOT's WITA-Omni Preview Tops Daily-Omni Audio-Visual Reasoning Benchmark

SHANGHAI, July 28, 2026 /PRNewswire/ -- AGIBOT today announced that WITA-Omni Preview, its multimodal foundation model for embodied interaction, has ranked first on the Daily-Omni audio-visual reasoning benchmark.

According to the latest published results, WITA-Omni Preview achieved an average accuracy of 85.21%, outperforming models including Qwen3.5-Omni-Plus, Gemini 3.1 Pro Preview and Doubao Seed 2.0 Lite.

The model recorded the highest scores in audio-visual alignment, comparison, event sequencing, and the benchmark's 30-second and 60-second video subsets. It also tied for first place in inference, ranking first or joint first in six of the eight metrics reported on the leaderboard.

WITA-Omni Preview ranks first on Daily-Omni with an average accuracy of 85.21%
WITA-Omni Preview ranks first on Daily-Omni with an average accuracy of 85.21%

Understanding Audio and Visual Information Together

Daily-Omni is a third-party benchmark designed to evaluate audio-visual reasoning and temporal alignment in everyday scenarios. It comprises 684 real-world videos and 1,197 multiple-choice questions across six task categories, including audio-visual alignment, event sequencing, inference and reasoning.

Unlike benchmarks that focus primarily on image-text alignment or short-form visual understanding, Daily-Omni tests whether a model can connect sounds with visible events, follow how situations unfold over time, and reason across audio and visual information.

These capabilities are particularly relevant to embodied AI. In dynamic, multi-person environments, a robot must be able to identify who is speaking, associate sounds with visible events, follow sequences of actions, and determine whether, when and to whom it should respond.

From Thinker–Talker to Thinker–Talker–Actor

Conventional robotic interaction systems often rely on separate modules for perception, speech and motion, making it difficult to coordinate understanding and response in real time.

WITA-Omni takes a different approach. Rather than simply integrating a conversational model into a robot, it extends the Thinker–Talker framework with an Actor component, treating physical movement and facial expression as native output streams alongside speech.

The Thinker–Talker–Actor architecture consists of three core components:

  • Thinker serves as the multimodal reasoning core. It maps text, images, audio and audio-visual inputs into a shared representation space and performs cross-modal reasoning and high-level interaction decision-making.
  • Talker generates natural speech in real time, conditioned on the Thinker's internal states.
  • Actor uses dedicated motion and facial-expression heads to generate physical movements and expressions alongside speech, extending the model from digital multimodal interaction to embodied output.

By bringing multimodal understanding, decision-making, speech, movement and facial expression generation into a unified framework, WITA-Omni is designed to coordinate what a robot perceives with how it responds.

Rather than passing information sequentially through independent perception, dialogue and action modules, the architecture maintains a shared model state and timeline across modalities. This allows a robot to continue observing its surroundings and receiving audio input while preparing its response and generating corresponding speech, movements and expressions.

Training for Embodied Multimodal Interaction

WITA-Omni's performance is driven not simply by model scale, but by a tiered data framework and training methods developed specifically for multimodal interaction in physical environments.

During continued training, WITA-Omni was trained on multimodal data at the scale of tens of millions of hours, extending its text and visual capabilities to support audio perception and joint reasoning across sound and visual information.

Publicly available audio-visual data, however, often lacks important signals found in real interactions, including conversational turn-taking, response timing, body movements and facial expressions.

To address this gap, AGIBOT developed a large-scale, human-centric multimodal interaction dataset spanning thousands of hours. The dataset preserves the temporal relationships among audio, video, language, body movement and facial expression, helping the model learn how real-world interactions unfold over time.

WITA-Omni was subsequently trained using a three-stage process comprising supervised fine-tuning, on-policy distillation and reinforcement learning.

  • Supervised fine-tuning, or SFT, establishes the model's core capabilities in multimodal understanding, interaction decision-making and coordinated multi-stream generation.
  • On-policy distillation, or OPD, uses a teacher version of the model with access to additional contextual information, such as relevant time intervals and target identities, to produce higher-quality outputs. These capabilities are then distilled into a student version that operates without the additional information, improving audio-visual reasoning while preserving the model's response style.
  • Reinforcement learning, or RL, further improves interaction decision-making, including response accuracy, timing, target-person selection, and whether the model should wait or respond. Group Relative Policy Optimization, or GRPO, is used to optimize the model's policy.

Through this process, the model learns not only to generate accurate responses, but also to determine whether a response is appropriate, when it should respond, and to whom the response should be directed.

Advancing Embodied Interaction

WITA-Omni Preview's performance on Daily-Omni demonstrates its capabilities in audio-visual temporal understanding and cross-modal reasoning—two areas that are important to human-robot interaction in dynamic physical environments.

WITA-Omni is designed to move embodied interaction away from a sequence of loosely connected modules and toward a continuous process in which perception, decision-making and expression are coordinated within a shared model state and timeline.

As the foundation of AGIBOT's Interaction Intelligence, WITA-Omni will continue to evolve alongside the company's Manipulation Intelligence and Locomotion Intelligence under its "Three Intelligences in One" architecture.

AGIBOT will continue advancing the WITA family of multimodal models and integrating these capabilities into its robotic platforms, with the goal of enabling more natural, timely and context-aware interaction between people and embodied AI.

To view the full Daily-Omni leaderboard, please visit: https://lliar-liar.github.io/Daily-Omni/#leaderboard

For more information, please visit AGIBOT.com and follow AGIBOT on:

AGIBOT is an Embodied AI foundation model company developing both the intelligence layer and the corresponding robotic embodiments needed to bring general intelligence into the physical world. AGIBOT's "Three Intelligences in One" architecture integrates Locomotion Intelligence, Interaction Intelligence, and Manipulation Intelligence into a unified embodied system. Its portfolio spans humanoid robots, quadrupeds, dexterous systems, and commercial cleaning solutions. In June 2026, AGIBOT announced that its 15,000th robot had rolled off the production line.

Latest
Qisda Group's AEWIN Launches Comprehensive Server Portfolio Featuring AMD EPYC™ 9006 Series Server CPUs
LabAsia Returns to Borneo Amid Rising Demand for Laboratory Technologies
Radisson Hotel Group and Accenture Redefine Travel Discovery on ChatGPT
i2Cool Advances Electricity-Free Cooling Through UNESCO-Endorsed FIRST Programme at Global Conference in Paris
Exclusive | Taiwan Lung Expert: Get Scanned at 40, Smoker or Not
Kumamoto 7.1 Quake Triggers TSMC Evacuation, Tsunami Warnings
Taiwan City Diplomacy | Tainan Courts Poland for Tropical Fruit Push
Taiwan City Diplomacy | Tainan Forges AI Robotics Alliance in Poland
GMMTV Fan Row: Queue Dispute Sparks China-Thailand Backlash
SC Introduces SCOPE Collection: A Curated Portfolio of Bangkok's Finest Luxury Residences
HBM4 Reshapes AI Memory — and Taiwan Is Winning
Silencing Satire Won't Solve Taiwan's Carcinogenic Oil Crisis
Israel's Pacific Comeback Puts Taiwan's Three Allies Under Its Watch
INFiLED Launches GeminiLED Series: a Dual-Sided LED Display for Public Spaces
Taiwan Food Safety Crisis 2026: Tainted Cooking Oil Scandal Explained
Agoda Data Shows Booking Flexibility is a Growing Priority for Asian Travelers in 2026
Wacom Yuify Comes to ibisPaint, Bringing Authorship Protection to Digital Art Creation
Articura Launches a New Chapter for Sovereign Healthcare AI in Indonesia
Ascletis Announces Once-Monthly Subcutaneously Administered GLP-1R/GIPR/GCGR Triple Peptide Agonist, ASC37, Demonstrated Superior Weight Loss in a Diet-Induced Obese Mouse Model
Opinion | Kimi K3 and the AI Cold War's Costlier New Battlefield
Beijing Watch | Kimi K3 Built on Stolen Anthropic Tech, White House Charges
Kimi K3 Release Shifts the US-China AI Contest From Power to Scale
China Has Written Its Own Prescription. It Cannot Fill It.
Exclusive | A 1993 Taiwan Policy Sparked the South China Sea Dispute
Taiwan's Drills Simulate Blackouts; Expert Urges Military Comms Reform
IHH Healthcare Expands 'IHH Catalyst' to North Asia to Accelerate Clinical Adoption of Healthcare Innovation
Opinion | Sovereign AI Is an Economic Battle Taiwan Cannot Sit Out
Opinion | The World Runs on Taiwan's Chips. But Taiwan Has No AI of Its Own