The evolution of Artificial Intelligence continues to push boundaries, moving beyond text and static images into the dynamic realm of video. Understanding video content, with its temporal complexity and rich data streams, is a significant challenge for AI models. A recent ZDNet comparison put leading AI models – Gemini, ChatGPT, and Claude – to the test to see how well they could "grok" the contents of videos, both from YouTube and local files.
What Happened
ZDNet's David Gewirtz conducted a series of tests, feeding each AI model a diverse set of three videos to assess their comprehension without additional metadata or hints:
- A YouTube Video: A published video on the scientific process of annealing. The AIs were tested on their ability to understand the content and suggest improved thumbnails.
- A Local MP4 File: A silent motion test for a DJI Neo 2 drone, featuring gesture controls. The goal was to see if the AIs could understand the actions in the absence of audio.
- A Local MOV File: The original, unedited MOV file of a YouTube strategy walk-and-talk, deliberately used to avoid any YouTube-provided metadata or transcripts.
The test utilized the latest available paid versions: ChatGPT Plus ($20/month), Gemini Pro ($20/month), and the $100/month Claude model.
The results were quite telling:
- Gemini Pro: Emerged as the clear leader, successfully processing and understanding content from all three video types. It demonstrated the ability to interpret events in the drone video, grasp the concepts in the annealing video, and analyze the local MOV file without external cues.
- ChatGPT Plus: Showed more limited direct video capabilities, requiring "Codex help for deeper video work." This suggests that while it can engage with video, a more robust or integrated approach, possibly via plugins or specialized tools, is needed for comprehensive analysis.
- Claude: Was unable to process video directly at all, indicating a current gap in its multimodal capabilities concerning video input.
Image 1: I tested whether Gemini, ChatGPT, and Claude can analyze videos - this one wins: image omitted due to site embedding policy; open the original article (ZDNet) (opens in a new tab) to view it. Photo/source: ZDNet (opens in a new tab).
Why It Matters
Gemini's strong performance in direct video analysis represents a significant leap forward for multimodal AI and has concrete implications across various sectors:
- For Developers: The ability to feed raw video files or YouTube URLs directly into an AI model streamlines development workflows enormously. It means less time and effort spent on pre-processing video (e.g., frame extraction, transcription, object detection with separate models) and more on building sophisticated applications. Developers can now consider directly integrating advanced video understanding into content creation tools, surveillance systems, educational platforms, or accessibility solutions.
- For Enterprises: This capability unlocks new levels of automation and insight for businesses dealing with large volumes of video content. Use cases range from automated content moderation, enhanced video search and indexing, and generating summaries for training videos, to analyzing marketing campaign performance and even supporting autonomous systems by interpreting real-world visual data. It can significantly reduce manual review and analysis costs.
- For the AI Landscape: Video understanding is inherently more complex than processing text or static images, requiring the AI to grasp temporal relationships, track objects, and interpret dynamic scenes. Gemini's success in this domain highlights the rapid progress in developing more 'perceptive' AI models that can better understand and interact with our visually rich world. It pushes the frontier for genuinely multimodal intelligence.
What To Watch
While Gemini currently leads, the AI race is fast-paced. Here are key areas to monitor:
- Catch-Up from Competitors: Expect Claude and ChatGPT to rapidly enhance their direct video processing capabilities, possibly through native improvements or more seamless integrations with specialized tools.
- API Accessibility and Features: For developers, the availability and robustness of API access to these advanced video analysis features will be crucial. What kind of outputs can be expected? Will there be support for real-time video streams?
- New Use Cases: We'll likely see an explosion of innovative applications leveraging direct AI video analysis, from sophisticated video editing suggestions to real-time analytics for live events and personalized content recommendations.
- Depth and Nuance of Understanding: As these capabilities mature, the focus will shift from simply 'what is in the video' to understanding subtle emotional cues, complex non-verbal interactions, and predicting future actions.
This ZDNet comparison underscores that direct video understanding is no longer a futuristic concept but a burgeoning reality. For developers and IT professionals, embracing these new multimodal capabilities will be key to building the next generation of intelligent applications.