Google has officially announced a new Active Video Understanding feature for its Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models, aimed at improving video analysis accuracy while significantly reducing token usage and costs.
Features and Benefits of Active Video Understanding
This new feature enables the models to dynamically scan video segments, resulting in up to an 88% reduction in token usage and a 66% decrease in costs while improving accuracy. A senior product manager at Google DeepMind stated that this feature will revolutionize video analysis.
“The Active Video Understanding feature we are launching today significantly reduces analysis costs while improving accuracy.”
Google
How to Enable Active Video Understanding
Users can enable this feature by configuring the API settings to "agentic" in Google AI Studio or the Gemini Enterprise Agent Platform. This feature is particularly well-suited for processing long videos, whether they are 10-minute tutorials or 90-minute lectures, significantly improving analysis efficiency.
“Active Video Understanding allows Gemini models to proactively decide what to watch, at what speed, and how, thereby greatly reducing development burdens.”
Google
Diverse Application Scenarios
The Active Video Understanding feature changes the way developers handle long video content, supporting various application scenarios, including rapid instant retrieval, anomaly detection, and object counting. These capabilities make the automation of video editing and complex queries more feasible.
“This technology demonstrates significant efficiency improvements in long video analysis without consuming large amounts of tokens.”
Google
In the future, this feature will be gradually rolled out to billions of users across Google products and is planned to enable higher-quality answers in YouTube's "Ask YouTube" feature.
Source: Google Official Announcement

