Automated Video Scene Classification via Audio Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Viewers face frustration in finding specific scenes in video content, such as movies or TV programs, due to the lack of efficient methods for segmenting and classifying scenes, especially when they cannot recall the timing or visual appearance of the scene, leading to labor-intensive and error-prone manual classification processes.

Innovation Solution

The implementation of an automated system that segments video content using audio analysis, combined with visual understanding and machine learning models, to classify scenes into categories like comedic, inspirational, or musical, allowing for the generation of playlists of specific segments, enabling faster-than-real-time processing and accurate classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification of video scenes is performed, then classification accuracy can be maintained, but processing time and labor intensity increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments video content into atomic audio segments and processes them in parallel using multiple machine learning models. This segmentation allows the system to handle large video files efficiently while maintaining classification accuracy through specialized models for different scene types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces manual mechanical classification with automated machine learning models that analyze audio and visual features. This substitution eliminates human labor while maintaining or improving classification accuracy through consistent, algorithmic processing of video content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated processing is implemented, then processing speed increases, but system complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments video processing into distinct analytical components handled by specialized machine learning models. This modular segmentation enables parallel processing that increases speed while keeping each individual model relatively simple and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal processing framework that handles multiple scene types and classifications through a single integrated system. This multi-functional approach increases processing speed by avoiding multiple separate systems while the modularity keeps complexity manageable.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If comprehensive scene analysis is performed, then classification accuracy improves, but computational resources required increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments video content into atomic audio segments and processes them in parallel using multiple specialized machine learning models. This segmentation enables comprehensive analysis to be distributed across multiple smaller, more efficient processing units, reducing overall computational resource requirements while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different levels of analysis to different segments based on their characteristics. By selectively applying comprehensive analysis only where needed and using simpler classification for other segments, the system maintains high overall accuracy while reducing total computational resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11120839B1Segmenting and classifying video content using conversation
Publication Date: 2021.09.14 AMAZON TECH INC
  • US11120839B1 patent drawing
  • US11120839B1 patent drawing
  • US11120839B1 patent drawing

AI summary

Disclosed are various embodiments for segmenting and classifying video content using conversation. In one embodiment, a plurality of segments of a video content item are generated by analyzing audio accompanying the video content item. A subset of the plurality of segments that correspond to conversation segments are selected. Individual segments of the subset of the plurality of segments are processed to determine whether a classification applies to the individual segments. A list of segments of the video content item to which the classification applies is generated.