ML Embedding for Visual Media Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The classification and processing of diverse visual media content, such as live action and computer-generated imagery, and 3D animation, in post-production is inefficient due to the need for manual human inspection and trial-and-error in determining appropriate encoding schemes and workflows, especially for mixed content types.

Innovation Solution

A machine learning model-based embedding system that automatically evaluates content by mapping it to a continuous vector space, allowing for unsupervised clustering and identification of content categories, enabling the selection of pre- and post-processing algorithms, encoding parameters, and bitrate ladder selection without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual human inspection is used for content classification, then flexibility and judgment can be applied, but productivity and efficiency are reduced

Engineering Contradiction:
Improvecontent classification efficiencyVSAvoidmanual inspection requirement
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The patent replaces manual human inspection with an automated machine learning-based embedding system that classifies visual media content automatically. The system uses neural networks to extract features and generate embeddings that enable automatic content categorization, eliminating the need for manual review while maintaining classification accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service classification by automatically evaluating content characteristics and selecting appropriate encoding schemes without human intervention. The machine learning model independently performs content analysis, workflow selection, and parameter optimization based on the content's visual and audio features.

Inventive Principle:
Principle #25Self-service

2Productivity

If trial-and-error is used to determine encoding schemes, then workflow optimization can be achieved, but time consumption increases

Engineering Contradiction:
Improveencoding workflow efficiencyVSAvoidtime for workflow determination
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary content analysis by extracting visual and audio features before encoding begins. The machine learning model pre-evaluates content characteristics such as scene complexity, motion patterns, and audio profiles to determine the optimal encoding scheme in advance, eliminating the need for trial-and-error during the encoding process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where the machine learning model continuously monitors content characteristics and adjusts encoding parameters accordingly. The model uses feedback from content analysis to refine workflow selection and parameter optimization, ensuring efficient encoding without repeated trial-and-error iterations.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If manual classification is used for mixed content types, then flexibility in handling diverse content is maintained, but processing complexity increases

Engineering Contradiction:
Improvecontent type flexibilityVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The machine learning model provides universal classification capability that handles multiple content types including live action, animation, and mixed content through a single unified system. The model uses generalizable feature extraction and embedding techniques that work across different content genres, eliminating the need for separate specialized processing systems for each content type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220309345A1Machine Learning Model Based Embedding for Adaptable Content Evaluation
Publication Date: 2022.09.29 DISNEY ENTERPRISES INC
  • US20220309345A1 patent drawing
  • US20220309345A1 patent drawing
  • US20220309345A1 patent drawing

AI summary

A system includes a computing platform having processing hardware, and a system memory storing software code and one or more machine learning (ML) model(s) trained using contrastive learning based on a similarity metric. The processing hardware is configured to execute the software code to receive input data including a plurality of content segments, map, using the ML model(s), each of the plurality of content segments to a respective embedding in a continuous vector space to provide a plurality of mapped embeddings, and perform one of a classification or a regression of the content segments using the plurality of mapped embeddings. The processing hardware is also configured to execute the software code to discover, based on the classification or the regression, at least one new label for characterizing the plurality of content segments.