ML Embedding for Visual Media Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The classification and processing of diverse visual media content, such as live action and computer-generated imagery, and 3D animation, in post-production is inefficient due to the need for manual human inspection and trial-and-error in determining appropriate encoding schemes and workflows, especially for mixed content types.
Innovation Solution
A machine learning model-based embedding system that automatically evaluates content by mapping it to a continuous vector space, allowing for unsupervised clustering and identification of content categories, enabling the selection of pre- and post-processing algorithms, encoding parameters, and bitrate ladder selection without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual human inspection is used for content classification, then flexibility and judgment can be applied, but productivity and efficiency are reduced
Solution Approach 1:
The patent replaces manual human inspection with an automated machine learning-based embedding system that classifies visual media content automatically. The system uses neural networks to extract features and generate embeddings that enable automatic content categorization, eliminating the need for manual review while maintaining classification accuracy.
Solution Approach 2:
The system enables self-service classification by automatically evaluating content characteristics and selecting appropriate encoding schemes without human intervention. The machine learning model independently performs content analysis, workflow selection, and parameter optimization based on the content's visual and audio features.
2Productivity
If trial-and-error is used to determine encoding schemes, then workflow optimization can be achieved, but time consumption increases
Solution Approach 1:
The system performs preliminary content analysis by extracting visual and audio features before encoding begins. The machine learning model pre-evaluates content characteristics such as scene complexity, motion patterns, and audio profiles to determine the optimal encoding scheme in advance, eliminating the need for trial-and-error during the encoding process.
Solution Approach 2:
The system incorporates feedback mechanisms where the machine learning model continuously monitors content characteristics and adjusts encoding parameters accordingly. The model uses feedback from content analysis to refine workflow selection and parameter optimization, ensuring efficient encoding without repeated trial-and-error iterations.
3Adaptability or versatility
If manual classification is used for mixed content types, then flexibility in handling diverse content is maintained, but processing complexity increases
Solution Approach 1:
The machine learning model provides universal classification capability that handles multiple content types including live action, animation, and mixed content through a single unified system. The model uses generalizable feature extraction and embedding techniques that work across different content genres, eliminating the need for separate specialized processing systems for each content type.
Data Source
AI summary
A system includes a computing platform having processing hardware, and a system memory storing software code and one or more machine learning (ML) model(s) trained using contrastive learning based on a similarity metric. The processing hardware is configured to execute the software code to receive input data including a plurality of content segments, map, using the ML model(s), each of the plurality of content segments to a respective embedding in a continuous vector space to provide a plurality of mapped embeddings, and perform one of a classification or a regression of the content segments using the plurality of mapped embeddings. The processing hardware is also configured to execute the software code to discover, based on the classification or the regression, at least one new label for characterizing the plurality of content segments.


