Machine Learning Detection of Embedded Video Frames in Composite Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content sharing platforms struggle to detect and identify embedded video frames within composite images, as conventional detection technologies are ineffective in discerning reduced video frames within specific portions of the screen.

Innovation Solution

A machine learning model is trained using composite images with embedded video frames to identify the position of the frames within the composite images. The model processes pixel data to generate outputs indicating the level of confidence and spatial area containing the embedded frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional detection technologies are used to detect video frames, then the system is simple and easy to operate, but the detection precision is insufficient for embedded frames within composite images

Engineering Contradiction:
Improvedetection precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces conventional mechanical/image processing detection systems with a machine learning-based system. The machine learning model is trained to recognize patterns and characteristics of embedded video frames within composite images, enabling accurate detection without relying on traditional image processing mechanisms. This substitution allows the system to achieve high detection precision for embedded frames while maintaining operational simplicity through automated learning-based detection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If the video frame is reduced in size within the composite image, then the embedding is more effective and harder to detect, but the detection capability of conventional technologies becomes even more insufficient

Engineering Contradiction:
Improveembedding effectivenessVSAvoiddetection difficulty
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent employs machine learning models that can adapt to various parameter changes in the embedded frames, including size reductions. The model is trained to recognize frames at different scales and resolutions, allowing it to maintain effective detection capability even when frames are significantly reduced in size within composite images. This parameter adaptation enables the system to overcome the increased detection difficulty while preserving embedding effectiveness.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If machine learning model is trained to detect embedded frames, then the detection accuracy improves, but the training process and model complexity increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a preliminary training phase where the machine learning model is extensively trained on composite images containing embedded video frames. During this preliminary action, the model learns to recognize patterns, characteristics, and spatial relationships of embedded frames. Once trained, the model can accurately detect embedded frames in new composite images without requiring complex real-time processing, thus achieving high detection accuracy while managing model complexity through pre-computed learning parameters.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12217142B2Using machine learning to detect which part of the screen includes embedded frames of an uploaded video
Publication Date: 2025.02.04 GOOGLE LLC
  • US12217142B2 patent drawing
  • US12217142B2 patent drawing
  • US12217142B2 patent drawing

AI summary

A system and methods are disclosed for using a trained machine learning model to identify constituent images within composite images. A method may include providing data identifying a first image as input to a machine learning model trained using training data identifying a plurality of composite images that each include one or more constituent images, and determining, using one or more outputs of the trained machine learning model, that the first image is a composite image that includes a first constituent image, wherein at least a portion of the first constituent image is in a spatial area of the first image, and wherein the first constituent image corresponds to a frame of a video embedded into the first image.