Video Frame Saliency Maps for Automated Secondary Content Placement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing media asset creation tasks, such as generating personalized promotional assets, are labor-intensive, costly, and unscalable due to the reliance on human artistic perception for identifying salient and non-salient regions in video frames.

Innovation Solution

A computer-implemented method using a machine learning model that generates pixel-wise score maps based on human-perception inspired saliency cues to automate media asset creation tasks by identifying salient and non-salient regions in video frames, allowing for the automated placement of secondary content without obstructing important information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If human artistic perception is used to identify salient and non-salient regions in video frames, then the quality of media asset creation is maintained, but the process becomes labor-intensive, costly, and unscalable

Engineering Contradiction:
Improvemedia asset creation throughputVSAvoidautomation of salient region identification
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The patent creates synthetic saliency maps that copy and simulate human artistic perception through machine learning. These synthetic maps replicate the visual importance assessment that human artists would make, allowing automated systems to identify salient and non-salient regions without human intervention while maintaining quality standards

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical human visual inspection process with an automated machine learning system. The system uses trained models to process video frames and generate saliency maps, substituting human artistic perception with algorithmic analysis that can operate at scale without labor constraints

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If manual methods are used for media asset creation, then quality control is maintained, but time consumption and costs increase significantly

Engineering Contradiction:
Improvetime for media asset creationVSAvoidquality of salient region identification
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent performs preliminary training of machine learning models using synthetic saliency maps before actual media asset creation. This pre-training phase establishes the system's ability to accurately identify salient regions, ensuring quality control is built into the automated process from the start rather than requiring manual quality checks during production

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses synthetic saliency maps as training data to copy human artistic judgment patterns. By training on these synthetic maps that represent human-perceived important regions, the system learns to maintain quality standards automatically, matching the precision of human artists while operating much faster

Inventive Principle:
Principle #26Copying

3Productivity

If automated systems are introduced to increase productivity, then scalability improves, but the accuracy of identifying important visual regions may deteriorate

Engineering Contradiction:
Improvescalability of media asset creationVSAvoidaccuracy of salient region detection
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary training with synthetic saliency maps to establish accurate detection capabilities before deployment. This advance preparation ensures the automated system achieves human-level precision in identifying salient regions, eliminating the trade-off between automation and accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the detection task into a parameter-based classification problem using saliency score maps. By converting visual importance into quantifiable score parameters, the system can accurately identify salient regions through numerical thresholds while maintaining high productivity and scalability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12477160B1Computer-implemented methods for determining salient and non-salient regions in video frames
Publication Date: 2025.11.18 AMAZON TECH INC
  • US12477160B1 patent drawing
  • US12477160B1 patent drawing
  • US12477160B1 patent drawing

AI summary

Techniques for a computer-implemented service that utilizes a machine learning model to identify the salient and/or non-salient regions in a video frame are described. According to some embodiments, a computer-implemented method includes receiving a frame of a video at a content delivery service, generating, by a machine learning model of the content delivery service, a per pixel salience score map for the frame, and inserting, by the content delivery service, secondary content into the frame based at least in part on the per pixel salience score map for the frame.