Script-Based Visual Content Detection for Media Programs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for determining visual content in media programs, such as object detection algorithms, are computationally expensive and may miss masked or partially obscured content, requiring efficient alternatives for compliance with rating standards and content moderation.

Innovation Solution

The system parses production scripts and subtitles to identify visual content descriptions, using natural language processing to associate script elements with timestamps and detect regulated themes, potentially augmenting or replacing vision-based algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If object detection algorithms are used to determine visual content, then detection capability is provided, but computational cost becomes excessively high

Engineering Contradiction:
Improvevisual content detection capabilityVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent introduces a script as an intermediary element that contains descriptive information about visual content. Instead of directly analyzing video frames through computationally expensive object detection algorithms, the system uses the script's textual descriptions as a mediator to identify visual content, thereby reducing computational costs while maintaining detection capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical vision-based object detection system with a text-processing system. By substituting image analysis with script parsing and natural language processing, the system achieves visual content determination without the high computational burden of frame-by-frame analysis of thousands of images.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If vision-based algorithms analyze every image frame, then complete visual content identification is achieved, but processing time becomes excessively long

Engineering Contradiction:
Improvevisual content identification completenessVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts the essential visual content information from the video by relying on the script's descriptive text rather than analyzing every image frame. This extraction approach maintains identification completeness for regulated content while dramatically reducing processing time by skipping the time-consuming frame-by-frame analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The script, which contains pre-written descriptions of visual content, actions, and scenes, serves as a preliminary resource that provides visual content information before actual video analysis would be needed. This preliminary action allows the system to identify regulated content quickly without performing extensive post-production analysis.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If object detection algorithms are used, then visual content can be identified, but masked or partially obscured content may escape recognition

Engineering Contradiction:
Improvevisual content recognition accuracyVSAvoiddetection failure on masked content
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The script acts as an intermediary that describes visual content including masked or obscured elements. Since the script contains textual descriptions of actions and content that may not be fully visible in frames, the system can identify regulated content through the script's descriptions even when vision-based algorithms fail to detect masked or partially obscured material.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11917222B1Determining visual content of media programs from scripts
Publication Date: 2024.02.27 AMAZON TECH INC
  • US11917222B1 patent drawing
  • US11917222B1 patent drawing
  • US11917222B1 patent drawing

AI summary

Visual content of media programs is recognized using text-based scripts of the media programs. A script of a media program includes sets of words to be spoken by actors during the media program, and also descriptions of points of interest of the media program. A file of subtitles or captions of a media program includes sets of words actually spoken by actors during the media program along with marks or stamps of times when such words were spoken. The sets of words of a script and in subtitles or captions are processed to determine where such words align. The sets of words of the script are then marked or stamped with times from the subtitles. The descriptions may then be processed to predict the visual content from such descriptions, and times at which the visual content appears within the media program are determined from the sets of words.