AV Narration Insertion Using Dialogue and Music Gap Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating narration for audio-video content without interfering with dialog or music is a time-consuming and cumbersome process.

Innovation Solution

A processor system identifies gaps in dialog and music within AV content, generates narration segments that fit within these gaps, and synchronizes the narration with the AV content's loudness and emotion to ensure seamless integration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If narration is generated and inserted into AV content, then accessibility and enjoyment are improved, but the process is time-consuming and cumbersome

Engineering Contradiction:
ImproveaccessibilityVSAvoidtime-consuming process
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically analyzing the AV content to identify gaps in dialog and music before generating narration. It pre-processes the audio track to detect silence periods and temporal gaps, then generates narration segments that fit these pre-identified gaps, eliminating the need for manual timing and placement by the user.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system serves itself by automatically detecting dialog gaps, generating appropriate narration, and synchronizing it with the AV content without requiring manual intervention. The automated pipeline handles the entire process from gap detection to narration insertion, making the system self-sufficient and eliminating the cumbersome manual process.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If narration is inserted into AV content, then accessibility is improved, but the narration may interfere with dialog or music

Engineering Contradiction:
ImproveaccessibilityVSAvoidinterference with dialog or music
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system extracts and removes existing dialog and music segments to identify gaps where narration can be inserted without interference. By analyzing the audio track and isolating silence periods and temporal gaps between dialog segments, the system creates dedicated spaces for narration that do not overlap with or disrupt the original content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies local quality by placing narration only in specific local regions where gaps exist in the dialog or music. Rather than uniformly adding narration throughout, it strategically positions narration segments only in temporal gaps identified in the audio analysis, ensuring that narration is added only where it will not interfere with existing content.

Inventive Principle:
Principle #3Local quality

3Reliability

If narration segments are generated to fit gaps, then seamless integration is achieved, but the process complexity increases

Engineering Contradiction:
Improveseamless integrationVSAvoidprocess complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the narration generation process into distinct modular steps: (1) analyzing the audio track to identify gaps, (2) generating narration segments for each gap, (3) adjusting narration timing and duration to fit the gaps precisely, and (4) synthesizing the final output. This segmentation breaks down the complex task into manageable, sequential operations that can be automated independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs parameter changes by dynamically adjusting narration duration, speed, and timing based on the identified gap characteristics. It modifies narration parameters such as playback speed and segment length to precisely match the temporal gaps in the AV content, ensuring seamless integration while automating the complex synchronization process.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12549824B2Machine narration
Publication Date: 2026.02.10 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12549824B2 patent drawing
  • US12549824B2 patent drawing
  • US12549824B2 patent drawing

AI summary

A technique for generating and inserting voice narration about action in audio-video (AV) content such as a movie or computer game includes generating the narration, e.g., from dialog in the AV content, and determining how and when to insert portions of the narration into the AV content so as not to interfere with dialog in the content.