Voice Segment Detection for Mixed Audio Audibility Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The manual editing tasks required to ensure voice is audible in video content, such as raising or lowering volumes and changing equalization, are costly and inefficient.

Innovation Solution

A signal processing device and method that automates the detection and determination of voice segments in mixed audio signals using machine learning, allowing for automated adjustments to improve voice audibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual editing tasks are performed to ensure voice audibility, then voice clarity is improved, but production cost increases

Engineering Contradiction:
Improvevoice audibilityVSAvoidproduction cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system automatically detects voice segments and determines audibility without human intervention. The voice detection unit identifies time segments containing voice sounds, and the voice determination unit automatically assesses whether the voice is easy to hear by comparing detection results with label information, eliminating the need for manual editing while maintaining voice clarity assessment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical editing operations are replaced by an automated signal processing system using machine learning models. The voice detection unit and voice determination unit substitute human editors by automatically analyzing audio signals, detecting voice segments, and determining audibility through computational processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual editing tasks are performed to ensure voice audibility, then voice clarity is improved, but productivity decreases

Engineering Contradiction:
Improvevoice audibilityVSAvoidcontent production efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The automated system performs voice detection and audibility determination independently without requiring manual editing operations. The voice detection unit automatically identifies voice segments, and the voice determination unit autonomously assesses audibility by comparing detection results with reference label information, significantly improving content production efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual editing operations are replaced by automated machine learning-based signal processing. The system rapidly detects voice segments and determines audibility through computational algorithms, eliminating the time-consuming manual editing process while maintaining accurate voice clarity assessment

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated voice detection is implemented, then productivity is improved, but device complexity increases

Engineering Contradiction:
Improvecontent production efficiencyVSAvoidsignal processing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The signal processing system is divided into distinct functional modules: a voice detection unit that identifies voice segments in mixed audio signals, and a voice determination unit that assesses audibility. This segmentation allows each module to perform a specific function, managing system complexity through modular design while maintaining high productivity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Label information serves as an intermediary reference that bridges the voice detection process and the audibility determination process. The voice determination unit compares detection results with pre-prepared label information to objectively assess whether voice is easy to hear, providing a structured method that manages system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12505853B2Signal processing device and method
Publication Date: 2025.12.23 SONY GROUP CORP
  • US12505853B2 patent drawing
  • US12505853B2 patent drawing
  • US12505853B2 patent drawing

AI summary

Provided is a signal processing device that includes a voice detection unit that, based on a mixed audio signal containing a sound of a target sound source and a sound of a non-target sound source different from the target sound source, detects a time segment of the sound of the target sound source from the mixed audio signal; and a voice determination unit that, based on label information indicating the time segment of the sound of the target sound source in an audio signal of the target sound source and a detection result for the time segment of the sound of the target sound source, performs determination processing for determining whether the sound of the target sound source in the mixed audio signal is easy to hear.