Voice Segment Detection for Mixed Audio Audibility Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The manual editing tasks required to ensure voice is audible in video content, such as raising or lowering volumes and changing equalization, are costly and inefficient.
Innovation Solution
A signal processing device and method that automates the detection and determination of voice segments in mixed audio signals using machine learning, allowing for automated adjustments to improve voice audibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual editing tasks are performed to ensure voice audibility, then voice clarity is improved, but production cost increases
Solution Approach 1:
The system automatically detects voice segments and determines audibility without human intervention. The voice detection unit identifies time segments containing voice sounds, and the voice determination unit automatically assesses whether the voice is easy to hear by comparing detection results with label information, eliminating the need for manual editing while maintaining voice clarity assessment
Solution Approach 2:
Manual mechanical editing operations are replaced by an automated signal processing system using machine learning models. The voice detection unit and voice determination unit substitute human editors by automatically analyzing audio signals, detecting voice segments, and determining audibility through computational processing
2Measurement precision
If manual editing tasks are performed to ensure voice audibility, then voice clarity is improved, but productivity decreases
Solution Approach 1:
The automated system performs voice detection and audibility determination independently without requiring manual editing operations. The voice detection unit automatically identifies voice segments, and the voice determination unit autonomously assesses audibility by comparing detection results with reference label information, significantly improving content production efficiency
Solution Approach 2:
Manual editing operations are replaced by automated machine learning-based signal processing. The system rapidly detects voice segments and determines audibility through computational algorithms, eliminating the time-consuming manual editing process while maintaining accurate voice clarity assessment
3Productivity
If automated voice detection is implemented, then productivity is improved, but device complexity increases
Solution Approach 1:
The signal processing system is divided into distinct functional modules: a voice detection unit that identifies voice segments in mixed audio signals, and a voice determination unit that assesses audibility. This segmentation allows each module to perform a specific function, managing system complexity through modular design while maintaining high productivity
Solution Approach 2:
Label information serves as an intermediary reference that bridges the voice detection process and the audibility determination process. The voice determination unit compares detection results with pre-prepared label information to objectively assess whether voice is easy to hear, providing a structured method that manages system complexity
Data Source
AI summary
Provided is a signal processing device that includes a voice detection unit that, based on a mixed audio signal containing a sound of a target sound source and a sound of a non-target sound source different from the target sound source, detects a time segment of the sound of the target sound source from the mixed audio signal; and a voice determination unit that, based on label information indicating the time segment of the sound of the target sound source in an audio signal of the target sound source and a detection result for the time segment of the sound of the target sound source, performs determination processing for determining whether the sound of the target sound source in the mixed audio signal is easy to hear.


