Audio Processing Apparatus Voice Section Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems face challenges in improving voice recognition accuracy due to the inclusion of non-voice sections in the audio signal processing, which degrades the performance of voice recognition processes.
Innovation Solution
An audio processing apparatus and method that includes a first-section detection unit to identify sections with high spatial spectrum power, a speech state determination unit to assess speech presence, a likelihood calculation unit to determine voice and non-voice likelihoods, and a second-section detection unit to accurately identify voice sections based on these likelihoods, thereby isolating voice sections for enhanced voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recognition is performed on the separated sound from speaker sound source, then voice recognition can be conducted, but non-voice sections are included which degrade voice recognition accuracy
Solution Approach 1:
The audio signal is divided into multiple frames, and each frame is analyzed to determine whether it belongs to a voice section or non-voice section based on spectral power comparison with threshold values. This segmentation allows the system to process only relevant voice portions, improving recognition accuracy while excluding non-voice content.
Solution Approach 2:
The invention extracts and identifies voice sections from the mixed audio signal by comparing spectral power characteristics against predetermined thresholds. By separating and isolating only the voice-containing frames, the system removes non-voice sections that would otherwise degrade recognition performance.
2Measurement precision
If spectral power comparison with predetermined threshold is used for each frame, then voice sections can be detected, but voice recognition accuracy is degraded due to non-voice section inclusion
Solution Approach 1:
Before performing voice recognition, the system pre-processes the audio signal by analyzing each frame's spectral power and comparing it with predetermined thresholds to identify voice sections. This preliminary classification ensures that only frames containing actual voice content are passed to the voice recognition engine, preventing non-voice sections from degrading accuracy.
Data Source
AI summary
An audio processing apparatus includes a first-section detection unit configured to detect a first section that is a section in which the power of a spatial spectrum in a sound source direction is higher than a predetermined amount of power on the basis of an audio signal of a plurality of channels, a speech state determination unit configured to determine a speech state on the basis of an audio signal within the first section, a likelihood calculation unit configured to calculate a first likelihood that a type of sound source according to an audio signal within the first section is voice and a second likelihood that the type of sound source is non-voice, and a second-section detection unit configured to determine whether or not a second section in which power is higher than average the power of a speech section is a voice section on the basis of the first likelihood and the second likelihood within the second section.


