Dynamic Speech Section Detection Level Adjustment for Voice Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice interaction systems face inefficiencies in noise reduction and power consumption when performing voice recognition in noisy environments, as existing speech section detection methods fail to accurately distinguish speech from noise, leading to increased communication costs and power usage.
Innovation Solution
A control apparatus that dynamically adjusts the identification level of a speech section detector based on noise levels and distance, lowering the detection level when it's likely that a target person is speaking to improve accuracy while minimizing unnecessary data transmission and power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If speech section detection is performed with a high identification level to reduce noise transmission, then communication cost and power consumption are reduced, but voice recognition accuracy deteriorates due to false identification of speech as noise
Solution Approach 1:
The identification level of the speech section detector is dynamically adjusted based on the operational state of the voice interaction apparatus. When voice recognition is actively performed, the identification level is lowered to improve speech detection accuracy; when voice recognition is not active, the identification level is raised to reduce noise transmission and save power. This dynamic adjustment resolves the contradiction between power efficiency and recognition accuracy.
Solution Approach 2:
The system changes the identification level parameter of the speech section detector according to different operational conditions. By adjusting this parameter based on whether voice recognition is being performed, the system optimizes both power consumption and voice recognition accuracy for different states of operation.
2Measurement precision
If speech section detection is performed with a low identification level to improve speech detection accuracy, then voice recognition accuracy is improved, but communication cost and power consumption increase due to transmission of noise data
Solution Approach 1:
The identification level is dynamically adjusted based on operational needs. During active voice recognition, the level is lowered to capture all speech; during inactive periods, the level is raised to filter noise and conserve power, thus resolving the contradiction between detection sensitivity and energy efficiency.
Solution Approach 2:
The system modifies the identification level parameter according to the operational state, lowering it during voice recognition activities to improve accuracy and raising it during idle periods to reduce power consumption and communication overhead.
3Measurement precision
If constant communication with voice recognition server is performed to ensure accurate voice recognition, then voice recognition accuracy is improved, but communication cost and power consumption are wastefully increased during non-speech periods
Solution Approach 1:
The system extracts and transmits only the relevant speech sections to the voice recognition server by using the speech section detector to identify and separate actual speech from noise. This selective transmission reduces the volume of communication data while maintaining voice recognition accuracy.
Solution Approach 2:
Instead of constant communication, the system performs periodic speech section detection and transmits data only during periods when speech is detected. This periodic approach reduces communication frequency and data volume during non-speech periods while ensuring accurate voice recognition when needed.
4Quantity of substance
If speech section detection is performed to reduce data transmission, then communication cost is reduced, but voice recognition accuracy deteriorates due to insufficient or false speech detection
Solution Approach 1:
The speech section detector extracts actual speech from the mixed audio signal by comparing against noise profiles. This extraction process identifies genuine speech sections for transmission while filtering out noise, thus reducing communication data volume without compromising voice recognition accuracy.
Solution Approach 2:
The system uses feedback from the voice recognition server to adjust the speech section detection parameters. When the server indicates successful recognition, the system continues current detection settings; when recognition fails, the system adjusts detection sensitivity to improve future speech identification accuracy.
Data Source
AI summary
The control apparatus includes: a calculation unit configured to control a voice interaction apparatus including a speech section detector, the speech section detector being configured to identify whether an acquired voice includes a speech made by a target person by a set identification level and perform speech section detection, in which the calculation unit instructs, when an estimation result indicating that it is highly likely that the speech made by the target person is included in the acquired voice has been acquired from a voice recognition server, the voice interaction apparatus to change a setting in such a way as to lower the identification level of the speech section detector, and to perform communication with the voice recognition server by speech section detection in accordance with the identification level after the change.


