Voice Segment Detection via Client-Server Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In server-client voice recognition systems, client terminals with limited resources often fail to accurately detect voice segments, leading to missed voice transmission due to insufficient processing capacity compared to server devices.
Innovation Solution
A voice segment detection system that employs a cooperative mechanism between a voice starting end detection apparatus and a voice terminal end detection apparatus, where the voice starting end detection apparatus transmits input signals subsequent to detected voice starts to the voice terminal end detection apparatus, which has sufficient resources for accurate terminal end detection, thereby reducing missed transmissions and communication load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If the client terminal executes voice segment detection independently, then the communication amount is reduced, but the detection accuracy deteriorates due to limited resources
Solution Approach 1:
The voice segment detection process is divided into two separate detection tasks: voice start detection performed by the client terminal and voice end detection performed by the server device. This segmentation allows each device to perform only the detection task it is best suited for, reducing communication overhead while maintaining high detection accuracy through server-side verification.
Solution Approach 2:
The server device acts as an intermediary that receives audio data from the client terminal, performs voice end detection on the received data, and sends control signals back to the client terminal. This intermediary role enables the server to leverage its superior resources for accurate detection while the client maintains low communication overhead.
2Device complexity
If the client terminal has limited processing resources, then the device complexity is reduced, but the voice segment detection accuracy deteriorates
Solution Approach 1:
The detection functionality is segmented between client and server: the client performs simple voice start detection with low computational requirements, while the server performs the more computationally intensive voice end detection. This segmentation allows the client terminal to maintain low device complexity while achieving high overall detection accuracy through server assistance.
3Measurement precision
If more voice data is transmitted to ensure complete capture, then the detection accuracy is improved, but the communication overhead increases
Solution Approach 1:
The server device sends voice end detection results as feedback to the client terminal, which then stops transmitting audio data when the end of speech is detected. This feedback mechanism ensures complete voice segment capture with high accuracy while minimizing communication overhead by preventing unnecessary transmission of post-speech data.
Data Source
AI summary
A voice starting end detection apparatus includes a first detector that detects a starting end of a voice segment from input signals that are input in a time series, a first transmitting unit that transmits, when the starting end is detected, input signals subsequent to the starting end, and a first receiving unit that receives a terminal end detection signal indicating that a terminal end of the voice segment has been detected. The voice terminal end detection apparatus includes a second receiving unit that receives input signals subsequent to the starting end, a second detector that detects the terminal end from the received input signals, a second transmitting unit that transmits, when the terminal end is detected, the terminal end detection signal. The first transmitting unit stops transmitting the input signals when the first receiving unit receives the terminal end detection signal.


