Voice Recognition End-of-Speech Detection via Dynamic Silence Duration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems struggle to accurately determine the end of speech input without pre-setting a category for search items, leading to inappropriate determination of input completion, especially when user features such as voice characteristics are not considered.
Innovation Solution
A voice recognition device and method that extracts features from input voice data to set a duration for determining the end of speech, allowing for flexible determination based on categories like addresses, facility names, telephone numbers, recognition errors, user age, and speech speed, ensuring accurate input completion detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a fixed duration is set for determining end of speech based on pre-set categories, then the determination process is simple, but the flexibility and accuracy deteriorate when user features or voice characteristics are not considered
Solution Approach 1:
The patent makes the duration setting dynamic by adjusting it based on extracted voice features (pitch, volume, timbre) and user characteristics (age, speech habits) rather than using a fixed predetermined value. The processor dynamically determines appropriate duration based on real-time voice analysis, allowing the system to adapt to different users and speaking conditions while maintaining operational simplicity through automated feature extraction.
Solution Approach 2:
The patent changes the parameter of duration from a fixed predetermined value to a variable determined by multiple factors including voice features (pitch, volume, timbre), user age, and speech habits. This parameter transformation allows the system to optimize end-of-speech determination accuracy for different users and contexts without increasing operational complexity, as the changes are automatically calculated based on extracted features.
2Ease of operation
If voice recognition determines end of speech without considering user features, then the system operation is simple, but the accuracy of input completion detection deteriorates
Solution Approach 1:
The system performs self-service by automatically extracting voice features (pitch, volume, timbre) and analyzing user characteristics (age, speech habits) to determine optimal duration settings without requiring manual configuration. The processor autonomously adjusts the end-of-speech determination parameters based on real-time voice analysis, maintaining ease of operation while significantly improving detection accuracy through adaptive feature-based determination.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors voice features and user responses to refine duration settings. By analyzing pitch variations, volume changes, and speech patterns, the system receives feedback on actual speech characteristics and adjusts the end-of-speech determination threshold accordingly, improving accuracy while maintaining simple operation through automated closed-loop control.
3Productivity
If a predetermined category is set for search items, then the voice recognition processing is efficient, but the adaptability to different user needs and speech patterns deteriorates
Solution Approach 1:
The patent creates a universal voice recognition system that handles multiple user needs and speech patterns through a single adaptive framework. By extracting general voice features (pitch, volume, timbre) and analyzing various user characteristics (age, speech habits, input speed), the system provides multi-functional adaptation across different users and contexts without requiring separate processing paths, maintaining efficiency while enhancing versatility.
Solution Approach 2:
The system dynamically adapts to different user needs by adjusting duration settings based on real-time analysis of voice features and user characteristics rather than relying on fixed category-based processing. This dynamic approach allows the system to efficiently handle diverse search items and user preferences through automated parameter adjustment, maintaining processing efficiency while significantly improving adaptability to individual user patterns.
Data Source
AI summary
A voice recognition device includes a memory and a processor including hardware. The processor is configured to extract a feature of input voice data and set a duration of a silent state after transition of the voice data to the silent state. The duration is used for determining that an input of the voice data is completed.


