Adaptive Voice Endpoint Detection for User-Specific Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition systems face challenges in accurately determining the end-point detection time, leading to incomplete recognition of user inputs when the interval is too short or increased misrecognition due to noise when it is too long, necessitating a more precise method for enhancing voice recognition performance.
Innovation Solution
A voice end-point detection system that varies the end-point detection time for each user and domain, using a processor to set and adjust detection times based on recognition results, including multiple end-point detection times for different speech scenarios, and a database to store and manage these settings for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If the end-point detection time is set to be short, then the recognition time is reduced, but the user input may be incompletely recognized
Solution Approach 1:
The patent applies dynamics by making the end-point detection time adjustable and adaptive rather than fixed. The system dynamically changes the detection time based on user identity and domain, allowing optimization between speed and accuracy for different scenarios.
Solution Approach 2:
The patent changes the parameter of end-point detection time based on user characteristics and domain. By storing multiple detection time values and selecting appropriate ones, the system optimizes both recognition speed and accuracy for different users and contexts.
2Reliability
If the end-point detection time is set to be long, then the complete user input can be captured, but misrecognition may increase due to noise
Solution Approach 1:
The patent applies local quality by setting different end-point detection times for different users and domains rather than using a uniform setting. This allows optimization for each specific context, capturing complete input where needed while minimizing noise impact where possible.
Solution Approach 2:
The system changes the detection time parameter based on user identity and domain characteristics. By selecting appropriate detection time values from stored options, the system balances complete input capture with noise reduction for different scenarios.
3Ease of operation
If a fixed end-point detection time is used, then the system is simple to operate, but it cannot adapt to different user speech patterns
Solution Approach 1:
The patent makes the system adaptive by dynamically selecting end-point detection times based on user identity and domain. The system automatically adjusts parameters without requiring manual configuration, maintaining ease of operation while improving adaptability.
Solution Approach 2:
The system serves itself by automatically selecting appropriate detection time parameters based on stored user profiles and domain information. No manual intervention is needed - the system self-adjusts to optimize performance for each user and context.
Data Source
AI summary
A voice end-point detection device, a system and a method are provided. The voice end-point detection system includes a processor that is configured to determine an end-point detection time to detect an end-point of speaking of a user that varies vary for each user and for each domain. The voice end-point detection system is configured to perform voice recognition and a database (DB) is configured to store data for the voice recognition by the processor.


