Speech Recognition Service Refinement via Experience Points
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The initial operation of a speech recognition service on electronic devices often results in a low speech recognition rate due to limited user and device information, leading to reduced efficiency and reliability.
Innovation Solution
An integrated intelligent system that includes a user terminal, intelligence server, and personal information server to collect and utilize user and device information, with an experience point calculation mechanism for the artificial intelligence assistant to refine its speech recognition capabilities based on user interactions and feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the speech recognition service is operated with limited user and device information in initial operation, then the system complexity is reduced, but the speech recognition rate and reliability deteriorate
Solution Approach 1:
The system performs preliminary information collection during initial operation, gathering user and device data in advance to build recognition models. This preliminary action ensures that sufficient information is available before formal speech recognition begins, resolving the contradiction between starting with simple systems and achieving high recognition accuracy.
Solution Approach 2:
The speech recognition system dynamically adapts its information collection and processing capabilities based on operational stage. In initial operation, it collects information gradually; as it matures, it utilizes accumulated data to improve recognition rates, making the system's complexity and performance dynamic rather than static.
2Reliability
If user and device information is collected continuously to improve speech recognition, then the speech recognition rate improves, but the information processing complexity increases
Solution Approach 1:
The information collection and processing system is segmented into distinct modules: information collection module, information processing module, and speech recognition module. Each module handles specific tasks independently, reducing overall system complexity while enabling continuous information gathering to improve recognition rates.
Solution Approach 2:
An information processing intermediary layer is introduced between raw data collection and speech recognition. This intermediary processes and structures user and device information before feeding it to the recognition engine, reducing the complexity burden on the core recognition system while maintaining continuous improvement capabilities.
3Ease of operation
If the speech recognition service operates without sufficient user information, then the ease of operation is maintained, but the recognition accuracy and reliability worsen
Solution Approach 1:
The speech recognition system provides self-service by automatically collecting and processing user information without requiring manual configuration or user intervention. This maintains ease of operation while gradually accumulating data to improve recognition accuracy over time.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
An electronic device is provided. The electronic device includes a communication module, a microphone receiving a voice input according to user speech, a memory storing information about an operation of the speech recognition service, a display, and a processor electrically connected with the communication module, the microphone, the memory, and the display. The processor is configured to calculate a specified numerical value associated with the operation of the speech recognition service, to transmit information about the numerical value to a first external device processing the voice input, and to transmit a request for a function, which corresponds to the calculated numerical value, of at least one function associated with the speech recognition service stepwisely provided from the first external device depending on a numerical value, to the first external device to refine a function of the speech recognition service supported by the electronic device.