Speech Recognizer Weight Learning for Accuracy and Complexity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in effectively selecting an optimum word string when multiple speech recognizers generate common errors, leading to incorrect answers and high computational complexity, and reducing the number of recognizers compromises performance and choice.
Innovation Solution
A recognizer weight learning apparatus that selects and learns weight values for a subset of speech recognizers based on characteristics, using a majority decision to minimize recognition errors and efficiently generate an optimum recognition result.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple speech recognizers are used to improve recognition accuracy through majority decision, then recognition reliability is improved, but device complexity and computational load increase
Solution Approach 1:
The patent changes the parameter of recognizer selection from fixed to dynamic by introducing weight values that are learned and updated based on performance. This allows the system to adaptively adjust which recognizers are used and with what influence, resolving the contradiction by maintaining reliability through selective use of recognizers while reducing complexity by not always using all recognizers at full capacity
Solution Approach 2:
The system transitions from a static configuration where all recognizers are treated equally to a dynamic configuration where weight values are continuously learned and adjusted. This dynamic adaptation allows the system to optimize the balance between using multiple recognizers for accuracy and managing system complexity by learning when to rely more or less on specific recognizers
2Measurement precision
If all speech recognizers are used for recognition processing, then recognition precision is improved, but computational load increases
Solution Approach 1:
The patent introduces weight values as a new parameter that controls the contribution of each recognizer. By learning optimal weight values, the system can dynamically adjust the computational resources allocated to each recognizer, maintaining high recognition precision while reducing overall computational load by not fully utilizing all recognizers in all situations
Solution Approach 2:
The system implements partial action by using weight values to selectively emphasize or de-emphasize certain recognizers based on their learned performance. This allows the system to achieve high precision by focusing computational resources on the most reliable recognizers for each specific input, rather than uniformly processing all recognizers
3Device complexity
If the number of speech recognizers is reduced to decrease computational complexity, then device complexity is reduced, but recognition performance deteriorates
Solution Approach 1:
Instead of changing the number of recognizers, the patent changes the parameter of their utilization through weight values. This allows the system to maintain a large pool of recognizers for potential use while controlling computational complexity by learning to rely more heavily on a subset of high-performing recognizers, thus preserving recognition performance without requiring all recognizers to be actively processed
4Reliability
If weight values are learned for speech recognizers to minimize error rate, then recognition reliability is improved, but learning time and processing time increase
Solution Approach 1:
The patent applies preliminary action by pre-learning weight values during a training phase before actual recognition operations. This separates the time-consuming learning process from the operational phase, allowing the system to quickly apply learned weights during recognition without adding significant time to the operational workflow, thus improving reliability while minimizing time loss in practice
Data Source
AI summary
A speech recognition apparatus (110) selects an optimum recognition result from recognition results output from a set of speech recognizers (s1-sM) based on a majority decision. This decision is implemented with taking into account weight values, as to the set of the speech recognizers, learned by a learning apparatus (100). The learning apparatus includes a unit (103) selecting speech recognizers corresponding to characteristics of speech for learning (101), a unit (104) finding recognition results of the speech for learning by using the selected speech recognizers, a unit (105) unifying the recognition results and generating a word string network, and a unit (106) finding weight values concerning a set of the speech recognizers by implementing learning processing. When finding weight values, the learning apparatus selects a word from each arc set in the word string network based on a majority decision which is taken into account candidates of weight value, and outputs weight value candidates which minimize a recognition error rate of a word string formed of the selected words, as a learning result.


