Speech Recognition Semiconductor Circuit Dynamic Pattern Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face challenges in efficiently updating option information during execution and achieving improved recognition rates due to constant recognition accuracy settings, which are not adaptable to varying numbers of options or similar patterns.
Innovation Solution
A semiconductor integrated circuit device and method that dynamically adjust the range of standard patterns and recognition accuracy parameters based on the number of options and similarity of patterns, using a conversion list to narrow down comparisons and facilitate updates, thereby improving recognition rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the range of standard patterns to be checked is not restricted, then the recognition rate improves, but the processing time increases significantly
Solution Approach 1:
The patent segments the speech recognition database into multiple categories (e.g., food categories, drink categories) and further divides each category into subcategories. This hierarchical segmentation allows the system to check standard patterns within restricted subcategories rather than the entire database, improving processing speed while maintaining recognition accuracy for category-specific speech inputs.
Solution Approach 2:
The patent performs preliminary classification of speech inputs into categories before detailed pattern matching. By pre-categorizing the recognition database and determining the relevant category based on contextual information, the system narrows down the search scope in advance, reducing the number of patterns to compare without sacrificing recognition quality.
2Measurement precision
If the recognition accuracy parameter is set to a fixed high value, then the measurement precision improves, but the number of matches decreases
Solution Approach 1:
The patent dynamically adjusts the recognition accuracy parameter (threshold value) based on the category and subcategory of the speech input. Different categories have different optimal threshold values that balance precision and recall requirements. This dynamic adjustment allows the system to maintain high accuracy where needed while accepting more matches in categories where broader matching is appropriate.
Solution Approach 2:
The patent changes the recognition threshold parameter according to the specific category being processed. By storing multiple threshold values associated with different categories and subcategories, the system can select the appropriate parameter setting for each recognition task, optimizing the balance between measurement precision and the quantity of valid matches.
3Ease of operation
If text information is converted into candidate data for speech recognition, then the ease of operation improves, but the processing load increases
Solution Approach 1:
The patent extracts only the necessary text information (titles, names, or other identification data) from the database and converts it into candidate data for speech recognition. By selectively extracting only the relevant information needed for recognition rather than processing all database contents, the system reduces the processing load while still providing convenient text-based candidate generation for users.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances speech recognition accuracy and efficiency by adaptively restricting the range of options and adjusting recognition parameters, allowing for better handling of deep hierarchical menus and improving overall recognition rates.
Implementation Method 1
a signal processing unit that extracts frequency components of an input speech signal by computing a Fourier transform of the speech signal
Data Source
AI summary
A semiconductor integrated circuit device for speech recognition includes a conversion candidate setting unit that receives text data indicating words or sentences together with a command and sets the text data in a conversion list in accordance with the command; a standard pattern extracting unit that extracts, from a speech recognition database, a standard pattern corresponding to at least a part of the words or sentences indicated by the text data that is set in the conversion list; a signal processing unit that extracts frequency components of an input speech signal and generates a feature pattern indicating distribution of the frequency components; and a match detecting unit that detects a match between the feature pattern generated from at least a part of the speech signal and the standard pattern and outputs a speech recognition result.


