Speech Recognition Device Using Reading Information Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition devices require large storage capacity and cannot perform real-time speech recognition across multiple languages without pre-generated phonetic information for unsupported languages.
Innovation Solution
A speech recognition device that converts reading information of words from different languages into a predetermined language using a reading information conversion database, allowing real-time speech recognition without the need for pre-stored phonetic information for each language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If phonetic information for multiple languages is stored in advance, then speech recognition accuracy for multiple languages is improved, but storage capacity requirements increase and real-time processing capability deteriorates
Solution Approach 1:
The patent extracts only the essential phonetic information (reading information) needed for speech recognition from complete language data, storing minimal necessary data in the recognition dictionary while maintaining high recognition accuracy. This allows the system to support multiple languages without storing excessive phonetic information for each language.
Solution Approach 2:
The speech recognition engine is designed to universally process multiple languages using a single engine that can handle different languages by referencing language-specific entries in the recognition dictionary. This eliminates the need for separate speech recognition engines for each language, reducing storage requirements while maintaining multi-language capability.
2Adaptability or versatility
If phonetic information for multiple languages is stored in advance, then speech recognition capability for multiple languages is improved, but processing time increases and real-time performance deteriorates
Solution Approach 1:
The recognition dictionary is pre-prepared with essential language-specific entries (writing information, reading information, and language identification) before speech recognition begins. This preliminary preparation allows the speech recognition engine to quickly identify and process the appropriate language without performing complex language detection during real-time processing, thus maintaining real-time performance while supporting multiple languages.
Solution Approach 2:
The system applies different processing strategies to different languages by storing language-specific characteristics in the recognition dictionary. Each language entry contains localized phonetic information and language identification markers, allowing the engine to efficiently process each language according to its specific properties without uniformly processing all languages the same way, thereby reducing overall processing time.
3Device complexity
If a single speech recognition engine is used for multiple languages, then device complexity is reduced, but speech recognition accuracy for each language may deteriorate
Solution Approach 1:
The patent changes the parameters stored in the recognition dictionary to include language identification information and language-specific phonetic characteristics. By modifying the data structure to include language tags and language-specific reading information, a single speech recognition engine can adapt to different languages by referencing these parameters, maintaining high accuracy without requiring separate engines for each language.
Data Source
AI summary
A speech recognition device includes: a speech recognition unit 23a that performs speech recognition for input speech; a reading information conversion data base in which a reading information conversion rule L is registered; a reading information conversion unit 27a that converts reading information of the word among the languages based on the rule L; and a speech recognition control unit 24a that performs control such that, when a word in a different language that is different from a predetermined language is included in a recognition subject vocabulary in which a speech recognition unit 23a refers to recognition subject word information E, the unit 27a converts the reading information in the different language into reading information in the predetermined language, and that the unit 23a performs the speech recognition that makes reference to the recognition subject word information of the corresponding word, including the converted reading information in the predetermined language.


