Speech-to-Text Conversion with Dynamic Language Priority Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information handling devices often inaccurately perform speech-to-text conversions, particularly when dealing with multiple language variations, as they typically recognize only one primary language and fail to correctly process dialects or combinations of languages.

Innovation Solution

An apparatus and method that determine a priority ranking for multiple language variations, including languages and dialects, to accurately convert audible inputs to text, with the ability to detect and correct errors by switching to a higher-priority language variation based on user selection or geographic location, using a processor, sensor, and memory to execute code that handles language processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system recognizes only one primary language, then the device complexity is reduced, but the accuracy of speech-to-text conversion deteriorates when dealing with multiple language variations

Engineering Contradiction:
Improveaccuracy of speech-to-text conversionVSAvoidlanguage processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the language processing into multiple ranked language models, where each model corresponds to a specific language variation (e.g., Standard English, British English, Australian English). The system processes speech through these segmented language models in sequence, stopping when a satisfactory transcription is achieved. This segmentation allows the system to handle multiple language variations without requiring a single monolithic complex processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic language model selection based on priority rankings. The system dynamically switches between different language models depending on the priority ranking assigned to each language variation. This dynamic approach enables the system to adapt to different language contexts and user preferences, improving accuracy without requiring all language models to operate simultaneously with equal complexity.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If the system processes multiple language variations simultaneously, then the adaptability improves, but the processing time increases

Engineering Contradiction:
Improvelanguage variation handling capabilityVSAvoidspeech-to-text processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-establishing priority rankings for different language variations before processing speech. The system pre-orders language models based on user preferences, geographic location, or other factors. This preliminary sorting allows the system to process languages in a predetermined sequence, avoiding the need to evaluate all language variations simultaneously and thus reducing processing time while maintaining high adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by processing language variations in a ranked sequence rather than all at once. The system attempts transcription using the highest-priority language model first, and only proceeds to lower-priority models if the initial attempt fails or produces unsatisfactory results. This partial processing approach maintains adaptability to multiple languages while significantly reducing the time required compared to simultaneous processing of all variations.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If the system uses automatic language detection based on geographic location, then the ease of operation improves, but the measurement precision may deteriorate when users are traveling or using devices in different locations

Engineering Contradiction:
Improvelanguage selection convenienceVSAvoidlanguage detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms that allow users to review and correct automatically detected language preferences. The system provides users with the opportunity to verify the detected language variation and adjust the priority rankings as needed. This feedback loop maintains ease of operation by reducing manual language selection requirements while preserving measurement precision by allowing users to correct automatic detection errors, especially useful when traveling or using devices in different locations.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances the accuracy of speech-to-text conversions by prioritizing language variations based on user selection and geographic location, ensuring correct processing of language and dialects, thereby improving the handling of multiple language inputs.

Implementation Method 1

The code, in certain embodiments, is executable by the processor to detect, by use of the sensor, an audible input

Methodology Applied
Scientific EffectSound wave detection: Sound

Data Source

PatentUS11093720B2Apparatus, method, and program product for converting multiple language variations
Publication Date: 2021.08.17 LENOVO SWITZERLAND INTERNATIONAL GMBH
  • US11093720B2 patent drawing
  • US11093720B2 patent drawing
  • US11093720B2 patent drawing

AI summary

Apparatuses, methods, and program products are disclosed for converting multiple language variations. One apparatus includes a processor, a sensor, and a memory that stores code executable by the processor. The code is executable by the processor to: determine a priority ranking corresponding to each language variation of multiple language variations, wherein each language variation includes a language, a dialect, or a combination thereof; detect, by use of the sensor, an audible input; convert the audible input to text based on a first language variation, wherein the priority ranking of the first language variation is a highest priority; and in response to a portion of the text being incorrect, convert the audible input corresponding to the portion of the text to a revised portion of the text based on a second language variation, wherein the priority ranking of the second language variation is a second highest priority.