Personalized Speech Recognition via Dynamic Profile Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems use a single, generic profile for multiple users, failing to account for individual speaker characteristics, location, and device-specific factors, leading to suboptimal recognition accuracy.

Innovation Solution

A system that creates and selects speaker-specific profiles based on identified parameters such as speaker identity, location, and microphone type, using hierarchical structures to dynamically adjust speech recognition settings for personalized performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a single generic profile is used for multiple users, then the system complexity is reduced and ease of operation is improved, but speech recognition accuracy deteriorates

Engineering Contradiction:
Improveease of operationVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the speech recognition system into multiple speaker-specific profiles, each tailored to individual acoustic characteristics, vocabularies, and preferences. This segmentation allows the system to maintain simplicity in profile selection (user just needs to identify themselves) while achieving high accuracy through personalized recognition parameters.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically changes recognition parameters based on the selected speaker profile, including acoustic models, vocabulary lists, and grammar rules. This parameter adaptation enables the system to optimize recognition accuracy for each speaker without requiring complex manual configuration, thus maintaining ease of operation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speaker-specific profiles are created and maintained, then speech recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-configuring speaker profiles with acoustic models, vocabularies, and preferences before actual speech recognition occurs. Profile creation and maintenance happen in advance, allowing the recognition engine to simply select and apply the appropriate profile during operation, thereby managing complexity efficiently.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The profile management system serves multiple functions: storing acoustic characteristics, maintaining vocabulary lists, managing user preferences, and selecting appropriate grammar rules. This multi-functionality consolidates what could be separate complex systems into a unified profile structure, reducing overall device complexity while maintaining high recognition accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If multiple user profiles with different settings are maintained, then adaptability to different speakers is improved, but loss of information increases due to managing multiple configurations

Engineering Contradiction:
ImproveadaptabilityVSAvoidinformation loss
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system uses feedback mechanisms to continuously refine and update speaker profiles based on recognition performance and user corrections. This feedback loop ensures that the most current and accurate information about each speaker's characteristics is maintained, preventing information loss while adapting to changing speaker patterns over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9966063B2System and method for personalization in speech recognition
Publication Date: 2018.05.08 AT&T INTELLECTUAL PROPERTY I L P

AI summary

Systems, methods, and computer-readable storage devices are for identifying a user profile for speech recognition. The user profile is selected from one of several user profiles which are all associated with a speaker, and can be selected based on the identity of the speaker, the location of the speaker, the device the speaker is using, or other relevant parameters. Such parameters can be hierarchical, having multiple layers, and can also be dependent or independent from one another. Using the parameters identified, the user profile is selected and used to recognize speech.