Acoustic Model Generation via Voice Authentication Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Developing speech recognition systems for new linguistic markets is time-consuming and costly due to the need for large datasets of language and accent-specific speech samples, which are often unavailable.

Innovation Solution

A method for automatically configuring a speech recognition system by processing speech samples obtained during voice authentication to generate and update acoustic models, allowing for the creation of language and accent-specific models without the need for extensive data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems are developed for new linguistic markets using traditional methods, then recognition accuracy can be achieved, but the process becomes time-consuming and costly due to the need for large datasets

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevelopment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses speech samples naturally provided by users during voice authentication processes to automatically build and update acoustic models. This self-service approach eliminates the need for separate, time-consuming data collection campaigns, as the system configures itself using readily available user speech data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The speech samples collected for voice authentication serve dual purposes: both security verification and speech recognition model training. This multi-functionality allows the same data to address multiple system requirements simultaneously, reducing overall development time and resource needs

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If speech recognition systems are developed for new linguistic markets using traditional methods, then recognition accuracy can be achieved, but the cost increases due to extensive data collection requirements

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevelopment cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system automatically configures its own acoustic models using speech data from authentication processes, eliminating the need for expensive external data collection services and manual processing, thereby significantly reducing development costs

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

By making the speech samples serve dual purposes (authentication and model training), the system eliminates redundant data collection activities, reducing overall development costs while maintaining recognition accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If language and accent-specific acoustic models are created for each linguistic market, then speech recognition accuracy improves, but the complexity of system configuration increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system dynamically adapts its acoustic models based on the speech samples it receives during authentication. Rather than requiring static pre-configuration for each linguistic market, the system evolves its models automatically, reducing configuration complexity while maintaining accuracy

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system continuously updates acoustic models using feedback from authentication speech samples. This automated feedback loop allows the system to adapt to different linguistic markets without manual reconfiguration, simplifying the deployment process while preserving recognition accuracy

Inventive Principle:
Principle #23Feedback

4Reliability

If extensive speech data collection is performed for each linguistic market, then acoustic model quality improves, but the productivity of system deployment decreases

Engineering Contradiction:
Improveacoustic model qualityVSAvoidsystem deployment speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs automatic model configuration using authentication data, eliminating the need for slow, manual data collection and processing steps. This self-service approach maintains model quality while dramatically accelerating deployment speed

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system prepares acoustic models in advance using speech samples collected during normal authentication operations. By continuously pre-processing available data, the system ensures high-quality models are ready when needed, improving deployment speed without sacrificing model quality

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9424837B2Voice authentication and speech recognition system and method
Publication Date: 2016.08.23 AURAYA
  • US9424837B2 patent drawing
  • US9424837B2 patent drawing
  • US9424837B2 patent drawing

AI summary

A method for configuring a speech recognition system comprises obtaining a speech sample utilized by a voice authentication system in a voice authentication process. The speech sample is processed to generate acoustic models for units of speech associated with the speech sample. The acoustic models are stored for subsequent use by the speech recognition system as part of a speech recognition process.