Speaker-Adaptation Speech Recognition Terminal Reducing Server Load

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech-recognition systems face data overload and inefficiency due to the need for pre-learning processes and speaker ID management, which are cumbersome and require significant server resources.

Innovation Solution

A speech-recognition system where statistical variables for speaker-adaptation are extracted and accumulated on a terminal, allowing the terminal to generate conversion parameters for the speech-recognition server, reducing the need for pre-learning and speaker ID management, and thereby minimizing data transmission and server burden.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pre-learning process and speaker ID management are implemented, then speaker-adaptation performance is improved, but server resource consumption and data overload increase

Engineering Contradiction:
Improvespeaker-adaptation performanceVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts the speaker-adaptation functionality from the server side and relocates it to the terminal side. The terminal now performs statistical variable accumulation and conversion parameter generation locally, while the server only needs to provide basic speech recognition services. This extraction eliminates the need for server-side speaker ID management and pre-learning processes, thereby reducing data overload on the server while maintaining speaker-adaptation performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The terminal is empowered to perform speaker-adaptation independently through self-service mechanisms. It automatically accumulates statistical variables from recognized speech, generates conversion parameters, and applies them to improve recognition accuracy for the current speaker. This self-service capability eliminates the need for server-side pre-learning processes and speaker ID management, reducing both data volume and server resource consumption.

Inventive Principle:
Principle #25Self-service

2Reliability

If pre-learning process is performed, then speaker-adaptation accuracy is improved, but operation complexity and time consumption increase

Engineering Contradiction:
Improvespeaker-adaptation accuracyVSAvoidoperation complexity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs preliminary accumulation of statistical variables continuously in the background during normal speech recognition operations. Instead of requiring a separate pre-learning phase, the terminal gradually builds up speaker-specific statistical data from everyday speech interactions. This preliminary action occurs automatically without user intervention, maintaining high speaker-adaptation accuracy while eliminating operation complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The terminal automatically performs speaker-adaptation without requiring user-initiated pre-learning processes. It self-manages the accumulation of statistical variables, generation of conversion parameters, and application of speaker-specific adaptations throughout normal operation. This self-service approach maintains high accuracy while completely eliminating the need for complex manual pre-learning operations.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If speaker ID management is implemented, then speaker identification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent removes the speaker ID management function from the system architecture. Instead of implementing speaker identification and management mechanisms, the system directly performs speaker-adaptation using speech content analysis. The terminal extracts speaker-specific statistical variables from the speech itself without requiring separate speaker identification processes, thereby maintaining adaptation accuracy while reducing system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces statistical variables as an intermediary mechanism between speech input and speaker-adaptation output. Rather than directly implementing speaker ID management, the terminal uses accumulated statistical variables derived from speech content to enable speaker-specific adaptation. This intermediary approach maintains measurement precision for speaker characteristics while avoiding the complexity of direct speaker ID management systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9530403B2Terminal and server of speaker-adaptation speech-recognition system and method for operating the system
Publication Date: 2016.12.27 ELECTRONICS & TELECOMM RES INST
  • US9530403B2 patent drawing
  • US9530403B2 patent drawing
  • US9530403B2 patent drawing

AI summary

Provided are a terminal and server of a speaker-adaptation speech-recognition system and a method for operating the system. The terminal in the speaker-adaptation speech-recognition system includes a speech recorder which transmits speech data of a speaker to a speech-recognition server, a statistical variable accumulator which receives a statistical variable including acoustic statistical information about speech of the speaker from the speech-recognition server which recognizes the transmitted speech data, and accumulates the received statistical variable, a conversion parameter generator which generates a conversion parameter about the speech of the speaker using the accumulated statistical variable and transmits the generated conversion parameter to the speech-recognition server, and a result displaying user interface which receives and displays result data when the speech-recognition server recognizes the speech data of the speaker using the transmitted conversion parameter and transmits the recognized result data.