Voice Model Correction via Continuous Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice transcription software relies on voice-dependent models and has limited continuous learning capabilities, which restricts its accuracy and adaptability to transcribing audio messages from multiple speakers.

Innovation Solution

A system that includes a transcription server, translation server, and training server, allowing for the reception of audio data from various sources, transcription based on a voice-independent model, and correction interface for multiple users to improve transcription accuracy by modifying the voice model using corrected text data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice-dependent models are used for transcription, then accuracy for trained speakers is improved, but adaptability to new speakers deteriorates

Engineering Contradiction:
Improvetranscription accuracyVSAvoidadaptability to new speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the learning process into continuous small updates rather than requiring complete retraining. Each correction provides incremental training data that is processed immediately, allowing the model to adapt to new speakers gradually without losing performance on previously trained speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice model is made dynamic through continuous learning capabilities. The model can be modified in real-time as new correction data becomes available, allowing it to adapt its parameters and structure continuously rather than remaining static between training sessions.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If continuous learning is implemented, then adaptability is improved, but system complexity increases

Engineering Contradiction:
Improvecontinuous learning capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements feedback loops where user corrections are automatically captured and fed back into the training process. This continuous feedback mechanism enables the model to learn from errors and improve automatically without requiring complex manual intervention or retraining procedures.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-training through automatic processing of correction data. When users correct transcription errors, the system automatically uses these corrections to retrain and update the voice model without requiring external training sessions or complex administrative intervention.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If multiple users provide corrections, then learning diversity is improved, but data processing time increases

Engineering Contradiction:
Improvelearning from multiple speakersVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system processes correction data continuously as it is received from multiple users, rather than batching it for later processing. This continuous processing approach ensures that learning opportunities are captured immediately, reducing delays and enabling the model to adapt in real-time to diverse speaker inputs.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11594211B2Methods and systems for correcting transcribed audio files
Publication Date: 2023.02.28 III HOLDINGS 1 LLC
  • US11594211B2 patent drawing
  • US11594211B2 patent drawing
  • US11594211B2 patent drawing

AI summary

Methods and systems for correcting transcribed text. One method includes receiving audio data from one or more audio data sources and transcribing the audio data based on a voice model to generate text data. The method also includes making the text data available to a plurality of users over at least one computer network and receiving corrected text data over the at least one computer network from the plurality of users. In addition, the method can include modifying the voice model based on the corrected text data.