Corpus Augmenting System for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Collecting and recording large corpora of speech from individuals with articulation disorders is time-consuming and burdensome, affecting the accuracy of voice conversion and automatic speech recognition models.

Innovation Solution

A method and system that utilize a conversion model to transform normal speech feature data into augmented corpora simulating the speech patterns of individuals with articulation disorders, combining target and training speech feature data to enhance model training and recognition capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech corpora are collected by recording from individuals with articulation disorders, then the accuracy of voice conversion and speech recognition models is improved, but the time consumption and physical burden increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidrecording time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies copying by using a conversion model to generate synthetic speech samples that replicate the characteristics of articulation disorders. Instead of recording actual speech from individuals, the system copies the disordered speech patterns through AI-generated synthetic data, maintaining model training effectiveness while eliminating the time-consuming recording process

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements preliminary action by pre-training a conversion model on existing articulation disorder speech data. This pre-trained model can then generate synthetic speech samples on-demand without requiring new recordings, allowing the system to prepare training corpora in advance and reduce real-time recording requirements

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If speech corpora are collected by recording from individuals with articulation disorders, then the accuracy of voice conversion and speech recognition models is improved, but the physical and emotional burden on speakers increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidphysical and emotional burden
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system copies articulation disorder speech characteristics through AI synthesis rather than requiring actual speakers to produce disordered speech. This eliminates the physical strain and emotional distress associated with prolonged reading sessions while maintaining the authenticity of disordered speech patterns in the training data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The conversion model serves as an intermediary between normal speech data and the required articulation disorder speech training data. This intermediary transforms standard speech into synthetic disordered speech, removing the need for actual individuals with articulation disorders to participate in the recording process and thereby eliminating their physical and emotional burden

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If large corpora of speech from individuals with articulation disorders are collected, then the accuracy of speech recognition models is improved, but the complexity and difficulty of data acquisition increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata acquisition difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent uses copying to generate synthetic speech corpora that replicate articulation disorder characteristics without requiring actual collection from speakers. The conversion model produces large volumes of synthetic training data that copy the essential features of disordered speech, eliminating the logistical challenges of recruiting, recording, and managing real speaker sessions

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of physical recording sessions with an automated computational system. The conversion model automatically generates synthetic speech samples through algorithmic processing, substituting the manual, labor-intensive recording process with an automated digital system that can produce large corpora without human intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12159624B2Method of forming augmented corpus related to articulation disorder, corpus augmenting system, speech recognition platform, and assisting device
Publication Date: 2024.12.03 APREVENT MEDICAL INC(CN)
  • US12159624B2 patent drawing
  • US12159624B2 patent drawing
  • US12159624B2 patent drawing

AI summary

A method of forming augmented corpus related to articulation disorder includes acquiring target speech feature data from a target corpus; acquiring training speech feature data from training corpora; training a conversion model to make it capable of converting training speech feature data into a respective output that is similar to the target speech feature data; receiving an augmenting source corpus and acquiring augmenting source speech feature data therefrom; converting, by the conversion model thus trained, the augmenting source speech feature data into converted speech feature data; and synthesizing the augmented corpus based on the converted speech feature data.