Corpus Augmenting System for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting and recording large corpora of speech from individuals with articulation disorders is time-consuming and burdensome, affecting the accuracy of voice conversion and automatic speech recognition models.
Innovation Solution
A method and system that utilize a conversion model to transform normal speech feature data into augmented corpora simulating the speech patterns of individuals with articulation disorders, combining target and training speech feature data to enhance model training and recognition capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech corpora are collected by recording from individuals with articulation disorders, then the accuracy of voice conversion and speech recognition models is improved, but the time consumption and physical burden increase significantly
Solution Approach 1:
The patent applies copying by using a conversion model to generate synthetic speech samples that replicate the characteristics of articulation disorders. Instead of recording actual speech from individuals, the system copies the disordered speech patterns through AI-generated synthetic data, maintaining model training effectiveness while eliminating the time-consuming recording process
Solution Approach 2:
The patent implements preliminary action by pre-training a conversion model on existing articulation disorder speech data. This pre-trained model can then generate synthetic speech samples on-demand without requiring new recordings, allowing the system to prepare training corpora in advance and reduce real-time recording requirements
2Measurement precision
If speech corpora are collected by recording from individuals with articulation disorders, then the accuracy of voice conversion and speech recognition models is improved, but the physical and emotional burden on speakers increases
Solution Approach 1:
The system copies articulation disorder speech characteristics through AI synthesis rather than requiring actual speakers to produce disordered speech. This eliminates the physical strain and emotional distress associated with prolonged reading sessions while maintaining the authenticity of disordered speech patterns in the training data
Solution Approach 2:
The conversion model serves as an intermediary between normal speech data and the required articulation disorder speech training data. This intermediary transforms standard speech into synthetic disordered speech, removing the need for actual individuals with articulation disorders to participate in the recording process and thereby eliminating their physical and emotional burden
3Measurement precision
If large corpora of speech from individuals with articulation disorders are collected, then the accuracy of speech recognition models is improved, but the complexity and difficulty of data acquisition increases
Solution Approach 1:
The patent uses copying to generate synthetic speech corpora that replicate articulation disorder characteristics without requiring actual collection from speakers. The conversion model produces large volumes of synthetic training data that copy the essential features of disordered speech, eliminating the logistical challenges of recruiting, recording, and managing real speaker sessions
Solution Approach 2:
The patent replaces the mechanical process of physical recording sessions with an automated computational system. The conversion model automatically generates synthetic speech samples through algorithmic processing, substituting the manual, labor-intensive recording process with an automated digital system that can produce large corpora without human intervention
Data Source
AI summary
A method of forming augmented corpus related to articulation disorder includes acquiring target speech feature data from a target corpus; acquiring training speech feature data from training corpora; training a conversion model to make it capable of converting training speech feature data into a respective output that is similar to the target speech feature data; receiving an augmenting source corpus and acquiring augmenting source speech feature data therefrom; converting, by the conversion model thus trained, the augmenting source speech feature data into converted speech feature data; and synthesizing the augmented corpus based on the converted speech feature data.


