Speech Recognition Model Construction for Simultaneous Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition techniques struggle with accurately discriminating simultaneous speech segments in monaural sound recordings, leading to inferior accuracy and high manual labor costs for data labeling, as they fail to effectively handle overlapping speech in mixed speaker environments.
Innovation Solution
A method for constructing a speech recognition model that aligns and synthesizes speech from multiple speakers, replacing overlapping segments with a unit representing simultaneous speech, and building acoustic and language models based on this aligned data to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional speech recognition techniques are used on monaural sound, then data volume is reduced, but recognition accuracy deteriorates due to inability to discriminate simultaneous speech segments
Solution Approach 1:
The patent segments the monaural sound into multiple speaker channels using speaker diarization technology. This allows the system to separate mixed speech from different speakers, enabling accurate identification of simultaneous speech segments while maintaining the space-efficient monaural recording format.
Solution Approach 2:
The patent introduces speaker diarization as an intermediary process between monaural recording and speech recognition. This intermediary technology analyzes the monaural sound to identify which speaker is speaking at each time point, thereby enabling the system to handle simultaneous speech segments without requiring stereo recording.
2Reliability
If manual transcription of monaural sound is performed, then usable transcript data can be obtained, but manual labor cost increases significantly
Solution Approach 1:
The patent enables the system to automatically generate aligned transcript data by itself using speaker diarization. The system processes monaural sound recordings to automatically identify speaker turns and generate corresponding transcripts, eliminating the need for manual transcription while maintaining high accuracy through the use of alignment information between speech segments and transcripts.
3Speed
If speech recognition is performed on overlapping speech portions, then continuous recognition can be maintained, but burst errors occur affecting preceding and following portions
Solution Approach 1:
The patent performs preliminary speaker identification and segmentation before speech recognition. By using speaker diarization to pre-process the monaural sound and identify which speaker is speaking at each moment, the system can prepare alignment information in advance. This allows the speech recognition process to handle simultaneous speech segments appropriately without causing burst errors in preceding or following portions.
Data Source
AI summary
A construction method for a speech recognition model, in which a computer system includes; a step of acquiring alignment between speech of each of a plurality of speakers and a transcript of the speaker; a step of joining transcripts of the respective ones of the plurality of speakers along a time axis, creating a transcript of speech of mixed speakers obtained from synthesized speech of the speakers, and replacing predetermined transcribed portions of the plurality of speakers overlapping on the time axis with a unit which represents a simultaneous speech segment; and a step of constructing at least one of an acoustic model and a language model which make up a speech recognition model, based on the transcript of the speech of the mixed speakers.


