Speech Recognition Text Concatenation for Overlapping Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition techniques, such as those described in Patent Document 1, may inaccurately recognize words when a correct word is detected from only one of the overlapping sections, leading to errors in the recognition result.
Innovation Solution
A speech recognition apparatus and method that converts a source audio signal into a text string and generates a concatenated text by overlapping and adjusting adjacent text segments, specifically by eliminating trailing portions of preceding texts and leading portions of succeeding texts during concatenation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed on overlapping audio sections and words are selected from both sections, then coverage of recognized words is improved, but recognition accuracy deteriorates when correct words are detected from only one section
Solution Approach 1:
The patent applies preliminary action by performing speech recognition on overlapping audio sections in advance, generating candidate words from each section before final selection. This allows the system to prepare multiple potential recognitions and then apply selection criteria to determine the most accurate result, resolving the contradiction between coverage and accuracy.
Solution Approach 2:
The patent implements feedback by comparing recognition results from overlapping sections and using selection criteria to determine the final recognized word. The system feeds back the comparison results and adjusts the selection based on which section provides more reliable recognition, thereby maintaining accuracy while utilizing multiple sections for coverage.
2Productivity
If audio signals are divided into multiple sections for speech recognition, then processing efficiency is improved, but errors occur when correct words are detected from only one section
Solution Approach 1:
The patent applies segmentation by dividing the audio signal into multiple overlapping sections for parallel speech recognition processing. This segmentation improves processing efficiency by allowing simultaneous analysis of different sections while the overlap ensures that correct words detected from only one section are not lost, thereby maintaining reliability.
Solution Approach 2:
The patent implements beforehand cushioning by creating overlapping sections where the overlap acts as a buffer or cushion. This ensures that if a correct word is detected from only one section, the overlapping region provides additional coverage and validation, preventing recognition errors while maintaining processing efficiency through parallel section analysis.
3Measurement precision
If trailing portions of preceding texts and leading portions of succeeding texts are eliminated during concatenation, then accuracy of concatenated text is improved, but information loss occurs in the eliminated portions
Solution Approach 1:
The patent applies taking out by extracting and eliminating the trailing portion of the preceding text and the leading portion of the succeeding text during concatenation. This extraction removes redundant or potentially erroneous information from the overlap regions, improving the accuracy of the final concatenated text while minimizing information loss by only removing necessary portions.
Solution Approach 2:
The patent implements partial action by eliminating only the necessary trailing and leading portions rather than entire sections. This partial elimination is sufficient to improve concatenation accuracy by removing redundant information, while avoiding excessive action that would cause unnecessary information loss. The elimination is precisely controlled to maintain the balance between accuracy and information preservation.
Data Source
AI summary
A speech recognition apparatus (2000) acquires source data (10) representing an audio signal including an utterance. The speech recognition apparatus (2000) converts the source data (10) into a text string (30). The speech recognition apparatus (2000) generates a concatenated text (40) representing a content of an utterance by concatenating a text (32) included in the text string (30). Herein, texts (32) adjacent to each other in the text string (30) are such that parts of associated audio signals overlap each other on a time axis. At a time of concatenating texts (32) adjacent to each other, the speech recognition apparatus (2000) eliminates a trailing portion of a preceding text (32) and a leading portion of a succeeding text (32).


