Speech Recognition Text Concatenation for Overlapping Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition techniques, such as those described in Patent Document 1, may inaccurately recognize words when a correct word is detected from only one of the overlapping sections, leading to errors in the recognition result.

Innovation Solution

A speech recognition apparatus and method that converts a source audio signal into a text string and generates a concatenated text by overlapping and adjusting adjacent text segments, specifically by eliminating trailing portions of preceding texts and leading portions of succeeding texts during concatenation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition is performed on overlapping audio sections and words are selected from both sections, then coverage of recognized words is improved, but recognition accuracy deteriorates when correct words are detected from only one section

Engineering Contradiction:
Improvecoverage of recognized wordsVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing speech recognition on overlapping audio sections in advance, generating candidate words from each section before final selection. This allows the system to prepare multiple potential recognitions and then apply selection criteria to determine the most accurate result, resolving the contradiction between coverage and accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback by comparing recognition results from overlapping sections and using selection criteria to determine the final recognized word. The system feeds back the comparison results and adjusts the selection based on which section provides more reliable recognition, thereby maintaining accuracy while utilizing multiple sections for coverage.

Inventive Principle:
Principle #23Feedback

2Productivity

If audio signals are divided into multiple sections for speech recognition, then processing efficiency is improved, but errors occur when correct words are detected from only one section

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidrecognition reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies segmentation by dividing the audio signal into multiple overlapping sections for parallel speech recognition processing. This segmentation improves processing efficiency by allowing simultaneous analysis of different sections while the overlap ensures that correct words detected from only one section are not lost, thereby maintaining reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements beforehand cushioning by creating overlapping sections where the overlap acts as a buffer or cushion. This ensures that if a correct word is detected from only one section, the overlapping region provides additional coverage and validation, preventing recognition errors while maintaining processing efficiency through parallel section analysis.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Measurement precision

If trailing portions of preceding texts and leading portions of succeeding texts are eliminated during concatenation, then accuracy of concatenated text is improved, but information loss occurs in the eliminated portions

Engineering Contradiction:
Improveaccuracy of concatenated textVSAvoidinformation loss in eliminated portions
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies taking out by extracting and eliminating the trailing portion of the preceding text and the leading portion of the succeeding text during concatenation. This extraction removes redundant or potentially erroneous information from the overlap regions, improving the accuracy of the final concatenated text while minimizing information loss by only removing necessary portions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements partial action by eliminating only the necessary trailing and leading portions rather than entire sections. This partial elimination is sufficient to improve concatenation accuracy by removing redundant information, while avoiding excessive action that would cause unnecessary information loss. The elimination is precisely controlled to maintain the balance between accuracy and information preservation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12340807B2Speech recognition apparatus, control method, and non-transitory storage medium
Publication Date: 2025.06.24 NEC CORP
  • US12340807B2 patent drawing
  • US12340807B2 patent drawing
  • US12340807B2 patent drawing

AI summary

A speech recognition apparatus (2000) acquires source data (10) representing an audio signal including an utterance. The speech recognition apparatus (2000) converts the source data (10) into a text string (30). The speech recognition apparatus (2000) generates a concatenated text (40) representing a content of an utterance by concatenating a text (32) included in the text string (30). Herein, texts (32) adjacent to each other in the text string (30) are such that parts of associated audio signals overlap each other on a time axis. At a time of concatenating texts (32) adjacent to each other, the speech recognition apparatus (2000) eliminates a trailing portion of a preceding text (32) and a leading portion of a succeeding text (32).