Parallel Speech Recognition with Temporal Ordering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In speech recognition systems for multiple speakers, existing technologies face challenges in providing real-time output of recognized speech due to variations in processing times across different speech recognition engines, leading to potential confusion when results are displayed out of order.

Innovation Solution

A system that includes multiple speech recognition units, each processing speech from a specific speaker, with a display processing unit that associates and displays speaker information and text data in the order of speech emission, ensuring correct temporal alignment of recognition results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple speech recognition engines process speeches in parallel, then the speech recognition results can be output quickly, but the results may be displayed in the wrong order causing user confusion

Engineering Contradiction:
Improvespeech recognition processing speedVSAvoidtemporal order of speech emission
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary actions by assigning unique identification information to each speech signal before processing, and pre-establishes the correspondence between identification information and speaker information. This allows the display processing unit to correctly order the results after parallel processing without losing temporal sequence information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces identification information as an intermediary element that bridges the speech signal and the speaker information. This intermediary carries the temporal ordering information through the parallel processing engines, enabling the display unit to reconstruct the correct sequence of speech emissions from multiple concurrent processing streams.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If a single speech recognition engine processes all speeches, then the processing is simpler to manage, but the output time is delayed and real-time recognition is difficult

Engineering Contradiction:
Improvespeech recognition system structureVSAvoidspeech recognition output time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent divides the speech recognition system into multiple independent speech recognition engines, each capable of processing speech signals in parallel. This segmentation increases processing capacity and reduces output time while the identification information mechanism maintains manageable complexity by providing clear tracking of each processed speech segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of processing capacity by deploying multiple speech recognition engines simultaneously, transforming a single-threaded sequential processing model into a multi-threaded parallel processing model. This parameter change directly addresses the time delay issue while the identification information system manages the increased complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8606574B2Speech recognition processing system and speech recognition processing method
Publication Date: 2013.12.10 IP WAVE PTE LTD
  • US8606574B2 patent drawing
  • US8606574B2 patent drawing
  • US8606574B2 patent drawing

AI summary

The present invention provides a speech recognition processing system in which speech recognition processing is executed parallelly by plural speech recognizing units. Before text data as the speech recognition result is output from each of the speech recognizing units, information indicating each speaker is parallelly displayed on a display in emission order of each speech. When the text data is output from each of the speech recognizing units, the text data is associated with the information indicating each speaker and the text data is displayed.