Speech-to-text matching unit preserving time info during editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech-to-text systems fail to match edited text with speech data when recognition result text is edited or new text is created without using recognition result text, leading to a loss of time information and reduced operating efficiency.

Innovation Solution

A speech-to-text system that includes a speech recognition unit, a text editor unit, and a matching unit to collate edit result text with recognition result information, using time information and sub-word conversion to ensure accurate matching and synchronization of edit result text with speech data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech recognition is used to automatically transform speech data into text, then writing-out efficiency is improved, but time information is lost when the recognition result text is edited

Engineering Contradiction:
Improvewriting-out efficiencyVSAvoidtime information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary action by storing time information associated with each word in the recognition result text before editing occurs. The edit history management unit preserves the correspondence between edited words and their original time positions, ensuring time information is retained even after text modification.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy of the time information associated with each word in the recognition result text. When editing occurs, the edit history management unit maintains this copied time information in the edit history, allowing the original time correspondence to be preserved and referenced later.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If the recognition result text is edited to correct errors, then text accuracy is improved, but the correspondence with speech data is lost

Engineering Contradiction:
Improvetext accuracyVSAvoidcorrespondence with speech data
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system implements feedback by providing a playback function that allows operators to listen to the original speech data while viewing the edited text. The playback control unit uses the preserved time information from edit history to synchronize speech playback with the edited text, enabling operators to verify accuracy while maintaining correspondence with the original speech data.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If manual editing of recognition results is performed, then text quality is improved, but operating time increases

Engineering Contradiction:
Improvetext qualityVSAvoidoperating time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system creates a copy of the time information associated with each word in the recognition result text. When editing occurs, the edit history management unit maintains this copied time information in the edit history, allowing the original time correspondence to be preserved and referenced later.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system replaces the mechanical process of manually tracking time positions during editing with an automated information processing system. The edit history management unit automatically records and manages the correspondence between edited words and their time positions, eliminating the need for manual time tracking and reducing operating time.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8155958B2Speech-to-text system, speech-to-text method, and speech-to-text program
Publication Date: 2012.04.10 NEC CORP
  • US8155958B2 patent drawing
  • US8155958B2 patent drawing
  • US8155958B2 patent drawing

AI summary

[PROBLEMS] To provide a speech-to-text system and the like capable of matching edit result text acquired by editing recognition result text or edit result text which is newly-written text information with speech data. [MEANS FOR SOLVING PROBLEMS] A speech-to-text system (1) includes a matching unit (27) which collates edit result text acquired by a text editor unit (22) with speech recognition result information having time information created by a speech recognition unit (11) to thereby match the edit result text and speech data.