Speech Recognition Text Editing Interface for Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies require users to re-input entire voice content when errors occur, leading to low efficiency and poor user experience, especially when only a small number of letters or words are incorrect, and often fail to produce accurate results even after multiple attempts.

Innovation Solution

An information input method and device that allows users to manually correct speech recognition results in a text editing format, enabling users to revise recognition errors without re-inputting the entire voice content, and feeds the edited results back to the server for improved model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users re-input entire voice content when recognition errors occur, then recognition accuracy can be improved, but voice input efficiency deteriorates significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidvoice input efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the voice input process into two distinct modes: voice input mode for initial recognition and text editing mode for error correction. This segmentation allows users to switch between automatic recognition and manual editing based on the recognition quality, thereby maintaining high input efficiency while achieving accurate results when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic input mechanism that allows flexible switching between voice input and text editing modes. The system adapts to user needs by enabling transition from automatic recognition to manual correction, optimizing both efficiency and accuracy based on the specific situation.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If users re-input entire voice content when recognition errors occur, then correct recognition can be achieved, but user experience deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser experience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent divides the input process into distinct phases: automatic voice recognition phase and manual text editing phase. This segmentation allows users to correct only the erroneous portions rather than re-inputting everything, significantly improving ease of operation and user experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables users to self-correct recognition errors through text editing functionality. Users can independently identify and correct mistakes without requiring complete re-input, making the system more user-friendly and improving overall ease of operation.

Inventive Principle:
Principle #25Self-service

3Device complexity

If speech recognition technology is not significantly improved, then current system complexity is maintained, but recognition accuracy deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces a text editing interface as an intermediary between voice input and final result. This intermediary layer allows users to manually correct errors that the recognition system cannot resolve, improving accuracy without requiring complex improvements to the speech recognition technology itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a dynamic system that combines automatic recognition with manual editing capabilities. This dynamic approach allows the system to maintain current technological complexity while achieving better accuracy through user intervention when needed.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10796699B2Method, apparatus, and computing device for revision of speech recognition results
Publication Date: 2020.10.06 ALIBABA GROUP HOLDING LTD
  • US10796699B2 patent drawing
  • US10796699B2 patent drawing

AI summary

The present disclosure discloses an information input method and device, and a computing apparatus. The information input method comprises receiving a voice input of a user, acquiring a recognition result on the received voice input, and enabling editing of the acquired recognition result in a text format. With the information input mechanism, according to the present invention, a user is able to choose to revise an automatic speech recognition result in a text editing format, particularly in the case where a small amount of errors occurs to the contents of speech recognition. As a result, the trouble that all contents of a voice input need to be input again is avoided, the speech recognition efficiency is increased, and the user experience is improved.