Speech Recognition Audio Input Editing Modes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies face challenges with low accuracy and limited richness in audio input processing, leading to a poor user experience due to immature models and frequent errors in recognition results.

Innovation Solution

An audio input method and terminal device that incorporate two modes: an audio-input mode for initial recognition and an editing mode for processing corrections and additions, allowing users to switch between modes to enhance accuracy and richness of input content, with features like command execution, error correction, and content element addition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition technology is used to convert audio information to text, then audio input functionality is provided, but recognition accuracy is low and user experience is poor

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser experience
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the audio input process into two distinct modes: audio-input mode for initial recognition and editing mode for correction. This segmentation allows different processing strategies to be applied at different stages, improving overall recognition accuracy while maintaining user experience through targeted error correction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a feedback mechanism where recognition results are displayed to users, who can then provide corrections in editing mode. This feedback loop enables continuous improvement of recognition accuracy by learning from user corrections, thereby enhancing both measurement precision and reliability.

Inventive Principle:
Principle #23Feedback

2Productivity

If simple text display is used for audio recognition output, then processing speed is fast, but content richness is limited

Engineering Contradiction:
Improveprocessing speedVSAvoidcontent richness
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent makes the audio input system multi-functional by enabling it to handle not only text display but also editing operations, command execution, and content enrichment. The same audio input interface serves multiple purposes, allowing fast processing while expanding content richness through integrated editing and command capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces dynamic mode switching between audio-input mode and editing mode. This dynamic adaptability allows the system to adjust its functionality based on user needs, transitioning from simple text display to comprehensive content editing, thereby maintaining processing speed while enhancing content richness.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If manual selection of content for correction is required, then editing precision is high, but operation complexity increases

Engineering Contradiction:
Improveediting precisionVSAvoidoperation complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent enables self-service editing where the system automatically presents recognition results that users can directly review and correct. This approach maintains high editing precision by allowing user oversight while reducing operation complexity by eliminating the need for manual content selection - users simply review and correct what is already presented to them.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10923118B2Speech recognition based audio input and editing method and terminal device
Publication Date: 2021.02.16 BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
  • US10923118B2 patent drawing
  • US10923118B2 patent drawing
  • US10923118B2 patent drawing

AI summary

An audio input method includes: in an audio-input mode, receiving a first audio input by a user, recognizing the first audio to generate a first recognition result, and displaying corresponding verbal content to the user based on the first recognition result; and in an editing mode, receiving a second audio input by the user and recognizing and generating a second recognition result, converting the second recognition result to an editing instruction, and executing a corresponding operation based on the editing operation. The audio-input mode and the editing mode are switchable.