Real-Time Audio Error Correction for Hands-Free Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in providing accurate audio input due to background noise, speech impediments, or foreign languages, which can lead to errors in controlling information processing apparatuses and communicating effectively.

Innovation Solution

An apparatus comprising an acquiring unit to capture audio, an identification unit to detect errors, a selecting unit to choose correction processing, and a generating unit to produce corrected audio content, enabling real-time correction of speech errors and improving accessibility for audio command usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If audio input is used to control information processing apparatus, then accessibility and hands-free operation are improved, but accuracy of audio recognition deteriorates in noisy environments

Engineering Contradiction:
Improvehands-free operationVSAvoidaudio recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary correction mechanism that detects and corrects recognition errors by analyzing audio characteristics (such as pauses, repetitions, or atypical sound patterns) and substituting likely incorrect words with correct ones based on context and language models, thereby maintaining high accuracy without requiring manual input

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by continuously monitoring the audio stream for error indicators (such as unexpected silences, repeated phrases, or contradictory words) and dynamically adjusting the recognition output, allowing real-time correction of misrecognized words while maintaining hands-free operation

Inventive Principle:
Principle #23Feedback

2Productivity

If real-time audio correction is implemented, then communication effectiveness is improved, but processing time and system complexity increase

Engineering Contradiction:
Improvecommunication effectivenessVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The correction process is segmented into independent modules: error detection module (identifying potential errors based on audio patterns), correction decision module (determining whether correction is needed), and correction execution module (applying the correction). This modular approach reduces overall system complexity by allowing each component to operate independently and be optimized separately

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial correction by only correcting portions of the audio stream that contain errors while leaving correct segments unchanged, and by using excessive computation only when error indicators are detected, thereby maintaining real-time performance without requiring full-system analysis at all times

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240203435A1Information processing method, apparatus and computer program
Publication Date: 2024.06.20 SONY INTERTACTIVE ENTERTAINMENT INC
  • US20240203435A1 patent drawing
  • US20240203435A1 patent drawing
  • US20240203435A1 patent drawing

AI summary

An information processing method, of generating corrected audio content in which a portion of first audio content has been corrected, comprises acquiring first audio content from an audio receiving device, identifying a target portion of the first audio content having a predetermined characteristic, selecting correction processing to be performed on the target portion of the first audio content in accordance with the predetermined characteristic of the target portion, and generating corrected audio content in which the target portion of the first audio content has been corrected by performing the selected correction processing on the target portion of the first audio content.