Voice Extraction Unit Noise Removal via Sound Source Direction Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice interactive systems face challenges in distinguishing user speech from noise generated by external devices, leading to erroneous processing due to increased processing load and time, as existing methods require complex audio recognition and subtraction of common data from multiple signals.

Innovation Solution

An information processing device and method that analyze sound source directions and frequency characteristics of external apparatus output sounds, registering these in a database to remove noise from user speech inputs, thereby isolating clear user speech for accurate recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If complex audio recognition processing is performed on multiple audio signals to distinguish user speech from noise, then the accuracy of speech recognition is improved, but the processing load and processing time increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary audio recognition processing on a reference audio signal (from the reference microphone) to identify and extract noise components before the main speech recognition process. This pre-processing step creates a noise profile that is then used to filter the target audio signal, reducing the complexity of the main recognition process while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention extracts noise components from the audio signal by comparing the target audio (from target microphone) with reference audio (from reference microphone). The noise portion is identified and separated from the target speech signal, allowing clean speech recognition without complex processing of the entire mixed signal.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If multiple audio signals are processed individually to remove noise, then the quality of user speech extraction is improved, but the processing cost increases

Engineering Contradiction:
Improvespeech extraction qualityVSAvoidprocessing cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The system combines multiple audio signals (target audio from target microphone and reference audio from reference microphone) into a single processed output. By merging these signals and performing unified noise cancellation processing, the system achieves high-quality speech extraction without the need for separate independent processing of each signal, thereby reducing processing cost.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If the system processes all input sounds as user speech, then no speech is lost, but erroneous processing increases

Engineering Contradiction:
Improvespeech detection reliabilityVSAvoiderroneous processing
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The reference microphone acts as an intermediary that captures ambient noise and device output sounds separately from the target speech. This reference signal serves as a mediator to identify and remove noise components from the target audio, allowing the system to reliably distinguish between actual user speech and noise without erroneous processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12119017B2Information processing device, information processing system and information processing method
Publication Date: 2024.10.15 SONY GROUP CORP
  • US12119017B2 patent drawing
  • US12119017B2 patent drawing
  • US12119017B2 patent drawing

AI summary

Provided is a device that includes a user spoken voice extraction unit that extracts a user spoken voice from a microphone input sound. The user spoken voice extraction unit analyzes a sound source direction of an input sound, determines whether the input sound includes an external apparatus output sound on the basis of sound source directions of external apparatus output sounds recorded in a database, and removes a sound signal corresponding to a feature amount, for example, a frequency characteristic of the external apparatus output sound recorded in the database, from the input sound to extract a user spoken voice from which the external apparatus output sound has been removed upon determining that the input sound includes the external apparatus output sound.