Multi-Speaker Voice Recognition via Lip-Reading in Noisy Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In noisy environments like automobiles, existing voice recognition systems struggle to accurately distinguish between multiple speakers, leading to confused responses and misinterpretation of intentions.

Innovation Solution

An information processing device that acquires voices and videos of multiple speakers, uses lip-reading technology to specify each speaker, and generates responses based on recognized attributes and properties of the utterances, prioritizing responses to determine optimal interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice recognition is performed in a noisy environment with multiple speakers, then the system can process audio inputs, but the recognition accuracy deteriorates due to voice confusion and inability to distinguish intentions

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidnoise and voice confusion
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system segments the audio signal by identifying and separating individual speaker voices from the mixed audio stream. By detecting voice attributes and tracking speakers over time, the system divides the complex audio input into distinct speaker segments, enabling accurate recognition of each speaker's intentions without confusion from other speakers or background noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer that analyzes voice attributes, speaker identification, and utterance context before final recognition. This intermediary layer acts as a mediator between the raw audio signal and the recognition output, filtering out noise and disambiguating speaker intentions to improve overall recognition accuracy in noisy environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system processes multiple utterances from different speakers, then the system can provide comprehensive responses, but the complexity of determining speaker priority and appropriate responses increases

Engineering Contradiction:
Improvemulti-speaker response capabilityVSAvoidspeaker priority determination logic
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically determines speaker priority and response generation based on real-time analysis of voice attributes, utterance content, and interaction context. Rather than using fixed priority rules, the system adapts its response generation strategy by continuously monitoring the speaking situation, allowing flexible and appropriate responses to multiple speakers without requiring complex predetermined logic for every possible scenario.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250006200A1Information processing device, information processing method, and information processing program
Publication Date: 2025.01.02 SONY GROUP CORP
  • US20250006200A1 patent drawing
  • US20250006200A1 patent drawing
  • US20250006200A1 patent drawing

AI summary

An information processing device (100) according to the present disclosure includes an acquisition unit (131) that acquires voices generated by a plurality of utterers and a video in which a state where the utterer generates an utterance is imaged, a specification unit (132) that specifies each of the plurality of utterers, on the basis of the acquired voice and video, a recognition unit (133) that recognizes an utterance generated by each specified utterer and an attribute of each utterer or a property of the utterance, and a generation unit (134) that generates a response to the recognized utterance, on the basis of the recognized attribute of each utterer or property of the utterance.