Neural Network Main Speech Detection for Wearable Microphones

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in accurately recognizing and transcribing speech from multiple speakers using wearable distributed microphones, particularly in environments where speeches overlap, leading to inaccurate transcription and difficulty in distinguishing main speech from crosstalk.

Innovation Solution

A signal processing system employing a neural network-based main speech detection unit to identify and isolate main speech from crosstalk, using multi-label classification to differentiate between speakers and handle overlapping speeches, and integrating crosstalk reduction to enhance voice recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If wearable distributed microphones are used to collect speech from multiple speakers, then speech recognition can be performed for each speaker, but crosstalk between speakers causes inaccurate transcription and difficulty in distinguishing main speech from crosstalk

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcrosstalk interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the speech signal detection task by assigning a dedicated main speech detection unit to each speaker's microphone. Each unit independently processes its assigned signal through neural network-based voice activity detection, separating the detection of main speech from crosstalk interference. This segmentation allows accurate identification of which speaker is speaking at each time frame, resolving the crosstalk problem while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If neural network-based main speech detection is implemented, then accurate identification of main speech can be achieved, but computational complexity and processing time increase

Engineering Contradiction:
Improvemain speech identification accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the computational task into separate main speech detection units, each handling one speaker's signal independently. This segmentation allows the system to process multiple speakers simultaneously without requiring a single complex unified model, reducing overall computational burden while maintaining high accuracy through specialized neural networks for each speaker.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs voice activity detection using neural networks to identify main speech segments before subsequent speech recognition processing. This preliminary action filters out crosstalk and non-speech portions, reducing the amount of data that requires full recognition processing and thereby decreasing overall computational complexity while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If frame information indicating presence or absence of main speech is output, then accurate transcription can be generated, but additional processing steps and time are required

Engineering Contradiction:
Improvetranscription generation speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The main speech detection unit performs voice activity detection and generates frame information indicating presence or absence of main speech as an intermediate output. This preliminary classification of speech frames enables subsequent speech recognition units to process only relevant segments, reducing overall processing time and increasing transcription generation speed despite the additional detection step.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges the main speech detection function with the voice activity detection function, combining multiple objectives into a single integrated unit. This merging eliminates redundant processing steps and allows the system to generate frame information efficiently, reducing the time penalty associated with additional processing while maintaining high productivity in transcription generation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12148432B2Signal processing device, signal processing method, and signal processing system
Publication Date: 2024.11.19 SONY GROUP CORP
  • US12148432B2 patent drawing
  • US12148432B2 patent drawing
  • US12148432B2 patent drawing

AI summary

Provided is a signal processing device including a main speech detection unit that detects, by using a neural network, whether or not a signal input to a sound collection device assigned to each of at least two speakers includes a main speech that is a voice of the corresponding speaker, and outputs frame information indicating presence or absence of the main speech.