Multi-Speaker Voice Signal Processing with Conflict Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition technologies face challenges in accurately processing and distinguishing speech signals from multiple speakers simultaneously without conflicts, leading to inefficiencies in recognizing speech intentions and performing operations on electronic devices.

Innovation Solution

A method involving a machine learning module or language understanding module to determine relations between speech content from multiple speakers, adjusting or mediating the content to prevent conflicts, and prioritizing or combining speech intentions for appropriate device operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing speech recognition technologies process speech signals from multiple speakers, then the system can recognize speech intentions, but conflicts arise between different speakers causing reduced recognition accuracy

Engineering Contradiction:
Improvemulti-speaker speech recognition capabilityVSAvoidspeech intention recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the mixed speech signal into individual speaker components by detecting speaker-specific features (pitch, tone, rhythm) and separating them into distinct speech intent groups. This segmentation allows the system to process each speaker's intent independently, resolving conflicts and improving overall recognition accuracy in multi-speaker environments.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If the system processes speech signals without conflict resolution, then processing is simpler, but recognition rate decreases due to overlapping speech content

Engineering Contradiction:
Improvespeech processing complexityVSAvoidspeech recognition reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces an intermediary processing layer that analyzes the relationships between different speech contents before executing operations. This intermediary module mediates conflicts by determining which speech intent should be prioritized or combined, thereby maintaining high recognition reliability while managing processing complexity through structured conflict resolution.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the system adjusts or mediates speech content from multiple speakers, then recognition rate improves, but processing time increases

Engineering Contradiction:
Improvespeech intention recognition accuracyVSAvoidspeech processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis of speech features (pitch, tone, rhythm) and identifies speaker distinctions before the actual speech recognition process. This preliminary action prepares the speech data in advance, organizing it into separated intent groups that can be processed more efficiently, thereby reducing overall processing time while maintaining high recognition accuracy through pre-organized conflict-free speech representations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12118996B2Method for processing voice signals of multiple speakers, and electronic device according thereto
Publication Date: 2024.10.15 SAMSUNG ELECTRONICS CO LTD
  • US12118996B2 patent drawing
  • US12118996B2 patent drawing
  • US12118996B2 patent drawing

AI summary

Disclosed is an electronic device. The electronic device includes a processor configured to execute one or more instructions stored in a memory to: control a receiver to receive a speech signal; determine whether the received speech signal includes speech signals of a plurality of different speakers; when the received speech signal includes the speech signals of the plurality of different speakers, detect feature information from a speech signal of each speaker; determine relations between pieces of speech content of the plurality of different speakers, based on the detected feature information; determine a response method based on the determined relations between the pieces of speech content; and control the electronic device such that an operation of the electronic device is performed according to the determined response method.