Multi-Speaker Voice Signal Processing with Conflict Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face challenges in accurately processing and distinguishing speech signals from multiple speakers simultaneously without conflicts, leading to inefficiencies in recognizing speech intentions and performing operations on electronic devices.
Innovation Solution
A method involving a machine learning module or language understanding module to determine relations between speech content from multiple speakers, adjusting or mediating the content to prevent conflicts, and prioritizing or combining speech intentions for appropriate device operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing speech recognition technologies process speech signals from multiple speakers, then the system can recognize speech intentions, but conflicts arise between different speakers causing reduced recognition accuracy
Solution Approach 1:
The patent segments the mixed speech signal into individual speaker components by detecting speaker-specific features (pitch, tone, rhythm) and separating them into distinct speech intent groups. This segmentation allows the system to process each speaker's intent independently, resolving conflicts and improving overall recognition accuracy in multi-speaker environments.
2Device complexity
If the system processes speech signals without conflict resolution, then processing is simpler, but recognition rate decreases due to overlapping speech content
Solution Approach 1:
The patent introduces an intermediary processing layer that analyzes the relationships between different speech contents before executing operations. This intermediary module mediates conflicts by determining which speech intent should be prioritized or combined, thereby maintaining high recognition reliability while managing processing complexity through structured conflict resolution.
3Measurement precision
If the system adjusts or mediates speech content from multiple speakers, then recognition rate improves, but processing time increases
Solution Approach 1:
The patent performs preliminary analysis of speech features (pitch, tone, rhythm) and identifies speaker distinctions before the actual speech recognition process. This preliminary action prepares the speech data in advance, organizing it into separated intent groups that can be processed more efficiently, thereby reducing overall processing time while maintaining high recognition accuracy through pre-organized conflict-free speech representations.
Data Source
AI summary
Disclosed is an electronic device. The electronic device includes a processor configured to execute one or more instructions stored in a memory to: control a receiver to receive a speech signal; determine whether the received speech signal includes speech signals of a plurality of different speakers; when the received speech signal includes the speech signals of the plurality of different speakers, detect feature information from a speech signal of each speaker; determine relations between pieces of speech content of the plurality of different speakers, based on the detected feature information; determine a response method based on the determined relations between the pieces of speech content; and control the electronic device such that an operation of the electronic device is performed according to the determined response method.


