Voice Separation Model With Automatic Speaker Count Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice separation models require pre-determination of the number of speaking objects in voice information, leading to complex and inefficient separation processes.

Innovation Solution

A model determining method that includes an initial quantity determining module to analyze the number of speaking objects and an initial voice separation module to perform separation based on this quantity, allowing for automatic voice separation without prior object recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the quantity of speaking objects is determined in advance before voice separation, then the voice separation model can obtain accurate separation results, but the voice separation process becomes complex and requires high requirements on information input

Engineering Contradiction:
Improvevoice separation accuracyVSAvoidvoice separation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the quantity determination function and voice separation function into a single integrated voice separation model. The model simultaneously performs both tasks: determining the number of speaking objects and separating their voices, eliminating the need for separate preprocessing steps and reducing overall system complexity while maintaining separation accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The voice separation model is designed with multi-functionality, enabling it to perform both quantity determination and voice separation tasks. This universal model can adapt to different numbers of speakers automatically without requiring separate modules or preprocessing, simplifying the information input requirements while preserving accurate separation results.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the quantity of speaking objects is determined in advance, then accurate voice separation can be achieved, but preprocessing is required which reduces efficiency

Engineering Contradiction:
Improvevoice separation accuracyVSAvoidvoice separation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The model incorporates preliminary action by integrating the quantity determination capability directly within the voice separation model structure. This allows the model to automatically determine the number of speakers as part of the separation process itself, eliminating the need for separate preprocessing steps and improving overall efficiency while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By combining quantity determination and voice separation into a single unified model, the patent eliminates redundant preprocessing steps. The model processes the input signal once, simultaneously determining speaker quantity and performing separation, thereby improving productivity without sacrificing separation accuracy.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If a separate quantity determination step is added before voice separation, then speaking objects can be accurately identified, but the overall process complexity increases

Engineering Contradiction:
Improvespeaking object information accuracyVSAvoidprocess complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges the quantity determination step with the voice separation step into a single integrated model. This unified approach ensures that speaking object information is accurately identified while avoiding the increased complexity that would result from adding a separate preprocessing module. The model achieves both goals simultaneously through its multi-functional architecture.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4708285A1Model determination method, model application method, and related device
Publication Date: 2026.03.11 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP4708285A1 patent drawingFigure 1
  • EP4708285A1 patent drawingFigure 2
  • EP4708285A1 patent drawingFigure 3

AI summary

Embodiments of the present application disclose a model determination method, a model application method, and a related device. An initial voice separation model comprises an initial number determination module used for analyzing the number of voice producing objects, and an initial voice separation module used for performing voice separation on the basis of the number of voice producing objects determined by the initial number determination module, and sample voice information only needs to be inputted, so that a voice separation result can be obtained by means of separation by the model. On the basis of the difference between an accurate voice separation result corresponding to the sample voice information and a model output, the accuracy of the model for analysis of the number of the voice producing objects and the accuracy of the model for the separation of the voice information can be reflected. Therefore, parameter adjustment is performed on the initial voice separation model on the basis of the difference, so that the model can learn at the same time how to accurately analyze the number of the voice producing objects and accurately separate the voice information, the obtained voice separation model can realize accurate voice separation without other information input than voice information to be separated, thereby improving the voice separation efficiency.