Voice separation and recognition method and device for multiple speakers, terminal equipment and storage medium

By extracting the acoustic, pronunciation and semantic features of the multi-speaker voice signal in the speech separation recognition model, and combining the CTC layer and WFST algorithm for text recognition, the problem of speech signal separation and recognition in the multi-speaker scenario is solved, and high-precision speech separation and text recognition effects are achieved.

CN120048282APending Publication Date: 2025-05-27GUANGDONG POWER GRID CO LTD CUSTOMER SERVICE CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275183.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing speech recognition system is difficult to accurately distinguish and identify the voice signals of different speakers in scenarios of multiple speakers, especially in cases where background noise and interference are not performed well.

Method used

The preset speech separation recognition model is adopted to separate the speech signal by extracting acoustic features, pronunciation features and semantic features, and text recognition is performed using the CTC layer and WFST algorithm. Combined with self-supervised training and data augmentation technology, the joint loss function of the speech separation and text recognition model is optimized to improve the recognition accuracy.

Benefits of technology

It realizes efficient voice signal separation and text recognition in multi-speaker scenarios, significantly improving the recognition accuracy and robustness in complex environments, and is suitable for applications such as meeting records and conference calls that require high-precision voice processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048282A_ABST
    Figure CN120048282A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-speaker voice separation and recognition method and device, terminal equipment and a storage medium, and the method comprises the steps: obtaining to-be-separated and recognized multi-speaker voice signals, inputting the to-be-separated and recognized multi-speaker voice signals to a preset voice separation and recognition model to extract voice features, and separating the multi-speaker voice signals according to the voice features, a plurality of single-person voice signals are obtained; inputting the single-person voice signals and the voice features into a voice text recognition model in a preset voice separation recognition model, recognizing probability distribution of text characters corresponding to each voice frame in the single-person voice signals, and performing weighted calculation according to a WFST algorithm to obtain text information of each single-person voice signal; and finally, according to the single-person voice signal and the corresponding text information, obtaining a separation recognition result of the voice signals of the multiple speakers. According to the invention, the method can achieve the separation and recognition of a mixed voice signal containing a plurality of speakers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice separation, and particularly to a method, device, terminal device and storage medium for separating and recognizing voices of multiple speakers. Background Art

[0002] In the field of speech recognition and processing, speech recognition and speaker separation in multi-speaker scenarios have long been extremely challenging research directions. Nowadays, with the continuous development of intelligent voice interaction technology, speech recognition systems are widely used in scenarios such as smart homes and smart offices, and people's expectations for their performance in complex environments are also getting higher and higher.

[0003] Existing speech recognition systems usually show good single-speaker recognition effects. However, in the case of multiple speakers speaking simultaneously, since speech signals will be superimposed on each other, increasing background noise and interference, it is difficult for traditional speech recognition models to accurately distinguish and recognize speech signals of different speakers. Summary of the Invention

[0004] The present invention provides a method, device, terminal device and storage medium for separating and recognizing voices of multiple speakers, which can separate and recognize mixed voice signals containing multiple speakers.

[0005] An embodiment of the present invention provides a method for separating and recognizing voices of multiple speakers, including:

[0006] Obtaining multi-speaker voice signals to be separated and recognized;

[0007] Inputting the multi-speaker voice signals into a preset voice separation and recognition model, so that the voice separation model in the preset voice separation and recognition model extracts voice features according to the multi-speaker voice signals, and separates the multi-speaker voice signals according to the voice features to obtain several single-speaker voice signals; wherein, the voice features include: acoustic features, prosodic features and semantic features;

[0008] Inputting the single-speaker voice signals and the voice features into the voice text recognition model in the preset voice separation and recognition model, so that the CTC layer in the voice text recognition model, for each single-speaker voice signal, recognizes the probability distribution of the text characters corresponding to each voice frame in the single-speaker voice signal according to the voice features and the single-speaker voice signal, and performs weighted calculation on the probability distribution according to the WFST algorithm to obtain the text information of each single-speaker voice signal;

[0009] Obtaining the separation and recognition result of the multi-speaker voice signals according to the single-speaker voice signals and the corresponding text information.

[0010] Further, the training of the above-mentioned preset voice separation and recognition model includes:

[0011] Obtain the first multi-person mixed voice sample data with the first true label; wherein, the first true label is used to represent several actual single-person voice signals in the first multi-person mixed voice sample data and the actual character information corresponding to each voice frame in each actual single-person voice signal;

[0012] Input the first multi-person mixed voice sample data into the pre-trained voice separation and recognition model, so that the pre-trained voice separation and recognition model extracts voice features from the first multi-person mixed voice sample data according to the built-in pre-trained voice separation model, obtains the first voice feature, and separates several predicted first single-person voice signals from the first multi-person mixed voice sample data according to the first voice feature;

[0013] Input the predicted first single-person voice signal into the pre-trained voice text recognition model built in the voice separation and recognition model, so that the pre-trained voice text recognition model performs text recognition on each predicted first single-person voice signal according to the voice feature, and obtains the first character information of each voice frame in the predicted first single-person voice signal;

[0014] Calculate the joint loss function value according to the predicted first single-person voice signal, the first character information and the first true label;

[0015] When each joint loss function value is obtained, judge whether the joint loss function value converges; if it converges, the training of the voice separation and recognition model is completed, and the above-mentioned preset voice separation and recognition model is obtained; otherwise, after adjusting the parameters of the voice separation and recognition model, continue the training.

[0016] Further, the pre-training of the above-mentioned voice separation model includes:

[0017] Obtain the first multi-person mixed voice sample, the noisy single-person voice sample, and the second multi-person mixed voice sample with the second true label; wherein, the second true label is used to represent several actual single-person voice signals in the second multi-person mixed voice signal;

[0018] Perform self-supervised training on the voice base model to be trained according to the first multi-person mixed voice sample to extract voice features, and obtain the trained voice base model;

[0019] Input the above noisy single-person speech sample into the speech separation model to be trained, so that the speech separation model performs self-supervised training based on the above noisy single-person speech sample for noise separation, and obtain the trained first speech separation model; wherein, the speech separation model to be trained is constructed based on the trained speech base model;

[0020] Input the above second multi-person mixed speech sample into the above first speech separation model, so that the first speech separation model extracts speech features from the above second multi-person mixed speech sample to obtain second speech features;

[0021] Perform speech signal separation on the above second multi-person mixed speech sample according to the above second speech features to obtain a number of predicted second single-person speech signals;

[0022] Calculate the first loss function according to the predicted second single-person speech signal and the above second true label;

[0023] For each obtained first loss function, determine whether the first loss function converges; if it converges, the pre-training of the speech separation model is completed, and the pre-trained speech separation model is obtained; otherwise, after adjusting the parameters in the first speech separation model, continue to train the first speech separation model.

[0024] Further, the above-mentioned self-supervised training of the speech base model to be trained according to the above first multi-person mixed speech sample to extract speech features to obtain the trained speech base model includes:

[0025] Divide the above first multi-person mixed speech sample into a number of short-time frames, and each short-time frame contains a preset number of audio sampling points;

[0026] Input the above short-time frames into the speech base model to be trained, so that the encoder in the speech base model to be trained captures the context relationship between adjacent short-time frames and hierarchically extracts different speech features to obtain the trained speech base model.

[0027] Further, the above speech text recognition model includes: a trained speech base model and a CTC layer;

[0028] The pre-training of the above speech text recognition model includes:

[0029] Obtain a single-person speech sample with a third true label; wherein, the third true label is used to represent the actual character information of each speech frame in the single-person speech sample;

[0030] Input the above single - speaker speech sample into the speech - text recognition model to be trained, so that the trained speech base model extracts features from the above single - speaker speech sample to obtain the third speech feature corresponding to the above single - speaker speech sample;

[0031] Input the above third speech feature into the above CTC layer, so that the CTC layer recognizes each speech frame in the above single - speaker speech sample according to the above third speech feature to obtain the second character information corresponding to each speech frame;

[0032] Calculate the CTC loss function value according to the above second character information and the above third true label;

[0033] Every time a CTC loss function value is obtained, determine whether the above CTC loss function value converges; if so, the pre - training of the above speech - text recognition model is completed to obtain the pre - trained speech - text recognition model; otherwise, after adjusting the parameters in the speech - text recognition model, continue to pre - train the speech - text recognition model.

[0034] Based on the above method - item embodiments, the present invention correspondingly provides device - item embodiments;

[0035] The present invention provides a multi - speaker speech separation and recognition device, including:

[0036] A speech signal acquisition module, a speech signal separation module, a text information recognition module, and a speech separation and recognition result acquisition module;

[0037] The above - mentioned speech signal acquisition module is used to acquire the multi - speaker speech signal to be separated and recognized;

[0038] The above - mentioned speech signal separation module is used to input the above multi - speaker speech signal into a preset speech separation and recognition model, so that the speech separation model in the above preset speech separation and recognition model extracts speech features according to the above multi - speaker speech signal, and separates the above multi - speaker speech signal according to the above speech features to obtain several single - speaker speech signals; wherein, the above speech features include: acoustic features, prosodic features, and semantic features;

[0039] The above - mentioned text information recognition module is used to input the above single - speaker speech signal and the above speech features into the speech - text recognition model in the above preset speech separation and recognition model, so that the CTC layer in the above speech - text recognition model, for each single - speaker speech signal, recognizes the probability distribution of the text characters corresponding to each speech frame in the single - speaker speech signal according to the above speech features and the above single - speaker speech signal, and performs weighted calculation on the above probability distribution according to the WFST algorithm to obtain the text information of each single - speaker speech signal;

[0040] The above-mentioned speech separation and recognition result acquisition module is used to obtain the separation and recognition result of the multi-speaker speech signal according to the above-mentioned single-speaker speech signal and the corresponding text information.

[0041] Furthermore, the above-mentioned speech signal separation module includes:

[0042] A speech sample data acquisition unit, a single-speaker speech signal prediction unit, a single-speaker speech signal recognition unit, a joint loss function calculation unit, and a joint loss function convergence determination unit;

[0043] The above-mentioned speech sample data acquisition unit is used to acquire first multi-speaker mixed speech sample data with a first true label; wherein, the above-mentioned first true label is used to represent several actual single-speaker speech signals in the above-mentioned first multi-speaker mixed speech sample data and the actual character information corresponding to each speech frame in each actual single-speaker speech signal;

[0044] The above-mentioned single-speaker speech signal prediction unit is used to input the above-mentioned first multi-speaker mixed speech sample data into the pre-trained speech separation and recognition model, so that the pre-trained speech separation and recognition model extracts speech features from the above-mentioned first multi-speaker mixed speech sample data according to the built-in pre-trained speech separation model, obtains first speech features, and separates several predicted first single-speaker speech signals from the above-mentioned first multi-speaker mixed speech sample data according to the above-mentioned first speech features;

[0045] The above-mentioned single-speaker speech signal recognition unit is used to input the predicted first single-speaker speech signal into the pre-trained speech text recognition model built in the above-mentioned speech separation and recognition model, so that the pre-trained speech text recognition model performs text recognition on each predicted first single-speaker speech signal according to the above-mentioned speech features, and obtains the first character information of each speech frame in the predicted first single-speaker speech signal;

[0046] The above-mentioned joint loss function calculation unit is used to calculate a joint loss function value according to the predicted first single-speaker speech signal, the above-mentioned first character information, and the above-mentioned first true label;

[0047] The above-mentioned joint loss function convergence determination unit is used to determine whether the above-mentioned joint loss function value converges every time a joint loss function value is obtained; if it converges, the training of the above-mentioned speech separation and recognition model is completed, and the above-mentioned preset speech separation and recognition model is obtained; otherwise, after adjusting the parameters of the above-mentioned speech separation and recognition model, continue the training.

[0048] Furthermore, the above-mentioned single-speaker speech signal prediction unit includes:

[0049] A speech data acquisition subunit, a model self-supervised training subunit, a noise separation training subunit, a feature extraction subunit, a speech sample signal separation subunit, a loss function calculation subunit, and a loss function determination subunit;

[0050] The above-mentioned speech data acquisition subunit is used to acquire a first multi-person mixed speech sample, a noisy single-person speech sample, and a second multi-person mixed speech sample with a second true label; wherein, the above-mentioned second true label is used to represent a plurality of actual single-person speech signals in the above-mentioned second multi-person mixed speech signal;

[0051] The above-mentioned model self-supervised training subunit is used to perform self-supervised training on the speech base model to be trained according to the above-mentioned first multi-person mixed speech sample to extract speech features and obtain a trained speech base model;

[0052] The above-mentioned noise separation training subunit is used to input the above-mentioned noisy single-person speech sample into the speech separation model to be trained, so that the speech separation model performs self-supervised training according to the above-mentioned noisy single-person speech sample to perform noise separation and obtain a trained first speech separation model; wherein, the above-mentioned speech separation model to be trained is constructed based on the trained speech base model;

[0053] The above-mentioned feature extraction subunit is used to input the above-mentioned second multi-person mixed speech sample into the above-mentioned first speech separation model, so that the above-mentioned first speech separation model extracts speech features from the above-mentioned second multi-person mixed speech sample to obtain second speech features;

[0054] The above-mentioned speech sample signal separation subunit is used to separate the speech signal of the above-mentioned second multi-person mixed speech sample according to the above-mentioned second speech features to obtain a plurality of predicted second single-person speech signals;

[0055] The above-mentioned loss function calculation subunit is used to calculate a first loss function according to the predicted second single-person speech signal and the above-mentioned second true label;

[0056] The above-mentioned loss function determination subunit is used to determine whether the above-mentioned first loss function converges every time a first loss function is obtained; if it converges, the pre-training of the above-mentioned speech separation model is completed to obtain a pre-trained speech separation model; otherwise, after adjusting the parameters in the above-mentioned first speech separation model, continue to train the above-mentioned first speech separation model.

[0057] Based on the above method embodiment, the present invention correspondingly provides a terminal device embodiment;

[0058] The present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, a method for separating and recognizing voices of multiple speakers according to any embodiment of the present invention is implemented.

[0059] Based on the above method embodiment, the present invention correspondingly provides a storage medium embodiment;

[0060] The present invention provides a storage medium, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, a method for separating and recognizing voices of multiple speakers according to any embodiment of the present invention is implemented.

[0061] Embodiments of the present invention have the following beneficial effects:

[0062] The present invention provides a method, a device, a terminal device, and a storage medium for separating and recognizing voices of multiple speakers. The method includes: first, obtaining a voice signal of multiple speakers to be separated and recognized; then inputting the voice signal of multiple speakers into a preset voice separation and recognition model, so that the voice separation model in the preset voice separation and recognition model extracts voice features according to the voice signal of multiple speakers, and separates the voice signal of multiple speakers according to the voice features to obtain a plurality of single-speaker voice signals; wherein the voice features include: acoustic features, prosodic features, and semantic features; then inputting the single-speaker voice signals and the voice features into the voice text recognition model in the preset voice separation and recognition model, so that the CTC layer in the voice text recognition model, for each single-speaker voice signal, recognizes the probability distribution of the text characters corresponding to each voice frame in the single-speaker voice signal according to the voice features and the single-speaker voice signal, and performs weighted calculation on the probability distribution according to the WFST algorithm to obtain the text information of each single-speaker voice signal; finally, according to the single-speaker voice signals and the corresponding text information, obtaining the separation and recognition result of the voice signal of multiple speakers. Therefore, the present invention can separate and recognize the voice signal of multiple speakers by first separating the voice signal of multiple speakers to obtain a plurality of single-speaker voice signals, and then performing text recognition on each single-speaker voice signal. Description of the Drawings

[0063] Figure 1 It is a schematic flowchart of a method for separating and recognizing voices of multiple speakers provided by an embodiment of the present invention.

[0064] Figure 2 It is a schematic structural diagram of a device for separating and recognizing voices of multiple speakers provided by an embodiment of the present invention. Detailed implementation manners

[0065] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0066] As Figure 1 shown, a multi-speaker speech separation and recognition method provided by an embodiment of the present invention includes:

[0067] Step S101: Obtain a multi-speaker speech signal to be separated and recognized;

[0068] Specifically, it can be obtained by mixing multiple independent speech samples according to different signal-to-noise ratios, time delays, and spatial position information from a public dataset platform (such as LibriSpeech, CHiME, MUSAN, etc.), or by means of speech synthesis software and audio editing tools to obtain the above multi-speaker speech signal.

[0069] Step S102: Input the above multi-speaker speech signal into a preset speech separation and recognition model, so that the speech separation model in the preset speech separation and recognition model extracts speech features according to the above multi-speaker speech signal, and separates the above multi-speaker speech signal according to the above speech features to obtain several single-speaker speech signals; wherein, the above speech features include: acoustic features, prosodic features, and semantic features;

[0070] In a preferred embodiment, the training of the above preset speech separation and recognition model includes:

[0071] Obtain first multi-person mixed speech sample data with a first true label; wherein, the above first true label is used to represent several actual single-speaker speech signals in the above first multi-person mixed speech sample data and the actual character information corresponding to each speech frame in each actual single-speaker speech signal;

[0072] Input the above first multi-person mixed speech sample data into a pre-trained speech separation and recognition model, so that the pre-trained speech separation and recognition model extracts speech features from the above first multi-person mixed speech sample data according to the built-in pre-trained speech separation model to obtain first speech features, and separates several predicted first single-speaker speech signals from the above first multi-person mixed speech sample data according to the above first speech features;

[0073] Input the predicted first single-person speech signal into the pre-trained speech text recognition model built in the above speech separation and recognition model, so that the above pre-trained speech text recognition model performs text recognition on each predicted first single-person speech signal according to the above speech features, and obtains the first character information of each speech frame in the predicted first single-person speech signal;

[0074] Calculate the joint loss function value according to the predicted first single-person speech signal, the above first character information, and the above first true label;

[0075] Every time a joint loss function value is obtained, determine whether the joint loss function value converges; if it converges, the training of the above speech separation and recognition model is completed, and the above preset speech separation and recognition model is obtained; otherwise, after adjusting the parameters of the above speech separation and recognition model, continue the training.

[0076] Specifically, use the joint loss function to optimize the speech separation and recognition model, that is, optimize the speech separation model and the speech text recognition model in the speech separation and recognition model at the same time. The joint loss function comprehensively considers the accuracy of text recognition and the effect of speech separation, aiming to improve the overall performance of the system in a multi-speaker scenario. Through this joint training, the speech separation and recognition model can better coordinate the information flow between speech separation and text recognition, so that the speech separation model can provide more accurate independent single-person speech signals, and at the same time the speech text recognition model can also perform more accurate text conversion on relatively clear single-person speech signals.

[0077] Specifically, in the initial stage of the above joint training, there may be more text recognition errors or poor speech separation effects. Therefore, the training process needs to be iterated multiple times. Each iteration will adjust the parameters according to the output of the current model, gradually improving the overall performance of the speech separation and recognition model. This method of gradual optimization can ensure that the final speech separation and recognition model has high robustness and accuracy in practical applications.

[0078] Preferably, separating the multi-person mixed speech into independent single-person speech signals first can reduce the interference between speakers, enabling subsequent text recognition to be performed on clearer single-person speech signals, reducing the error of text recognition results, and improving the recognition efficiency.

[0079] Preferably, through joint training, the two tasks of speech separation and text recognition are effectively combined, significantly improving the performance in complex multi-speaker scenarios, especially having a wide range of application prospects in applications that require high-precision speech processing such as meeting records and conference calls.

[0080] In this preferred embodiment, the speech separation and recognition model is jointly trained with the first multi-person mixed speech sample data with the first true label to obtain a trained preset speech separation and recognition model.

[0081] In another preferred embodiment, the pre-training of the above speech separation model includes:

[0082] Obtain a first multi-person mixed speech sample, a noisy single-person speech sample, and a second multi-person mixed speech sample with a second true label; wherein, the second true label is used to represent a plurality of actual single-person speech signals in the second multi-person mixed speech signal;

[0083] Perform self-supervised training on the speech base model to be trained according to the first multi-person mixed speech sample to extract speech features and obtain a trained speech base model;

[0084] Input the noisy single-person speech sample into the speech separation model to be trained, so that the speech separation model performs self-supervised training according to the noisy single-person speech sample to perform noise separation and obtain a trained first speech separation model; wherein, the speech separation model to be trained is constructed based on the trained speech base model;

[0085] Specifically, since the speech separation model to be trained is constructed based on the trained speech base model, and the speech base model already has a high sensitivity and adaptability to speech features during the previous self-supervised training process. On this basis, the single-speaker speech signal is mixed with a randomly sampled noise signal to obtain the above-mentioned noisy single-person speech sample, and the speech separation model is subjected to self-supervised training, using the contrastive learning strategy to enable the model to separate the speech signal from the noise and enhance the model's speech feature separation ability.

[0086] Input the second multi-person mixed speech sample into the first speech separation model, so that the first speech separation model extracts speech features from the second multi-person mixed speech sample to obtain second speech features;

[0087] Separate the speech signal of the second multi-person mixed speech sample according to the second speech feature to obtain a plurality of predicted second single-person speech signals;

[0088] Calculate a first loss function according to the predicted second single-person speech signal and the second true label;

[0089] For each obtained first loss function, determine whether the above first loss function converges; if it converges, the pre-training of the above speech separation model is completed, and the pre-trained speech separation model is obtained; otherwise, after adjusting the parameters in the above first speech separation model, continue to train the above first speech separation model.

[0090] Specifically, when using the second multi-person mixed speech sample to supervise the training of the first speech separation model, through the normalized mean square error (MSE) loss function, the second single-person speech signal predicted by the model is compared with the true label annotated in the second multi-person mixed speech sample. By minimizing the difference between the two, the speech separation quality of the model is enhanced. When the normalized mean square error loss function does not converge, continue to train the first speech separation model with the second multi-person mixed speech sample.

[0091] Specifically, the first speech separation model adopts a convolutional enhanced transformer architecture. It not only uses the multi-head self-attention mechanism in the transformer architecture, which is good at capturing long-distance dependencies and global information in the sequence, but also uses a convolutional module to capture local context information in the speech signal, which can significantly improve the speech separation accuracy.

[0092] Preferably, when training the first speech separation model, to improve the robustness and generalization ability of the model, data augmentation techniques are also used to perform operations such as transformation and noise addition on the second multi-person mixed speech sample, so as to ensure its reliability and stability in practical applications, be able to handle various complex multi-speaker scenarios, and provide high-quality input data for subsequent text recognition tasks.

[0093] In this preferred embodiment, the first speech separation model is trained with the first multi-person mixed speech sample, the noisy single-person speech sample, and the second multi-person mixed speech sample with the second true label to obtain the pre-trained speech separation model.

[0094] In another preferred embodiment, the above speech base model to be trained is self-supervised trained according to the above first multi-person mixed speech sample to extract speech features, and the trained speech base model is obtained, including:

[0095] Divide the above first multi-person mixed speech sample into several short-time frames, and each short-time frame contains a preset number of audio sampling points;

[0096] Input the above short-time frames into the speech base model to be trained, so that the encoder in the speech base model to be trained captures the context relationship between adjacent short-time frames and hierarchically extracts different speech features to obtain the trained speech base model.

[0097] Specifically, the voice base model uses a Transformer neural network as its core architecture, aiming to extract rich voice features from a large number of unlabeled first multi-person mixed voice samples through self-supervised training methods, including acoustic features, prosodic features, and semantic features. First, the first multi-person mixed voice samples are divided into short-time frames, each frame containing a certain number of audio sampling points, and these short-time frames serve as the input data for the model. The Transformer neural network processes these frames through multiple encoder layers, each encoder layer containing a self-attention mechanism and a feed-forward neural network. The encoder effectively captures the context relationship between frames and performs hierarchical feature extraction on the input short-time voice frames. The lower-layer encoder extracts acoustic features, the middle-layer encoder extracts prosodic features, and the upper-layer encoder extracts semantic features.

[0098] Specifically, a contrastive learning strategy is adopted for self-supervised training of the voice base model. Specifically, for each short-time frame, the model not only generates its representation but also compares it with short-time frames at other positions in the same voice sequence. By maximizing the similarity between the currently selected short-time frame and its correct context while minimizing its similarity with randomly selected other short-time frames through a contrastive loss function, the model learns to distinguish relevant and irrelevant voice features. This contrastive learning strategy enables the voice base model to learn rich voice features without explicit labels, thereby enhancing the generalization ability of the model.

[0099] Preferably, this self-supervised training process also includes a noise injection technique, which further improves the robustness of the model by adding different types and levels of noise to the first multi-person mixed voice samples. This method enables the voice base model to better adapt to variable audio environments, including background noise and aliasing effects, laying a solid foundation for subsequent voice separation and text recognition tasks. After such self-supervised training, the voice base model can serve as a powerful feature extractor, providing high-quality voice representations for downstream tasks.

[0100] In this preferred embodiment, the voice base model to be trained is subjected to self-supervised training using the first multi-person mixed voice samples to extract voice features, resulting in a trained voice base model.

[0101] In another preferred embodiment, the above voice text recognition model includes: a trained voice base model and a CTC layer;

[0102] The pre-training of the above voice text recognition model includes:

[0103] Obtaining single-person voice samples with third ground truth labels; wherein the above third ground truth labels are used to represent the actual character information of each voice frame in the above single-person voice samples;

[0104] Input the above single - person speech sample into the speech - text recognition model to be trained, so that the trained speech base model extracts features from the above single - person speech sample, and obtains the third speech feature corresponding to the above single - person speech sample;

[0105] Input the above third speech feature into the above CTC layer, so that the CTC layer recognizes each speech frame in the above single - person speech sample according to the above third speech feature, and obtains the second character information corresponding to each speech frame;

[0106] Calculate the CTC loss function value according to the above second character information and the above third true label;

[0107] Every time a CTC loss function value is obtained, determine whether the above CTC loss function value converges; if so, the pre - training of the speech - text recognition model is completed, and the pre - trained speech - text recognition model is obtained; otherwise, after adjusting the parameters in the speech - text recognition model, continue the pre - training of the speech - text recognition model.

[0108] Specifically, during the training process, the speech - text recognition model receives a single - person speech sample as input, extracts feature representations through the internal speech base model, and then inputs these features into the CTC layer for processing. The CTC layer, based on the extracted third speech feature, generates a probability distribution for each speech frame, representing the distribution of possible characters or words corresponding to each speech frame, and calculates the loss between the probability distribution generated by the CTC layer and the true label using the CTC loss function.

[0109] Preferably, the CTC loss function, that is, the Connectionist Temporal Classification loss function, is very effective in text recognition of speech signals because it does not require precise alignment information between speech frames and true labels. CTC automatically aligns the input speech sequence and label sequence through a dynamic programming algorithm and calculates the optimal matching path. This feature enables the model to maintain flexibility when facing variable - length input speech sequences and can also handle noisy speech signals.

[0110] Preferably, in order to further improve the text recognition ability of the speech - text recognition model, data augmentation techniques are introduced in the training. These techniques include processing the input single - person speech sample such as noise injection, speed perturbation, and volume change, aiming to simulate environmental changes in various actual application scenarios. Through such data augmentation methods, the model can better adapt to various noise conditions, improve the robustness of recognition, and provide a solid foundation for speech separation and text recognition in multi - speaker scenarios.

[0111] In this preferred embodiment, the speech text recognition model is pre-trained with a single-person speech sample with a third true label to obtain a pre-trained speech text recognition model.

[0112] Step S103: Input the above single-person speech signal and the above speech features into the speech text recognition model in the above preset speech separation and recognition model, so that the CTC layer in the above speech text recognition model, for each single-person speech signal, according to the above speech features and the above single-person speech signal, recognizes the probability distribution of the text characters corresponding to each speech frame in the single-person speech signal, and performs weighted calculation on the above probability distribution according to the WFST algorithm to obtain the text information of each single-person speech signal;

[0113] Specifically, in the actual application process, the weighted finite state transducer (WFST) algorithm is used to find the most likely path to determine the text information corresponding to the single-person speech signal, thereby forming the recognition result of the model.

[0114] Step S104: Obtain the separation and recognition result of the above multi-speaker speech signal according to the above single-person speech signal and the corresponding text information.

[0115] Specifically, the preset speech separation and recognition model outputs the separated independent single-person speech signals and the text information corresponding to each single-person speech signal, and the separation and recognition result of the multi-speaker speech signal to be separated and recognized can be obtained.

[0116] Based on the above method item embodiment, the present invention correspondingly provides a device item embodiment.

[0117] As Figure 2 shown, an embodiment of the present invention provides a multi-speaker speech separation and recognition device, including:

[0118] A speech signal acquisition module, a speech signal separation module, a text information recognition module, and a speech separation and recognition result acquisition module;

[0119] The above speech signal acquisition module is used to acquire a multi-speaker speech signal to be separated and recognized;

[0120] The above speech signal separation module is used to input the above multi-speaker speech signal into a preset speech separation and recognition model, so that the speech separation model in the above preset speech separation and recognition model extracts speech features according to the above multi-speaker speech signal, and separates the above multi-speaker speech signal according to the above speech features to obtain a plurality of single-person speech signals; wherein, the above speech features include: acoustic features, prosodic features, and semantic features;

[0121] The above-mentioned text information recognition module is used to input the above-mentioned single-person speech signal and the above-mentioned speech features into the speech text recognition model in the above-mentioned preset speech separation and recognition model, so that the CTC layer in the above-mentioned speech text recognition model, for each single-person speech signal, according to the above-mentioned speech features and the above-mentioned single-person speech signal, recognizes the probability distribution of the text characters corresponding to each speech frame in the single-person speech signal, and performs weighted calculation on the above-mentioned probability distribution according to the WFST algorithm to obtain the text information of each single-person speech signal;

[0122] The above-mentioned speech separation and recognition result acquisition module is used to obtain the separation and recognition result of the above-mentioned multi-speaker speech signal according to the above-mentioned single-person speech signal and the corresponding text information.

[0123] In a preferred embodiment, the above-mentioned speech signal separation module includes:

[0124] A speech sample data acquisition unit, a single-person speech signal prediction unit, a single-person speech signal recognition unit, a joint loss function calculation unit, and a joint loss function convergence determination unit;

[0125] The above-mentioned speech sample data acquisition unit is used to acquire first multi-speaker mixed speech sample data with a first true label; wherein, the above-mentioned first true label is used to represent several actual single-person speech signals in the above-mentioned first multi-speaker mixed speech sample data and the actual character information corresponding to each speech frame in each actual single-person speech signal;

[0126] The above-mentioned single-person speech signal prediction unit is used to input the above-mentioned first multi-speaker mixed speech sample data into the pre-trained speech separation and recognition model, so that the pre-trained speech separation and recognition model extracts speech features from the above-mentioned first multi-speaker mixed speech sample data according to the built-in pre-trained speech separation model to obtain first speech features, and separates several predicted first single-person speech signals from the above-mentioned first multi-speaker mixed speech sample data according to the above-mentioned first speech features;

[0127] The above-mentioned single-person speech signal recognition unit is used to input the predicted first single-person speech signal into the pre-trained speech text recognition model built in the above-mentioned speech separation and recognition model, so that the pre-trained speech text recognition model performs text recognition on each predicted first single-person speech signal according to the above-mentioned speech features to obtain the first character information of each speech frame in the predicted first single-person speech signal;

[0128] The above-mentioned joint loss function calculation unit is used to calculate the joint loss function value according to the predicted first single-person speech signal, the above-mentioned first character information, and the above-mentioned first true label;

[0129] The above-mentioned combined loss function convergence determination unit is used to determine whether the combined loss function value converges every time a combined loss function value is obtained; if it converges, the above-mentioned voice separation and recognition model training is completed, and the above-mentioned preset voice separation and recognition model is obtained; otherwise, after adjusting the parameters of the above-mentioned voice separation and recognition model, continue the training.

[0130] In another preferred embodiment, the above-mentioned single-person voice signal prediction unit includes:

[0131] A voice data acquisition subunit, a model self-supervised training subunit, a noise separation training subunit, a feature extraction subunit, a voice sample signal separation subunit, a loss function calculation subunit, and a loss function determination subunit;

[0132] The above-mentioned voice data acquisition subunit is used to acquire a first multi-person mixed voice sample, a noisy single-person voice sample, and a second multi-person mixed voice sample with a second true label; wherein, the above-mentioned second true label is used to represent several actual single-person voice signals in the above-mentioned second multi-person mixed voice signal;

[0133] The above-mentioned model self-supervised training subunit is used to perform self-supervised training on the voice base model to be trained according to the above-mentioned first multi-person mixed voice sample to extract voice features and obtain a trained voice base model;

[0134] The above-mentioned noise separation training subunit is used to input the above-mentioned noisy single-person voice sample into the voice separation model to be trained, so that the voice separation model performs self-supervised training according to the above-mentioned noisy single-person voice sample to perform noise separation and obtain a first voice separation model that has been trained; wherein, the above-mentioned voice separation model to be trained is constructed based on the trained voice base model;

[0135] The above-mentioned feature extraction subunit is used to input the above-mentioned second multi-person mixed voice sample into the above-mentioned first voice separation model, so that the above-mentioned first voice separation model extracts voice features from the above-mentioned second multi-person mixed voice sample to obtain second voice features;

[0136] The above-mentioned voice sample signal separation subunit is used to separate the voice signals of the above-mentioned second multi-person mixed voice sample according to the above-mentioned second voice features to obtain several predicted second single-person voice signals;

[0137] The above-mentioned loss function calculation subunit is used to calculate a first loss function according to the predicted second single-person voice signal and the above-mentioned second true label;

[0138] The above loss function determination subunit is used to determine, for each obtained first loss function, whether the first loss function converges; if it converges, the pre-training of the speech separation model is completed, and the pre-trained speech separation model is obtained; otherwise, after adjusting the parameters in the first speech separation model, continue to train the first speech separation model.

[0139] It should be noted that the device embodiments described above are merely illustrative. The modules described as separation components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts. The above schematic diagram is only an example of a multi-speaker speech separation and recognition device, and does not constitute a limitation on a multi-speaker speech separation and recognition device. It may include more or fewer components than shown in the figure, or combine some components, or different components.

[0140] Based on the above method embodiment, the present invention correspondingly provides a terminal device embodiment.

[0141] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the multi-speaker speech separation and recognition method in any one of the embodiments of the present invention.

[0142] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more module elements can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the device.

[0143] The above terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor and a memory.

[0144] The so-called processor may be a central processing module (Central Processing Unit, CPU), or may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above-mentioned processor is the control center of the above-mentioned device, and connects various parts of the entire device through various interfaces and circuits;

[0145] The above-mentioned memory can be used to store the above-mentioned computer programs and / or modules. The above-mentioned processor realizes various functions of the above-mentioned device by running or executing the computer programs and / or modules stored in the above-mentioned memory, and calling the data stored in the memory. The above-mentioned memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; in addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (Smart Media Card, SMC), secure digital (Secure Digital, SD) card, flash card (Flash Card), at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0146] Based on the above method item embodiments, the present invention correspondingly provides storage medium item embodiments.

[0147] Another embodiment of the present invention provides a storage medium. The above storage medium includes a stored computer program, wherein when the above computer program runs, it controls the device where the above storage medium is located to execute the multi-speaker voice separation and recognition method according to any one of the embodiments of the present invention.

[0148] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0149] Compared with the prior art, by implementing the above various embodiments of the present invention, a mixed speech signal containing multiple speakers can be separated and recognized.

[0150] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for separating and recognizing speech of multiple speakers, characterized in that: include: Acquire multi-speaker speech signals to be separated and recognized; Inputting the multi-speaker speech signal into a preset speech separation recognition model, so that the speech separation model in the preset speech separation recognition model extracts speech features according to the multi-speaker speech signal, and separates the multi-speaker speech signal according to the speech features to obtain a plurality of single-speaker speech signals; wherein the speech features include: acoustic features, prosodic features, and semantic features; Inputting the single-person voice signal and the voice feature into the speech-to-text recognition model in the preset speech separation and recognition model, so that the CTC layer in the speech-to-text recognition model recognizes the probability distribution of text characters corresponding to each voice frame in the single-person voice signal according to the voice feature and the single-person voice signal for each single-person voice signal, and performs weighted calculation on the probability distribution according to the WFST algorithm to obtain the text information of each single-person voice signal; According to the single-speaker voice signal and the corresponding text information, the separation and recognition results of the multi-speaker voice signals are obtained.

2. The method for separating and recognizing speech of multiple speakers according to claim 1, characterized in that: The training of the preset speech separation and recognition model includes: Acquire first multi-person mixed voice sample data with a first real label; wherein the first real label is used to represent a number of actual single-person voice signals in the first multi-person mixed voice sample data and actual character information corresponding to each voice frame in each actual single-person voice signal; Inputting the first multi-person mixed speech sample data into a pre-trained speech separation recognition model, so that the pre-trained speech separation recognition model extracts speech features from the first multi-person mixed speech sample data according to a built-in pre-trained speech separation model to obtain first speech features, and separates a plurality of predicted first single-person speech signals from the first multi-person mixed speech sample data according to the first speech features; Inputting the predicted first single-person voice signal into the pre-trained voice-to-text recognition model built into the voice separation and recognition model, so that the pre-trained voice-to-text recognition model performs text recognition on each predicted first single-person voice signal according to the voice features, and obtains the first character information of each voice frame in the predicted first single-person voice signal; Calculate a joint loss function value according to the predicted first single-person speech signal, the first character information, and the first true label; Every time a joint loss function value is obtained, it is determined whether the joint loss function value converges; if converged, the speech separation and recognition model training is completed and the preset speech separation and recognition model is obtained; otherwise, the parameters of the speech separation and recognition model are adjusted and the training continues.

3. The method for separating and recognizing speech of multiple speakers according to claim 2, characterized in that: The pre-training of the speech separation model includes: Acquire a first multi-person mixed voice sample, a noisy single-person voice sample, and a second multi-person mixed voice sample with a second real label; wherein the second real label is used to represent a number of actual single-person voice signals in the second multi-person mixed voice signal; Performing self-supervisory training on the speech base model to be trained according to the first multi-person mixed speech sample to extract speech features and obtain a trained speech base model; Inputting the noisy single-person speech sample into the speech separation model to be trained, so that the speech separation model performs self-supervisory training according to the noisy single-person speech sample to perform noise separation, and obtains a trained first speech separation model; wherein the speech separation model to be trained is constructed based on the trained speech base model; Inputting the second multi-person mixed speech sample into the first speech separation model, so that the first speech separation model extracts speech features from the second multi-person mixed speech sample to obtain second speech features; Performing voice signal separation on the second multi-person mixed voice sample according to the second voice feature to obtain a plurality of predicted second single-person voice signals; Calculate a first loss function according to the predicted second single-person speech signal and the second true label; Every time a first loss function is obtained, it is determined whether the first loss function converges; if it converges, the pre-training of the speech separation model is completed, and the pre-trained speech separation model is obtained; otherwise, after adjusting the parameters in the first speech separation model, the first speech separation model continues to be trained.

4. The method for separating and recognizing speech of multiple speakers according to claim 3, characterized in that: The self-supervised training of the speech base model to be trained according to the first multi-person mixed speech sample to extract speech features to obtain a trained speech base model includes: Dividing the first multi-person mixed speech sample into a plurality of short time frames, wherein each short time frame includes a preset number of audio sampling points; The short-time frames are input into the speech base model to be trained, so that the encoder in the speech base model to be trained can capture the contextual relationship between adjacent short-time frames, extract different speech features in layers, and obtain a trained speech base model.

5. The method for separating and recognizing speech of multiple speakers according to claim 4, characterized in that: The speech-to-text recognition model includes: a trained speech base model and a CTC layer; The pre-training of the speech-to-text recognition model includes: Acquire a single-person voice sample with a third real label; wherein the third real label is used to represent actual character information of each voice frame in the single-person voice sample; Inputting the single-person voice sample into the speech-to-text recognition model to be trained, so that the trained speech base model performs feature extraction on the single-person voice sample to obtain a third voice feature corresponding to the single-person voice sample; Inputting the third speech feature into the CTC layer, so that the CTC layer recognizes each speech frame in the single-person speech sample according to the third speech feature, and obtains second character information corresponding to each speech frame; Calculate a CTC loss function value according to the second character information and the third true label; Each time a CTC loss function value is obtained, it is determined whether the CTC loss function value converges; if so, the pre-training of the speech-text recognition model is completed, and the pre-trained speech-text recognition model is obtained; otherwise, after adjusting the parameters in the speech-text recognition model, the speech-text recognition model continues to be pre-trained.

6. A multi-speaker speech separation and recognition device, characterized in that: include: A voice signal acquisition module, a voice signal separation module, a text information recognition module, and a voice separation and recognition result acquisition module; The speech signal acquisition module is used to acquire multi-speaker speech signals to be separated and recognized; The speech signal separation module is used to input the multi-speaker speech signal into a preset speech separation recognition model, so that the speech separation model in the preset speech separation recognition model extracts speech features according to the multi-speaker speech signal, and separates the multi-speaker speech signal according to the speech features to obtain a plurality of single-speaker speech signals; wherein the speech features include: acoustic features, prosodic features and semantic features; The text information recognition module is used to input the single-person voice signal and the voice feature into the voice-text recognition model in the preset voice separation and recognition model, so that the CTC layer in the voice-text recognition model recognizes the probability distribution of text characters corresponding to each voice frame in the single-person voice signal according to the voice feature and the single-person voice signal for each single-person voice signal, and performs weighted calculation on the probability distribution according to the WFST algorithm to obtain the text information of each single-person voice signal; The speech separation and recognition result acquisition module is used to obtain the separation and recognition results of the multi-speaker speech signals according to the single-speaker speech signal and the corresponding text information.

7. The multi-speaker speech separation and recognition device according to claim 6, characterized in that: The speech signal separation module comprises: A speech sample data acquisition unit, a single-person speech signal prediction unit, a single-person speech signal recognition unit, a joint loss function calculation unit, and a joint loss function convergence determination unit; The speech sample data acquisition unit is used to acquire first multi-person mixed speech sample data with a first real label; wherein the first real label is used to represent a number of actual single-person speech signals in the first multi-person mixed speech sample data and actual character information corresponding to each speech frame in each actual single-person speech signal; The single-person speech signal prediction unit is used to input the first multi-person mixed speech sample data into the pre-trained speech separation recognition model, so that the pre-trained speech separation recognition model performs speech feature extraction on the first multi-person mixed speech sample data according to the built-in pre-trained speech separation model to obtain a first speech feature, and separates a plurality of predicted first single-person speech signals from the first multi-person mixed speech sample data according to the first speech feature; The single-person speech signal recognition unit is used to input the predicted first single-person speech signal into the pre-trained speech-to-text recognition model built into the speech separation and recognition model, so that the pre-trained speech-to-text recognition model performs text recognition on each predicted first single-person speech signal according to the speech features, and obtains the first character information of each speech frame in the predicted first single-person speech signal; The joint loss function calculation unit is used to calculate a joint loss function value according to the predicted first single-person speech signal, the first character information and the first true label; The joint loss function convergence judgment unit is used to determine whether the joint loss function value converges each time a joint loss function value is obtained; if converged, the speech separation and recognition model training is completed and the preset speech separation and recognition model is obtained; otherwise, the parameters of the speech separation and recognition model are adjusted and the training continues.

8. The multi-speaker speech separation and recognition device according to claim 7, characterized in that: The single-person speech signal prediction unit comprises: Speech data acquisition subunit, model self-supervision training subunit, noise separation training subunit, feature extraction subunit, speech sample signal separation subunit, loss function calculation subunit and loss function determination subunit; The speech data acquisition subunit is used to acquire a first multi-person mixed speech sample, a noisy single-person speech sample, and a second multi-person mixed speech sample with a second real label; wherein the second real label is used to represent a number of actual single-person speech signals in the second multi-person mixed speech signal; The model self-supervised training subunit is used to perform self-supervised training on the speech base model to be trained according to the first multi-person mixed speech sample to extract speech features and obtain a trained speech base model; The noise separation training subunit is used to input the noisy single-person speech sample into the speech separation model to be trained, so that the speech separation model performs self-supervisory training according to the noisy single-person speech sample to perform noise separation, and obtain a trained first speech separation model; wherein the speech separation model to be trained is constructed based on the trained speech base model; The feature extraction subunit is used to input the second multi-person mixed speech sample into the first speech separation model, so that the first speech separation model extracts speech features from the second multi-person mixed speech sample to obtain second speech features; The voice sample signal separation subunit is used to perform voice signal separation on the second multi-person mixed voice sample according to the second voice feature to obtain a plurality of predicted second single-person voice signals; The loss function calculation subunit is used to calculate a first loss function according to the predicted second single-person speech signal and the second true label; The loss function determination subunit is used to determine whether a first loss function converges each time it is obtained; if it converges, the pre-training of the speech separation model is completed and the pre-trained speech separation model is obtained; otherwise, after adjusting the parameters in the first speech separation model, the first speech separation model continues to be trained.

9. A terminal device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for speech separation and recognition of multiple speakers as claimed in any one of claims 1 to 5 is implemented.

10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the multi-speaker speech separation and recognition method according to any one of claims 1 to 5.