Voice extraction method and device, equipment and medium

Through the combination of cross attention mechanism and language model, the problems of complexity and low speech quality in the prior art are solved, and high-precision and high-quality speech reconstruction are achieved.

CN119993130AActive Publication Date: 2025-05-13PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510253089.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing target speaker's voice extraction technology is complex and the voice quality is not high, making it difficult to effectively capture the long-term dependence between voice tokens.

Method used

The cross attention mechanism is used to fuse the discrete token sequences of reference speech and mixed speech, combine the language model to predict the target human speech, and output the token sequence with high probability distribution through a linear classifier, and finally reconstruct it into a speech waveform.

Benefits of technology

Simplified model training, transform complex audio generation problems into classification problems, effectively capture the long-term dependence between speech tokens, improve the accuracy of speech recognition and prediction, and generate high-quality prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993130A_ABST
    Figure CN119993130A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a voice extraction method, device and equipment and a medium, and the method comprises the steps: firstly obtaining the reference voice of a target speaker and the mixed voice of all speakers; the reference voice and the mixed voice are preprocessed and coded, and two discrete token sequences are generated; fusing the two discrete token sequences to form a fused discrete token sequence; predicting the fused discrete token sequence by using a language model, and generating a candidate discrete token sequence of the target speaker; calculating probability distribution of the candidate token sequences through a linear classifier, and selecting a sequence with high probability as a target discrete token sequence; and reconstructing the target discrete token sequence into a voice waveform to obtain the voice of the target speaker. According to the method, a complex audio generation problem is converted into a classification problem, and model training is simplified; and capturing the long-term dependency relationship between the speech tokens by using the sequence modeling capability of the language model to realize high-quality speech reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a speech extraction method, device, equipment and medium. Background Art

[0002] Target speaker extraction technology aims to accurately extract the speech of a specific speaker from a mixed speech of multiple speakers. This technology is particularly important in speech recognition in financial and medical scenarios. For example, in the telephone customer service system scenario of the financial industry, this technology can help the system accurately identify and extract the customer's voice, ensuring the clarity and accuracy of the call content even in a noisy environment, thereby improving customer experience and service quality. For example, in the medical field, when doctors and patients communicate remotely, this technology can effectively separate the doctor's instructions or the patient's symptom description, assist in medical diagnosis, and ensure the accurate transmission of information.

[0003] Currently, target speaker extraction methods are mainly divided into two categories: discriminative models and generative models. Discriminative models usually adopt a masking strategy to directly minimize the difference between the estimated speech and the pure speech. However, this type of method is difficult to generalize when faced with unseen data and may introduce unnecessary distortion. In contrast, generative models are committed to learning the latent distribution of the target speaker's speech and using the learned knowledge to generate the target speaker's speech from the mixed speech. However, the current target speaker speech extraction process is relatively complex and cannot effectively capture the long-term dependencies between speech tokens, resulting in low quality of the acquired speech. Summary of the invention

[0004] The present invention provides an artificial intelligence speech extraction method, device, computer equipment and medium to solve the technical problems that the existing target speaker speech extraction technology is complex and the speech quality is not high.

[0005] In a first aspect, a speech extraction method is provided, comprising:

[0006] Acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0007] Preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0008] Using a cross attention mechanism to fuse the first discrete token sequence and the second discrete token sequence to obtain a fused discrete token sequence;

[0009] Using a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker;

[0010] A linear classifier is used to predict the probability distribution of each token sequence in the candidate discrete token sequence, and a token sequence with a high probability distribution is output as a target discrete token sequence;

[0011] The target discrete token sequence is reconstructed into a speech waveform to obtain the speech of the target speaker.

[0012] In a second aspect, a speech extraction device is provided, comprising:

[0013] An acquisition module, used to acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0014] An encoding module, used for preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0015] A sequence fusion module, configured to fuse the first discrete token sequence and the second discrete token sequence using a cross attention mechanism to obtain a fused discrete token sequence;

[0016] A candidate sequence prediction module, used to use a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker;

[0017] A target sequence output module, used to predict the probability distribution of each token sequence in the candidate discrete token sequence using a linear classifier, and output the token sequence with a high probability distribution as a target discrete token sequence;

[0018] The speech reconstruction module is used to reconstruct the target discrete token sequence into a speech waveform to obtain the speech of the target speaker.

[0019] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned speech extraction method when executing the computer program.

[0020] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned speech extraction method are implemented.

[0021] In the above-mentioned speech extraction method, device, equipment and medium, the scheme implemented by the speech extraction method: first obtain the reference speech of the target speaker and the mixed speech of all speakers; preprocess and encode the reference speech and the mixed speech to generate two discrete token sequences; use the cross-attention mechanism to fuse the two discrete token sequences to form a fused discrete token sequence; use the language model to predict the fused discrete token sequence to generate a candidate discrete token sequence of the target speaker; calculate the probability distribution of the candidate token sequence through a linear classifier, and select the sequence with high probability as the target discrete token sequence; then reconstruct the target discrete token sequence into a speech waveform to obtain the speech of the target speaker. The present invention converts the complex audio generation problem into a classification problem, which simplifies the model training; uses the sequence modeling ability of the language model to effectively capture the long-term dependency between speech tokens, realizes speech reconstruction, effectively solves the problem of target speech prediction in the case of multi-speaker speech mixing, and improves the accuracy of speech recognition and prediction. It can not only accurately capture the speech characteristics of a specific speaker in a complex multi-person dialogue environment, but also generate high-quality prediction results with high speech quality and intelligibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0023] Figure 1 is a schematic diagram of an application environment of a speech extraction method according to an embodiment of the present invention;

[0024] Figure 2 is a flow chart of a speech extraction method according to an embodiment of the present invention;

[0025] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S202;

[0026] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S203;

[0027] Figure 5 yes Figure 2 A schematic flow chart of a specific implementation of step S204;

[0028] Figure 6 yes Figure 2A schematic flow chart of a specific implementation of step S205;

[0029] Figure 7 yes Figure 2 A schematic flow chart of a specific implementation of step S206;

[0030] Figure 8 is a structural schematic diagram of a speech extraction device in one embodiment of the present invention;

[0031] Fig. 9 is a schematic diagram of a structure of a computer device in one embodiment of the present invention;

[0032] Fig.10 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] The speech extraction method provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can receive the reference speech of the target speaker and the mixed speech of all speakers through the client and generate a target speaker speech extraction task, and send it to the server, which preprocesses and encodes the reference speech and the mixed speech respectively to obtain the corresponding first discrete token sequence and second discrete token sequence; the first discrete token sequence and the second discrete token sequence are fused by a cross-attention mechanism to obtain a fused discrete token sequence; the fused discrete token sequence is predicted by a language model to obtain a candidate discrete token sequence of the target speaker; a linear classifier is used to predict the probability distribution of each token sequence in the candidate discrete token sequence, and the token sequence with a high probability distribution is output as a target discrete token sequence; the target discrete token sequence is reconstructed into a speech waveform to obtain the speech of the target speaker. The present invention converts a complex audio generation problem into a classification problem, simplifies model training; and uses the sequence modeling capability of the language model to effectively capture the long-term dependencies between speech tokens, thereby achieving high-quality speech reconstruction. Among them, the client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server end can be implemented by an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0035] See also Figure 2 As shown, Figure 2 A flow chart of a speech extraction method provided by an embodiment of the present invention includes the following steps:

[0036] S201, obtaining a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0037] In this step, a reference voice of the target speaker (eg, a registered voice) is obtained in advance; the reference voice and the mixed voice may be concatenated to update the mixed voice, and the updated mixed voice may enhance the attention of subsequent models to the target speaker.

[0038] S202, preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0039] S203, using a cross attention mechanism to fuse the first discrete token sequence and the second discrete token sequence to obtain a fused discrete token sequence;

[0040] S204, using a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker;

[0041] S205, using a linear classifier to predict the probability distribution of each token sequence in the candidate discrete token sequence, and outputting the token sequence with a high probability distribution as the target discrete token sequence;

[0042] S206: Reconstruct the target discrete token sequence into a speech waveform to obtain the speech of the target speaker.

[0043] In this embodiment, the target speaker extraction method mainly includes three stages: encoding, modeling and decoding. Encoding stage: the reference speech and the mixed speech are preprocessed and encoded respectively to obtain the corresponding discrete token sequence. Modeling stage: the reference speech information is integrated into the mixed speech token sequence using the cross-attention mechanism, and the dependency between the token sequences is modeled using the language model to predict the token sequence of the target speaker's speech. Decoding stage: the HiFi-GAN vocoder can be used to reconstruct the token sequence of the target speaker's speech into a speech waveform.

[0044] For example, in a specific scenario in the financial field, a financial institution needs to perform voice recognition on customer service staff's calls to extract key information. Due to the large number of customer service staff and the complex call background, it is difficult for traditional voice recognition methods to accurately extract the target speaker's voice. At this time, the target speaker extraction method in this embodiment can be used. First, obtain the reference voice of the customer service staff and the mixed voice to be extracted. Then, the reference voice and the mixed voice are preprocessed and encoded to obtain the corresponding first discrete token sequence and second discrete token sequence. Next, the cross-attention mechanism is used to integrate the reference voice information into the mixed voice token sequence, and the language model is used to model the dependency between token sequences to predict the token sequence of the customer service staff's voice. Finally, the HiFi-GAN vocoder is used to reconstruct the predicted token sequence into a speech waveform to obtain the voice of the customer service staff. In this way, financial institutions can accurately extract the voice of customer service staff and provide strong support for subsequent information processing.

[0045] The present invention converts the complex audio generation problem into a classification problem, thus simplifying model training; it utilizes the sequence modeling capability of the language model to effectively capture the long-term dependencies between speech tokens, thus achieving speech reconstruction, and effectively solving the problem of target speech prediction in the case of multi-speaker speech mixture, thereby improving the accuracy of speech recognition and prediction. It can not only accurately capture the speech features of a specific speaker in a complex multi-person conversation environment, but also generate high-quality prediction results with high speech quality and intelligibility.

[0046] In one embodiment, if Figure 3 As shown, step S202 includes:

[0047] S301, inputting the reference speech into the WavLM model to extract features of multiple hidden layers, obtaining outputs of the multiple hidden layers as first speech features, and using the K-means clustering algorithm to perform feature quantization processing on the first speech features to obtain a first discrete token sequence;

[0048] S302, input the mixed speech into the WavLM model to extract features of multiple hidden layers, obtain the outputs of the multiple hidden layers as the second speech features, and use the K-means clustering algorithm to quantize the second speech features to obtain a second discrete token sequence.

[0049] In this embodiment, the WavLM model is used to extract features from the input reference speech and mixed speech respectively. The WavLM model has a powerful sequence modeling capability and can capture the long-term dependencies between speech tokens, which is crucial for subsequent speech reconstruction and target speech prediction. Through the processing of multiple hidden layers of the WavLM model, we can obtain feature representations containing rich speech information, which are the basis for subsequent feature quantization processing and discrete token sequence generation. In the device, the WavLM model works in conjunction with other components to achieve efficient extraction of the target speaker's speech.

[0050] For example, in a specific scenario in the medical field, it is necessary to extract the voice of the doctor communicating with the patient in a noisy environment. The WavLM model can be used to accurately extract the features of the doctor's voice. First, the doctor's voice is input into the WavLM model as a reference voice. The model will perform an in-depth analysis and extract key feature information. Subsequently, the mixed voice of the doctor and the patient is also input into the WavLM model, and the model will perform feature extraction again. In these two steps, the WavLM model will use its powerful sequence modeling capabilities to capture the long-term dependencies between the doctor's voice tokens, thereby ensuring that the extracted features can truly reflect the characteristics of the doctor's speech.

[0051] In one embodiment, if Figure 4 As shown, step S203 includes:

[0052] S401, extracting features from the first discrete token sequence and the second discrete token sequence respectively to obtain corresponding first sequence features and second sequence features;

[0053] S402, performing information interaction on the first sequence features and the second sequence features through a cross calculation module to generate an interaction feature matrix;

[0054] S403: Use the attention mechanism to assign weights to the interactive feature matrix to determine the association weights between the first sequence features and the second sequence features.

[0055] S404, performing sequence alignment on the first sequence feature and the second sequence feature according to the association weight to generate an alignment feature vector;

[0056] S405, mapping the aligned feature vector to a unified feature space to obtain a mapped feature vector;

[0057] S406: Reconstruct the mapped feature vector to generate a fused discrete token sequence.

[0058] In this embodiment, the process of steps S401-S406 respectively extracts features from the first discrete token sequence and the second discrete token sequence to obtain corresponding first sequence features and second sequence features. The cross calculation module is responsible for information interaction between these features and generating an interactive feature matrix. Next, the attention mechanism module assigns weights to the interactive feature matrix to determine the association weights between the first sequence features and the second sequence features. The sequence alignment module performs sequence alignment on the features according to these association weights to generate an alignment feature vector. The feature space mapping module maps the alignment feature vector to a unified feature space to obtain a mapping feature vector. Finally, the sequence reconstruction module reconstructs the mapping feature vector to generate a fused discrete token sequence, thereby efficiently extracting the speech information of the target speaker.

[0059] In one embodiment, if Figure 5 As shown, step S204 includes:

[0060] S501, using a pre-trained language model to perform feature extraction on the fused discrete token sequence, and generating a total vector representation containing speech features of each speaker;

[0061] S502, using a pre-trained language model to perform feature extraction on the first discrete token sequence to generate a reference vector representation of the target speaker's speech features;

[0062] S503: Match the total vector representation with the reference vector representation through a matching algorithm, calculate the matching score, and select feature sequences with high matching degree with the reference vector from the total vector representation as candidate discrete token sequences of the target speaker.

[0063] In this embodiment, the process of steps S501-S503 extracts features from the fused discrete token sequence and the first discrete token sequence through a pre-trained language model, and generates a total vector representation and a reference vector representation, respectively. These two vector representations contain the speech features of each speaker and the speech features of the target speaker, respectively. Subsequently, a matching algorithm is used to match the total vector representation with the reference vector representation, and a feature sequence with a high degree of match with the reference vector is screened out by calculating the matching score. These feature sequences are used as candidate discrete token sequences for the target speaker, providing key information for subsequent processing. Through such a design, the speech extraction process can efficiently and accurately extract the speech information of the target speaker, providing strong support for applications in the fields of speech recognition and speech analysis.

[0064] In one embodiment, if Figure 6 As shown, step S205 includes:

[0065] S601, extracting features from each token sequence in the candidate discrete token sequence to form a feature vector set;

[0066] S602, constructing a linear classifier model and setting model parameters including a weight vector and a bias term;

[0067] S603, input the feature vector set into the linear classifier, and obtain the probability distribution of each token sequence through weight calculation and bias adjustment;

[0068] S604: Select a token sequence whose probability distribution is higher than a preset classification threshold as a target discrete token sequence.

[0069] In this embodiment, the process of steps S601-S604 further analyzes and processes the candidate discrete token sequence through a specific processing module. First, feature extraction is performed on each token sequence. This step can deeply mine the key information in each token sequence and convert it into a feature vector set to provide data support for subsequent classification processing. Then, a linear classifier model is constructed, and model parameters including weight vectors and bias terms are set. The setting of these parameters has a crucial impact on the performance of the classifier.

[0070] After the model is built, the feature vector set is input into the linear classifier. Through weight calculation and bias adjustment, the classifier can output a probability distribution for each token sequence. This probability distribution reflects the possibility that each token sequence belongs to the target speaker's speech. Finally, according to the preset classification threshold, the token sequence with a probability distribution higher than the threshold is selected as the target discrete token sequence. These target discrete token sequences accurately reflect the speech characteristics of the target speaker and provide a reliable basis for subsequent processing and application.

[0071] Through the design of steps S601-S604, the speech extraction process can efficiently filter out the speech information of the target speaker, providing strong support for applications in the fields of speech recognition, speech analysis, etc. At the same time, the device also has high accuracy and robustness, and can adapt to the speech extraction needs in different scenarios.

[0072] In one embodiment, if Figure 7 As shown, step S206 includes:

[0073] S701, extracting acoustic features of a target discrete token sequence;

[0074] S702, input the acoustic features into the HiFi-GAN vocoder for sequence reconstruction to generate an initial speech waveform;

[0075] S703: Perform noise reduction processing on the initial speech waveform and output the speech waveform of the target speaker.

[0076] In this embodiment, steps S701-S703 together constitute a complete speech reconstruction and optimization process. Among them, S701 provides key input information for subsequent steps by extracting the acoustic features of the target discrete token sequence. These acoustic features accurately reflect the speech characteristics of the target speaker and are the basis for speech reconstruction. S702 uses the HiFi-GAN vocoder to reconstruct the sequence and generate an initial speech waveform. The HiFi-GAN vocoder has a wide range of applications in the field of speech synthesis due to its powerful generation capability and high-quality output. Through this vocoder, we can convert acoustic features into realistic speech waveforms, providing high-quality input for subsequent processing. S703 performs noise reduction on the initial speech waveform to eliminate possible background noise and interference. This step is crucial to improving the clarity and purity of the output speech. After noise reduction, we finally obtained the speech waveform of the target speaker, which accurately reflects the speech characteristics of the target speaker and provides a reliable basis for subsequent applications and processing.

[0077] This embodiment combines the above steps S701-S703 to achieve efficient and accurate speech extraction and reconstruction. The device not only has high accuracy and robustness, but also can adapt to speech extraction requirements in different scenarios, providing strong support for applications in the fields of speech recognition, speech analysis, etc.

[0078] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0079] In some other embodiments, the present invention can be effectively applied to dual tasks of speech synthesis and identity verification.

[0080] For example, in a telephone customer service scenario in the financial field, when a customer initiates a business consultation through the IVR system, the standard response text of the customer service representative (the first voice feature) and the customer's real-time voice (the second voice feature) are first converted into a discrete token sequence through the Mel spectrum extraction module. The cross-computation module uses a multi-head cross-modal attention mechanism to establish a dynamic association between the prosodic features of the customer service response (average fundamental frequency 230Hz, speech rate 4.2 words / second) and the personalized features of the customer's voice (specific voiceprint fingerprint).

[0081] In the identity verification phase, the present invention compares the interactive feature matrix frame by frame through the attention weight allocation module. When the resonance peak characteristics of the customer's voice (such as F1 = 500Hz, F2 = 1500Hz) are detected and the matching degree with the reserved voiceprint template reaches the preset threshold, the sequence alignment module will generate a fusion feature vector with time synchronization. After the vector is reconstructed by the WaveGlow vocoder, it can retain the standard business language of the customer service end (such as "Your account balance is...") and embed the customer's unique timbre characteristics to ensure that the synthesized voice meets financial compliance requirements.

[0082] In one embodiment, a speech extraction device is provided, which corresponds one-to-one to the speech extraction method in the above embodiment. Figure 8 As shown, the speech extraction device includes an acquisition module 801, an encoding module 802, a sequence fusion module 803, a candidate sequence prediction module 804, a target sequence output module 805 and a speech reconstruction module 806. The functional modules are described in detail as follows:

[0083] An acquisition module 801 is used to acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0084] The encoding module 802 is used to pre-process and encode the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0085] A sequence fusion module 803 is used to fuse the first discrete token sequence and the second discrete token sequence using a cross attention mechanism to obtain a fused discrete token sequence;

[0086] A candidate sequence prediction module 804 is used to use a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker;

[0087] A target sequence output module 805 is used to predict the probability distribution of each token sequence in the candidate discrete token sequence using a linear classifier, and output the token sequence with a high probability distribution as the target discrete token sequence;

[0088] The speech reconstruction module 806 is used to reconstruct the target discrete token sequence into a speech waveform to obtain the speech of the target speaker.

[0089] In one embodiment, the encoding module 802 is specifically configured to:

[0090] Input the reference speech into the WavLM model to extract features of multiple hidden layers, obtain the outputs of the multiple hidden layers as the first speech features, and use the K-means clustering algorithm to perform feature quantization processing on the first speech features to obtain a first discrete token sequence;

[0091] The mixed speech is input into the WavLM model to extract features of multiple hidden layers, and the outputs of the multiple hidden layers are obtained as the second speech features. The second speech features are quantized using the K-means clustering algorithm to obtain a second discrete token sequence.

[0092] In one embodiment, the sequence fusion module 803 is specifically configured to:

[0093] Perform feature extraction on the first discrete token sequence and the second discrete token sequence respectively to obtain corresponding first sequence features and second sequence features;

[0094] Perform information interaction on the first sequence features and the second sequence features through a cross calculation module to generate an interactive feature matrix;

[0095] The attention mechanism is used to assign weights to the interaction feature matrix and determine the association weights between the first sequence features and the second sequence features.

[0096] Perform sequence alignment on the first sequence feature and the second sequence feature according to the association weight to generate an alignment feature vector;

[0097] Mapping the aligned feature vector to a unified feature space to obtain a mapped feature vector;

[0098] The mapped feature vector is reconstructed to generate a fused discrete token sequence.

[0099] In one embodiment, the candidate sequence prediction module 804 is specifically configured to:

[0100] Use the pre-trained language model to extract features from the fused discrete token sequence and generate a total vector representation containing the speech features of each speaker;

[0101] Using a pre-trained language model to extract features from the first discrete token sequence, generating a reference vector representation of the target speaker's speech features;

[0102] The total vector representation is matched with the reference vector representation through a matching algorithm, the matching score is calculated, and the feature sequence with a high matching degree with the reference vector is screened out from the total vector representation and used as the candidate discrete token sequence of the target speaker.

[0103] In one embodiment, the target sequence output module 805 is specifically used to:

[0104] Perform feature extraction on each token sequence in the candidate discrete token sequence to form a feature vector set;

[0105] Build a linear classifier model and set the model parameters including weight vector and bias term;

[0106] Input the feature vector set into the linear classifier, and obtain the probability distribution of each token sequence through weight calculation and bias adjustment;

[0107] Select the token sequence whose probability distribution is higher than the preset classification threshold as the target discrete token sequence.

[0108] In one embodiment, the speech reconstruction module 806 is specifically configured to:

[0109] Extract the acoustic features of the target discrete token sequence;

[0110] The acoustic features are input into the HiFi-GAN vocoder for sequence reconstruction to generate the initial speech waveform;

[0111] The initial speech waveform is subjected to noise reduction processing and the speech waveform of the target speaker is output.

[0112] The present invention provides a speech extraction device, in which a target speaker extraction method mainly includes three stages: encoding, modeling and decoding. Encoding stage: preprocessing and encoding the reference speech and the mixed speech respectively to obtain the corresponding discrete token sequence. Modeling stage: using the cross-attention mechanism to integrate the reference speech information into the mixed speech token sequence, and using the language model to model the dependency between the token sequences, and predicting the token sequence of the target speaker's speech. Decoding stage: using the HiFi-GAN vocoder to reconstruct the token sequence of the target speaker's speech into a speech waveform.

[0113] The present invention converts the complex audio generation problem into a classification problem, thus simplifying model training; it utilizes the sequence modeling capability of the language model to effectively capture the long-term dependencies between speech tokens, thus achieving speech reconstruction, and effectively solving the problem of target speech prediction in the case of multi-speaker speech mixture, thereby improving the accuracy of speech recognition and prediction. It can not only accurately capture the speech features of a specific speaker in a complex multi-person conversation environment, but also generate high-quality prediction results with high speech quality and intelligibility.

[0114] For the specific definition of the speech extraction device, please refer to the definition of the speech extraction method above, which will not be repeated here. Each module in the above-mentioned speech extraction device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0115] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a speech extraction method server side.

[0116] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Fig.10As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the client side of a speech extraction method

[0117] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0118] Acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0119] Preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0120] The first discrete token sequence and the second discrete token sequence are fused by using a cross attention mechanism to obtain a fused discrete token sequence;

[0121] Use the language model to predict the target speaker’s speech on the fused discrete token sequence to obtain the candidate discrete token sequence of the target speaker;

[0122] A linear classifier is used to predict the probability distribution of each token sequence in the candidate discrete token sequence, and the token sequence with a high probability distribution is output as the target discrete token sequence;

[0123] The target discrete token sequence is reconstructed into a speech waveform to obtain the speech of the target speaker.

[0124] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0125] Acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates;

[0126] Preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence;

[0127] The first discrete token sequence and the second discrete token sequence are fused by using a cross attention mechanism to obtain a fused discrete token sequence;

[0128] Use the language model to predict the target speaker’s speech on the fused discrete token sequence to obtain the candidate discrete token sequence of the target speaker;

[0129] A linear classifier is used to predict the probability distribution of each token sequence in the candidate discrete token sequence, and the token sequence with a high probability distribution is output as the target discrete token sequence;

[0130] The target discrete token sequence is reconstructed into a speech waveform to obtain the speech of the target speaker.

[0131] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0132] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0133] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0134] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A speech extraction method, characterized in that: include: Acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates; Preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence; Using a cross attention mechanism to fuse the first discrete token sequence and the second discrete token sequence to obtain a fused discrete token sequence; Using a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker; A linear classifier is used to predict the probability distribution of each token sequence in the candidate discrete token sequence, and a token sequence with a high probability distribution is output as a target discrete token sequence; The target discrete token sequence is reconstructed into a speech waveform to obtain the speech of the target speaker.

2. The speech extraction method according to claim 1, wherein: The preprocessing and encoding of the reference speech and the mixed speech respectively to obtain the corresponding first discrete token sequence and second discrete token sequence includes: Input the reference speech into the WavLM model to extract features of multiple hidden layers, obtain outputs of the multiple hidden layers as first speech features, and use a K-means clustering algorithm to perform feature quantization processing on the first speech features to obtain a first discrete token sequence; The mixed speech is input into the WavLM model to extract features of multiple hidden layers, and the outputs of the multiple hidden layers are obtained as the second speech features. The second speech features are quantized using the K-means clustering algorithm to obtain a second discrete token sequence.

3. The speech extraction method according to claim 1, wherein: The method of using a cross attention mechanism to fuse the first discrete token sequence and the second discrete token sequence to obtain a fused discrete token sequence includes: Performing feature extraction on the first discrete token sequence and the second discrete token sequence respectively to obtain corresponding first sequence features and second sequence features; Perform information interaction on the first sequence features and the second sequence features through a cross calculation module to generate an interactive feature matrix; Using an attention mechanism to assign weights to the interaction feature matrix, and determining association weights between the first sequence feature and the second sequence feature; Performing sequence alignment on the first sequence feature and the second sequence feature according to the association weight to generate an alignment feature vector; Mapping the aligned feature vector to a unified feature space to obtain a mapped feature vector; The mapped feature vector is reconstructed to generate a fused discrete token sequence.

4. The speech extraction method according to claim 2, wherein: The method of using a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker includes: Using a pre-trained language model to perform feature extraction on the fused discrete token sequence to generate a total vector representation containing speech features of each speaker; Using a pre-trained language model to perform feature extraction on the first discrete token sequence to generate a reference vector representation of the target speaker's speech features; The total vector representation is matched with the reference vector representation through a matching algorithm, a matching score is calculated, and a feature sequence with a high matching degree with the reference vector is screened out from the total vector representation and used as a candidate discrete token sequence of the target speaker.

5. The speech extraction method according to claim 1, characterized in that: The using a linear classifier to predict the probability distribution of each token sequence in the candidate discrete token sequence, and outputting the token sequence with a high probability distribution as the target discrete token sequence, includes: Performing feature extraction on each token sequence in the candidate discrete token sequence to form a feature vector set; Build a linear classifier model and set the model parameters including weight vector and bias term; Input the feature vector set into a linear classifier, and obtain the probability distribution of each token sequence through weight calculation and bias adjustment; Select the token sequence whose probability distribution is higher than the preset classification threshold as the target discrete token sequence.

6. The speech extraction method according to claim 1, characterized in that: The step of reconstructing the target discrete token sequence into a speech waveform to obtain the speech of the target speaker includes: Extracting acoustic features of the target discrete token sequence; Inputting the acoustic features into a HiFi-GAN vocoder for sequence reconstruction to generate an initial speech waveform; The initial speech waveform is subjected to noise reduction processing, and a speech waveform of a target speaker is output.

7. The speech extraction method according to claim 1, characterized in that: Also includes: The reference speech and the mixed speech are concatenated to update the mixed speech.

8. A speech extraction device, characterized in that: include: An acquisition module, used to acquire a reference speech of a target speaker and a mixed speech to be extracted in which the target speaker participates; An encoding module, used for preprocessing and encoding the reference speech and the mixed speech respectively to obtain a corresponding first discrete token sequence and a second discrete token sequence; A sequence fusion module, configured to fuse the first discrete token sequence and the second discrete token sequence using a cross attention mechanism to obtain a fused discrete token sequence; A candidate sequence prediction module, used to use a language model to perform target speaker speech prediction on the fused discrete token sequence to obtain a candidate discrete token sequence of the target speaker; A target sequence output module, used to predict the probability distribution of each token sequence in the candidate discrete token sequence using a linear classifier, and output the token sequence with a high probability distribution as a target discrete token sequence; The speech reconstruction module is used to reconstruct the target discrete token sequence into a speech waveform to obtain the speech of the target speaker.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the speech extraction method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the speech extraction method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Target speaker extraction system based on voice discretization and vocoder

    CN117912469A

  • Voice generation method and device, equipment and medium

    CN119360819A

  • Training speech recognition systems using word sequences

    US10672383B1