Method, device, medium and product for extracting target speaker based on historical conversation text

By combining historical dialogue text and deep learning technology, the mixed speech signals in complex multispeaker scenes were successfully extracted, and the voice of the target speaker was solved, which solved the recognition difficulties of the existing technology when dealing with overlapping speech and noise, and achieved a more efficient speech extraction effect.

CN119626211BActive Publication Date: 2025-05-30TRUE SPACE (ZHUHAI) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162330.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-30
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In complex multispeaker scenarios, existing speech enhancement and speech separation methods are difficult to accurately extract the target speaker's voice, especially when dealing with overlapping speech and strong background noise.

Method used

By using historical dialogue text as a clue, combining speech encoder and text encoder to process mixed speech signals and text prompt information, fusing spectrum features and text features, and using mask estimator and echo prompt module to extract the speech signal of the target speaker.

Benefits of technology

It realizes efficiently extracting the voice of the target speaker without pre-registering, overcomes the limitations of existing methods when dealing with complex acoustic scenarios, and improves the recognition effect and widespreadness of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119626211B_ABST
    Figure CN119626211B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, medium and product for extracting a target speaker based on historical dialogue texts. The method includes: processing a mixed speech signal through a voice encoder to obtain a spectral feature embedding; wherein, the mixed speech signal includes the speech signal of the target speaker, the speech signal of the interfering speaker, and a noise signal; processing text prompt information through a text encoder to obtain a text feature embedding; the text prompt information is associated with the text content corresponding to the speech signal of the target speaker; fusing the spectral feature embedding and the text feature embedding through a fusion layer to obtain a fused feature; processing the fused feature through a mask estimator to obtain a target mask corresponding to the fused feature; and extracting the speech signal of the target speaker according to the target mask and the mixed speech signal. The present invention can improve the recognition effect of the target speaker and does not require pre-registration of speech, which is convenient to implement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular to a method, device, medium and product for extracting a target speaker based on historical conversation texts. Background Art

[0002] In complex multi-speaker scenarios, extracting the speech of a target speaker remains a great challenge due to severe speaker overlap and significant background noise. This is particularly important in dialogue-based applications, such as artificial intelligence voice assistants, where accurately understanding human intent is crucial.

[0003] However, existing deep learning-based speech enhancement (SE) and speech separation (SS) methods perform poorly in dealing with complex acoustic scenarios, especially in differentiating similar sound sources, maintaining speech integrity, and solving the identity permutation problem.

[0004] Traditional speech enhancement methods such as spectral subtraction and Wiener filtering are difficult to cope with rapidly changing noise, resulting in speech distortion. Recent deep learning methods, such as the Multi-Parallel Spectral Subtraction and Enhancement Network (MP-SENet) and the Full Residual Co-Attention Network (FRCRN) for noise reduction, perform well in denoising but do not handle interfering speakers. To address multi-speaker scenarios, speech separation methods include TF-GridNet, Conv-TasNet, DPRNN, DPTNet, SPGM, Sep-Former, and MossFormer2. However, these methods require prior knowledge of the number of speakers and are difficult to accurately associate the separated signals with speaker identities, i.e., the ID permutation problem.

[0005] Although Permutation Invariant Training (PIT) was initially proposed to solve the identity permutation problem in speech separation tasks, it still cannot accurately identify which separated speech corresponds to the target speaker. As an alternative, Target Speaker Extraction (TSE) uses auxiliary information to extract the target speech from mixed and noisy speech, thus overcoming the limitations of speech enhancement and speech separation methods.

[0006] Target speaker extraction methods aim to extract the speech of a target speaker without prior knowledge. However, due to the unavailability of pre-registered speech, this method cannot be well applied in dialogue-based artificial intelligence voice assistants. Summary of the Invention

[0007] The first object of the present invention is to provide a method for extracting a target speaker based on historical conversation texts with better performance and more convenient implementation.

[0008] The second object of the present invention is to provide a computer device for implementing the above method for extracting a target speaker based on historical conversation texts.

[0009] The third object of the present invention is to provide a computer-readable storage medium for implementing the above method for extracting a target speaker based on historical conversation texts.

[0010] The fourth object of the present invention is to provide a computer program product for implementing the above method for extracting a target speaker based on historical conversation texts.

[0011] To achieve the above first object, the present invention provides a method for extracting a target speaker based on historical conversation texts, which includes the following steps: processing a mixed speech signal through a voice encoder to obtain a spectral feature embedding; wherein, the mixed speech signal includes the speech signal of the target speaker, the speech signal of the interfering speaker, and a noise signal; processing text prompt information through a text encoder to obtain a text feature embedding; the text content corresponding to the speech signal of the target speaker matches the text prompt information; fusing the spectral feature embedding and the text feature embedding through a fusion layer to obtain a fused feature; processing the fused feature through a mask estimator to obtain a target mask corresponding to the fused feature; and extracting the speech signal of the target speaker according to the target mask and the mixed speech signal.

[0012] As can be seen from the above solution, the present invention processes the speech dialogue scenario of overlapping speech and strong background noise, uses the historical conversation text as a clue to explore the semantic association between the question and the human answer, and uses the prior knowledge of context prompts as a clue, which can effectively extract the speech of the target speaker. Since no voice pre-registration is required, it is more convenient to implement than the existing solutions and has a wider range of usage scenarios.

[0013] A further solution is that the mask estimator includes a plurality of DPRNN modules. When processing the fused feature through the mask estimator, it includes: dividing the fused feature into segments, connecting all the segments to form a block, the block is sequentially processed through a plurality of DPRNN modules and a plurality of echo prompt blocks of an echo prompt module, and finally the block is reconstructed through an aggregation layer, and a target mask is generated through an activation layer.

[0014] Thus, by designing an echo prompt module, noise with a low signal-to-noise ratio is effectively removed at multiple granularity levels, improving the recognition effect.

[0015] A further solution is that the number of echo prompt blocks is 3.

[0016] It can be seen that when the number of echo prompt blocks is 3, a better recognition effect can be obtained.

[0017] A further solution is that when obtaining the spectral feature embedding by processing the mixed speech signal through a voice encoder, it includes: processing the mixed speech signal through a one-dimensional convolutional layer, and then applying a ReLU layer to obtain the spectral feature embedding.

[0018] A further solution is that when obtaining the text feature embedding by processing the text prompt information through a text encoder, it includes: generating the text feature embedding corresponding to the text prompt information through a contrastive language audio pre-trained model.

[0019] It can be seen that the text feature embedding corresponding to the text prompt information can be obtained more accurately.

[0020] A further solution is that the fusion layer, the mask estimator, and the echo prompt module are obtained through a deep learning training method, including: setting a data set, the data set includes a plurality of speech segments and a plurality of noise segments, and each of the speech segments is set with a corresponding uniquely matched text generation problem, and two of the speech segments and one of the noise segments are randomly selected and mixed using a random SNR mixing method during training.

[0021] A further solution is that both the speech segments and most of the noise segments are sampled to 8 kHz.

[0022] It can be seen that it can better conform to the usage scenario of an artificial intelligence voice assistant.

[0023] To achieve the above-mentioned second object, a computer device provided by the present invention includes a processor and a memory, wherein: a computer program is stored on the memory, and when the computer program is executed by the processor, the above-mentioned method for extracting a target speaker based on historical conversation texts is implemented.

[0024] To achieve the above-mentioned third object, a computer-readable storage medium provided by the present invention has a computer program stored thereon, wherein: when the computer program is executed by a processor, the above-mentioned method for extracting a target speaker based on historical conversation texts is implemented.

[0025] To achieve the above-mentioned fourth object, a computer program product includes computer instructions, wherein: when the computer instructions are executed by a processor, the above-mentioned method for extracting a target speaker based on historical conversation texts is implemented. Description of the Drawings

[0026] Figure 1 It is a schematic diagram of the application of the present invention in the scenario of an artificial intelligence voice assistant.

[0027] Figure 2It is a flowchart of the first embodiment of the method for extracting the target speaker based on historical dialogue text of the present invention.

[0028] Figure 3 It is a structural diagram of the TPEech model in the first embodiment of the method for extracting the target speaker based on historical dialogue text of the present invention.

[0029] Figure 4 It is a comparison chart of the scores of the TPEech model and other models of the first embodiment of the method for extracting the target speaker based on historical dialogue text of the present invention under different metrics.

[0030] Figure 5 It is a schematic diagram of the results after the ablation experiment of the TPEech model in the first embodiment of the method for extracting the target speaker based on historical dialogue text of the present invention.

[0031] Figure 6 It is a schematic diagram of the difference in normalized metrics of the TPEech model in the first embodiment of the method for extracting the target speaker based on historical dialogue text of the present invention under different signal-to-noise ratios.

[0032] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. Specific embodiments

[0033] The method for extracting the target speaker based on historical dialogue text of the present invention takes the historical dialogue text as a clue and extracts the speech signal of the target speaker from the mixed speech signal in a scenario with multiple speakers and obvious background noise. The present invention also provides a computer device, a computer-readable storage medium, and a computer program product for implementing the above method for extracting the target speaker based on historical dialogue text.

[0034] See Figure 1 , Figure 1Illustrates the application of the present invention in the application scenario of an artificial intelligence voice assistant. In this application scenario, the user interacts with the AI voice assistant. At this time, the AI voice assistant asks the question "What actions did he take after these disagreements?", and the user answers according to the question "He happened to leave his name...". The AI voice assistant needs to receive the content of the user's answer for further feedback. However, since the user is in a noisy environment, the AI voice assistant will receive a mixed voice signal that includes a noise signal and the voice signal of an interfering speaker, and it is impossible to extract the user's answer from the received answer, that is, it is impossible to extract the voice signal of the target speaker. The present invention uses the question "What actions did he take after these disagreements?" proposed by the AI voice assistant as text prompt information to extract the target speaker. Since the question proposed by the AI voice assistant is related to the text content corresponding to the voice signal of the target speaker, while the noise and the voice signal of the interfering speaker are not related to the question proposed by the AI voice assistant, using the text prompt information as a clue, the voice signal of the target speaker can be extracted from the mixed voice signal. Thus, the AI voice assistant can respond based on the voice signal of the target speaker extracted by the present invention, enabling the AI voice assistant to still be reliably used in a noisy environment.

[0035] First Embodiment of the Method for Extracting the Target Speaker Based on the Historical Conversation Text:

[0036] Let s(τ) represent the voice signal of the target speaker, v i (τ) represent the interfering voice signal of the i-th interfering speaker, and n(τ) be the noise signal, where i ∈ {1, 2,..., I}, I is the total number of interfering speakers, and τ ∈ {1, 2,..., T}, τ represents the time index. Then the mixed voice signal x(τ) can be expressed as:

[0037] , where α and β are scalar variables used to adjust the signal-to-noise ratio (SNR) of the interfering voice signal v i (τ) and the noise signal n(τ), respectively.

[0038] This embodiment realizes the extraction of the voice signal s(τ) of the target speaker from the mixed voice signal x(τ) by executing a computer program. Refer to Figure 2 , which specifically includes the following steps:

[0039] S1: Obtain the mixed voice signal and the text prompt information.

[0040] Among them, the text prompt information is the question proposed by the current interactive AI voice assistant. For example, Figure 1 "What actions did he take after these disagreements" in Figure 1Waveform after superposition of medium noise signal, speech signal of interfering speaker, and speech signal of target speaker.

[0041] S2: Process the mixed speech signal and the text prompt information to obtain the speech signal of the target speaker.

[0042] Among them, the text content corresponding to the speech signal of the target speaker in the mixed speech signal is associated with the text prompt information. "Association" means that the text content corresponding to the speech signal of the target speaker is directly given in response to the text prompt information. For example, Figure 1 the text content corresponding to the speech signal of the target speaker in [example] "He happened to leave his name..." is given in response to the text prompt information "What actions did he take after these disagreements?"

[0043] S3: Output the speech signal of the target speaker.

[0044] The above steps S1 to S3 are implemented based on a parameter-trainable model, hereinafter referred to as the TPEech model.

[0045] See Figure 3 , the TPEech model includes a speech encoder 11, a text encoder 12, a fusion layer 13, a mask estimator 14, an echo prompt module 15, and a speech decoder 16.

[0046] The speech encoder 11 is used to convert the mixed speech signal x(τ) into a spectral feature embedding X(t). The speech encoder 11 processes the mixed speech signal x(τ) through a one-dimensional convolutional layer Conv1D, and then applies the activation function of the ReLU layer to obtain the spectral feature embedding X(t). The mixed speech signal x(τ) is converted into a spectral feature embedding X(t) in the latent space:

[0047] , where speech represents a series of operations applied to the input. The one-dimensional convolution is configured with input and output dimensions of 1 and D respectively. It uses a kernel of size L to capture local patterns in the data, and a stride of L / 2 can achieve overlapping receptive fields.

[0048] The text encoder 12 is used to convert the text prompt information pt into a text feature embedding P text . In this embodiment, the text encoder 12 generates a text feature embedding P corresponding to the speech through a Contrastive Language - Audio Pretraining (CLAP) model text :

[0049] ,

[0050] where text(·) represents the text encoder, and P text is a D-dimensional text embedding. Among them, the Contrastive Language-Audio Pretrained (CLAP) model does not participate in the training of the TPEech model, that is, all the parameters of the CLAP model are fixed, Figure 3 and the snowflake logo is used to indicate that all the parameters of this model are frozen.

[0051] In some embodiments, if the text prompt information pt is not presented in the form of speech but directly in the form of text, there is no need to convert the text prompt information pt into a text feature embedding P through the CLAP model text .

[0052] The fusion layer 13 is used to fuse the spectral feature embedding X(t) and the text feature embedding P text to generate a fused feature F(t). Layer normalization (LN) is applied to the spectral feature embedding X(t) to stabilize and normalize the speech embedding. In particular, the FiLM (Feature-wise Linear Modulation) layer modulation combines text and speech features, where the text feature embedding P text modulates the normalized spectral feature embedding X(t) through scaling and shifting. Since the text feature embedding P text is 512-dimensional and the spectral feature embedding X(t) is 256-dimensional, the text feature embedding P text is resampled through a resampling layer, and the final fused feature F(t) is expressed as:

[0053] , γ(P text ) and β(P text ) are learnable parameters, which scale and shift the spectral feature embedding X(t) according to the text feature embedding P text , and ⊙ represents element-wise multiplication. This process can be repeated multiple times to more deeply integrate text and speech, enabling the TPEech model to better incorporate text information into speech. In this embodiment, Figure 3 "×2" at the fusion layer 13 means repeating 2 times.

[0054] The mask estimator 14 is used to obtain the target mask M(t) corresponding to the fused feature F(t). The target mask M(t) is a mask used to filter out irrelevant noise and interference in the spectral feature embedding X(t). First, the fused feature F(t) passes through a one-dimensional convolutional layer (Conv1D), and then is segmented by a segmentation layer into segments with a window size of K and a stride size of K / 2. To minimize the total input length, K≈2T is set seq, and then all segments are concatenated to form block W. The DPRNN (Dual-Path Recurrent Neural Network) module is repeated R times, processing block W through bidirectional and cross-sequence processing, and introducing a residual connection mechanism, namely the echo hint module 15. The output of each DPRNN module is called dual-path feature, denoted as E dp . Finally, the Aggregation layer reconstructs block W back to the shape of F(t), and the subsequent PReLU activation layer generates the target mask M(t).

[0055] The echo hint module 15 includes multiple echo hint blocks ECB, which are used to connect to the mask estimator 14 to enhance the performance of target speaker extraction. By setting the echo hint module 15, the text feature embedding P text and the hierarchical integration of the spectral feature embedding X(t) enhance the output of each DPRNN block. Given that the outputs of different DPRNN blocks represent features of different granularities, this design improves the noise suppression ability. It should be noted that the text feature embedding P text and the spectral feature embedding X(t) are respectively output to each echo hint block ECB, that is Figure 3 the blue dashed line in indicates that each echo hint block ECB receives the spectral feature embedding X(t), and the red dashed line indicates that each echo hint block ECB receives the text feature embedding P text , and the text feature embedding P received by each echo hint block ECB text and the spectral feature embedding X(t) are independent of other echo hint blocks ECB.

[0056] The echo hint block ECB contains a cross-attention layer. In this layer, the query (Q ech ), key (K ech ), and value (V ech ) matrices are calculated by combining the text feature embedding P text and the spectral feature embedding X(t), and are weighted and calculated by their respective weight matrices W s and W t respectively. This cross-attention mechanism focuses on retaining the semantic content of the spectral feature embedding X(t), making it very suitable for filtering out semantically irrelevant noise. The dual-path feature E dp in the mask estimator 14 is enhanced by a residual connection, named H, where the attention weights are calculated by applying the softmax function to the scaled dot product of the query matrix and the key matrix. The overall design can be expressed as:

[0057] ,

[0058] where, W s represents the text feature embedding Ptext weight matrix, W t The weight matrix representing the spectral feature embedding X(t), where d represents the dimensions of the query, key, and value matrices.

[0059] The speech decoder 16 is used to reconstruct the speech signal s(τ) of the target speaker according to the target mask M(t). The speech decoder 16 outputs the masked speech embedding through element-wise multiplication between the target mask M(t) and the output of the spectral feature embedding X(t). :

[0060] ,

[0061] Then, the Linear layer maps the masked features from high dimension back to 1 dimension, facilitating the subsequent operations of the OnA layer. By performing the overlap-and-add operation (OnA), the masked speech embedding can be converted to the speech signal s(τ) of the target speaker.

[0062] In different embodiments, the above TPEech model can be integrated into an artificial intelligence voice intelligent assistant as a part of the artificial intelligence voice intelligent assistant, or can be independent of the artificial intelligence voice intelligent assistant and perform data transmission with the artificial intelligence voice intelligent assistant.

[0063] The process of training the TPEech model:

[0064] In terms of dataset setting, the self-designed MNSpeech dataset is adopted. The MNSpeech dataset includes 272.28 hours of audio data from 417 speakers. The composition of the audio data comes from the LibriHeavy-small dataset. Among them, each speech segment is padded or cropped to 8 seconds long.

[0065] Automatic Speech Recognition (ASR) is performed through the Whisper-large-v3 model to obtain the corresponding text for each speech segment. Then, the Mistral-7BInstruct-v0.2 model is used to generate questions based on the text obtained from ASR. The question answers generated by the Mistral-7B-Instructions-v0.2 model correspond to the text obtained from the ASR results. To better adapt to the real scenario, all speech segments are sampled to 8kHz.

[0066] See Figure 1 , the question is presented in the form of the text prompt information "What actions did he take after these disagreements?" This question is based on the ASR result of the target speech "He happened to leave his name..." Therefore, each speech segment corresponds to a uniquely matched text.

[0067] The noise components in the dataset come from six datasets: the MS-SNSD dataset, the MUSAN dataset, the Nonspeech dataset, the QUT-NOISE dataset, the UrbanSound 8k dataset, and the WHAM! dataset. All noise segments are resampled to a sampling rate of 8 kHz to match the speech segments. To generate overlapping speech, two samples are randomly selected from the MNSpeech dataset and mixed using the random SNR mixing method, with the SNR ranging from -5 to 10 decibels. Then, a noise sample is added, and its SNR value is also selected from the same range. This mixing process is dynamically generated during training.

[0068] The dataset is divided into a training set, a validation set, and a test set, containing 200,000, 5,000, and 1,000 samples respectively, with no overlap in speaker IDs between the sets, that is, the training set, the validation set, and the test set do not contain the voices of duplicate individuals.

[0069] For the implementation device, the original DPRNN module design is adopted, with the input channel set to 256, the output channel set to 64, the hidden channel set to 128, and a total of six modules. Different from the original kernel size of 2, a kernel size of 40 is used. According to the discussion in the ablation study section, we empirically selected 3 echo hint modules.

[0070] For the evaluation and training metrics, the signal-to-noise ratio (SDR), the perceptual evaluation of speech quality (PESQ), and the short-time objective intelligibility (STOI) are used as metrics. For each metric, the higher the value, the better the performance (↑). The Adam optimizer is adopted, with an initial learning rate of 5×10 −4 。

[0071] The negative-scale inverse transformation signal distortion ratio improvement (SI-SDRi) is adopted as the loss function to evaluate the signal quality of the speech signal for extracting the target speaker. SI-SDRi is an improved version of the scale inverse transformation signal distortion ratio (SI-SDR):

[0072] ,

[0073] ,

[0074] For the sake of simplicity, (τ) is omitted in the loss function.

[0075] See Figure 4 , Figure 4The figure shows the scores of the TPEech model compared to various SE models on different metrics. SE models such as the MP-SENet model and the FRCRN model demonstrated the ability to reduce background noise and partially suppress the voices of non-primary speakers, achieving scores of 2.076 dB and 2.135 dB respectively for the SI-SDRi metric. However, in cases where the target speaker is not the dominant speaker, especially under low signal-to-noise ratio mixing conditions, their performance significantly degrades. On the other hand, the SGMSE model performed poorly in handling multi-speaker scenarios, achieving only a score of 1.721 dB for the SI-SDRi metric, slightly inferior to the above models.

[0076] Figure 4 Multiple speech separation methods were also evaluated. In this context, PIT association involves calculating and manually selecting the optimal speech output according to the SI-SDR metric to align with the speech of the target speaker. Models such as the DPTNet model and the SPGM model performed poorly in effectively separating speech in noisy environments, achieving scores of 3.554 dB and 4.323 dB respectively for the SI-SDRi metric. On the other hand, the SepFormer model and the MossFormer2 model performed well in handling noisy speech. However, during the speech separation process, although the first output can effectively remove noise, the second output often retains residual noise. When using PIT association, the target speech may correspond to the second output, resulting in a decrease in overall average performance. This problem limits their effectiveness, as can be seen from their relatively low SI-SDRi metric scores of 6.114 dB and 6.292 dB. Figure 4 The TPEech model of this embodiment was also compared with another TSE model, AudioSep-CLAP, which was fine-tuned using the MNSpeech dataset of this embodiment. The AudioSep-CLAP model is also an audio extraction model that uses text as a cue. However, it focuses more on audio rather than speech, making it less effective in accurately removing background noise and extracting the speech of the target speaker. Therefore, the AudioSep-CLAP model is inferior to the TPEech model of this embodiment in terms of performance, with a score 7.969 dB lower for the SI-SDRi metric.

[0077] See Figure 5 , Figure 5The ablation experiment is shown. The ablation experiment reveals the number of echo prompt blocks and the necessity of text prompt embedding. The results include: (1) When using 6 echo prompt blocks, compared with the complete model, the performance of the TPEech model slightly decreases, with a decrease of 0.272 dB in the SI-SDRi metric score. (2) When using 2 echo prompt blocks, the performance degradation is more obvious, with a decrease of 1.09 dB in the SI-SDRi score. (3) When not using echo prompt blocks, the performance significantly decreases, with a reduction of 2.961 dB in the SI-SDRi score. When using DPRNN (PIT association), TPEech performs better than DPRNN itself when using 3 echo prompt blocks (i.e., the complete model), with an improvement of 0.733 dB in the SI-SDRi metric. (4) When deleting the text prompt information, the TPEech model is completely unable to extract the target speech, resulting in an SI-SDRi metric score of -6.227 dB.

[0078] See Figure 6 , Figure 6 reveals the impact of SNR mixing and SNR noise on the TSE performance. The experiment is conducted under two settings using 500 fixed combinations of speakers, interfering sources, and noises. The results show that when the signal-to-noise ratio is negative, the impact of noise on all metrics is greater than that of interfering speech at the same signal-to-noise ratio level. On the contrary, when the signal-to-noise ratio is positive, the impact of interfering speech exceeds that of noise. This indicates that in extreme cases, the impact of noise is greater.

[0079] Second Embodiment of the Target Speaker Extraction Method Based on Historical Conversation Texts:

[0080] The difference between this embodiment and the above first embodiment is that this embodiment is applied to the voice call scenario. In the voice call scenario, there are a first interlocutor and a second interlocutor who need to converse with each other, and the first interlocutor and the second interlocutor make a call through an existing communication terminal. The communication terminal can be a smartphone, a tablet computer, or a smartwatch. The TPEech model of the above first embodiment runs in the communication terminal of the first interlocutor and / or the second interlocutor.

[0081] At this time, the text prompt information is the current speech content of one interlocutor collected, and the mixed voice signal is the reply of the communication terminal of the other interlocutor in response to the current speech content. Thus, one interlocutor can use the TPEech model to extract the voice signal of the target speaker from the reply of the other interlocutor, thereby improving the voice call quality.

[0082] For example, assume that the first interlocutor and the second interlocutor are having a voice call in a strong noise environment. During this process, the first interlocutor says, "Where are you going this afternoon?" At this time, "Where are you going this afternoon?" is the extracted text prompt information. The second interlocutor says, "I'm going to location A to watch a show this afternoon." "I'm going to location A to watch a show this afternoon" is the text content corresponding to the voice signal of the target speaker. It can be seen that this text content is semantically related to the text prompt information. The TPEech model running on the communication terminal of the first interlocutor processes and combines the extracted text prompt information (corresponding to "Where are you going this afternoon?"), as well as the mixed voice signal obtained from the second interlocutor, and outputs the voice signal of the second interlocutor. At this time, the mixed voice signal includes the voice signal of the second interlocutor (the corresponding text content is "I'm going to location A to watch a show this afternoon"), the ambient noise around the second interlocutor, and the voice signals of interfering speakers around the second interlocutor.

[0083] Then, the first interlocutor says, "What a coincidence. I'm also going to location A to watch a show this afternoon." Similarly, at this time, "I'm going to location A to watch a show this afternoon" is the extracted text prompt information, and "What a coincidence. I'm also going to location A to watch a show this afternoon" is the text content corresponding to the voice signal of the target speaker. The TPEech model running on the communication terminal of the second interlocutor combines the extracted text prompt information (corresponding to "I'm going to location A to watch a show this afternoon"), as well as the mixed voice signal obtained from the call terminal of the first interlocutor, and outputs the voice signal of the first interlocutor. At this time, the mixed voice signal includes the voice signal of the first interlocutor (the corresponding text content is "What a coincidence. I'm also going to location A to watch a show this afternoon"), the ambient noise around the first interlocutor, and the voice signals of interfering speakers around the first interlocutor.

[0084] In summary, the present invention processes the voice conversation scenario of overlapping voices and strong background noise, uses the historical conversation text as a clue to explore the semantic relationship between the question and the human answer, and uses the prior knowledge of context prompts as a clue to effectively extract the voice of the target speaker. In addition, by designing an echo prompt module, low signal-to-noise ratio noise is effectively removed at multiple granularity levels. The present invention improves the effect of extracting the voice signal of the target recognition speaker, is easy to implement, and has a wide application range.

[0085] Embodiment of computer device:

[0086] The computer device of this embodiment includes a processor and a memory. The memory stores a computer program. When the processor executes the computer program, it implements the above-mentioned embodiment of the method for extracting the target speaker based on historical conversation text.

[0087] A computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that a computer device may include more or fewer components, or combine certain components, or different components. For example, a computer device may also include input / output devices, network access devices, buses, etc.

[0088] For example, the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microcontroller or the processor may also be any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines.

[0089] The memory can be used to store computer programs and / or modules. The controller realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. For example, the memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a voice reception function, a voice-to-text conversion function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, text data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0090] Examples of computer-readable storage media:

[0091] If the modules integrated in the computer device of the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the process of implementing the embodiment of the method for extracting the target speaker based on historical conversation texts can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a controller, the steps of the above embodiment of the method for extracting the target speaker based on historical conversation texts can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0092] Embodiment of computer program product:

[0093] The computer program product of this embodiment includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes each step of the above embodiment of the method for extracting the target speaker based on historical conversation texts.

[0094] Finally, it should be emphasized that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A target speaker extraction method based on historical dialogue text, characterized in that: The following steps are involved: Processing a mixed speech signal through a speech encoder to obtain spectral feature embedding; wherein the mixed speech signal includes a speech signal of a target speaker, a speech signal of an interfering speaker, and a noise signal; Processing the text prompt information through a text encoder to obtain text feature embedding; associating the text content corresponding to the speech signal of the target speaker with the text prompt information; fusing the spectral feature embedding and the text feature embedding through a fusion layer to obtain a fusion feature; Processing the fused features through a mask estimator to obtain a target mask corresponding to the fused features; Extracting a speech signal of the target speaker according to the target mask and the mixed speech signal; The mask estimator includes a plurality of DPRNN modules. When the fusion feature is processed by the mask estimator, the mask estimator includes: The fusion feature is divided into segments, all segments are connected to form blocks, and the blocks are processed by multiple echo prompt blocks of the multiple DPRNN modules and the echo prompt module in turn, and finally the blocks are reconstructed by the aggregation layer, and the target mask is generated by the activation layer; wherein the output of each DPRNN module is called a dual path feature, denoted by E dp ; The text feature embedding and the spectral feature embedding are output to each of the echo prompt blocks respectively, and the text feature embedding and the spectral feature embedding received by each echo prompt block are independent of other echo prompt blocks; The echo prompt block includes a cross attention layer, in which the query (Q ech ), key (K ech ) and value (V ech ) matrix is ​​obtained by combining the text feature embedding and the spectral feature embedding, and respectively by their respective weight matrices and weighted calculations; the dual path feature E in the mask estimator dp Enhanced by residual connection, the connection is named H and is expressed as , where d represents the dimensions of the query, key, and value matrices.

2. The target speaker extraction method based on historical dialogue text according to claim 1, characterized in that: The number of the echo prompt blocks is 3.

3. The target speaker extraction method based on historical conversation text according to claim 1, characterized in that: When the mixed speech signal is processed by the speech encoder to obtain the spectral feature embedding, the method includes: The mixed speech signal is processed through a one-dimensional convolutional layer and then a ReLU layer is applied to obtain the spectral feature embedding.

4. The target speaker extraction method based on historical dialogue text according to claim 1, characterized in that: When the text prompt information is processed by the text encoder to obtain the text feature embedding, the method includes: The text feature embedding corresponding to the text prompt information is generated by comparing the language audio pre-training model.

5. The target speaker extraction method based on historical dialogue text according to claim 1, characterized in that: The fusion layer, the mask estimator and the echo prompt module are obtained through a deep learning training method, including: setting a data set, the data set includes multiple voice segments and multiple noise segments, each of the voice segments is provided with a corresponding unique matching text generation problem, and during training, two of the voice segments and one of the noise segments are randomly selected and mixed using a random SNR mixing method.

6. The method for extracting a target speaker based on historical conversation text according to claim 5, characterized in that: The speech segment and the noise segment are both sampled to 8 kHz.

7. A computer device comprising a processor and a memory, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the method for extracting a target speaker based on historical conversation text according to any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for extracting a target speaker based on historical conversation text as described in any one of claims 1 to 6 is implemented.

9. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by the processor, the method for extracting a target speaker based on historical conversation text as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Voice interaction method and device, robot, electronic equipment and storage medium

    CN118782027A

  • Speaker extraction method and system

    CN118865940A