Speech extraction method, apparatus, device, and medium

CN122531395APending Publication Date: 2026-08-07PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-05-28
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]本发明提供一种语音提取方法、装置、计算机设备及介质,以解决如何提升语音提取准确性的技术问题

Benefits of technology

[0008]上述语音提取方法、装置、设备及介质所实现的方案中,可以通过获取目标用户的参考语音和混合语音,并分别进行时频变换,得到参考语音频谱特征和混合语音频谱特征;进一步获取混合语音的语音场景信息,并根据语音场景信息对混合语音频谱特征进行噪声抑制处理,得到噪声抑制频谱特征;再对噪声抑制频谱特征和参考语音频谱特征进行特征交互,得到用户语音引导特征,并根据用户语音引导特征进行语音频谱提取,得到目标用户语音频谱信息,最终通过逆时频变换得到目标用户语音。在本发明中,通过引入语音场景信息对混合语音频谱特征进行噪声抑制处理,能够降低环境噪声及干扰说话人对目标用户语音定位过程的影响,并通过参考语音频谱特征与噪声抑制频谱特征之间的特征交互,提升目标用户语音内容定位的准确性,从而有效克服现有技术在含噪场景下语音提取准确性较低的缺陷,显著提升了在复杂含噪场景下目标语音提取的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531395A_ABST
    Figure CN122531395A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses a voice extraction method, device and equipment and a medium, and belongs to the technical field of artificial intelligence, and is applied to financial and medical technology scenes, and comprises the following steps: obtaining noise-free reference voice of a target user and mixed voice containing voice of the target user and noise; performing noise suppression processing on spectral features of the mixed voice according to voice scene information; performing feature interaction according to the noise-suppressed spectral features after the suppression processing and spectral features of the reference voice to obtain user voice guide features; and determining voice spectral information of the target user according to the user voice guide features to extract the voice of the target user. In the application, the voice scene information is introduced to perform noise suppression processing on the spectral features of the mixed voice, the feature interaction between the spectral features of the reference voice and the noise-suppressed spectral features is combined, the voice of the target user is extracted through the voice guide features after the interaction, and the accuracy of target voice extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, applicable to fintech and medical technology scenarios, and particularly to a speech extraction method, apparatus, device, and medium. Background Technology

[0002] Currently, traditional speech extraction methods primarily use neural network models (such as VoiceFilter) to jointly model the reference speech of the target speaker and the mixed speech containing the target speaker's speech to be identified, in order to extract the target speaker's speech from the mixed speech. For example, in financial voice verification scenarios, if the current call recording contains the customer's own voice, customer service voice, and environmental noise, the VoiceFilter model can extract the customer's historical voice embedding features and extract the recording spectrum features of the current call recording. Based on the customer's voice embedding features, the customer's own voice can be extracted from the recording spectrum features. In medical telemedicine scenarios, if the consultation recording contains the patient's voice, the accompanying person's voice, and environmental noise, the patient's pre-reserved voice and the current consultation recording can be input into the VoiceFilter model to extract the patient's own voice. However, this method directly extracts the target speech from the mixed speech, which does not fully utilize the speech context information. This can easily lead to noise interference in guiding the target user's speech from the reference speech, resulting in low accuracy of speech extraction in noisy scenarios. Therefore, improving the accuracy of speech extraction has become an urgent problem to be solved. Summary of the Invention

[0003] This invention provides a speech extraction method, apparatus, computer equipment, and medium to address the technical problem of improving the accuracy of speech extraction.

[0004] Firstly, a speech extraction method is provided, including: Acquire a reference speech of the target user and acquire a mixed speech; wherein the reference speech is noise-free, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user; The reference speech is subjected to time-frequency transformation by a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and the mixed speech is subjected to time-frequency transformation to obtain the spectral features of the mixed speech. Acquire the speech scene information of the mixed speech, and perform noise suppression processing on the spectral features of the mixed speech based on the speech scene information to obtain noise-suppressed spectral features; The noise suppression spectral features and the reference speech spectral features are interacted to obtain user speech guidance features; wherein, the user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech; Based on the user's voice guidance features, voice spectrum extraction is performed to obtain the target user's voice spectrum information; The target user's speech spectrum information is subjected to inverse time-frequency transformation to obtain the target user's speech.

[0005] Secondly, a speech extraction device is provided, comprising: A reference speech and mixed speech acquisition module is used to acquire a reference speech of a target user and acquire a mixed speech; wherein, the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user; The time-frequency transformation module is used to perform time-frequency transformation on the reference speech using a pre-trained target speech extraction model to obtain the reference speech spectral features, and to perform time-frequency transformation on the mixed speech to obtain the mixed speech spectral features. The noise suppression processing module is used to acquire the speech scene information of the mixed speech, and perform noise suppression processing on the spectral features of the mixed speech according to the speech scene information to obtain noise suppression spectral features; The feature interaction module is used to perform feature interaction between the noise suppression spectral features and the reference speech spectral features to obtain user speech guidance features; wherein, the user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech; The speech spectrum extraction module is used to extract the speech spectrum based on the user's speech guidance features to obtain the target user's speech spectrum information. The speech extraction module is used to perform inverse time-frequency transformation on the speech spectrum information of the target user to obtain the speech of the target user.

[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described target speech extraction method.

[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described target speech extraction method.

[0008] The aforementioned speech extraction method, apparatus, device, and medium can acquire the target user's reference speech and mixed speech, and perform time-frequency transformation on them respectively to obtain the reference speech spectral features and mixed speech spectral features. Further, speech scene information of the mixed speech is acquired, and noise suppression processing is applied to the mixed speech spectral features based on the speech scene information to obtain noise-suppressed spectral features. Then, feature interaction is performed between the noise-suppressed spectral features and the reference speech spectral features to obtain user speech guidance features. Speech spectrum extraction is then performed based on the user speech guidance features to obtain the target user's speech spectrum information. Finally, the target user's speech is obtained through inverse time-frequency transformation. In this invention, by introducing speech scene information to perform noise suppression processing on the mixed speech spectral features, the influence of environmental noise and interfering speakers on the target user's speech localization process can be reduced. Furthermore, through feature interaction between the reference speech spectral features and the noise-suppressed spectral features, the accuracy of target user speech content localization is improved. This effectively overcomes the shortcomings of existing technologies in terms of low speech extraction accuracy in noisy scenes and significantly improves the accuracy of target speech extraction in complex noisy scenes. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of an application environment for the speech extraction method in one embodiment of the present invention; Figure 2 This is a schematic flowchart of a speech extraction method according to an embodiment of the present invention; Figure 3 This is another flowchart illustrating the speech extraction method in one embodiment of the present invention; Figure 4 yes Figure 3 A flowchart illustrating another specific implementation of step S207; Figure 5 yes Figure 2 A schematic diagram of a specific implementation method for step S102; Figure 6 yes Figure 2 A schematic diagram of a specific implementation method for step S103; Figure 7 yes Figure 2 A schematic diagram of a specific implementation of step S104; Figure 8 yes Figure 2A schematic diagram of a specific implementation of step S105; Figure 9 This is a schematic diagram of the structure of a speech extraction device in one embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 11 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] The speech extraction method provided in this application relates to the field of artificial intelligence technology.

[0012] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0013] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0014] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] The speech extraction method provided in this embodiment of the invention can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. The server can acquire the target user's reference speech and mixed speech, and perform time-frequency transformation on them respectively to obtain the reference speech spectral features and mixed speech spectral features. It further acquires the speech scene information of the mixed speech, and performs noise suppression processing on the mixed speech spectral features based on the speech scene information to obtain noise-suppressed spectral features. Then, it performs feature interaction between the noise-suppressed spectral features and the reference speech spectral features to obtain user speech guidance features, and extracts the speech spectrum based on the user speech guidance features to obtain the target user's speech spectrum information. Finally, it obtains the target user's speech through inverse time-frequency transformation. In this invention, by introducing speech scene information to perform noise suppression processing on the mixed speech spectral features, the impact of environmental noise and interfering speakers on the target user's speech localization process can be reduced. Furthermore, through feature interaction between the reference speech spectral features and the noise-suppressed spectral features, the accuracy of target user speech content localization is improved. This effectively overcomes the shortcomings of existing technologies in terms of low speech extraction accuracy in noisy scenes, and significantly improves the accuracy of target speech extraction in complex noisy scenes. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0017] like Figure 2 As shown, Figure 2 A flowchart illustrating the speech extraction method provided in this embodiment of the invention may include, but is not limited to, steps S101 to S106: Step S101: Obtain the reference speech of the target user and obtain the mixed speech; wherein the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user.

[0018] Step S102: The reference speech is subjected to time-frequency transformation by a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and the mixed speech is subjected to time-frequency transformation to obtain the spectral features of the mixed speech.

[0019] Step S103: Obtain speech scene information of mixed speech, and perform noise suppression processing on the spectral features of mixed speech based on the speech scene information to obtain noise-suppressed spectral features.

[0020] Step S104: Perform feature interaction between the noise suppression spectral features and the reference speech spectral features to obtain user speech guidance features; wherein, the user speech guidance features are used to guide the target speech extraction model to locate the target user's speech content from the mixed speech.

[0021] Step S105: Extract the speech spectrum based on the user's voice guidance features to obtain the target user's speech spectrum information.

[0022] Step S106: Perform inverse time-frequency transformation on the target user's speech spectrum information to obtain the target user's speech.

[0023] Steps S101 to S106 of this embodiment involve acquiring the target user's reference speech and mixed speech, performing time-frequency transformation on each, and obtaining the reference speech spectral features and mixed speech spectral features. Further, speech scene information of the mixed speech is acquired, and noise suppression processing is applied to the mixed speech spectral features based on the speech scene information to obtain noise-suppressed spectral features. Then, feature interaction is performed between the noise-suppressed spectral features and the reference speech spectral features to obtain user speech guidance features. Speech spectrum extraction is performed based on the user speech guidance features to obtain the target user's speech spectrum information. Finally, the target user's speech is obtained through inverse time-frequency transformation. In this invention, by introducing speech scene information to perform noise suppression processing on the mixed speech spectral features, the impact of environmental noise and interfering speakers on the target user's speech localization process can be reduced. Furthermore, through feature interaction between the reference speech spectral features and the noise-suppressed spectral features, the accuracy of target user speech content localization is improved. This effectively overcomes the shortcomings of existing technologies in terms of low speech extraction accuracy in noisy scenes and significantly improves the accuracy of target speech extraction in complex noisy scenes.

[0024] In step S101 of some embodiments, specifically, the reference speech refers to a clean speech signal of the target user that is free from background noise and interference from the speaker, and the speaker of the reference speech is the target user, and the speech is used to provide the user's voiceprint identity features.

[0025] Specifically, mixed speech refers to the speech signal to be processed that is obtained in a real business environment and includes the speech of the target user, the speech of non-target speakers, and background noise.

[0026] For example, in a financial context, the reference voice can be a pure voice segment of the customer in a historical call recording stored during the business transaction, while the mixed voice can be a complete voice signal in the current telephone customer service call that includes the customer, customer service personnel, and background keyboard sounds; in a medical telemedicine scenario, the reference voice can be a voice sample stored during the patient's first visit, while the mixed voice can be a comprehensive audio recording of the current consultation that includes the patient, accompanying family members, consulting doctor, and medical equipment prompts.

[0027] Specifically, retained voice recordings from target user registrations, historical identity authentication voice recordings, or pre-collected voice samples can be obtained from the business backend to determine reference voice recordings. Further mixed voice recordings can be obtained from real-time call streams, audio recordings, or online voice interaction data.

[0028] In this embodiment, by acquiring noise-free target user reference speech and mixed speech containing noise and the speaker of the speech containing the target user, the acoustic identity of the target user and the mixed speech containing the target user's speech in the complex acoustic scene to be separated are clarified, providing the necessary data foundation for accurate positioning of the target user's speech in the noisy environment.

[0029] like Figure 3 As shown, in some embodiments, the speech extraction method further includes, but is not limited to, steps S201 to S208: Step S201: Obtain the target training set; wherein, the target training set includes training mixed speech generated for the same target user under at least two different noise background conditions, the user's real speech corresponding to each training mixed speech, the target user's training reference speech, and the training scene information corresponding to each training mixed speech.

[0030] Step S202: The training reference speech is subjected to time-frequency transformation using a preset original speech extraction model to obtain the spectral features of the training reference speech. The training mixed speech under various noise background conditions is also subjected to time-frequency transformation to obtain the spectral features of the training mixed speech.

[0031] Step S203: Perform noise suppression processing on the spectral features of each training mixed speech according to the training scenario information to obtain the training noise-suppressed spectral features.

[0032] Step S204: Perform feature interaction between the training noise suppression spectral features and the training reference speech spectral features to obtain the training user speech guidance features.

[0033] Step S205: Extract the speech spectrum based on the speech guidance features of the training user to obtain the speech spectrum information of the training user.

[0034] Step S206: Perform inverse time-frequency transformation on the training user speech spectrum information to obtain the predicted user speech under various noise background conditions.

[0035] Step S207: Calculate the loss value based on the predicted user speech and the user's actual speech under various noise background conditions to obtain the target loss value.

[0036] Step S208: Train the original speech extraction model based on the target loss value to obtain the target speech extraction model.

[0037] In step S201 of some embodiments, specifically, the target training set refers to the sample set used to train the original speech extraction model. The target training set includes training mixed speech generated for the same target user under at least two different noise background conditions, real user speech corresponding to each training mixed speech, training reference speech of the target user, and training scene information corresponding to each training mixed speech. Among them, the at least two different noise background conditions may include any two or three of the following: single speaker plus noise condition, two speaker condition, and two speaker plus noise condition. The training mixed speech refers to the speech signal generated for the same target user under at least two different noise background conditions during the training phase. The real user speech refers to the real clean speech corresponding to the target user in the training mixed speech. The training reference speech refers to the reference speech used to characterize the speech features of the target user during the training phase. The training scene information refers to the information reflecting the noise background conditions, interfering speaker conditions, or environment types of each training mixed speech in different business scenarios.

[0038] For example, in a financial scenario, the target training set can include training samples formed by different combinations of the customer's own voice, insurance customer service voice, and environmental noise in an insurance business hall. The user's real voice can be the customer's clean voice, and the training reference voice can be the customer's reserved voice. The training scenario information can correspond to scenario data such as insurance customer service calls and background noise in an insurance business hall. In a medical scenario, the target training set can include training samples formed by different combinations of the patient's own voice, doctor's voice, accompanying person's voice, and noise from medical equipment. The user's real voice can be the patient's clean voice, and the training reference voice can be the patient's historical reserved voice. The training scenario information can correspond to scenario data such as consultation in a clinic, remote consultation, and interference from accompanying persons.

[0039] In step S202 of some embodiments, specifically, the original speech extraction model refers to an untrained neural network model, which includes a time-frequency transformation network, a lightweight noise suppression network, a feature interaction network, a speech spectrum extraction network, an inverse time-frequency transformation network, and a loss optimization network. Specifically, the time-frequency transformation network performs time-frequency transformation on the training reference speech and the training mixed speech to obtain the spectral features of the training reference speech and the spectral features of the training mixed speech, respectively. The lightweight noise suppression network performs noise suppression processing on the spectral features of the training mixed speech based on training scene information to obtain the noise-suppressed spectral features of the training. The feature interaction network performs feature interaction between the noise-suppressed spectral features and the spectral features of the training reference speech to obtain the training user speech guidance features. The speech spectrum extraction network extracts the speech spectrum based on the training user speech guidance features to obtain the training user speech spectrum information. The inverse time-frequency transformation network performs inverse time-frequency transformation on the training user speech spectrum information to obtain the predicted user speech. The loss optimization network calculates the loss value based on the predicted user speech and the user's actual speech under various noise background conditions, and updates the parameters of the original speech extraction model based on the target loss value to obtain the target speech extraction model.

[0040] Specifically, the spectral features of the training reference speech refer to the spectral representation of the training reference speech after time-frequency transformation.

[0041] Specifically, the training mixed speech spectral features refer to the spectral representation obtained by time-frequency transformation of the training mixed speech for the same target user under various noise background conditions. The spectral features can also be complex spectral features, and the feature representation is formed by concatenating the real and imaginary parts. In a further embodiment, the spectral amplitude can also be compressed to reduce the impact of high dynamic range fluctuations on subsequent feature interactions.

[0042] For example, a short-time Fourier transform can be performed on the training reference speech of the target customer to obtain the spectral features of the customer's training reference speech. A short-time Fourier transform can also be performed on the training mixed speech of the customer's own speech, customer service speech, and insurance business hall noise to obtain the corresponding training mixed speech spectral features. In medical application scenarios, a short-time Fourier transform can be performed on the training reference speech of the patient to obtain the spectral features of the patient's training reference speech. A short-time Fourier transform can also be performed on the training mixed speech of the patient's speech, doctor's speech, and medical testing equipment noise to obtain the corresponding training mixed speech spectral features.

[0043] In step S203 of some embodiments, specifically, training noise suppression spectral features refers to the spectral representation obtained after training the mixed speech spectral features and processing them through a noise suppression network, which is used to reduce the influence of environmental noise and non-target interference speakers on the extraction of the target user's speech.

[0044] Specifically, a lightweight noise suppression network can be a gated temporal convolutional recurrent network (GTCRN). The GTCRN network can suppress noise components and non-target speech components in the training mixed speech spectral features based on the training scenario information. Furthermore, under noise-free or low-noise conditions, the GTCRN network can also output the training mixed speech spectral features with only minor adjustments to avoid unnecessary damage to the effective speech components of the target user.

[0045] For example, in a financial scenario, if the training scenario information represents the current training mixed speech as including the customer's own voice, the insurance customer service voice, and background noise from the insurance branch, the GTCRN network can be used to perform noise suppression processing on the spectral features of the training mixed speech under this training scenario information. This reduces the background noise from the insurance branch and decreases the interference of the insurance customer service voice on the extraction of the customer's own voice. In a medical scenario, if the training scenario information represents the current training mixed speech as including the patient's own voice, the doctor's voice, and noise from the diagnostic equipment, the GTCRN network can be used to perform noise suppression processing on the spectral features of the training mixed speech under this training scenario information. This reduces equipment noise and the voice of non-target speakers, improving the discriminability of the spectral components of the patient's own voice.

[0046] In step S204 of some embodiments, specifically, training user speech guidance features refers to a guidance feature representation that integrates training reference speech spectrum features and training noise suppression spectrum features, which is used to guide the speech spectrum extraction network to locate the target user speech content in the training mixed speech.

[0047] Specifically, the feature interaction network can be implemented using a cross-attention network based on an attention mechanism. By using the training reference speech spectrum features as prior information of the target user and calculating the correlation with the training noise suppression spectrum features, the training user speech guidance features corresponding to the target user's speech are obtained.

[0048] In this embodiment, by inputting the noise suppression spectrum features that reduce noise pollution into the cross-attention network to perform feature fusion between the training reference speech spectrum features and the noise suppression spectrum features, the training user speech guidance features are obtained, which can ensure that the entire speech extraction link remains structurally consistent.

[0049] For example, in financial scenarios, the customer's training reference speech spectrum features and the target customer's training mixed speech spectrum features can be input into a cross-attention network for feature fusion to obtain the customer's training user voice guidance features, thereby highlighting the time-frequency region consistent with the customer's own voice. In medical scenarios, the patient's training reference speech spectrum features and the patient's training mixed speech spectrum features can be input into a cross-attention network for feature fusion to obtain the patient's training user voice guidance features, thereby strengthening the time-frequency region corresponding to the patient's own voice.

[0050] In step S205 of some embodiments, specifically, training user speech spectrum information refers to target user speech spectrum data recovered from the spectrum representation corresponding to the training mixed speech.

[0051] Specifically, the speech spectrum extraction network can be a target speech spectrum extraction network based on a convolutional encoder and a temporal modeling layer to decode and recover the speech guidance features of the training user, so as to output the speech spectrum information of the training user. Specifically, the speech spectrum extraction network can encode the speech guidance features of the training user into a time-frequency mask to obtain the time-frequency mask of the target user, and then use the time-frequency mask of the target user to decode the training noise suppression spectrum features to obtain the speech spectrum information of the training user.

[0052] In step S206 of some embodiments, specifically, predicting user speech refers to the time-domain target user speech signal recovered by inverse time-frequency transformation of the trained user speech guidance features.

[0053] Specifically, the inverse time-frequency transform network can be an inverse short-time Fourier transform algorithm used to recover the spectral information of the training user's speech into a time-domain waveform.

[0054] For example, in a financial scenario, the customer's training user speech spectrum information obtained under conditions of customer's own voice and insurance business hall noise can be restored to the first predicted customer speech through inverse short-time Fourier transform, and the customer's training user speech spectrum information obtained under conditions of customer's own voice, insurance customer service voice, and insurance business hall noise can be restored to the second predicted customer speech; in a medical scenario, the patient's training user speech spectrum information obtained under conditions of patient's own voice and medical equipment noise can be restored to the first predicted patient speech, and the patient's training user speech spectrum information obtained under conditions of patient's own voice, doctor's voice, and medical equipment noise can be restored to the second predicted patient speech.

[0055] like Figure 4 As shown, in some embodiments, step S207 includes, but is not limited to, steps S301 to S303: Step S301: Calculate the speech loss based on the predicted user speech and the actual user speech under each noise background condition to obtain the user speech loss value.

[0056] Step S302: Perform cross-conditional loss calculation on the predicted user speech under various noise background conditions to obtain the cross-conditional speech loss value.

[0057] Step S303: Calculate the loss based on the user's speech loss value and the cross-conditional speech loss value to obtain the target loss value.

[0058] In step S301 of some embodiments, specifically, the user speech loss value is used to characterize the degree of difference between the predicted user speech and the user's actual speech under various noise background conditions.

[0059] Specifically, for the predicted user speech obtained under different noise background conditions for the same target user, the scale-invariant signal distortion ratio loss between each predicted user speech and the user's real speech can be calculated separately, and the loss terms can be summed to obtain the user speech loss value.

[0060] Specifically, the user's speech loss value can be calculated using the following formula:

[0061] in, Indicates the user's voice loss value. Let s represent the predicted user speech under the i-th noisy background condition, and s represent the user's actual speech. Represents the i-th predicted user voice The scale-invariant signal distortion ratio between s and the user's actual speech.

[0062] Furthermore, since a larger scale-invariant signal-to-distortion ratio indicates that the predicted user speech is closer to the user's actual speech, a negative value is taken as the loss term during the training of the original speech extraction model, so as to optimize the model parameters by minimizing the loss.

[0063] For example, if the same customer's voice is predicted under the condition of "customer's own voice plus background noise in the business hall", and the customer's voice is predicted under the condition of "customer's own voice plus customer service voice plus background noise in the business hall", then the scale-invariant signal distortion ratio between the first predicted customer's voice and the customer's actual voice, and the scale-invariant signal distortion ratio between the second predicted customer's voice and the customer's actual voice can be calculated separately. Then, the summation and negation can be performed according to the above formula to obtain the user voice loss value corresponding to the customer.

[0064] In step S302 of some embodiments, specifically, the cross-conditional speech loss value is used to characterize the degree of consistency between the predicted user speech output by the same target user under different noise background conditions.

[0065] Specifically, the difference between predicted user speech obtained under different noise background conditions can be measured at time-domain sampling points or frame-level representations to calculate the degree of deviation between the output results of the same target user under different noise background conditions.

[0066] Specifically, the cross-conditional speech loss value can be calculated using the following formula:

[0067] in, represents the cross-conditional speech loss value, w represents the weight of the cross-conditional speech loss term, and T represents the total number of sampling points or total time steps for predicting user speech. This represents the predicted user speech amplitude at time step t under the first noise background condition. This represents the predicted user speech amplitude at time step t under the second noise background condition.

[0068] Furthermore, the aforementioned cross-condition loss calculation method enables the model to focus not only on the speech recovery quality under a single noise background condition, but also on the consistency of the output results of the same target user under multiple noise background conditions, thereby enhancing the robustness of the original speech extraction model.

[0069] In step S303 of some embodiments, specifically, the target loss value refers to the total loss value of the model that integrates the user speech loss value and the cross-conditional speech loss value.

[0070] Specifically, the target loss value can be calculated using the following formula:

[0071] in, Indicates the target loss value. Indicates the user's voice loss value. This represents the cross-conditional speech loss value.

[0072] By calculating the user speech loss value and the cross-conditional speech loss value separately through steps S301 to S303, and summing the two to obtain the target loss value, the target speech extraction model can take into account both the speech recovery accuracy under a single noise background condition and the output consistency under different noise background conditions during the training process. Furthermore, by introducing cross-conditional speech loss constraints, the robustness of the target speech extraction model in complex noisy, multi-speaker environments such as financial and medical scenarios and the accuracy of target user speech extraction can be further improved.

[0073] In step S208 of some embodiments, specifically, the target speech extraction model refers to a neural network model that, after being trained on a target training set, is capable of extracting the target user's speech based on reference speech and mixed speech.

[0074] Specifically, the backpropagation algorithm can be used to jointly update the lightweight noise suppression network, feature interaction network, speech spectrum extraction network and related parameters in the original speech extraction model end-to-end based on the target loss value until the target loss value converges or reaches the preset training rounds (such as 10 times), thereby obtaining the target speech extraction model.

[0075] Through steps S201 to S208, the original speech extraction model can sequentially learn scene-aware noise suppression, target user speech localization under reference speech guidance, target user speech spectrum recovery, temporal speech reconstruction, and consistency optimization capabilities under cross-noise background conditions during the training phase. Compared with the method of training the target speech extraction model using only a single scene training sample, by introducing training scene information, training reference speech spectrum features, training noise suppression spectrum features, and cross-condition consistency loss constraints, the accuracy of target user speech extraction in complex noisy, multi-speaker scenarios can be improved.

[0076] like Figure 5 As shown, in some embodiments, step S102 includes, but is not limited to, steps S401 to S404: Step S401: Perform a short-time Fourier transform on the reference speech to obtain the complex spectrum of the reference speech.

[0077] Step S402: Obtain the reference amplitude information and reference phase information of the complex spectrum of the reference speech.

[0078] Step S403: Perform amplitude compression processing on the reference amplitude information to obtain compressed speech amplitude information.

[0079] Step S404: Generate reference speech spectral features based on compressed speech amplitude information and reference phase information.

[0080] In step S401 of some embodiments, specifically, the reference speech complex spectrum refers to the spectrum data containing the real and imaginary parts obtained after the short-time Fourier transform.

[0081] Specifically, the reference speech can be segmented into frames according to a preset frame length and frame shift using a time-frequency transformation network. Then, a preset window function is applied to each frame of speech, and a fast Fourier transform is performed on each frame of speech after windowing to obtain the complex spectrum value on the corresponding time-frequency unit. The complex spectrum values ​​of each time-frequency unit are integrated to obtain the complex spectrum of the reference speech.

[0082] For example, in a financial scenario, the 3-second reference speech reserved by the customer during the remote account opening process can be processed. At a sampling rate of 16kHz, the 3-second reference speech contains a total of 48,000 sampling points. It can be divided into frames according to a preset frame length of 512 sampling points and a frame shift of 256 sampling points. A Hamming window of length 512 is applied to each frame, and then a short-time Fourier transform is performed to obtain the complex spectrum of the customer's reference speech.

[0083] In step S402 of some embodiments, specifically, the reference amplitude information is used to characterize the energy distribution value of the reference speech in each time-frequency unit.

[0084] Specifically, the reference phase information is used to characterize the phase value of the reference speech in each time-frequency unit.

[0085] Specifically, the magnitude of each complex spectral value in the complex spectrum of the reference speech can be calculated to obtain reference amplitude information, and the phase angle of each complex spectral value can be calculated to obtain reference phase information.

[0086] For example, if the complex value of a time-frequency unit in the complex spectrum of the reference speech is 3+4j, its corresponding reference amplitude value can be 5, and the reference phase value corresponding to (4 / 3) calculated by the tangent function is approximately 0.927 radians.

[0087] In step S403 of some embodiments, specifically, compressed speech amplitude information refers to amplitude feature information obtained after nonlinear amplitude compression processing of reference amplitude information, which is used to characterize the energy distribution of the reference speech spectrum and has a relatively reduced dynamic range.

[0088] Specifically, compressed speech amplitude information can be obtained by nonlinearly transforming each amplitude value in the reference amplitude information according to a preset compression index using power-law compression.

[0089] For example, if the preset compression index is 0.5, when compressing the reference amplitude value of 25 of a time-frequency unit, it can be compressed according to A to the power of 0.5, and the compressed speech amplitude information is 5.

[0090] In step S404 of some embodiments, specifically, the reference speech spectral features refer to the spectral representation used to characterize the time-frequency features of the target user's reference speech.

[0091] Specifically, compressed speech amplitude information can be combined with reference phase information to recover the corresponding complex spectrum form; or, real and imaginary features can be calculated based on compressed speech amplitude information and reference phase information respectively, and the real and imaginary features can be concatenated to obtain the reference speech spectrum features.

[0092] For example, if the compressed speech amplitude of a certain time-frequency unit is 5 and the reference phase is 0.927 radians, the real part of the time-frequency unit can be calculated to be approximately 3 and the imaginary part approximately 4 based on the polar coordinate transformation relationship, thereby recovering the corresponding complex spectrum representation. Alternatively, the real and imaginary parts can be directly used as the dual-channel spectrum feature output of the time-frequency unit.

[0093] Through steps S401 to S404, the reference speech spectrum features and the training mixed speech spectrum features can be kept consistent in terms of expression. While retaining the key speaker features and time-frequency structure information of the target user's speech, the impact of spectrum dynamic range fluctuations on subsequent feature interaction and target speech extraction is reduced, thereby improving the target speech extraction model's target user speech guidance effect and target speech extraction accuracy in complex and noisy environments.

[0094] In step S102 of some embodiments, specifically, the mixed speech spectral features are used to characterize the spectral distribution of the mixed speech in each time-frequency unit.

[0095] For example, in financial scenarios, mixed speech spectrum features corresponding to the account opening scenario can be generated based on the mixed speech complex spectrum in the insurance customer account opening environment; in medical scenarios, mixed speech spectrum features corresponding to the patient scenario can be generated based on the mixed speech complex spectrum in the patient consultation environment.

[0096] In this embodiment, by performing time-frequency transformation on the mixed speech, the spectral features of the mixed speech are obtained, which can more clearly characterize the distribution features of the target user's speech, background noise and interference speech at different time and frequency positions, thereby providing a more effective input basis for subsequent target speech extraction based on reference speech guidance.

[0097] In step S103 of some embodiments, specifically, the voice scene information refers to information reflecting the noise background conditions, interference conditions, or environmental types of mixed voice in different business scenarios.

[0098] For example, in the financial sector, voice scene information can correspond to scenario data such as insurance customer service calls and background noise in insurance business halls; in the medical sector, voice scene information can correspond to scenario data such as consultations in clinics, remote consultations, and interference from accompanying patients.

[0099] In this embodiment, by acquiring the speech scene information of mixed speech, the speech scene information can be combined to provide the subsequent target speech extraction model with scene data that can distinguish different noise environments, sound pickup conditions and business interaction situations, thereby improving the scene adaptability of the target speech extraction process and further improving the target user speech extraction accuracy in complex application environments.

[0100] like Figure 6As shown, in some embodiments, step S103 may include, but is not limited to, steps S501 to S504: Step S501: Generate scene modulation parameters based on the voice scene information.

[0101] Step S502: Perform spectral modulation processing on the mixed speech spectral features according to the scene modulation parameters to obtain scene spectral modulation features.

[0102] Step S503: Perform noise masking on the scene spectrum modulation features to obtain masked spectrum modulation features.

[0103] Step S504: Based on the mask spectrum modulation features, the mixed speech spectrum features are reconstructed to obtain the noise suppression spectrum features.

[0104] In step S501 of some embodiments, specifically, the scene modulation parameter refers to the control parameter generated based on the speech scene information, which is used to adjust the response intensity or feature weight of different frequency bands in the mixed speech spectral features.

[0105] For example, in financial scenarios, when voice scene information corresponds to insurance customer service calls or background noise in insurance business halls, scene modulation parameters can be used to characterize the suppression requirements of the steady-state noise frequency band in the business hall and the interference frequency band of customer service voices; in medical scenarios, when voice scene information corresponds to consultations in clinics, interference from accompanying persons, or noise from medical equipment, scene modulation parameters can be used to characterize the adjustment requirements of the noise frequency band of equipment operation and the area of ​​interference from accompanying persons' voices.

[0106] Specifically, the speech scene information can first be processed by scene encoding to obtain scene representation vectors corresponding to scene category, environment type and interference type. Then, the scene representation vectors can be mapped based on a preset parameter mapping layer to obtain scene modulation parameters.

[0107] In step S502 of some embodiments, specifically, the scene spectrum modulation feature refers to the speech spectrum feature representation that enhances the speech-related spectrum region of the target user and suppresses irrelevant interference spectrum regions under the speech scene information.

[0108] Specifically, the scene modulation parameters can be extended along the time and frequency dimensions and then multiplied element-wise with the mixed speech spectral features to obtain the scene spectral modulation features.

[0109] For example, in a financial scenario, if the number of channels, time frames, and frequency points of the mixed speech spectrum features of the target insurance customer are 64, the number of time frames is 100, and the number of frequency points is 257, the scene modulation parameters of length 64 can be expanded into a modulation matrix of 64x100x257, and multiplied element-wise with the mixed speech spectrum features to obtain scene spectrum modulation features that reduce customer service voice and background noise, and highlight the voice of the target insurance customer.

[0110] In step S503 of some embodiments, specifically, the masked spectrum modulation feature refers to the spectrum representation of the scene spectrum modulation feature after noise masking.

[0111] For example, in financial scenarios, for mixed speech consisting of background noise in insurance sales halls and insurance customer service voices, noise masks can be used to assign higher retention weights to time-frequency units where the customer's voice is dominant, and lower retention weights to time-frequency units where customer service voice interference and sales hall noise are dominant. In medical scenarios, for mixed speech consisting of the patient's voice, the voice of the accompanying person, and noise from medical equipment, noise masks can be used to prioritize the retention of time-frequency regions related to the patient's voice, while suppressing the regions corresponding to the interference from the accompanying person and the noise from the equipment.

[0112] Specifically, the scene spectrum modulation features can be input into a noise mask of the same size as the scene spectrum modulation features in a gated temporal convolutional recurrent network (GTCRN), and the noise mask can be multiplied element-wise with the scene spectrum modulation features to obtain the mask spectrum modulation features.

[0113] For example, if the size of the scene spectrum modulation feature is 64x100x257, GTCRN can output a 64x100x257 noise mask with a value range of 0 to 1; if the mask value corresponding to a certain time-frequency unit is 0.2, it means that the time-frequency unit is strongly suppressed, and if the mask value corresponding to a certain time-frequency unit is 0.9, it means that the time-frequency unit is largely preserved.

[0114] In step S504 of some embodiments, specifically, the noise suppression spectral feature refers to the spectral representation obtained after feature reconstruction in which the noise component and the non-target speech component are weakened.

[0115] For example, in the financial sector, the mixed speech spectrum features of customers can be reconstructed based on the masked spectrum modulation features to reduce background noise in insurance business halls and interference from insurance customer service voices, thus obtaining the noise-suppressed spectrum features corresponding to the customer; in the medical sector, the mixed speech spectrum features of patients can be reconstructed based on the masked spectrum modulation features to reduce the impact of accompanying persons' voices and noise from medical equipment, thus obtaining the noise-suppressed spectrum features corresponding to the patient.

[0116] Specifically, the masked spectral modulation features and the mixed speech spectral features can be weighted and summed according to preset weights to obtain the noise suppression spectral features.

[0117] For example, the masked spectral modulation features and their corresponding masked feature weights (e.g., 0.7) can be weighted and summed with the mixed speech spectral features and their mixed feature weights (e.g., 0.3) to determine the noise suppression spectral features.

[0118] Through steps S501 to S504, compared to the unified noise suppression method in traditional methods that does not incorporate speech scene information, the lightweight noise suppression network in the target speech extraction network can combine the noise background conditions, interfering speaker conditions, and environment types under different business scenarios to suppress noise more effectively. This is beneficial to improving the accuracy of target user speech extraction model in complex noisy, multi-speaker scenarios.

[0119] like Figure 7 As shown, in some embodiments, step S104 includes, but is not limited to, steps S601 to S603: Step S601: Map the noise suppression spectral features to query features, and map the reference speech spectral features to key features and value features.

[0120] Step S602: Attention fusion is performed based on query features, key features, and value features to obtain target interaction features.

[0121] Step S603: The target interaction features and noise suppression spectrum features are fused to obtain user voice guidance features.

[0122] In step S601 of some embodiments, specifically, the query feature, key feature, and value feature refer to the attention calculation features obtained by mapping different input spectral features in the feature interaction network. Specifically, the query feature is used to characterize the retrieval requirements of the current speech content to be extracted, the key feature is used to characterize the matching index information in the reference speech, and the value feature is used to characterize the target user feature information in the reference speech that can be guided for extraction.

[0123] Specifically, the noise-suppressed spectral features and the reference speech spectral features can be projected and transformed using linear mapping layers, convolutional mapping layers, or fully connected mapping layers in the feature interaction network of the target speech extraction network, respectively, to obtain query features, key features, and value features for the corresponding feature dimensions. The noise-suppressed spectral features can be used as the query input because they contain the current time-frequency distribution information of the speech to be separated; the reference speech spectral features can be used as the key-value input because they contain stable speaker prior information for the target user.

[0124] In step S602 of some embodiments, specifically, the target interaction feature refers to the vector representation of the integrated query feature, key feature and value feature. The target interaction feature is used to reflect the matching relationship between the noise suppression spectrum feature and the reference speech spectrum feature, so as to highlight the time-frequency region in the mixed speech that is consistent with the speaking features of the target user.

[0125] Specifically, the similarity matrix between the query features and the key features can be calculated first through the scaling dot product attention mechanism. Then, the similarity matrix is ​​normalized to obtain the attention weights. Finally, the value features are weighted and summed using the attention weights to obtain the target interaction features.

[0126] In step S603 of some embodiments, specifically, the user voice guidance feature refers to a vector representation that integrates the target interaction feature and the noise suppression spectrum feature, which is used to guide the target voice extraction model to locate the target user's voice content from the mixed speech.

[0127] Specifically, the target interaction features and noise suppression spectral features can be concatenated first, and then the user's voice guidance features can be output through linear transformation or convolutional transformation.

[0128] For example, if the target interaction feature and the noise suppression spectrum feature have dimensions of 32x100x257 and 64x100x257 respectively, they can be first concatenated into a fusion feature of 96x100x257, and then the user voice guidance feature can be obtained through a 3x3 convolution with 64 output channels.

[0129] By introducing reference speech spectral features into attention interaction through steps S601 to S603, the time-frequency region belonging to the target user in mixed speech can be located more accurately, reducing the impact of residual components of non-target speakers on the extracted speech results. Furthermore, since the user speech guidance feature integrates the current speech content information in the noise-suppressed spectral feature and the prior information of the target user identity in the reference speech spectral feature, the accuracy of the target speech extraction model in complex noisy, multi-speaker application environments can be improved.

[0130] like Figure 8As shown, in some embodiments, step S105 includes, but is not limited to, steps S701 to S703: Step S701: Extract frequency features from the user's voice guidance features to obtain voice frequency features.

[0131] Step S702: Extract time features from speech frequency features to obtain speech time features; Step S703: Extract the speech spectrum from the pronunciation time features based on the reference speech spectrum features to obtain the speech spectrum information of the target user.

[0132] In step S701 of some embodiments, specifically, speech frequency features refer to the feature representation obtained after frequency feature extraction that highlights the local structure of the target user's speech spectrum and the frequency band dependency, and is used to enhance the separability of the target user's speech in the frequency dimension.

[0133] For example, in financial scenarios, the target speech delivered by customers during insurance customer service calls or in noisy environments like insurance business halls typically contains a stable main frequency band distribution and harmonic structure. After extracting frequency features from the user's voice guidance features, the frequency band energy concentration area of ​​the customer's own voice can be further highlighted, and the aliasing effect of residual customer service voice and environmental noise in the frequency dimension can be weakened. In medical scenarios, the target speech of patients during consultations in clinics, remote consultations, or in scenarios with interference from accompanying persons also has a relatively stable frequency band structure. After extracting frequency features from the user's voice guidance features, the effective spectrum pattern of the patient's own voice can be highlighted, and the frequency band disturbances caused by the voice of accompanying persons and noise from medical equipment can be suppressed.

[0134] Specifically, the user's voice guidance features can be locally extracted along the frequency direction using a two-dimensional convolutional layer in the speech spectrum extraction network of the target speech extraction model, thus obtaining the speech frequency features. For example, when the feature size of the user's voice guidance features is 64x100x257, a two-dimensional convolutional layer with a kernel size of 1x5, 64 output channels, and a stride of 1 can be used to extract the frequency features of the user's voice guidance features, outputting speech frequency features with a size of 64x100x257.

[0135] In step S702 of some embodiments, specifically, the pronunciation time feature refers to the feature representation obtained after time feature extraction, which characterizes the pronunciation evolution law, speech activity continuity and contextual dependence of the target user's speech at different time steps, and is used to reflect the dynamic change features of the target user's speech in the continuous speech stream, so as to support the subsequent recovery of the target user's speech spectrum information.

[0136] For example, in financial scenarios, a customer's pronunciation during insurance business communication has continuity and contextual relevance. Even with interruptions from insurance customer service representatives or background noise in the business hall, the customer's own voice still has strong temporal continuity between adjacent time frames. After extracting temporal features from the speech frequency features, the inter-frame correlation of the customer's own voice can be further strengthened and the sudden impact of short-term interfering speech can be suppressed. In medical scenarios, a patient's pronunciation when describing their condition during a consultation also has obvious temporal continuity. After extracting temporal features from the speech frequency features, the differences between the patient's continuous expression and the interruptions from accompanying personnel or intermittent noise from equipment can be better distinguished.

[0137] Specifically, speech frequency features can be extracted into temporal features through the bidirectional long short-term memory layer in the speech spectrum extraction network to obtain the pronunciation time features.

[0138] In step S703 of some embodiments, specifically, the target user speech spectrum information refers to the target user speech spectrum data extracted from the spectrum representation of the mixed speech.

[0139] For example, in financial scenarios, the customer's own speech spectrum components in mixed speech can be extracted based on the customer's pronunciation time characteristics and the customer's reference speech spectrum characteristics. This allows for the extraction of the target user's speech spectrum information in insurance customer service calls and insurance business hall background noise environments.

[0140] Specifically, the pronunciation time features and reference speech spectrum features can be concatenated along the channel dimension and input into the speech spectrum extraction head for convolutional decoding to output the target user's speech spectrum information.

[0141] For example, if the number of channels for the pronunciation time feature is 128 and the number of channels for the reference speech spectrum feature is 64, the two can be concatenated into a fusion feature with 192 channels. Then, the target user's speech spectrum information can be generated through a 3x3 convolutional layer with 2 output channels, where the two output channels correspond to the real and imaginary parts of the target user's speech spectrum, respectively.

[0142] Through steps S701 to S703, the frequency band structure of the target user's speech can be enhanced by frequency feature extraction, the continuous pronunciation pattern of the target user's speech can be enhanced by time feature extraction, and finally, the speech spectrum extraction under speaker consistency constraints can be performed by combining the reference speech spectrum features. This allows for a more accurate distinction between the target user's speech and non-target speaker speech as well as background noise, thereby improving the target speech extraction model's accuracy in recovering the target user's speech spectrum in complex, noisy, multi-speaker environments.

[0143] In step S106 of some embodiments, specifically, the target user's speech refers to the speech signal of the target user extracted from the mixed speech.

[0144] Specifically, the inverse short-time Fourier transform of the target user's speech spectrum information can be performed through the inverse time-frequency transform network in the target speech extraction model, and time-domain reconstruction can be performed by combining the frame length, frame shift and window function parameters consistent with the preceding time-frequency transform, so as to restore the target user's speech spectrum components in each time-frequency unit to a continuous time-domain waveform.

[0145] For example, in a financial setting, the inverse short-time Fourier transform can be performed on the speech spectrum information of the target user corresponding to the customer to obtain the customer's speech; in a medical setting, the inverse short-time Fourier transform can be performed on the speech spectrum information of the target user corresponding to the patient to obtain the patient's speech.

[0146] In this embodiment, by performing inverse time-frequency transformation on the target user's speech spectrum information, the target user's speech is obtained. This can restore the target user's speech spectrum information in the frequency domain into playable, storable, and further analyzable target user speech, effectively overcoming the shortcomings of existing technologies in terms of low speech extraction accuracy in noisy scenarios, and significantly improving the accuracy of target speech extraction in complex noisy scenarios.

[0147] As can be seen, in the above scheme, by acquiring the target user's reference speech and mixed speech, and performing time-frequency transformation on them respectively, the spectral features of the reference speech and the spectral features of the mixed speech are obtained. Further, the speech scene information of the mixed speech is acquired, and noise suppression processing is applied to the spectral features of the mixed speech based on the speech scene information to obtain noise-suppressed spectral features. Then, feature interaction is performed between the noise-suppressed spectral features and the reference speech spectral features to obtain user speech guidance features. Speech spectrum extraction is then performed based on the user speech guidance features to obtain the target user's speech spectrum information. Finally, the target user's speech is obtained through inverse time-frequency transformation. In this invention, by introducing speech scene information to perform noise suppression processing on the mixed speech spectral features, the influence of environmental noise and interfering speakers on the target user's speech localization process can be reduced. Furthermore, through feature interaction between the reference speech spectral features and the noise-suppressed spectral features, the accuracy of target user speech content localization is improved. This effectively overcomes the shortcomings of existing technologies in terms of low speech extraction accuracy in noisy scenes and significantly improves the accuracy of target speech extraction in complex noisy scenes.

[0148] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0149] Please see Figure 9This application also provides a speech extraction device that can implement the above-described speech extraction method. The device includes: The reference speech and mixed speech acquisition module is used to acquire the reference speech of the target user and acquire the mixed speech; wherein, the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user. The time-frequency transformation module is used to perform time-frequency transformation on the reference speech using a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and to perform time-frequency transformation on the mixed speech to obtain the spectral features of the mixed speech. The noise suppression processing module is used to acquire speech scene information of mixed speech and perform noise suppression processing on the spectral features of mixed speech based on the speech scene information to obtain noise-suppressed spectral features. The feature interaction module is used to perform feature interaction between noise suppression spectral features and reference speech spectral features to obtain user speech guidance features; among them, the user speech guidance features are used to guide the target speech extraction model to locate the target user's speech content from the mixed speech. The speech spectrum extraction module is used to extract the speech spectrum based on the user's voice guidance characteristics to obtain the target user's speech spectrum information; The speech extraction module is used to perform inverse time-frequency transformation on the target user's speech spectrum information to obtain the target user's speech.

[0150] In one embodiment, the speech extraction device further includes: Obtain the target training set; wherein, the target training set includes training mixed speech generated for the same target user under at least two different noise background conditions, the user's real speech corresponding to each training mixed speech, the target user's training reference speech, and the training scene information corresponding to each training mixed speech; The training reference speech is subjected to time-frequency transformation by a preset original speech extraction model to obtain the spectral features of the training reference speech. The training mixed speech under various noise background conditions is also subjected to time-frequency transformation to obtain the spectral features of the training mixed speech. Based on the training scenario information, noise suppression processing is performed on the spectral features of each training mixed speech to obtain the training noise-suppressed spectral features; By performing feature interaction between the training noise suppression spectral features and the training reference speech spectral features, the training user speech guidance features are obtained. Speech spectrum is extracted based on the speech guidance features of the training user to obtain the speech spectrum information of the training user; The inverse time-frequency transformation is performed on the training user speech spectrum information to obtain the predicted user speech under various noise background conditions; The target loss value is obtained by calculating the loss value based on the predicted user speech and the user's actual speech under various noise background conditions. The original speech extraction model is trained based on the target loss value to obtain the target speech extraction model.

[0151] In one embodiment, the speech extraction device further includes: Speech loss is calculated based on the predicted user speech and the actual user speech under various noise background conditions to obtain the user speech loss value. Cross-conditional loss is calculated for the predicted user speech under various noise background conditions to obtain the cross-conditional speech loss value; The target loss value is obtained by calculating the loss based on the user's speech loss value and the cross-conditional speech loss value.

[0152] In one embodiment, the time-frequency conversion module is specifically used for: Perform a short-time Fourier transform on the reference speech to obtain the complex spectrum of the reference speech; Obtain reference amplitude and reference phase information of the complex spectrum of the reference speech; The reference amplitude information is subjected to amplitude compression processing to obtain compressed speech amplitude information; Reference speech spectral features are generated based on compressed speech amplitude information and reference phase information.

[0153] In one embodiment, the noise suppression processing module is specifically used for: Generate scene modulation parameters based on speech scene information; The mixed speech spectral features are obtained by performing spectral modulation processing on the scene modulation parameters; The scene spectrum modulation features are subjected to noise masking to obtain the masked spectrum modulation features; Based on the masked spectral modulation characteristics, the spectral features of the mixed speech are reconstructed to obtain the noise-suppressed spectral features.

[0154] In one embodiment, the feature interaction module is specifically used for: The noise-suppressed spectral features are mapped to query features, and the reference speech spectral features are mapped to key and value features; Attention fusion is performed based on query features, key features, and value features to obtain target interaction features; The user voice guidance features are obtained by fusing the target interaction features with the noise suppression spectrum features.

[0155] In one embodiment, the speech spectrum extraction module is specifically used for: Frequency features are extracted from the user's voice guidance features to obtain voice frequency features; Temporal features are extracted from the speech frequency features to obtain the pronunciation time features; Based on the spectral features of the reference speech, the speech spectrum of the pronunciation time features is extracted to obtain the speech spectrum information of the target user.

[0156] Specific limitations regarding the speech extraction device can be found in the limitations of the speech extraction method described above, and will not be repeated here. Each module in the aforementioned speech extraction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0157] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a speech extraction method on the server side.

[0158] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a speech extraction method on the client side.

[0159] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire the reference speech of the target user and acquire the mixed speech; wherein the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user. The reference speech is subjected to time-frequency transformation by a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and the mixed speech is subjected to time-frequency transformation to obtain the spectral features of the mixed speech. Acquire speech scene information of mixed speech, and perform noise suppression processing on the spectral features of mixed speech based on the speech scene information to obtain noise-suppressed spectral features; By performing feature interaction between the noise-suppressed spectral features and the reference speech spectral features, user speech guidance features are obtained; among them, user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech. Speech spectrum extraction is performed based on user voice guidance characteristics to obtain the target user's speech spectrum information; The target user's speech spectrum information is subjected to inverse time-frequency transformation to obtain the target user's speech.

[0160] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire the reference speech of the target user and acquire the mixed speech; wherein the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user. The reference speech is subjected to time-frequency transformation by a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and the mixed speech is subjected to time-frequency transformation to obtain the spectral features of the mixed speech. Acquire speech scene information of mixed speech, and perform noise suppression processing on the spectral features of mixed speech based on the speech scene information to obtain noise-suppressed spectral features; By performing feature interaction between the noise-suppressed spectral features and the reference speech spectral features, user speech guidance features are obtained; among them, user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech. Speech spectrum extraction is performed based on user voice guidance characteristics to obtain the target user's speech spectrum information; The target user's speech spectrum information is subjected to inverse time-frequency transformation to obtain the target user's speech.

[0161] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0163] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0164] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0165] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A speech extraction method, characterized in that, include: Acquire a reference speech of the target user and acquire a mixed speech; wherein the reference speech is noise-free, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user; The reference speech is subjected to time-frequency transformation by a pre-trained target speech extraction model to obtain the spectral features of the reference speech, and the mixed speech is subjected to time-frequency transformation to obtain the spectral features of the mixed speech. Acquire the speech scene information of the mixed speech, and perform noise suppression processing on the spectral features of the mixed speech based on the speech scene information to obtain noise-suppressed spectral features; The noise suppression spectral features and the reference speech spectral features are interacted to obtain user speech guidance features; wherein, the user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech; Based on the user's voice guidance features, voice spectrum extraction is performed to obtain the target user's voice spectrum information; The target user's speech spectrum information is subjected to inverse time-frequency transformation to obtain the target user's speech.

2. The method as described in claim 1, characterized in that, The step of performing time-frequency transformation on the reference speech using a pre-trained target speech extraction model to obtain the spectral features of the reference speech includes: Perform a short-time Fourier transform on the reference speech to obtain the complex spectrum of the reference speech; Obtain the reference amplitude information and reference phase information of the complex spectrum of the reference speech; The reference amplitude information is subjected to amplitude compression processing to obtain compressed speech amplitude information; The reference speech spectral features are generated based on the compressed speech amplitude information and the reference phase information.

3. The method as described in claim 1, characterized in that, The step of performing noise suppression processing on the mixed speech spectral features based on the speech scene information to obtain noise-suppressed spectral features includes: Generate scene modulation parameters based on the speech scene information; The mixed speech spectral features are subjected to spectral modulation processing based on the scene modulation parameters to obtain scene spectral modulation features; The scene spectrum modulation features are subjected to noise masking to obtain masked spectrum modulation features; Based on the mask spectrum modulation features, the mixed speech spectrum features are reconstructed to obtain the noise suppression spectrum features.

4. The method as described in claim 1, characterized in that, The step of performing feature interaction between the noise suppression spectral features and the reference speech spectral features to obtain user voice guidance features includes: The noise suppression spectral features are mapped to query features, and the reference speech spectral features are mapped to key features and value features; Attention fusion is performed based on the query features, the key features, and the value features to obtain the target interaction features; The target interaction features and the noise suppression spectrum features are fused to obtain the user voice guidance features.

5. The method as described in claim 1, characterized in that, The step of extracting the speech spectrum based on the user's voice guidance features to obtain the target user's speech spectrum information includes: Frequency features are extracted from the user voice guidance features to obtain voice frequency features; The speech frequency features are subjected to time feature extraction to obtain the pronunciation time features; Based on the reference speech spectrum features, the speech spectrum of the pronunciation time features is extracted to obtain the speech spectrum information of the target user.

6. The method as described in claim 1, characterized in that, Before performing time-frequency transformation on the reference speech using a pre-trained target speech extraction model to obtain the spectral features of the reference speech, the method further includes: Obtain a target training set; wherein, the target training set includes training mixed speech generated for the same target user under at least two different noise background conditions, the user's real speech corresponding to each training mixed speech, the target user's training reference speech, and training scene information corresponding to each training mixed speech; The training reference speech is subjected to time-frequency transformation by a preset original speech extraction model to obtain the spectral features of the training reference speech. The training mixed speech under various noise background conditions is also subjected to time-frequency transformation to obtain the spectral features of the training mixed speech. Based on the training scenario information, noise suppression processing is performed on the spectral features of each training mixed speech to obtain training noise-suppressed spectral features; The training noise suppression spectral features and the training reference speech spectral features are interacted to obtain the training user speech guidance features; Based on the training user's voice guidance features, the speech spectrum is extracted to obtain the training user's speech spectrum information; The inverse time-frequency transform is performed on the training user speech spectrum information to obtain the predicted user speech under various noise background conditions; The target loss value is obtained by calculating the loss value based on the predicted user speech and the actual user speech under various noise background conditions. The original speech extraction model is trained based on the target loss value to obtain the target speech extraction model.

7. The method as described in claim 6, characterized in that, The step of calculating the target loss value based on the predicted user speech and the actual user speech under various noise background conditions includes: Speech loss is calculated based on the predicted user speech and the actual user speech under various noise background conditions to obtain the user speech loss value. Cross-conditional loss is calculated for the predicted user speech under various noise background conditions to obtain cross-conditional speech loss values; The target loss value is obtained by calculating the loss based on the user's speech loss value and the cross-conditional speech loss value.

8. A speech extraction device, characterized in that, The device includes: A reference speech and mixed speech acquisition module is used to acquire a reference speech of a target user and acquire a mixed speech; wherein the reference speech is free of noise, the mixed speech contains noise, the speaker of the reference speech is the target user, and the speaker of the mixed speech includes the target user; The time-frequency transformation module is used to perform time-frequency transformation on the reference speech through a pre-trained target speech extraction model to obtain the reference speech spectral features, and to perform time-frequency transformation on the mixed speech to obtain the mixed speech spectral features; The noise suppression processing module is used to acquire the speech scene information of the mixed speech, and perform noise suppression processing on the spectral features of the mixed speech according to the speech scene information to obtain noise suppression spectral features; The feature interaction module is used to perform feature interaction between the noise suppression spectral features and the reference speech spectral features to obtain user speech guidance features; wherein, the user speech guidance features are used to guide the target speech extraction model to locate the speech content of the target user from the mixed speech; The speech spectrum extraction module is used to extract the speech spectrum based on the user's speech guidance features to obtain the target user's speech spectrum information. The speech extraction module is used to perform inverse time-frequency transformation on the speech spectrum information of the target user to obtain the speech of the target user.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the target speech extraction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the target speech extraction method as described in any one of claims 1 to 7.