Speech enhancement method based on discrimination-generation joint model

Through the speech enhancement method based on the discriminant-generating joint model, the problems of poor voice enhancement effect and high computing resource consumption in the prior art are solved, and higher quality speech enhancement and lower computing complexity are achieved.

CN120220708APending Publication Date: 2025-06-27UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389187.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing speech enhancement methods have a large difference from the actual situation when outputting enhanced speech, and consume a lot of computing resources.

Method used

A speech enhancement method based on a discriminant-generating joint model is adopted, and the speech signal to be processed is obtained, input to the discriminant-generating joint model, generating predicted frequency domain information and predicted fraction function, and finally generating enhanced speech signal. The model includes a speech discriminant network, a speech interaction network and a speech generation network, which is used to fuse hidden features to generate predicted fractional functions.

Benefits of technology

The quality of speech reconstruction is improved, the computational complexity is reduced, the computing resources required for speech enhancement is reduced, and the generated enhanced speech signal is closer to the actual speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220708A_ABST
    Figure CN120220708A_ABST
Patent Text Reader

Abstract

The invention provides a voice enhancement method based on a discrimination-generation joint model, and the method comprises the steps: obtaining a to-be-processed voice signal which represents a voice signal with noise; to-be-processed voice signals are input into a discrimination-generation joint model, prediction frequency domain information and a prediction fractional function are obtained, and the discrimination-generation joint model comprises a voice discrimination network, a voice interaction network and a voice generation network; the voice interaction network is used for fusing hidden features in the voice discrimination network and the voice generation network, so that the voice generation network generates a prediction fractional function according to the fused hidden features; and generating an enhanced speech signal according to the predicted frequency domain information and the predicted fractional function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech processing, and more particularly, to a speech enhancement method, device, electronic device, computer-readable storage medium, and computer program product based on a discriminative-generative joint model. Background Art

[0002] Speech Enhancement (SE) aims to recover a clean speech signal from an audio signal disturbed by various degradation types (including background noise, room reverberation, codec artifacts, etc.), and it is widely used as a front-end module for applications such as human-computer interaction and remote conferencing.

[0003] Many existing SE methods are usually task-driven and are designed for denoising, dereverberation, or speech super-resolution tasks respectively. However, the enhanced speech output by the speech enhancement method of the related technology is quite different from the actual one, and more computing resources are used. Summary of the Invention

[0004] In view of this, the present application provides a speech enhancement method, device, electronic device, computer-readable storage medium, and computer program product based on a discriminative-generative joint model.

[0005] One aspect of the present application provides a speech enhancement method based on a discriminative-generative joint model, including:

[0006] Obtaining a speech signal to be processed, where the speech signal to be processed represents a noisy speech signal;

[0007] Inputting the speech signal to be processed into a discriminative-generative joint model to obtain predicted frequency-domain information and a predicted score function, where the discriminative-generative joint model includes a speech discriminative network, a speech interaction network, and a speech generation network, and the speech interaction network is used to fuse the hidden features in the speech discriminative network and the speech generation network so that the speech generation network generates the predicted score function according to the fused hidden features;

[0008] Generating an enhanced speech signal according to the predicted frequency-domain information and the predicted score function.

[0009] Another aspect of the present application provides a speech enhancement device, including:

[0010] An obtaining module, configured to obtain a speech signal to be processed, where the speech signal to be processed represents a noisy speech signal;

[0011] A processing module, configured to input the to-be-processed voice signal into a discriminative-generation joint model to obtain predicted frequency-domain information and a predicted score function, wherein the discriminative-generation joint model includes a voice discrimination network, a voice interaction network, and a voice generation network, and the voice interaction network is configured to fuse hidden features in the voice discrimination network and the voice generation network so that the voice generation network generates the predicted score function according to the fused hidden features;

[0012] A generation module, configured to generate an enhanced voice signal according to the predicted frequency-domain information and the predicted score function.

[0013] Another aspect of the present application provides an electronic device, including:

[0014] One or more processors;

[0015] A memory, configured to store one or more programs,

[0016] wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method as described above.

[0017] Another aspect of the present application provides a computer-readable storage medium, storing computer-executable instructions, which are used to implement the method as described above when executed.

[0018] Another aspect of the present application provides a computer program product, the computer program product including computer-executable instructions, which are used to implement the method as described above when executed. Description of the Drawings

[0019] Through the following description of the embodiments of the present application with reference to the drawings, the above and other objects, features, and advantages of the present application will become clearer. In the drawings:

[0020] Figure 1 An exemplary system architecture to which the voice enhancement method according to the embodiments of the present application can be applied is shown;

[0021] Figure 2 A flowchart of the voice enhancement method according to the embodiments of the present application is shown;

[0022] Figure 3 A flowchart of the training method of the discriminative-generation joint model according to the embodiments of the present application is shown;

[0023] Figure 4 A schematic diagram of the model structure of the discriminative-generation joint model according to the embodiments of the present application is shown;

[0024] Figure 5Shows a schematic structural diagram of a subband down / upsampling block according to an embodiment of the present application;

[0025] Figure 6 Shows a schematic structural diagram of a voice interaction network according to an embodiment of the present application;

[0026] Figure 7 Shows a schematic structural diagram of a dual-path recurrent attention network according to an embodiment of the present application;

[0027] Figure 8 Shows a schematic diagram of the performance comparison between a discriminative-generative joint model and a baseline model on the TIMIT-UNI dataset according to an embodiment of the present application;

[0028] Figure 9 Shows a block diagram of a voice enhancement device according to an embodiment of the present application; and

[0029] Figure 10 Shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present application. Detailed implementation manners

[0030] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, obviously, one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0031] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0032] Existing SE methods can be roughly divided into discriminative and generative methods. Discriminative methods are usually based on supervised learning, treating SE as a regression task and learning the best mapping from the degraded signal to the target signal under certain optimization criteria. Benefiting from the development of deep neural networks, both time-domain models that directly operate on speech waveforms and time-frequency domain models that operate in the Short-Time Fourier Transform (STFT) domain have shown good performance in improving speech quality. However, from the perspective of Universal Speech Enhancement (USE), generative methods intuitively have greater potential because some types of degradation require the model to generate signals from scratch, such as clipping and bandwidth limitation. Generative models aim to integrate the inherent distribution of data into the latent space and generate samples from it, compensating for missing information by learning the prior. Different from discriminative methods that provide deterministic predictions, generative methods can generate many reasonable candidate samples, which is consistent with the property that the signal reconstruction method is not unique.

[0033] The diffusion models proposed by relevant scholars currently have become the state-of-the-art generative method paradigm. The embodiments of this application use score-based diffusion models, including a forward diffusion process and a reverse sampling process. The former gradually adds noise to the data to make its distribution tend to a tractable prior, and the latter reverses this process to restore the original distribution of the data. This process can be described by a Stochastic Differential Equation (SDE) defined from time 0 to T.

[0034] To improve both the speech reconstruction quality and reduce the computational complexity simultaneously, the embodiments of this application provide a speech enhancement method based on a discriminative-generative joint model, including obtaining a speech signal to be processed, where the speech signal to be processed represents a noisy speech signal; inputting the speech signal to be processed into the discriminative-generative joint model to obtain predicted frequency-domain information and a predicted score function, where the discriminative-generative joint model includes a speech discriminative network, a speech interaction network, and a speech generation network, and the speech interaction network is used to fuse the hidden features in the speech discriminative network and the speech generation network so that the speech generation network generates a predicted score function according to the fused hidden features; generating an enhanced speech signal according to the predicted frequency-domain information and the predicted score function.

[0035] In the embodiments of this application, in aspects such as the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.

[0036] In the embodiments of the present application, before obtaining or collecting user personal information (such as voice signals), user authorization or consent is obtained.

[0037] Figure 1 FIG. 100 shows an exemplary system architecture to which the voice enhancement method according to the embodiments of the present application can be applied. It should be noted that, Figure 1 The figure shown is only an example of the system architecture to which the embodiments of the present application can be applied, to help those skilled in the art understand the technical content of the present application, but it does not mean that the embodiments of the present application cannot be used in other devices, systems, environments or scenarios.

[0038] As Figure 1 shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0039] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).

[0040] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0041] The server 105 may be a server providing various services, such as a background management server (only for example) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0042] It should be noted that the voice enhancement method provided by the embodiments of the present application can generally be executed by the server 105. Correspondingly, the voice enhancement device provided by the embodiments of the present application can generally be set in the server 105. The voice enhancement method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the voice enhancement device provided by the embodiments of the present application can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0043] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0044] Figure 2 shows a flowchart of the voice enhancement method according to an embodiment of the present application.

[0045] As Figure 2 shown, the voice enhancement method based on the discriminative-generative joint model includes operations S201 to S203.

[0046] In operation S201, a voice signal to be processed is obtained, where the voice signal to be processed represents a voice signal with noise;

[0047] In operation S202, the voice signal to be processed is input into the discriminative-generative joint model to obtain predicted frequency-domain information and a predicted score function, where the discriminative-generative joint model includes a voice discriminative network, a voice interaction network, and a voice generation network, and the voice interaction network is used to fuse the hidden features in the voice discriminative network and the voice generation network so that the voice generation network generates a predicted score function according to the fused hidden features;

[0048] In operation S203, an enhanced voice signal is generated according to the predicted frequency-domain information and the predicted score function.

[0049] According to an embodiment of the present application, a voice signal with noise may refer to a voice segment including background noise, room reverberation, or artifacts of an encoder or a decoder.

[0050] According to an embodiment of the present application, performing Fourier transform processing on the to-be-processed speech signal with noise can obtain the to-be-processed frequency-domain information of the to-be-processed speech signal. The to-be-processed frequency-domain information includes the real part, imaginary part, and amplitude spectrum of the complex spectrum. The state variable of the amplitude can be obtained from the amplitude spectrum, and then the to-be-processed frequency-domain information and the state variable of the amplitude are input into the discriminative-generative joint model.

[0051] According to an embodiment of the present application, during the process of the speech discriminative network in the discriminative-generative joint model processing the to-be-processed frequency-domain information to generate predicted frequency-domain information, some hidden features will be generated. At the same time, some hidden features will also be generated during the process of the speech generative network generating the predicted score function. The above hidden features characterize the correlation relationship of the to-be-processed speech signal in the time dimension and frequency dimension.

[0052] According to an embodiment of the present application, the speech interaction network performs weighted fusion processing on the hidden features output by the speech discriminative network and the speech generative network, and the output fused hidden features are input into the speech generative network. Thus, the speech generative network can output the final predicted score function. Finally, the enhanced speech signal containing clean speech information can be obtained according to the predicted frequency-domain information and the predicted score function.

[0053] According to an embodiment of the present application, by inputting the to-be-processed speech signal into the discriminative-generative joint model to obtain the predicted frequency-domain information and the predicted score function, and generating the enhanced speech signal according to the predicted frequency-domain information and the predicted score function. Since the speech interaction network in the discriminative-generative joint model can fuse the hidden features output by the speech discriminative network and the speech generative network, it is possible to make the speech generative network combine the fused hidden features for feature processing to obtain a predicted score function with a smaller possibility of artifacts. Finally, a more accurate enhanced speech signal can be obtained by combining the predicted frequency-domain information, and the computational complexity is small. Thus, the computational resources required for speech enhancement can be reduced.

[0054] According to an embodiment of the present application, generating the enhanced speech signal according to the predicted frequency-domain information and the predicted score function includes the following operations:

[0055] Processing the predicted score function based on the reverse stochastic differential equation of the score diffusion model to obtain the predicted amplitude information ; according to the predicted amplitude information and the predicted frequency-domain information , generating the predicted amplitude spectrum , as shown in formula (1):

[0056] (1)

[0057] (2)

[0058] Among them, is the fusion coefficient, is the predicted frequency-domain information, is the imaginary part of the complex spectrum in the predicted frequency-domain information, is the real part of the complex spectrum in the predicted frequency-domain information.

[0059] According to the embodiments of the present application, based on formula (2), the predicted frequency-domain information can be generated according to the features output by the speech discrimination network , for the predicted frequency-domain information and the predicted amplitude spectrum perform inverse Fourier transform to obtain the enhanced speech signal.

[0060] Figure 3 shows a flowchart of the training method of the discrimination-generation joint model according to the embodiments of the present application.

[0061] As Figure 3 shown, the training method of the discrimination-generation joint model includes operations S301~S305.

[0062] In operation S301, obtain a speech training set, where the speech training set includes multiple speech training samples and the labeled speech information corresponding to each speech training sample, and the speech training samples include the initial frequency-domain information and amplitude state variables of the training speech.

[0063] In operation S302, for each speech training sample, based on the attention mechanism, use the speech discrimination network to process the amplitude and phase in the initial frequency-domain information to obtain the target discrimination hidden feature and the target frequency-domain information.

[0064] In operation S303, use the speech interaction network to perform weighted fusion on the target discrimination hidden feature and the initial generation hidden feature to obtain the target generation hidden feature, where the initial generation hidden feature is obtained by the speech generation network processing the amplitude state variables, and both the target discrimination hidden feature and the target generation hidden feature characterize the correlation relationship of the speech training sample in the time dimension and the frequency dimension.

[0065] In operation S304, based on the attention mechanism, use the speech generation network to process the amplitude state variables and the target generation hidden feature to obtain the target score function, where the initial enhancement model includes a speech discrimination network, a speech generation network, and a speech interaction network.

[0066] In operation S305, iteratively adjust the model parameters of the initial enhancement model according to the target score function, the target frequency-domain information, and the labeled speech information to obtain the trained discrimination-generation joint model.

[0067] According to an embodiment of the present application, the training speech may be a speech signal with noise, and the initial frequency domain information may refer to the amplitude information and phase information obtained by performing a Fourier transform on the training speech. For example, it can be represented by the real part, imaginary part of the complex spectrum, and the amplitude spectrum. The amplitude state variable is determined from the amplitude information obtained by the Fourier transform.

[0068] According to an embodiment of the present application, for each speech training sample, the speech discrimination network processes the amplitude and phase in the initial frequency domain information to obtain a target discrimination hidden feature and target frequency domain information. For example, at the same time, the target frequency domain information may be the real part of the complex spectrum , imaginary part . During the process of processing the amplitude state variable, the speech generation network outputs an initial generation hidden feature. The initial generation hidden feature and the target discrimination hidden feature are simultaneously input into the speech interaction network to perform weighted fusion on the target discrimination hidden feature and the initial generation hidden feature, obtaining a target generation hidden feature. The target generation hidden feature is input into the speech generation network, so that the speech generation network combines the target generation hidden feature, thereby outputting a target score function.

[0069] According to an embodiment of the present application, by iteratively adjusting the model parameters of the initial enhancement model according to the target score function, target frequency domain information, and labeled speech information, a trained discriminative - generative joint model can be obtained.

[0070] Figure 4 shows a schematic diagram of the model structure of the discriminative - generative joint model according to an embodiment of the present application.

[0071] According to an embodiment of the present application, as Figure 4 shown, based on the attention mechanism, using the speech discrimination network to process the amplitude and phase in the initial frequency domain information to obtain a target discrimination hidden feature and target frequency domain information, including: using a discriminative encoder to encode the initial frequency domain information to obtain discriminative convolutional features; using a discriminative recurrent attention network to process the discriminative convolutional features to obtain multiple discriminative attention features, where the target discrimination hidden feature includes the discriminative convolutional features or discriminative attention features, and the discriminative recurrent attention network includes multiple first dual - path recurrent attention networks; using a discriminative decoder to process the discriminative attention features and discriminative convolutional features output by the last first dual - path recurrent attention network to obtain the target frequency domain information.

[0072] According to an embodiment of the present application, the first dual - path recurrent attention network may be a Dual - Path Recurrent Attention (DPRA) network.

[0073] According to an embodiment of the present application, after inputting the initial frequency domain information into the speech discrimination network, the discrimination encoder in the speech discrimination network first encodes the initial frequency domain information to obtain discrimination convolution features. Among them, the discrimination convolution features are input into the speech interaction network as a target discrimination hidden feature to combine with an initial generation hidden feature to generate a target generation hidden feature, where the number of the target discrimination hidden feature and the target generation hidden feature is related to the number of the first dual-path recurrent attention networks.

[0074] According to an embodiment of the present application, the discrimination attention features output by the last one of the multiple first dual-path recurrent attention networks and the discrimination convolution features output by the discrimination encoder are synchronously input into the discrimination decoder to encode the discrimination attention features and the discrimination convolution features, so as to obtain the target frequency domain information, which may include, for example, the real part of the complex spectrum and the imaginary part .

[0075] According to an embodiment of the present application, as Figure 4 shown, the amplitude state variable and the target generation hidden feature are input into the speech generation network, and for the amplitude state variable, a target score function is output, including: encoding the amplitude state variable by using a generation encoder to obtain generation convolution features, and determining the generation convolution features as an initial generation hidden feature; processing multiple target generation hidden features by using a generation recurrent attention network to obtain target generation attention features, where the generation recurrent attention network includes multiple second dual-path recurrent attention networks; transmitting the target generation attention features to the speech interaction network, so that the speech interaction network generates a target hidden feature according to the target generation attention features and the discrimination attention features output by the last one of the first dual-path recurrent attention networks; and processing the target hidden feature and the generation convolution features by using a generation decoder to obtain the target score function.

[0076] According to an embodiment of the present application, the network structure of the second dual-path recurrent attention network is the same as that of the first dual-path recurrent attention network, and both can be a dual-path recurrent attention (DPRA) network.

[0077] According to an embodiment of the present application, for the amplitude state variable , first, the amplitude state variable is encoded by using a generation encoder to obtain generation convolution features, and the generation convolution features can be input into the speech interaction network as an initial generation hidden feature, so that the speech interaction network generates a target generation hidden feature according to the generation convolution features and the discrimination convolution features as the target discrimination hidden feature.

[0078] According to an embodiment of the present application, a generative recurrent attention network is used to process the target to generate a hidden feature to generate a new target-generated hidden feature and input it into the generative recurrent attention network. The generative recurrent attention network generates a new target-generated hidden feature again according to the new target-generated hidden feature and a discriminative attention feature output by the discriminative recurrent attention network, and finally outputs a target hidden feature through iteration.

[0079] According to an embodiment of the present application, a generative decoder is used to perform feature encoding on the target hidden feature and the generative convolutional feature output by the generative encoder, so as to obtain a target score function.

[0080] According to an embodiment of the present application, as Figure 4 shown, the discriminative encoder and the generative encoder perform data processing in the following manner: the first convolutional layer is used to process the first input data to obtain a first convolutional feature, where the first input data includes initial frequency domain information or amplitude state variables; m sub-band downsampling blocks are used to process the first convolutional feature to obtain a first output feature, where the first output feature includes a discriminative convolutional feature or a generative convolutional feature.

[0081] According to an embodiment of the present application, for the discriminative encoder and the generative encoder, their data processing flows are the same, and the only difference lies in the different input data and output data, that is, the first input data of the discriminative encoder is the initial frequency domain information, and the first output feature is the discriminative convolutional feature, and the first input data of the generative encoder is the amplitude state variable, and the first output feature is the generative convolutional feature. The data processing of the discriminative encoder is described below by way of example.

[0082] According to an embodiment of the present application, for the first input data of the initial frequency domain information, first, the first convolutional layer in the discriminative encoder performs feature convolution processing on the initial frequency domain information to obtain a first convolutional feature. Then the first convolutional feature is input into m sub-band downsampling blocks, and the m sub-band downsampling blocks are cascaded in sequence, and thus a discriminative convolutional feature can be obtained. Wherein, the value of m can be specifically set according to requirements, for example, m can be 3.

[0083] According to an embodiment of the present application, as Figure 4 shown, the discriminative decoder and the generative decoder perform data processing in the following manner: n sub-band upsampling blocks are used to process the third input data to obtain an upsampled feature, where the third input data includes a first discriminative feature or a first generative feature, the first discriminative feature includes a discriminative attention feature and a discriminative convolutional feature, and the first generative feature includes a target hidden feature and a generative convolutional feature; the second convolutional layer is used to process the upsampled feature to obtain a second convolutional feature, where the second convolutional feature represents target frequency domain information or a target score function.

[0084] According to an embodiment of the present application, for the discriminative decoder and the generative decoder, their data processing flows are the same, and the only difference lies in the different input data and output data. That is, the third input data of the discriminative decoder is the discriminative attention feature, and the second convolutional feature is the target frequency domain information. The third input data of the generative decoder is the first generation feature, and the second convolutional feature is the target score function. The data processing of the discriminative decoder is exemplarily described below.

[0085] According to an embodiment of the present application, for the discriminative decoder, the discriminative attention feature output by the last first dual-path recurrent attention network is input into n sub-band upsampling blocks, and the n sub-band upsampling blocks are cascaded in sequence, thereby obtaining an upsampled feature. The value of n can be specifically set according to requirements. For example, n can be 3.

[0086] According to an embodiment of the present application, the upsampled feature output by the last sub-band upsampling block is input into the second convolutional layer for feature convolution, thereby obtaining the target frequency domain information.

[0087] According to an embodiment of the present application, the discriminative decoder (or generative decoder) and the discriminative encoder (or generative encoder) are in a mirror image relationship, and the two are connected by a skip connection to facilitate the backpropagation of gradients during training.

[0088] Figure 5 FIG. shows a schematic structural diagram of a sub-band down / upsampling block according to an embodiment of the present application.

[0089] According to an embodiment of the present application, as Figure 5 shown, the first convolutional feature is processed by m sub-band downsampling blocks to obtain a first output feature, including: for the i-th sub-band downsampling block, performing band splitting processing on the second input feature to obtain a first frequency feature and a second frequency feature, where the second input feature includes the first convolutional feature or the (i-1)-th normalized feature; performing downsampling processing on the first frequency feature to obtain a downsampled feature; performing convolutional processing on the second frequency feature to obtain a convolutional feature; performing merging processing on the downsampled feature, the convolutional feature, and the time embedding feature to obtain a merged feature; performing layer normalization processing on the merged feature to obtain the i-th normalized feature, where the i-th normalized feature represents the first output feature when i = m.

[0090] According to an embodiment of the present application, the network structure of the sub-band downsampling block is the same as that of the sub-band upsampling block. The data processing flow of one sub-band downsampling block is exemplarily described in the embodiment of the present application.

[0091] According to an embodiment of the present application, for the first sub-band downsampling block, first, perform band splitting processing on the second input feature to obtain a high-frequency first frequency feature and a low-frequency second frequency feature. Perform downsampling processing on the high-frequency first frequency feature, and perform two-dimensional convolution processing on the low-frequency second frequency feature, so as to obtain a downsampled feature and a convolution feature respectively. Combine the downsampled feature, the convolution feature, and the time embedding feature t emb to perform merging processing to obtain a merged feature. Perform layer normalization processing on the merged feature, and activate it based on the activation function PReLU to obtain the first normalized feature.

[0092] According to an embodiment of the present application, for the i-th sub-band downsampling block, where i = 2 to m - 1, perform the same operations as above on the (i - 1)-th normalized feature to obtain the i-th normalized feature. For the m-th sub-band downsampling block, perform the same operations as above on the (m - 1)-th normalized feature to obtain the m-th normalized feature, and this m-th normalized feature is the first output feature.

[0093] In the above embodiment, the sub-band downsampling block only performs downsampling operations on the high-frequency first frequency feature, thereby maintaining the resolution of the low-frequency part, and thus can improve the auditory perception.

[0094] According to an embodiment of the present application, the number of the first dual-path recurrent attention networks and the second dual-path recurrent attention networks is L, and the number of L can be specifically set according to the situation, for example, it can be 3.

[0095] According to an embodiment of the present application, use the discriminative recurrent attention network to process the discriminative convolutional feature to obtain a plurality of discriminative attention features, including: for the i-th first dual-path recurrent attention network, use the i-th first dual-path recurrent attention network to process the fourth input data and output the i-th discriminative feature. Among them, when i = 1, the fourth input data represents the discriminative convolutional feature, and when i ≠ 1, the fourth input data represents the (i - 1)-th discriminative feature. The target discriminative hidden feature includes the discriminative convolutional feature and the i-th discriminative feature.

[0096] According to an embodiment of the present application, for the discriminative recurrent attention network, the first first dual-path recurrent attention network processes the discriminative convolutional feature output by the discriminative encoder based on the attention mechanism to obtain the first discriminative feature. This first discriminative feature is used as a target discriminative hidden feature and input into the speech interaction network to combine with the target generation attention feature output by the first second dual-path recurrent attention network to generate the input data for the second second dual-path recurrent attention network.

[0097] According to an embodiment of the present application, the second first dual-path recurrent attention network processes the first discriminative feature output by the first first dual-path recurrent attention network based on the attention mechanism to obtain a second discriminative feature, and this second discriminative feature is input into the speech interaction network as a target discriminative hidden feature to combine with the target generation attention feature output by the second second dual-path recurrent attention network to generate input data for the third second dual-path recurrent attention network.

[0098] According to an embodiment of the present application, the third (last) first dual-path recurrent attention network processes the second discriminative feature output by the second first dual-path recurrent attention network based on the attention mechanism to obtain a third discriminative feature and input it into the discriminative decoder. At the same time, this third discriminative feature is input into the speech interaction network as a target discriminative hidden feature to combine with the target generation attention feature output by the third second dual-path recurrent attention network to generate input data for the generation decoder.

[0099] According to an embodiment of the present application, using the generation recurrent attention network to process the generation convolutional feature and multiple target generation hidden features to obtain the target generation attention feature includes: for the i-th second dual-path recurrent attention network, using the i-th second dual-path recurrent attention network to process the fifth input data and output the i-th generation feature. Among them, when i = 1, the fifth input data includes the target generation hidden feature generated by the speech interaction network according to the discriminative convolutional feature and the generation convolutional feature; when i ≠ 1, the fifth input data includes the target generation hidden feature generated by the speech interaction network according to the (i - 1)-th discriminative feature and the (i - 1)-th generation feature. Among them, when i = L, the i-th generation feature represents the target generation attention feature.

[0100] According to an embodiment of the present application, the first second dual-path recurrent attention network processes the target generation hidden feature generated by the speech interaction network according to the discriminative convolutional feature and the generation convolutional feature based on the attention mechanism to obtain a first generation feature, and this first generation feature is input into the speech interaction network as a target generation attention feature to combine with the first discriminative feature output by the first first dual-path recurrent attention network to generate input data for the second second dual-path recurrent attention network.

[0101] According to an embodiment of the present application, the second second dual-path recurrent attention network processes the target generation hidden feature generated by the speech interaction network according to the first generation feature and the first discriminative feature based on the attention mechanism to obtain a second generation feature, and this second generation feature is input into the speech interaction network as a target generation attention feature to combine with the second discriminative feature output by the second first dual-path recurrent attention network to generate input data for the third second dual-path recurrent attention network.

[0102] According to an embodiment of the present application, the third (last) second dual-path cyclic attention network processes the target generated latent feature generated by the speech interaction network based on the second generated feature and the second discriminant feature according to the attention mechanism, so as to obtain the third generated feature, and the third generated feature is input into the speech interaction network as a target generated attention feature to combine with the third discriminant feature output by the third first dual-path cyclic attention network to generate the input data for the generation decoder.

[0103] It should be noted that the above example only illustrates the case where L is 3, and any value of L can be set according to actual needs.

[0104] Figure 6 The structural schematic diagram of the speech interaction network according to an embodiment of the present application is shown.

[0105] According to an embodiment of the present application, as Figure 6 shown, the target discriminant latent feature and the initial generated latent feature are weighted and fused by using the speech interaction network to obtain the target generated latent feature, including:

[0106] Generate a first fusion feature according to the target discriminant latent feature and the initial generated latent feature;

[0107] Perform convolution processing on the first fusion feature to obtain a first fusion convolution feature;

[0108] Generate a second fusion convolution feature according to the first fusion convolution feature and the time embedding feature;

[0109] Perform layer normalization processing on the second fusion convolution feature to obtain a fusion normalization feature;

[0110] Generate the target generated latent feature according to the target discriminant latent feature, the initial generated latent feature and the normalization feature.

[0111] According to an embodiment of the present application, after the speech interaction network receives the target discriminant latent feature and the initial generated latent feature, it first performs feature fusion on the two to obtain a first fusion feature, and then performs feature convolution on the first fusion feature to obtain a first fusion convolution feature.

[0112] According to an embodiment of the present application, according to the first fusion convolution feature and the time embedding feature t emb , generate a second fusion convolution feature, then perform layer normalization processing on the second fusion convolution feature, and activate it based on the activation function Sigmoid to obtain the fusion normalization feature. Finally, generate the target generated latent feature according to the target discriminant latent feature, the initial generated latent feature and the normalization feature.

[0113] Figure 7Shows a schematic structural diagram of a dual-path cyclic attention network according to an embodiment of the present application.

[0114] According to an embodiment of the present application, as Figure 7 shown, any one of the first dual-path cyclic attention network and the second dual-path cyclic attention network processes data in the following manner:

[0115] Perform dimensionality reorganization processing on the sixth input data to obtain a first reorganized feature, where the sixth input data includes discriminative convolutional features or target generation latent features; process the first reorganized feature using multiple long short-term attention networks to obtain a second reorganized feature, where the long short-term attention network includes a normalization layer, a bidirectional long short-term memory network, a multi-head attention mechanism layer, and a dimensionality reorganization layer; process the second reorganized feature using a convolutional gated linear network to obtain a first linear feature, where the convolutional gated linear network includes a first linear layer and a depth convolutional layer; process the second reorganized feature and the first linear feature using a linear block to obtain a second linear feature; generate a target output feature based on the second reorganized feature and the second linear feature, where the target output feature includes discriminative attention features or target generation attention features.

[0116] According to an embodiment of the present application, an exemplary description is given with the sixth input data being discriminative convolutional features. Based on , , and perform dimensionality reorganization processing on the sixth input data to obtain a first reorganized feature, then process the first reorganized feature using multiple long short-term attention networks to obtain a second reorganized feature, then process the second reorganized feature using a convolutional gated linear network to obtain a first linear feature, then process the second reorganized feature and the first linear feature using a linear block to obtain a second linear feature, and finally generate a target output feature based on the second reorganized feature and the second linear feature, and the target output feature is discriminative attention features, where , , and respectively represent the batch size, the number of time frames, the downsampled frequency dimension, and the network embedding dimension.

[0117] According to an embodiment of the present application, as Figure 7As shown, for each long short-term attention network, first, layer normalization is performed on the first reorganized feature input to the long short-term attention network to obtain a normalized vector. A bidirectional LSTM (Long Short-Term Memory) is used to process the normalized vector to obtain a long short-term memory vector. Then, a multi-head attention mechanism layer is used to process the long short-term memory vector to obtain an attention vector. Based on the attention vector and the first reorganized feature, a fusion vector is generated, and feature reorganization is performed on the fusion vector to obtain a second reorganized feature. Among them, multiple long short-term attention networks are cascaded in sequence, and the number can be specifically set according to actual needs, for example, it can be 2.

[0118] According to an embodiment of the present application, as Figure 7 shown, for the convolutional gated linear network, first, a first linear layer is used to process the second reorganized feature to obtain a linear vector, and then a depth convolutional layer (such as a depthwise convolutional layer) is used to process the linear vector and activate it based on the Mish activation function to obtain a first linear feature.

[0119] According to an embodiment of the present application, as Figure 7 shown, using a linear block to process the second reorganized feature and the first linear feature to obtain a second linear feature includes: using a second linear layer to process the second reorganized feature to obtain a third linear feature; generating a third reorganized feature according to the first linear feature and the third linear feature; using a third linear layer to process the third reorganized feature to obtain a second linear feature.

[0120] According to an embodiment of the present application, for the linear block, first, a second linear layer is used to linearly process the second reorganized feature to obtain a third linear feature, then feature fusion is performed on the third linear feature and the first linear feature output by the convolutional gated linear network to obtain a third reorganized feature, and then a third linear layer is used to process the third reorganized feature to obtain a second linear feature.

[0121] According to an embodiment of the present application, the first dual-path recurrent attention network and the second dual-path recurrent attention network efficiently capture feature dependencies in the time and frequency dimensions by using a combination of a bidirectional long short-term memory network and a multi-head attention mechanism layer, and use a convolutional gated linear unit composed of a linear layer, depthwise convolution, and Mish activation function to ensure the fusion of fine-grained channel attention information, enhancing the temporal modeling ability of the model.

[0122] According to an embodiment of the present application, in order to ensure the diffusion time-dependence of the speech generation network, Fourier embedding is used to integrate the time information (i.e., the amplitude state variable) of the diffusion process into the discriminative-generative joint model, specifically by fusing in the convolutional layers of the generative encoder and the generative decoder. The speech discriminative network can provide clues about the degraded speech as conditions to guide the generation process of the speech generation network, and the mapping process from the degraded speech to the clean speech (i.e., the enhanced speech signal) can facilitate the score estimation of the current diffusion state because the theoretical mean of the diffusion state variable is the weighted sum of the degraded speech amplitude spectrum and the clean speech amplitude spectrum. Therefore, a speech interaction network is used to connect between the speech discriminative network and the speech generation network in the recurrent attention network to transfer the complementary information of the speech discriminative network and the speech generation network. The hidden features from the speech discriminative network and the hidden features from the speech generation network are combined and a mask is generated through the Sigmoid function. The target discriminative hidden features are filtered through the mask and fused with the initial generative hidden features to obtain the final target generative hidden features.

[0123] According to an embodiment of the present application, the target frequency domain information includes target amplitude information and target complex number information, and the labeled speech information includes labeled amplitude information and labeled complex number information.

[0124] According to an embodiment of the present application, the model parameters of the initial enhancement model are iteratively adjusted according to the target score function, the target frequency domain information, and the labeled speech information to obtain a trained discriminative-generative joint model, including: calculating a discriminative loss result according to the target amplitude information, the target complex number information, the labeled amplitude information, and the labeled complex number information; calculating a generative loss result according to the target score function and the state parameters; generating a target loss result according to the discriminative loss result and the generative loss result; and iteratively adjusting the model parameters of the initial enhancement model according to the target loss result to obtain the discriminative-generative joint model.

[0125] According to an embodiment of the present application, the discriminative loss result is calculated according to the target amplitude information, the target complex number information, the labeled amplitude information, and the labeled complex number information , as shown in formula (3):

[0126] (3)

[0127] Wherein, is the mean squared error between the target amplitude information and the labeled amplitude information, is the mean squared error between the target complex number information and the labeled complex number information, is a weighting coefficient, which can be 0.5.

[0128] According to an embodiment of the present application, the generative loss result is calculated according to the target score function and the state parameters , as shown in formula (4):

[0129]

[0130] (4)

[0131] Wherein, is the Gaussian white noise in the voice training sample, and are the state parameters of the amplitude state variable at the diffusion time , that is, the theoretical mean and variance.

[0132] According to the embodiment of the present application, a target loss result is generated according to the discriminant loss result and the generation loss result , as shown in formula (5):

[0133] (5)

[0134] In a specific embodiment, during the voice enhancement process, when the voice generation network processes the state variable of the voice signal to be processed, the prediction score function obtained for the first time can be used to calculate the state variable of the next time based on this prediction score function, so as to iteratively use the voice generation network to perform operations in combination with the target discriminant hidden features output by the voice discriminant network. When the maximum number of iterations (for example, T = 15 - 20) is reached, the state variable calculated in the last iteration is combined with the target frequency domain information output by the voice discriminant network to generate an enhanced voice signal.

[0135] Figure 8 Shows a schematic diagram of the performance comparison between the discriminant-generation joint model and the baseline model according to the embodiment of the present application on the TIMIT-UNI dataset.

[0136] In a specific embodiment, a Hanning window with a window length of 512 points and a frame shift of 192 points are used for the audio with a sampling rate of 16 kHz to obtain the STFT spectrum, and the exponential compression factor is set to 0.3. The AdamW optimizer is used during training, the learning rate is set to 0.001, and the training period is 200 rounds. At the same time, in order to optimize the neural network weights, the exponential moving average method is used, and the decay rate is set to 0.999. The L2 norm of the gradient is clipped to 5.0.

[0137] According to the embodiment of the present application, multiple datasets are used for training and testing, including:

[0138] 1) WSJ0-UNI: The clean speech samples (i.e., labeled speech information) are from a regional daily newspaper dataset (WSJ0), and the real noise samples (i.e., speech training samples) are from the WHAM dataset. Multiple degradation types are introduced to each clean speech sample, and the degradation types include noise, reverberation, microphone frequency response, analog-to-digital converter effect, automatic gain control, and data transmission impact; specific degradations include additive noise, room impulse response convolution, band filtering, bit-depth adjustment, clipping, volume adjustment, resampling, and Global System for Mobile Communications data transmission compression;

[0139] 2) VBDMD: A widely used noise reduction evaluation dataset that only contains additive noise degradation. The original 48 kHz sampled audio is all resampled to 16 kHz for evaluating the noise reduction ability of the model.

[0140] 3) VBDMD-REVERB: Uses the clean speech samples from VBDMD and the reverberation generation method from WSJ0-UNI, with an average reverberation time of about 0.4 seconds, for evaluating the dereverberation ability of the model.

[0141] 4) VBDMD-SR: Applies a 12th-order Butterworth low-pass filter with a cut-off frequency of 4 kHz to the clean speech samples of the VBDMD test set to generate bandwidth-limited degraded speech, for evaluating the speech super-resolution ability of the model.

[0142] 5) TIMIT-UNI: Uses the same degradation scheme as WSJ0-UNI, but the clean speech samples are from the TIMIT dataset. Since the speech transcriptions of the TIMIT dataset are available, the impact of the model on the downstream Automatic Speech Recognition (ASR) performance can be evaluated.

[0143] According to the embodiments of the present application, WSJ0-UNI is used as the speech training set for training the model, and the test sets are from various datasets.

[0144] According to the embodiments of the present application, in the performance evaluation, multiple metrics are used to measure the overall quality of the enhanced speech signal, including: Perceptual Evaluation of Speech Quality (PESQ), Extended Short-Time Objective Intelligibility (ESTOI), three Composite Mean Opinion Scores (CSIG, CBAK, COVL), non-intrusive metrics based on the pre-trained wav2vec2.0 model (WV-MOS), Virtual Speech Quality Objective Listener (ViSQOL), Log-Spectral Distance (LSD), Structural Similarity Index Measure (SSIM), and Word Error Rate (WER) based on the pre-trained squeezeformer model. Except that the lower the LSD and WER are, the better, the higher the other metrics are, the better.

[0145] Table 1 shows the performance comparison between the voice enhancement method of this embodiment and multiple baseline models on the WSJ0-UNI dataset. The baseline models include 3 discriminative baseline models (Conv-TasNet, MANNER, CMGAN) and 4 generative baseline models (CDiffuSE, SGMSE+, StoRM, UNIVERSE++). The comparison items include model parameters (Para.), the number of multiply-accumulate operations (MACs), and voice quality evaluation metrics. It can be seen from Table 1 that compared with the discriminative baseline models, the voice enhancement method of this embodiment has improvements in all metrics. Compared with Conv-TasNet with the lowest computational complexity, the discriminative-generative joint model of this embodiment shows obvious advantages in performance. Compared with the discriminative baseline model CMGAN, the discriminative-generative joint model of this embodiment still has better performance, even though CMGAN has been specially optimized with a PESQ discriminator. Compared with diffusion-based generative methods, the discriminative-generative joint model of this embodiment has better performance and a lighter computational burden, indicating the efficiency of neural network design and the efficiency and effectiveness of the truncated diffusion scheme.

[0146] Table 1

[0147]

[0148] Table 2 shows the performance of the discriminative-generative joint model of this embodiment and the baseline models on the VBDMD dataset to verify the denoising ability of the model. CMGAN is superior to the generative baselines in most metrics, indicating that the discriminative model is sufficient for a single denoising task because it does not involve a large amount of information loss. It can be seen from Table 2 that the discriminative-generative joint model of this embodiment has achieved the best performance, demonstrating the potential of combining discriminative and generative modeling in a single denoising task.

[0149] Table 2

[0150]

[0151] Table 3 shows the performance of the discriminative-generative joint model of this embodiment and the baseline models on the VBDMD-REVERB dataset to evaluate the dereverberation ability of the model. The performance of the diffusion baseline SGMSE+ is better than that of CMGAN, which indicates that the diffusion model is effective in detecting the correlation between specific time-frequency intervals and the corresponding dry speech regions. The discriminative-generative joint model of this embodiment still shows leading performance in most metrics, demonstrating its applicability to the dereverberation task and its robustness and generalization to unseen data.

[0152] Table 3

[0153]

[0154] Table 4 shows the performance of the discriminative - generative joint model and the baseline model of this embodiment on the VBDMD - SR dataset to evaluate the speech bandwidth expansion ability of the model. The LSD and SSIM metrics indicate that the present invention more accurately restores the spectral structure compared to other baselines, while the CSIG and COVL metrics indicate an improvement in the perceived speech quality. The PESQ metric has the highest value on the degraded speech because PESQ is not designed specifically for speech super - resolution, and the low - frequency region, which is more important for the human hearing experience, dominates the PESQ score.

[0155] Table 4

[0156]

[0157] Figure 8 shows the speech enhancement and backend speech recognition performance of the discriminative - generative joint model and the baseline model of this embodiment on the TIMIT - UNI dataset. The discriminative - generative joint model of this embodiment obtains the highest PESQ score and significantly reduces the WER of the degraded speech. Compared with the discriminative baseline CMGAN, the generative baseline shows a higher WER, which is attributed to the pronunciation artifacts and speech confusion caused by the generative model. The discriminative - generative joint model of this embodiment can effectively combine the discriminative and generative modeling capabilities to improve the reconstruction accuracy and reduce artifacts, achieving a WER comparable to that of the discriminative baseline CMGAN.

[0158] Figure 9 shows a block diagram of a speech enhancement device according to an embodiment of the present application.

[0159] As Figure 9 shown, the speech enhancement device 900 includes an acquisition module 910, a processing module 920, and a generation module 930.

[0160] The acquisition module 910 is configured to acquire a speech signal to be processed, where the speech signal to be processed represents a speech signal with noise.

[0161] The processing module 920 is configured to input the speech signal to be processed into the discriminative - generative joint model to obtain predicted frequency - domain information and a predicted score function, where the discriminative - generative joint model includes a speech discriminative network, a speech interaction network, and a speech generation network. The speech interaction network is configured to fuse the hidden features in the speech discriminative network and the speech generation network so that the speech generation network generates a predicted score function according to the fused hidden features.

[0162] The generation module 930 is configured to generate an enhanced speech signal according to the predicted frequency - domain information and the predicted score function.

[0163] According to an embodiment of the present application, by inputting a voice signal to be processed into a discriminative - generative joint model, predicted frequency - domain information and a predicted score function are obtained. Based on the predicted frequency - domain information and the predicted score function, an enhanced voice signal is generated. Since the voice interaction network in the discriminative - generative joint model can fuse the hidden features output by the voice discrimination network and the voice generation network, the voice generation network is enabled to combine the fused hidden features for feature processing to obtain a predicted score function with a lower possibility of artifacts. Finally, by combining the predicted frequency - domain information, a relatively accurate enhanced voice signal can be obtained, and the computational complexity is small. Thus, the computing resources required for voice enhancement can be reduced.

[0164] In an embodiment of the present application, any plurality of modules, sub - modules, units, and sub - units, or at least part of the functions of any of them can be implemented in one module. Any one or more of the modules, sub - modules, units, and sub - units according to the embodiments of the present application can be split into multiple modules for implementation. Any one or more of the modules, sub - modules, units, and sub - units according to the embodiments of the present application can be at least partially implemented as a hardware circuit, such as a field - programmable gate array (FPGA), a programmable logic array (PLA), a system - on - chip, a system - on - substrate, a system - on - package, an application - specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub - modules, units, and sub - units according to the embodiments of the present application can be at least partially implemented as a computer program module, and when the computer program module runs, the corresponding functions can be executed.

[0165] It should be noted that the voice enhancement device part in the embodiments of the present application corresponds to the voice enhancement method part in the embodiments of the present application. For the description of the voice enhancement device part, please refer to the voice enhancement method part specifically, and details will not be repeated here.

[0166] Figure 10 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present application is shown. Figure 10 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0167] As Figure 10As shown, the electronic device 1000 according to an embodiment of the present application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 1001 may also include on-board memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0168] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to an embodiment of the present application by executing the program in the ROM 1002 and / or the RAM 1003. It should be noted that the program may also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 may also perform various operations of the method flow according to an embodiment of the present application by executing the program stored in the one or more memories.

[0169] According to an embodiment of the present application, the electronic device 1000 may further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The electronic device 1000 may further include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.

[0170] According to an embodiment of the present application, the method flow according to the embodiments of the present application can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above functions defined in the system of the embodiments of the present application are executed. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0171] The present application also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0172] An embodiment of the present application further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method provided by the embodiments of the present application. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the method provided by the embodiments of the present application.

[0173] When the computer program is executed by the processor 1001, the above functions defined in the system / apparatus of the embodiments of the present application are executed. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0174] The above describes the embodiments of the present application. However, these embodiments are only for illustrative purposes and not for limiting the scope of the present application. Although the above embodiments are described separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A speech enhancement method based on a joint discriminant-generative model, characterized in that: include: Acquire a speech signal to be processed, wherein the speech signal to be processed represents a speech signal with noise; Inputting the speech signal to be processed into a joint discriminant-generator model to obtain predicted frequency domain information and a predicted score function, wherein the joint discriminant-generator model includes a speech discriminant network, a speech interaction network and a speech generation network, and the speech interaction network is used to fuse latent features in the speech discriminant network and the speech generation network so that the speech generation network generates the predicted score function according to the fused latent features; An enhanced speech signal is generated according to the predicted frequency domain information and the predicted score function.

2. The method according to claim 1, characterized in that The joint discriminative-generative model is trained as follows: Acquire a speech training set, wherein the speech training set includes a plurality of speech training samples and label speech information corresponding to each of the speech training samples, and the speech training samples include initial frequency domain information and amplitude state variables of the training speech; For each of the speech training samples, based on the attention mechanism, a speech discrimination network is used to process the amplitude and phase in the initial frequency domain information to obtain target discrimination latent features and target frequency domain information; The target discriminant latent feature and the initial generation latent feature are weightedly fused by using a speech interaction network to obtain the target generation latent feature, wherein the initial generation latent feature is obtained by processing the amplitude state variable by the speech generation network, and the target discriminant latent feature and the target generation latent feature both represent the correlation between the speech training samples in the time dimension and the frequency dimension; Based on the attention mechanism, the amplitude state variable and the target generation latent feature are processed by the speech generation network to obtain a target score function, wherein the initial enhancement model includes the speech discrimination network, the speech generation network and the speech interaction network; The model parameters of the initial enhancement model are iteratively adjusted according to the target score function, the target frequency domain information and the label speech information to obtain a trained discriminant-generative joint model.

3. The method according to claim 2, characterized in that Based on the attention mechanism, the amplitude and phase in the initial frequency domain information are processed using the speech discrimination network to obtain target discrimination latent features and target frequency domain information, including: Using a discriminative encoder to encode the initial frequency domain information to obtain a discriminative convolution feature; Processing the discriminative convolutional features using a discriminative recurrent attention network to obtain a plurality of discriminative attention features, wherein the target discriminative latent features include the discriminative convolutional features or the discriminative attention features, and the discriminative recurrent attention network includes a plurality of first dual-path recurrent attention networks; The discriminative decoder is used to process the discriminative attention feature and the discriminative convolution feature output by the last of the first dual-path recurrent attention network to obtain the target frequency domain information.

4. The method according to claim 3, characterized in that The amplitude state variable and the target generation latent feature are input into the speech generation network, and a target score function is output for the amplitude state variable, including: Using a generative encoder to encode the amplitude state variable to obtain a generative convolution feature, and determining the generative convolution feature as an initial generative latent feature; Processing a plurality of target generated latent features using a generative recurrent attention network to obtain target generated attention features, wherein the generative recurrent attention network includes a plurality of second dual-path recurrent attention networks; Transmitting the target generated attention feature to the speech interaction network, so that the speech interaction network generates a target latent feature according to the target generated attention feature and the last discriminant attention feature output by the first dual-path recurrent attention network; The target latent features and the generated convolutional features are processed by using a generative decoder to obtain the target score function.

5. The method according to claim 4, characterized in that The discriminative encoder and the generative encoder perform data processing in the following manner: Processing first input data using a first convolution layer to obtain a first convolution feature, wherein the first input data includes the initial frequency domain information or the amplitude state variable; Processing the first convolution feature using m sub-band downsampling blocks to obtain a first output feature, wherein the first output feature includes the discriminant convolution feature or the generative convolution feature; The discriminant decoder and the generative decoder perform data processing in the following manner: Processing the third input data using n sub-band upsampling blocks to obtain upsampled features, wherein the third input data includes a first discriminant feature or a first generative feature, the first discriminant feature includes a discriminant attention feature and the discriminant convolution feature, and the first generative feature includes the target latent feature and the generative convolution feature; The up-sampled features are processed using a second convolutional layer to obtain second convolutional features, wherein the second convolutional features represent the target frequency domain information or the target score function.

6. The method according to claim 5, characterized in that Processing the first convolution feature using m sub-band downsampling blocks to obtain a first output feature includes: For the i-th subband downsampling block, perform frequency band segmentation processing on the second input feature to obtain a first frequency feature and a second frequency feature, wherein the second input feature includes the first convolution feature or the i-1-th normalized feature; Downsampling the first frequency feature to obtain a downsampled feature; Performing convolution processing on the second frequency feature to obtain a convolution feature; Merging the down-sampling feature, the convolution feature and the time embedding feature to obtain a merged feature; The merged features are subjected to layer normalization processing to obtain an i-th normalized feature, wherein when i=m, the i-th normalized feature represents the first output feature.

7. The method according to claim 4, characterized in that The number of the first dual-path recurrent attention network and the second dual-path recurrent attention network is L; The discriminative convolutional features are processed using a discriminative recurrent attention network to obtain multiple discriminative attention features, including: For the i-th first dual-path recurrent attention network, use the i-th first dual-path recurrent attention network to process the fourth input data and output the i-th discriminant feature, wherein when i=1, the fourth input data represents the discriminant convolution feature, when i≠1, the fourth input data represents the i-1th discriminant feature, and the target discriminant latent feature includes the discriminant convolution feature and the i-th discriminant feature; The generated convolutional features and the plurality of target generated latent features are processed by a generated recurrent attention network to obtain target generated attention features, including: For the i-th second dual-path recurrent attention network, use the i-th second dual-path recurrent attention network to process the fifth input data and output the i-th generated feature, wherein when i=1, the fifth input data includes the target generated latent feature generated by the voice interaction network according to the discriminant convolutional feature and the generated convolutional feature, and when i≠1, the fifth input data includes the target generated latent feature generated by the voice interaction network according to the i-1th discriminant feature and the i-1th generated feature; Among them, when i=L, the i-th generated feature represents the target generated attention feature.

8. The method according to claim 2 or 4, characterized in that: The target discriminant latent features and the initial generated latent features are weightedly fused by using a speech interaction network to obtain the target generated latent features, including: Generate a first fusion feature according to the target discriminant latent feature and the initial generated latent feature; Performing convolution processing on the first fused features to obtain first fused convolution features; Generate a second fused convolution feature according to the first fused convolution feature and the time embedding feature; Performing layer normalization processing on the second fused convolutional features to obtain fused normalized features; The target generated latent feature is generated according to the target discriminant latent feature, the initial generated latent feature and the normalized feature.

9. The method according to claim 4, characterized in that Any of the first dual-path recurrent attention network and the second dual-path recurrent attention network performs data processing in the following manner: Performing dimension reorganization processing on the sixth input data to obtain a first reorganized feature, wherein the sixth input data includes the discriminant convolution feature or the target generated latent feature; Processing the first reorganized features using multiple long short-term attention networks to obtain second reorganized features, wherein the long short-term attention networks include a normalization layer, a bidirectional long short-term memory network, a multi-head attention mechanism layer, and a dimension reorganization layer; Processing the second reorganized feature using a convolutional gated linear network to obtain a first linear feature, wherein the convolutional gated linear network includes a first linear layer and a deep convolutional layer; Processing the second recombined feature and the first linear feature using a linear block to obtain a second linear feature; Generate a target output feature according to the second recombined feature and the second linear feature, wherein the target output feature includes the discriminative attention feature or the target generated attention feature; The step of processing the second recombined feature and the first linear feature by using a linear block to obtain a second linear feature includes: Processing the second recombined features using a second linear layer to obtain a third linear feature; generating a third recombinant feature according to the first linear feature and the third linear feature; The third recombined features are processed using a third linear layer to obtain the second linear features.

10. The method according to claim 2, characterized in that The target frequency domain information includes target amplitude information and target complex number information, and the label voice information includes label amplitude information and label complex number information; The model parameters of the initial enhancement model are iteratively adjusted according to the target score function, the target frequency domain information and the label speech information to obtain a trained discriminant-generative joint model, including: Calculating a discrimination loss result according to the target amplitude information, the target complex number information, the label amplitude information and the label complex number information; Calculate and generate a loss result according to the target score function and the state parameter; Generate a target loss result according to the discrimination loss result and the generation loss result; The model parameters of the initial enhancement model are iteratively adjusted according to the target loss result to obtain the discriminant-generative joint model.