Speech Enhancement

By generating target time-frequency representation and combining attention mechanisms to utilize frequency and time correlation information, the problem of failure to fully utilize time-domain and frequency-domain information in the prior art is solved, and a more efficient voice enhancement effect is achieved and a pure voice signal is obtained.

CN113870878BActive Publication Date: 2025-07-04MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010617322.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-07-04
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

The existing speech enhancement methods fail to make full use of the correlation information of audio signals in the time and frequency domains, resulting in insufficient speech enhancement performance and it is difficult to effectively distinguish between speech components and noise components.

Method used

By generating the target time-frequency representation, the attention mechanism combines frequency correlation and time-dependent information to generate the target feature representation, the mask generation layer is used to enhance the speech components, and the decoder is used to convert it into the output audio signal.

Benefits of technology

Improves the performance of voice enhancement, obtains a purer voice signal, improves signal distortion ratio and objective voice quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870878B_ABST
    Figure CN113870878B_ABST
Patent Text Reader

Abstract

According to an implementation of the present disclosure, a scheme for speech enhancement is proposed. In this scheme, at least a target time-frequency representation indicating the intensity of an input audio signal varying with time at different frequencies is obtained. The input audio signal includes a speech component and a noise component. Frequency correlation information and time correlation information of the input audio signal are determined based on the target time-frequency representation. A target feature representation is generated based on the frequency correlation information, the time correlation information, and the target time-frequency representation. The target feature representation is used to distinguish the speech component and the noise component. An output audio signal is generated based on the target feature representation and the target time-frequency representation. In the output audio signal, the speech component is enhanced relative to the noise component. In this way, the performance of speech enhancement can be improved, which helps to obtain completely pure speech.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In voice recording or voice communication, voice is usually interfered by noise in the environment. That is, the recorded or transmitted audio signal will include both voice and noise. Voice enhancement aims to restore the noise-interfered voice to pure voice. Many fields such as video processing, audio processing, video conferencing, Voice over Internet Protocol (VoIP), speech recognition, hearing aids have a demand for voice enhancement technology. Existing voice enhancement methods can be classified into time-frequency (T-F) domain (hereinafter simply referred to as time-frequency domain) methods and time domain methods according to the signal domain in which they work. Summary of the Invention

[0002] According to an implementation of the present disclosure, a solution for voice enhancement is proposed. In this solution, at least a target time-frequency representation indicating the intensity of the input audio signal varying with time at different frequencies is obtained. The input audio signal includes a voice component and a noise component. The frequency correlation information and time correlation information of the input audio signal are determined based on the target time-frequency representation. A target feature representation is generated based on the frequency correlation information, time correlation information, and the target time-frequency representation. The target feature representation is used to distinguish the voice component and the noise component. An output audio signal is generated based on the target feature representation and the target time-frequency representation. In the output audio signal, the voice component is enhanced relative to the noise component. This solution can make full use of the correlation information of the audio signal in the time domain and the frequency domain. In this way, the performance of voice enhancement can be improved, which helps to obtain completely pure voice.

[0003] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Brief Description of the Drawings

[0004] Figure 1 A block diagram of a computing device capable of implementing multiple implementations of the present disclosure is shown;

[0005] Figure 2 An architecture diagram of a system for voice enhancement according to an implementation of the present disclosure is shown;

[0006] Figure 3 A block diagram of a correlation unit according to some implementations of the present disclosure is shown;

[0007] Figure 4 A block diagram of an attention block according to some implementations of the present disclosure is shown;

[0008] Figure 5 A block diagram of the change of data flow according to some implementations of the present disclosure is shown;

[0009] Figure 6 A block diagram showing the generation of a feature representation according to some implementations of the present disclosure; and

[0010] Figure 7 A flowchart showing a method for speech enhancement according to an implementation of the present disclosure.

[0011] In these figures, the same or similar reference signs are used to denote the same or similar elements. Detailed Description

[0012] The present disclosure will now be described with reference to several example implementations. It should be understood that these implementations are described only to enable those of ordinary skill in the art to better understand and thus implement the present disclosure, and do not imply any limitation on the scope of the present disclosure.

[0013] As used herein, the term "comprising" and its variants are to be construed as open-ended terms meaning "including but not limited to". The term "based on" is to be construed as "at least partially based on". The terms "one implementation" and "an implementation" are to be construed as "at least one implementation". The term "another implementation" is to be construed as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.

[0014] As used herein, a "neural network" is capable of processing an input and providing a corresponding output, and generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. A neural network used in deep learning applications typically includes many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence such that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also referred to as processing nodes or neurons), and each node processes the input from the previous layer. In this document, the terms "neural network", "network" and "neural network model" may be used interchangeably.

[0015] As used herein, "speech enhancement" refers to the task of restoring noisy speech to clean speech. "Speech enhancement" can be achieved by improving the quality of the speech signal itself, eliminating or reducing the noise signal and their combinations. Thus, herein, similar expressions such as "speech is enhanced relative to noise" may refer to the elimination or reduction of noise.

[0016] Example environment

[0017] Figure 1FIG. 100 is a block diagram of a computing device 100 capable of implementing various implementations of the present disclosure. It should be understood that Figure 1 the computing device 100 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described in the present disclosure. As Figure 1 shown, the computing device 100 includes a computing device 100 in the form of a general-purpose computing device. The components of the computing device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.

[0018] In some implementations, the computing device 100 may be implemented as various user terminals or service terminals having computing capabilities. The service terminal may be a server, a large computing device, etc. provided by various service providers. The user terminal may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the computing device 100 can support any type of user interface (such as a "wearable" circuit, etc.).

[0019] The processing unit 110 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 100. The processing unit 110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0020] The computing device 100 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 120 can include audio processing modules 122, and these program modules are configured to execute the functions of various implementations described herein. The audio processing modules 122 can be accessed and run by the processing unit 110 to implement the corresponding functions.

[0021] The storage device 130 can be removable or non-removable media and can include machine-readable media that can be used to store information and / or data and can be accessed within the computing device 100. The computing device 100 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 1 a disk drive for reading from or writing to a removable, non-volatile disk and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.

[0022] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 100 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the computing device 100 can operate in a networked environment using a logical connection with one or more other servers, personal computers (PCs), or another general network node.

[0023] The input device 150 can be one or more various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 160 can be one or more output devices, such as a display, speaker, printer, etc. The computing device 100 can also communicate with one or more external devices (not shown) as needed via the communication unit 140, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 100, or communicate with any device that enables the computing device 100 to communicate with one or more other computing devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0024] In some implementations, in addition to being integrated on a single device, some or all of the various components of computing device 100 may also be provided in the form of a cloud computing architecture. In a cloud computing architecture, these components may be remotely located and may work together to implement the functions described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services that do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network (such as the Internet). For example, a cloud computing provider provides applications over a wide area network, and they can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be consolidated at a remote data center location or they may be dispersed. The cloud computing infrastructure may provide services through a shared data center, even though they appear as a single access point for the user. Thus, the components and functions described herein may be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they may also be provided from a conventional server, or they may be installed directly or otherwise on a client device.

[0025] Computing device 100 may be used to implement speech enhancement in various implementations of this disclosure. As Figure 1 shown, computing device 100 may receive an input audio signal 170 through input device 150, which includes a speech component 171 and a noise component 172. The input audio signal 170 is also referred to as the audio signal to be processed. As shown, the input audio signal 170 is presented as a waveform in the time domain. In some implementations, the input audio signal 170 may be a pre-recorded and stored segment of an audio signal, such as a part of an audio file. This implementation is also referred to as an offline implementation in this document. In other implementations, the input audio signal 170 may be a segment of an audio signal generated in real time, such as audio data generated in real time during a video conference.

[0026] Computing device 100 may implement the speech enhancement schemes described herein to generate an output audio signal 180. In the output audio signal 180, the speech component 171 is enhanced relative to the noise component 172. In Figure 1 the example of, only the speech component 171 is retained and the noise component 172 is eliminated.

[0027] As mentioned above, existing speech enhancement methods include time-frequency domain methods and time domain methods. The time domain methods use an encoder-decoder architecture to directly model the waveform in the time domain that includes both the speech component and the noise component, and perform separation of the speech component and the noise component on the output of the encoder.

[0028] The time-frequency domain method is based on the two-dimensional spectrogram. In a typical time-frequency domain method, the short-time Fourier transform (STFT) is used to convert the original audio signal into a spectrogram. Then, a separation network predicts a T-F mask, and this T-F mask is applied to the spectrogram to generate an output spectrogram. Finally, the output spectrogram is converted into an audio signal in the time domain, i.e., a waveform in the time domain, through the inverse STFT (ISTFT). Compared with the time domain method, the time-frequency domain method can benefit from the rich auditory pattern information in the spectrogram.

[0029] In some traditional time-frequency domain schemes, a frequency transformation block (FTB) is used at the front end of the separation network to capture the harmonic correlation in the frequency domain. In the later part of the separation network, a bidirectional long short-term memory network (Bi-LSTM) is used to capture the temporal correlation. In some other traditional time-frequency domain schemes, the spectrogram is processed as a time series, and a transformer model is adopted to capture the long-range temporal correlation.

[0030] The inventors of the present application have realized that the correlation of speech components in the frequency domain (also known as frequency correlation) is different from that of noise components in the frequency domain. Similarly, the correlation of speech components in the time domain (also known as temporal correlation) is also different from that of noise components in the frequency domain. Therefore, the correlations in the frequency domain and the time domain are important for denoising performance. However, in the above traditional schemes, the correlations in the frequency domain and the time domain are considered and learned separately. The separation of these two types of correlations results in the inability to fully fuse the information in the time domain and the frequency domain, which hinders the improvement of speech enhancement performance.

[0031] The above discusses some problems existing in the current speech enhancement schemes. According to the implementation of the present disclosure, a speech enhancement scheme is provided, aiming to solve one or more of the above problems and other potential problems. In this scheme, a time-frequency representation (e.g., spectrogram) of an input audio signal including speech components and noise components is obtained. Based on this time-frequency representation, according to the attention mechanism, the frequency correlation information and the temporal correlation information of the input audio signal are determined. Then, based on the temporal correlation information, the frequency correlation information, and the time-frequency representation, a target feature representation for distinguishing speech components and noise components is generated. Based on the target feature representation and the time-frequency representation, an output audio signal is generated. In the output audio signal, the speech components are enhanced relative to the noise components. The speech enhancement scheme proposed herein can fully utilize the correlation information of the audio signal in the time domain and the frequency domain. In this way, the performance of speech enhancement can be improved, which helps to obtain completely pure speech.

[0032] The following further describes various exemplary implementations of this scheme in detail in combination with the accompanying drawings.

[0033] System architecture

[0034] Figure 2 shows an architectural diagram of a system 200 for speech enhancement according to an implementation of the present disclosure. The system 200 can be implemented in Figure 1 a computing device 100. For example, in some implementations, the system 200 can be implemented as Figure 1 at least a part of the audio processing module 122 of the computing device 100, that is, implemented as a computer program module. As Figure 2 shown, the system 200 generally can include a separation network 250 and a decoder 240. The separation network 250 can include at least one correlation unit. Figure 2 Shows a series of correlation units 220-1, 220-2... 220-N (collectively or individually referred to as "correlation unit 220", where N≥1). It should be understood that the structure and function of the system 200 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure. Implementations of the present disclosure can also be implemented in different structures and / or functions.

[0035] As Figure 2 shown, first, a time-frequency representation 201 is obtained, which is also referred to as the "target time-frequency representation" herein. The time-frequency representation 201 at least indicates the intensity of the input audio signal 170 varying with time at different frequencies. In Figure 2 it, an example of the time-frequency representation 201 is shown in the form of a spectrogram. In some implementations, the system 200 can directly obtain the time-frequency representation 201 without processing the input audio signal 170.

[0036] In some implementations, the system 200 can process the input audio signal 170 to generate at least a part of the time-frequency representation 201. Although not shown, the system 200 can use the STFT to generate the time-frequency representation of the input audio signal 170. In some implementations, for example, in the offline implementation mentioned above, the input audio signal 170 can be a pre-captured audio signal segment with a predetermined length (e.g., 3 seconds). In this implementation, the time-frequency representation of the input audio signal 170 is the time-frequency representation 201.

[0037] In some implementations, such as the online implementation mentioned above, the input audio signal 170 can be an audio signal generated in real time. In such an implementation, to meet the real-time requirement, the duration of the input audio signal 170 is usually short, for example, only a few tens of milliseconds. At the same time, to obtain high-quality speech enhancement, the correlation of the audio signal over a longer period of time needs to be utilized. Therefore, in such an implementation, in addition to the time-frequency representation of the input audio signal 170 (also referred to as the "second time-frequency representation" in this article), the time-frequency representation 201 can also include the time-frequency representation of the processed audio signal (also referred to as the "first time-frequency representation" in this article). That is, the time-frequency representation 201 is a combination of the first time-frequency representation and the second time-frequency representation.

[0038] The processed audio signal can occur before the input audio signal 170, for example, immediately before the input audio signal 170 in time. The input audio signal 170 and the processed audio signal can have a predetermined total length. It should be understood that the time lengths mentioned above are only illustrative and are not intended to limit the scope of the present disclosure.

[0039] The time-frequency representation 201 with complex values can be represented by , where T represents the number of time segments in the time domain and F represents the number of frequency bands in the frequency domain. As Figure 2 shown, the time-frequency representation 201 is input into the separation network 250.

[0040] In some implementations, the separation network 250 can include one or more convolutional layers 210, such as a 1×7 convolutional layer and a 7×1 convolutional layer. In this article, a "m×n convolutional layer" refers to a convolutional layer with a convolutional kernel size of m×n. The time-frequency representation 201 is fed into the one or more convolutional layers 210 to generate a feature representation where C represents the number of channels. In some other implementations, the separation network 250 may not include one or more convolutional layers 210. The time-frequency representation 201 can be directly input into the correlation unit 220-1.

[0041] The correlation unit 220-1 can generate a feature representation S 0 based on the time-frequency representation 201 or the feature representation S 1 . The other correlation units 220-2 to 220-N can generate new feature representations based on the feature representation output by the previous correlation unit. Generally, the feature representation generated by the i-th correlation unit 220 can be represented by S i , where i ∈ {1, 2,..., N}. In some implementations, such as in the offline implementation, the feature representation S i can have the same dimension as the feature representation S 0 , that is, In some other implementations, such as in the online implementation, the feature representation Si can be less than the feature representation S in the time dimension 0 , for example, will be described below with reference to Figure 5 .

[0042] One or more correlation units 220 can generate the feature representation S by utilizing the frequency correlation and time correlation of the audio signal i . In some implementations, the frequency correlation and time correlation can be utilized simultaneously in each correlation unit 220. In some implementations, the frequency correlation and time correlation can be utilized simultaneously in at least one correlation unit, and only the frequency correlation is utilized in other correlation units. For example, the frequency correlation and time correlation can be utilized simultaneously in at least the first correlation unit 220-1, and only the frequency correlation is utilized in other correlation units 220-2 to 220-N.

[0043] In this article, the feature representations S 1 , S 2 ......S N-1 generated by the first (N-1) correlation units are also referred to as "intermediate feature representations", and the feature representation S N generated by the last correlation unit 220-N is also referred to as the "target feature representation". The intermediate feature representation and the target feature representation are used to distinguish the speech component and the noise component.

[0044] The feature representation S N generated by the last correlation unit 220-N is fed to the mask generation layer 230. The mask generation layer 230 generates a mask M∈T×F×C N based on the feature representation S m , where C m is the number of channels. As an example, the mask generation layer can be a head layer, which can be implemented as a 1×1 convolutional layer and C m = 2.

[0045] As Figure 2 shown, the separation network 250 generates the mask M as an output. Then, the masked time-frequency representation S out can be determined by applying the mask M to at least a part of the time-frequency representation 201. It can be understood that in the masked time-frequency representation S out , the speech component is enhanced relative to the noise component. Therefore, the masked time-frequency representation S out can also be referred to as the enhanced time-frequency representation.

[0046] In some implementations, such as in an offline implementation, the mask M can be applied to the complete time-frequency representation 201. As an example, if the mask generation layer 230 is implemented as a 1×1 convolutional layer and C m = 2, the masked time-frequency representation S out can be determined according to the following formula

[0047] S out = f(M) ⊙ g(S in ) (1)

[0048] where f(M) = tanh(|M|) * M / |M|, g(S in ) = S in and ⊙ represents complex multiplication.

[0049] In other implementations, such as in an online implementation, the time-frequency representation 201 can include a first time-frequency representation of the processed audio signal and a second time-frequency representation of the input audio signal 170, as described above. Thus, in such an implementation, the mask M can be applied only to the part of the time-frequency representation 201 corresponding to the second time-frequency representation.

[0050] The masked time-frequency representation S out is fed to the decoder 240 and is converted by the decoder 240 into the output audio signal 180. In some implementations, the decoder 240 can be a decoder based on a fixed transform. For example, the decoder 240 can implement the inverse short-time Fourier transform (ISTFT) corresponding to the short-time Fourier transform (STFT) used to generate the time-frequency representation of the input audio signal 170.

[0051] In some implementations, the decoder 240 can be a trained decoder. Such a decoder 240 can be trained to convert the audio signal intensity varying over time at different frequencies into a waveform varying over time. Although the decoder does not have to perform exactly the opposite operation of the encoder, in some implementations the following method can be used to ensure the convergence of the encoder network. For example, the network structure of the decoder 240 can be configured to be the same as the convolutional implementation of the ISTFT. Specifically, the encoder 240 can be implemented as a one-dimensional transposed convolutional layer, the kernel size and stride of which are the window length and hop size used in the STFT.

[0052] In such an implementation, the decoder 240 can be trained together with the separation network 250. The system 200 can be implemented as a fully end-to-end learnable system. In this way, the learnability of the system 200 is improved, which helps to improve the performance of speech enhancement.

[0053] Attention mechanism

[0054] The overall architecture of the system 200 for speech enhancement has been described above. Some implementations of the correlation unit 220 therein will be described below. Figure 3 A block diagram of the correlation unit 220 according to some implementations of the present disclosure is shown. The correlation unit 220 can determine the output feature representation S based on the input feature representation S i-1 to determine the output feature representation S i .

[0055] Generally, the correlation unit 220 can include two parts. The first part includes a set of convolutional layers 311, 312, 313, 314, and the second part includes an attention block 320. During the training process of the system 200, the convolutional layers 311, 312, 313, 314 can learn the local correlations of the audio signal. Therefore, in processing the input audio signal 170, the local correlations of the input audio signal 170 can be utilized through the convolutional layers 311, 312, 313, 314. As an example, the convolutional layers 311, 312, 313, 314 can be implemented as 3×3 convolutional layers.

[0056] As Figure 3 shown, in some implementations, a skip connection (also known as a residual connection) can be employed between every two convolutional layers. The use of skip connections can improve the performance of speech enhancement by increasing the network depth.

[0057] The number of convolutional layers, connection methods, and convolutional kernel sizes described above are merely illustrative and are not intended to limit the scope of the present disclosure. In implementations according to the present disclosure, any suitable number of convolutional layers and convolutional kernels of suitable sizes can be employed. In some implementations, the number of convolutional layers in different correlation units 220 can also be different.

[0058] The second part of the correlation unit 220 can include an attention block 320. During the training process of the system 200, the attention block 320 can learn the long-range correlations (also known as global correlations) of the audio signal based on the attention mechanism. Therefore, in processing the input audio signal 170, the long-range correlations of the input audio signal 170 can be utilized through the attention block 320. Specifically, the attention block 230 receives the feature representation S generated by utilizing the local correlations from the convolutional layer 314 con , and utilizes the attention mechanism to determine the feature representation S based on the feature representation S con and the long-range correlations. i .

[0059] In a time-frequency representation (e.g., spectrogram), there are long-range correlations both in the time domain (i.e., along the time axis) and in the frequency domain (i.e., along the frequency axis). It can be understood that an audio signal, as a time series, contains global correlations along the time axis; along the frequency axis, there are harmonic correlations. In this article, the correlation along the time axis is also referred to as time correlation, and the correlation along the frequency axis is also referred to as frequency correlation.

[0060] The attention block 320 can utilize time correlation and frequency correlation in parallel. This kind of attention is also referred to as "Dual-Path Attention Block (DAB)". In order to fully utilize time correlation and frequency correlation without incurring excessive computational consumption, a lightweight DAB solution is proposed herein.

[0061] Figure 4 A block diagram of the attention block 320 according to some implementations of the present disclosure is shown. In Figure 4 the example, the feature representation 401 (also denoted as S con ) from the convolutional layer 314 has dimensions of T×F×C. As described above, T means that the time-frequency representation 201 is associated with T frequency bands in the time domain, F means that the time-frequency representation 201 is associated with F frequency bands in the frequency domain, and C represents the number of channels. It can be understood that for the correlation unit 220-1, the feature representation 401 is the convolved time-frequency representation 201; for the other correlation units 220-2 to 220-N, the feature representation 401 is the feature representation generated by the previous correlation unit, i.e., the intermediate feature representation as referred to herein.

[0062] On the frequency path, the two-dimensional feature representation 401 with C channels is decomposed into TC (i.e., T×C) frequency-domain feature representations 410-1 to 410-8, which are collectively or individually referred to as the frequency-domain feature representations 410. Each frequency-domain feature representation 410 corresponds to one of the T time periods and includes multiple features on the F frequency bands. Each frequency-domain feature representation 410 can be implemented as a vector of dimension 1×F.

[0063] As shown by the arrow 420, an attention mechanism is applied to each frequency-domain feature representation 410 along the frequency axis, resulting in the updated frequency-domain feature representation 430. In this article, the set of the updated frequency-domain feature representations by applying the attention mechanism is referred to as the frequency correlation information 430. If there are multiple correlation units 220, the frequency correlation information in the attention block 320 of the second to the Nth correlation units 220 can also be referred to as the updated frequency correlation information.

[0064] In some implementations, a self-attention mechanism can be used to learn an attention map for each frequency-domain feature representation (which can be regarded as a sample) 410. Each element in the attention map indicates the correlation between corresponding frequency bands of the audio signal.

[0065] The inventors of the present application have realized that the harmonic correlation of an audio signal in the frequency domain is sample-independent, and in some classical speech enhancement schemes, a uniform non-linear function is usually applied along the frequency axis to reproduce harmonics. Therefore, in some implementations, a sample-independent attention mechanism can be applied along the frequency axis. Specifically, in such an implementation, the application of the attention mechanism as shown by arrow 420 can be achieved through a fully-connected layer. The parameters of this fully-connected layer can be learned during the process of training system 200 using the reference audio signal. In other words, the parameters of this fully-connected layer indicate the correlation between the reference audio signal among F frequency bands. The process of applying the sample-independent attention mechanism to a certain frequency-domain feature representation 410 is equivalent to the process of multiplying a 1×F vector by an F×F weight matrix. In this article, the parameters of this fully-connected layer representing frequency correlation are also referred to as frequency-domain weighted information.

[0066] Similarly, in the time path, the two-dimensional feature representation 401 with C channels is decomposed into FC (i.e., F×C) time-domain feature representations 440-1 to 440-8, which can also be collectively or individually referred to as time-domain feature representations 440. Each time-domain feature representation 440 corresponds to one of the F frequency bands and includes multiple features in T time periods. Each time-domain feature representation 440 can be implemented as a vector of dimension T×1.

[0067] As shown by arrow 450, an attention mechanism is applied to each time-domain feature representation 440 along the time axis to obtain an updated time-domain feature representation 440. In this article, the set of time-domain feature representations updated by applying the attention mechanism is referred to as time correlation information 460. If there are multiple correlation units 220, the time correlation information in the attention blocks 320 of the second to the Nth correlation units 220 is also referred to as updated time correlation information.

[0068] In some implementations, a self-attention mechanism can be used to learn an attention map for each time-domain feature representation (which can be regarded as a sample) 440. Each element in the attention map indicates the correlation between corresponding time periods of the audio signal.

[0069] Similar to the fact that the harmonic correlation is sample-independent, the inventors of the present application realized that the temporal correlation is time-invariant. Therefore, in some implementations, a sample-independent attention mechanism can also be utilized along the time axis. Specifically, in such an implementation, the application of the attention mechanism as shown by arrow 450 can be achieved through a fully-connected layer. The parameters of this fully-connected layer can be learned during the process of training system 200 using the reference audio signal. In other words, the parameters of this fully-connected layer indicate the correlation between the reference audio signals over T time periods. The process of applying a sample-independent attention mechanism to a time-domain feature representation 440 is equivalent to the process of multiplying a T×1 vector by a T×T weight matrix. In this document, the parameters of this fully-connected layer representing temporal correlation are also referred to as time-domain weighted information.

[0070] In some implementations, a sample-independent attention mechanism can be applied both on the frequency path and the time path. In such an implementation, the attention block 320 can also be referred to as a sample-independent dual-path block (SDAB). The application of the sample-independent attention mechanism can reduce the computational complexity of the network of system 200 without degrading the speech enhancement performance.

[0071] Continuing to refer to Figure 4 , after obtaining the temporal correlation information 430 and the frequency correlation information 460, the feature representation S is determined based on the frequency correlation information 430, the temporal correlation information 460, and the feature representation 401 i . Specifically, the frequency correlation information 430 including TC frequency-domain feature representations can be reconstructed into a feature representation with dimensions of T×F×C, and the temporal correlation information 460 including FC time-domain feature representations can be reconstructed into a feature representation with dimensions of T×F×C. The two reconstructed feature representations are combined with the feature representation 401 to form the feature S i .

[0072] As an example, in Figure 4 , the two reconstructed feature representations and the feature representation 401 can be concatenated by channel. Then, a convolutional layer 470 can be applied to the concatenated feature representation, thereby obtaining a feature representation S with dimensions of T×F×C i . As another example, the feature representation S with dimensions of T×F×C can also be obtained by directly adding the features at the same positions in the two reconstructed feature representations and the feature representation 401 i . The implementations of the present disclosure are not limited in this regard.

[0073] The above referred to Figure 4 the DAB. Continuing to refer to Figure 2. In some implementations, each of the correlation units 220-1 to 220-N may include a DAB. In other words, temporal correlation and frequency correlation can be utilized in each correlation unit 220. In some other implementations, at least one of the correlation units 220-1 to 220-N may include a DAB, and the other correlation units may utilize the attention mechanism only on the frequency path.

[0074] As an example, the dual-path attention mechanism can be utilized only in the first correlation unit 220-1, while the attention mechanism is applied only on the frequency path in the subsequent correlation units 220-2 to 220-N. Utilizing temporal correlation in the front part of the classification network 250 helps improve speech enhancement performance.

[0075] As mentioned above, in some implementations, the dimension of the target feature representation in the time domain can be smaller than the dimension of the time-frequency representation 201 in the time domain. For example, in an online implementation, to meet the real-time requirement, the duration of the input audio signal 170 is usually short. To apply long-range correlation in the time domain, the processed audio signal needs to be utilized. Figure 5 FIG. 500 is a block diagram showing changes in the data flow in the system 200 according to some implementations of the present disclosure. The following will be described in conjunction with Figure 2 to describe Figure 5 . As Figure 5 shown, the time-frequency representation 201 includes a first time-frequency representation 501 of the processed audio signal and a second time-frequency representation 502 of the input audio signal 170. In the Figure 5 example, the unprocessed input audio signal 170 involves 4 time segments in the time domain, which corresponds to the second time-frequency representation 502 shown in shadow.

[0076] The convolutional layer 210 can generate a feature representation 511 based on the time-frequency representation 201, where the shaded part 503 is associated with the second time-frequency representation 502. That is, the part 503 is determined based on the second time-frequency representation 502. The part not shown in shadow is not updated, so the data for processing the previous input audio signal can be used.

[0077] The feature representation 511 is fed to the first correlation unit 220-1. The correlation unit 220-1 generates a feature representation 512 by applying one or more convolutional layers (e.g., the convolutional layers 311 to 314 shown in Figure 3 ) to the feature representation 511. In the feature representation 512, the shaded part 504 is associated with the second time-frequency representation 502. That is, the part 504 is determined based on the second time-frequency representation 502. The part not shown in shadow is not updated, so the data for processing the previous input audio signal can be used.

[0078] Next, the correlation unit 220-1 generates a feature representation 513 by applying DAB to the feature representation 512. The feature representation 513 may include only the portion associated with the second time-frequency representation 502. Figure 5 It can be seen that, relative to feature representation 512, feature representation 513 is shortened in the time domain.

[0079] Figure 6 A block diagram 600 for generating feature representations according to some implementations of the present disclosure is shown. Figure 4 and Figure 5 To describe Figure 6 . Feature representation 512 can be regarded as an implementation of feature representation 401. On the time path, feature representation 512 is decomposed into multiple time-domain feature representations, and by applying an attention mechanism along the time axis, time correlation information 660 is determined. Time correlation information 660 can be regarded as an implementation of time correlation information 460. A portion of the information in time correlation information 660 (also referred to as first portion of information 661) is associated with the second time-frequency representation 502. That is, the determination of the first portion of information 661 is at least partially associated with the second time-frequency representation 502.

[0080] On the frequency path, the feature representation 512 is decomposed into a plurality of frequency domain feature representations, and by applying an attention mechanism along the frequency axis, the frequency correlation information 630 is determined. The frequency correlation information 630 can be regarded as an implementation of the frequency correlation information 430. A portion of the frequency correlation information 630 (also referred to as the second portion of information 631) is associated with the second time-frequency representation 502. For example, the second portion of information 631 corresponds to the first portion of information 661 on the time axis.

[0081] Next, the feature representation 513 shown in gray may be determined based on the first part of information 661, the second part of information 631, and the feature representation 512 (specifically, the portion of the feature representation 512 corresponding to the first part of information 661 and / or the second part of information 661 on the time axis). It should be understood that the portion shown in the dotted box may be determined or not output. The determination of the feature representation 513 is similar to the above reference numeral 514. Figure 4 The description is similar and will not be repeated here.

[0082] Continue to refer Figure 5 , compared with the time-frequency representation 201, the feature representation 513 generated by the first correlation unit 220-1 is shortened in the time domain, but remains unchanged in the frequency domain. Therefore, the attention mechanism will only be applied in the frequency domain in the subsequent correlation units 220-2 to 220-6, that is, only the frequency correlation will be utilized.

[0083] Specifically, the correlation unit 220-2 generates the feature representation 514 by applying one or more convolutional layers (e.g., the convolutional layers 311 to 314 shown in Figure 3 ) to the feature representation 513. In the example of Figure 5 , the feature representation 514 is shortened in the time domain due to the convolutional operation with respect to the feature representation 513. Next, the correlation unit 220-2 generates the feature representation 515 by applying an attention mechanism in the frequency domain.

[0084] The correlation unit 220-3 generates the feature representation 516 by applying one or more convolutional layers (e.g., the convolutional layers 311 to 314 shown in Figure 3 ) to the feature representation 515. In the example of Figure 5 , the feature representation 516 is shortened in the time domain due to the convolutional operation with respect to the feature representation 515. Next, the correlation unit 220-3 generates the feature representation 517 by applying an attention mechanism in the frequency domain.

[0085] In a correlation unit not shown, similar operations are performed. The last correlation unit 220-6 applies the attention mechanism in the frequency domain to the feature representation 519 to determine the feature representation 520 as the target feature representation. As can be seen from Figure 6 , the feature representation 520 is associated with 4 time periods in the time domain, which corresponds to the duration of the input audio signal 170.

[0086] The above describes an example implementation of processing a short-duration input audio signal using a processed audio signal. It should be understood that the number of correlation units described, the number of time periods associated in the time domain, etc. are merely illustrative and are not intended to limit the scope of the present disclosure. In this way, temporal correlation and frequency correlation can be utilized in parallel to improve the speech enhancement effect without affecting the timeliness requirements in an online scenario. Figure 5 and Figure 6

[0087] Training of the voice enhancement system

[0088] Figure 2 Continuing to refer to Figure 2 , the system 200 can be trained using a reference audio signal including a reference speech signal and a reference noise signal. The reference speech signal can be a clean speech signal. The reference speech signal and the reference noise signal are synthesized (e.g., by weighting) into a reference audio signal to train the system 200.

[0089] In some implementations, the reference noise signal can include various types of noise. The system 200 trained in this way is versatile and applicable to speech enhancement in various scenarios.

[0090] In some implementations, the reference noise signal may include a specific type or types of noise (e.g., the sound of an engine, etc.). The system 200 trained thereby is particularly suitable for speech enhancement in a specific scenario. There is a specific type or types of noise in this specific scenario.

[0091] In some implementations, the system 200 can be trained under the architecture of the T-F domain. That is, the loss function can be calculated between the time-frequency representation S in and the masked time-frequency representation S out .

[0092] In some implementations, a cross-domain training process can be implemented. That is, the loss function can be calculated between the input audio signal 170 and the output audio signal 180. In this implementation, the speech enhancement performance of the system 200 can be further improved. For example, both the signal distortion ratio (SDR) and the perceptual evaluation of speech quality (PESQ) can be improved.

[0093] Example method and example implementation

[0094] Figure 7 FIG. shows a flowchart of a method 700 for speech enhancement according to some implementations of the present disclosure. The method 700 can be implemented by the computing device 100, for example, at the audio processing module 122 that can be implemented in the memory 120 of the computing device 100.

[0095] As Figure 7 shown, at block 710, the computing device 100 obtains a target time-frequency representation that at least indicates the intensity of the input audio signal varying with time at different frequencies. The input audio signal includes a speech component and a noise component. At block 720, the computing device 100 determines frequency correlation information and time correlation information of the input audio signal based on the target time-frequency representation. At block 730, the computing device 100 generates a target feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the target time-frequency representation. At block 740, the computing device 100 generates an output audio signal based on the target feature representation and the target time-frequency representation. In the output audio signal, the speech component is enhanced relative to the noise component.

[0096] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation and the updated frequency correlation information.

[0097] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information and the time correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation, the updated frequency correlation information, and the updated time correlation information.

[0098] In some implementations, obtaining the target time-frequency representation includes: obtaining a first time-frequency representation of a processed audio signal before the input audio signal occurs, the first time-frequency representation indicating the intensity of the processed audio signal varying with time at different frequencies; determining a second time-frequency representation of the input audio signal, the second time-frequency representation indicating the intensity of the input audio signal varying with time at different frequencies; and combining the first time-frequency representation and the second time-frequency representation into the target time-frequency representation.

[0099] In some implementations, generating the target feature representation includes: determining a first part of information in the time correlation information associated with the second time-frequency representation; determining a second part of information in the frequency correlation information associated with the second time-frequency representation; and determining the target feature representation based on the first part of information, the second part of information, and the target time-frequency representation.

[0100] In some implementations, the target time-frequency representation is associated with multiple frequency bands in the frequency domain and multiple time periods in the time domain. Determining the frequency correlation information and the time correlation information includes: obtaining a frequency-domain feature representation and a time-domain feature representation by processing the time-frequency representation, the frequency-domain feature representation including multiple features in multiple frequency bands within one of the multiple time periods, and the time-domain feature representation including multiple features in multiple time periods on one of the multiple frequency bands; determining the frequency correlation information based on the frequency-domain feature representation and frequency-domain weighting information, the frequency-domain weighting information indicating the degree of correlation between reference audio signals in the multiple frequency bands; and determining the time correlation information based on the time-domain feature representation and time-domain weighting information, the time-domain weighting information indicating the degree of correlation between reference audio signals in the multiple time periods.

[0101] In some implementations, the frequency-domain weighting information and the time-domain weighting information are determined based on the reference audio signal including a reference speech signal and a reference noise signal.

[0102] In some implementations, generating the output audio signal includes: determining a masked time-frequency representation by applying a mask generated based on the target feature representation to at least a portion of the target time-frequency representation; and converting the masked time-frequency representation into the output audio signal.

[0103] In some implementations, converting the masked time-frequency representation into the output audio signal includes: converting the masked time-frequency representation into the output audio signal by applying the masked time-frequency representation to a trained decoder, where the trained decoder is configured to convert the audio signal intensities varying over time at different frequencies into a waveform varying over time.

[0104] It can be seen from the above description that the speech enhancement scheme according to the implementations of the present disclosure can make full use of the correlation information of the audio signal in the time domain and the frequency domain. In this way, the performance of speech enhancement can be improved, which helps to obtain completely pure speech.

[0105] Some example implementations of the present disclosure are listed below.

[0106] In one aspect, the present disclosure provides a computer-implemented method. The method includes: obtaining a target time-frequency representation that at least indicates the intensities of an input audio signal varying over time at different frequencies, where the input audio signal includes a speech component and a noise component; determining the frequency correlation information and the time correlation information of the input audio signal based on the target time-frequency representation; generating a target feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the target time-frequency representation; and generating an output audio signal based on the target feature representation and the target time-frequency representation, where the speech component is enhanced relative to the noise component in the output audio signal.

[0107] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation and the updated frequency correlation information.

[0108] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information and the time correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation, the updated frequency correlation information, and the updated time correlation information.

[0109] In some implementations, obtaining the target time-frequency representation includes: obtaining a first time-frequency representation of a processed audio signal before the input audio signal occurs, the first time-frequency representation indicating the intensity of the processed audio signal varying with time at different frequencies; determining a second time-frequency representation of the input audio signal, the second time-frequency representation indicating the intensity of the input audio signal varying with time at different frequencies; and combining the first time-frequency representation and the second time-frequency representation into the target time-frequency representation.

[0110] In some implementations, generating the target feature representation includes: determining a first part of information in the time correlation information associated with the second time-frequency representation; determining a second part of information in the frequency correlation information associated with the second time-frequency representation; and determining the target feature representation based on the first part of information, the second part of information, and the target time-frequency representation.

[0111] In some implementations, the target time-frequency representation is associated with multiple frequency bands in the frequency domain and multiple time periods in the time domain. Determining the frequency correlation information and the time correlation information includes: obtaining a frequency-domain feature representation and a time-domain feature representation by processing the time-frequency representation, the frequency-domain feature representation including multiple features in multiple frequency bands within one of the multiple time periods, and the time-domain feature representation including multiple features in multiple time periods on one of the multiple frequency bands; determining the frequency correlation information based on the frequency-domain feature representation and frequency-domain weighting information, the frequency-domain weighting information indicating the degree of correlation between reference audio signals in the multiple frequency bands; and determining the time correlation information based on the time-domain feature representation and time-domain weighting information, the time-domain weighting information indicating the degree of correlation between reference audio signals in the multiple time periods.

[0112] In some implementations, the frequency-domain weighting information and the time-domain weighting information are determined based on the reference audio signal including a reference speech signal and a reference noise signal.

[0113] In some implementations, generating the output audio signal includes: determining a masked time-frequency representation by applying a mask generated based on the target feature representation to at least a portion of the target time-frequency representation; and converting the masked time-frequency representation into the output audio signal.

[0114] In some implementations, converting the masked time-frequency representation into the output audio signal includes: converting the masked time-frequency representation into the output audio signal by applying the masked time-frequency representation to a trained decoder, where the trained decoder is configured to convert the audio signal intensity varying over time at different frequencies into a waveform varying over time.

[0115] In another aspect, the present disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, which, when executed by the processing unit, cause the device to perform actions, the actions including: obtaining a target time-frequency representation that at least indicates the intensity of an input audio signal varying over time at different frequencies, the input audio signal including a speech component and a noise component; determining frequency correlation information and time correlation information of the input audio signal based on the target time-frequency representation; generating a target feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the target time-frequency representation; and generating an output audio signal based on the target feature representation and the target time-frequency representation, in which the speech component is enhanced relative to the noise component.

[0116] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation and the updated frequency correlation information.

[0117] In some implementations, generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; updating the frequency correlation information and the time correlation information based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation, the updated frequency correlation information, and the updated time correlation information.

[0118] In some implementations, obtaining the target time-frequency representation includes: obtaining a first time-frequency representation of a processed audio signal before the input audio signal occurs, the first time-frequency representation indicating the intensity of the processed audio signal varying with time at the different frequencies; determining a second time-frequency representation of the input audio signal, the second time-frequency representation indicating the intensity of the input audio signal varying with time at the different frequencies; and combining the first time-frequency representation and the second time-frequency representation into the target time-frequency representation.

[0119] In some implementations, generating the target feature representation includes: determining a first portion of information in the time correlation information associated with the second time-frequency representation; determining a second portion of information in the frequency correlation information associated with the second time-frequency representation; and determining the target feature representation based on the first portion of information, the second portion of information, and the target time-frequency representation.

[0120] In some implementations, the target time-frequency representation is associated with a plurality of frequency bands in the frequency domain and a plurality of time periods in the time domain. Determining the frequency correlation information and the time correlation information includes: obtaining a frequency-domain feature representation and a time-domain feature representation by processing the time-frequency representation, the frequency-domain feature representation including a plurality of features in a plurality of the frequency bands during one of the plurality of time periods, the time-domain feature representation including a plurality of features on one of the plurality of frequency bands during the plurality of time periods; determining the frequency correlation information based on the frequency-domain feature representation and frequency-domain weighting information, the frequency-domain weighting information indicating the degree of correlation of a reference audio signal among the plurality of frequency bands; and determining the time correlation information based on the time-domain feature representation and time-domain weighting information, the time-domain weighting information indicating the degree of correlation of the reference audio signal among the plurality of time periods.

[0121] In some implementations, the frequency-domain weighting information and the time-domain weighting information are determined based on the reference audio signal including a reference speech signal and a reference noise signal.

[0122] In some implementations, generating the output audio signal includes: determining a masked time-frequency representation by applying a mask generated based on the target feature representation to at least a portion of the target time-frequency representation; and converting the masked time-frequency representation into the output audio signal.

[0123] In some implementations, converting the masked time-frequency representation into the output audio signal includes: converting the masked time-frequency representation into the output audio signal by applying the masked time-frequency representation to a trained decoder, where the trained decoder is configured to convert the intensity of the audio signal varying with time at the different frequencies into a waveform varying with time.

[0124] In another aspect, the present disclosure provides a computer program product tangibly stored in a non-transitory computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform the methods of the above aspects.

[0125] In another aspect, the present disclosure provides a computer-readable medium having stored thereon machine-executable instructions that, when executed by a device, cause the device to perform the methods of the above aspects.

[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0127] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0128] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0129] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation may also be implemented separately or in any suitable subcombination in multiple implementations.

[0130] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A computer-implemented method, comprising: obtaining a target time-frequency representation that at least indicates the intensity of an input audio signal varying over time at different frequencies, the input audio signal including a speech component and a noise component; determining frequency correlation information and time correlation information of the input audio signal based on the target time-frequency representation; generating a target feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the target time-frequency representation; and generating an output audio signal based on the target feature representation and the target time-frequency representation, in which the speech component is enhanced relative to the noise component, wherein the target time-frequency representation is associated with a plurality of frequency bands in the frequency domain and a plurality of time periods in the time domain, and determining the frequency correlation information and the time correlation information includes: obtaining a frequency-domain feature representation and a time-domain feature representation by processing the time-frequency representation, the frequency-domain feature representation including a plurality of features in a plurality of the frequency bands within one of the plurality of time periods, and the time-domain feature representation including a plurality of features in a plurality of the time periods on one of the plurality of frequency bands; determining the frequency correlation information based on the frequency-domain feature representation and frequency-domain weighting information, the frequency-domain weighting information indicating the degree of correlation between reference audio signals in the plurality of frequency bands; and determining the time correlation information based on the time-domain feature representation and time-domain weighting information, the time-domain weighting information indicating the degree of correlation between reference audio signals in the plurality of time periods, wherein the frequency-domain weighting information and the time-domain weighting information are based on an attention mechanism.

2. The method according to claim 1, wherein generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation, wherein the intermediate feature representation is an intermediate calculation result generated by one of a plurality of correlation units; updating the frequency correlation information by applying an attention mechanism based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation and the updated frequency correlation information.

3. The method according to claim 1, wherein generating the target feature representation includes: generating an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation, wherein the intermediate feature representation is an intermediate calculation result generated by one of a plurality of correlation units; updating the frequency correlation information and the time correlation information by applying an attention mechanism based on the intermediate feature representation; and determining the target feature representation based on the intermediate feature representation, the updated frequency correlation information, and the updated time correlation information.

4. The method according to claim 1, wherein obtaining the target time-frequency representation includes: Obtain a first time-frequency representation of a processed audio signal before the input audio signal occurs, the first time-frequency representation indicating the intensity of the processed audio signal varying with time at the different frequencies; Determine a second time-frequency representation of the input audio signal, the second time-frequency representation indicating the intensity of the input audio signal varying with time at the different frequencies; and Combine the first time-frequency representation and the second time-frequency representation into the target time-frequency representation.

5. The method according to claim 4, wherein generating the target feature representation includes: Determine a first part of information in the time correlation information associated with the second time-frequency representation; Determine a second part of information in the frequency correlation information associated with the second time-frequency representation; And Based on the first part of information, the second part of information and the target time-frequency representation, determine the target feature representation.

6. The method according to claim 1, wherein the frequency-domain weighting information and the time-domain weighting information are determined based on the reference audio signal including a reference speech signal and a reference noise signal.

7. The method according to claim 1, wherein generating the output audio signal includes: Determine a masked time-frequency representation by applying a mask generated based on the target feature representation to at least a part of the target time-frequency representation; And Convert the masked time-frequency representation into the output audio signal.

8. The method according to claim 7, wherein converting the masked time-frequency representation into the output audio signal includes: Convert the masked time-frequency representation into the output audio signal by applying the masked time-frequency representation to a trained decoder, wherein the trained decoder is configured to convert the intensity of the audio signal varying with time at the different frequencies into a waveform varying with time.

9. An electronic device, comprising: A processing unit; And A memory, coupled to the processing unit and containing instructions stored thereon, the instructions when executed by the processing unit cause the device to perform actions, the actions including: Obtain a target time-frequency representation at least indicating the intensity of an input audio signal varying with time at different frequencies, the input audio signal including a speech component and a noise component; Based on the target time-frequency representation, determine the frequency correlation information and the time correlation information of the input audio signal; Based on the frequency correlation information, the time correlation information and the target time-frequency representation, generate a target feature representation for distinguishing the speech component and the noise component; and Based on the target feature representation and the target time-frequency representation, generate an output audio signal in which the speech component is enhanced relative to the noise component, Wherein the target time-frequency representation is associated with a plurality of frequency bands in the frequency domain and a plurality of time periods in the time domain, and determining the frequency correlation information and the time correlation information includes: By processing the time-frequency representation, a frequency-domain feature representation and a time-domain feature representation are obtained. The frequency-domain feature representation includes a plurality of features in a plurality of frequency bands within one of the plurality of time periods, and the time-domain feature representation includes a plurality of features in a plurality of time periods on one of the plurality of frequency bands; Based on the frequency-domain feature representation and frequency-domain weighting information, the frequency correlation information is determined, and the frequency-domain weighting information indicates the degree of correlation of the reference audio signal among the plurality of frequency bands; and Based on the time-domain feature representation and time-domain weighting information, the time correlation information is determined, and the time-domain weighting information indicates the degree of correlation of the reference audio signal among the plurality of time periods, wherein the frequency-domain weighting information and the time-domain weighting information are based on an attention mechanism.

10. The apparatus according to claim 9, wherein generating the target feature representation includes: Based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation, an intermediate feature representation for distinguishing the speech component and the noise component is generated, where the intermediate feature representation is an intermediate calculation result generated by one of a plurality of correlation units; Based on the intermediate feature representation, the frequency correlation information is updated by applying an attention mechanism; and Based on the intermediate feature representation and the updated frequency correlation information, the target feature representation is determined.

11. The apparatus according to claim 9, wherein generating the target feature representation includes: Based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation, an intermediate feature representation for distinguishing the speech component and the noise component is generated, where the intermediate feature representation is an intermediate calculation result generated by one of a plurality of correlation units; Based on the intermediate feature representation, the frequency correlation information and the time correlation information are updated by applying an attention mechanism; and Based on the intermediate feature representation, the updated frequency correlation information, and the updated time correlation information, the target feature representation is determined.

12. The apparatus according to claim 12, wherein obtaining the target time-frequency representation includes: Obtaining a first time-frequency representation of a processed audio signal before the input audio signal occurs, where the first time-frequency representation indicates the intensity of the processed audio signal changing with time at different frequencies; Determining a second time-frequency representation of the input audio signal, where the second time-frequency representation indicates the intensity of the input audio signal changing with time at different frequencies; and Combining the first time-frequency representation and the second time-frequency representation into the target time-frequency representation.

13. The apparatus according to claim 12, wherein generating the target feature representation includes: Determining a first part of information in the time correlation information associated with the second time-frequency representation; Determining a second part of information in the frequency correlation information associated with the second time-frequency representation; and Determine the target feature representation based on the first part of information, the second part of information, and the target time-frequency representation.

14. The apparatus according to claim 9, wherein the frequency-domain weighting information and the time-domain weighting information are determined based on the reference audio signal including a reference speech signal and a reference noise signal.

15. The apparatus according to claim 9, wherein generating the output audio signal comprises: Determining a masked time-frequency representation by applying a mask generated based on the target feature representation to at least a part of the target time-frequency representation; and Converting the masked time-frequency representation into the output audio signal.

16. The apparatus according to claim 15, wherein converting the masked time-frequency representation into the output audio signal comprises: Converting the masked time-frequency representation into the output audio signal by applying the masked time-frequency representation to a trained decoder, wherein the trained decoder is configured to convert audio signal intensities varying with time at different frequencies into waveforms varying with time.

17. A computer program product tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions that, when executed by a device, cause the device to perform operations, the operations comprising: Obtaining a target time-frequency representation that at least indicates intensities of an input audio signal varying with time at different frequencies, the input audio signal including a speech component and a noise component; Determining frequency correlation information and time correlation information of the input audio signal based on the target time-frequency representation; Generating a target feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the target time-frequency representation; and Generating an output audio signal based on the target feature representation and the target time-frequency representation, in which the speech component is enhanced relative to the noise component, wherein the target time-frequency representation is associated with a plurality of frequency bands in the frequency domain and a plurality of time periods in the time domain, and determining the frequency correlation information and the time correlation information comprises: Obtaining a frequency-domain feature representation and a time-domain feature representation by processing the time-frequency representation, the frequency-domain feature representation including a plurality of features in a plurality of the frequency bands during one of the plurality of time periods, and the time-domain feature representation including a plurality of features in a plurality of the time periods on one of the plurality of frequency bands; Determining the frequency correlation information based on the frequency-domain feature representation and frequency-domain weighting information, the frequency-domain weighting information indicating a degree of correlation between the plurality of frequency bands of the reference audio signal; and Determining the time correlation information based on the time-domain feature representation and time-domain weighting information, the time-domain weighting information indicating a degree of correlation between the plurality of time periods of the reference audio signal, wherein the frequency-domain weighting information and the time-domain weighting information are based on an attention mechanism.

18. The computer program product according to claim 17, wherein generating the target feature representation comprises: Generate an intermediate feature representation for distinguishing the speech component and the noise component based on the frequency correlation information, the time correlation information, and the convolved target time-frequency representation; Update the frequency correlation information based on the intermediate feature representation; and Determine the target feature representation based on the intermediate feature representation and the updated frequency correlation information.