Audio signal processing method, device, equipment and storage medium
The audio signal is processed through a complex dual-path encoder and a decoder, and combined with complex ratio masking and spectrum mapping, the poor speech recognition effect caused by noise interference in the prior art is solved, and the speech quality and recognition effect are improved.
Patent Information
- Application Number
- CN202211702084.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing voice enhancement technologies cannot effectively reduce noise in audio signals, resulting in poor voice recognition.
The audio signal is encoded by a complex dual-path encoder. Through the combination of complex ratio masking and complex spectrum mapping, a target audio signal after speech enhancement is generated. The time series and frequency sequence are processed by a complex dual-path masking decoder and spectrum decoder respectively to achieve simultaneous enhancement of the amplitude spectrum and phase.
It significantly improves the voice quality and speech recognition effect, solves the problem of noise interference, and improves the effect of voice enhancement.
Smart Images

Figure CN116229999B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular to an audio signal processing method, apparatus, device, and storage medium. Background Art
[0002] Current speech enhancement technologies include noise reduction, loudness enhancement, echo cancellation, and dereverberation.
[0003] However, after performing speech enhancement on audio signals using current speech enhancement technology, good noise reduction results cannot be obtained. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an audio signal processing method, apparatus, device and storage medium to solve the noise interference problem and improve the speech recognition effect.
[0005] In a first aspect, an embodiment of the present disclosure provides an audio signal processing method, including:
[0006] Get the original audio signal;
[0007] Encoding a first complex spectrogram corresponding to the original audio signal to obtain a first feature map;
[0008] Processing the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map;
[0009] predicting a complex ratio mask and a complex spectral map based on the second feature map;
[0010] A second complex spectrogram is generated according to the complex ratio mask and the complex spectrum mapping, and a speech-enhanced target audio signal is generated according to the second complex spectrogram.
[0011] In a second aspect, an embodiment of the present disclosure provides an audio signal processing device, including:
[0012] An acquisition module, used to obtain the original audio signal;
[0013] an encoding module, configured to encode a first complex spectrum graph corresponding to the original audio signal to obtain a first feature graph;
[0014] a processing module, configured to process the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map;
[0015] a prediction module, configured to predict a complex ratio mask and a complex spectrum map based on the second feature map;
[0016] A generating module is configured to generate a second complex spectrogram according to the complex ratio masking and the complex spectrum mapping, and generate a speech-enhanced target audio signal according to the second complex spectrogram.
[0017] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0018] Memory;
[0019] processor; and
[0020] computer programs;
[0021] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0022] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described in the first aspect.
[0023] The audio signal processing method, apparatus, device, and storage medium provided by the embodiments of the present disclosure encode the first complex spectrogram corresponding to the original audio signal to obtain a first feature map, and separately process the time series and frequency series corresponding to the first feature map to obtain a second feature map. Based on the second feature map, complex ratio masking and complex spectrum mapping are simultaneously learned. Thus, by combining masking prediction and spectrum prediction, the amplitude spectrum and phase of the original audio signal are simultaneously enhanced, thereby improving the speech enhancement effect. Speech enhancement can greatly improve speech quality, solve noise interference problems, and enhance speech recognition effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0025] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0026] Figure 1 A flow chart of the audio signal processing method provided in an embodiment of the present disclosure;
[0027] Figure 2 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0028] Figure 3This is an architectural diagram of D2Former provided in an embodiment of the present disclosure;
[0029] Figure 4 A schematic diagram of the structure of a complex dual-path encoder provided in an embodiment of the present disclosure;
[0030] Figure 5 A schematic structural diagram of a complex dual-path converter block provided by another embodiment of the present disclosure;
[0031] Figure 6 A schematic structural diagram of a complex dual-path masked decoder provided by another embodiment of the present disclosure;
[0032] Figure 7 A schematic structural diagram of a complex dual-path spectrum decoder provided by another embodiment of the present disclosure;
[0033] Figure 8 A schematic structural diagram of an audio signal processing device provided in another embodiment of the present disclosure;
[0034] Figure 9 A schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. DETAILED DESCRIPTION
[0035] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0036] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0037] It should be noted that the audio signals involved in this application (including but not limited to audio signals collected by the terminal, audio signals pre-stored in the terminal or server, etc.) are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0038] Typically, speech enhancement technologies include noise reduction, loudness enhancement, echo cancellation, and dereverberation. However, current speech enhancement technologies fail to achieve effective noise reduction results when applied to audio signals. To address this issue, the present disclosure provides an audio signal processing method, which is described below with reference to specific embodiments.
[0039] Figure 1 A flow chart of the audio signal processing method provided in an embodiment of the present disclosure. The method can be executed by an audio signal processing device, which can be implemented in software and / or hardware, and the device can be configured in an electronic device, such as a server or a terminal, wherein the terminal specifically includes a mobile phone, a computer or a tablet computer. In addition, the audio signal processing method described in this embodiment can be applied to application scenarios such as front-end voice processing of audio and video conferences, noise processing of long-distance sound pickup, and speech recognition in conference scenes. For example, in an audio and video conference, the microphone is located in the middle of the conference room. Since the speaker may be far away from the microphone, the speaker's voice is small and the noise is large in the audio signal collected by the microphone. In this case, the method described in this embodiment can be used to perform noise reduction processing on the audio signal, thereby achieving voice enhancement. For another example, in a long-distance sound pickup scenario, the recorder of the audio and video may be outdoors, and the person being recorded is far away from the recording device (such as a camera), resulting in large noise in the audio signal recorded by the recording device. The method described in this embodiment can be used to perform noise reduction processing on the audio signal. For another example, in a conference scenario, after the audio signal is subjected to noise reduction processing, speech recognition can be further performed on the noise-reduced audio signal, thereby improving the accuracy of speech recognition. In addition, the method described in this embodiment can be performed by Figure 2 The terminal 21 or server 22 shown is executed. If executed by the terminal 21, the terminal 21 can use this method to process the original audio signal it collects. If executed by the server 22, the terminal 21 can send the original audio signal it collects to the server 22, and the server 22 uses this method to process the original audio signal. In some embodiments, the server 22 can not only process the original audio signal collected by the terminal 21, but also process the original audio signal pre-stored in the server 22, and the original audio signal from other terminals or other servers. That is to say, when the method is executed by the server 22, the source of the original audio signal to be processed is not limited. This embodiment aims to extract the target speech signal, that is, the useful component in the original audio signal, such as the user's voice, from the original audio signal corrupted by noise, thereby achieving monophonic speech enhancement. This method can be widely used in voice communications to obtain better perceptual quality and intelligibility, as well as in automatic speech recognition (ASR) systems with more robust speech recognition performance. As Figure 1 As shown, the specific steps of this method are as follows:
[0040] S101: Acquire an original audio signal.
[0041] For example, the server 22 may receive the original audio signal from the terminal 21, another terminal, or another server. Alternatively, the server 22 may receive audio and video information, such as a real-time audio and video stream, from the terminal 21, another terminal, or another server. Furthermore, the server 22 extracts the original audio signal from the audio and video information, such as extracting the audio stream from the real-time audio and video stream.
[0042] S102: Encode a first complex spectrum graph corresponding to the original audio signal to obtain a first feature graph.
[0043] For example, the original audio signal is a noisy speech signal, and the waveform y of the speech signal can be expressed as y=s+n, where s represents a useful signal, such as a noise-free speech signal, and n represents noise. In this embodiment, a completely complex dual-path dual-decoder conformer network (D2Former) can be deployed in the server 22. The architecture diagram of D2Former is shown in FIG. Figure 3 As shown, specifically, D2Former includes a complex dual-path encoder (Complex Dual-Path Encoder), a complex dual-path converter block (Complex Dual-Path Conformer Block) repeated N times, and two independent complex dual-path decoders, the two independent complex dual-path decoders are a complex dual-path masking decoder (Complex Dual-Path Masking Decoder) and a complex dual-path spectral decoder (Complex Dual-Path Spectral Decoder). Among them, the complex dual-path masking decoder is used to learn complex ratio masking, and the complex dual-path spectral decoder is used to learn complex spectrum mapping. Specifically, the original audio signal can be used as the input of D2Former, and the input waveform y is first converted into a complex spectrogram. For example, the input waveform y is subjected to a short-time Fourier transform (STFT) to obtain a complex spectrogram, which is recorded as a first complex spectrogram. The complex spectrogram is expressed as Y = STFT(y)∈C T×F , where C represents the complex domain, T represents the time dimension of Y, and F represents the frequency dimension of Y. For example, Y includes T time series and F frequency series, where the time series and frequency series are both complex sequences. Furthermore, the first complex spectrogram is encoded to obtain a first feature map.
[0044] like Figure 3 As shown, the first complex spectrum graph Y is input into the complex dual-path encoder to extract a high-level complex feature sequence, which can be recorded as a first feature graph.
[0045] S103: Process the time series and frequency series corresponding to the first feature map respectively to obtain a second feature map.
[0046] For example, a first feature map is input into a complex dual-path converter block that is repeated N times for processing, so that the complex dual-path converter block processes the time series and frequency series corresponding to the first feature map separately to obtain a second feature map. Specifically, the complex dual-path converter block includes a time path and a frequency path. In the complex dual-path converter block, subband information in the time path and full-band information in the frequency path are alternately modeled and processed to obtain the second feature map.
[0047] S104: Predict a complex ratio mask and a complex spectrum map according to the second feature map.
[0048] Specifically, the second feature map can be respectively as follows Figure 3 The input of the complex dual path mask decoder and the input of the complex dual path spectrum decoder are shown. Specifically, the complex dual path mask decoder can predict the complex ratio mask according to the second feature map, and the complex dual path spectrum decoder can predict the complex spectrum mapping according to the second feature map.
[0049] Optionally, based on the second feature map, predicting a complex ratio mask and a complex spectrum map includes: using the second feature map as the input of a mask decoder, and outputting the complex ratio mask through the mask decoder; using the second feature map as the input of a spectrum decoder, and outputting the complex spectrum map through the spectrum decoder.
[0050] For example, the second feature map is used as an input of a complex dual-path mask decoder, so that the complex dual-path mask decoder outputs a complex ratio mask. The second feature map is used as an input of a complex dual-path spectrum decoder, so that the complex dual-path spectrum decoder outputs a complex spectrum map.
[0051] S105 : Generate a second complex spectrogram according to the complex ratio masking and the complex spectrum mapping, and generate a speech-enhanced target audio signal according to the second complex spectrogram.
[0052] For example, after the complex ratio mask and the complex spectrum map are decoded, they are weighted and summed to generate a second complex spectrogram. After the second complex spectrogram undergoes an inverse short-time Fourier transform (ISTFT), the resulting audio signal is the target audio signal after speech enhancement. This target audio signal after speech enhancement can be used as the output of D2Former. This target audio signal after speech enhancement can also be called the denoised target audio signal.
[0053] Optionally, generating a second complex spectrum diagram based on the complex ratio mask and the complex spectrum mapping includes: generating an enhanced spectrum diagram based on the first complex spectrum diagram and the complex ratio mask; generating the second complex spectrum diagram based on the enhanced spectrum diagram and the complex spectrum mapping.
[0054] For example, the complex ratio mask is denoted as M, M = M R +jM I , M R represents the real part of M, M I represents the imaginary part of M. Further, the complex ratio masking is applied to the first complex spectrum Y to generate an enhanced spectrum S′, S′=S R ′+S I ′=M⊙Y, where ⊙ represents element-by-element complex multiplication, S R ' represents the real part of S', S I ' represents the imaginary part of S'. In addition, the complex spectrum mapping is recorded as S", S" = S R ″+jS I ″.S R ″ represents the real part of S″, S I "" represents the imaginary part of S". Further, a weighted sum is performed on S' and S", generating a second complex spectrum graph, which is denoted as S. S can be expressed as the following formula (1):
[0055] S=αS′+βS″=(αS R +βS R ″)+j(αS I ′+βS I ″) (1)
[0056] Among them, α and β represent weight factors respectively.
[0057] After the second complex spectrogram S is subjected to inverse short-time Fourier transform (ISTFT), the waveform of the target audio signal after speech enhancement can be recorded as Therefore, if Figure 3 The D2Former shown can process the noisy speech waveform y=s+n into an enhanced speech waveform
[0058] The disclosed embodiment encodes the first complex spectrogram corresponding to the original audio signal to obtain a first feature graph, and then processes the time series and frequency series corresponding to the first feature graph to obtain a second feature graph. Based on the second feature graph, complex ratio masking and complex spectrum mapping are simultaneously learned. Thus, by combining masking prediction and spectrum prediction, the amplitude spectrum and phase of the original audio signal are simultaneously enhanced, thereby improving the speech enhancement effect. Speech enhancement can greatly improve speech quality, resolve noise interference problems, and enhance speech recognition.
[0059] Below Figure 3 The complex dual-path encoder, complex dual-path converter block, complex dual-path mask decoder, and complex dual-path spectrum decoder shown are introduced in detail respectively.
[0060] Optionally, encoding the first complex spectrum graph corresponding to the original audio signal includes: using a complex dual-path encoder to encode the first complex spectrum graph corresponding to the original audio signal, the complex dual-path encoder including a first convolution module, a second convolution module and a first dual-path hole module.
[0061] like Figure 4 As shown, the complex dual-path encoder includes two complex two-dimensional convolution modules (ComplexConv2DModule) and a complex dual-path dilated module (Complex Dual-Path Dilated Module). Complex two-dimensional convolution module 41 is denoted as the first convolution module, and complex two-dimensional convolution module 42 is denoted as the second convolution module. Complex dual-path dilated module 43 is denoted as the first dual-path dilated module. Specifically, the complex dual-path encoder is used to encode the first complex spectrogram Y.
[0062] Specifically, the internal structure of the complex two-dimensional convolution module 41 is the same as the internal structure of the complex two-dimensional convolution module 42. Taking the complex two-dimensional convolution module 41 as an example, the complex two-dimensional convolution module 41 includes a complex two-dimensional convolution layer (Complex 2D Convolution), a complex instance normalization (Complex InstanceNorm), and a complex activation function (ParametricRectified Linear Unit, PReLU). Among them, the complex two-dimensional convolution layer is a complex layer with trainable parameters. Assume that the complex layer with trainable parameters is denoted as H, and the input of the complex layer is a complex sequence, and the complex sequence is denoted as Z=Z R +Z I , Z R is the real part of Z, Z Iis the imaginary part of Z. Specifically, Z can be Y as described above, or Z is the result of Y being processed by a module. The output of the complex layer H with trainable parameters is denoted as H(Z). H(Z) can be expressed as the following formula (2):
[0063] H(Z)=[H R (Z R )-H I (Z I )]+j[H R (Z I )+H I (Z R )] (2)
[0064] Among them, H R is a real layer in the complex layer H with trainable parameters, H I It is an imaginary layer in the complex layer H with trainable parameters. The real layer and the imaginary layer can respectively operate or calculate on real values, so that the operation or calculation results constitute the complex number shown in formula (2).
[0065] In addition, complex instance regularization is a complex layer without trainable parameters. The complex operations that can be implemented by complex instance regularization are normalization, regularization, or standardization. Assuming that the function that implements normalization, regularization, or standardization is recorded as N, complex instance regularization can normalize, regularize, or standardize the real part and imaginary part of H(Z) as described above, respectively. For example, the output of complex instance regularization is recorded as N(H(Z)), and N(H(Z)) can be expressed as the following formula (3):
[0066] N(H(Z))=N(H(Z) R )+jN(H(Z) I ) (3)
[0067] Among them, H(Z) R Indicates [H R (Z R )- I (Z I )],H(Z) I express[ R (Z I )+ I ( R )].
[0068] Optionally, the first convolution module is used to expand the first complex spectrum graph from one feature channel to multiple feature channels; the second convolution module is used to reduce the frequency dimension corresponding to the first complex spectrum graph; and the first dual-path hole module is used to perform feature extraction on the first complex spectrum graph in the time dimension and frequency dimension corresponding to the first complex spectrum graph respectively.
[0069] In some embodiments, the first complex spectrum graph Y can be recorded as the initial feature graph, and the size of the feature graph can be expressed as follows: Figure 4 As shown, B×1×T×F×2, where B represents B parallel audio signals. 1 represents the number of feature channels. 2 represents the real and imaginary parts of the complex sequence. The complex two-dimensional convolution module 41 can expand the first complex spectrogram Y from one feature channel to C feature channels. The complex two-dimensional convolution module 42 can halve the frequency dimension from F to F / 2 to achieve reasonable complexity. The complex dual-path hole module 43 is used to extract features from the first complex spectrogram Y in both the time and frequency dimensions.
[0070] Optionally, the first dual-path hole module includes a plurality of dual-path blocks connected sequentially, and in the connection direction, the hole factor of a subsequent dual-path block is twice the hole factor of a previous dual-path block.
[0071] like Figure 4 As shown, the complex dual-path hole module 43 has a depth of four levels. Specifically, the complex dual-path hole module 43 includes four dual-path blocks (DPBlocks), each DPBlock including a complex constant pad (ComplexConstantPad), a complex two-dimensional convolution module (Complex Conv2D Module), and a complex feedforward sequential memory network (CFSMN) module. Because the complex two-dimensional convolution module in each DPBlock reduces the time dimension, it is necessary to pre-increase the time dimension by padding with zeros using the complex constant pad. This allows the time dimension of the complex two-dimensional convolution module output to be maintained at the value T. In addition, the complex two-dimensional convolution module in each DPBlock is responsible for feature extraction of the first complex spectrum graph in the time dimension.
[0072] Specifically, such as Figure 4The four dual-path blocks shown can be connected sequentially, and the hole factors d of the four dual-path blocks are 1, 2, 4, and 8, respectively, that is, the hole factor of the subsequent dual-path block is twice the hole factor of the previous dual-path block. In addition, the hole factor d of the complex two-dimensional convolution module in each dual-path block is consistent with the hole factor d of the dual-path block. For example, if the hole factor d of the first dual-path block is 1, then the hole factor d of the complex two-dimensional convolution module in the first dual-path block is also 1. Similarly, if the hole factor d of the fourth dual-path block is 8, then the hole factor d of the complex two-dimensional convolution module in the fourth dual-path block is also 8. In addition, in some other embodiments, the complex two-dimensional convolution module in each dual-path block can be recorded as a hole convolution, wherein the complex two-dimensional convolution module with d greater than 1 is recorded as an expanded convolution. When d is greater than 1, the receptive field of the complex two-dimensional convolution module can be increased. For example, when d=1, the complex two-dimensional convolution module can extract features from 10 time series. When d = 2, the complex two-dimensional convolution module can extract features from 20 time series, and so on. It is understood that the internal structure of the complex two-dimensional convolution module in each dual-path block is similar to the internal structure of the complex two-dimensional convolution module 41 or the complex two-dimensional convolution module 42, and will not be repeated here. The CFSMN module includes a CFSMN layer to learn long-range frequency correlation. Specifically, the CFSMN module in each DPBlock is responsible for extracting features from the first complex spectrogram in the frequency dimension.
[0073] Optionally, the input of any other dual path block among the multiple dual path blocks except the first dual path block is a concatenation result of the input of the first dual path block and the output of the previous dual path block of any dual path block.
[0074] like Figure 4As shown, in the four dual-path blocks, the outputs of the first three dual-path blocks and their inputs are connected by short-circuiting. Specifically, starting from the second dual-path block, the input of any dual-path block is the splicing result of the output of the previous dual-path block and the input of the first dual-path block. For example, the splicing result obtained by splicing the output of the first dual-path block and the input of the first dual-path block is the input of the second dual-path block. The splicing result obtained by splicing the output of the second dual-path block and the input of the first dual-path block is the input of the third dual-path block. The splicing result obtained by splicing the output of the third dual-path block and the input of the first dual-path block is the input of the fourth dual-path block. In addition, it can be understood that since the splicing process will increase the number of feature channels, starting from the second dual-path block, the number of feature channels in the input information of any dual-path block will be greater than C, resulting in the number of feature channels in the input information of the complex two-dimensional convolution module in any dual-path block being greater than C. However, by configuring the complex two-dimensional convolution module, the number of feature channels in the information output by the complex two-dimensional convolution module can be restored to the C value. Therefore, the number of feature channels in the input information of the complex dual-path hole module 43 and the number of feature channels in its output information are respectively maintained at the C value. Figure 4 As shown, the size of the output information of the complex dual-path hole module 43 is B×C×T×F×2. The feature map generated after the output information of the complex dual-path hole module 43 is processed by the complex two-dimensional convolution module 42 is a three-dimensional time-frequency map with a size of B×C×T×F / 2×2. The three-dimensional time-frequency map is the first feature map obtained after the complex dual-path encoder encodes the first complex spectrum map.
[0075] In this embodiment, the dual-path design of the complex dual-path encoder not only improves the feature representation in the frequency dimension, but also increases the receptive field of the complex two-dimensional convolution module through the dilated convolution.
[0076] Considering that the feature map generated by the complex dual-path encoder is a three-dimensional time-frequency map with a size of B×C×T×F / 2×2, this embodiment applies a dual-path self-attention mechanism on the three-dimensional time-frequency map to independently and alternately process the time series and frequency series respectively.
[0077] Optionally, the time series and frequency series corresponding to the first feature map are processed separately, including: adjusting the first feature map to a time series; processing the time series using a self-attention mechanism through a time converter to obtain the output of the time converter; obtaining a frequency sequence based on the output of the time converter; processing the frequency sequence using a self-attention mechanism through a frequency converter to obtain the output of the frequency converter.
[0078] like Figure 5As shown, the complex dual-path converter block includes a complex time converter (Complex Time Conformer) and a complex frequency converter (Complex Frequency Conformer). Both the complex time converter and the complex frequency converter are based on the original converter and extended to the complex domain.
[0079] For example, the three-dimensional time-frequency map with the size of B×C×T×F / 2×2 output by the complex dual-path encoder is adjusted to a time series with the size of Figure 5 As shown in BF / 2×T×C×2, this time series can be used as the input of the complex time converter, and the complex time converter can perform the self-attention mechanism along the time series. In some embodiments, the output of the complex time converter can be adjusted to a frequency sequence with a size of Figure 5 As shown in FIG. BT×F / 2×C×2, this frequency sequence can be used as the input of a complex frequency converter, which can perform a self-attention mechanism along the frequency sequence. In other embodiments, a time series of size BF / 2×T×C×2 can be resized to a frequency series of size BT×F / 2×C×2.
[0080] Optionally, obtaining a frequency sequence based on the output of the time converter includes: connecting the output of the time converter and the input of the time converter to obtain a first connection result; adjusting the first connection result to obtain the frequency sequence; accordingly, processing the frequency sequence through the frequency converter using a self-attention mechanism to obtain the output of the frequency converter, the method also includes: connecting the output of the frequency converter and the input of the frequency converter to obtain a second connection result; adjusting the second connection result to obtain an updated time series.
[0081] For example, in some other embodiments, Figure 5The input and output of the complex time converter are connected to obtain a first connection result, which is then reshaped to obtain a frequency sequence of size BT×F / 2×C×2. After the complex frequency converter processes the frequency sequence, the input and output of the complex frequency converter can be connected to obtain a second connection result, which is then reshaped to obtain an updated time sequence of size BF / 2×T×C×2. The updated time sequence can serve as a new input to the complex time converter, so that the complex time converter processes the updated time sequence to obtain a new output of the complex time converter. Furthermore, the new input and output of the complex time converter are connected to obtain a new first connection result, which is then reshaped to obtain a new frequency sequence. The new frequency sequence can be used as a new input of the complex frequency converter. After the complex frequency converter processes the new frequency sequence, a new output of the complex frequency converter is obtained. Further, the new input of the complex frequency converter and the new output of the complex frequency converter are connected to obtain a new second connection result. After adjusting the new second connection result, a new time sequence is obtained. The new time sequence is input into the complex time converter again. Figure 5 The process shown continues with subsequent processing, and so on, and is repeated N times, so that the time series and the frequency series are processed N times respectively. In each processing process, the time series and the frequency series are processed alternately, and the time series and the frequency series are updated in each processing process.
[0082] like Figure 5 As shown, the internal structure of the complex time converter is the same as the internal structure of the complex frequency converter. For example, taking the complex time converter as an example, the complex time converter includes a complex feed forward module (Complex Feed Forward Module), a complex multi-head self attention module (Complex Multi-Head Self Attention Module), a complex convolution module (Complex Convolution Module), and a complex LayerNorm function. Among them, the complex operations in the complex feed forward module and the complex operations in the complex convolution module can refer to the formula (2) mentioned above. The complex multi-head self attention module in the complex time converter can be recorded as time series attention, and the complex multi-head self attention module in the complex frequency converter can be recorded as frequency series attention. The time series attention and the frequency series attention use the same mechanism. For example, taking the time series attention as an example, the input of the time series attention is a complex sequence, and the complex sequence is recorded as Z=Z R +ZI , Z R is the real part of Z, Z I is the imaginary part of Z. The complex parameters in the complex multi-head self-attention module include Q 、W K 、W V , in, It's W Q The real part, W I Q It's W Q The imaginary part of . It's W K The real part, W I K It's W K The imaginary part of . It's W V The real part, W I V It's W V The imaginary part of . The queries in the multi-head self-attention module are denoted as Q, the keys as K, and the values as V. Q, K, and V are calculated using the following formula (4):
[0083] Q=ZW Q ,K=ZW K ,V=ZW V (4)
[0084] Furthermore, the complex self-attention ComplexAttention(Q,K,V) can be calculated by the following formula (5):
[0085]
[0086] Among them, |·| represents the absolute value, d k Represents the dimension of the feature channel, the complex matrix product QK T It can be calculated by the following formula (6):
[0087]
[0088] Among them, Q R is the real part of Q, Q I is the imaginary part of Q. K R is the real part of K, K I is the imaginary part of K. K T is the transpose of K.
[0089] In this embodiment, the multi-head attention used by the complex multi-head self-attention module can specifically be four attention heads. In addition, by connecting the input of the complex time converter to the output of the complex time converter, and connecting the input of the complex frequency converter to the output of the complex frequency converter, model training of the D2Former can be facilitated.
[0090] like Figure 3 As shown, the disclosed embodiment provides a dual-decoder solution that uses complex ratio masking and complex spectral mapping to simultaneously estimate the amplitude and phase spectra, i.e., the spectrogram of the denoised target audio signal. Because the two decoders output more information, this dual-decoder solution can learn more and more versatile feature representations. Furthermore, while information output by one decoder may be lost, the information output by the two decoders can compensate for each other, thereby compensating for any lost information and improving the performance ceiling.
[0091] Optionally, the masked decoder includes a second dual-path hole module, a complex transpose layer, a first complex convolution layer, and a second complex convolution layer; the complex transpose layer is used to reset the frequency dimension; the first complex convolution layer is used to compress the second feature map from multiple feature channels to one feature channel; the second complex convolution layer is used to project the output of its previous layer to the masked domain.
[0092] like Figure 6As shown, the complex dual-path masked decoder includes a complex dual-path hole module 60 with a depth of 4, a complex transposed layer (Complex 2D Transposed Conv), a complex 2D convolution layer 61, a complex instance regularization, a complex activation function (Leaky Rectified Linear Unit, LeakyReLU), a complex 2D convolution layer 62, and a tanh activation function. Among them, the complex dual-path hole module 60 is recorded as the second dual-path hole module, and the internal structure of the complex dual-path hole module 60 is consistent with the internal structure of the complex dual-path hole module described above, which will not be repeated here. The complex 2D convolution layer 61 is recorded as the first complex convolution layer, and the complex 2D convolution layer 62 is recorded as the second complex convolution layer. Among them, the complex dual-path hole module 60 is used to aggregate information from the time dimension and the frequency dimension. For example, the output of the complex dual-path converter block repeated N times is adjusted to obtain a second feature map, and the size of the second feature map is B×C×T×F / 2×2. The second feature map can be used as the input of the complex dual-path masked decoder and the input of the complex dual-path spectrum decoder respectively. The complex dual-path hole module 60 in the complex dual-path masked decoder can perform feature extraction on the second feature map in the time dimension and the frequency dimension respectively. The dimension of the output of the complex dual-path hole module 60 is consistent with the dimension of its input. The complex transpose layer is used to reset the frequency dimension, for example, to reset the frequency dimension from F / 2 to F. The complex two-dimensional convolution layer 61 is used to compress the second feature map from C feature channels to 1 feature channel. The output of the complex two-dimensional convolution layer 61 is then subjected to complex instance regularization and complex activation function LeakyReLU operations. The complex two-dimensional convolution layer 62 can project the output of the complex activation function LeakyReLU to the masked domain, and then use the tanh activation function to limit the range of the cIRM estimation value.
[0093] like Figure 7 As shown, the complex dual-path spectrum decoder includes a complex dual-path hole module 70, a complex transpose layer, a complex activation function (Parametric Rectified Linear Unit, PReLU), a complex instance regularization, and a complex two-dimensional convolution layer. The complex transpose layer is used to reset the frequency dimension, for example, to reset the frequency dimension from F / 2 to F. The complex two-dimensional convolution layer is used to compress the second feature map from C feature channels to 1 feature channel.
[0094] This embodiment uses a fully complex D2Former network model to perform monophonic speech enhancement. Unlike existing models, this D2Former network model can simultaneously learn complex ratio masking and complex spectral mapping, that is, it combines the advantages of these two training objectives through a joint learning framework, thereby improving the performance upper limit of each of these two training objectives. In addition, by combining complex ratio masking and complex spectral mapping for monophonic speech enhancement using a fully complex D2Former network model, the speech quality after enhancement can be improved. In the D2Former network model, since the D2Former network model can be extended to the complex domain and is designed as a dual-path complex time-frequency self-attention architecture (for example, a complex dual-path converter block including a time converter and a frequency converter), it can effectively model complex time and frequency series. In addition, frequency recursion is performed based on time-dependent complex dilated convolution and a complex feedforward sequence memory network (CFSMN), and the dual-path learning structure is used to further improve the time-frequency feature representation in the encoder and decoder. Therefore, the D2Former network model can fully utilize complex operations, dual-path processing, and joint training objectives. Compared with existing models, the D2Former network model achieves the best denoising results on the VoiceBank+Demand benchmark and has the fewest parameters, for example, 0.87M.
[0095] Figure 8 The audio signal processing device provided by the embodiment of the present disclosure can execute the processing flow provided by the embodiment of the audio signal processing method, such as Figure 8 As shown, the audio signal processing device 80 includes:
[0096] An acquisition module 81 is used to acquire an original audio signal;
[0097] An encoding module 82 is configured to encode a first complex spectrum corresponding to the original audio signal to obtain a first feature map;
[0098] A processing module 83 is configured to process the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map;
[0099] A prediction module 84, configured to predict a complex ratio mask and a complex spectrum map based on the second feature map;
[0100] The generating module 85 is configured to generate a second complex spectrogram according to the complex ratio masking and the complex spectrum mapping, and generate a speech-enhanced target audio signal according to the second complex spectrogram.
[0101] Optionally, when the encoding module 82 encodes the first complex spectrum graph corresponding to the original audio signal, it is specifically used to: use a complex dual-path encoder to encode the first complex spectrum graph corresponding to the original audio signal, and the complex dual-path encoder includes a first convolution module, a second convolution module and a first dual-path hole module.
[0102] Optionally, the first convolution module is used to expand the first complex spectrum from one feature channel to multiple feature channels;
[0103] The second convolution module is used to reduce the frequency dimension corresponding to the first complex spectrum graph;
[0104] The first dual-path hole module is used to extract features from the first complex spectrum graph in a time dimension and a frequency dimension corresponding to the first complex spectrum graph.
[0105] Optionally, the first dual-path hole module includes a plurality of dual-path blocks connected sequentially, and in the connection direction, the hole factor of a subsequent dual-path block is twice the hole factor of a previous dual-path block.
[0106] Optionally, the input of any other dual path block among the multiple dual path blocks except the first dual path block is a concatenation result of the input of the first dual path block and the output of the previous dual path block of any dual path block.
[0107] Optionally, when the processing module 83 processes the time series and frequency series corresponding to the first feature graph respectively, it is specifically configured to:
[0108] Adjusting the first feature map into a time series;
[0109] Processing the time series using a self-attention mechanism through a time transformer to obtain an output of the time transformer;
[0110] Obtaining a frequency sequence according to the output of the time converter;
[0111] The frequency sequence is processed by a frequency converter using a self-attention mechanism to obtain an output of the frequency converter.
[0112] Optionally, when the processing module 83 obtains a frequency sequence based on the output of the time converter, it is specifically used to: connect the output of the time converter and the input of the time converter to obtain a first connection result; adjust the first connection result to obtain the frequency sequence; accordingly, the processing module 83 processes the frequency sequence through the frequency converter using a self-attention mechanism to obtain the output of the frequency converter, and is also used to: connect the output of the frequency converter and the input of the frequency converter to obtain a second connection result; adjust the second connection result to obtain an updated time series.
[0113] Optionally, when the prediction module 84 predicts the complex ratio mask and the complex spectrum mapping according to the second feature map, it is specifically configured to:
[0114] Using the second feature map as an input of a mask decoder, and outputting a complex ratio mask through the mask decoder;
[0115] The second feature map is used as an input of a spectrum decoder, and the spectrum decoder outputs a complex spectrum map.
[0116] Optionally, the masked decoder includes a second dual-path hole module, a complex transpose layer, a first complex convolution layer, and a second complex convolution layer;
[0117] The complex transpose layer is used to reset the frequency dimension;
[0118] The first complex convolutional layer is used to compress the second feature map from multiple feature channels into one feature channel;
[0119] The second complex convolutional layer is used to project the output of the previous layer into the masked domain.
[0120] Optionally, when the generating module 85 generates the second complex spectrum diagram according to the complex ratio masking and the complex spectrum mapping, it is specifically configured to:
[0121] generating an enhanced spectrogram based on the first complex spectrogram and the complex ratio mask;
[0122] The second complex spectrum map is generated according to the enhanced spectrum map and the complex spectrum mapping.
[0123] Figure 8 The audio signal processing device of the illustrated embodiment can be used to implement the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail here.
[0124] The internal functions and structure of the audio signal processing device are described above. The device can be implemented as an electronic device. Figure 9This is a schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. Figure 9 As shown, the electronic device includes a memory 91 and a processor 92 .
[0125] The memory 91 is used to store programs. In addition to the aforementioned programs, the memory 91 may also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, images, videos, and the like.
[0126] The memory 91 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0127] The processor 92 is coupled to the memory 91 and executes the program stored in the memory 91 to:
[0128] Get the original audio signal;
[0129] Encoding a first complex spectrogram corresponding to the original audio signal to obtain a first feature map;
[0130] Processing the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map;
[0131] predicting a complex ratio mask and a complex spectral map based on the second feature map;
[0132] A second complex spectrogram is generated according to the complex ratio mask and the complex spectrum mapping, and a speech-enhanced target audio signal is generated according to the second complex spectrogram.
[0133] Further, if Figure 9 As shown, the electronic device may further include: a communication component 93, a power component 94, an audio component 95, a display 96 and other components. Figure 9 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 9 Components shown.
[0134] The communication component 93 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 93 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 93 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0135] The power supply assembly 94 provides power to various components of the electronic device. The power supply assembly 94 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0136] The audio component 95 is configured to output and / or input audio signals. For example, the audio component 95 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 91 or transmitted via the communication component 93. In some embodiments, the audio component 95 also includes a speaker for outputting audio signals.
[0137] The display 96 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0138] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the audio signal processing method described in the above embodiment.
[0139] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0140] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing an audio signal, wherein: The method comprises: Get the original audio signal; Encoding a first complex spectrogram corresponding to the original audio signal to obtain a first feature map; Processing the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map; predicting a complex ratio mask and a complex spectral map based on the second feature map; generating a second complex spectrogram according to the complex ratio mask and the complex spectrum mapping, and generating a speech-enhanced target audio signal according to the second complex spectrogram; Among them, encoding the first complex spectrum graph corresponding to the original audio signal to obtain a first feature graph includes: using a complex dual-path encoder to expand the feature channel of the first complex spectrum graph; using the complex dual-path encoder to extract features of the expanded first complex spectrum graph in the time dimension and frequency dimension respectively; using the complex dual-path encoder to reduce the frequency dimension of the first complex spectrum graph after feature extraction to obtain the first feature graph.
2. The method according to claim 1, wherein The complex dual-path encoder includes a first convolution module, a second convolution module and a first dual-path hole module, wherein the first convolution module is used to expand the first complex spectrogram from one feature channel to multiple feature channels; The second convolution module is used to reduce the frequency dimension corresponding to the first complex spectrum; The first dual-path hole module is used to perform feature extraction on the first complex spectrum graph in the time dimension and the frequency dimension corresponding to the first complex spectrum graph respectively.
3. The method according to claim 2, wherein: The first dual-path hole module includes a plurality of dual-path blocks connected in sequence, and in the connection direction, a hole factor of a subsequent dual-path block is twice a hole factor of a previous dual-path block.
4. The method according to claim 3, wherein: The input of any other dual path block among the multiple dual path blocks except the first dual path block is a concatenation result of the input of the first dual path block and the output of the previous dual path block of any dual path block.
5. The method according to claim 1, wherein The time series and frequency series corresponding to the first feature graph are processed separately, including: Adjusting the first feature map into a time series; Processing the time series using a self-attention mechanism through a time transformer to obtain an output of the time transformer; obtaining a frequency sequence according to an output of the time converter; The frequency sequence is processed by a frequency converter using a self-attention mechanism to obtain an output of the frequency converter.
6. The method according to claim 5, wherein: According to the output of the time converter, a frequency sequence is obtained, including: connecting the output of the time converter and the input of the time converter to obtain a first connection result; Adjusting the first connection result to obtain the frequency sequence; Accordingly, after the frequency converter processes the frequency sequence using a self-attention mechanism to obtain the output of the frequency converter, the method further includes: connecting the output of the frequency converter and the input of the frequency converter to obtain a second connection result; The second connection result is adjusted to obtain an updated time series.
7. The method according to claim 1, wherein Predicting a complex ratio mask and a complex spectral map based on the second feature map, comprising: Using the second feature map as an input of a mask decoder, and outputting a complex ratio mask through the mask decoder; The second feature map is used as an input of a spectrum decoder, and the spectrum decoder outputs a complex spectrum map.
8. The method according to claim 7, wherein: The masked decoder includes a second dual-path hole module, a complex transpose layer, a first complex convolutional layer, and a second complex convolutional layer; The complex transpose layer is used to reset the frequency dimension; The first complex convolutional layer is used to compress the second feature map from multiple feature channels into one feature channel; The second complex convolutional layer is used to project the output of the previous layer into the masked domain.
9. The method according to claim 1, wherein: Generating a second complex spectrum map according to the complex ratio mask and the complex spectrum mapping, comprising: generating an enhanced spectrogram based on the first complex spectrogram and the complex ratio mask; The second complex spectrum map is generated according to the enhanced spectrum map and the complex spectrum mapping.
10. An audio signal processing device, wherein: include: An acquisition module, used to obtain the original audio signal; An encoding module, configured to encode a first complex spectrum graph corresponding to the original audio signal to obtain a first feature graph; a processing module, configured to process the time series and the frequency series corresponding to the first feature map respectively to obtain a second feature map; a prediction module, configured to predict a complex ratio mask and a complex spectrum map based on the second feature map; a generating module, configured to generate a second complex spectrogram according to the complex ratio masking and the complex spectrum mapping, and generate a speech-enhanced target audio signal according to the second complex spectrogram; In which, the encoding module is used to perform the following steps to obtain the first feature map: using a complex dual-path encoder to expand the feature channel of the first complex spectrum map; using the complex dual-path encoder to extract features of the expanded first complex spectrum map in the time dimension and frequency dimension respectively; using the complex dual-path encoder to reduce the frequency dimension of the first complex spectrum map after feature extraction to obtain the first feature map.
11. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Speech enhancement method, device and equipment
CN114694672A
Audio enhancement method and device, electronic equipment and readable storage medium
CN114974292A