Voice enhancement method, device, electronic device and storage medium
By introducing the gated loop unit and binary gate processing blocks in the voice enhancement network, the calculation complexity of the voice enhancement process is reduced and real-time improvement is improved, and the problem of insufficient real-time in the prior art is solved.
Patent Information
- Application Number
- CN202510053357.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Existing speech enhancement technologies have challenges in computational complexity and real-time performance, especially when dealing with acoustic echoes and ambient noise in real-time communication systems.
A voice-enhancing network including an encoder and a decoder is adopted, which consists of a plurality of sequentially cascaded processing blocks, which include a gated loop unit and a binary gate for frame mixing mapping, band mixing mapping and band compression, reducing computational complexity and improving real-timeness.
By reducing the computational complexity in the speech enhancement process, the real-time performance of speech enhancement is significantly improved, and noise and echoes in real-time communication systems can be handled more effectively.
Smart Images

Figure CN119541522B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech processing technology, and in particular to a speech enhancement method, device, electronic device and storage medium. Background Art
[0002] Acoustic echoes and environmental noise are common in real-time communication systems, which reduce the voice quality during communication. With the development of artificial intelligence technology, speech enhancement can be achieved through the architecture of deep neural networks in related technologies. However, the computational complexity of this method is often very high, and the real-time performance of speech enhancement needs to be improved. Summary of the invention
[0003] The following is a summary of the subject matter of the detailed description of the present disclosure. This summary is not intended to limit the scope of the claims.
[0004] The embodiments of the present disclosure provide a speech enhancement method, device, electronic device and storage medium, which can reduce the computational complexity during speech enhancement and improve the real-time performance of speech enhancement.
[0005] On the one hand, an embodiment of the present disclosure provides a method for speech enhancement, comprising:
[0006] Acquire a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the processing blocks include a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate updating the hidden state features of the gated recurrent unit or keeping the hidden state features unchanged;
[0007] Based on the first processing block in the encoder, frame hybrid mapping is performed on the initial signal feature to obtain a first mapping feature, band hybrid mapping is performed on the first mapping feature to obtain a second mapping feature, band compression is performed on the second mapping feature to obtain the output of the first processing block, and the output of the first processing block is input to the next processing block until the encoding result of the encoder is obtained;
[0008] Decoding the encoding result of the encoder based on the decoder to obtain target signal characteristics;
[0009] An enhanced speech signal is restored from the speech signal to be processed based on the target signal feature.
[0010] On the other hand, the present disclosure also provides a speech enhancement device, including:
[0011] An input module is used to obtain a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the processing blocks include a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate whether to update the hidden state features of the gated recurrent unit or keep the hidden state features unchanged;
[0012] An encoding module, configured to perform frame hybrid mapping on the initial signal feature to obtain a first mapping feature based on the first processing block in the encoder, perform band hybrid mapping on the first mapping feature to obtain a second mapping feature, perform band compression on the second mapping feature to obtain an output of the first processing block, and input the output of the first processing block to the next processing block until an encoding result of the encoder is obtained;
[0013] A decoding module, used for decoding the encoding result of the encoder based on the decoder to obtain target signal characteristics;
[0014] The restoration module is used to restore the to-be-processed speech signal to obtain an enhanced speech signal based on the target signal feature.
[0015] Furthermore, the encoding module is also used to:
[0016] Normalizing the initial signal features, and performing frame mixing mapping on the normalized initial signal features through the gated cycle unit;
[0017] Fully connect the output of the gated recurrent unit to obtain the first mapping feature.
[0018] Furthermore, the encoding module is also used to:
[0019] For any current time step among the multiple time steps, obtaining the initial state update probability and the initial state update probability change value of the gated recurrent unit in the previous time step of the current time step, and determining the target state update probability of the gated recurrent unit in the current time step according to the initial state update probability and the initial state update probability change value;
[0020] Rounding the target state update probability to obtain a gating parameter corresponding to the binary gate in the current time step;
[0021] The hidden state feature output by the gated cyclic unit in the current time step is determined according to the gating parameter, and frame mixing mapping is performed on the normalized initial signal feature through the gated cyclic unit based on the hidden state feature.
[0022] Furthermore, the encoding module is also used to:
[0023] Determining the initial state update probability corresponding to the current time step according to the gating parameter corresponding to the current time step;
[0024] The initial state update probability change value corresponding to the current time step is determined according to the gating parameter corresponding to the current time step and the target state update probability.
[0025] Furthermore, the encoding module is also used to:
[0026] Acquire a sample speech signal and a sample label signal corresponding to the sample speech signal, and call the speech enhancement network to perform speech enhancement on the sample speech signal;
[0027] The logarithmic variance of the enhancement result of the sample speech signal is determined by the logarithmic variance estimation module, the target loss is determined according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance, and the speech enhancement network is trained based on the target loss.
[0028] Furthermore, the encoding module is also used to:
[0029] Determine a target loss for a first training phase according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, wherein the target loss for the first training phase is used to constrain the accuracy of the logarithmic variance;
[0030] Determine a target loss for a second training phase according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, wherein the target loss for the second training phase is used to assign corresponding weights to the time-frequency units according to the uncertainty of the time-frequency units when determining the enhanced result of the sample speech signal and the difference between the sample label signal;
[0031] The speech enhancement network is trained sequentially based on the target loss of the first training stage and the target loss of the second training stage.
[0032] Furthermore, the encoding module is also used to:
[0033] Determine a real part difference of a complex spectrum, an imaginary part difference of a complex spectrum, and an amplitude spectrum difference between the enhancement result of the sample speech signal and the sample label signal;
[0034] The real part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the real part of the complex spectrum as an exponent, the imaginary part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the imaginary part of the complex spectrum as an exponent, and a first loss is determined according to the sum of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, the weighted real part difference of the complex spectrum, and the weighted imaginary part difference of the complex spectrum;
[0035] Weighting the amplitude spectrum difference according to a natural exponential function with the amplitude spectrum logarithmic variance as an exponent, and determining a second loss according to the sum of the amplitude spectrum logarithmic variance and the weighted amplitude spectrum difference;
[0036] The first loss and the second loss are weightedly summed to obtain the target loss of the first training stage.
[0037] Furthermore, the encoding module is also used to:
[0038] Determine a signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine an update rate loss according to a difference between an update rate of the binary gate and a preset update rate threshold;
[0039] The first loss and the second loss are weightedly summed to obtain a first uncertainty loss, and the first uncertainty loss, the signal-to-noise ratio loss, and the update rate loss are weightedly summed to obtain a target loss for the first training stage.
[0040] Furthermore, the encoding module is also used to:
[0041] Determine a real part difference of a complex spectrum, an imaginary part difference of a complex spectrum, and an amplitude spectrum difference between the enhancement result of the sample speech signal and the sample label signal;
[0042] Weighting the real part difference of the complex spectrum according to the logarithmic variance of the real part of the complex spectrum, weighting the imaginary part difference of the complex spectrum according to the logarithmic variance of the imaginary part of the complex spectrum, and determining a third loss according to the sum of the weighted real part difference of the complex spectrum and the weighted imaginary part difference of the complex spectrum;
[0043] weighting the amplitude spectrum difference according to the amplitude spectrum logarithmic variance, and determining the weighted amplitude spectrum difference as a fourth loss;
[0044] The target loss of the second training stage is obtained by weighted summing the third loss and the fourth loss.
[0045] Furthermore, the encoding module is also used to:
[0046] Determine a signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine an update rate loss according to a difference between an update rate of the binary gate and a preset update rate threshold;
[0047] The third loss and the fourth loss are weightedly summed to obtain a second uncertainty loss, and the second uncertainty loss, the signal-to-noise ratio loss and the update rate loss are weightedly summed to obtain the target loss of the second training stage.
[0048] Furthermore, the encoding module is also used to:
[0049] Segmenting the first mapping feature along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature;
[0050] Normalizing the first intra-band relationship sub-feature, and convolving the normalized first intra-band relationship sub-feature in a sub-band dimension along a time axis to obtain a first convolution feature;
[0051] A first product of the first convolution feature and the second in-band relation sub-feature is determined, and the second mapping feature is obtained based on a sum of the first product and the second in-band relation sub-feature.
[0052] Furthermore, the encoding module is also used to:
[0053] Convolving the first mapping feature in the channel dimension along the time axis to obtain a second convolution feature;
[0054] The second convolution feature is normalized, and the normalized second convolution feature is segmented along the channel dimension to obtain the first in-band relationship sub-feature and the second in-band relationship sub-feature.
[0055] Furthermore, the encoding module is also used to:
[0056] Convolving the first product and the sum of the second in-band relationship sub-features in the channel dimension along the time axis to obtain the second mapping feature.
[0057] Furthermore, the encoding module is also used to:
[0058] Normalizing the second mapping features, and performing parallel convolution on the normalized second mapping features to obtain compressed third convolution features and fourth convolution features;
[0059] Performing nonlinear activation on the fourth convolutional feature to obtain an attention weight;
[0060] The third convolution feature is weighted based on the attention weight to obtain the output of the first processing block.
[0061] On the other hand, an embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned speech enhancement method when executing the computer program.
[0062] On the other hand, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned speech enhancement method.
[0063] On the other hand, the embodiment of the present disclosure further provides a computer program product, which includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned speech enhancement method.
[0064] The disclosed embodiment includes at least the following beneficial effects: by acquiring a speech signal to be processed, extracting features of the speech signal to be processed, obtaining initial signal features, inputting the initial signal features into a speech enhancement network, performing frame hybrid mapping on the initial signal features in a first processing block in an encoder based on the speech enhancement network to obtain a first mapping feature, performing band hybrid mapping on the first mapping feature to obtain a second mapping feature, performing band compression on the second mapping feature to obtain the output of the first processing block, and inputting the output of the first processing block into the next processing block until the encoding result of the encoder is obtained, and the computational complexity can be reduced by band compression after performing frame hybrid mapping and band hybrid mapping, and since the processing block includes a gated cyclic unit, the gated cyclic unit is configured with a binary gate, and the binary gate is used to indicate updating the hidden state features of the gated cyclic unit or keeping the hidden state features unchanged, so that the hidden state features of the gated cyclic unit can be flexibly controlled by the binary gate, further reducing the computational complexity, so that the encoding result of the encoder is subsequently decoded based on a decoder to obtain a target signal feature, and when an enhanced speech signal is restored from the speech signal to be processed based on the target signal feature, the real-time performance of speech enhancement can be improved.
[0065] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0067] Figure 1A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;
[0068] Figure 2 An optional flow chart of the speech enhancement method provided in the embodiment of the present disclosure;
[0069] Figure 3 A schematic diagram of an optional structure of a processing block in a speech enhancement network provided in an embodiment of the present disclosure;
[0070] Figure 4 The embodiments of the present disclosure provide Figure 3 An optional schematic diagram of a frame mixing module in;
[0071] Figure 5 An optional schematic diagram of a logarithmic variance estimation module provided in an embodiment of the present disclosure;
[0072] Figure 6 The embodiments of the present disclosure provide Figure 3 An optional structural diagram of a mid-band hybrid module;
[0073] Figure 7 The embodiments of the present disclosure provide Figure 3 An optional schematic diagram of a neutron band compression / decompression module;
[0074] Figure 8 An optional schematic diagram of frequency band merging provided in an embodiment of the present disclosure;
[0075] Fig. 9 Another optional structural diagram of a processing block in a speech enhancement network provided by an embodiment of the present disclosure;
[0076] Fig.10 An optional schematic diagram of performing frequency band segmentation provided in an embodiment of the present disclosure;
[0077] Fig.11 The embodiments of the present disclosure provide Fig.10 An optional schematic diagram of frequency band segmentation of signal sub-features in ;
[0078] Fig.12 Another optional schematic diagram of performing frequency band segmentation provided in an embodiment of the present disclosure;
[0079] Fig.13 An optional overall flow chart of the speech enhancement method provided in the embodiment of the present disclosure;
[0080] Fig.14 The embodiments of the present disclosure provide Fig.13 An optional structural diagram of a speech enhancement model in FIG.
[0081] Fig.15 The embodiments of the present disclosure provide Fig.13 Another optional structural diagram of the speech enhancement model;
[0082] Fig.16 A schematic diagram of the structure of a speech enhancement device provided in an embodiment of the present disclosure;
[0083] Fig.17 A partial structural block diagram of a terminal provided in an embodiment of the present disclosure;
[0084] Fig.18 A partial structural block diagram of a server provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0085] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0086] It should be noted that in various specific embodiments of the present disclosure, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present disclosure needs to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data used to enable the normal operation of the embodiment of the present disclosure will be obtained.
[0087] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0088] To facilitate understanding of the technical solution provided by the embodiments of the present disclosure, some key terms used in the embodiments of the present disclosure are explained here:
[0089] Frequency bin: It is used to describe the signal in the frequency domain, which can be understood as the interval or resolution between two adjacent points in the spectrum diagram.
[0090] Multilayer Perceptron (MLP): is a deep learning model based on a feedforward neural network. It is a multi-layer structure composed of multiple neurons, and each neuron layer is connected to all neurons in the previous layer. The multilayer perceptron receives external data through the input layer, extracts features through a series of hidden layers, and outputs the results through the output layer.
[0091] Acoustic Echo Cancellation (AEC): A key signal processing technology that uses an adaptive filter to simulate the echo path and estimates and eliminates the echo through an algorithm. It is mainly used to eliminate the echo generated by the speaker sound being picked up by the microphone during voice communication.
[0092] In real-time communication systems, acoustic echo and environmental noise are two common problems that reduce the voice quality during communication. With the development of artificial intelligence technology, especially the progress in the field of deep learning, the related technologies usually use the architecture of deep neural networks to achieve speech enhancement. Deep neural networks have powerful nonlinear modeling capabilities and can learn complex speech features, thereby eliminating echoes and suppressing noise while retaining the useful information in the original speech signal as much as possible. However, due to the complexity of the deep neural network structure, a large number of parameters and the number of layers, the computational complexity of this method is often very high. High computational complexity not only increases resource consumption, but also may cause processing delays, thereby affecting the real-time performance of speech enhancement.
[0093] Based on this, the embodiments of the present disclosure provide a speech enhancement method, device, electronic device and storage medium, which can reduce the computational complexity during speech enhancement and improve the real-time performance of speech enhancement.
[0094] Reference Figure 1 , Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure, the implementation environment includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected via a communication network.
[0095] Exemplarily, a speech signal to be processed is obtained in the terminal 101, and the speech signal to be processed is sent to the server 102. In the server 102, feature extraction is performed on the speech signal to be processed to obtain initial signal features, and the initial signal features are input into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, and the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence. In the first processing block in the encoder, the initial signal features are subjected to frame mixed mapping to obtain a first mapping feature, the first mapping feature is subjected to band mixed mapping to obtain a second mapping feature, the second mapping feature is subjected to band compression to obtain the output of the first processing block, and the output of the first processing block is input into the next processing block until the encoding result of the encoder is obtained, the encoding result of the encoder is decoded based on the decoder to obtain the target signal features, and the enhanced speech signal is restored from the speech signal to be processed based on the target signal features, and the enhanced speech signal is sent to the terminal 101 to complete the speech enhancement process.
[0096] It is understandable that the speech enhancement method provided in the embodiment of the present disclosure may also be independently executed in the terminal 101.
[0097] Server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, server 102 may also be a node server in a blockchain network.
[0098] The terminal 101 may be a mobile phone, a computer, an intelligent voice interaction device, an intelligent wearable device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present disclosure.
[0099] Reference Figure 2 , Figure 2 An optional flow chart of a voice enhancement method provided in an embodiment of the present disclosure. The voice enhancement method can be executed by a terminal, can be executed by a server, or can be executed by a terminal and a server in cooperation. The voice enhancement method includes but is not limited to the following steps S201 to S204.
[0100] Step S201: Obtain a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network.
[0101] Among them, the speech signal to be processed includes a mixed speech signal and a reference speech signal. The mixed speech signal is a near-end speech signal containing a target speech signal and interferences such as an echo signal, and is a speech signal captured by a sound capture device such as a microphone. The target speech signal is the speech signal as the main body in the mixed speech signal, and can be the speech signal emitted by the current speaker collected by the microphone. For example, when speaker A is talking to speaker B, for speaker A's microphone, speaker A's speech signal is the target speech signal; the reference speech signal is a speech signal used as a benchmark for echo prediction, and can be a far-end speech signal. The far-end speech signal is the speech signal sent by the other end of the call. For example, continuing with the above example, the speech signal transmitted by speaker B to speaker A is the far-end speech signal; the echo signal is a signal collected by the microphone when collecting the target speech signal due to the presence of an echo when the speaker plays the reference speech signal, and is obtained based on echo prediction; the error speech signal is used to indicate the difference between the mixed speech signal and the echo signal. The difference between is a signal obtained based on a mixed speech signal and an echo signal, for example, it can be a signal obtained after removing the echo signal from the mixed speech signal; the initial signal feature is a signal feature obtained after feature extraction of the speech signal to be processed, the speech enhancement network includes an encoder and a decoder connected to each other, the speech enhancement network can be a cyclic UNet network, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the encoder is used to encode the initial signal feature or other features obtained based on the initial signal feature, the decoder is used to decode the encoded initial signal feature or other features obtained based on the encoded initial signal feature, the encoder and the decoder are constructed symmetrically, for example, three processing blocks cascaded in sequence are provided in the encoder, and correspondingly, three processing blocks cascaded in sequence are also provided in the decoder; the processing block includes a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate updating the hidden state feature of the gated recurrent unit or keeping the hidden state feature unchanged.
[0102] Specifically, a mixed voice signal and a reference voice signal corresponding to a target voice signal in the mixed voice signal are obtained from a microphone. The mixed voice signal and the reference voice signal are input into an echo prediction model for echo prediction, and the echo component in the mixed voice signal is identified by analyzing the frequency, amplitude, phase and other characteristics of the mixed voice signal and the reference voice signal to obtain a predicted echo signal. Then, the echo signal and the mixed voice signal are input into an echo cancellation model, and an error voice signal is obtained based on the difference between the mixed voice signal and the echo signal. Echo cancellation is performed based on the mixed voice signal and the reference voice signal to obtain an echo signal and an error voice signal, so that a high-quality voice signal can be obtained, which provides good data for subsequent noise suppression.
[0103] Next, the mixed speech signal, the reference speech signal, the error speech signal and the echo signal are respectively converted into the time-frequency domain, wherein the time-frequency domain conversion is used to convert the mixed speech signal, the reference speech signal, the error speech signal and the echo signal from the time domain to the frequency domain, and the time-frequency domain conversion can be a short-time Fourier transform. The four signals are converted from the time domain to the frequency domain to obtain a mixed speech spectrum corresponding to the mixed speech signal, a reference speech spectrum corresponding to the reference speech signal, an error speech spectrum corresponding to the error speech signal and an echo spectrum corresponding to the echo signal. In the process of performing the time-frequency domain conversion, the mixed speech signal, the reference speech signal, the error speech signal and the echo signal are represented in complex form in the frequency domain, so the mixed speech spectrum, the reference speech spectrum, the error speech spectrum and the echo spectrum all have corresponding real and imaginary parts. The real and imaginary parts of the mixed speech spectrum, the reference speech spectrum, the error speech spectrum and the echo spectrum are respectively spliced along the channel dimension to obtain the speech signal to be processed, wherein the characteristic dimension of the speech signal to be processed can be , B is the batch size, T is the number of frames, and F is the number of frequency units. Feature extraction is performed on the speech signal to be processed, and the real part feature of the complex spectrum and the imaginary part feature of the complex spectrum are extracted from the speech signal to be processed. The real part feature of the complex spectrum and the imaginary part feature of the complex spectrum are used as initial signal features, and the initial signal features are input into the speech enhancement network for speech enhancement.
[0104] Alternatively, the feature dimension of the speech signal to be processed can also be , C is the number of channels, which can be 8 at this time. In the process of extracting features from the processed speech signal, the processed speech signal can also be input into a two-dimensional convolution for feature extraction. The obtained features are normalized and then passed through an activation function to obtain the initial signal features. The feature dimension of the initial signal features can be , E is the embedding dimension. The initial signal features are input into the speech enhancement network for speech enhancement. Converting the channel dimension to the embedding dimension through two-dimensional convolution can be understood as a dimensionality reduction process. In the subsequent steps, processing based on the initial signal features can remove some redundant or unimportant features, making the subsequent processing steps simpler, thereby reducing the amount of calculation and complexity of the subsequent steps, and effectively improving the real-time performance of the enhanced speech signal.
[0105] Step S202: Based on the first processing block in the encoder, perform frame mixing mapping on the initial signal feature to obtain a first mapping feature, perform band mixing mapping on the first mapping feature to obtain a second mapping feature, perform band compression on the second mapping feature to obtain the output of the first processing block, and input the output of the first processing block to the next processing block until the encoding result of the encoder is obtained.
[0106] Among them, the processing block of the encoder is used to perform frame mixing mapping, band mixing mapping and sub-band compression operations on the input features, and the processing block of the decoder is used to perform frame mixing mapping, band mixing mapping and sub-band decompression operations on the input features. The processing block includes a frame mixing module, a band mixing module and a sub-band compression / decompression module; frame mixing mapping is modeling in the time dimension, and band mixing mapping is modeling in the sub-band dimension; the first mapping feature is a feature obtained after frame mixing mapping, and the second mapping feature is a feature obtained after the first mapping feature is subjected to band mixing mapping. The second mapping feature can be regarded as an inter-band relationship feature used to indicate the relationship between bands.
[0107] Reference Figure 3 , Figure 3 A schematic diagram of an optional structure of a processing block in a speech enhancement network provided in an embodiment of the present disclosure, wherein the first processing block of the encoder of the speech enhancement network has an input of an initial signal feature. X . The initial signal characteristics X Input into the frame mixing module for frame mixing mapping to obtain the first mapping feature , the first mapping feature With the initial signal characteristics X Add together to get the first fusion feature . The first fusion feature Input to the band mixing module for band mixing mapping to obtain the second mapping feature , the second mapping feature With the first fusion feature Add them together to get the second fusion feature Finally, the second fusion feature Input to the sub-band compression / decompression module for band compression to obtain the output of the processing block .
[0108] In one possible implementation, in the process of performing frame mixed mapping on the initial signal features to obtain the first mapping features, the initial signal features can be normalized, frame mixed mapping is performed on the normalized initial signal features through a gated cycle unit, and the output of the gated cycle unit is fully connected to obtain the first mapping features.
[0109] Specifically, the initial signal features are normalized so that the initial signal features have the same data distribution. The normalized initial signal features are input into the gated recurrent unit for feature mapping along the time dimension, and the sequence information of the speech signal is extracted from the initial signal features by the gated recurrent unit or the time-dependent features of the speech signal are captured. These information and features reflect the characteristics of the speech signal changing with time steps. Since the feature dimension of the mapping feature output by the gated recurrent unit may be inconsistent with the feature dimension of the initial signal feature, the mapping feature is input into the fully connected layer for further feature extraction to learn complex feature combinations, and the feature dimension of the mapping feature is mapped by adjusting the number of output neurons of the fully connected layer. When performing feature mapping, the mapping feature is mapped to a new feature space, and the feature dimension of the new feature space is consistent with the feature dimension of the initial signal feature, so as to obtain a first mapping feature consistent with the feature dimension of the initial signal feature. By adopting the gated recurrent unit, the noise in the initial signal feature can be filtered out while capturing the timing information and time dependency in the initial signal feature, thereby improving the robustness of the encoder and helping to improve the accuracy of the decoder in predicting the speech signal.
[0110] Reference Figure 4 , Figure 4 The embodiments of the present disclosure provide Figure 3 An optional schematic diagram of the frame mixing module in FIG. X Input to the frame mixing module, the initial signal characteristics X First, the normalization operation is performed through the normalization layer to obtain the normalized initial signal features, and then the normalized initial signal features are input into the jump gated recurrent unit for feature mapping to obtain the mapping features, and finally the mapping features are input into the fully connected layer for feature mapping again to obtain the first mapping features .
[0111] In one possible implementation, the speech enhancement network is configured to iterate through multiple time steps. In the process of performing frame mixing mapping on the normalized initial signal features through the gated recurrent unit, specifically, for any current time step among the multiple time steps, the initial state update probability and the initial state update probability change value of the gated recurrent unit in the previous time step of the current time step are obtained, and the target state update probability of the gated recurrent unit in the current time step is determined according to the initial state update probability and the initial state update probability change value. Among them, the initial state update probability of the current time step is the state update probability of the final output after passing through the binary gate in the previous time step. For example, the state update probability of the final output of the previous time step is , then the probability of updating the initial state at the current time step is ; The initial state update probability change value of the current time step is the state update probability change value finally determined after passing through the binary gate in the previous time step; The target state update probability is the state update probability without passing through the binary gate in the current time step.
[0112] Specifically, for any current time step among multiple time steps, the initial state update probability and the initial state update probability change value of the current time step are obtained, and the probability difference between the initial state update probability and the total probability is calculated, where the value of the total probability can be 1. The one with the smallest value between the initial state update probability change value and the probability difference is selected as the state update probability change value of the previous time step, and the state update probability change value is added to the initial state update probability to obtain the target state update probability of the gated recurrent unit in the current time step. The current time step is obtained. t The probability of updating the target state The process can be expressed by the following formula:
[0113] ,
[0114] in, For the time step The probability of state update after passing through the binary gate, For the time step The state update probability change value after passing through the binary gate, is the minimum function, which is used to find the minimum value of a set of data. It should also be noted that the target state update probability of the current time step increases by the state update probability change value that did not pass through the binary gate in the previous time step, and the state update probability change value that did not pass through the binary gate in the previous time step It can be obtained by the following formula:
[0115]
[0116] in, represents the proportionality coefficient, represents the weight coefficient, Represents the scale offset, Represents the sigmoid function. Weight coefficient and scale offset Can be a parameter in a fully connected layer.
[0117] Next, the target state update probability is rounded to get the gating parameter corresponding to the binary gate in the current time step , according to the gating parameters, the hidden state features output by the gated recurrent unit in the current time step are determined, and the normalized initial signal features are frame mixed and mapped through the gated recurrent unit based on the hidden state features. Specifically, the target state update probability is rounded based on the rounding rule to obtain the gating parameters corresponding to the binary gate in the current time step, where the gating parameters are used to control whether the hidden state features of the current time step contain the hidden state features of the previous time step. The value of the gating parameters is The hidden state features output by the gated recurrent unit in the current time step are determined according to the gating parameters, and the normalized initial signal features are frame mixed and mapped through the gated recurrent unit based on the hidden state features. Therefore, the current time step can be obtained by the following formula: t The gating parameters ,
[0118] ,
[0119] in, is the current time step t The target state update probability, is a rounding function used to update the target state probability Converted to an integer. Also, the current time step t The hidden state features It can be obtained by the following formula:
[0120] ,
[0121] in, is the initial hidden state feature of time step t without passing through the binary gate, the initial hidden state feature It can be expressed by the formula get, is the time step after passing through the binary gate The hidden state features of For the time step t The initial signal features input to the speech enhancement network when = 1, the hidden state feature is the gating parameter and the initial hidden state features The product of , that is, to update the hidden state features of the current time step; when = 0, the hidden state feature Equal to the time step The hidden state features , that is, the hidden state features of the current time step are not updated, and the previous time step is used The hidden state features .
[0122] The gated parameters corresponding to the binary gates are used to control the hidden state features output by the gated recurrent unit in the current time step, so that the hidden state features do not need to be updated in each time step, especially in the time step where the hidden state features do not change much. This effectively reduces the computational complexity in the reasoning process and helps to improve the real-time performance of the speech enhancement process.
[0123] In one possible implementation, after rounding the target state update probability to obtain the gating parameter corresponding to the current time step, the initial state update probability change value corresponding to the current time step can be determined based on the gating parameter corresponding to the current time step and the target state update probability. Specifically, the final state update probability output after the current time step passes through the binary gate is determined based on the gating parameter corresponding to the current time step and the target state update probability; then, the final state update probability change value output after the current time step passes through the binary gate is obtained based on the gating parameter corresponding to the current time step, the initial state update probability change value, and the final state update probability output in the previous time step. At time step t The probability of updating the state and the state update probability change value at time step t It can be obtained by the following formula:
[0124] ,
[0125] ,
[0126] in, is the gating parameter at time step t, is the target state update probability at time step t, is the time step t The state update probability change value that does not pass through the binary gate, is the time step The state update probability change value after passing through the binary gate. Since the binary gate only processes parameters of the two states of 1 or 0, the initial state update probability and the initial state update probability change value of the current time step are updated through the gating parameters of the binary gate. Even if the input initial signal characteristics are interfered by a certain degree of noise, the gated recurrent unit can output a relatively stable state update probability and state update probability change value, which improves the stability of the processing block processing features, thereby improving the robustness of the speech enhancement network to noise and interference.
[0127] In a possible implementation, the speech enhancement network is further provided with a logarithmic variance estimation module. Before the initial signal feature is input into the speech enhancement network, a sample speech signal and a sample label signal corresponding to the sample speech signal are obtained, the speech enhancement network is called to perform speech enhancement on the sample speech signal, the logarithmic variance of the enhancement result of the sample speech signal is determined by the logarithmic variance estimation module, the target loss is determined according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance, and the speech enhancement network is trained based on the target loss. The logarithmic variance estimation module is used to estimate the uncertainty in the enhancement result; the sample speech signal is a speech signal for training, including a sample mixed speech signal, a sample reference speech signal, a sample echo signal, and a sample error speech signal; the sample label signal is used as a reference speech signal for comparison with the enhancement result of the sample speech signal; the enhancement result of the sample speech signal includes a sample enhanced real signal corresponding to the sample real signal, a sample enhanced imaginary signal corresponding to the sample imaginary signal, and a sample enhanced speech signal corresponding to the sample speech signal.
[0128] Specifically, a sample speech signal and a sample label signal corresponding to the sample speech signal are obtained, and the sample speech signal is feature extracted to obtain a sample real signal and a sample imaginary signal. The sample real signal, the sample imaginary signal and the sample speech signal are input into a speech enhancement network, and the sample speech signal is speech enhanced to obtain a sample enhanced real signal, a sample enhanced imaginary signal and a sample enhanced speech signal. The logarithmic variance of the sample enhanced real signal, the sample enhanced imaginary signal and the sample enhanced speech signal is determined by a logarithmic variance estimation module, and the target loss is determined according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance, and the speech enhancement network is guided to perform training based on the target loss. By introducing a logarithmic variance estimation module to estimate the uncertainty of the enhancement result of the sample speech signal for training, the speech enhancement network can focus on the content pointed to by the uncertainty of the enhancement result of the sample speech signal, thereby guiding the training direction of the speech enhancement network and making the training of the speech enhancement network more targeted.
[0129] Among them, refer to Figure 5 , Figure 5 An optional schematic diagram of a logarithmic variance estimation module provided in an embodiment of the present disclosure, wherein the enhanced result of the sample speech signal is The input is sent to the logarithmic variance estimation module for uncertainty estimation. The logarithmic variance estimation module consists of three submodules cascaded in sequence. For the first submodule, the enhanced result of the sample speech signal is Input to the convolution layer for feature extraction to obtain intermediate features f , and then the intermediate features fThe output of the first submodule is input into the nonlinear activation function for nonlinear activation to introduce the nonlinear characteristics of the intermediate features, and the output of the first submodule is input into the second submodule until the output of the third submodule is obtained. At this time, the output of the third submodule is the variance corresponding to the enhanced result of the sample speech signal. θ, The nonlinear activation function may be an exponential linear unit function (ELU). Finally, the variance corresponding to the enhancement result of the sample speech signal is θ Perform natural logarithmic transformation to obtain the logarithmic variance corresponding to the enhanced result of the sample speech signal v .
[0130] In a possible implementation, in the process of performing band mixed mapping on the first mapping feature to obtain the second mapping feature, the first mapping feature can be specifically segmented along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature, the first in-band relationship sub-feature is normalized, the normalized first in-band relationship sub-feature is convolved in the sub-band dimension along the time axis to obtain a first convolution feature, a first product of the first convolution feature and the second in-band relationship sub-feature is determined, and the second mapping feature is obtained based on the sum of the first product and the second in-band relationship sub-feature.
[0131] Specifically, when modeling the inter-band relationship of the first mapping feature, first, channel mapping is performed on the first mapping feature, and the first mapping feature after channel mapping is divided along the channel dimension to obtain a first intra-band relationship sub-feature and a second intra-band relationship sub-feature, wherein the first intra-band relationship sub-feature is used to control the information flow in the first mapping feature, and the second intra-band relationship sub-feature is used to directly transmit the information in the first mapping feature. The first intra-band relationship sub-feature is normalized, and when the normalized first intra-band relationship sub-feature is subjected to frequency band mapping, the intra-band relationship modeling is performed along the time axis, that is, a convolution operation is performed on the sub-band dimension through a first convolution to obtain a first convolution feature, wherein the first convolution is used for frequency band mapping, the convolution kernel size may be 1×3, and the step size of the convolution operation may be 3. Then, the first convolution feature and the second intra-band relationship sub-feature are element-wise multiplied to obtain a first product, and a second mapping feature for indicating the intra-band relationship is obtained based on the sum of the first product and the second intra-band relationship sub-feature. By dividing the first mapping feature into two sub-features, it is helpful to capture the complex relationship in the first mapping feature, and the speech enhancement network can gradually introduce sub-band information in the process of learning features, so that the speech enhancement network can process input data more flexibly during encoding and decoding, thereby improving the quality of the enhanced speech signal. Furthermore, element-wise multiplication of the first convolution feature and the second intra-band relationship sub-feature can further enhance the common features in the first convolution feature and the second intra-band relationship sub-feature, ignore or reduce the attention to irrelevant features or features with weak correlation, and effectively improve the ability to enhance speech signals.
[0132] In a possible implementation, in the process of segmenting the first mapping feature along the channel dimension to obtain the first in-band relationship sub-feature and the second in-band relationship sub-feature, the first mapping feature can be convolved along the time axis in the channel dimension to obtain the second convolution feature, the second convolution feature is normalized, and the normalized second convolution feature is segmented along the channel dimension to obtain the first in-band relationship sub-feature and the second in-band relationship sub-feature.
[0133] Specifically, channel mapping is performed on the first mapping feature along the time axis, and a convolution operation is performed on the first mapping feature in the channel dimension through a second convolution to increase the number of channels of the first mapping feature to obtain a second convolution feature, wherein the second convolution is used for channel mapping, and the convolution kernel size can be 1×1. For example, the feature dimension of the first mapping feature is , then the feature dimension of the second convolutional feature can be Then, the second convolution feature is normalized, and the normalized second convolution feature is segmented along the channel dimension to obtain the first in-band sub-feature and the second in-band sub-feature. For example, the feature dimension of the normalized second convolution feature can be , after segmentation along the channel dimension, the channel dimension of the sub-features in the first band can be , the channel dimension of the sub-features in the second band can be , or, the channel dimension of the sub-features in the first band can be , the channel dimension of the sub-features in the second band can be ,make + =2E. By increasing the channel dimension of the first mapping feature, the channel dimensions of the first in-band relation sub-feature and the second in-band relation sub-feature are made consistent with the first mapping feature, which is helpful for subsequent feature fusion operations such as addition or multiplication.
[0134] Alternatively, the first mapping feature may be normalized first, channel mapping may be performed on the normalized first mapping feature along the time axis, and a convolution operation may be performed on the normalized first mapping feature in the channel dimension through a second convolution to increase the number of channels of the normalized first mapping feature to obtain a second convolution feature.
[0135] In a possible implementation, the first product of the first convolution feature and the second in-band relation sub-feature is determined, and when the second mapping feature is obtained based on the sum of the first product and the second in-band relation sub-feature, the sum of the first product and the second in-band relation sub-feature is convolved along the time axis in the channel dimension to obtain an inter-band relation feature for indicating the inter-band relation. Specifically, the first product and the second in-band relation sub-feature of each frame are added element by element along the time axis to obtain a fused feature, the fused feature is channel mapped, and a convolution operation is performed on the channel dimension through a second convolution to obtain the second mapping feature. By adding the second in-band relation sub-feature split along the channel dimension and the first product after element-wise multiplication, the speech signal information in the two feature maps of the second in-band relation sub-feature and the first product can be balanced, ensuring that the fused feature contains both the fused information of the first product and the original information of the second in-band relation sub-feature, so that the speech enhancement network can learn more feature combination methods, thereby enhancing the generalization ability of the speech enhancement network. In addition, the addition operation can effectively alleviate the gradient vanishing problem that may occur during the training process, and effectively improve the training stability of the speech enhancement network.
[0136] Reference Figure 6 , Figure 6 The embodiments of the present disclosure provide Figure 3 An optional structural diagram of the mid-band mixing module, which transforms the intra-band relationship features of the transposed Input to the multi-layer perception layer of inter-band modeling, first perform channel mapping through a one-dimensional convolution, and transform the transposed intra-band relationship features The number of channels is doubled, and the in-band relationship characteristics of the increased number of channels Perform normalization to obtain intermediate features . The intermediate features Input to the gated unit module, which is used to capture intermediate features In the gated unit module, the intermediate features are first segmented along the channel dimension through a segmentation head. Split into first intra-band relation sub-features and the second intra-band relation sub-feature , the first intra-band relation sub-feature The input is sent to the normalization layer for normalization, and the normalized first intra-band relation sub-feature is transformed along the time axis. Perform frequency band mapping to obtain the first convolution feature , the first convolution feature The second intra-band relationship sub-feature Perform element-wise multiplication to get the first product , and then multiply the first product The second intra-band relationship sub-feature Add together to get the fusion feature Finally, the fusion features are constructed through a one-dimensional convolution Perform channel mapping, change the feature dimension of the fusion feature, and obtain the inter-band relationship feature used to indicate the inter-band relationship The multi-layer perception layer performs inter-band relationship modeling along the sub-band dimension through inter-band modeling, which helps to understand the interaction and mutual influence between different sub-bands. In the process of enhancing the speech signal, it can accurately eliminate noise in each sub-band, effectively improving the quality of the enhanced speech signal.
[0137] In a possible implementation, in the process of performing compression on the second mapping feature to obtain the output of the first processing block, the second mapping feature may be normalized, the normalized second mapping feature is convolved in parallel to obtain the compressed third convolution feature and the fourth convolution feature, the fourth convolution feature is nonlinearly activated to obtain the attention weight, and the third convolution feature is weighted based on the attention weight to obtain the output of the first processing block. The compression operation is performed in the sub-band compression / decompression module of the processing block in the encoder, and the decompression operation is performed in the sub-band compression / decompression module of the processing block in the decoder.
[0138] Specifically, the second mapping feature is normalized, and the normalized second mapping feature is convolved in parallel to obtain the compressed third convolution feature and the fourth convolution feature, wherein the convolution process passes through two convolution layers, the first convolution layer can be a standard convolution, and the second convolution layer can be a transposed convolution. The fourth convolution feature is input into a nonlinear activation function for nonlinear activation to obtain the attention weight corresponding to the second mapping feature, and the attention weight is weighted with the third convolution feature to obtain the output of the first processing block, wherein the nonlinear activation function can be a Sigmoid function. Reference Figure 7 , Figure 7 The embodiments of the present disclosure provide Figure 3 An optional schematic diagram of a neutron band compression / decompression module for the second mapping feature Perform normalization operation and normalize the second mapping feature The first convolutional layer and the second convolutional layer are input respectively for parallel convolution. After the convolution operation of the first convolutional layer and the second convolutional layer, the fourth convolution feature is obtained. , the fourth convolution feature As the input of the nonlinear activation function, the second mapping feature is obtained The corresponding attention weight W Then, the lower path undergoes convolution operations of the first convolution layer and the second convolution layer to obtain the third convolution feature , based on the attention weight W For the third convolution feature Weighted, that is, the attention weight W And the third convolution feature Perform element-wise multiplication to get the output of the first processing block By performing band compression in the encoder, it is possible to reduce the interference of redundant information while retaining important spectral information, so that the processing block can adaptively emphasize important frequency band features. In addition, normalizing the second mapping feature before band compression can ensure the stability of feature distribution, thereby improving the stability of speech enhancement network enhanced speech signal.
[0139] Step S203: Decode the encoding result of the encoder based on the decoder to obtain target signal characteristics.
[0140] Among them, during the decoding process of the decoder, the processing block of the decoder is used to perform frame mixing, band mixing and sub-band decompression operations; the target signal feature is a feature corresponding to the target speech signal in the mixed speech signal, which is a feature used for speech enhancement.
[0141] In a possible implementation, in the process of decoding the encoding result based on the decoder to obtain the target signal feature, the decoder may be called to decode the encoding result based on the encoder, the decoding result may be segmented along the frequency dimension to obtain multiple decoding sub-features, the channel dimension of each decoding sub-feature may be restored to be consistent with the channel dimension of the speech signal to be processed, and the multiple restored decoding sub-features may be merged along the frequency dimension to obtain the target signal feature. The decoding sub-features are obtained by segmentation based on sub-bands, and the number thereof is the same as the number of sub-bands. For example, the feature dimension of the compression feature is When Q is the number of sub-bands after frequency compression, the number of decoded sub-features is also Q.
[0142] Specifically, the decoder is called to decode based on the encoding result of the encoder, and the decoding result is divided along the frequency dimension. The frequency dimension represents the number of sub-bands after frequency compression, and the decoding result is divided into multiple decoding sub-features according to the number of sub-bands. Then, the channel dimension of each decoding sub-feature is restored to be consistent with the channel dimension of the speech signal to be processed through a convolution operation or a full connection operation, and the multiple restored decoding sub-features are merged along the frequency dimension to obtain the target signal feature. Alternatively, after obtaining multiple decoding sub-features, multiple decoding sub-features can also be merged along the frequency dimension, and then the channel dimension of the merged decoding feature is restored to be consistent with the channel dimension of the speech signal to be processed to obtain the target signal feature. By restoring the channel dimension of each decoding sub-feature to be consistent with the channel dimension of the speech signal to be processed, it is convenient for subsequent feature fusion processes such as feature addition or feature multiplication, and the smooth progress of the feature fusion process is guaranteed.
[0143] Reference Figure 8 , Figure 8 An optional schematic diagram of performing frequency band merging provided in an embodiment of the present disclosure, assuming that the feature dimension corresponding to the decoding result U is , the feature dimension corresponding to the speech signal to be processed is , based on the frequency dimension Q, the decoding result U is segmented to obtain Q decoding sub-features ,in, . The decoded sub-feature The channel dimension of is restored to be consistent with the channel dimension of the speech signal to be processed, and the decoding sub-feature is obtained , all decoded sub-features Merge along the frequency dimension to obtain the target feature signal The above process can be expressed by the following formula:
[0144] ,
[0145] ,
[0146] ,
[0147] in, Indicates sub-band partitioning.
[0148] In a possible implementation, in the process of restoring the channel dimension of each decoding sub-feature to be consistent with the initial signal feature, each decoding sub-feature may be normalized, and the channel dimension of the normalized decoding sub-feature is restored to be consistent with the channel dimension of the speech signal to be processed by a multi-layer perceptron. Specifically, for any decoding sub-feature, the decoding sub-feature is normalized, and the channel dimension of the normalized decoding sub-feature is restored to be consistent with the channel dimension of the speech signal to be processed by adjusting the number of output neurons in the multi-layer perceptron, and the sub-time-frequency mask corresponding to the decoding sub-feature is predicted by the multi-layer perceptron, wherein the sub-time-frequency mask may be a mask with complex values applied in the time-frequency domain. Then, the sub-time-frequency masks corresponding to each decoding sub-feature are merged along the time axis to obtain the time-frequency mask, i.e., the target signal feature. The multi-layer perceptron can learn the complex linear relationship between each decoding sub-feature and has strong adaptability to different types of noise and interference. The output sub-time-frequency mask can weight the signal in the time-frequency domain to suppress noise or eliminate part of the echo, thereby retaining the enhanced speech signal part (i.e., the target signal feature), effectively improving the clarity and quality of the enhanced speech signal.
[0149] Step S204: restoring the enhanced speech signal from the speech signal to be processed based on the target signal feature.
[0150] In the above steps, by acquiring the speech signal to be processed, the feature extraction of the speech signal to be processed is performed to obtain the initial signal feature, and the initial signal feature is divided into frequency bands to obtain the segmentation mapping feature. The segmentation mapping feature is input into the speech enhancement network to perform operations such as noise removal. The speech enhancement network includes an encoder and a decoder connected to each other. The encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence. In the processing block of the encoder, the input is subjected to frame mixed mapping to obtain the first mapping feature, the first mapping feature is subjected to band mixed mapping to obtain the second mapping feature, the second mapping feature is subjected to band compression to obtain the output of the processing block, and the output of the processing block is input into the next processing block until the encoding result of the encoder is obtained. The encoding result is input into the decoder for decoding, and the decoding result is subjected to frequency band merging to obtain the target signal feature. Based on the target signal feature, the enhanced speech signal is restored from the speech signal to be processed to obtain the enhanced speech signal, wherein the enhanced speech signal is the target speech signal after the echo signal is removed from the mixed speech signal and the interference such as noise is suppressed. The restoration process can be understood as the time-frequency domain inverse conversion of the target signal feature, and the target signal feature is converted from the frequency domain to the time domain. The time-frequency domain inverse conversion can be a short-time Fourier inverse transform.
[0151] In a possible implementation, in the process of restoring the enhanced speech signal from the speech signal to be processed based on the target signal feature, specifically, a second product between the target signal feature and the speech signal to be processed is determined, and the second product is filtered to restore the enhanced speech signal from the speech signal to be processed. The second product is the product obtained by element-wise multiplication of the target speech signal and the speech signal to be processed.
[0152] Specifically, after obtaining the target signal feature, some residual noise may still exist in the target signal feature. In order to further suppress the participating noise, the target signal feature can be post-processed. First, the target signal feature and the speech signal to be processed are element-wise multiplied along the channel dimension to obtain the second product. The above element-wise multiplication process can be expressed by the following formula:
[0153] ,
[0154] in, C represents the target signal characteristics and the number of channels of the speech signal to be processed (in the above steps, the number of channels of the target signal characteristics has been adjusted to be consistent with the number of channels of the speech signal to be processed), i Indicates i channels, I represents the speech signal to be processed, G Represents the target signal feature. This formula indicates that the target signal feature and the speech signal to be processed are multiplied element-wise on a channel-by-channel basis, and the multiplication results of each channel are added to obtain the second product .
[0155] Then, the second product and the speech signal to be processed are input into multiple gated recurrent units to further learn complex sequence features and long-term dependencies, and intermediate features with long-term dependencies are obtained. The intermediate features are then linearly transformed through a set of linear layers to restore the enhanced speech signal from the speech signal to be processed. The second product is obtained by element-wise multiplication of the target signal feature and the speech signal to be processed, so that the post-processing process can pay more attention to the feature information in the speech signal to be processed that is more relevant to the target signal feature. On this basis, the second product is deeply filtered through gated recurrent units and linear layers, which can reduce the number of parameters while maintaining the processing capacity of the post-processing process, which helps to restore the enhanced speech signal from the speech signal to be processed, effectively improves the accuracy of speech separation, and thus ensures the quality and reliability of the enhanced speech signal.
[0156] In a possible implementation, before calling the encoder to encode based on the initial signal features, the encoder and decoder need to be trained, specifically, a sample speech signal is obtained, an enhanced sample speech signal is restored from the sample speech signal based on the encoder and decoder, a label speech signal and a speech activity label corresponding to the sample speech signal are obtained, a first loss is determined according to the difference between the sample speech signal and the label speech signal, the sample speech signal is adjusted according to the speech activity label to obtain a second loss, the first loss and the second loss are weighted and summed to obtain a target loss, and the parameters of the encoder and decoder are adjusted based on the target loss. The sample speech signal is a speech signal used for training, including a sample mixed speech signal, a sample reference speech signal, a sample echo signal, and a sample error speech signal; the label speech signal is used as a reference speech signal for comparison with the predicted speech signal; the speech activity label is used to indicate whether the speech of the target object exists in the sample speech signal or not; when the speech activity label is 1, it indicates that the speech of the target object exists in the sample semantic signal; when the speech activity label is 0, it indicates that the speech of the target object does not exist in the sample semantic signal; the speech of the target object is a sample mixed speech signal, and the sample mixed speech signal can be a near-end signal.
[0157] Specifically, a sample speech signal is obtained, and an encoder and a decoder are trained based on the speech signal, so that the encoder and the decoder can restore the enhanced speech signal from the sample speech signal. During the training process, the label real part, label imaginary part and label amplitude corresponding to the label speech signal are first obtained, and then the sample real part, sample imaginary part and sample amplitude of the predicted spectrum obtained during the training process are obtained. The first loss is determined based on the average absolute error between the label real part and the sample real part, the average absolute error between the label imaginary part and the sample imaginary part, and the average absolute error between the label amplitude and the sample amplitude. According to the above description, the first loss It can be expressed by the following formula:
[0158] ,
[0159] in, represents the real part of the sample, represents the imaginary part of the sample, represents the sample amplitude, represents the real part of the label, represents the imaginary part of the label, Represents the label amplitude. The loss function calculated by the mean absolute error can intuitively measure the difference between the content in the predicted spectrum and the actual label, and understand the prediction accuracy and error distribution of the encoder and decoder.
[0160] Next, the speech activity label is obtained, and the norm is calculated based on the product of the speech activity label and the sample amplitude to obtain the first norm. Then, the logarithm of the first norm is taken to obtain the second loss. According to the above description, the second loss It can be expressed by the following formula:
[0161] ,
[0162] in, is the voice active label, To predict the spectrum, it is the near-end speech signal without noise and echo. is a hyperparameter used to prevent excessive suppression of the target speech signal in the sample mixed speech signal. When it is 0, the calculated first norm value is large, and the penalty intensity increases with the increase of the first norm value. A value of 0 indicates that there is no target object's speech in the sample speech signal, which can be understood as the near-end speech is missing. In this case, the second loss The value of will gradually tend to 0, and as the second loss The value of decreases and its rate of change gradually increases, based on which the penalty for near-end speech loss is increased. In order to deal with the near-end speech loss, the encoder and decoder will suppress the output as much as possible, that is, avoid unnecessary spectrum prediction, ensuring that the encoder and decoder can make a more accurate and reasonable response when facing near-end speech loss, thereby improving the overall speech processing effect.
[0163] Then, configure the weight for the second loss, perform the weighted sum of the first loss and the second loss to get the target loss, the target loss L It can be expressed by the following formula:
[0164] ,
[0165] in, β It is a preset weight (hyperparameter) used to enhance the ability of preserving near-end speech.
[0166] In addition, in order to effectively suppress the echo signal, the third loss is calculated based on the frame rate, frequency, amplitude and phase of the sample speech signal. Used to evaluate the echo suppression loss, it can be expressed by the following formula,
[0167] ,
[0168] in, is the amplitude loss, which can be calculated based on the difference between the sample amplitude and the label amplitude; is the echo weighting coefficient, which can adjust the weights of different frequency units according to the power ratio of the echo signal; is the phase loss, which can be calculated from the difference between the sample phase and the label phase.
[0169] After obtaining the third loss, the target loss can be obtained by weighted summing the first loss, the second loss and the third loss. L It can be expressed by the following formula:
[0170] ,
[0171] in, β Set to balance the capabilities between echo suppression and near-end speech preservation.
[0172] Finally, the parameters of the encoder and decoder are adjusted based on the target loss so that the encoder and decoder can effectively suppress noise and echo, output the target signal features without noise and echo, and restore the clear enhanced speech signal based on the target signal features. The joint loss function can evaluate the encoder and decoder from aspects such as predicted spectrum, enhanced near-end speech preservation, and echo suppression loss, thereby providing targeted optimization directions for the encoder and decoder, which helps the encoder and decoder improve their ability to suppress noise.
[0173] In a possible implementation, a target loss is determined according to the difference and logarithmic variance between the enhanced result of the sample speech signal and the sample label signal. In the process of training the speech enhancement network based on the target loss, the target loss of the first training stage can be determined according to the difference and logarithmic variance between the enhanced result of the sample speech signal and the sample label signal, and the target loss of the second training stage can be determined according to the enhanced result of the sample speech signal and the difference and logarithmic variance between the sample label signal. The speech enhancement network is trained in sequence based on the target loss of the first training stage and the target loss of the second training stage. The target loss of the first training stage is used to constrain the accuracy of the logarithmic variance, and the target loss of the second training stage is used to assign corresponding weights to the time-frequency units according to the uncertainty of the time-frequency units when determining the difference between the enhanced result of the sample speech signal and the sample label signal; the sample label signal includes a sample label real signal, a sample label imaginary signal and a sample label pure signal, the sample label real signal is a sample label corresponding to the sample real signal, the sample label imaginary signal is a sample label corresponding to the sample imaginary signal, and the sample label pure signal is a sample label corresponding to the sample speech signal.
[0174] Specifically, in the first training stage, the sample enhanced real signal and the sample enhanced imaginary signal are input into the logarithmic variance estimation module to obtain the real variance, imaginary variance and amplitude variance. In the second training stage, all parameters in the logarithmic variance estimation module are frozen, and the real variance, imaginary variance and amplitude variance are used to prioritize the time-frequency units with high uncertainty. Then, the real variance, imaginary variance and amplitude variance are logarithmically operated to obtain the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum and the logarithmic variance of the amplitude spectrum. For example, the calculated real variance is θ=e , according to the formula Perform logarithmic operations to obtain the logarithmic variance of the real part of the complex spectrum . Then, the target loss of the first training stage and the target loss of the second training stage are determined according to the difference between the enhancement result of the sample speech signal and the sample label signal, as well as the logarithmic variance, and the speech enhancement network is trained in sequence based on the target loss of the first training stage and the target loss of the second training stage. By training the speech enhancement network in stages, different training goals and optimization strategies can be set for training at different stages, and the speech enhancement network can be gradually optimized, which helps the speech enhancement network converge faster during training.
[0175] In a possible implementation, the logarithmic variance includes the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, and the logarithmic variance of the amplitude spectrum. In the process of determining the target loss of the first training stage according to the difference and logarithmic variance between the enhanced result of the sample speech signal and the sample label signal, the difference of the real part of the complex spectrum, the difference of the imaginary part of the complex spectrum, and the difference of the amplitude spectrum between the enhanced result of the sample speech signal and the sample label signal can be specifically determined, the difference of the real part of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the real part of the complex spectrum as an exponent, the difference of the imaginary part of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the imaginary part of the complex spectrum as an exponent, and the first loss is determined according to the sum of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, the weighted difference of the real part of the complex spectrum, and the weighted difference of the imaginary part of the complex spectrum. Among them, the first loss is used to predict the uncertainty of the logarithmic variance of the real part of the complex spectrum and the logarithmic variance of the imaginary part of the complex spectrum, and the second loss is used to predict the uncertainty of the logarithmic variance of the amplitude spectrum.
[0176] Specifically, the real part difference of the complex spectrum is determined according to the sample enhanced real part signal and the sample label real part signal, and the imaginary part difference of the complex spectrum is determined according to the sample enhanced imaginary part signal and the sample label imaginary part signal, and the norm values of the real part difference of the complex spectrum and the imaginary part difference of the complex spectrum are calculated respectively. Next, an exponential operation is performed on the logarithmic variance of the real part of the complex spectrum to obtain a first exponential result, and an exponential operation is performed on the logarithmic variance of the imaginary part of the complex spectrum to obtain a second exponential result. According to the first exponential result, the norm value of the real part difference of the complex spectrum is weighted to obtain a first weighted result, and according to the second exponential result, the norm value of the imaginary part difference of the complex spectrum is weighted to obtain a second weighted result. Finally, the first loss is determined according to the sum of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, the first weighted result, and the second weighted result. Determine the first loss The process can be expressed by the following formula:
[0177] ,
[0178] in, represents the logarithmic variance of the real part of the complex spectrum, represents the logarithmic variance of the imaginary part of the complex spectrum, represents the sample enhanced real part signal, represents the sample enhanced imaginary signal, represents the real part signal of the sample label, represents the imaginary signal of the sample label, represents the natural exponential function.
[0179] Next, the amplitude spectrum difference is weighted according to a natural exponential function with the amplitude spectrum logarithmic variance as an exponent, and the second loss is determined according to the sum of the amplitude spectrum logarithmic variance and the weighted amplitude spectrum difference. Specifically, the amplitude spectrum difference is determined according to the sample enhanced speech signal and the sample label pure signal, and the norm value of the amplitude spectrum difference is calculated. Then, an exponential operation is performed on the amplitude spectrum logarithmic variance to obtain a third exponential result, and the norm value of the amplitude spectrum difference is weighted according to the third exponential result to obtain a third weighted result, and the second loss is determined according to the third weighted result and the sum of the amplitude spectrum logarithmic variance. Determine the second loss The process can be expressed by the following formula:
[0180] ,
[0181] in, represents the logarithmic variance of the amplitude spectrum, represents the sample enhanced speech signal, Indicates the sample label pure signal.
[0182] Next, the first loss and the second loss are weighted and summed to obtain the target loss of the first training stage. Specifically, a first weight factor is configured for the second loss, the second loss is weighted based on the first weight factor, and the target loss of the first training stage is obtained according to the sum of the first loss and the weighted second loss. Target loss of the first training stage It can be expressed by the following formula, where α is a fixed weight coefficient used to control the contribution of the second loss during the training process.
[0183]
[0184] By performing weighted summation of the real part difference of the complex spectrum, the imaginary part difference of the complex spectrum and the amplitude spectrum difference based on the exponential values of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum and the logarithmic variance of the amplitude spectrum, the speech enhancement network can pay attention to the variance and characteristic distribution of each sub-band during the training process, so that the speech enhancement network can effectively remove interference such as noise while minimizing the initial signal characteristics, thereby improving the speech enhancement network's ability to enhance speech signals.
[0185] In a possible implementation, in the process of weighting the first loss and the second loss to obtain the target loss of the first training stage, the signal-to-noise ratio loss can be determined based on the enhancement result of the sample speech signal and the sample label signal, the update rate loss can be determined based on the difference between the update rate of the binary gate and the preset update rate threshold, the first loss and the second loss can be weighted to obtain the first uncertain loss, and the first uncertain loss, the signal-to-noise ratio loss and the update rate loss can be weighted to obtain the target loss of the first training stage. Among them, the signal-to-noise ratio loss is used to evaluate the influence of noise in the sample enhanced speech signal, and the update rate loss is used to train the gated recurrent unit that performs the frame mixing mapping process so that the gated recurrent unit has a higher update rate. The update rate is the frequency of updating the hidden state features in the gated recurrent unit, and the update rate threshold is a preset parameter used to measure and constrain the update frequency of the hidden state features.
[0186] Specifically, the enhancement result of the sample speech signal also includes a sample enhanced speech waveform, and the sample label signal also includes a sample label pure speech waveform. A decentralized sample waveform is obtained based on the sample enhanced speech waveform, and a decentralized sample label waveform is obtained based on the sample label pure speech waveform. A norm operation is performed on the decentralized sample label waveform, and a norm operation is performed on the difference between the decentralized sample waveform and the decentralized sample label waveform. The signal-to-noise ratio loss is obtained by taking the logarithm of the ratio of the norm value of the decentralized sample label waveform to the norm value of the difference between the decentralized sample waveform and the decentralized sample label waveform. Signal-to-noise ratio loss It can be expressed by the following formula:
[0187] ,
[0188] in, represents the decentralized sample enhanced waveform, Represents the pure waveform of decentralized samples, Represents a logarithmic function. Centered sample enhanced waveform It can be expressed by the formula Calculated, Indicates sample enhanced waveform and decentralized sample pure waveform It can be expressed by the formula Calculated, Represents a sample pure waveform. It is the average function, which is used to calculate the average value of the sample enhanced waveform and the sample pure waveform.
[0189] Next, the update rate loss is determined based on the difference between the update rate of the binary gate and the preset update rate threshold. It can be expressed by the following formula, where g represents the update rate, u Indicates the update rate threshold.
[0190]
[0191] Next, the weighted sum of the first loss and the second loss is used to obtain the first uncertainty loss. It can be expressed by the following formula:
[0192] ,
[0193] A second weight factor is configured for the signal-to-noise ratio loss, a third weight factor is configured for the update rate loss, and a first uncertainty loss is calculated based on the second weight factor and the third weight factor. , signal-to-noise ratio loss and update rate loss Perform weighted summation to obtain the target loss of the first training stage, the target loss of the first training stage It can be expressed by the following formula:
[0194] ,
[0195] in, β is the second weight factor, the second weight factor β The value of can be 0.02, γ is the third weight factor, the third weight factor γ The value of can be 0.2. The first stage of training the speech enhancement network through the first uncertainty loss, signal-to-noise ratio loss and update rate loss enables the logarithmic variance estimation module to more accurately predict the uncertainty in the target signal characteristics, thereby helping to improve the quality of the enhanced speech signal output by the speech enhancement network.
[0196] In one possible implementation, in the process of determining the target loss of the second training stage based on the difference and logarithmic variance between the enhanced results of the sample speech signal and the sample label signal, specifically, the real part difference of the complex spectrum is weighted according to the logarithmic variance of the real part of the complex spectrum, the imaginary part difference of the complex spectrum is weighted according to the logarithmic variance of the imaginary part of the complex spectrum, the third loss is determined according to the sum of the weighted real part difference of the complex spectrum and the weighted imaginary part difference of the complex spectrum, the amplitude spectrum difference is weighted according to the logarithmic variance of the amplitude spectrum, and the weighted amplitude spectrum difference is determined as the fourth loss.
[0197] Specifically, the norm values of the real part difference of the complex spectrum, the imaginary part difference of the complex spectrum, and the amplitude spectrum difference are calculated respectively, and the norm value of the real part difference of the complex spectrum is weighted according to the logarithmic variance of the real part of the complex spectrum to obtain a fourth weighted result, and the norm value of the imaginary part difference of the complex spectrum is weighted according to the logarithmic variance of the imaginary part of the complex spectrum to obtain a fifth weighted result, and the norm value of the imaginary part difference of the amplitude spectrum is weighted according to the logarithmic variance of the amplitude spectrum to obtain a sixth weighted result. The third loss is determined according to the sum of the fourth weighted result and the fifth weighted result, and the fourth loss is determined according to the sixth weighted result. Third loss 4. Loss It can be expressed by the following formula.
[0198]
[0199]
[0200] Next, a first weight factor is configured for the fourth loss, and the third loss and the fourth loss are weighted and summed based on the first weight factor to obtain the target loss of the second training stage. It can be expressed by the following formula.
[0201]
[0202] By respectively performing complex spectrum real part difference, complex spectrum imaginary part difference and amplitude spectrum difference based on the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum and the logarithmic variance of the amplitude spectrum, it is possible to assign larger weights to time-frequency units under high uncertainty, so that the speech enhancement network can better focus on highly uncertain content, and thus can better focus on subtle features, thereby improving the enhanced quality of speech enhanced speech.
[0203] In a possible implementation, in the process of weighting the third loss and the fourth loss to obtain the target loss of the second training stage, the signal-to-noise ratio loss can be determined according to the enhancement result of the sample speech signal and the sample label signal, the update rate loss can be determined according to the difference between the update rate of the binary gate and the preset update rate threshold, the third loss and the fourth loss are weighted to obtain the second uncertainty loss, and the second uncertainty loss, the signal-to-noise ratio loss and the update rate loss are weighted to obtain the target loss of the second training stage. Specifically, the third loss and the fourth loss are weighted to obtain the second uncertainty loss, and the second uncertainty loss is It can be expressed by the following formula:
[0204] ,
[0205] A second weight factor is configured for the signal-to-noise ratio loss, a third weight factor is configured for the update rate loss, and a second uncertainty loss is calculated based on the second weight factor and the third weight factor. , signal-to-noise ratio loss and update rate loss Perform weighted summation to obtain the target loss of the second training stage, the target loss of the second training stage It can be expressed by the following formula.
[0206]
[0207] In addition, the processing block of the speech enhancement network can also be used for causal time sampling, wherein causal time sampling is performed under the premise of maintaining data causality, that is, the order of data points and time relativity remain unchanged during the sampling process, and causal time sampling can include causal time downsampling and causal time upsampling. Different processing blocks have different time compression ratios when performing causal time sampling, and the time compression ratio of causal time sampling in any processing block is the same, so the frame rate of the input feature and the output feature can be kept unchanged after downsampling and upsampling in any processing block; the target signal feature is the feature corresponding to the target speech signal in the mixed speech signal, which is the feature used for speech enhancement.
[0208] In a possible implementation, in the process of calling the encoder to encode based on the initial signal feature, the initial signal feature may be input into the first processing block in the encoder, the initial signal feature may be downsampled to obtain the downsampled feature, the downsampled feature may be mapped to obtain the intra-band relationship feature for indicating the intra-band relationship, the first mapping feature may be mapped to obtain the inter-band relationship feature for indicating the inter-band relationship, the inter-band relationship feature may be upsampled to obtain the output of the first processing block and then input into the next processing block until the output of the last processing block is obtained as the encoding result of the encoder. The downsampling operation is used to reduce the frame rate of the initial signal feature, and the downsampling operation may be a time causal downsampling operation; the upsampling operation is used to increase the frame rate of the inter-band relationship feature, and the upsampling operation may be a causal time upsampling operation.
[0209] Specifically, the initial signal feature is input to the first processing block in the encoder, in which the initial signal feature is first input to the downsampling layer, and the initial signal feature is downsampled in the downsampling layer, wherein the downsampling layer includes a first causal convolution, which can be composed of a one-dimensional causal convolution, and mainly acts on the frame rate of the initial signal feature. When performing the downsampling operation, the convolution kernel size and step size of the first causal convolution used are the same as the value of the time compression ratio, and the frame rate of the initial signal feature is reduced according to the time compression ratio to obtain a downsampled feature with a low frame rate. For example, the feature dimension of the initial signal feature is , Q is the number of sub-bands after compression, then the feature dimension of the down-sampled feature can be , λ is the time compression ratio. Then, mapping is performed based on the downsampled features to obtain the intra-band relationship features for indicating the intra-band relationship, and then the intra-band relationship features are transposed. The transposition operation may be performed by using the transpose function or the reshape function. The first mapping feature is mapped to obtain the inter-band relationship features for indicating the inter-band relationship. The inter-band relationship features are input into the upsampling layer, and the inter-band relationship features are upsampled in the upsampling layer, wherein the upsampling layer includes a second causal convolution, and the second causal convolution may be a one-dimensional point convolution, which mainly acts on the frame rate of the inter-band relationship features. When performing the upsampling operation, the convolution kernel size and step size of the second causal convolution used are the same as the value of the time compression ratio. The inter-band relationship features are first interpolated along the time dimension, and the frame rate of the interpolated inter-band relationship features is increased in the second causal convolution according to the time compression ratio to obtain the output of the first processing block with the same frame rate as the initial signal feature. For example, the feature dimension of the initial signal feature is , when the feature dimension of the inter-relation feature is When , the output features of the first output block can be The output of the first processing block is input to the next processing block to obtain the output of the next processing block, and the process is iterated in each cascaded processing block until the output of the last processing block is obtained, and the output of the last processing block is used as the encoding result of the encoder.
[0210] It should be noted that the time compression ratio configured for each processing block in the encoder is different. Since the encoder and decoder are constructed symmetrically, the time compression ratio of each processing block in the decoder is the same as the time compression ratio of the processing block symmetrical to the position in the encoder. For example, there are 6 processing blocks cascaded in sequence in the encoder, and the time compression ratios of these 6 processing blocks cascaded in sequence can be 1, 2, 4, 8, 16, 32 respectively. Then the time compression ratios of the 6 processing blocks cascaded in sequence in the decoder are 32, 16, 8, 4, 2, 1. Since the time compression ratio configured in each processing block is different, upsampling and downsampling based on the time compression ratio can effectively control the computational complexity of the processing block, thereby improving the training speed and inference speed of the encoder and decoder, and then improving the real-time performance of speech enhancement.
[0211] Reference Fig. 9 , Fig. 9 Another optional structural diagram of a processing block in a speech enhancement network provided in an embodiment of the present disclosure, assuming that the processing block is the first processing block, and the input signal input to the processing block is the initial signal feature . The initial signal characteristics Input to the downsampling layer for downsampling, through the time compression ratio Reduce initial signal characteristics The frame rate is , and the down-sampling features of the low frame rate are obtained , and then downsample the features Input into the in-band relationship model for feature mapping to obtain the in-band relationship features . Through the Reshape function, the intra-band relationship features The frame dimension (T) and frequency dimension (Q) are transformed to realize feature transposition and obtain the transposed in-band relationship feature . Transpose the in-band relationship features Input to the multi-layer perception layer for inter-band modeling to perform the second feature mapping to obtain the inter-band relationship features Finally, the inter-band relationship features Input to the upsampling layer, through the time compression ratio The relationship characteristics between The frame rate is restored to the original signal characteristics The frame rate is consistent with that of the processing block, and the output The above process can be expressed by the following formula:
[0212] ,
[0213] ,
[0214] ,
[0215] ,
[0216] in, represents downsampling, FC represents the fully connected layer, GRU represents a gated recurrent unit, Norm represents the normalization layer, Represents the multi-layer perception layer of inter-band modeling, Indicates upsampling.
[0217] In a possible implementation, in the process of mapping based on downsampled features to obtain in-band relationship features for indicating in-band relationships, the downsampled features may be normalized, the normalized downsampled features may be mapped through a gated cyclic unit to obtain an N-1th mapping feature, the N-1th mapping feature may be fully connected to obtain an Nth mapping feature, and the Nth mapping feature may be summed with the upsampled feature to obtain an in-band relationship feature for indicating in-band relationships. The N-1th mapping feature is a feature mapped by a gated cyclic unit, and the Nth mapping feature is a feature fully connected.
[0218] Specifically, the downsampled features are normalized so that they have the same data distribution. The normalized downsampled features are input into the gated cyclic unit for feature mapping. The gated cyclic unit is used to extract the sequence information of the speech signal or capture the time-dependent features of the speech signal in the downsampled features. These information and features reflect the characteristics of the speech signal changing over time. A reset gate and an update gate are configured in the gated cyclic unit. The feature information with weak correlation with the speech signal in the downsampled features is forgotten by the reset gate, and the important feature information related to the speech signal is retained by the update gate to obtain the N-1th mapping feature with time dependence. At this time, the feature dimension of the N-1th mapping feature may be inconsistent with the feature dimension of the downsampled feature. Then, the N-1th mapping feature is input into the fully connected layer for further feature extraction to learn complex feature combinations. When the feature dimension of the N-1th mapping feature is inconsistent with the feature dimension of the downsampled feature, the feature dimension of the N-1th mapping feature is feature mapped by adjusting the number of output neurons of the fully connected layer. When performing feature mapping, the N-1th mapping feature is mapped to a new feature space, and the feature dimension of the new feature space is consistent with the feature dimension of the downsampled feature, so as to obtain the Nth mapping feature that is consistent with the feature dimension of the downsampled feature. Finally, the Nth mapping feature is added element by element with the upsampled feature to fuse features of different scales and levels, and obtain the multi-scale fused intra-band relationship feature. For example, Fig. 9 As shown in the intra-band relationship model in, the downsampled features The input is sent to the normalization layer for normalization to obtain the normalized down-sampled features, the normalized down-sampled features are sent to the gated recurrent unit for feature mapping to obtain the N-1th mapping features, and the N-1th mapping features are sent to the fully connected layer for feature mapping again to obtain the Nth mapping features, and finally the Nth mapping features are combined with the down-sampled features. Add them together to get the intra-band relationship features.
[0219] By adopting the gated recurrent unit, the noise in the downsampled features can be filtered out while capturing the timing information and time dependency in the downsampled features, effectively improving the robustness of the encoder and decoder. Furthermore, by combining normalization, gated recurrent units, fully connected layers, and feature fusion to map and combine the downsampled features, a richer feature representation can be obtained, which helps improve the accuracy of the decoder in predicting speech signals.
[0220] In addition, when the feature dimension of the N-1th mapping feature is much higher or much lower than the feature dimension of the downsampled feature, feature mapping can be performed through a convolution operation to avoid excessive parameters or information loss that may be caused by feature mapping only through a fully connected layer.
[0221] In a possible implementation, the output of each processing block is input to the remaining processing blocks, and the input of the nth processing block is obtained by normalizing the outputs of the first to n-1th processing blocks and summing them, where n is a positive integer and n>2. Specifically, the output of each processing block is connected to the input after the processing block by cross-scale jump connection, that is, starting from the second processing block, the input of the nth processing block is related to the outputs of all the first n-1 processing blocks, and the outputs of the first n-1 processing blocks are normalized and summed to obtain the input of the nth processing block. For example, the encoder includes 3 processing blocks (processing block 1, processing block 2, processing block 3), and the decoder includes 3 processing blocks (processing block 4, processing block 5, processing block 6). The initial signal feature is input into processing block 1 of the encoder to obtain the first intermediate feature of processing block 1, and the first intermediate feature is input into processing block 2 to obtain the second intermediate feature. The second intermediate feature and the first intermediate feature are normalized respectively, and the normalized second intermediate feature and the first intermediate feature are added to obtain the first fusion feature. The first fusion feature is input into processing block 3 to obtain the third intermediate feature of processing block 3. The third intermediate feature, the first intermediate feature, and the first fusion feature are normalized respectively and added to obtain the second fusion feature. The second fusion feature is input into processing block 4 to obtain the fourth intermediate feature. The fourth intermediate feature, the first fusion feature, the second fusion feature, and the first intermediate feature are normalized respectively and added to obtain the third fusion feature. The third fusion feature is input into processing block 5 to obtain the fifth intermediate feature. The fifth intermediate feature, the first fusion feature, the second fusion feature, the third fusion feature, and the first intermediate feature are normalized respectively and added to obtain the fourth fusion feature. The fourth fusion feature is input into processing block 6. The output of processing block 6 is the output of the decoder, and subsequent speech signal enhancement processing is performed based on the output. By connecting multiple processing blocks through cross-scale jump connections, the encoder and decoder can adapt to input data of different scales while calibrating the data input by the previous processing block, which promotes the encoder and decoder to capture and integrate feature information from different scales. In addition, cross-scale jump connections can extract richer feature representations without increasing additional computational consumption, improve the robustness and generalization ability of the encoder and decoder, and thus improve the effect of enhancing speech signals.
[0222] In a possible implementation, in the process of calling the encoder to encode based on the initial signal feature, the initial signal feature can be divided into multiple signal sub-features according to multiple frequency ranges, and each signal sub-feature is frequency compressed to obtain the compressed sub-feature corresponding to each signal sub-feature, and the multiple compressed sub-features are merged along the frequency dimension to obtain the compressed feature, and the encoder is called to encode based on the compressed feature. Among them, the frequency range is divided according to the frequency unit, a frequency range can be composed of at least one frequency unit, and a frequency unit corresponds to a specific frequency sub-range; the signal sub-feature is obtained by dividing the initial signal feature according to multiple frequency ranges, that is, one frequency range corresponds to one signal sub-feature; the compressed sub-feature is obtained after the signal sub-feature is compressed by frequency.
[0223] Specifically, first, multiple frequency ranges are set, each frequency range contains different frequency units, and the initial signal features are divided into multiple signal sub-features along the frequency dimension according to the multiple frequency ranges. Each signal sub-feature is frequency compressed based on the convolution group, and a frequency units in the frequency range corresponding to the signal sub-feature are compressed into b sub-frequency bands, where a>b, to obtain compressed sub-features corresponding to each signal sub-feature. Finally, multiple compressed sub-features are merged along the frequency dimension to obtain compressed features, and the encoder is called to encode based on the compressed features. By compressing multiple frequency units into a smaller number of sub-frequency bands along the frequency dimension, a dimensionality reduction operation of the frequency dimension is achieved, which reduces the computational cost for subsequent operations based on compressed features and improves the efficiency of speech enhancement.
[0224] Reference Fig.10 , Fig.10 An optional schematic diagram of frequency band segmentation provided in an embodiment of the present disclosure, wherein the initial signal feature is segmented into P frequency ranges along the frequency dimension, and the signal sub-features corresponding to each frequency range are A set of two-dimensional convolutions with different convolution kernel sizes and the same output channels are used to perform frequency band segmentation, and the features processed by different convolution kernel sizes are concatenated to obtain a compressed three-dimensional representation. ,in, For example, the signal sub-signature A set of convolution kernels with a size of , the step length is The two-dimensional convolution is used to perform frequency band segmentation processing, and the features processed by different convolution kernel sizes are spliced to obtain signal sub-features. The corresponding compressed sub-feature Then, the compression sub-features corresponding to each signal sub-feature are merged to obtain the compression feature The above process can be expressed by the following formula:
[0225] ,
[0226] ,
[0227] ,
[0228] in, represents the frequency range-based segmentation operation, Represents a splicing operation, Represents a merge operation.
[0229] For example, assuming that the initial signal feature contains 161 frequency units, the feature dimension of the initial signal feature can be ,when P =3, the 161 frequency units are divided into 3 frequency ranges, and the initial signal features are divided into 3 signal sub-features according to the 3 frequency ranges. Among them, the first frequency range contains 20 frequency units, so the feature dimension of the first signal sub-feature corresponding to the first frequency range can be ; The second frequency range contains 60 frequency units, so the feature dimension of the second signal sub-feature corresponding to the second frequency range can be ; The third frequency range contains 81 frequency units, so the feature dimension of the third signal sub-feature corresponding to the third frequency range can be The frequency compression of these three signal sub-features is performed respectively, and the feature dimension of the first compressed sub-feature corresponding to the first signal sub-feature can be , 0, the feature dimension of the second compression sub-feature corresponding to the second signal sub-feature can be , 0, the feature dimension of the third compression sub-feature corresponding to the third signal sub-feature can be , 81. Finally, these three compressed sub-features are combined along the frequency dimension to obtain the compressed feature. The feature dimension of the compressed feature can be ,in, ,The compressed features are input into the encoder of the speech enhancement network for encoding.
[0230] In a possible implementation, in the process of performing frequency compression on each signal sub-feature to obtain compressed sub-features corresponding to each signal sub-feature, it can be specifically that for any signal sub-feature, frequency compression is performed through multiple convolutional layers to obtain multiple fifth convolutional features corresponding to the signal sub-feature, and the multiple fifth convolutional features corresponding to the same signal sub-feature are merged along the channel dimension to obtain compressed sub-features corresponding to each signal sub-feature. Among them, the multiple convolutional layers are respectively configured with different convolution kernel sizes and the same number of output channels, and the convolution kernel size is related to the frequency range corresponding to the signal sub-feature, and the convolution kernel sizes corresponding to different frequency ranges are all different.
[0231] Specifically, for any signal sub-feature, multiple convolution layers are used to perform frequency compression on the signal sub-feature, and the convolution kernel sizes of each convolution layer are different, but the step size is the same. After the signal sub-feature passes through each convolution layer, the fifth convolution feature corresponding to each convolution layer is obtained. The fifth convolution features corresponding to each convolution layer are merged along the channel dimension to obtain the compressed sub-feature corresponding to the signal sub-feature. By using a group of convolutions with different convolution kernel sizes to perform frequency compression on the same signal sub-feature, it is possible to capture features of different scales in the signal sub-feature, thereby enhancing the ability to understand and represent the signal sub-feature. On this basis, the process of merging each fifth convolution feature promotes the fusion and integration of features of different scales, which helps to form a richer and more accurate feature representation, thereby improving the ability to enhance speech signals in subsequent steps.
[0232] For example, refer to Fig.11 , Fig.11 The embodiments of the present disclosure provide Fig.10 An optional schematic diagram of frequency band segmentation of signal sub-features in , Fig.11 Each frequency unit is divided into three frequency ranges, and the initial signal characteristics are First signal sub-signal , the second signal sub-feature And the third signal sub-signal For the first signal sub-feature , three groups of convolutions can be used to convolve the first signal sub-feature For frequency compression, the convolution kernel size of convolution A is , the step length is , that is, use The convolution kernel is applied to the first signal sub-feature The frame rate dimension is based on the step size , in the frequency dimension with a step size of The convolution operation is performed in the form of The convolution kernel size of convolution B is , the step length is , that is, use The convolution kernel is applied to the first signal sub-feature The frame rate dimension is based on the step size , in the frequency dimension with a step size of The convolution operation is performed in the form of The convolution kernel size of convolution C is , the step length is , that is, use The convolution kernel is applied to the first signal sub-feature The frame rate dimension is based on the step size , in the frequency dimension with a step size of The convolution operation is performed in the form of Then, the convolution feature , convolutional features And convolutional features Splice along the channel dimension to obtain the first signal sub-feature The corresponding first compressed sub-feature The second signal sub-signature And the third signal sub-signal The frequency compression and the above characteristics of the first signal sub It is worth noting that for signal sub-features in the same frequency range, frequency compression is performed using convolution operations with the same step length but different convolution kernel sizes; for signal sub-features in different frequency ranges, frequency compression is performed using convolution operations with different step lengths and different convolution kernel sizes.
[0233] In addition to using multiple convolutional layers to perform frequency compression on signal sub-features, a single convolution can also be used to perform frequency compression on signal sub-features, either by using a one-dimensional convolution with a step size of at least 2 in the frequency dimension, or by using a convolution kernel size of The two-dimensional convolution of Frequency compression is performed with a step size of .
[0234] In one possible implementation, in the process of merging multiple compressed sub-features along the frequency dimension to obtain the compressed feature, the multiple compressed sub-features may be merged along the frequency dimension, and the channel dimension of the merged result is restored to be consistent with the channel dimension of the initial signal feature and then normalized to obtain the compressed feature. Specifically, multiple compressed sub-features are merged along the frequency dimension to obtain a merged feature, and then the merged feature is input into a convolutional layer, the channel dimension of the merged feature is reduced by the convolutional layer, and the channel dimension of the merged feature is restored to be consistent with the channel dimension of the initial signal feature. Finally, the merged feature with the restored channel dimension is normalized to obtain the compressed feature. For example, the feature dimension of the initial signal feature is , multiple compressed sub-features are merged along the frequency dimension, and the feature dimension of the merged feature can be , M is the convolution scale. The merged features are input into a two-dimensional convolution, which reduces the channel dimension from ( M × E ) is reduced to E , that is, restore to the same channel dimension as the initial signal feature, and finally normalize the merged features of the restored channel dimension to obtain the compressed feature. The feature dimension of the compressed feature can be By keeping the channel dimension of the output compressed features consistent with the channel dimension of the input initial signal features, it helps to reduce the computational cost of the frequency compression process, thereby reducing the complexity of the frequency compression process, improving the efficiency of frequency compression, and thus improving the efficiency of enhancing the speech signal.
[0235] In a possible implementation, in the process of calling the encoder to encode based on the initial signal feature, the initial signal feature can also be divided into multiple signal sub-features along the frequency dimension, each signal sub-feature is normalized, and the normalized signal sub-feature is feature mapped to obtain multiple mapping sub-features. The multiple mapping sub-features are merged along the frequency dimension to obtain segmentation mapping features, and the segmentation mapping features are input into the encoder of the speech enhancement network for encoding. Fig.12 , Fig.12 Another optional schematic diagram of performing frequency band segmentation provided in the embodiment of the present disclosure is to divide the initial signal characteristics into X Split into Q signal sub-features along the frequency dimension ,in, , Indicates q The frequency division unit of the sub-band satisfies Next, the Q signal sub-features Perform normalization operations respectively and convert the normalized signal sub-features Input into the convolutional layer for feature mapping to obtain multiple mapping sub-features , , the convolution layer can be a one-dimensional convolution layer. Finally, multiple mapping sub-features Merge along the frequency dimension to obtain segmentation mapping features H The normalized signal sub-features are feature mapped through a convolution layer, which further simplifies the process of frequency band segmentation of the initial signal features and effectively reduces the amount of calculation in the frequency band segmentation process, thereby improving the efficiency of speech enhancement and helping to improve the real-time performance of speech enhancement.
[0236] In one possible implementation, the speech enhancement method provided by the embodiment of the present disclosure can be applied in a communication scenario, and a mixed voice signal is obtained through a terminal device such as a telephone, and a voice signal to be processed is obtained based on the mixed voice signal. The voice signal to be processed is input into the speech enhancement model provided by the embodiment of the present disclosure for echo elimination and noise suppression, so as to reduce background noise during the communication process and make the call clearer, thereby improving the call experience and ensuring the call quality.
[0237] In another possible implementation, the speech enhancement method provided in the embodiment of the present disclosure can be applied to a conference system. Usually, the venue for meetings is large and echoes are prone to occur. Therefore, by collecting mixed speech signals in the conference venue, echo elimination and noise suppression are performed based on the speech enhancement method provided in the embodiment of the present disclosure to improve the clarity of the conference speech and the participation of each participant.
[0238] Reference Fig.13 and Fig.14 , Fig.13 An optional overall flow chart of the speech enhancement method provided in the embodiment of the present disclosure is provided. Fig.14 The embodiments of the present disclosure provide Fig.13 An optional structural diagram of the speech enhancement model is shown below. The principle of the speech enhancement method in the embodiment of the present disclosure is fully described as a whole as follows:
[0239] The speech enhancement method of the embodiment of the present disclosure can be implemented through a speech enhancement model, which includes a feature extraction layer, a frequency band separation layer, a speech enhancement network, a frequency band merging layer and a post-processing network.
[0240] First, a sample speech signal is obtained, and a joint loss obtained by weighted summation of at least two loss functions is used to guide the training of the encoder and decoder to obtain a trained speech enhancement model.
[0241] Next, obtain the mixed speech signal and the reference speech signal , the mixed speech signal and the reference speech signal Input into the echo prediction model to obtain the echo signal and the error speech signal The mixed speech signal is transformed by short-time Fourier transform. , reference speech signal , echo signal and the error speech signal Perform time-frequency domain conversion to obtain the mixed speech spectrum D , Reference speech spectrum X , echo spectrum Y and the error speech spectrum E , so that these four signals are represented in complex form in the frequency domain. Then, the mixed speech spectrum D , Reference speech spectrum X , echo spectrum Y and the error speech spectrum E Splice along the channel dimension to obtain the speech signal to be processed I , the speech signal to be processed I Input to the feature extraction layer for feature extraction, perform feature extraction through two-dimensional convolution, and normalize the extracted features to obtain the initial signal features R .
[0242] Next, the initial signal features R Input to the speech enhancement model. Initial signal features R First input to the frequency band separation layer (such as Fig.10 ), the initial signal characteristics are converted into R Divide it into P signal sub-features, use a set of convolutions with different convolution kernel sizes and the same output channels to perform frequency compression on each signal sub-feature, obtain the compressed sub-features corresponding to each signal sub-feature, merge each compressed sub-feature along the frequency dimension to obtain the compressed feature H Then compress the feature H The compressed features are encoded by a plurality of sequentially cascaded processing blocks, and the encoded results are input to the decoder. The decoder decodes the encoded results by a plurality of sequentially cascaded primary blocks to obtain the decoded results. U .
[0243] In each processing block (such as Fig. 9 ), for the features of the input processing block, first downsample the input features according to the time compression ratio to obtain downsampled features, perform intra-band relationship modeling on the downsampled features along the frame rate dimension to obtain intra-band relationship features, and perform inter-band relationship modeling along the frequency dimension based on the intra-band relationship features (inter-band relationship modeling is as follows Figure 6), and finally upsample the inter-band relationship features according to the time compression ratio to obtain the output of the processing block.
[0244] Next, the decoding result U Input to the band merging layer (such as Figure 8 ), the decoding result is divided into multiple decoding sub-features along the frequency dimension, each decoding sub-feature is normalized and input into the multi-layer perception layer to predict the sub-time-frequency mask, and then each sub-time-frequency mask is merged along the frequency dimension to obtain the target speech signal G . The target speech signal G With the speech signal to be processed I Perform element-wise multiplication to obtain fused features When there is residual noise in the target speech signal, the fusion feature The signal is then input into the post-processing network, and the noise is further suppressed through the gated recurrent unit and linear transformation to obtain enhanced signal features. .
[0245] Finally, the signal features are enhanced by inverse short-time Fourier transform. Perform inverse time-frequency conversion to obtain enhanced speech signal .
[0246] Reference Fig.15 , Fig.15 The embodiments of the present disclosure provide Fig.13 Another optional structural diagram of the speech enhancement model in FIG. 1 is a schematic diagram of the speech enhancement model, which includes a frequency band separation layer, a speech enhancement network, a frequency band merging layer, and a logarithmic variance estimation module. The following is an overall diagram based on Fig.15 The principle of the speech enhancement method in the embodiment of the present disclosure is fully described as follows:
[0247] First, a sample speech signal is obtained, and the speech enhancement network is trained in the first stage based on the sample speech signal, the first uncertainty loss, the signal-to-noise ratio loss, and the update rate loss. Then, the speech enhancement network is trained in the second stage based on the sample speech signal, the second uncertainty loss, the signal-to-noise ratio loss, and the update rate loss to obtain a trained speech enhancement network. The logarithmic variance estimation module is used to predict the uncertainty of the enhancement result of the sample speech signal during the training process.
[0248] Next, the mixed speech signal, the reference speech signal, the echo signal and the error speech signal are obtained as the speech signal to be processed. I , from the speech signal to be processed I Extract the initial signal features X , where the initial signal characteristics X Including the real part of the complex spectrum and the imaginary part of the complex spectrum .
[0249] Next, the initial signal features X Input to the band splitting layer (such as Fig.12 ), the initial signal features are transformed into X Split into Q Sub-bands, each sub-band corresponds to a signal sub-feature, feature mapping is performed on multiple signal sub-features to obtain mapping sub-features corresponding to each signal sub-feature, and each mapping sub-feature is merged to obtain a segmentation mapping feature H Then the segmentation mapping feature H The input is sent to the encoder of the speech enhancement network, and the segmentation mapping features are processed through multiple sequentially cascaded processing blocks. H Encode to get the encoding result, then input the encoding result into the decoder for decoding to get the decoding result U .
[0250] In each processing block (such as Figure 3 ), for the features of the processing block of the input encoder, firstly, the frame mixing module performs frame mixing mapping along the time dimension to obtain the first mapping features. The first mapping features are input into the band mixing module to perform sub-band mixing mapping along the sub-band dimension to obtain the second mapping features. Finally, the second mapping features are input into the sub-band compression / decompression module to perform band compression operation to obtain the output of the processing block. For the features of the processing block of the input decoder, after obtaining the second mapping features, the second mapping features are input into the sub-band compression / decompression module to perform band decompression operation to obtain the output of the processing block.
[0251] Next, the decoding result U Input to the band merging layer (such as Figure 8 ), the decoding result is divided into multiple decoding sub-features along the frequency dimension, each decoding sub-feature is normalized and input into the multi-layer perception layer to predict the sub-time-frequency mask, and then each sub-time-frequency mask is merged along the frequency dimension to obtain the target speech signal G . The target speech signal G and initial signal characteristics X Perform element-wise multiplication to obtain the enhanced real part characteristics of the complex spectrum and enhanced complex spectrum imaginary part features .
[0252] Finally, the enhanced real part of the complex spectrum is transformed by short-time inverse Fourier transform. and enhanced complex spectrum imaginary part features Perform inverse time-frequency conversion to obtain enhanced speech signal .
[0253] The speech enhancement method provided by the embodiment of the present disclosure performs frequency band segmentation on the initial speech signal, performs frequency compression on the frequency unit along the frequency dimension, obtains the compressed features, and then passes the compressed features through the codec network to suppress noise to obtain the decoding result. Then, the decoding result is segmented into multiple decoding sub-features along the frequency dimension, each decoding sub-feature is normalized and input into the prediction sub-time-frequency mask in the multi-layer perception layer, and then each sub-time-frequency mask is merged along the frequency dimension to obtain the target speech signal, and the target speech signal is further subjected to noise suppression through the post-processing network to obtain the enhanced signal features. The enhanced speech signal is obtained based on the enhanced signal features, which can effectively reduce the computational complexity of speech enhancement and improve the real-time performance of speech enhancement.
[0254] It is to be understood that, although the steps in the above-mentioned various flow charts are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in the present embodiment, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flow charts can include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.
[0255] Reference Fig.16 , Fig.16 This is a schematic diagram of the structure of a speech enhancement device provided in an embodiment of the present disclosure. The speech enhancement device 1600 includes:
[0256] 1601 input module, used to obtain a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the processing block includes a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate updating the hidden state features of the gated recurrent unit or keeping the hidden state features unchanged;
[0257] 1602 encoding module, used for performing frame hybrid mapping on the initial signal feature to obtain a first mapping feature based on the first processing block in the encoder, performing band hybrid mapping on the first mapping feature to obtain a second mapping feature, performing band compression on the second mapping feature to obtain an output of the first processing block, and inputting the output of the first processing block to the next processing block until an encoding result of the encoder is obtained;
[0258] 1603 decoding module, used for decoding the encoding result of the encoder based on the decoder to obtain the target signal characteristics;
[0259] 1604 restoration module, used for restoring the enhanced speech signal from the speech signal to be processed based on the target signal feature.
[0260] Furthermore, the encoding module 1602 is also used for:
[0261] Normalizing the initial signal features, and performing frame mixing mapping on the normalized initial signal features through a gated recurrent unit;
[0262] The output of the gated recurrent unit is fully connected to obtain the first mapping feature.
[0263] Furthermore, the encoding module 1602 is also used for:
[0264] For any current time step among the multiple time steps, obtaining the initial state update probability and the initial state update probability change value of the gated recurrent unit in the previous time step of the current time step, and determining the target state update probability of the gated recurrent unit in the current time step according to the initial state update probability and the initial state update probability change value;
[0265] Round the target state update probability to get the gating parameter corresponding to the binary gate in the current time step;
[0266] The hidden state features output by the gated recurrent unit in the current time step are determined according to the gating parameters, and frame mixing mapping is performed on the normalized initial signal features through the gated recurrent unit based on the hidden state features.
[0267] Furthermore, the encoding module 1602 is also used for:
[0268] According to the gating parameters corresponding to the current time step, the probability of updating the initial state corresponding to the current time step is determined;
[0269] According to the gating parameters corresponding to the current time step and the target state update probability, the initial state update probability change value corresponding to the current time step is determined.
[0270] Furthermore, the encoding module 1602 is also used for:
[0271] Obtain a sample speech signal and a sample label signal corresponding to the sample speech signal, and call a speech enhancement network to perform speech enhancement on the sample speech signal;
[0272] The logarithmic variance of the enhanced result of the sample speech signal is determined by the logarithmic variance estimation module, the target loss is determined according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, and the speech enhancement network is trained based on the target loss.
[0273] Furthermore, the encoding module 1602 is also used for:
[0274] Determine the target loss of the first training stage according to the enhancement result of the sample speech signal and the difference between the sample label signal and the logarithmic variance, wherein the target loss of the first training stage is used to constrain the accuracy of the logarithmic variance;
[0275] Determine the target loss of the second training stage according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, wherein the target loss of the second training stage is used to assign corresponding weights to the time-frequency units according to the uncertainty of the time-frequency units when determining the enhanced result of the sample speech signal and the difference between the sample label signal;
[0276] The speech enhancement network is trained sequentially based on the target loss of the first training stage and the target loss of the second training stage.
[0277] Furthermore, the encoding module 1602 is also used for:
[0278] Determine the difference in the real part of the complex spectrum, the difference in the imaginary part of the complex spectrum, and the difference in the amplitude spectrum between the enhancement result of the sample speech signal and the sample label signal;
[0279] The real part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the real part of the complex spectrum as an exponent, the imaginary part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the imaginary part of the complex spectrum as an exponent, and the first loss is determined according to the sum of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, the weighted real part difference of the complex spectrum, and the weighted imaginary part difference of the complex spectrum;
[0280] The amplitude spectrum difference is weighted according to a natural exponential function with the amplitude spectrum logarithmic variance as an exponent, and a second loss is determined according to the sum of the amplitude spectrum logarithmic variance and the weighted amplitude spectrum difference;
[0281] The weighted sum of the first loss and the second loss is used to obtain the target loss of the first training stage.
[0282] Furthermore, the encoding module 1602 is also used for:
[0283] Determine the signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine the update rate loss according to the difference between the update rate of the binary gate and the preset update rate threshold;
[0284] The first uncertainty loss is obtained by weighted summing of the first loss and the second loss, and the target loss of the first training stage is obtained by weighted summing of the first uncertainty loss, the signal-to-noise ratio loss and the update rate loss.
[0285] Furthermore, the encoding module 1602 is also used for:
[0286] The real part difference of the complex spectrum is weighted according to the logarithmic variance of the real part of the complex spectrum, the imaginary part difference of the complex spectrum is weighted according to the logarithmic variance of the imaginary part of the complex spectrum, and the third loss is determined according to the sum of the weighted real part difference of the complex spectrum and the weighted imaginary part difference of the complex spectrum;
[0287] weighting the amplitude spectrum difference according to the amplitude spectrum logarithmic variance, and determining the weighted amplitude spectrum difference as the fourth loss;
[0288] The weighted sum of the third loss and the fourth loss is used to obtain the target loss of the second training stage.
[0289] Furthermore, the encoding module 1602 is also used for:
[0290] Determine the signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine the update rate loss according to the difference between the update rate of the binary gate and the preset update rate threshold;
[0291] The third loss and the fourth loss are weightedly summed to obtain the second uncertainty loss, and the second uncertainty loss, the signal-to-noise ratio loss, and the update rate loss are weightedly summed to obtain the target loss of the second training stage.
[0292] Furthermore, the encoding module 1602 is also used for:
[0293] The first mapping feature is segmented along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature;
[0294] Normalizing the first intra-band relationship sub-feature, and convolving the normalized first intra-band relationship sub-feature in the sub-band dimension along the time axis to obtain a first convolution feature;
[0295] A first product of the first convolution feature and the second in-band relation sub-feature is determined, and a second mapping feature is obtained based on a sum of the first product and the second in-band relation sub-feature.
[0296] Furthermore, the encoding module 1602 is also used for:
[0297] Convolve the first mapping feature in the channel dimension along the time axis to obtain a second convolution feature;
[0298] The second convolution feature is normalized, and the normalized second convolution feature is segmented along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature.
[0299] Furthermore, the encoding module 1602 is also used for:
[0300] The first product and the sum of the second intra-band relationship sub-features are convolved in the channel dimension along the time axis to obtain an inter-band relationship feature for indicating the inter-band relationship.
[0301] Furthermore, the encoding module 1602 is also used for:
[0302] Normalizing the second mapping features, and performing parallel convolution on the normalized second mapping features to obtain compressed third convolution features and fourth convolution features;
[0303] Perform nonlinear activation on the fourth convolution feature to obtain the attention weight;
[0304] The third convolutional feature is weighted based on the attention weight to obtain the output of the first processing block.
[0305] The electronic device for executing the above-mentioned voice enhancement method provided in the embodiment of the present disclosure may be a terminal. Fig.17 , Fig.17 This is a partial structural block diagram of a terminal provided in an embodiment of the present disclosure, and the terminal includes: a camera assembly 1710, a first memory 1720, an input unit 1730, a display unit 1740, a sensor 1750, an audio circuit 1760, a wireless fidelity (WiFi) module 1770, a first processor 1780, and a first power supply 1790. Those skilled in the art can understand that Fig.17 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0306] The camera assembly 1710 can be used to capture images or videos. Optionally, the camera assembly 1710 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting, and the VR (Virtual Reality) shooting function or other fusion shooting functions.
[0307] The first memory 1720 may be used to store software programs and modules. The first processor 1780 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 1720 .
[0308] The input unit 1730 may be used to receive input digital or character information and generate key signal input related to the terminal's settings and function control. Specifically, the input unit 1730 may include a touch panel 1731 and other input devices 1732 .
[0309] The display unit 1740 may be used to display input information or provided information and various menus of the terminal. The display unit 1740 may include a display panel 1741 .
[0310] The audio circuit 1760, the speaker 1761, and the microphone 1762 may provide an audio interface.
[0311] The first power source 1790 may be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0312] The number of sensors 1750 may be one or more, and the one or more sensors 1750 include but are not limited to: acceleration sensors, gyroscope sensors, pressure sensors, optical sensors, etc. Among them:
[0313] The acceleration sensor can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of gravity acceleration on the three coordinate axes. The first processor 1780 can control the display unit 1740 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for collecting game or user motion data.
[0314] The gyroscope sensor can detect the body direction and rotation angle of the terminal, and the gyroscope sensor can cooperate with the acceleration sensor to collect the user's 3D actions on the terminal. The first processor 1780 can implement the following functions based on the data collected by the gyroscope sensor: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0315] The pressure sensor can be set in the side frame of the terminal and / or the lower layer of the display unit 1740. When the pressure sensor is set in the side frame of the terminal, the user's holding signal of the terminal can be detected, and the first processor 1780 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor. When the pressure sensor is set in the lower layer of the display unit 1740, the first processor 1780 controls the operability controls on the UI interface according to the user's pressure operation on the display unit 1740. The operability control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0316] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1780 can control the display brightness of the display unit 1740 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1740 is increased; when the ambient light intensity is low, the display brightness of the display unit 1740 is decreased. In another embodiment, the first processor 1780 can also dynamically adjust the shooting parameters of the camera assembly 1710 according to the ambient light intensity collected by the optical sensor.
[0317] In this embodiment, the first processor 1780 included in the terminal can execute the speech enhancement method of the previous embodiment.
[0318] The electronic device for executing the above-mentioned voice enhancement method provided in the embodiment of the present disclosure may also be a server. Fig.18 , Fig.18 Partial structural block diagram of a server provided in an embodiment of the present disclosure. The server may have relatively large differences due to different configurations or performances, and may include one or more second processors 1810 and a second memory 1830, and one or more storage media 1840 (e.g., one or more mass storage devices) storing application programs 1843 or data 1842. Among them, the second memory 1830 and the storage medium 1840 may be short-term storage or persistent storage. The program stored in the storage medium 1840 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the second processor 1810 may be configured to communicate with the storage medium 1840 to execute a series of instruction operations in the storage medium 1840 on the server.
[0319] The server may also include one or more second power supplies 1820, one or more wired or wireless network interfaces 1850, one or more input and output interfaces 1860, and / or one or more operating systems 1841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0320] The second processor 1810 in the server may be configured to execute the speech enhancement method.
[0321] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the speech enhancement method of each of the aforementioned embodiments.
[0322] The present disclosure also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device implements the above-mentioned speech enhancement method.
[0323] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate to describe the embodiments of the present disclosure, such as being able to be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0324] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0325] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood as not including the number, and above, below, within, etc. are understood as including the number.
[0326] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0327] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0328] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0329] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0330] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0331] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions under the shared conditions without violating the spirit of the present disclosure. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A speech enhancement method, characterized in that: include: Acquire a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the processing blocks include a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate updating the hidden state features of the gated recurrent unit or keeping the hidden state features unchanged; Based on the first processing block in the encoder, frame hybrid mapping is performed on the initial signal feature to obtain a first mapping feature, band hybrid mapping is performed on the first mapping feature to obtain a second mapping feature, band compression is performed on the second mapping feature to obtain the output of the first processing block, and the output of the first processing block is input to the next processing block until the encoding result of the encoder is obtained; Decoding the encoding result of the encoder based on the decoder to obtain target signal characteristics; An enhanced speech signal is restored from the speech signal to be processed based on the target signal feature.
2. The speech enhancement method according to claim 1, characterized in that: The performing frame hybrid mapping on the initial signal feature to obtain a first mapping feature includes: Normalizing the initial signal features, and performing frame mixing mapping on the normalized initial signal features through the gated cycle unit; Fully connect the output of the gated recurrent unit to obtain the first mapping feature.
3. The speech enhancement method according to claim 2, characterized in that: The speech enhancement network is configured to iterate through multiple time steps, and the gated cycle unit performs frame mixing mapping on the normalized initial signal features, including: For any current time step among the multiple time steps, obtaining the initial state update probability and the initial state update probability change value of the gated recurrent unit in the previous time step of the current time step, and determining the target state update probability of the gated recurrent unit in the current time step according to the initial state update probability and the initial state update probability change value; Rounding the target state update probability to obtain a gating parameter corresponding to the binary gate in the current time step; The hidden state feature output by the gated cyclic unit in the current time step is determined according to the gating parameter, and frame mixing mapping is performed on the normalized initial signal feature through the gated cyclic unit based on the hidden state feature.
4. The speech enhancement method according to claim 3, characterized in that: After rounding the target state update probability to obtain the gating parameter corresponding to the binary gate in the current time step, the speech enhancement method further includes: Determining the initial state update probability corresponding to the current time step according to the gating parameter corresponding to the current time step; The initial state update probability change value corresponding to the current time step is determined according to the gating parameter corresponding to the current time step and the target state update probability.
5. The speech enhancement method according to claim 1, characterized in that: The speech enhancement network is further provided with a logarithmic variance estimation module. Before the initial signal feature is input into the speech enhancement network, the speech enhancement method further comprises: Acquire a sample speech signal and a sample label signal corresponding to the sample speech signal, and call the speech enhancement network to perform speech enhancement on the sample speech signal; The logarithmic variance of the enhancement result of the sample speech signal is determined by the logarithmic variance estimation module, the target loss is determined according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance, and the speech enhancement network is trained based on the target loss.
6. The speech enhancement method according to claim 5, characterized in that: The step of determining a target loss according to the enhancement result of the sample speech signal, the difference between the sample label signal and the logarithmic variance, and training the speech enhancement network based on the target loss comprises: Determine a target loss for a first training phase according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, wherein the target loss for the first training phase is used to constrain the accuracy of the logarithmic variance; Determine a target loss for a second training phase according to the difference between the enhanced result of the sample speech signal and the sample label signal and the logarithmic variance, wherein the target loss for the second training phase is used to assign corresponding weights to the time-frequency units according to the uncertainty of the time-frequency units when determining the enhanced result of the sample speech signal and the difference between the sample label signal; The speech enhancement network is trained sequentially based on the target loss of the first training stage and the target loss of the second training stage.
7. The speech enhancement method according to claim 6, characterized in that: The logarithmic variance includes the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, and the logarithmic variance of the amplitude spectrum, and determining the target loss of the first training stage according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance includes: Determine a real part difference of a complex spectrum, an imaginary part difference of a complex spectrum, and an amplitude spectrum difference between the enhancement result of the sample speech signal and the sample label signal; The real part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the real part of the complex spectrum as an exponent, the imaginary part difference of the complex spectrum is weighted according to a natural exponential function with the logarithmic variance of the imaginary part of the complex spectrum as an exponent, and a first loss is determined according to the sum of the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, the weighted real part difference of the complex spectrum, and the weighted imaginary part difference of the complex spectrum; Weighting the amplitude spectrum difference according to a natural exponential function with the amplitude spectrum logarithmic variance as an exponent, and determining a second loss according to the sum of the amplitude spectrum logarithmic variance and the weighted amplitude spectrum difference; The first loss and the second loss are weightedly summed to obtain the target loss of the first training stage.
8. The speech enhancement method according to claim 7, characterized in that: The step of obtaining a target loss of the first training stage by weighted summing of the first loss and the second loss includes: Determine a signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine an update rate loss according to a difference between an update rate of the binary gate and a preset update rate threshold; The first loss and the second loss are weightedly summed to obtain a first uncertainty loss, and the first uncertainty loss, the signal-to-noise ratio loss, and the update rate loss are weightedly summed to obtain a target loss for the first training stage.
9. The speech enhancement method according to claim 6, characterized in that: The logarithmic variance includes the logarithmic variance of the real part of the complex spectrum, the logarithmic variance of the imaginary part of the complex spectrum, and the logarithmic variance of the amplitude spectrum, and determining the target loss of the second training stage according to the difference between the enhancement result of the sample speech signal and the sample label signal and the logarithmic variance includes: Determine a real part difference of a complex spectrum, an imaginary part difference of a complex spectrum, and an amplitude spectrum difference between the enhancement result of the sample speech signal and the sample label signal; Weighting the real part difference of the complex spectrum according to the logarithmic variance of the real part of the complex spectrum, weighting the imaginary part difference of the complex spectrum according to the logarithmic variance of the imaginary part of the complex spectrum, and determining a third loss according to the sum of the weighted real part difference of the complex spectrum and the weighted imaginary part difference of the complex spectrum; weighting the amplitude spectrum difference according to the amplitude spectrum logarithmic variance, and determining the weighted amplitude spectrum difference as a fourth loss; The target loss of the second training stage is obtained by weighted summing the third loss and the fourth loss.
10. The speech enhancement method according to claim 9, characterized in that: The step of obtaining the target loss of the second training stage by weighted summing of the third loss and the fourth loss includes: Determine a signal-to-noise ratio loss according to the enhancement result of the sample speech signal and the sample label signal, and determine an update rate loss according to a difference between an update rate of the binary gate and a preset update rate threshold; The third loss and the fourth loss are weightedly summed to obtain a second uncertainty loss, and the second uncertainty loss, the signal-to-noise ratio loss and the update rate loss are weightedly summed to obtain the target loss of the second training stage.
11. The speech enhancement method according to claim 1, characterized in that: The performing mixed mapping on the first mapping feature to obtain the second mapping feature includes: Segmenting the first mapping feature along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature; Normalizing the first intra-band relationship sub-feature, and convolving the normalized first intra-band relationship sub-feature in a sub-band dimension along a time axis to obtain a first convolution feature; A first product of the first convolution feature and the second in-band relation sub-feature is determined, and the second mapping feature is obtained based on a sum of the first product and the second in-band relation sub-feature.
12. The speech enhancement method according to claim 11, characterized in that: The segmenting of the first mapping feature along the channel dimension to obtain a first in-band relationship sub-feature and a second in-band relationship sub-feature comprises: Convolving the first mapping feature in the channel dimension along the time axis to obtain a second convolution feature; The second convolution feature is normalized, and the normalized second convolution feature is segmented along the channel dimension to obtain the first in-band relationship sub-feature and the second in-band relationship sub-feature.
13. The speech enhancement method according to claim 11, characterized in that: The obtaining the second mapping feature based on the sum of the first product and the second in-band relationship sub-feature includes: Convolving the first product and the sum of the second in-band relationship sub-features in the channel dimension along the time axis to obtain the second mapping feature.
14. The method for speech enhancement according to claim 1, characterized in that: The step of performing band compression on the second mapping feature to obtain the output of the first processing block comprises: Normalizing the second mapping features, and performing parallel convolution on the normalized second mapping features to obtain compressed third convolution features and fourth convolution features; Performing nonlinear activation on the fourth convolutional feature to obtain an attention weight; The third convolution feature is weighted based on the attention weight to obtain the output of the first processing block.
15. A speech enhancement device, characterized in that: include: An input module is used to obtain a speech signal to be processed, perform feature extraction on the speech signal to be processed, obtain initial signal features, and input the initial signal features into a speech enhancement network, wherein the speech enhancement network includes an encoder and a decoder connected to each other, the encoder and the decoder are both provided with a plurality of processing blocks cascaded in sequence, the processing blocks include a gated recurrent unit, the gated recurrent unit is configured with a binary gate, and the binary gate is used to indicate whether to update the hidden state features of the gated recurrent unit or keep the hidden state features unchanged; An encoding module, configured to perform frame hybrid mapping on the initial signal feature to obtain a first mapping feature based on the first processing block in the encoder, perform band hybrid mapping on the first mapping feature to obtain a second mapping feature, perform band compression on the second mapping feature to obtain an output of the first processing block, and input the output of the first processing block to the next processing block until an encoding result of the encoder is obtained; A decoding module, used for decoding the encoding result of the encoder based on the decoder to obtain target signal characteristics; The restoration module is used to restore the to-be-processed speech signal to obtain an enhanced speech signal based on the target signal feature.
16. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the speech enhancement method according to any one of claims 1 to 14 is implemented.
17. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech enhancement method according to any one of claims 1 to 14 is implemented.
18. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the speech enhancement method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
Voice enhancement network model and single-channel speech enhancement method and system
CN112509593A
Voice noise reduction method and device based on artificial intelligence, equipment and storage medium
CN114694674A