Voice enhancement system, method, device and equipment
By collecting multi-channel speech data through a microphone array and utilizing complex domain calculation and a mask estimation model of the attention module, the problem of insufficient speech enhancement performance in far-field environments is solved, and high-accuracy speech masking and background noise suppression in far-field environments are achieved, thereby improving speech quality and recognition accuracy.
Patent Information
- Application Number
- CN202110656378.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-11
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-06-11
AI Technical Summary
Existing technologies have low speech enhancement performance in far-field environments and cannot effectively suppress background noise interference. In particular, it is difficult to separate the target speech in far-field environments with low speech signal-to-noise ratio, reverberation, and multiple people speaking.
A microphone array is used to collect multi-channel speech data, and speech mask estimation is performed through a mask estimation model. Combined with complex domain calculation and attention module, enhanced speech data is generated, taking into account the correlation between amplitude spectrum, phase spectrum and multi-channel spectral features.
It effectively improves the accuracy of voice masking, especially in far-field environments, and can effectively suppress background noise interference, thereby improving the overall voice quality and recognition accuracy.
Smart Images

Figure CN115472153B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and specifically to speech enhancement systems, methods and devices, mask estimation model methods and devices, conference systems, online education systems, live broadcast systems, speech recognition systems, speech recognition text editing systems, user recognition systems, and electronic devices. Background Art
[0002] Speech enhancement (noise reduction) technology has been applied to scenarios in multiple fields, such as audio and video conferencing and online education in the field of real-time communication (RTC), and speech recognition and speaker recognition in the field of machine recognition. Its purpose is to suppress other noises besides the target speech, such as howling, mobile phone ringtones, keyboard sounds, and background voices.
[0003] With the application of deep learning technology, intelligent speech noise reduction has significant advantages over traditional signal processing noise reduction, especially in terms of model learning capabilities and handling non-stationary noise. Currently, a typical deep learning-based speech enhancement solution uses a complex convolutional recurrent network (CCRN) for single-channel speech enhancement.
[0004] However, in the process of implementing the present invention, the inventors found that the above solution has at least the following problems: the performance is limited in a far-field environment and cannot meet the requirements of speech enhancement performance in a far-field environment. Summary of the Invention
[0005] This application provides a speech enhancement system to address the existing problem of low speech enhancement performance in far-field environments. This application also provides a speech enhancement method and apparatus, a mask estimation model method and apparatus, a conference system, an online education system, a live broadcast system, a speech recognition system, a speech recognition text editing system, a user recognition system, and electronic equipment.
[0006] This application provides a speech enhancement system, comprising:
[0007] A voice acquisition module is used to collect multi-channel voice data through a microphone array;
[0008] The speech enhancement module is used to determine the spectral feature data of multi-channel speech data; determine the speech mask based on the spectral feature data of the multi-channel speech data through a mask estimation model; and generate enhanced speech data based on the speech mask and the spectral feature data of the single-channel speech data.
[0009] This application also provides a speech enhancement method, comprising:
[0010] Collect multi-channel voice data;
[0011] Determining a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model;
[0012] Enhanced speech data is generated according to the speech mask and the spectrum feature data of the single-channel speech data.
[0013] Optionally, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder;
[0014] The encoder includes a channel attention module corresponding to the encoding layer;
[0015] The decoder includes a first spectral attention module corresponding to the decoding layer;
[0016] The method of determining a speech mask based on spectral feature data of multi-channel speech data by using a mask estimation model comprises:
[0017] Determining, by a channel attention module, channel feature attention weights based on the multi-channel spectral feature data to be encoded, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weights;
[0018] Encoding the adjusted multi-channel spectrum feature data to be encoded through the encoding layer;
[0019] as well as,
[0020] determining, by a first spectrum attention module, a first spectrum attention weight according to the multi-channel spectrum feature data to be decoded, and adjusting the multi-channel spectrum feature data to be decoded according to the first spectrum feature attention weight;
[0021] The adjusted multi-channel spectrum feature data to be decoded is decoded through the decoding layer.
[0022] Optionally, the mask estimation model further includes a skip-layer connection between the encoding layer and the decoding layer, wherein the skip-layer connection includes a second spectral attention module;
[0023] The method further comprises determining the speech mask according to the spectral feature data of the multi-channel speech data by using the mask estimation model:
[0024] determining, by a second spectral attention module, a second spectral attention weight based on the multi-channel spectral feature encoded data output by the encoding layer, and adjusting the multi-channel spectral feature encoded data based on the second spectral feature attention weight;
[0025] The multi-channel spectrum feature data to be decoded of the next decoding layer is determined according to the output data of the decoding layer corresponding to the encoding layer and the adjusted multi-channel spectrum feature encoding data.
[0026] Optionally, also include:
[0027] The channel attention module determines the channel feature attention weight according to the multi-channel spectral feature data to be encoded, including:
[0028] Compressing the frequency and time dimension data of the multi-channel spectral feature data to be encoded as first compressed data through the compression submodule in the channel attention module;
[0029] The channel attention weight determination layer in the channel attention module determines the channel feature attention weight according to the first compressed data.
[0030] Optionally, also include:
[0031] The first spectrum attention module determines the first spectrum attention weight according to the multi-channel spectrum feature data to be decoded, including:
[0032] Compressing the channel dimension data of the multi-channel spectrum feature data to be decoded as the second compressed data through the compression submodule in the first spectrum attention module;
[0033] The first spectrum attention weight is determined according to the second compressed data through the spectrum attention weight determination layer in the first spectrum attention module.
[0034] Optionally, the encoder includes four encoding layers, and the decoder includes four decoding layers.
[0035] Optionally, the frequency spectrum feature data of the single-channel speech data is determined in the following manner:
[0036] The frequency spectrum feature data of the single-channel speech data with a high signal-to-noise ratio is selected from the frequency spectrum feature data of the multi-channel speech data.
[0037] Optionally, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder;
[0038] The encoder includes a channel attention module corresponding to the encoding layer;
[0039] The method of determining a speech mask based on spectral feature data of multi-channel speech data by using a mask estimation model comprises:
[0040] Determining, by a channel attention module, channel feature attention weights based on the multi-channel spectral feature data to be encoded, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weights;
[0041] The adjusted multi-channel spectrum feature data to be encoded is encoded through the encoding layer.
[0042] Optionally, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder;
[0043] The decoder includes a first spectral attention module corresponding to the decoding layer;
[0044] The method of determining a speech mask based on spectral feature data of multi-channel speech data by using a mask estimation model comprises:
[0045] determining, by a first spectrum attention module, a first spectrum attention weight according to the multi-channel spectrum feature data to be decoded, and adjusting the multi-channel spectrum feature data to be decoded according to the first spectrum feature attention weight;
[0046] The adjusted multi-channel spectrum feature data to be decoded is decoded through the decoding layer.
[0047] Optionally, the mask estimation model further includes a skip-layer connection between the encoding layer and the decoding layer, wherein the skip-layer connection includes a second spectral attention module;
[0048] The method further comprises determining the speech mask according to the spectral feature data of the multi-channel speech data by using the mask estimation model:
[0049] determining, by a second spectral attention module, a second spectral attention weight based on the multi-channel spectral feature encoded data output by the encoding layer, and adjusting the multi-channel spectral feature encoded data based on the second spectral feature attention weight;
[0050] The multi-channel spectrum feature data to be decoded of the next decoding layer is determined according to the output data of the decoding layer corresponding to the encoding layer and the adjusted multi-channel spectrum feature encoding data.
[0051] This application also provides a speech enhancement method, comprising:
[0052] The client collects multi-channel voice data through a microphone array;
[0053] The multi-channel speech data is sent to a server so that the server determines spectral feature data of the multi-channel speech data; a speech mask is determined based on the spectral feature data of the multi-channel speech data through a mask estimation model; and enhanced speech data is generated based on the speech mask and the spectral feature data of the single-channel speech data.
[0054] This application also provides a speech enhancement method, comprising:
[0055] The server receives multi-channel voice data;
[0056] Determining spectral feature data of multi-channel speech data;
[0057] Determining a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model;
[0058] Enhanced speech data is generated according to the speech mask and the spectrum feature data of the single-channel speech data.
[0059] This application also provides a mask estimation model processing method, including:
[0060] Determine a training data set, wherein the training data includes: spectral feature data and a speech mask of multi-channel speech data;
[0061] Construct the network structure of the mask estimation model based on complex-valued networks;
[0062] The network parameters of the mask estimation model are trained according to the training data set.
[0063] This application also provides a speech recognition system, comprising:
[0064] A voice acquisition module is used to collect multi-channel voice data through a microphone array;
[0065] The speech enhancement module is used to determine the spectral feature data of the multi-channel speech data; determine the speech mask based on the spectral feature data of the multi-channel speech data through the mask estimation model; and generate enhanced speech data based on the speech mask and the spectral feature data of the single-channel speech data;
[0066] The speech recognition module is used to convert enhanced speech data into text through a speech recognition model.
[0067] This application also provides a conference system, including:
[0068] Voice collection module, used to collect multi-channel conference voice data through microphone array;
[0069] The speech enhancement module is used to determine the spectral feature data of the multi-channel conference speech data; determine the speech mask based on the spectral feature data of the multi-channel conference speech data through the mask estimation model; and generate enhanced conference speech data based on the speech mask and the spectral feature data of the single-channel conference speech data;
[0070] The voice playing module is used to play the enhanced conference voice data.
[0071] This application also provides a live broadcast system, including:
[0072] Voice acquisition module, used to collect multi-channel live voice data through microphone array;
[0073] The speech enhancement module is used to determine the spectral feature data of the multi-channel live speech data; determine the speech mask based on the spectral feature data of the multi-channel live speech data through the mask estimation model; and generate enhanced live speech data based on the speech mask and the spectral feature data of the single-channel live speech data;
[0074] The voice playing module is used to play the enhanced live voice data.
[0075] This application also provides an online education system, including:
[0076] The voice acquisition module is used to collect multi-channel online education voice data through a microphone array;
[0077] The speech enhancement module is used to determine the spectral feature data of the multi-channel online education speech data; determine the speech mask based on the spectral feature data of the multi-channel online education speech data through the mask estimation model; and generate enhanced online education speech data based on the speech mask and the spectral feature data of the single-channel online education speech data;
[0078] The voice playing module is used to play the enhanced online education voice data.
[0079] The present application also provides a speech enhancement device, comprising:
[0080] The client collects multi-channel voice data through a microphone array;
[0081] The multi-channel speech data is sent to a server so that the server determines spectral feature data of the multi-channel speech data; a speech mask is determined based on the spectral feature data of the multi-channel speech data through a mask estimation model; and enhanced speech data is generated based on the speech mask and the spectral feature data of the single-channel speech data.
[0082] The present application also provides a speech enhancement device, comprising:
[0083] The server receives multi-channel voice data;
[0084] Determining spectral feature data of multi-channel speech data;
[0085] Determining a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model;
[0086] Enhanced speech data is generated according to the speech mask and the spectrum feature data of the single-channel speech data.
[0087] The present application also provides a mask estimation model processing device, comprising:
[0088] Determine a training data set, wherein the training data includes: spectral feature data and a speech mask of multi-channel speech data;
[0089] Construct the network structure of the mask estimation model based on complex-valued networks;
[0090] The network parameters of the mask estimation model are trained according to the training data set.
[0091] The present application also provides a speech enhancement device, comprising:
[0092] Collect multi-channel voice data;
[0093] Determining a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model;
[0094] Enhanced speech data is generated according to the speech mask and the spectrum feature data of the single-channel speech data.
[0095] The present application also provides an electronic device, comprising:
[0096] A processor and a memory; the memory is used to store a program for implementing the above method, and the device is powered on and runs the program of the method through the processor.
[0097] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned various methods.
[0098] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to perform the above methods.
[0099] Compared with the prior art, this application has the following advantages:
[0100] The speech enhancement system provided in the embodiment of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; and generates enhanced speech data based on the speech mask and the spectral feature data of any channel. This processing method not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the speech mask, but also takes into account the correlation between the multi-channel spectral features, and can enhance the representation capability of the features that need to be restored to the target speech. This can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the overall speech quality.
[0101] The mask estimation model processing method provided in the embodiment of the present application determines a training data set, wherein the training data includes: spectral feature data of multi-channel speech data and a single-channel speech mask; constructs a network structure of a mask estimation model based on a complex-valued network; and trains the network parameters of the mask estimation model according to the training data set. By adopting this processing method, the constructed mask estimation model not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the speech mask, but also takes into account the correlation between the multi-channel spectral features, and can enhance the representation ability of the features that need to be restored to the target speech. This can effectively improve the accuracy of speech masking, especially the accuracy of speech masking in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the overall speech quality.
[0102] The speech recognition system provided in the embodiment of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; and converts the enhanced speech data into text through a speech recognition model. This processing method not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the speech mask, but also takes into account the correlation between the multi-channel spectral features, and can enhance the representation capability of the features that need to be restored to the target speech. This can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference, improving the overall speech quality, and thus improving the accuracy of speech recognition.
[0103] The conference (such as video conferencing, telephone conferencing, etc.) system provided by the embodiment of the present application collects multi-channel conference voice data through a microphone array; estimates the voice mask of the noisy conference voice based on the spectral feature data of the multi-channel conference voice data through a mask estimation model; generates enhanced conference voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced conference voice data. By adopting this processing method, when estimating the voice mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization capability of the features that need to be restored to the target voice can be enhanced, which can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with poor reception, thereby effectively suppressing background noise interference and improving the overall quality of the conference voice.
[0104] The online education system provided by the embodiment of the present application collects multi-channel online education voice data through a microphone array; estimates the voice mask of the noisy online education voice based on the spectral feature data of the multi-channel online education voice data through a mask estimation model; generates enhanced online education voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced online education voice data. By adopting this processing method, when estimating the voice mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization ability of the features that need to be restored to the target voice can be enhanced, which can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with unsatisfactory sound reception, thereby effectively suppressing background noise interference and improving the voice quality of online teaching.
[0105] The live broadcast system provided by the embodiment of the present application collects multi-channel live voice data through a microphone array; estimates the voice mask of the noisy live voice based on the spectral feature data of the multi-channel live voice data through a mask estimation model; generates enhanced live voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced live voice data. This processing method not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the voice mask, but also takes into account the correlation between the multi-channel spectral features, and can enhance the representation ability of the features that need to be restored to the target voice. This can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the quality of live voice.
[0106] The speech recognition text editing system provided in the embodiment of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of the noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; converts the enhanced speech data into text through the speech recognition model, and edits the transcribed text. By adopting this processing method, when estimating the speech mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization ability of the features that need to be restored to the target speech can be enhanced, which can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with unsatisfactory sound reception, thereby effectively suppressing background noise interference, improving the overall quality of the speech, and thereby improving the accuracy of speech recognition, and thereby improving the efficiency of speech recognition text editing.
[0107] The user recognition system provided in an embodiment of the present application collects multi-channel speech data using a microphone array; estimates the speech mask of noisy speech based on the spectral feature data of the multi-channel speech data using a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; and determines user information of the enhanced speech data using a user recognition model. This processing approach not only considers the correlation between the amplitude spectrum and phase spectrum when estimating the speech mask, but also the correlation between multi-channel spectral features, and enhances the representation capability of features that need to be restored to the target speech. This effectively improves the accuracy of speech masking, especially in far-field environments with poor sound reception. It can effectively suppress background noise interference, improve overall speech quality, and thus improve the accuracy of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] Figure 1 A schematic structural diagram of an embodiment of a speech enhancement system provided by this application;
[0109] Figure 2 Schematic diagram of application scenarios of an embodiment of the speech enhancement system provided in this application;
[0110] Figure 3 A schematic diagram of a mask estimation model of an embodiment of a speech enhancement system provided by the present application;
[0111] Figure 4 A schematic diagram of a plural domain attention module of an embodiment of a speech enhancement system provided by the present application;
[0112] Figure 5 A schematic diagram of a mask estimation model with residual connections in an embodiment of a speech enhancement system provided by the present application;
[0113] Figure 6 A flowchart of a speech enhancement process in an embodiment of a speech enhancement system provided by the present application;
[0114] Figure 7 A comparison diagram of spectral characteristics of an embodiment of the speech enhancement system provided by this application;
[0115] Figure 8 This application provides a flow chart of an embodiment of a speech enhancement method. DETAILED DESCRIPTION
[0116] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0117] This application provides a speech enhancement system, method, and apparatus, a mask estimation model method and apparatus, a conference system, an online education system, a live broadcast system, a speech recognition system, a speech recognition text editing system, a user identification system, and an electronic device. Each of these solutions is described in detail in the following embodiments.
[0118] First embodiment
[0119] The present invention provides a speech enhancement system that can separate target speech data from ambient noise. Figure 1 , which is a structural diagram of an embodiment of the speech enhancement system of the present application. The speech enhancement system provided in this embodiment includes: a speech acquisition module 1 and a speech enhancement module 2.
[0120] The voice acquisition module 1 and the voice enhancement module 2 can be deployed in different devices, such as the voice acquisition module 1 can be deployed on the client and the voice enhancement module 2 can be deployed on the server. The voice acquisition module 1 and the voice enhancement module 2 can also be deployed in the same device, such as the voice acquisition module 1 and the voice enhancement module 2 are both deployed in a smart speaker.
[0121] Please see Figure 2 , which is a schematic diagram of the application scenario of the embodiment of the voice enhancement system of the present application. In this embodiment, the voice acquisition module 1 is deployed on the client, and the voice enhancement module 2 is deployed on the server. The client includes but is not limited to smart speakers, and may also include terminal devices such as smart microphones, smart phones, personal computers, PADs, iPads, etc. The server can be a local area network server or a cloud server. Among them, the voice acquisition module 1 is used to collect multi-channel voice data through a microphone array, and the client sends the multi-channel voice data to the server; accordingly, the voice enhancement module 2 is used to determine the spectral feature data of the voice data of each channel; through the mask estimation model, the voice mask is determined according to the spectral feature data of the multi-channel voice data; based on the voice mask and the spectral feature data of the single-channel voice data, enhanced voice data is generated.
[0122] The speech enhancement system can be used for both far-field and near-field speech recognition. Far-field speech recognition occurs when the speaker is 3 to 5 meters from the microphone. Common scenarios include conference rooms, vehicles, and smart homes. In practical applications, if the speaker is not intentionally close to the microphone and is speaking naturally with the far-end microphone picking up their voice, this is considered far-field speech recognition.
[0123] like Figure 2As shown in the figure, in a far-field environment, there are several problems: (1) The distance is far and the signal-to-noise ratio is very low, resulting in poor sound reception; (2) At the same time, for example, in a closed or indoor environment, there will be a certain amount of reverberation, as well as various unidirectional noises such as background sound in the environment; (3) Furthermore, in a far-field recognition scenario, it means that there will be multiple people talking, such as in a conference room, in a car, or in a smart home in a house, so it is necessary to distinguish multiple target voice signals. This is often called the "cocktail party problem" or the "multi-source signal interference detection" problem; (4) Finally, there is the echo cancellation caused by the device. Echo refers to the sound played by the speaker that is transmitted back to the microphone. It can be seen that due to the low signal-to-noise ratio of the voice signal collected in the far-field environment, the existing single-channel voice enhancement processing technology cannot effectively separate the target voice data from the environmental noise.
[0124] The system provided in an embodiment of the present application collects multi-channel voice data in a far-field environment using a microphone array. Due to the presence of ambient noise, the collected multi-channel voice data may include noise data. The noisy voice data for any channel can be a one-dimensional array, the length of which can be determined by the audio length and the sampling rate. For example, a sampling rate Fs of 16 kHz indicates that 16,000 points are sampled in one second. If the audio length is 10 seconds, then the noisy voice data will contain 160,000 values, and the size of the values generally represents the amplitude.
[0125] The speech signal can be represented by various spectra in the time domain or frequency domain. In addition, the spectrogram (i.e., speech spectrum diagram, referred to as spectrogram or spectrogram) can simultaneously display information in the time domain and frequency domain. In this embodiment, the speech data collected by the microphone is a speech signal represented by a time domain waveform, and the spectral feature data of the speech data can be a spectrogram. The spectral feature data can be complex time-frequency domain feature data, also known as complex spectrum. Figure 2 As shown in the figure, the server converts the speech signal into a spectrogram through the spectrum feature extraction module. Figure 3 As shown, Er represents the real part of the spectrogram, and Ei represents the imaginary part of the spectrogram.
[0126] The spectral feature data can be the complex spectrum obtained by the short-time Fourier transform (STFT) of the speech waveform. STFT is a general tool for speech signal processing. It defines a time and frequency distribution class that specifies the complex amplitude of any signal varying with time and frequency. The process of calculating the short-time Fourier transform is to divide a longer time signal into shorter segments (speech frames) of the same length and calculate the Fourier transform, i.e., the Fourier spectrum, on each shorter segment. In other words, STFT is to perform a discrete Fourier transform (DCT) on a series of windowed data. The function of the Fourier transform is to convert the time domain signal into a frequency domain signal. By stacking the transformed frequency domain signals (spectrograms) of each frame in time, a spectrogram can be obtained. In specific implementation, the convolutional short-time Fourier transform (ConvSTFT) can be used to extract spectral features from the collected multi-channel speech data.
[0127] After extracting the spectral features from the noisy speech data of each channel, the spectral feature data of the multi-channel speech data can be used as the input data of the mask estimation model, and the speech mask can be determined through the mask estimation model. The speech mask (Mask, masking value or mask value) can also be called a time-frequency domain mask. The speech mask can define the ratio of clean speech signals to noisy speech signals in the time-frequency domain. The mask is used to filter the noise components of each frequency point in the spectrum. The mask can effectively retain the speech components and suppress the noise components in the time-frequency domain of noisy speech, thereby achieving the purpose of speech enhancement. The speech mask includes but is not limited to the Complex Ideal Ratio Mask (IRM) and the Complex Ideal Binary Mask (IBM). As Figure 3 As shown, the speech mask is a complex value, where Mr represents the real mask and Mi represents the imaginary mask.
[0128] The speech mask is the target data that needs to be learned and output by the mask estimation model. The mask estimation model can adopt a neural network structure (complex network) based on complex number calculation. The speech mask of the amplitude spectrum and phase spectrum of the noisy speech is estimated by a mask estimation model based on a complex value network, and the output data obtained from the mask estimation model are the real part and imaginary part of the complex value mask, respectively, which can improve the correlation between the amplitude spectrum and the phase spectrum. In the system provided in this embodiment, the speech mask is estimated based on the spectral feature data of the multi-channel speech data. When estimating the speech mask, the correlation between the multi-channel spectral features is also taken into account, and the characterization ability of the features that need to be restored to the target speech can be enhanced. This can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in the far-field environment, and performs well in suppressing background noise interference and improving the overall quality of the speech.
[0129] The mask estimation model can be based on a deep complex U-net (DCUnet). The mask estimation model based on DCUnet includes an encoder and a decoder. The encoder may include multiple encoding layers, and the decoder may include multiple decoding layers. The encoder may use a downsampling processing mechanism to extract high-order features, and the decoder may use upsampling to reconstruct the target spectrum (speech mask). In this embodiment, the encoder input is multi-channel data, such as Figure 3 The depth of the block. In the encoder, the depth of the block is increased through complex-valued convolution operations to extract high-level features. In the decoder, the depth of the block is reduced through complex-valued deconvolution operations to recover the feature representation of the target speech.
[0130] Different from the artificially designed spectral feature data, the spectral feature encoding data is an abstract data representation obtained by processing the spectral feature data through weighted operations of an encoder in a mask estimation model pre-trained based on prior knowledge (model training data). The spectral feature encoding data can be a high-order feature.
[0131] It should be noted that since the mask estimation model outputs a single-channel mask, in the last layer of complex-valued deconvolution operation (such as the complex-valued deconvolution operation before tanh), the convolution kernel of the complex-valued deconvolution operation is 2, and this output corresponds to the real and imaginary parts of the estimated complex mask, respectively. Therefore and The depth is 1, Figure 3 Shown in plan view.
[0132] like Figure 5As shown in the figure, in specific implementations, a long short-term memory (LSTM) layer can be added between the encoder and decoder to model temporal dependencies. This means that the encoding and decoding layers can use LSTM neural networks. This approach combines the decoding and encoding networks with the LSTM network to implement a deep complex-valued convolutional recurrent neural network. This allows for more accurate speech mask predictions based on contextual information (such as complex time-frequency feature data from adjacent frames).
[0133] like Figure 3 As shown in the figure, in one example, the mask estimation network is composed of an attention module and a DCUnet network. The encoder includes a channel attention module corresponding to the encoding layer, and each encoding layer may correspond to a channel attention module; the decoder may include a spectrum attention module (first spectrum attention module) corresponding to the decoding layer, and each decoding layer may correspond to a spectrum attention module. In this case, the encoder adopts the following encoding method:
[0134] Step S101: determining channel feature attention weights based on the multi-channel spectral feature data to be encoded through a channel attention module, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weights.
[0135] The complex domain channel attention module is used in the encoder part of the mask estimation network. Its main purpose is to better capture the correlation between multi-channel spectral features.
[0136] The input data of the channel attention module is called the multi-channel spectral feature data to be encoded. The input data of the channel attention module before the first coding layer can be the original spectral feature data of the multi-channel speech data. The input data of the channel attention module before subsequent coding layers is the spectral feature encoded data output by the adjacent previous coding layer. Different coding layers output spectral feature encoded data of different depth levels.
[0137] In practical applications, step S101 can be implemented as follows: the compression submodule in the channel attention module compresses the frequency and time dimensions of the multi-channel spectral feature data to be encoded as first compressed data; and the channel attention weight determination layer in the channel attention module determines the channel feature attention weights based on the first compressed data. This processing method determines the channel feature attention weights based on the compressed data, effectively improving speech enhancement performance.
[0138] like Figure 4As shown in Figure 1, the working principle of the channel attention module can be to first compress the frequency and time dimensions of the spectral features through complex-valued convolution operations and maximum pooling operations to obtain channel feature representations as the first compressed data (1*1*C); then further use the activation function of the neural network (such as the sigmoid function) on the channel feature representation to obtain the channel feature attention weights, which can be expressed as which channel feature representations need to be strengthened. By multiplying the channel feature attention weights with the multi-channel spectral features to be encoded point by point, the important feature representations between the channel spectral features can be obtained. Then, the important feature representations can be encoded through the encoding layer.
[0139] Step S102: encoding the adjusted multi-channel spectrum feature data to be encoded through the encoding layer.
[0140] Accordingly, the decoder adopts the following decoding method:
[0141] Step S201: Determine a first spectrum attention weight based on the multi-channel spectrum feature data to be decoded through a first spectrum attention module, and adjust the multi-channel spectrum feature data to be decoded based on the first spectrum feature attention weight.
[0142] The complex-domain spectral attention module is used in the decoder of the mask estimation network. Its main purpose is to enhance the feature representation of the frequency bins that need to be restored to the target speech. The complex-domain spectral attention module focuses on the feature representation of each frequency bin.
[0143] The input data of the first spectral attention module before the first decoding layer may include the spectral feature encoded data output by the last encoding layer. The input data of the first spectral attention module before each subsequent decoding layer may include the decoded data output by the immediately preceding decoding layer. Different decoding layers output speech mask decoded data at different depth levels.
[0144] In practical applications, step S103 can be implemented as follows: the compression submodule in the first spectral attention module compresses the channel dimension data of the multi-channel spectral feature data to be decoded as the second compressed data (F*T*1); the spectral attention weight determination layer in the first spectral attention module determines the first spectral attention weight based on the second compressed data. This processing method determines the first spectral attention weight based on the compressed data, thereby effectively improving the performance of speech enhancement.
[0145] like Figure 4As shown in Figure 2b, the working principle of the first spectral attention module can be to first compress the channel dimension of the spectral feature through operations such as complex-valued convolution to obtain the spectral feature of a single channel as the second compressed data. Then, a sigmoid function is applied to this feature to obtain the spectral attention weight. This weight can be used to represent which frequency points need to be restored to the target speech. This weight is multiplied point by point by the complex value of the multi-channel spectral feature data to be decoded to obtain the enhanced spectral feature, that is, the adjusted multi-channel spectral feature data to be decoded. The enhanced spectral feature can then be decoded by the decoding layer.
[0146] Step S202: decoding the adjusted multi-channel spectrum feature data to be decoded through a decoding layer.
[0147] like Figure 3 As shown, an activation function (such as tanh) can be used on the output of the decoder to enhance a [-1,1] boundary for the predicted mask. This makes it easier for the mask estimation network to converge during training.
[0148] The system provided in the embodiment of the present application can effectively capture the correlation between multi-channel features through the complex domain attention module (channel attention module, spectrum attention module), and enhance the representation ability of the features that need to be restored to the target speech.
[0149] like Figure 5 As shown, in another example, the mask estimation model is composed of an attention module and a DCCRN network, and the DCCRN network may include an encoder, a decoder, a skip connection (Skip Connection) and a long short-term memory model (LSTM). Among them, the encoding layer and the decoding layer may adopt a complex-valued convolutional network, that is, the network structure of the mask estimation model includes a complex-valued convolutional encoder and a complex-valued convolutional decoder. The skip connection includes a second spectral attention module. In this case, the speech mask is determined according to the spectral feature data of the multi-channel speech data through the mask estimation model, and the following steps may also be included:
[0150] Step S301: Determine a second spectral attention weight based on the multi-channel spectral feature encoding data output by the encoding layer through the second spectral attention module, and adjust the multi-channel spectral feature encoding data based on the second spectral feature attention weight.
[0151] In this case, the multi-channel spectrum feature data to be decoded at the next decoding layer can be determined based on the output data of the decoding layer corresponding to the coding layer and the adjusted multi-channel spectrum feature coding data, such as Figure 5 The structure and processing method of the second spectrum attention module are the same as those of the first spectrum attention module, so they will not be described here.
[0152] In this case, step S201 can be implemented in the following manner: through the first spectral attention module, based on the multi-channel spectral feature data to be decoded and the adjusted multi-channel spectral feature encoding data, a first spectral attention weight is determined, and based on the first spectral feature attention weight, the multi-channel spectral feature data to be decoded is adjusted.
[0153] Take the complex-valued deconvolution operation performed by the first decoding layer of the decoder as an example, that is, Figure 5 The first decoding operation after the LSTM layer. Figure 5 As shown in the figure, the output of the last encoding layer of the encoder is used as the input of the LSTM, and the LSTM will output a channel of features, which are distinguished by different grayscale colors in the figure. In the decoder, taking the first decoding layer as an example, the decoder input data (LSTM output features) and the output data of the corresponding encoding layer (the last encoding layer) (through the skip layer connection and the output features of the second spectral attention module) are superimposed on the channel dimension of the feature to form a large feature, as shown in the figure. Figure 5 The concatenated features are then passed through the first spectral attention module and the complex-valued deconvolution operation to obtain the output of the first layer of the decoder, which is also the input of the second layer of the decoder.
[0154] The speech enhancement system provided in the embodiment of the present application determines, through the second spectral attention module, a second spectral attention weight based on the multi-channel spectral feature encoding data output by the encoding layer, and adjusts the multi-channel spectral feature encoding data according to the second spectral feature attention weight; and determines, through the first spectral attention module, a first spectral attention weight based on the multi-channel spectral feature data to be decoded and the adjusted multi-channel spectral feature encoding data, and adjusts the multi-channel spectral feature data to be decoded according to the first spectral feature attention weight. This processing method can further improve the accuracy of speech mask estimation.
[0155] In another example, the DCCRN network may include four encoding layers (complex domain convolution) and four decoding layers (complex domain deconvolution). In specific implementations, the size of the convolution kernel can also be reduced. This processing approach can effectively reduce the number of network parameters and computational complexity, thereby effectively improving the real-time performance of speech enhancement.
[0156] In specific implementation, the network structure of the mask estimation model is not limited to the above Figure 3 and Figure 5 The network structure shown can also be other network structures, such as the mask estimation model only includes any one or two modules among the above-mentioned channel attention module, the first spectrum feature attention module, and the second spectrum feature attention module.
[0157] For example, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder; the encoder includes a channel attention module corresponding to the encoding layer; the mask estimation model is used to determine the speech mask according to the spectral feature data of the multi-channel speech data, including: through the channel attention module, based on the multi-channel spectral feature data to be encoded, determining the channel feature attention weight, and adjusting the multi-channel spectral feature data to be encoded according to the channel feature attention weight; through the encoding layer, encoding the adjusted multi-channel spectral feature data to be encoded.
[0158] For another example, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder; the decoder includes a first spectral attention module corresponding to the decoding layer; the mask estimation model is used to determine the speech mask according to the spectral feature data of the multi-channel speech data, including: through the first spectral attention module, based on the multi-channel spectral feature data to be decoded, determining the first spectral attention weight, and adjusting the multi-channel spectral feature data to be decoded according to the first spectral feature attention weight; through the decoding layer, the adjusted multi-channel spectral feature data to be decoded is decoded.
[0159] For another example, the mask estimation model also includes a skip-layer connection between the encoding layer and the decoding layer, and the skip-layer connection includes a second spectral attention module; the method of determining the speech mask based on the spectral feature data of the multi-channel speech data through the mask estimation model also includes: determining the second spectral attention weight through the second spectral attention module based on the multi-channel spectral feature encoding data output by the encoding layer, and adjusting the multi-channel spectral feature encoding data based on the second spectral feature attention weight; determining the multi-channel spectral feature data to be decoded of the next decoding layer based on the output data of the decoding layer corresponding to the encoding layer and the adjusted multi-channel spectral feature encoding data.
[0160] After the mask estimation model determines the speech mask based on the spectral feature data of the multi-channel speech data, the clean speech can be reconstructed based on the speech mask and the noisy speech. In specific implementations, the step of reconstructing the clean speech based on the speech mask and the noisy speech may include the following sub-steps:
[0161] Step S301: generating enhanced spectral feature data of the single-channel speech data according to the speech mask and the spectral feature data of the single-channel speech data.
[0162] The result of the operation (such as the product, n-fold product, etc.) of the spectral feature data of the single-channel speech data and the speech mask is used as the enhanced spectral feature data. The enhanced spectral feature data can be a spectrogram of the enhanced clean speech. In this embodiment, after the real mask and imaginary mask of the spectrum are predicted by the mask estimation model, the predicted speech mask can be multiplied point by point with a complex value by the original spectrum extracted from the collected speech data to obtain the output enhanced spectrum.
[0163] In a specific implementation, before the calculation result of the spectrum feature data of the single-channel speech data and the speech mask is used as the enhanced spectrum feature data, the speech mask may be normalized, which can further improve the accuracy of the enhanced spectrum feature data.
[0164] In one example, the spectrum feature data of the single-channel speech data with the highest signal-to-noise ratio is selected from the spectrum feature data of the multi-channel speech data. Figure 6 As shown, considering the spatial information contained in multi-channel speech data, differential beamforming can be used to select the spectral feature data of single-channel speech data with a high signal-to-noise ratio from the spectral feature data of the multi-channel speech data. In specific implementation, the selected spectral feature with a high signal-to-noise ratio can be multiplied point-by-point by the predicted speech mask using complex values to obtain enhanced spectral features.
[0165] In practical applications, depending on the environment, the space where sound needs to be picked up can be divided into n directions. For example, for a linear microphone array, if sound is picked up in a direction perpendicular to the linear array (the broadside direction) with a pickup range of 180 degrees, then n is less than or equal to 180. Differential beamforming technology is used to pick up sound in each direction. Differential beamforming inputs multi-channel data and outputs single-channel data after the sound is picked up. Because the target sound source exists at a certain location in space, differential beamforming outputs a higher signal-to-noise ratio (SNR) in the direction of the target sound source. The differential beamforming output with the highest SNR is selected as the target output and point-wise multiplied by the estimated speech mask. The spectrum selected by differential beamforming has a higher SNR, so point-by-point multiplication with the predicted speech mask further suppresses noise and improves speech quality.
[0166] like Figure 7 As shown in Figure 2, the spectral signal-to-noise ratio of the output using differential beamforming technology is higher. It should be noted that the input data of the mask estimation model is the spectral features of multi-channel speech data. For demonstration purposes, only the spectral features of the first channel (the original speech waveform) are selected here.
[0167] Step S303: generating the enhanced speech data according to the enhanced spectrum feature data.
[0168] The enhanced spectral feature data is converted into waveform speech to obtain enhanced speech data. For example, an inverse short-time Fourier transform (e.g., a convolutional inverse short-time Fourier transform (ConviSTFT)) can be used to output the enhanced speech data. Since mask-based speech separation is a relatively mature existing technology, it will not be described in detail here.
[0169] The mask estimation model is constructed using a supervised machine learning algorithm, hoping to learn a set of parameters from a large amount of noisy speech data to filter out noise and obtain clean speech. In specific implementation, the mask estimation model can be constructed using the following steps: 1) determining a training dataset, which includes spectral feature data and speech masks of multi-channel noisy speech data; 2) constructing a network structure for the mask estimation model based on complex-valued computation; and 3) training the network parameters of the mask estimation model based on the training dataset. Since supervised machine learning algorithms are relatively mature existing technologies, they will not be discussed in detail here.
[0170] like Figure 6 As shown, in this embodiment, different processing methods are used to determine the single-channel speech data used to generate enhanced speech in the model training and testing stages. The purpose of the training stage is to obtain a mask that can filter noise components at each frequency point. In the testing stage, differential beamforming technology is used to obtain a spectrum with a higher SNR value, and then the mask output by the trained mask estimation network is applied to the spectrum to obtain enhanced spectral features. In the training stage, the spectral features of any channel extracted by ConvSTFT are selected, such as the spectral features of the first channel, and the spectral features of the selected channel are multiplied by the mask predicted by the mask estimation network to generate enhanced spectral features. In the testing stage, a spectral feature with a higher signal-to-noise ratio selected by differential beamforming technology is multiplied by the predicted mask.
[0171] As can be seen from the above embodiments, the speech enhancement system provided by the embodiments of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of the noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; and generates enhanced speech data based on the speech mask and the spectral feature data of any channel. By adopting this processing method, when estimating the speech mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization capability of the features that need to be restored to the target speech can be enhanced. This can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the overall speech quality.
[0172] Second embodiment
[0173] This application provides a method for speech enhancement, the execution subject of which includes but is not limited to terminal devices such as smart speakers, smart pickups, etc. Figure 8 , which is a flow chart of an embodiment of the speech enhancement method of the present application. In this embodiment, the method may include the following steps:
[0174] Step S801: Collect multi-channel voice data.
[0175] Step S803: Determine a speech mask based on the spectral feature data of the multi-channel speech data through a mask estimation model.
[0176] In one example, the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder; the encoder includes a channel attention module corresponding to the encoding layer; the decoder includes a first spectrum attention module corresponding to the decoding layer; the encoder part may include the following sub-steps:
[0177] Step S901: determining, by a channel attention module, channel feature attention weights according to the multi-channel spectral feature data to be encoded, and adjusting the multi-channel spectral feature data to be encoded according to the channel feature attention weights;
[0178] Step S902: encoding the adjusted multi-channel spectrum feature data to be encoded through the encoding layer;
[0179] Accordingly, the decoder part may include the following sub-steps:
[0180] Step S1001: determining, by a first spectrum attention module, a first spectrum attention weight according to the multi-channel spectrum feature data to be decoded, and adjusting the multi-channel spectrum feature data to be decoded according to the first spectrum feature attention weight;
[0181] Step S1002: decoding the adjusted multi-channel spectrum feature data to be decoded through a decoding layer.
[0182] In one example, the mask estimation model may further include a skip-layer connection between the encoding layer and the decoding layer, wherein the skip-layer connection includes a second spectral attention module; and determining the speech mask based on the spectral feature data of the multi-channel speech data by the mask estimation model may further include the following sub-steps:
[0183] Step S1101: determining, by a second spectral attention module, a second spectral attention weight based on the multi-channel spectral feature encoded data output by the encoding layer, and adjusting the multi-channel spectral feature encoded data based on the second spectral feature attention weight;
[0184] Step S1102: determining the multi-channel spectrum feature data to be decoded of the next decoding layer according to the output data of the decoding layer corresponding to the coding layer and the adjusted multi-channel spectrum feature coding data.
[0185] Step S805: generating enhanced speech data according to the speech mask and the spectrum feature data of the single-channel speech data.
[0186] In one example, the frequency spectrum feature data of the single-channel speech data can be determined in the following manner: the frequency spectrum feature data of the single-channel speech data with a high signal-to-noise ratio is selected from the frequency spectrum feature data of the multi-channel speech data.
[0187] Third embodiment
[0188] In the above-mentioned embodiments, a speech enhancement method is provided. Correspondingly, this application also provides a speech enhancement device. This device corresponds to the above-mentioned method embodiment. Since the device embodiment is substantially similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the description of the method embodiment. The device embodiment described below is merely illustrative.
[0189] The present application further provides a speech enhancement device, comprising:
[0190] A data acquisition unit, used for collecting multi-channel voice data;
[0191] A speech mask prediction unit, configured to determine a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model;
[0192] The enhanced speech generation unit is used to generate enhanced speech data according to the speech mask and the spectrum feature data of the single-channel speech data.
[0193] Fourth embodiment
[0194] The present application provides a speech enhancement method, the execution subject of which is a speech enhancement device, which is usually deployed on the server side, but is not limited to the server side, and can also be any device that can implement the speech enhancement method.
[0195] The speech enhancement method of the present application may include the following steps:
[0196] Step 1: Receive multi-channel voice data;
[0197] Step 2: Determine the spectrum feature data of the multi-channel speech data;
[0198] Step 3: Determine the speech mask based on the spectral feature data of the multi-channel speech data through the mask estimation model;
[0199] Step 4: Generate enhanced speech data based on the speech mask and the spectral feature data of the single-channel speech data.
[0200] Fifth embodiment
[0201] In the above-mentioned embodiments, a speech enhancement method is provided. Correspondingly, this application also provides a speech enhancement device. This device corresponds to the above-mentioned method embodiment. Since the device embodiment is substantially similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the description of the method embodiment. The device embodiment described below is merely illustrative.
[0202] The present application further provides a speech enhancement device, comprising:
[0203] A data receiving unit, configured to receive multi-channel voice data;
[0204] A feature extraction unit, configured to determine spectral feature data of multi-channel speech data;
[0205] A mask estimation unit, configured to determine a speech mask based on spectral feature data of the multi-channel speech data using a mask estimation model;
[0206] The data generating unit is used to generate enhanced speech data according to the speech mask and the spectrum feature data of the single-channel speech data.
[0207] Sixth embodiment
[0208] This application provides a method for speech enhancement, which is performed by a speech acquisition device, which is usually deployed in a terminal device such as a smart speaker, a smart pickup, etc. The speech enhancement method of this application may include the following steps:
[0209] Step 1: Collect multi-channel speech data through a microphone array.
[0210] For example, multi-channel voice data is collected through the client (such as smart speakers, smart microphones, etc.).
[0211] Step 2: Send the multi-channel voice data so that the receiver can determine the spectral feature data of the multi-channel voice data; determine the voice mask based on the spectral feature data of the multi-channel voice data through the mask estimation model; and generate enhanced voice data based on the voice mask and the spectral feature data of the single-channel voice data.
[0212] For example, sending multi-channel voice data to the server.
[0213] Seventh embodiment
[0214] In the above-mentioned embodiments, a speech enhancement method is provided. Correspondingly, this application also provides a speech enhancement device. This device corresponds to the above-mentioned method embodiment. Since the device embodiment is substantially similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the description of the method embodiment. The device embodiment described below is merely illustrative.
[0215] The present application further provides a speech enhancement device, comprising:
[0216] A data acquisition unit, configured to acquire multi-channel voice data through a microphone array;
[0217] A data sending unit is used to send the multi-channel voice data so that a receiver can determine the spectral feature data of the multi-channel voice data; determine a voice mask based on the spectral feature data of the multi-channel voice data through a mask estimation model; and generate enhanced voice data based on the voice mask and the spectral feature data of the single-channel voice data.
[0218] Eighth embodiment
[0219] Corresponding to the above-mentioned speech enhancement method, the present application also provides a method for constructing a mask estimation model, the execution subject of which includes but is not limited to: a server. The parts of this embodiment that are the same as those in the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment. The speech enhancement model construction method provided in the present application includes:
[0220] Step 1: Determine a training data set; the training data includes: spectral feature data and speech mask of multi-channel speech data.
[0221] Step 2: Construct the network structure of the mask estimation model based on the complex-valued network.
[0222] The mask estimation model can be based on a Deep Complex U-net (DCUnet). The DCUnet-based speech mask prediction model includes an encoder and a decoder. The encoder can use a downsampling mechanism to extract high-order features, and the decoder can use upsampling to reconstruct the target spectrum.
[0223] In specific implementations, a long short-term memory (LSTM) layer can be added between the encoder and decoder to model temporal dependencies. This means that the encoding and decoding layers can employ LSTM neural networks. This approach combines the decoding and encoding networks with the LSTM network to implement a deep, complex-valued convolutional recurrent neural network. This allows for more accurate speech mask predictions based on contextual information (such as complex time-frequency feature data from adjacent frames).
[0224] Regarding the network structure of the model, please refer to the network structure description of the mask estimation model in Example 1, which will not be repeated here.
[0225] Step 3: Based on the training data set, train the network parameters of the mask estimation model.
[0226] As can be seen from the above embodiments, the mask estimation model processing method provided in the embodiments of the present application determines a training data set, wherein the training data includes: spectral feature data of multi-channel speech data and speech mask; constructs a network structure of a mask estimation model based on a complex-valued network; and trains the network parameters of the mask estimation model according to the training data set. By adopting this processing method, the constructed mask estimation model not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the speech mask, but also takes into account the correlation between multi-channel spectral features, and can enhance the characterization capability of the features that need to be restored to the target speech. This can effectively improve the accuracy of speech masking, especially the accuracy of speech masking in far-field environments with poor reception, thereby effectively suppressing background noise interference and improving the overall speech quality.
[0227] Ninth embodiment
[0228] In the above-mentioned embodiments, a mask estimation model processing method is provided. Correspondingly, this application also provides a speech mask prediction model processing device. This device corresponds to the above-mentioned method embodiment. Since the device embodiment is substantially similar to the method embodiment, the description is relatively brief. For relevant details, please refer to the description of the method embodiment. The device embodiment described below is merely illustrative.
[0229] The present application further provides a speech mask prediction model processing device, comprising:
[0230] A training data determination unit is used to determine a training data set, wherein the training data includes: complex time-frequency domain feature data including an amplitude spectrum and a phase spectrum of noisy speech data, and a speech mask of the time-frequency domain feature data;
[0231] A network construction unit, used to construct a network structure of a speech mask prediction model based on a complex-valued network;
[0232] A network training unit is used to train network parameters of the speech mask prediction model based on the training data set.
[0233] Tenth embodiment
[0234] In the above-mentioned embodiments, a speech enhancement method and a mask estimation model processing method are provided. Correspondingly, this application also provides an electronic device. This device corresponds to the above-mentioned method embodiments. Since the device embodiments are substantially similar to the method embodiments, the description is relatively simple. For relevant details, please refer to the description of the method embodiments. The device embodiments described below are merely illustrative.
[0235] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing any of the above methods. When the device is powered on, the program of the method is run by the processor.
[0236] Eleventh embodiment
[0237] Corresponding to the above-mentioned speech enhancement system, the present application also provides a speech recognition system. The parts of this embodiment that are identical to the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment. The speech recognition system provided in the present application includes: a speech acquisition module, a speech enhancement module, and a speech recognition module.
[0238] Among them, the voice acquisition module is used to collect multi-channel voice data through a microphone array; the voice enhancement module is used to determine the spectral feature data of the multi-channel voice data; through the mask estimation model, the voice mask is determined according to the spectral feature data of the multi-channel voice data; based on the voice mask and the spectral feature data of the single-channel voice data, enhanced voice data is generated; the voice recognition module is used to convert the enhanced voice data into text through the voice recognition model.
[0239] The voice acquisition module, voice enhancement module, and voice recognition module can be deployed in different devices, such as the voice acquisition module can be deployed on the client, and the voice enhancement module and voice recognition module can be deployed on the server. The voice acquisition module, voice enhancement module, and voice recognition module can also be deployed in the same device, such as the voice acquisition module, voice enhancement module, and voice recognition module are all deployed in a smart speaker.
[0240] The client includes but is not limited to smart speakers, and may also include terminal devices such as smart microphones, smart phones, personal computers, PADs, iPads, etc. The server may be a local area network server or a cloud server.
[0241] As can be seen from the above embodiments, the speech recognition system provided by the embodiments of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of the noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; and converts the enhanced speech data into text through the speech recognition model. This processing method not only takes into account the correlation between the amplitude spectrum and the phase spectrum when estimating the speech mask, but also takes into account the correlation between the multi-channel spectral features, and can enhance the representation ability of the features that need to be restored to the target speech. This can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference, improving the overall speech quality, and thus improving the accuracy of speech recognition.
[0242] Twelfth embodiment
[0243] Corresponding to the above-mentioned voice enhancement system, the present application also provides a conference system. The parts of this embodiment that are identical to the first embodiment are not repeated here; please refer to the corresponding parts in the first embodiment. The conference system provided in the present application includes: a voice acquisition module, a voice enhancement module, and a voice playback module. The conference system can be a video conferencing system or a telephone conferencing system.
[0244] Among them, the voice acquisition module is used to collect multi-channel conference voice data through a microphone array; the voice enhancement module is used to determine the spectral feature data of the multi-channel conference voice data; through the mask estimation model, the voice mask is determined according to the spectral feature data of the multi-channel conference voice data; based on the voice mask and the spectral feature data of the single-channel conference voice data, enhanced conference voice data is generated; the voice playback module is used to play the enhanced conference voice data.
[0245] In one example, the conference system is a video conferencing system, the voice collection module can be deployed on the first client, the voice enhancement module can be deployed on the server, and the voice playback module can be deployed on the second client. For example, multiple users conduct an online video conference through the DingTalk video conferencing system. The first client of the video conferencing system (such as user A's client) can collect the video conferencing voice data of the first client user and send the voice data to the video conferencing server. The video conferencing server first performs voice enhancement processing on the voice data to obtain enhanced voice data of the voice data. The enhanced voice data is then sent to the second client of the video conferencing (such as the client of user B and user X). The second client plays the enhanced voice data so that the second client user can hear the voice of the first client user after noise removal.
[0246] As can be seen from the above embodiments, the conference system provided by the embodiments of the present application collects multi-channel conference voice data through a microphone array; estimates the voice mask of the noisy conference voice based on the spectral feature data of the multi-channel conference voice data through a mask estimation model; generates enhanced conference voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced conference voice data. By adopting this processing method, when estimating the voice mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization capability of the features that need to be restored to the target voice can be enhanced, which can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with poor reception, thereby effectively suppressing background noise interference and improving the overall quality of the conference voice.
[0247] Thirteenth embodiment
[0248] Corresponding to the above-mentioned speech enhancement method, the present application also provides an online education system. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in Example 1. The online education system provided by the present application includes: a speech acquisition module, a speech enhancement module, and a speech playback module.
[0249] Among them, the voice acquisition module is used to collect multi-channel online education voice data through a microphone array; the voice enhancement module is used to determine the spectral feature data of the multi-channel online education voice data; through the mask estimation model, the voice mask is determined according to the spectral feature data of the multi-channel online education voice data; based on the voice mask and the spectral feature data of the single-channel online education voice data, enhanced online education voice data is generated; the voice playback module is used to play the enhanced online education voice data.
[0250] In one example, the voice collection module can be deployed on the first client, the voice enhancement module can be deployed on the server, and the voice playback module can be deployed on the second client. When a user of the first client conducts online education activities through the online education system, the first client can collect the first client's online education voice data and send the voice data to the online education server. The online education server first performs voice enhancement processing on the voice data to obtain enhanced voice data of the voice data. The enhanced voice data is then sent to the second online education client (such as the client of user B and user X). The second client plays the enhanced voice data, allowing the second client user to hear the first client's online education voice after noise removal.
[0251] As can be seen from the above embodiments, the online education system provided by the embodiments of the present application collects multi-channel live voice data through a microphone array; estimates the voice mask of the noisy live voice based on the spectral feature data of the multi-channel live voice data through a mask estimation model; generates enhanced live voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced live voice data. By adopting this processing method, when estimating the voice mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization ability of the features that need to be restored to the target voice can be enhanced. This can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the quality of live voice.
[0252] Fourteenth embodiment
[0253] Corresponding to the above-mentioned voice enhancement system, the present application also provides a live broadcast system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in Example 1. The online education system provided by the present application includes: a voice acquisition module, a voice enhancement module, and a voice playback module.
[0254] Among them, the voice acquisition module is used to collect multi-channel live voice data through a microphone array; the voice enhancement module is used to determine the spectral feature data of the multi-channel live voice data; through the mask estimation model, the voice mask is determined according to the spectral feature data of the multi-channel live voice data; based on the voice mask and the spectral feature data of the single-channel live voice data, enhanced live voice data is generated; the voice playback module is used to play the enhanced live voice data.
[0255] In one example, a voice acquisition module can be deployed on a first client, a voice enhancement module on a server, and a voice playback module on a second client. A livestreamer selling goods through a livestream can use the first client to collect the livestreamer's voice data and send it to the livestreaming server. The livestreaming server then performs voice enhancement processing on the voice data to obtain enhanced voice data. The enhanced voice data is then sent to a second client watching the livestream, which plays the enhanced voice data, allowing viewers to hear the livestreamer's voice after removing noise.
[0256] As can be seen from the above embodiments, the live broadcast system provided by the embodiments of the present application collects multi-channel live voice data through a microphone array; estimates the voice mask of the noisy live voice based on the spectral feature data of the multi-channel live voice data through a mask estimation model; generates enhanced live voice data based on the voice mask and the spectral feature data of any channel; and plays the enhanced live voice data. By adopting this processing method, when estimating the voice mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization capability of the features that need to be restored to the target voice can be enhanced. This can effectively improve the accuracy of the voice mask, especially the accuracy of the voice mask in a far-field environment with poor sound reception, thereby effectively suppressing background noise interference and improving the quality of live voice.
[0257] Fifteenth embodiment
[0258] Corresponding to the above-mentioned speech enhancement system, the present application also provides a speech recognition text editing system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in Example 1. The speech recognition text editing system provided in the present application includes: a client and a server.
[0259] Among them, the client is used to collect multi-channel voice data through a microphone array and send the multi-channel voice data to the server; and edit the voice transcription text recognized by the server; the server is used to estimate the voice mask of noisy speech based on the spectral feature data of the multi-channel voice data through a mask estimation model; generate enhanced voice data based on the voice mask and the spectral feature data of any channel; and convert the enhanced voice data into text through the voice recognition model.
[0260] Since speech recognition and editing of speech-transcribed text are relatively mature existing technologies, they will not be described in detail here.
[0261] As can be seen from the above embodiments, the speech recognition text editing system provided by the embodiments of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of the noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; and converts the enhanced speech data into text through a speech recognition model, and edits the transcribed text. By adopting this processing method, when estimating the speech mask, not only the correlation between the amplitude spectrum and the phase spectrum is taken into account, but also the correlation between the multi-channel spectral features, and the characterization ability of the features that need to be restored to the target speech can be enhanced, which can effectively improve the accuracy of the speech mask, especially the accuracy of the speech mask in a far-field environment with unsatisfactory sound reception, thereby effectively suppressing background noise interference, improving the overall quality of the speech, and thereby improving the accuracy of speech recognition, and thereby improving the efficiency of speech recognition text editing.
[0262] Sixteenth embodiment
[0263] Corresponding to the above-mentioned speech enhancement system, the present application also provides a user identification system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment. The speech recognition text editing system provided by the present application includes: a client and a server.
[0264] Among them, the client is used to collect multi-channel voice data through a microphone array and send the multi-channel voice data to the server; the server is used to estimate the voice mask of noisy speech based on the spectral feature data of the multi-channel voice data through a mask estimation model; generate enhanced voice data based on the voice mask and the spectral feature data of any channel; and determine the user information of the enhanced voice data through a user recognition model.
[0265] Since user identification is a relatively mature existing technology, it will not be described here in detail.
[0266] As can be seen from the above embodiments, the user recognition system provided by the embodiments of the present application collects multi-channel speech data through a microphone array; estimates the speech mask of noisy speech based on the spectral feature data of the multi-channel speech data through a mask estimation model; generates enhanced speech data based on the speech mask and the spectral feature data of any channel; and determines the user information of the enhanced speech data through a user recognition model. This processing method not only considers the correlation between the amplitude spectrum and phase spectrum when estimating the speech mask, but also considers the correlation between the multi-channel spectral features, and can enhance the representation capability of the features that need to be restored to the target speech. This can effectively improve the accuracy of speech masking, especially in far-field environments with poor sound reception, thereby effectively suppressing background noise interference, improving the overall speech quality, and thus improving the accuracy of speaker recognition.
[0267] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0268] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0269] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0270] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0271] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A speech enhancement system, characterized in that: include: A voice acquisition module is used to collect multi-channel voice data through a microphone array; A speech enhancement module, used to determine spectral feature data of multi-channel speech data; Determine a single-channel speech mask based on spectral feature data of multi-channel speech data through a mask estimation model; generate enhanced speech data based on the speech mask and the spectral feature data of the single-channel speech data; the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder, the encoder including a channel attention module corresponding to the encoding layer, and the decoder including a first spectral attention module corresponding to the decoding layer; Determining a speech mask of a single channel according to the spectral feature data of the multi-channel speech data by using a mask estimation model includes: determining a channel feature attention weight according to the multi-channel spectral feature data to be encoded by using a channel attention module, and adjusting the multi-channel spectral feature data to be encoded according to the channel feature attention weight; A first spectrum attention module is used to determine a first spectrum attention weight based on the multi-channel spectrum feature data to be decoded, and the multi-channel spectrum feature data to be decoded is adjusted based on the first spectrum attention weight.
2. A speech enhancement method, characterized in that: include: Collect multi-channel voice data; Determine the speech mask of a single channel based on the spectral feature data of the multi-channel speech data through a mask estimation model; generating enhanced speech data according to the speech mask and the spectrum feature data of the single-channel speech data; The mask estimation model is a model based on complex domain calculation, including an encoder and a decoder; The encoder includes a channel attention module corresponding to the encoding layer; The decoder includes a first spectral attention module corresponding to the decoding layer; The method of determining a single-channel speech mask based on spectral feature data of multi-channel speech data by using a mask estimation model includes: Determining, by a channel attention module, channel feature attention weights based on the multi-channel spectral feature data to be encoded, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weights; A first spectrum attention module is used to determine a first spectrum attention weight based on the multi-channel spectrum feature data to be decoded, and the multi-channel spectrum feature data to be decoded is adjusted based on the first spectrum attention weight.
3. The method according to claim 2, characterized in that The method further comprises determining a speech mask of a single channel based on the spectral feature data of the multi-channel speech data by using a mask estimation model: Encoding the adjusted multi-channel spectrum feature data to be encoded through the encoding layer; The adjusted multi-channel spectrum feature data to be decoded is decoded through the decoding layer.
4. The method according to claim 3, characterized in that The mask estimation model further includes a skip-layer connection between the encoding layer and the decoding layer, wherein the skip-layer connection includes a second spectral attention module; The method further comprises determining a speech mask of a single channel based on the spectral feature data of the multi-channel speech data by using a mask estimation model: determining, by a second spectral attention module, a second spectral attention weight based on the multi-channel spectral feature encoded data output by the encoding layer, and adjusting the multi-channel spectral feature encoded data based on the second spectral feature attention weight; The multi-channel spectrum feature data to be decoded of the next decoding layer is determined according to the output data of the decoding layer corresponding to the encoding layer and the adjusted multi-channel spectrum feature encoding data.
5. The method according to claim 2, characterized in that The frequency spectrum feature data of the single-channel speech data is determined in the following manner: The frequency spectrum feature data of the single-channel speech data with a high signal-to-noise ratio is selected from the frequency spectrum feature data of the multi-channel speech data.
6. A speech enhancement method, used in the speech acquisition module of the system according to claim 1, characterized in that: include: Collect multi-channel voice data through microphone array; Sending the multi-channel voice data so that a recipient determines frequency spectrum feature data of the multi-channel voice data; The mask estimation model is used to determine a speech mask based on the spectral feature data of the multi-channel speech data; and enhanced speech data is generated based on the speech mask and the spectral feature data of the single-channel speech data.
7. A speech enhancement method, used in the speech enhancement module of the system according to claim 1, characterized in that: include: Receive multi-channel voice data; Determining spectral feature data of multi-channel speech data; Determining a speech mask based on spectral feature data of multi-channel speech data using a mask estimation model; Enhanced speech data is generated according to the speech mask and the spectrum feature data of the single-channel speech data.
8. A method for processing a mask estimation model, wherein the mask estimation model is the mask estimation model included in the system of claim 1, characterized in that: include: Determine a training data set, wherein the training data includes: spectral feature data and a speech mask of multi-channel speech data; Construct the network structure of the mask estimation model based on complex-valued networks; The network parameters of the mask estimation model are trained according to the training data set.
9. A speech recognition system, characterized in that: include: A voice acquisition module is used to collect multi-channel voice data through a microphone array; A speech enhancement module, used to determine spectral feature data of multi-channel speech data; Determine a single-channel speech mask based on spectral feature data of multi-channel speech data through a mask estimation model; generate enhanced speech data based on the speech mask and the spectral feature data of the single-channel speech data; the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder, the encoder including a channel attention module corresponding to the encoding layer, and the decoder including a first spectral attention module corresponding to the decoding layer; Determining a speech mask for a single channel based on spectral feature data of multi-channel speech data through a mask estimation model, including: determining a channel feature attention weight based on the multi-channel spectral feature data to be encoded through a channel attention module, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weight; determining a first spectral attention weight based on the multi-channel spectral feature data to be decoded through a first spectral attention module, and adjusting the multi-channel spectral feature data to be decoded based on the first spectral attention weight; The speech recognition module is used to convert enhanced speech data into text through a speech recognition model.
10. A conference system, characterized in that: include: Voice collection module, used to collect multi-channel conference voice data through microphone array; A speech enhancement module, used to determine the spectral feature data of multi-channel conference speech data; Determine a speech mask based on spectral feature data of multi-channel conference speech data through a mask estimation model; generate enhanced conference speech data based on the speech mask and spectral feature data of single-channel conference speech data; the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder, the encoder including a channel attention module corresponding to the encoding layer, and the decoder including a first spectral attention module corresponding to the decoding layer; Determining a speech mask for a single channel based on spectral feature data of multi-channel speech data through a mask estimation model, including: determining a channel feature attention weight based on the multi-channel spectral feature data to be encoded through a channel attention module, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weight; determining a first spectral attention weight based on the multi-channel spectral feature data to be decoded through a first spectral attention module, and adjusting the multi-channel spectral feature data to be decoded based on the first spectral attention weight; The voice playing module is used to play the enhanced conference voice data.
11. A live broadcast system, characterized in that: include: Voice acquisition module, used to collect multi-channel live voice data through microphone array; A voice enhancement module, used to determine the spectral feature data of multi-channel live voice data; Determine the speech mask based on the spectral feature data of the multi-channel live speech data through the mask estimation model; generate enhanced live speech data based on the speech mask and the spectral feature data of the single-channel live speech data; The mask estimation model is a model based on complex domain calculation, including an encoder and a decoder, the encoder including a channel attention module corresponding to the encoding layer, and the decoder including a first spectrum attention module corresponding to the decoding layer; Determining a speech mask for a single channel based on spectral feature data of multi-channel speech data through a mask estimation model, including: determining a channel feature attention weight based on the multi-channel spectral feature data to be encoded through a channel attention module, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weight; determining a first spectral attention weight based on the multi-channel spectral feature data to be decoded through a first spectral attention module, and adjusting the multi-channel spectral feature data to be decoded based on the first spectral attention weight; The voice playing module is used to play the enhanced live voice data.
12. An online education system, characterized in that: include: The voice acquisition module is used to collect multi-channel online education voice data through a microphone array; A speech enhancement module, used to determine the spectral feature data of multi-channel online education speech data; Determine a single-channel speech mask based on the spectral feature data of the multi-channel online education speech data through a mask estimation model; generate enhanced online education speech data based on the speech mask and the spectral feature data of the single-channel online education speech data; the mask estimation model is a model based on complex domain calculation, including an encoder and a decoder, the encoder including a channel attention module corresponding to the encoding layer, and the decoder including a first spectral attention module corresponding to the decoding layer; Determining a speech mask for a single channel based on spectral feature data of multi-channel speech data through a mask estimation model, including: determining a channel feature attention weight based on the multi-channel spectral feature data to be encoded through a channel attention module, and adjusting the multi-channel spectral feature data to be encoded based on the channel feature attention weight; determining a first spectral attention weight based on the multi-channel spectral feature data to be decoded through a first spectral attention module, and adjusting the multi-channel spectral feature data to be decoded based on the first spectral attention weight; The voice playing module is used to play the enhanced online education voice data.
Citation Information
Patent Citations
Voice enhancing method, device, intelligent voice equipment and computer equipment
CN108615535A
Voice processing method, voice processing device and device for processing voice
CN110808063A
Noise estimation method and system for far-field call
CN111696567A
Microphone array-oriented channel attention weighted speech enhancement method
CN112151059A
Video deblurring method based on spectrum attention and feature shift
CN118710542A