A Multi-Channel Speech Enhancement Method and System Based on Neural Network
Through the multi-channel speech enhancement method based on neural networks, the beam and adaptive noise cancellation layer are formed using filters, which solves the problem of traditional methods relying on scene assumptions and array spatial information, and achieves high accuracy and high definition speech enhancement.
Patent Information
- Application Number
- CN202210870606.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-07-22
AI Technical Summary
The traditional multi-channel speech enhancement method relies on scene assumptions and array space information, resulting in reduced algorithm performance and insufficient accuracy.
A multi-channel speech enhancement method based on neural network is adopted. By receiving voice signals from multiple channels, a filter is used to form a beam, the target beam and wave arrival direction is determined, and the adaptive noise cancellation layer is used to enhance it, and the neural network model is trained to improve accuracy.
Highly accurate speech enhancement that does not depend on scene assumptions and array space information is achieved, improving speech clarity and intelligibility.
Smart Images

Figure CN115240695B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of signal processing, and in particular relates to a multi-channel speech enhancement method and system based on a neural network. Background Art
[0002] With the rapid development of science and technology worldwide and the increasing maturity of artificial intelligence (AI), voice has become increasingly important as a component of human-computer interaction. Devices equipped with voice processing systems are ubiquitous in our daily lives, such as conference systems, telephone equipment, and autonomous vehicles. Applying deep neural networks to speech enhancement technology leverages the complex nonlinear mapping and learning capabilities of neural network structures. Using extensive experimental data and spectral features as supervisory information, effective speech enhancement systems can be trained, improving speech clarity and intelligibility.
[0003] Traditional multi-channel enhancement methods have solid data theory support and are simple to implement, but their algorithmic effects rely on prior information such as scene assumptions, array spatial information, and parameter estimation. In actual use, they often cannot obtain accurate results and often can only rely on estimation, resulting in a significant decline in algorithmic performance. Summary of the Invention
[0004] In response to the deficiencies in the prior art, the present invention provides a multi-channel speech enhancement method and system based on a neural network, which has high accuracy, does not require scene assumptions, and does not rely on prior information such as array spatial information and parameter estimation.
[0005] In a first aspect, a multi-channel speech enhancement method based on a neural network comprises:
[0006] Receive voice signals from multiple channels;
[0007] The voice signal of each channel is processed using the filter of each channel to obtain the beam of the corresponding angle of each channel;
[0008] Determine the target beam and direction of arrival based on all beams;
[0009] Obtaining multiple reference noises according to the speech signals and directions of arrival of multiple channels;
[0010] The reference noise and target beam are input into the adaptive denoising layer to enhance the target beam.
[0011] Furthermore, the voice signals of the channels are processed using filters of the channels to obtain beams of corresponding angles of the channels. Specifically, the following steps are performed:
[0012] Select a channel as the reference channel;
[0013] In each channel, a filter is used to perform a convolution operation on the speech signal to obtain a beam with the corresponding angle for each channel.
[0014] Furthermore, the filter optimization method includes:
[0015] Calculate the cosine similarity between the beam obtained from each channel and the beam obtained from the reference channel to obtain the similarity of each channel;
[0016] Input the similarity of all channels into the fully connected network to obtain the first output data;
[0017] Performing affine transformation on the filters of all channels to obtain second output data;
[0018] The first output data and the second output data are added together and then input into a mapping function to obtain filters optimized for all channels.
[0019] Furthermore, determining the target beam and direction of arrival based on all beams specifically includes:
[0020] Assign weights to the beams of each channel;
[0021] All beams are selected using the weights of all channels to obtain the target beam and the direction of arrival in the target direction.
[0022] Furthermore, the weight optimization method includes:
[0023] Performing an affine transformation on the weights of all channels to obtain third output data;
[0024] The first output data and the third output data are added together and then input into a mapping function to obtain optimized weights of all channels.
[0025] Furthermore, the adaptive denoising layer includes an encoder, a 1×1 convolutional layer, and a decoder;
[0026] The encoder is configured to process the subframe to obtain a fourth output, the subframe being obtained by dividing the reference noise;
[0027] The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features;
[0028] The decoder is used to enhance the amplitude spectrum of the target beam using the noise characteristics.
[0029] In a second aspect, a multi-channel speech enhancement system based on a neural network includes:
[0030] Input layer: used to receive speech signals from multiple channels;
[0031] Fixed beamforming layer: used to process the voice signal of each channel using the filter of each channel to obtain the beam of the corresponding angle of each channel;
[0032] Beam direction selection unit: used to determine the target beam and direction of arrival based on all beams;
[0033] Noise blocking layer: used to obtain multiple reference noises based on the speech signals and directions of arrival of multiple channels;
[0034] Adaptive noise cancellation layer: used to receive reference noise and target beam, and output the enhanced signal of the target beam.
[0035] Furthermore, the fixed beamforming layer is specifically used for:
[0036] Select a channel as the reference channel;
[0037] In each channel, a filter is used to perform a convolution operation on the speech signal to obtain a beam with the corresponding angle for each channel.
[0038] Furthermore, the beam direction selection unit is specifically configured to:
[0039] Assign weights to the beams of each channel;
[0040] All beams are selected using the weights of all channels to obtain the target beam and the direction of arrival in the target direction.
[0041] Furthermore, the adaptive denoising layer includes an encoder, a 1×1 convolutional layer, and a decoder;
[0042] The encoder is configured to process the subframe to obtain a fourth output, the subframe being obtained by dividing the reference noise;
[0043] The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features;
[0044] The decoder is used to enhance the amplitude spectrum of the target beam using the noise characteristics.
[0045] It can be seen from the above technical solution that the multi-channel speech enhancement method and system provided by the present invention trains a neural network model based on historical data, and uses the trained neural network model to enhance the speech signal. It has high accuracy, does not require scene assumptions, and does not rely on prior information such as array spatial information and parameter estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.
[0047] Figure 1 The present invention provides a flowchart of a multi-channel speech enhancement method.
[0048] Figure 2 A flowchart of a target beam determination method provided in an embodiment.
[0049] Figure 3 Flowchart of the filter and weight optimization method provided in the embodiment.
[0050] Figure 4 This is a module block diagram of a multi-channel speech enhancement system provided in an embodiment. DETAILED DESCRIPTION
[0051] The following embodiments of the technical solution of the present invention are described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore only examples and are not intended to limit the scope of protection of the present invention. It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by those skilled in the art to which the present invention belongs.
[0052] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0053] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0054] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0055] Example:
[0056] A multi-channel speech enhancement method based on neural network, see Figure 1 ,include:
[0057] S1: Receives voice signals from multiple channels;
[0058] S2: Use the filters of each channel to process the voice signal of the channel to obtain the beam of the corresponding angle of each channel;
[0059] S3: Determine the target beam and direction of arrival based on all beams;
[0060] S4: obtaining multiple reference noises based on the speech signals and directions of arrival of multiple channels;
[0061] S5: Input the reference noise and target beam into the adaptive denoising layer to enhance the target beam.
[0062] In this embodiment, the speech signal may be a signal obtained by performing a short-time Fourier transform on the original speech signal. The original speech signal may include a multi-channel (e.g., M channels) noisy signal received through a microphone array. The speech signals of different channels may have the same or different angles. After obtaining the multi-channel speech signal, the multi-channel speech enhancement method processes the speech signal to obtain beams at angles corresponding to each channel. For example, after processing the speech signal, the method obtains frequency domain beam features at 0 degrees to 180 degrees. When obtaining beams at various angles, the method screens all beams and selects a target beam in a target direction, wherein the target beam can be screened according to a set direction or parameter. The method can also obtain the direction of arrival (DOA) of the target beam.
[0063] In this embodiment, the method can also obtain multiple reference noises based on the speech signals and directions of arrival of multiple channels. For example, the method can obtain reference noises for M channels. Finally, the method inputs the M reference noises and the target beam into the adaptive noise cancellation layer, which enhances the target beam. The adaptive noise cancellation layer can be obtained by training a DNN acoustic model. After obtaining the enhanced signal of the target beam, the method can re-input the enhanced signal into the DNN acoustic model to optimize the DNN acoustic model, thereby optimizing the adaptive noise cancellation layer.
[0064] This multi-channel speech enhancement method trains a neural network model based on historical data and uses the trained neural network model to enhance the speech signal. It has high accuracy, does not require scene assumptions, and does not rely on prior information such as array spatial information and parameter estimation.
[0065] Further, in some embodiments, see Figure 2 , using the filters of each channel to process the voice signal of the channel to obtain the beam of the corresponding angle of each channel. Specifically include:
[0066] S11: Select a channel as a reference channel;
[0067] S12: In each channel, a convolution operation is performed on the speech signal using a filter to obtain a beam corresponding to an angle of each channel.
[0068] In this embodiment, the method can define any one of the M channels as a reference channel, and use the filters of the m channels to perform convolution operations on the speech signals of each channel. For example, when used for the first time, the method can initialize the filters of each channel to obtain initial values for each filter. In this way, when used for the first time, the speech signal is convolved using the initial values of the filters. In subsequent use, the method can optimize the filter values based on historical data to obtain optimized values for each filter. In this way, in subsequent use, the speech signal is convolved using the optimized values of the filters to improve accuracy.
[0069] In this embodiment, for example, the set of all channel filters is represented as filter=[h0,h1,···,h m ], where filter is the collection name, h m is the value of the filter for the mth channel. Then the method uses the filter to perform a convolution operation on the speech signal to obtain the beam corresponding to the angle of each channel. where y m is the beam obtained for the mth channel, is a convolution operation, X is the set of speech signals of each channel, X=[x0,x1,···,x m ], x m is the speech signal of the mth channel.
[0070] Further, in some embodiments, see Figure 3 , the filter optimization methods include:
[0071] S21: Calculate the cosine similarity between the beam obtained by each channel and the beam obtained by the reference channel to obtain the similarity of each channel;
[0072] S22: Input the similarities of all channels into the fully connected network to obtain first output data;
[0073] S23: performing affine transformation on the filters of all channels to obtain second output data;
[0074] S24: After adding the first output data and the second output data, the sum is input into a mapping function to obtain optimized filters for all channels.
[0075] In this embodiment, in order to improve the accuracy, the filter value can be optimized according to the results during use. When performing filter optimization, the method first calculates the cosine similarity between the beam obtained by each channel and the beam obtained by the reference channel to obtain the similarity of each channel, where the set of all similarities can be expressed as sim = [sim0, sim1, ···, sim m ], sim is the name of the collection, sim m is the similarity of the mth channel, sim m =cos_sim(y m ,ref_channel), ref_channel is the beam obtained by the reference channel, and cos_sim is the cosine similarity function. Then the method inputs the similarity sim of all channels into the fully connected network net to obtain the first output data delta. Then, the filters of all channels are affine transformed to obtain the second output data res1. After adding the first output data delta and the second output data res1, the two are input into the mapping function f(·) to obtain the optimized filters of all channels. This method can obtain different beamformers by selecting different mapping functions f(·). For example, when the selected mapping function f(·) is the softmax(·) function, the beamformer obtained is the FSB beamformer, and the sum of the filter coefficients of each channel is 1. It can be expressed as:
[0076] delta = net(sim);
[0077] res1 = σ(filter);
[0078] filter = f(res1 + delta);
[0079] In this embodiment, the filter optimization method described above can gradually align the signals of each channel with the reference channel in time through repeated calculations. This uses a stacked structure to improve filter stability, resulting in a stable final filter value. The filter optimization method can employ the following loss function: Loss = MSE(sim, 1), where 1 indicates the goal is to achieve the highest cosine similarity.
[0080] Further, in some embodiments, see Figure 2 , determining the target beam and direction of arrival based on all beams specifically includes:
[0081] S13: assign weights to the beams of each channel;
[0082] S14: All beams are selected using the weights of all channels to obtain a target beam and a direction of arrival in the target direction.
[0083] In this embodiment, to achieve a better enhancement effect, the method can use an additional network to estimate the weight of each channel separately based on the FSB beamformer, assign different weights to each channel, and use the weights of all channels to select all beams to obtain the target beam in the target direction and the direction of arrival. For example, when used for the first time, the method can initialize the weight of each channel to obtain the initial value of each weight. In this way, when used for the first time, all beams are selected using the initial value of the weight. In subsequent use, the method can optimize the value of the weight based on historical data to obtain the optimized value of each weight. In this way, in subsequent use, all beams are selected using the optimized value of the weight, thereby improving accuracy.
[0084] Further, in some embodiments, see Figure 3 , weight optimization methods include:
[0085] S25: performing affine transformation on the weights of all channels to obtain third output data;
[0086] S26: After adding the first output data and the third output data, the sum is input into a mapping function to obtain optimized weights of all channels.
[0087] In this embodiment, the weight optimization method is similar to the filter optimization method. When performing the weight optimization method, the method first calculates the cosine similarity between the beam obtained by each channel and the beam obtained by the reference channel to obtain the similarity of each channel. Then the method inputs the similarity sim of all channels into the fully connected network net to obtain the first output data delta. Then, the weights of all channels are affine transformed to obtain the third output data res2. After adding the first output data delta and the third output data res2, the result is input into the mapping function f(·) to obtain the optimized weight channel_weight of all channels, which is expressed as:
[0088] delta = net(sim);
[0089] res2 = σ(filter);
[0090] channel_weight=f(res2+delta);
[0091] In this embodiment, the weight optimization method may adopt the following loss function: In the formula represents the filtered output, and y represents the enhanced signal of the target beam.
[0092] Further, in some embodiments, the adaptive denoising layer includes an encoder, a 1×1 convolutional layer, and a decoder;
[0093] The encoder is configured to process the subframe to obtain a fourth output, the subframe being obtained by dividing the reference noise;
[0094] The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features;
[0095] The decoder is used to enhance the amplitude spectrum of the target beam using the noise characteristics.
[0096] In this embodiment, the adaptive noise cancellation layer utilizes information between multiple channels. A 1×1 convolutional layer is added to the network to fuse the information between channels. Transposed convolution is then used for upsampling to restore pure speech, which facilitates multi-channel speech enhancement. The adaptive noise cancellation layer primarily consists of an encoder, a 1×1 convolutional layer (feature extraction layer), and a decoder. The adaptive noise cancellation layer can also incorporate long-hop connections, allowing the decoder to fully utilize the information in the encoder while also improving network training efficiency.
[0097] In this embodiment, the adaptive noise cancellation layer first divides the reference noise of C channels into T subframes, each of which has a length of F, and then sends it to the encoder for processing. The output of the encoder is subjected to a 1×1 convolutional layer to extract features and then sent to the decoder. The corresponding layers in the encoder and decoder can use far-hop connections, so that the data can be spliced together in the channel dimension, and finally the enhanced single-channel pure speech is output, that is, the enhanced target beam.
[0098] In this embodiment, the encoder and decoder can be composed of multiple corresponding blocks. This method fuses the input data with channel features through a 1×1 convolutional layer before performing a two-dimensional convolution. The ELU activation function is used after each convolution to provide nonlinearity, and GroupNorm is used for normalization after the two-dimensional convolution.
[0099] A multi-channel speech enhancement system based on neural network, see Figure 4 ,include:
[0100] Input layer: used to receive speech signals from multiple channels;
[0101] Fixed beamforming layer: used to process the voice signal of each channel using the filter of each channel to obtain the beam of the corresponding angle of each channel;
[0102] Beam direction selection unit: used to determine the target beam and direction of arrival based on all beams;
[0103] Noise blocking layer: used to obtain multiple reference noises based on the speech signals and directions of arrival of multiple channels;
[0104] Adaptive noise cancellation layer: used to receive reference noise and target beam, and output the enhanced signal of the target beam.
[0105] Furthermore, in some embodiments, the fixed beamforming layer is specifically used to:
[0106] Select a channel as the reference channel;
[0107] In each channel, a filter is used to perform a convolution operation on the speech signal to obtain a beam with the corresponding angle for each channel.
[0108] Furthermore, in some embodiments, the beam direction selecting unit is specifically configured to:
[0109] Assign weights to the beams of each channel;
[0110] All beams are selected using the weights of all channels to obtain the target beam and the direction of arrival in the target direction.
[0111] Further, in some embodiments, the adaptive denoising layer includes an encoder, a 1×1 convolutional layer, and a decoder;
[0112] The encoder is configured to process the subframe to obtain a fourth output, the subframe being obtained by dividing the reference noise;
[0113] The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features;
[0114] The decoder is used to enhance the amplitude spectrum of the target beam using the noise characteristics.
[0115] The system provided in the embodiment of the present invention is briefly described. For matters not mentioned in the embodiment part, reference may be made to the corresponding content in the aforementioned embodiment.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.
Claims
1. A multi-channel speech enhancement method based on neural network, characterized in that: include: Receive voice signals from multiple channels; Processing the speech signal of each channel using the filter of each channel to obtain a beam corresponding to the angle of each channel; determining a target beam and a direction of arrival based on all of the beams; Obtaining a plurality of reference noises according to the speech signals of the plurality of channels and the directions of arrival; Inputting the reference noise and the target beam into an adaptive noise reduction layer to enhance the target beam; The processing of the speech signal of each channel by using the filter of each channel to obtain the beam of the corresponding angle of each channel specifically includes: selecting one of the channels as a reference channel; In each channel, the voice signal is convolved with the filter to obtain a beam corresponding to an angle of each channel; The filter optimization method includes: Calculating the cosine similarity between the beam obtained by each channel and the beam obtained by the reference channel respectively to obtain the similarity of each channel; Inputting the similarities of all the channels into a fully connected network to obtain first output data; Performing affine transformation on the filters of all the channels to obtain second output data; The first output data and the second output data are added and input into a mapping function to obtain filters optimized for all channels.
2. The multi-channel speech enhancement method based on neural network according to claim 1, characterized in that: Determining the target beam and the direction of arrival based on all the beams specifically includes: Assign weights to the beams of each channel; All beams are selected using the weights of all channels to obtain the target beam and direction of arrival in the target direction.
3. The multi-channel speech enhancement method based on neural network according to claim 2, characterized in that: The weight optimization method includes: Performing an affine transformation on the weights of all the channels to obtain third output data; The first output data and the third output data are added together and input into a mapping function to obtain optimized weights of all channels.
4. The multi-channel speech enhancement method based on neural network according to claim 1, characterized in that: The adaptive denoising layer includes an encoder, a 1×1 convolutional layer and a decoder; The encoder is configured to process a subframe to obtain a fourth output, wherein the subframe is obtained by dividing the reference noise; The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features; The decoder is configured to enhance the amplitude spectrum of the target beam by utilizing the noise characteristics.
5. A multi-channel speech enhancement system based on neural network, characterized in that: include: Input layer: used to receive speech signals from multiple channels; Fixed beamforming layer: used to process the speech signal of each channel using the filter of each channel to obtain the beam of the corresponding angle of each channel; A beam direction selection unit is configured to determine a target beam and a direction of arrival based on all the beams; Noise blocking layer: used to obtain multiple reference noises according to the voice signals of the multiple channels and the directions of arrival; Adaptive noise cancellation layer: used to receive the reference noise and the target beam, and output the enhanced signal of the target beam; The fixed beamforming layer is specifically used for: selecting one of the channels as a reference channel; In each channel, the voice signal is convolved with the filter to obtain a beam corresponding to an angle of each channel; The filter optimization method includes: Calculating the cosine similarity between the beam obtained by each channel and the beam obtained by the reference channel respectively to obtain the similarity of each channel; Inputting the similarities of all the channels into a fully connected network to obtain first output data; Performing affine transformation on the filters of all the channels to obtain second output data; The first output data and the second output data are added and input into a mapping function to obtain filters optimized for all channels.
6. The multi-channel speech enhancement system based on neural network according to claim 5, characterized in that: The beam direction selection unit is specifically configured to: Assign weights to the beams of each channel; All beams are selected using the weights of all channels to obtain the target beam and direction of arrival in the target direction.
7. The multi-channel speech enhancement system based on neural network according to claim 5, characterized in that: The adaptive denoising layer includes an encoder, a 1×1 convolutional layer and a decoder; The encoder is configured to process a subframe to obtain a fourth output, wherein the subframe is obtained by dividing the reference noise; The 1×1 convolution layer is used to perform 1×1 convolution on the fourth output to extract features to obtain multiple noise features; The decoder is configured to enhance the amplitude spectrum of the target beam by utilizing the noise characteristics.
Citation Information
Patent Citations
Voice arrival direction estimation method and device
CN110261816A
Microphone array voice enhancement method and system based on deep neural network
CN110600050A