An electroencephalogram assisted speech enhancement method and system based on sparse channels
By fusing EEG and speech signals through sparse channel selection and cross-attention mechanisms, the problem of accurately extracting the speech of the target speaker in noisy environments has been solved, reducing equipment costs and promoting the popularization of EEG devices.
Patent Information
- Application Number
- CN202411647226.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing technologies struggle to accurately capture the voice of a target speaker in noisy environments, and the high cost of EEG equipment limits its widespread adoption.
By employing sparse channel selection technology, a small number of channels of EEG signals are selected through a channel selection method and then fused with speech signals using a cross-attention mechanism to achieve end-to-end speech enhancement.
While reducing the number of EEG channels, the voice enhancement effect is improved, the equipment cost and complexity are reduced, and the popularization of EEG equipment is promoted.
Smart Images

Figure CN119517058B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of speech enhancement, and particularly relates to an electroencephalogram auxiliary speech enhancement method and system based on a sparse channel. BACKGROUND
[0002] For speech enhancement, the prior art usually has speech enhancement technology, brain auxiliary speech enhancement technology, EEG-based target speaker extraction technology, and end-to-end network technology.
[0003] The existing speech enhancement based on brain auxiliary technology can realize extraction of a target speaker without relying on prior information of the target speaker, but brain auxiliary speech enhancement still faces challenges in performance and device cost.
[0004] Through the above analysis, the problems and defects of the prior art are that for hearing impairment, there is a group with decreased auditory attention, and the existing technology has poor effect in accurately focusing on the target speaker's voice in a noisy environment, and the device cost for extracting electroencephalogram signals is high. SUMMARY
[0005] To overcome the problems in the related art, the application discloses an electroencephalogram auxiliary speech enhancement method and system based on a sparse channel. In particular, it relates to an end-to-end auxiliary speech enhancement technology based on electroencephalogram sparse channel selection.
[0006] The technical solution is as follows: The electroencephalogram auxiliary speech enhancement method based on a sparse channel integrates channel selection technology into brain auxiliary speech enhancement, fuses time-domain speech signals and electroencephalogram signals based on a cross-attention mechanism, and realizes end-to-end enhancement and extraction of target speaker speech. Specifically, the method includes the following steps:
[0007] S1, EEG brain signals are screened by a channel selection method to obtain a small number of EEG channel signals, and then electroencephalogram features are extracted; speech signals are subjected to speech signal coding network to obtain time-series speech features; the speech signal coding network is composed of one-dimensional convolution, and the speech signal is subjected to one-dimensional convolution feature extraction to obtain time-series speech features;
[0008] S2, the two routes of the obtained speech features and the extracted electroencephalogram features are input into a multi-layer cross-attention mechanism module for feature fusion;
[0009] S3, based on a speech decoding network, the fused features are decoded into enhanced target speaker speech.
[0010] In step S1, the EEG brain signals are screened by a channel selection method to obtain a small number of EEG channel signals, including:
[0011] In the training process, for N EEG channels, K neurons are selected, the input and output size of each neuron is the number of EEG channels, and each neuron selects one channel from N EEG signals through the Gumbel Softmax function, and finally K neurons select K EEG channels;
[0012] In the inference process, the EEG signal is output after the EEG channel selection module selects a small amount of EEG channel signals.
[0013] Further, the training process includes that the EEG channel selection module adopts a residual structure, and the expression is:
[0014] z final = ( (1-r)z ori +rPadding(z choice )
[0015] In the formula, z final is the final EEG channel signal, r is a coefficient, z ori is the original EEG channel signal, Padding() is a padding operation, and z choice is the selected EEG channel.
[0016] After the original EEG channel signal passes through the EEG channel selection module, the selected EEG channel z choice is obtained, and after the padding operation, the coefficient r is multiplied, and the original EEG channel signal z ori is multiplied by the coefficient (1-r), and the two are added to obtain the final EEG channel signal z final , and finally the speech signal of multiple speakers is input into the main network to obtain the enhanced speech.
[0017] Further, the inference process includes that the residual process is removed, the EEG signal passes through the EEG channel selection network, is padded, and is input into the main network together with the speech signal of multiple speakers, and finally the enhanced speech is obtained by the main network.
[0018] In step S2, feature fusion is performed, including: by fusing the extracted speech features and EEG features with each other, after multi-layer cross-attention mechanism fusion, the two features are spliced to obtain the fused features.
[0019] Further, in the multi-layer cross-attention mechanism fusion, fusion is performed through a cross-attention unit module, which performs different processing according to the differences between the input speech features and the EEG features.
[0020] Further, for the speech feature, after the deep separable convolution feature extraction, the Q matrix is obtained, and the EEG feature obtains the K and V matrices; the attention coefficient matrix is obtained by multiplying the transpose of the K matrix by the Q matrix, and the attention coefficient matrix is multiplied by the V matrix to obtain the fusion feature after the cross attention of the speech feature to the EEG feature.
[0021] For the EEG feature, the Q matrix is obtained by the EEG feature, and the K and V matrices are obtained by the speech feature, and finally the fusion feature after the cross attention of the EEG feature to the speech feature is obtained.
[0022] Another object of the present application is to provide a sparse channel-based electroencephalogram-assisted speech enhancement system, which implements the sparse channel-based electroencephalogram-assisted speech enhancement method, and the system comprises:
[0023] An EEG channel selection module is configured to filter a small number of EEG channel signals through a channel selection method, and then extract EEG features, and configured to obtain time-series speech features through a speech signal coding network;
[0024] A training and reasoning process module is configured to train and reason the EEG and speech signals;
[0025] A multi-layer cross-attention mechanism module is configured to input the obtained speech features and the extracted EEG features into the multi-layer cross-attention mechanism module at the same time, and perform feature fusion;
[0026] A cross-attention unit module is configured to specifically perform speech feature and EEG feature fusion to obtain the fusion feature after the cross attention of the speech feature to the EEG feature;
[0027] A speech signal coding network module is configured to decode the fused features into enhanced target speaker speech.
[0028] Further, the system is mounted on a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the functions in the above-mentioned sparse channel-based electroencephalogram-assisted speech enhancement system.
[0029] Further, the system is mounted on a server, and the server is configured to provide a user input interface to implement the functions in the above-mentioned sparse channel-based electroencephalogram-assisted speech enhancement system when executed on an electronic device, and the electronic device is mounted on a headset
[0030] In combination with all the above technical solutions, the application has the beneficial effects that the speech signal and the electroencephalogram signal in the time domain are fused based on the cross attention mechanism to enhance and extract the target speaker voice, reducing the use of the number of electroencephalogram signal channels, and the fewer the electroencephalogram signal channels, the lower the cost of the device for extracting the electroencephalogram signal. For the hearing-impaired, there is a problem of decreased auditory attention and inability to accurately focus on the target speaker voice in a noisy environment. With the development of brain-computer interface technology, the application uses the electroencephalogram signal of the listener as additional modal information to extract the target speaker voice signal without relying on prior knowledge of the target speaker information. The application is applied to extracting the target speaker voice in a multi-speaker scene only relying on a small number of EEG channel signals. Through the scheme, the number of EEG channels is reduced, and the cost of the EEG signal collection device is reduced.
[0031] The traditional speech enhancement method usually relies on a large number of EEG channels, which not only increases the complexity and cost of the device, but also limits the popularity of the EEG device. The application can significantly reduce the number of required EEG channels while ensuring the speech enhancement effect through the innovative channel selection technology, thereby reducing the device cost and complexity. Although increasing the number of channels can improve the speech enhancement effect, it will lead to a significant increase in device cost and complexity, limiting the popularity of the EEG device. The application successfully solves this problem through the innovative channel selection technology, realizing high-quality speech enhancement effect while reducing the number of channels. The traditional concept believes that to achieve high-quality speech enhancement effect, a large number of EEG channels must be relied on, which leads to an increase in device cost and complexity. This bias limits the popularity and application of the EEG device. The application proves that effective speech enhancement effect can be achieved even with a reduced number of channels through the innovative channel selection technology. BRIEF DESCRIPTION OF DRAWINGS
[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure;
[0033] Figure 1 is a flow chart of the electroencephalogram assisted speech enhancement method based on sparse channels provided by the embodiment of the application;
[0034] Figure 2 is a principle diagram of the main network flow module provided by the embodiment of the application;
[0035] Figure 3 is a block diagram of the electroencephalogram assisted speech enhancement system based on sparse channels provided by the embodiment of the application;
[0036] Figure 4It is a multi-layer cross attention mechanism module schematic diagram provided by the embodiment of the application.
[0037] Figure 5 It is a cross attention unit module schematic diagram provided by the embodiment of the application.
[0038] Figure 6 It is an EEG channel selection module schematic diagram provided by the embodiment of the application.
[0039] Figure 7 It is a training process schematic diagram provided by the embodiment of the application.
[0040] Figure 8 It is a reasoning process schematic diagram provided by the embodiment of the application.
[0041] In the figure: 1, an EEG channel selection module; 2, a training and reasoning process module; 3, a multi-layer cross attention mechanism module; 4, a cross attention unit module; 5, a speech signal coding network module. DETAILED DESCRIPTION
[0042] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below. In the following description, a large number of specific details are set forth in order to facilitate a full understanding of the application. However, the application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the application, so the application is not limited by the specific embodiments disclosed below.
[0043] The innovation of the application lies in that the application can effectively reduce the dependence of speech enhancement using electroencephalogram signals on the number of EEG channels by proposing a channel selection technology. Generally, the more the number of channels, the better the speech enhancement effect, but the cost of the device is increased. The application can effectively reduce the number of EEG channels extracted while ensuring the enhancement of speech effect, thereby reducing the cost of the device and further promoting the popularization of EEG devices.
[0044] The speech signal and the electroencephalogram signal in the time domain are fused based on the cross attention mechanism, the target speaker's speech is extracted by end-to-end enhancement, and the target speaker's speech enhancement extraction in a noisy environment is realized only with a small number of electroencephalogram channels. Since the cost required by the electroencephalogram signal collection device increases with the number of channels, the application reduces the use cost of the electroencephalogram information collection device, and only a small number of channel electroencephalogram signal collection devices are required.
[0045] The EEG signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0046] S1, the EEG brain signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0047] S2, the two routes of the speech features and the extracted EEG features are input into a multi-layer cross-attention mechanism module for feature fusion;
[0048] S3, based on the speech decoding network, the fused features are decoded into enhanced target speaker speech.
[0049] In an example, the EEG signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0050] The main network flow module is as shown in Figure 2 The EEG signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0051] The two routes of the speech features and the extracted EEG features are input into a multi-layer cross-attention mechanism module for feature fusion. Finally, based on the speech decoding network, the fused features are decoded into enhanced target speaker speech.
[0052] In another example, as shown in Figure 3 The EEG signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0053] The EEG signal is filtered by the channel selection method to obtain a small number of channel signals, and then EEG features are extracted; the speech signal is encoded by a speech signal encoding network to obtain time sequence speech features; the speech signal is encoded by the speech signal encoding network, which is composed of one-dimensional convolution, and time sequence speech features are obtained after one-dimensional convolution feature extraction;
[0054] The training and inference process module 2 is used for training and inference of EEG and speech signals.
[0055] The multi-layer cross-attention mechanism module 3 is used for inputting the obtained speech features and the extracted EEG features into the multi-layer cross-attention mechanism module at the same time for feature fusion.
[0056] The cross-attention unit module 4 is used for specifically performing speech feature and EEG feature fusion to obtain the fusion features after cross-attention of the speech features to the EEG features.
[0057] The speech signal encoding network module 5 is used for decoding the fusion features into enhanced target speaker speech.
[0058] The EEG channel selection module 1 is as shown in the formula (2). Figure 6 At present, the mainstream EEG signal usually needs to sample 128 channels of EEG signals, and the EEG acquisition device is relatively expensive, and the more channels, the more time-consuming and high cost. The EEG channel selection module 1 is used to select a small number of EEG channels from 128 channels, and ensure that the reduction of the number of channels does not affect the speech enhancement effect of the model.
[0059] For example, the training and inference process module 2, in the training process, the specific screening and calculation process is as follows: for N EEG channels, K neurons are selected for screening, and the input and output sizes of each neuron are both the number of EEG channels. Through the Gumbel Softmax function, each neuron selects one channel from N EEG signals, and finally K neurons select K EEG channels. Gumbel Softmax is a technique for sampling in discrete distributions, especially when dealing with discrete variables, it can effectively propagate gradients. The core idea of Gumbel Softmax is to approximate discrete sampling by introducing Gumbel noise while maintaining differentiability.
[0060] In the inference process, the EEG signal is outputted after the screening of the small number of EEG channel signals by the EEG channel selection module.
[0061] For example, the training process of the present application is as shown in the formula (3). Figure 7 The mixed speech signals of multiple speakers and the collected EEG signals are inputted into the audio signal extraction module and the EEG feature channel selection module respectively. In the training process, the EEG channel selection module adopts a residual structure, as shown in the formula (1).
[0062] z final = ( 1-r)z ori +rPadding(z choice ) (1)
[0063] In the formula, z final is the final EEG channel signal, r is a coefficient, z ori is the original EEG channel signal, Padding() is a padding operation, z choice is the selected EEG channel;
[0064] After the original EEG channel signal passes through the EEG channel selection module, the selected EEG channel z choice is obtained, and after the padding operation, the selected EEG channel z ori is multiplied by the coefficient r, and the original EEG channel signal z final is added, and the two are added to obtain the final EEG channel signal z
[0065] The inference process of the present application is shown in Figure 8 , and the residual process is removed in the inference process. The EEG signal passes through the EEG channel selection network, and the residual result is used in the training process. In the inference process, the residual result is removed, the selected EEG features are expanded in the channel after padding, the retained channels are copied in the channel dimension, so that the EEG feature size after passing through the channel selection network remains the same as the original feature. The mixed speech signal of multiple speakers is input into the main network, and the enhanced speech is finally obtained by the main network.
[0066] Exemplarily, the multi-layer cross-attention mechanism module 3 is shown in Figure 4 . The module fuses the extracted speech features and EEG features with each other, and then splices the two features to obtain the fused features after multi-layer cross-attention mechanism fusion. The speech features and EEG features are fused through the multi-layer cross-attention layer, and in each layer, the speech features and EEG features are input into the cross-attention unit in Figure 5 of the module 4 for feature fusion. In each layer, the fusion is performed twice, the first time the speech features obtain the Q matrix, and the EEG features obtain the KV matrix, and the second time the speech features obtain the KV matrix, and the EEG features are taken as the Q matrix.
[0067] Exemplarily, the cross-attention unit module 4 is shown in Figure 5 . The cross-attention unit module inputs the speech features and EEG features. The cross-attention unit module has different processing procedures according to different input signals. For the speech features, the Q matrix is obtained after deep separable convolution feature extraction, and the EEG features obtain the K and V matrices. The attention coefficient matrix is obtained by multiplying the Q matrix by the transpose of the K matrix, and the attention coefficient matrix is multiplied by the V matrix to obtain the fusion features of the cross-attention between the speech features and the EEG features.
[0068] For the EEG feature, the Q matrix is obtained by the EEG brain wave feature, the K and V matrices are obtained by the speech feature, and finally the fusion feature of the EEG brain wave feature after cross attention of the speech feature is obtained. It can be known from the above embodiment that the present application can be used for realizing speech enhancement technology only by relying on a small number of EEG channels.
[0069] The present application is experimentally verified on the EEG and audio data set of Berlin Charite University, and based on the design of the present application, after reducing the channel number to 18 channels, the effect is less reduced compared with 128 channels, which proves that the scheme can effectively reduce the dependence on the number of channels, and the specific is shown in Table 1.
[0070] Table 1 Channel number monitoring table
[0071]
[0072] The above is only a more preferable specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any modification, equivalent replacement and improvement made by any person skilled in the art within the technical range disclosed by the present application and within the spirit and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A sparse-channel EEG-assisted speech enhancement method, characterized in that, This method integrates channel selection technology into brain-assisted speech enhancement, fusing temporal speech signals and EEG signals based on a cross-attention mechanism to achieve end-to-end enhancement and extraction of the target speaker's speech; specifically, it includes the following steps: S1, EEG signals are filtered through a small number of channels using a channel selection method, and then EEG features are extracted; speech signals are processed through a speech signal encoding network to obtain temporal speech features. The speech signal is processed through a speech signal encoding network, which consists of one-dimensional convolutions. The speech signal is processed through one-dimensional convolutional feature extraction to obtain temporal speech features. S2, inputs the two separately acquired speech features and extracted EEG features into the multi-layer cross-attention mechanism module for feature fusion; S3, based on a speech decoding network, decodes the fused features into enhanced target speaker speech; In step S1, the EEG signals are filtered through a small number of channels using a channel selection method, including: In the training process, for N EEG channels, K neurons are selected for screening. The input and output size of each neuron is the number of EEG channels. Through the Gumbel Softmax function, each neuron selects one channel from N EEG signals, and finally K neurons select K EEG channels. During the inference process, the EEG signal passes through the EEG channel selection module and outputs a small number of filtered EEG channel signals. The training process includes: The EEG channel selection module uses a residual structure, and the expression is: With final =(1-r)z ori +rPadding(z choice ) In the formula, z final The final EEG channel signal, r is the coefficient, z ori This is the original EEG channel signal, Padding() is the padding operation, z choice The selected EEG channel; The raw EEG channel signal passes through the EEG channel selection module to obtain the selected EEG channel z. choice After padding, the signal is multiplied by a coefficient r and then added to the original EEG channel signal z. ori Multiply by a coefficient of (1-r) and add the two together to obtain the final EEG channel signal z. final Finally, the speech signals from multiple speakers are mixed and input into the main network to obtain the enhanced speech; The inference process includes: removing residuals, selecting the network through the EEG channel, padding, and inputting it into the main network along with the speech signals from multiple speakers, and finally obtaining the enhanced speech from the main network.
2. The sparse channel-based EEG-assisted speech enhancement method according to claim 1, characterized in that, In step S2, feature fusion is performed, including fusing the extracted speech features and EEG features, fusing them through a multi-layer cross-attention mechanism, and then splicing the two features together to obtain fused features.
3. The sparse-channel-based EEG-assisted speech enhancement method according to claim 1, characterized in that, In the multi-layer cross-attention mechanism fusion, the fusion is carried out through the cross-attention unit module, which performs different processing according to the different input speech features and EEG features.
4. A sparse-channel EEG-assisted speech enhancement system, characterized in that, The system implements the sparse channel-based EEG-assisted speech enhancement method according to any one of claims 1-3, and the system comprises: EEG channel selection module (1) is used to filter a small number of EEG channel signals through the channel selection method and then extract EEG features; and is used to obtain temporal speech features from speech signals through a speech signal coding network. Training reasoning process module (2) is used to train reasoning on EEG brain signals and speech signals; The multi-layer cross-attention mechanism module (3) is used to simultaneously input the acquired speech features and the extracted EEG features into the multi-layer cross-attention mechanism module for feature fusion. The cross-attention unit module (4) is used to specifically perform the fusion of speech features and EEG features to obtain the fused features of speech features after cross-attention to EEG features; The speech signal coding network module (5) is used to decode the fused features into the enhanced target speaker's speech.
5. The sparse-channel-based EEG-assisted speech enhancement system according to claim 4, characterized in that, The system is mounted on a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it can realize the functions of the sparse channel-based EEG-assisted speech enhancement system described above.
6. The sparse-channel-based EEG-assisted speech enhancement system according to claim 4, characterized in that, The system is mounted on a server that, when executed on an electronic device, provides a user input interface to implement the functions of the sparse channel-based EEG-assisted speech enhancement system described above, which is mounted on an earphone.
Citation Information
Patent Citations
Multi-task target voice separation method and system fusing electroencephalogram signals
CN118737183A
Enhancement techniques for blind source separation (BSS)
WO2007130797A1