Electroencephalogram auditory speech extraction model training method and device, equipment and storage medium
By extracting features from EEG and speech data and fusing them to generate speech embedding features, the problem of hearing-impaired people having difficulty recognizing target speech in a multi-speaker environment is solved, and a fine-grained description of the listener's dynamic auditory attention state and accurate separation of the target speech are achieved.
Patent Information
- Application Number
- CN202511135363.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-14
AI Technical Summary
People with hearing impairments find it difficult to distinguish the target speaker's voice in a noisy multi-talker environment, and existing hearing aids cannot accurately identify the source of the voice that the listener is paying attention to.
By acquiring EEG data and mixed speech data, extracting audio coding features, EEG time features and EEG spatial features, fusing them to generate speech embedding features, using the EEG auditory speech extraction model to separate the target sound source from the mixed speech data, and training the model through the main model loss until the preset conditions are met.
It can accurately identify the speech source that the listener is paying attention to in a multi-sound source environment, has good generalization ability and stability, and can automatically infer the speech source of the listener's attention based on the listener's EEG response.
Smart Images

Figure CN120673751A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for training an EEG auditory speech extraction model. Background Art
[0002] In real life, humans are typically able to naturally focus their attention on a single target speaker and ignore other distracting sounds in a noisy, multi-talking environment. This ability is known as "selective auditory attention." For example, in a cocktail party setting, even if there are multiple people speaking, most people can still clearly discern the speech of the speaker they are focusing on.
[0003] However, for people with hearing impairments, their neural system's reduced ability to process auditory information makes it difficult to distinguish the intended speaker from multiple voice interferences, greatly limiting their communication efficiency and quality of life. Current hearing aid technology, while capable of amplifying sound, cannot accurately identify the source of the voice the listener is truly focused on. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for training an EEG auditory speech extraction model that can accurately identify the speech source that the listener is concerned about in order to solve the above technical problems.
[0005] In a first aspect, the present application provides a method for training an EEG auditory speech extraction model, comprising:
[0006] Acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0007] Performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used to characterize the change trend of the electroencephalogram signal data over time, and the electroencephalogram space features are used to characterize the change characteristics of the electroencephalogram signal data between different electroencephalogram channels;
[0008] fusing the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature;
[0009] The main model loss is calculated based on the first predicted audio data and the original audio data of the target sound source, and the EEG auditory speech extraction model is trained based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0010] In one embodiment, the method further comprises:
[0011] Dividing the speech embedding feature into a plurality of continuous feature segments; each of the feature segments corresponds to an original audio segment in the original audio data of the target sound source;
[0012] For each of the feature segments, a context fusion feature is obtained according to the feature segment and the previous feature segment;
[0013] Separating second predicted audio data of a target sound source from the mixed speech data according to the context fusion feature;
[0014] Calculating a causal layer loss based on the second predicted audio data and the original audio segment corresponding to the feature segment;
[0015] The method of training the EEG auditory speech extraction model according to the main model loss until a preset termination condition is met to obtain a trained EEG auditory speech extraction model includes:
[0016] The EEG auditory speech extraction model is trained according to the causal layer loss and the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0017] In one embodiment, separating the second predicted audio data of the target sound source from the mixed speech data according to the context fusion feature includes:
[0018] Get the initial sampling probability and attenuation coefficient;
[0019] Calculating the sampling probability of the characteristic segment according to the initial sampling probability and the attenuation coefficient;
[0020] Selecting the context fusion feature or the feature fragment as a target embedding feature according to the sampling probability;
[0021] Second predicted audio data of a target sound source is separated from the mixed speech data according to the target embedded feature.
[0022] In one embodiment, extracting features from the EEG data to obtain EEG temporal features and EEG spatial features includes:
[0023] Performing time position encoding on the electroencephalogram data to obtain a first electroencephalogram data sequence;
[0024] Performing convolution calculation on the time axis for the first EEG data sequence to obtain EEG time features;
[0025] performing channel position encoding on the electroencephalogram data to obtain a second electroencephalogram data sequence;
[0026] For the second EEG data sequence, convolution calculation is performed between different EEG channels to obtain EEG spatial features.
[0027] In one embodiment, the fusing of the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain the speech embedding feature includes:
[0028] Performing self-attention alignment on the EEG temporal feature and the EEG spatial feature on the time axis to obtain EEG signal features;
[0029] The EEG signal features and the audio coding features are fused to obtain speech embedding features.
[0030] In one embodiment, fusing the EEG signal feature and the audio coding feature to obtain a speech embedding feature includes:
[0031] Determining the number of speech frames according to the audio coding characteristics;
[0032] Performing linear interpolation calculation on the EEG signal feature according to the number of speech frames to update the EEG signal feature;
[0033] The updated EEG signal features are aligned and fused with the audio coding features to obtain speech embedding features.
[0034] In a second aspect, the present application also provides a device for training an EEG auditory speech extraction model, comprising:
[0035] A data acquisition module, configured to acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0036] A feature extraction module is used to perform feature extraction on the mixed speech data to obtain audio coding features, and to perform feature extraction on the EEG data to obtain EEG time features and EEG spatial features; wherein the EEG time features are used to characterize the change trend of the EEG signal data over time, and the EEG spatial features are used to characterize the change characteristics of the EEG signal data between different EEG channels;
[0037] a feature fusion module, configured to fuse the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separate first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature;
[0038] The model training module is used to calculate the main model loss based on the first predicted audio data and the original audio data of the target sound source, and train the EEG auditory speech extraction model based on the main model loss until a preset termination condition is met to obtain a trained EEG auditory speech extraction model.
[0039] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0040] Acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0041] Performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used to characterize the change trend of the electroencephalogram signal data over time, and the electroencephalogram space features are used to characterize the change characteristics of the electroencephalogram signal data between different electroencephalogram channels;
[0042] fusing the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature;
[0043] The main model loss is calculated based on the first predicted audio data and the original audio data of the target sound source, and the EEG auditory speech extraction model is trained based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0044] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0045] Acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0046] Performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used to characterize the change trend of the electroencephalogram signal data over time, and the electroencephalogram space features are used to characterize the change characteristics of the electroencephalogram signal data between different electroencephalogram channels;
[0047] fusing the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature;
[0048] The main model loss is calculated based on the first predicted audio data and the original audio data of the target sound source, and the EEG auditory speech extraction model is trained based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0049] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0050] Acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0051] Performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used to characterize the change trend of the electroencephalogram signal data over time, and the electroencephalogram space features are used to characterize the change characteristics of the electroencephalogram signal data between different electroencephalogram channels;
[0052] fusing the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature;
[0053] The main model loss is calculated based on the first predicted audio data and the original audio data of the target sound source, and the EEG auditory speech extraction model is trained based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0054] The above-mentioned EEG auditory speech extraction model training method, apparatus, computer device, storage medium, and computer program product obtain EEG data and mixed speech data. The mixed speech data includes multiple raw audio data from different sound sources; the EEG data includes EEG signal data from multiple EEG channels. Feature extraction is performed on the mixed speech data to obtain audio coding features. Feature extraction is also performed on the EEG data to obtain EEG temporal features and EEG spatial features. Feature extraction and semantic modeling are performed on these two types of data respectively, enabling the system to effectively perceive and decode the listener's auditory attention direction in a multi-source speech environment. The system obtains audio coding features representing the audio content, EEG temporal features that vary over time, and EEG spatial features that reflect the coordinated activity patterns of different brain regions, thereby achieving a fine-grained description of the listener's dynamic auditory attention state. On this basis, the audio coding features are jointly modeled with the EEG temporal features and EEG spatial features to further generate speech embedding features. This embedded feature serves as an intermediate representation of the "speech that the listener is paying attention to" and has semantic information that contains both speech content and attention cues. By guiding the separation of mixed speech data through speech embedding features, the target sound source can be accurately extracted from the mixed sound sources to obtain the first predicted audio data. Then, by comparing the first predicted audio data with the original audio data corresponding to the target sound source and calculating the main model loss function, effective training of the EEG auditory speech extraction model can be achieved. By continuously optimizing the model and completing the training when the loss function meets the preset termination condition, a trained EEG auditory speech extraction model can be obtained that can jointly judge the object of auditory attention based on EEG and speech data. This model has good generalization ability and stability, and can automatically infer the source of the speech that the listener is paying attention to based on the EEG response. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 A diagram illustrating an application environment of a method for training an EEG auditory speech extraction model in one embodiment;
[0057] Figure 2 1 is a flow chart of a method for training an EEG auditory speech extraction model in one embodiment;
[0058] Figure 3 is an architectural diagram of a MossFormer block in one embodiment;
[0059] Figure 4 Schematic diagram of a flow chart of a method for training an EEG auditory speech extraction model in another embodiment;
[0060] Figure 5 is an overall architecture diagram of a Mindflow model in one embodiment;
[0061] Figure 6 is an architectural diagram of a context fusion module in one embodiment;
[0062] Figure 7 1. It is a structural block diagram of a device for training an EEG auditory speech extraction model in one embodiment;
[0063] Figure 8 is a diagram of the internal structure of a computer device in one embodiment;
[0064] Figure 9 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0066] The EEG auditory speech extraction model training method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be used to obtain EEG data and mixed voice data, and send the EEG data and mixed voice data to the server 104 for processing. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The data storage system can be used to store mixed voice data, EEG data, and data in steps such as feature extraction, feature fusion, and audio prediction. Among them, the terminal 102 can be used to obtain EEG data, which can be an EEG data acquisition device, a head-mounted device, etc. The terminal 102 can also be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices for collecting mixed voice signals and converting the mixed voice signals into mixed voice data and sending them to the server 104. Among them, the Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart car device, and the portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0067] In an exemplary embodiment, Figure 2 As shown in the figure, a method for training an EEG auditory speech extraction model is provided, and the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps S202 to S208.
[0068] Step S202: Acquire electroencephalogram data and mixed speech data.
[0069] Mixed speech data includes multiple raw audio data from different sound sources, while EEG data includes EEG signal data from multiple EEG channels. Mixed speech data is the result of combining raw audio data from two or more different sound sources.
[0070] For example, in an actual application scenario, such as simulating a "cocktail party" environment, assuming that speaker A says "hello" and speaker B says "what time is it now" at the same time, the mixed speech may be the linear overlap of these two audio segments on the time axis. The server 104 can obtain the mixed speech data collected by the terminal 102 and regard the mixed speech as the original input signal x(t), which is an approximate reproduction of the actual auditory scene and a direct stimulus source for triggering the selective auditory attention response in the brain. is a noisy multi-speaker mixture signal in the time domain:
[0071]
[0072] in represents the speech signal of the target speaker selected by the user, represents the speech of M interfering speakers, Indicates ambient noise.
[0073] EEG data refers to EEG signals recorded by high-density EEG acquisition equipment (e.g., the 64-channel Bio Semi ActiveTwo system). This signal reflects the neural responses of different brain regions to speech stimulation during the subject's auditory task. Specifically, EEG data is a multi-channel voltage signal matrix that varies over time, where each channel represents an electrode acquisition point (e.g., Fz, Pz, Cz, etc.). These electrodes are distributed across different areas of the scalp surface and are used to reflect the activity of the cerebral cortex in different frequency bands. For example, when the target speaker is located to the right of the listener, the subject's left frontal lobe-related channels may produce more significant response waveforms. For example, given the real-time EEG signal E(t) encoding the user's auditory attention, server 104 can extract:
[0074]
[0075] in Indicates that it has parameters neural structure.
[0076] For example, server 104 can control the EEG acquisition device to start recording at the same time as the subject completes the auditory attention task, ensuring that the EEG sampling time base is fully synchronized with the speech playback. The raw EEG signal E_raw acquired during this process can be accompanied by a trigger signal (i.e., a trigger marker) to mark the time point when the mixed speech playback begins. Furthermore, regarding the sampling frequency, EEG acquisition can be set to 8196 Hz to capture high-precision neural potential changes, while the speech data sampling rate can be 44.1 kHz to preserve the sound quality and details of the speech.
[0077] Subsequently, the server 104 can save the raw EEG signal E_raw collected and sent by the terminal 102 as a 64-channel voltage sequence, with each channel corresponding to the activity record of a brain region. At the same time, the server 104 can also record the complete mixed signal of the voice playback. , where the sound sample value at each moment can simultaneously contain the speech information of multiple speakers and can also contain multiple types of environmental noise. In some embodiments, server 104 can also access a clean speech source (i.e., the original speech tracks of speakers A and B) as reference data for subsequent training and validation. Through this step, server 104 can obtain original input data representing auditory input (mixed speech) and auditory attention response (EEG) by synchronously collecting mixed speech data and EEG data.
[0078] Step S204: extract features from the mixed speech data to obtain audio coding features, and extract features from the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features.
[0079] Among them, the EEG time feature is used to characterize the changing trend of EEG signal data over time, and the EEG spatial feature is used to characterize the changing characteristics of EEG signal data between different EEG channels.
[0080] For example, for mixed speech data, the server 104 may use a one-dimensional convolutional encoder (Conv1DEncoder) and use a ReLU activation function to extract audio coding features. The encoder can perform convolution operations on the audio waveform in a sliding window manner and extract speech feature information in different frequency and timing modes through multiple channels. (where B is the batch size and Ts is the input length), the encoder application size is The kernel of , produces the encoded output , which is defined as:
[0081]
[0082] Where N represents the number of filters, , representing the reduced time dimension.
[0083] For example, server 104 may slice the audio into frames of 20 samples each, with a step size of 10, to form a sequence of overlapping frames. This is then transformed using a convolution kernel to compress the original signal while preserving the recognizable sound information. Each frame of audio features output by the convolutional encoder not only contains the time domain information of the original waveform but also possesses certain perceptual semantic capabilities.
[0084] Furthermore, for processing EEG data, the server 104 may employ a multi-stage structured encoding method to obtain EEG temporal features and EEG spatial features. For example, the server 104 may perform temporal position encoding on the EEG data to obtain a first EEG data sequence; perform convolution calculations on the time axis for the first EEG data sequence to obtain EEG temporal features; perform channel position encoding on the EEG data to obtain a second EEG data sequence; and perform convolution calculations between different EEG channels for the second EEG data sequence to obtain EEG spatial features.
[0085] For example, in order to preserve the dynamic characteristics of EEG signals changing over time, server 104 can enhance the time label of EEG input by introducing a time position encoding mechanism (i.e., position vector in the time dimension, Position Encoding in Time) to clarify the position of each frame of EEG signal on the time axis. Subsequently, server 104 can input the encoded EEG signal into the self-attention mechanism module (Self-Attention Module) to extract long-range dependencies in the time dimension. Through this operation, server 104 can identify the intrinsic patterns between different time points from the original multi-channel EEG waveform, such as whether the neural response generated by a certain speech stimulus will affect the neural activity of several frames later.
[0086] After completing the initial temporal modeling, server 104 can further apply a temporal convolution operation to the timing signal of each EEG channel. The convolution kernel slides only along the time direction to enhance the identification of fluctuation trends within the local time window. For example, server 104 may slide the convolution kernel over 25 consecutive time points to identify whether the channel has a rapid voltage rise or fall within a short period of time. These changes are often associated with speech stimulation. After the convolution output is pooled, normalized, and activated, server 104 obtains a clear EEG temporal feature, which is used to characterize the neural response trajectory of the brain that evolves over time during the auditory process.
[0087] For example, after temporal feature extraction is completed, server 104 can further apply spatial convolution to capture the synergistic relationship between different EEG channels. Cross-channel convolution processing can be performed between electrode channels, such as using a convolution kernel covering all 64 channels to identify whether there is a synchronous activation or complementary inhibition relationship between different brain regions at the same time. This spatial feature is the basis for the brain's neural processing of multi-source speech. For example, when a listener receives the main sound stimulus in the right ear, the EEG signal of the left temporal lobe may be enhanced, while the right side is relatively stable. Server 104 defines this part of the output as EEG spatial features, which are used to describe the organized and synergistic responses between brain regions.
[0088] For example, for EEG data (Where B is the batch size, C is the number of electrode channels, is the input length), the server 104 may apply time position coding The self-attention module (SA) encodes the temporal dependencies in the EEG signal. Its formula is as follows:
[0089]
[0090] Furthermore, the server 104 can extract temporal patterns using temporal and spatial convolution. and spatial distribution characteristics , as shown below:
[0091]
[0092]
[0093] Among them, TemporalConv2d() uses a size of (1, 25) Filters perform a two-dimensional convolution along the time axis. SpatialConv2d() performs convolution along the electrode channel dimension using a two-dimensional convolution filter of the same size as the EEG channel. Norm refers to BatchNorm2D and ELU refers to the activation function.
[0094] Through the above steps, server 104 obtains audio coding features, EEG temporal features (representing dynamic changes in EEG over time), and EEG spatial features (representing spatial difference patterns between different channels). These features not only preserve the important semantic structure of the original signal but also provide input data with unified shape and complementary semantics for subsequent multimodal fusion and speech embedding modeling, thereby inferring whether the listener is currently paying attention to a specific sound source in the speech.
[0095] Step S206: The audio coding features, the EEG time features, and the EEG space features are fused to obtain speech embedding features, and the first predicted audio data of the target sound source is separated from the mixed speech data according to the speech embedding features.
[0096] Exemplarily, the server 104 can perform shape alignment processing on the audio coding features, EEG time features and EEG space features. In the aforementioned steps, the EEG time features and EEG space features have been extracted through convolution and attention mechanisms and have been aligned to a uniform number of frames and channels, respectively, which are consistent with the audio coding features. The server 104 can splice the EEG time features and space features in the channel dimension, and map them to the same feature dimension as the audio coding features through 1D convolution, and finally form a unified EEG fusion feature vector. In order to enable the network to distinguish which channel or brain area such EEG features come from, the server 104 can also introduce channel position coding. , that is, attaching a vector label to each EEG channel to indicate the physical space location it represents (such as left frontal lobe, right occipital lobe, etc.).
[0097] For example, server 104 may perform self-attention alignment on the EEG temporal features and the EEG spatial features on the time axis to obtain EEG signal features; fuse the EEG signal features with the audio coding features to obtain speech embedding features. Furthermore, server 104 may determine the number of speech frames based on the audio coding features; perform linear interpolation calculations on the EEG signal features based on the number of speech frames to update the EEG signal features; and align and fuse the updated EEG signal features with the audio coding features to obtain speech embedding features.
[0098] Exemplarily, the server 104 can perform channel splicing operations on the audio coding features and the fused EEG features, and merge them into a new composite input feature tensor along the feature dimension. This tensor represents both the information of the speech itself and the listener's reaction state to the speech at this frame time. The server 104 can then perform channel compression on the spliced tensor through 1D convolution, compressing it from the original 2N channels back to N channels to prevent subsequent network operations from being too heavy, and retaining high-order feature patterns through nonlinear activation operations. Exemplarily, the server 104 can apply linear interpolation to strictly align the above-mentioned EEG features with the audio features, thereby achieving time-aligned fusion and generating clue features. , which is defined as:
[0099]
[0100] Furthermore, the fused multimodal features can be input into the speaker extraction module deployed by server 104, which can adopt the MossFormer structure that integrates local convolution and global attention mechanisms. MossFormer combines the gated single-head attention mechanism with the convolution-enhanced joint self-attention mechanism, and simultaneously models the local and global dependencies in the speech signal, significantly reducing the computational complexity while balancing performance and computational cost. Therefore, the MossFormer network has the ability to process semantic features of different scales in parallel. Server 104 can use multiple stacked blocks in this module, each of which includes substructures such as gated convolution, local attention, and global attention, which are used to model the local pronunciation features and overall vocalization pattern of the speaker's speech. For example, server 104 can obtain a mask representing the target speech in the network, that is, a weight matrix floating in the time and channel dimensions. The higher the value of the mask, the more server 104 believes that the channel information of the frame comes from the speaker that the listener is paying attention to. For example, the server 104 may use a speaker extraction module to estimate a mask M that only allows the speech of the concerned speaker to pass through. , while masked speech embedding The calculation formula is:
[0101]
[0102] Next, the server 104 may concatenate the mixed audio features along the channel dimension and EEG characteristics , and then apply a one-dimensional convolution to combine these features:
[0103]
[0104] The fused features are then fed into a speech separation model to extract the mask M.
[0105] Subsequently, the server 104 may further multiply the mask by the initially extracted audio coding features element-by-element, thereby filtering out unnoticed speech components and retaining only the feature regions most likely to belong to the target sound source, thereby forming an embedded feature vector of the target speech.
[0106] Furthermore, server 104 can feed this embedding vector into a decoder module. The decoder is a deconvolutional network with a symmetrical structure to the encoder, which restores the embedded features into continuous speech waveform data. Ultimately, the first predicted audio data output by server 104 is the target speaker's audio separated from the mixed speech based on the EEG attention state. For example, in a mixed speech containing male voice A and female voice B, if the listener's EEG response indicates that they are paying attention to female voice B, then the first predicted audio output by server 104 after the above processing is a restoration of the speech content of female voice B.
[0107] For example, the architecture diagram of the above MossFormer block is as follows: Figure 3 As shown in Figure 2, the MossFormer block consists of four convolutional modules, scaling and offset operations, a joint local and global single-head self-attention (SHSA) mechanism, and three gating operations.
[0108] Among them, the convolution module (ConvM) is used to normalize the input sequence, then project it to the linear layer, and finally use the SiLU activation function. Furthermore, one-dimensional depth convolution is applied to each feature for processing. The addition of skip connection and dropout algorithm helps to train and normalize the network.
[0109] Among them, the attention gating mechanism integrates the attention mechanism into the triple gating process to enhance the model capability. Given an input sequence , the convolutional module processes this sequence to generate two outputs and , as shown below:
[0110]
[0111] Using attention matrix , calculate the output sequence of this block
[0112] like:
[0113]
[0114]
[0115]
[0116] in, is the element-wise activation function, , .
[0117] For long sequences (i.e., large S), directly calculating the gated attention mechanism is computationally expensive. The server 104 can efficiently calculate the gated attention mechanism by combining the local and global attention mechanisms. and The server 104 may enter a sequence , projected into a shared representation via a convolutional neural network (ConvM):
[0118]
[0119] Next, the server 104 can apply the low-cost per-dimension scalar, offset, and RoPE (Rotary Position Embedding) to Z to generate the query , and key , , for local and global attention computation. For the global attention mechanism, the server 104 can adopt a linearized low-cost form to capture the long-range interaction of V and U:
[0120]
[0121] in, , is the scaling factor. For the local quadratic attention mechanism, V, U, Q, and K are divided into H non-overlapping segments of size P. The quadratic attention mechanism is calculated independently in each block:
[0122]
[0123] in, , is the scaling factor. Subsequently, the outputs of all blocks are concatenated along the time dimension to reconstruct the complete sequence. Finally, the outputs of the local and global attention mechanisms are summed to form the final joint attention mechanism:
[0124]
[0125] Furthermore, the speech decoder can be a one-dimensional transposed convolutional layer, using the same kernel size and stride as the encoder, for separating the feature sequence, which is finally decoded into a waveform by the decoder:
[0126]
[0127] The above module describes the single-step reasoning process of target speaker extraction based on EEG and audio. The model is represented as :
[0128]
[0129] Through the above steps, the server 104 can not only recognize the acoustic content of the speech, but also determine the user's auditory target with the help of EEG response, and accurately separate the sound source, thereby extracting the target of interest from both the auditory signal and the brain signal.
[0130] Step S208, calculating the main model loss based on the first predicted audio data and the original audio data of the target sound source, and training the EEG auditory speech extraction model based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0131] For example, server 104 can use the first predicted audio data output from the previous stage as the model's current output and compare it with the original audio data of the target sound source within the same time period. The original audio data of the target sound source is known clean speech in the training dataset, namely, a real speech track from a single speaker, unmixed and unprocessed. During the training process, server 104 pre-labeled each piece of data with the speaker of interest to the subject, allowing the corresponding target audio data to be directly retrieved from the annotated dataset as a reference. Server 104 can use the scale-invariant signal-to-noise ratio (SI-SDR), a signal-distortion ratio metric, as the basis for the loss function to measure the similarity between the model output and the real speech. This metric focuses on the consistency of speech shape, rather than overall loudness variations, making it particularly suitable for speech separation and enhancement tasks. Server 104 can input the first predicted audio data and the target audio data into the SI-SDR calculation module to generate a numerical score. The inverse of this score is then used as the main model loss (because the optimization goal is to maximize the SI-SDR, i.e., minimize negative values).
[0132] For example, the definition of SI-SDR may be:
[0133]
[0134] in, and Represent the extracted speech signal and the clean speech signal respectively. Main model loss function It can be:
[0135]
[0136] Furthermore, after obtaining the main model loss, server 104 can use this loss value as an optimization target to perform backpropagation in neural network training. This involves using automatic differentiation to propagate the error from the output layer back to each model module, including the audio encoder, EEG encoder, multimodal fusion network, and speaker mask predictor. Each backpropagation process updates the trainable parameters in these modules, such as convolution kernel weights, attention matrices, and gating coefficients, ensuring that the model's predictions in the next round are more accurate to the target audio. The training process is iterative, and server 104 can consider each parameter update as a training epoch. After each epoch, server 104 records the performance metrics of the current model on the validation set, such as the average SI-SDR, speech intelligibility, and speech quality metrics on the validation set. Server 104 can determine whether the model has reached convergence based on preset termination criteria. This termination criterion can include reaching a specified number of training epochs (e.g., 100 epochs), or the loss decreasing by less than a threshold (e.g., 0.001) for several consecutive epochs, or no further improvement in validation set metrics.
[0137] When the training process meets the termination conditions, the server 104 can solidify and save the model in the current parameter state, that is, obtain the trained EEG auditory speech extraction model. The model has the ability to accurately extract the target sound source voice from the input mixed voice and EEG data, and can be used for subsequent reasoning deployment scenarios. For example, when a user wears an EEG acquisition device and enters a multi-person conversation environment, as long as the real-time mixed voice and its EEG signal are input into the model, the server 104 can output the voice stream of the speaker that the user is paying attention to in real time, thereby enabling application devices such as smart hearing aids and brain-computer interface headphones, significantly improving their interactive capabilities in complex environments.
[0138] In the above-mentioned EEG auditory speech extraction model training method, EEG data and mixed speech data are acquired. The mixed speech data includes multiple raw audio data from different sound sources; the EEG data includes EEG signal data from multiple EEG channels. Feature extraction is performed on the mixed speech data to obtain audio coding features. Feature extraction is also performed on the EEG data to obtain EEG temporal features and EEG spatial features. Feature extraction and semantic modeling are performed on these two types of data respectively, enabling the system to effectively perceive and decode the listener's auditory attention direction in a multi-source speech environment. The system obtains audio coding features representing the audio content, EEG temporal features that vary over time, and EEG spatial features that reflect the coordinated activity patterns of different brain regions, thereby achieving a fine-grained description of the listener's dynamic auditory attention state. On this basis, the audio coding features are jointly modeled with the EEG temporal features and EEG spatial features to further generate speech embedding features. This embedded feature serves as an intermediate representation of the "speech that the listener is paying attention to" and has semantic information that contains both speech content and attention cues. By guiding the separation of mixed speech data through speech embedding features, the target sound source can be accurately extracted from the mixed sound sources to obtain the first predicted audio data. Then, by comparing the first predicted audio data with the original audio data corresponding to the target sound source and calculating the main model loss function, effective training of the EEG auditory speech extraction model can be achieved. By continuously optimizing the model and completing the training when the loss function meets the preset termination condition, a trained EEG auditory speech extraction model can be obtained that can jointly judge the object of auditory attention based on EEG and speech data. This model has good generalization ability and stability, and can automatically infer the source of the speech that the listener is paying attention to based on the EEG response.
[0139] In an exemplary embodiment, Figure 4 As shown, the above method may further include steps S302 to S310. In which:
[0140] Step S302: Divide the speech embedding features into a plurality of continuous feature segments.
[0141] Each feature segment corresponds to an original audio segment in the original audio data of the target sound source.
[0142] For example, server 104 can introduce a causal inference chain (CIC) to simultaneously predict multiple time steps while using hierarchical masks (parallel within each window, serial across windows) to strictly control the information flow. Server 104 can perform time division processing on the speech embedding features obtained above, dividing the entire speech embedding sequence into several continuous feature segments at equal intervals, each segment corresponding to a continuous speech signal in the original audio of the target sound source. For example, if the original target audio is a 9-second mixed conversation, server 104 can divide it into 9 feature segments, each corresponding to 1 second of content. These segments have a strict temporal sequence relationship, and their arrangement order is consistent with the actual playback order of the speech, ensuring that subsequent causal modeling conforms to the natural laws of semantic evolution. Each feature segment maintains a uniform number of time frames and feature channel dimensions, and retains the embedded semantic expression after the initial speech and EEG fusion.
[0143] Step S304 : For each feature segment in turn, a context fusion feature is obtained based on the feature segment and the previous feature segment.
[0144] For example, the server 104 can process each feature segment in turn to construct a contextual fusion feature. For the k-th feature segment, the server 104 can not only read the speech embedding of the segment itself, but also obtain the output embedding or processing status of the previous segment (k-1 segment). The server 104 can combine the embedding features of the current segment and the previous segment through a channel splicing operation to form a continuous contextual combination feature. Subsequently, this combination feature can be sent to a layer of Transformer module (which can be understood as a semantic fusion structure with global modeling capabilities), and the semantic dependency between the two segments is modeled through the self-attention mechanism, thereby outputting a contextual fusion feature with temporal continuity. This fusion feature not only contains the semantic information of the current segment, but also incorporates the judgment basis from the previous time segment, thereby providing a reference memory for target sound source separation.
[0145] For example, based on the above embodiments, Figure 5 The overall architecture of the proposed Mindflow model is shown in the figure, which consists of the main model and multiple causal reasoning chains The speaker extraction module uses electroencephalogram (EEG) as reference information to extract the audio features of the target speaker from the multi-speaker speech, which is then decoded into audio by the speech decoder. Among them, the context fusion module (CFM) refines the feature extraction of the current time step based on the extracted past speech. The speaker extraction module and the speech decoder module share parameters at each layer. CIC can be composed of a main model Modules and multiple Composition. With In comparison, each Each layer of the causal inference chain is trained based on the corresponding input. and Generate masked speech vector . Main Model You can directly Decode audio For the subsequent , the context information is integrated through CFM and then decoded into audio , context fusion module such as Figure 6 The calculation is as follows:
[0146]
[0147]
[0148]
[0149] in represents the parameters of the context fusion module, and TRM represents a single-layer Transformer module.
[0150] Step S306 : Separate the second predicted audio data of the target sound source from the mixed speech data according to the context fusion feature.
[0151] The above speech decoder can be a one-dimensional transposed convolution layer, using the same kernel size and stride as the encoder to separate the feature sequence, which is finally decoded into a waveform by the decoder:
[0152]
[0153] Exemplarily, the server 104 may obtain an initial sampling probability and an attenuation coefficient; calculate the sampling probability of the feature segment based on the initial sampling probability and the attenuation coefficient; select a context fusion feature or a feature segment as a target embedding feature based on the sampling probability; and separate the second predicted audio data of the target sound source from the mixed speech data based on the target embedding feature.
[0154] Furthermore, the server 104 can adopt adaptive teacher forcing to dynamically transition from supervised guidance to free prediction. During the training process, the next level receives the real speech embedding from the previous level. , and introduce a time-varying sampling probability To control the input source of the next time step. It follows an exponential decay law, as shown below:
[0155]
[0156]
[0157] in, Can be set to 0.9, each round to gradually transition from guided supervision to autonomous prediction.
[0158] For example, the server 104 may perform a second round of speech prediction and extraction operations on the target speaker based on the aforementioned context fusion features, and obtain an initial sampling probability (for example, set to = 0.9) and an attenuation coefficient (e.g. = 0.98), and then the actual sampling probability of this round of segments is calculated by exponential decay , that is, the probability that the segment uses the real embedding instead of the predicted embedding. The server 104 can use a random sampling mechanism to determine whether to use the current context fusion feature or directly use the feature segment itself as the target embedding feature. This random guidance mechanism constitutes a teacher forcing strategy, which can rely more on the real embedding in the early stage of training, and gradually delegate power to the model's own prediction in the later stage, thereby enhancing its generalization and adaptability. After determining the target embedding feature, the server 104 can send it to the speaker separation module, combine it with the original mixed speech, and use the mask mechanism to complete the second target sound source speech extraction, and output the second predicted audio data corresponding to the segment.
[0159] Step S308: Calculate the causal layer loss based on the second predicted audio data and the original audio segment corresponding to the feature segment.
[0160] For example, server 104 can calculate the difference between the second predicted audio data and the target original audio to obtain the prediction error for the causal segment. Server 104 can use the scale-invariant signal-to-noise ratio (SI-SDR) as an evaluation metric to calculate the semantic consistency between the model's prediction and the true target, and then back-transform this into a causal layer loss. Notably, to reflect the temporal importance of the causal inference chain, server 104 can assign different weights to different segments, for example, using an exponential weighting method so that later time periods contribute a larger proportion to the total loss, thereby strengthening the model's ability to learn long-term attention continuity.
[0161] Step S310: Train the EEG auditory speech extraction model according to the causal layer loss and the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0162] For example, the server 104 may combine the causal layer loss and the main model loss to form an overall loss function, and perform joint training optimization on all model parameters. For example, the definition of SI-SDR may be:
[0163]
[0164] in, and Represent the extracted speech signal and the clean speech signal respectively. Main model loss function It can be:
[0165]
[0166] Furthermore, to mitigate error propagation in the early stages of the causal inference chain, the server 104 may Introducing index weights , as shown below:
[0167]
[0168] Total loss Can be lost by the main model and hierarchical loss of causal chains The main model provides global predictions, while the causal chain refines local details.
[0169]
[0170] Through the backpropagation mechanism, server 104 performs end-to-end updates on all modules in the causal inference chain, including the feature segmentation module, context fusion module, mask prediction module, and speech decoding module. Server 104 repeats this process, traversing all training samples and all time segments. After each round of training, it evaluates the comprehensive performance indicators on the validation set to determine whether the preset termination conditions are met. When the total loss value stabilizes, the validation indicators show no significant improvement, or the maximum number of training rounds is reached, server 104 can save the current parameters as the final trained EEG auditory speech extraction model.
[0171] Through this causal inference chain mechanism, server 104 is able to not only extract the target sound source within a single time segment, but also maintain continuity and semantic coherence across the entire speech sequence, significantly improving the stability and robustness of target speaker extraction. In particular, server 104 can make more reasonable judgments based on the preceding and following context, especially when the subject's attention drifts slightly or the speech content changes. This causal inference capability is a key advantage that traditional static models lack.
[0172] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0173] Based on the same inventive concept, the embodiments of the present application also provide an EEG auditory speech extraction model training device for implementing the EEG auditory speech extraction model training method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more EEG auditory speech extraction model training device embodiments provided below can be found in the limitations of the EEG auditory speech extraction model training method above, and will not be repeated here.
[0174] In an exemplary embodiment, Figure 7 As shown, a device for training an EEG auditory speech extraction model is provided, comprising: a data acquisition module 702, a feature extraction module 704, a feature fusion module 706 and a model training module 708, wherein:
[0175] The data acquisition module 702 is used to acquire EEG data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels;
[0176] Feature extraction module 704 is used to extract features from the mixed speech data to obtain audio coding features, and to extract features from the EEG data to obtain EEG time features and EEG spatial features; wherein the EEG time features are used to characterize the change trend of the EEG signal data over time, and the EEG spatial features are used to characterize the change characteristics of the EEG signal data between different EEG channels;
[0177] A feature fusion module 706 is configured to fuse the audio coding features, the EEG temporal features, and the EEG spatial features to obtain a speech embedding feature, and to separate the first predicted audio data of the target sound source from the mixed speech data based on the speech embedding feature;
[0178] The model training module 708 is used to calculate the main model loss based on the first predicted audio data and the original audio data of the target sound source, and train the EEG auditory speech extraction model based on the main model loss until the preset termination condition is met to obtain a trained EEG auditory speech extraction model.
[0179] In one embodiment, the apparatus further comprises:
[0180] A speech processing module is used to divide the speech embedding feature into a plurality of continuous feature segments; each feature segment corresponds to an original audio segment of the original audio data of the target sound source;
[0181] A context fusion module is used to obtain context fusion features for each feature segment in turn based on the feature segment and the previous feature segment;
[0182] A causal processing module is configured to separate second predicted audio data of a target sound source from the mixed speech data based on the context fusion feature; and calculate a causal layer loss based on the second predicted audio data and the original audio segment corresponding to the feature segment;
[0183] The model training module 708 is specifically used to train the EEG auditory speech extraction model according to the causal layer loss and the main model loss until the preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0184] In one embodiment, the causal processing module is specifically used to: obtain an initial sampling probability and an attenuation coefficient; calculate the sampling probability of a feature segment based on the initial sampling probability and the attenuation coefficient; select a context fusion feature or a feature segment as a target embedding feature based on the sampling probability; and separate the second predicted audio data of the target sound source from the mixed speech data based on the target embedding feature.
[0185] In one embodiment, the feature extraction module 704 is specifically used to: perform time position encoding on the EEG data to obtain a first EEG data sequence; perform convolution calculation on the time axis for the first EEG data sequence to obtain EEG time features; perform channel position encoding on the EEG data to obtain a second EEG data sequence; perform convolution calculation between different EEG channels for the second EEG data sequence to obtain EEG spatial features.
[0186] In one embodiment, the feature fusion module 706 is specifically used to: perform self-attention alignment on the EEG time feature and the EEG spatial feature on the time axis to obtain the EEG signal feature; and fuse the EEG signal feature and the audio coding feature to obtain the speech embedding feature.
[0187] In one embodiment, the feature fusion module 706 is further used to: determine the number of speech frames based on the audio coding features; perform linear interpolation calculation on the EEG signal features based on the number of speech frames to update the EEG signal features; align and fuse the updated EEG signal features with the audio coding features to obtain speech embedding features.
[0188] Each module in the EEG auditory speech extraction model training device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0189] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store mixed speech data, electroencephalogram data, and data in steps such as feature extraction, feature fusion, and audio prediction. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for training an electroencephalogram and auditory speech extraction model is implemented.
[0190] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 9As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for training an EEG auditory speech extraction model. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0191] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0192] In an exemplary embodiment, a computer device is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: obtaining electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes multiple original audio data from different sound sources; the EEG data includes EEG signal data of multiple EEG channels; performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the EEG data to obtain EEG time features and EEG space features; wherein the EEG time features are used to characterize the changing trend of the EEG signal data over time, and the EEG space features are used to characterize the changing characteristics of the EEG signal data between different EEG channels; fusing the audio coding features, the EEG time features, and the EEG space features to obtain speech embedding features, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding features; calculating a main model loss based on the first predicted audio data and the original audio data of the target sound source, and training an EEG auditory speech extraction model based on the main model loss until a preset termination condition is met to obtain a trained EEG auditory speech extraction model.
[0193] In one embodiment, when the processor executes the computer program, the following steps are also implemented: dividing the speech embedding feature into a plurality of continuous feature segments; each feature segment corresponds to an original audio segment in the original audio data of the target sound source; obtaining a context fusion feature for each feature segment in turn based on the feature segment and the previous feature segment; separating the second predicted audio data of the target sound source from the mixed speech data based on the context fusion feature; calculating the causal layer loss based on the second predicted audio data and the original audio segment corresponding to the feature segment; training the EEG auditory speech extraction model based on the causal layer loss and the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
[0194] In one embodiment, when the processor executes the computer program, it also implements the following steps: obtaining an initial sampling probability and an attenuation coefficient; calculating the sampling probability of a feature segment based on the initial sampling probability and the attenuation coefficient; selecting a context fusion feature or a feature segment as a target embedding feature based on the sampling probability; and separating the second predicted audio data of the target sound source from the mixed speech data based on the target embedding feature.
[0195] In one embodiment, when the processor executes the computer program, it also implements the following steps: time-position encoding of the EEG data to obtain a first EEG data sequence; performing convolution calculation on the time axis for the first EEG data sequence to obtain EEG time features; channel-position encoding of the EEG data to obtain a second EEG data sequence; performing convolution calculation between different EEG channels for the second EEG data sequence to obtain EEG spatial features.
[0196] In one embodiment, when the processor executes the computer program, it also implements the following steps: performing self-attention alignment on the EEG temporal features and the EEG spatial features on the time axis to obtain EEG signal features; and fusing the EEG signal features and the audio coding features to obtain speech embedding features.
[0197] In one embodiment, when the processor executes the computer program, the processor also implements the following steps: determining the number of speech frames based on the audio coding features; performing linear interpolation calculation on the EEG signal features based on the number of speech frames to update the EEG signal features; aligning and fusing the updated EEG signal features with the audio coding features to obtain speech embedding features.
[0198] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0199] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0201] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0202] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0203] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for training an EEG auditory speech extraction model, characterized in that: The method comprises: Acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels; Performing feature extraction on the mixed speech data to obtain audio coding features, and performing feature extraction on the electroencephalogram data to obtain electroencephalogram time features and electroencephalogram space features; wherein the electroencephalogram time features are used to characterize the change trend of the electroencephalogram signal data over time, and the electroencephalogram space features are used to characterize the change characteristics of the electroencephalogram signal data between different electroencephalogram channels; fusing the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separating first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature; The main model loss is calculated based on the first predicted audio data and the original audio data of the target sound source, and the EEG auditory speech extraction model is trained based on the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
2. The method according to claim 1, characterized in that The method further comprises: Dividing the speech embedding feature into a plurality of continuous feature segments; each of the feature segments corresponds to an original audio segment in the original audio data of the target sound source; For each of the feature segments, a context fusion feature is obtained according to the feature segment and the previous feature segment; Separating second predicted audio data of a target sound source from the mixed speech data according to the context fusion feature; Calculating a causal layer loss based on the second predicted audio data and the original audio segment corresponding to the feature segment; The method of training the EEG auditory speech extraction model according to the main model loss until a preset termination condition is met to obtain a trained EEG auditory speech extraction model includes: The EEG auditory speech extraction model is trained according to the causal layer loss and the main model loss until a preset termination condition is met, thereby obtaining a trained EEG auditory speech extraction model.
3. The method according to claim 2, characterized in that The separating the second predicted audio data of the target sound source from the mixed speech data according to the context fusion feature includes: Get the initial sampling probability and attenuation coefficient; Calculating the sampling probability of the characteristic segment according to the initial sampling probability and the attenuation coefficient; Selecting the context fusion feature or the feature fragment as a target embedding feature according to the sampling probability; Second predicted audio data of a target sound source is separated from the mixed speech data according to the target embedded feature.
4. The method according to any one of claims 1 to 3, characterized in that The feature extraction of the EEG data to obtain EEG time features and EEG spatial features includes: Performing time position encoding on the electroencephalogram data to obtain a first electroencephalogram data sequence; Performing convolution calculation on the time axis for the first EEG data sequence to obtain EEG time features; performing channel position encoding on the electroencephalogram data to obtain a second electroencephalogram data sequence; For the second EEG data sequence, convolution calculation is performed between different EEG channels to obtain EEG spatial features.
5. The method according to claim 4, characterized in that The fusing of the audio coding feature, the EEG time feature, and the EEG spatial feature to obtain a speech embedding feature includes: Performing self-attention alignment on the EEG temporal feature and the EEG spatial feature on the time axis to obtain EEG signal features; The EEG signal features and the audio coding features are fused to obtain speech embedding features.
6. The method according to claim 5, characterized in that The fusing of the EEG signal feature and the audio coding feature to obtain a speech embedding feature includes: Determining the number of speech frames according to the audio coding characteristics; Performing linear interpolation calculation on the EEG signal feature according to the number of speech frames to update the EEG signal feature; The updated EEG signal features are aligned and fused with the audio coding features to obtain speech embedding features.
7. A device for training an EEG auditory speech extraction model, characterized in that: The device comprises: A data acquisition module, configured to acquire electroencephalogram (EEG) data and mixed speech data; wherein the mixed speech data includes a plurality of original audio data from different sound sources; and the EEG data includes EEG signal data of a plurality of EEG channels; A feature extraction module is used to perform feature extraction on the mixed speech data to obtain audio coding features, and to perform feature extraction on the EEG data to obtain EEG time features and EEG spatial features; wherein the EEG time features are used to characterize the change trend of the EEG signal data over time, and the EEG spatial features are used to characterize the change characteristics of the EEG signal data between different EEG channels; a feature fusion module, configured to fuse the audio coding feature, the EEG temporal feature, and the EEG spatial feature to obtain a speech embedding feature, and separate first predicted audio data of a target sound source from the mixed speech data based on the speech embedding feature; The model training module is used to calculate the main model loss based on the first predicted audio data and the original audio data of the target sound source, and train the EEG auditory speech extraction model based on the main model loss until a preset termination condition is met to obtain a trained EEG auditory speech extraction model.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Voice mixed signal separation method and device, storage medium and electronic equipment
CN113903354A
Multi-task target voice separation method and system fusing electroencephalogram signals
CN118737183A
Target person voice extraction method and device fusing voice and electroencephalogram signals
CN119049495A
Systems and methods for brain-informed speech separation
EP4226370A1
Speech chain apparatus, computer program, and DNN speech recognition / synthesis cross-learning method
JP2019120841A
Cited By
Streaming spatial audio separation method, streaming spatial audio separation equipment and vehicle-mounted audio system
CN121687094A