Multi-channel SSVEP (Steady-State Visual Evoked Potential) generation method fusing neuron group model and diffusion model
Through the multi-channel SSVEP generation method that integrates the neuron group model and the diffusion model, the accuracy and efficiency of SSVEP signal generation in the prior art are solved. The generated signals are highly authentic and reliable, and support brain-computer interfaces and other applications.
Patent Information
- Application Number
- CN202510217894.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art faces accuracy and efficiency issues when generating SSVEP signals, especially in terms of mapping between complex neurodynamics and signal generation, resulting in scarce high-quality EEG data.
A multi-channel SSVEP generation method using a multi-channel coupled neuron group model and diffusion model is used to build a multi-channel coupled neuron group model and spatiotemporal feature extraction model, and combined with a denoising probability diffusion model, a multi-channel SSVEP signal with physiological significance is generated.
The generated signals are highly similar to the real signal in terms of frequency characteristics, time domain waveform and spatial distribution, improving the authenticity and reliability of the signal, solving the problem of scarcity of high-quality EEG signals, and supporting brain-computer interfaces and other applications.
Smart Images

Figure CN120143977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of brain-computer interfaces and electroencephalogram (EEG) signal processing, and particularly to a multi-channel SSVEP generation method that combines a neural population model and a diffusion model. Background Art
[0002] The acquisition methods of electroencephalogram (EEG) signals can be divided into two categories: invasive and non-invasive. The invasive method implants sensors through craniotomy. Although high-quality signals can be obtained, due to surgical risks and ethical controversies, its current application is relatively limited. The non-invasive method collects signals by placing electrodes on the scalp surface, and the common device is an electrode cap. Among them, the electrodes are divided into dry electrodes and wet electrodes: dry electrodes do not require conductive media and are convenient to operate, but the signal quality is low; wet electrodes improve the signal quality by applying conductive gel, but it may cause discomfort to the subjects. Despite the continuous progress of EEG acquisition technology, high-quality EEG data is still very scarce, mainly for the following reasons:
[0003] (1) Constructing a professional and efficient EEG experimental environment requires a large amount of capital investment, and the later maintenance cost cannot be ignored.
[0004] (2) Application research for EEG-BCI often requires a large amount of repeated experimental data. However, the experimental process is relatively complex, and multiple repetitions will take a large amount of time of the subjects, which limits the number of experimental rounds and results in insufficient scale of the existing publicly available EEG datasets.
[0005] (3) The physiological state, psychological state, environmental adaptability of the subjects, as well as external environmental factors will all affect the quality of the collected signals, making it extremely challenging to obtain high-quality and repeatable EEG data.
[0006] In the face of the problem of scarce EEG data, researchers have tried various data generation methods. For example, GAN, VAE, DDPM. Although the generation methods based on deep learning can generate high-quality signals, these methods often regard EEG signals as ordinary time series data and ignore the spatial topological characteristics between different brain regions. Therefore, although the generated signals are similar to real data in statistical characteristics, they lack the necessary neurophysiological significance.
[0007] The traditional NMM (neural mass model) describes the average behavior of neuron populations through mathematical equations and simulates the neural activities of the cerebral cortex. It is mainly used for the research of EEG rhythm generation mechanisms and the modeling of epileptic seizure dynamics. However, this model rarely considers the spatial conduction characteristics between different brain regions during the simulation process, and it is difficult to accurately depict the complex correlations between multi-channel EEG signals, thus limiting its application in the accurate simulation of multi-channel EEG signals.
[0008] In recent years, diffusion models have shown great potential in the field of signal processing. By constructing a forward diffusion process from raw data to pure noise and a corresponding reverse denoising process, this model can generate high-quality data samples. This method has achieved remarkable results in fields such as computer vision. However, applying diffusion models to EEG signal generation is still in the exploratory stage, mainly facing two major challenges: one is how to ensure that the generated signals have reasonable neurophysiological interpretability; the other is how to effectively model the spatial dependencies between multi-channel EEG signals. Summary of the Invention
[0009] To solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a multi-channel SSVEP generation method that integrates a neural population model and a diffusion model, combines the diffusion model with neurophysiological characteristic modeling, comprehensively considers the temporal dynamics, spatial topology, and physiological characteristics of signals, so as to generate more realistic, reliable, and physiologically meaningful EEG signals, thereby solving the accuracy and efficiency problems faced by the prior art in generating SSVEP signals, especially solving the difficulties in mapping between complex neural dynamics and signal generation. This not only helps to expand the existing EEG dataset but also provides better data support for applications such as brain-computer interfaces.
[0010] To achieve the above purpose, the present invention provides the following solutions:
[0011] A multi-channel SSVEP generation method that integrates a neural population model and a diffusion model, comprising:
[0012] Obtain preliminary multi-channel SSVEP signals and real SSVEP signals, perform data preprocessing on the real SSVEP signals, and extract spatial features and time-frequency features;
[0013] Use the multi-channel SSVEP signals to train a spatio-temporal denoising probability diffusion model, and optimize the model parameters in combination with a loss function to generate a trained spatio-temporal denoising probability diffusion model; the spatio-temporal denoising probability diffusion model includes: a forward noise-adding sub-model and a backward denoising sub-model;
[0014] During the process of training the spatio-temporal denoising probability diffusion model, input the multi-channel SSVEP signals into the forward noise-adding sub-model, add Gaussian noise to the multi-channel SSVEP signals, and calculate the corresponding noise level using randomly sampled time steps to generate noise-added samples. Input the noise-added samples into the backward denoising sub-model, extract the spatial features and time-frequency features of the real SSVEP signals based on a graph convolutional network and a long short-term memory network to form spatio-temporal features. Use a gate mechanism to fuse the noise-added samples, the spatio-temporal features, and the time step length, and denoise the fusion result.
[0015] Obtain the multi-channel SSVEP signals to be measured, and input the multi-channel SSVEP signals to be measured into the trained spatio-temporal denoising probability diffusion model to generate multi-channel SSVEP signals.
[0016] Optionally, obtaining the preliminary multi-channel SSVEP signals includes:
[0017] Obtain a neural population model, perform parallel weighted control on different cell subgroups in the neural population model with PSP blocks having dynamic characteristics to form a multi-dynamic neural population model, couple multiple multi-dynamic neural population models, obtain a multi-channel coupled neural population model, and use the multi-channel coupled neural population model to simulate the activities of the cerebral cortex region to generate preliminary multi-channel SSVEP signals.
[0018] Optionally, preprocessing the real SSVEP signals includes: normalizing the real SSVEP signals and suppressing noise for the characteristics of the SSVEP signals.
[0019] Optionally, extracting the spatial features includes:
[0020] Input the preprocessed real SSVEP signals into a graph convolutional network model, calculate the spectral information of each channel, and construct an adjacency matrix based on the cosine similarity between channels:
[0021]
[0022] where, X i (t), X i (t) represent the frequency features of channels i and j at time step t respectively;
[0023] Update the features of each channel by aggregating the information of neighboring nodes to generate the spatial features:
[0024]
[0025] where, represents the normalized adjacency matrix, W 0 and W 1 represent trainable weight matrices, σ represents the Sigmoid activation function, Relu represents the activation function, H t is the spatial feature, X t is the node feature matrix, representing the matrix composed of the frequency features of all channels at time step t.
[0026] Optionally, extracting the time-frequency features includes:
[0027] Input the spatial features at each time step into the long short-term memory network model to process the time dependence and generate the time-frequency features;
[0028] Processing the time dependence includes:
[0029] r t = σ(W r [H t , h t-1 + b r )
[0030] z t = σ(W z [H t , h t-1 + b z )
[0031]
[0032] where r t represents the reset gate, which determines how to utilize the memory of the previous moment; z t represents the update gate, which controls the weighted combination of the memory of the previous moment and the candidate state at the current moment; represents the candidate state, which is determined by the current input information and the memory information of the previous moment; h t represents the hidden state at the current moment, which is a weighted combination of the memory of the previous moment and the current candidate state, b r represents the bias term of the reset gate, b z represents the bias term of the update gate, b h represents the bias term of the candidate state, W r represents the trainable weight matrix of the reset gate, W z represents the trainable weight matrix of the update gate, W h represents the trainable weight matrix of the candidate state.
[0033] Optionally, forming the spatio-temporal features includes:
[0034]
[0035] TE(Δt) = cos(Δt / scale)
[0036] where O′ is the average value of the output of the multi-head attention mechanism, is the average value after adding time encoding, TE(·) is the time encoding function, Δt is the time step, and scale represents the hyperparameter.
[0037] Optionally, obtaining the average value of the output of the multi-head attention mechanism includes:
[0038]
[0039] Among them, L is the L heads, and α(Q (l) (i), K (l) (j)) represents the similarity between the query vector and the key vector, and O (l) (i) is the i-th query vector of the l-th head, and O (l) (i) is the i-th query vector of the l-th head, and V (l) (j) is the j-th value vector of the l-th head.
[0040] Optionally, fusing the noisy sample, the spatio-temporal feature, and the time step by using a gating mechanism includes:
[0041]
[0042] x′ t = W P x t ⊙σ(W c c) + W b c
[0043] Among them, PE(·) is the position embedding layer, t is the diffusion step, W P , W c , W b are learnable parameters, c is the extracted spatio-temporal feature vector, and σ(W c c) is the gating mechanism, x t is the noisy sample at the t-th step in the diffusion process, that is, the noisy signal at the current time step, and x′ t represents the updated state after fusing the spatio-temporal feature, that is, the new sample combined with the spatio-temporal feature condition in the denoising process.
[0044] The beneficial effects of the present invention are:
[0045] The present invention generates multi-channel SSVEP signals with neurophysiological interpretability and suitable for brain-controlled intelligent devices by fusing the neural population model and the spatio-temporal diffusion model.
[0046] The present invention first constructs a multi-channel coupled neural population model, starting from the neurophysiological mechanism, and accurately simulates the oscillatory activities and dynamic interaction processes between different regions of the cerebral cortex. On this basis, a conditional diffusion model is innovatively introduced to optimize the output of the neural population, and the spatio-temporal features extracted from real SSVEP data are used as conditional information to guide the signal generation process.
[0047] The present invention combines the advantages of the neural population model and the diffusion model: on the one hand, it fully exploits the simulation ability of the neural population model for the brain dynamics mechanism, ensuring that the generated signal has a reliable physiological interpretation; on the other hand, it utilizes the excellent modeling ability of the diffusion model for the statistical characteristics of real data, and significantly enhances the ability to depict the complex spatio-temporal characteristics of SSVEP signals through conditional constraints on spatio-temporal features. This complementary design not only guarantees the biological rationality of the generated signal, but also greatly improves the authenticity and reliability of the signal.
[0048] The SSVEP signal generation method of the present invention organically combines the biological rationality of the neural population model and the statistical characteristics of real data, and can generate real, reliable and high-quality SSVEP signals, solving the difficulties faced in the traditional EEG signal acquisition process. Especially in the case of scarce high-quality EEG signals, it can be widely applied in the field of brain-computer interfaces to meet the urgent need for high-quality EEG signals in this field. Brief Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic flowchart of a multi-channel SSVEP generation method that combines a neural population model and a diffusion model in an embodiment of the present invention;
[0051] Figure 2 It is a structural diagram of a basic neural population model in an embodiment of the present invention;
[0052] Figure 3 It is a schematic structural diagram of a multi-dynamic neural population model in an embodiment of the present invention;
[0053] Figure 4 It is a schematic structural diagram of a multi-channel coupled neural population model in an embodiment of the present invention;
[0054] Figure 5 It is a schematic structural diagram of a spatio-temporal feature extraction module in an embodiment of the present invention;
[0055] Figure 6 It is a schematic structural diagram of a conditional diffusion model in an embodiment of the present invention. Detailed Embodiments
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0058] This embodiment discloses a multi-channel SSVEP generation method that combines a neural population model and a diffusion model. This method uses a multi-channel coupled neural population model as the basis of the electrophysiological source model, combines a spatio-temporal feature extraction model, and guides the denoising probability diffusion model to generate SSVEP signals. Specifically, it includes the following steps: First, use the neural population model to simulate the dynamic behavior and complex interactions between different regions of the cerebral cortex, especially to model the frequency-specific responses in SSVEP signals; Second, analyze the real multi-channel EEG signals through the spatio-temporal feature extraction model, extract the time-domain and space-domain features of the signals, and obtain the frequency oscillation patterns and inter-regional functional connection characteristics of SSVEP signals; Finally, use the extracted spatio-temporal features as prior conditions to guide the reverse generation process of the denoising probability diffusion model, and through multi-step iterative denoising, gradually convert the output of the neural population model into the target SSVEP signal. The present invention can generate artificial SSVEP signals that are highly similar to real signals in terms of frequency characteristics, time-domain waveforms, and spatial distributions, providing effective data support for brain-computer interface research.
[0059] Specifically: construct a multi-channel coupled neural population model to simulate the dynamic behavior and complex interactions between different regions of the cerebral cortex; construct a spatio-temporal feature extraction model to extract spatio-temporal features from real multi-channel EEG signals; construct a denoising probability diffusion model, design a forward noise addition process and a backward denoising process, and use spatio-temporal features as conditions to realize signal generation through a denoising network and multi-step iteration, that is, the forward noise addition is a known process, and the backward denoising process trains the denoising network by fusing time steps, noisy data, and spatio-temporal features to gradually denoise.
[0060] Further, the construction process of the neural population model includes: based on the basic neural populations of excitatory neurons, pyramidal cells, and inhibitory neurons, form a basic neural population model; simulate the neural population interactions between regions of the cerebral cortex by coupling multiple basic neural population models, and model the frequency-specific responses in SSVEP signals.
[0061] Furthermore, the construction process of the spatio-temporal feature extraction model includes: performing time-frequency analysis on real SSVEP data to extract frequency response features; calculating the functional connection strength between different channels to construct a spatial feature matrix; and fusing the time-frequency features and spatial features to form a spatio-temporal feature representation.
[0062] Furthermore, the construction process of the denoising probability diffusion model includes: designing a forward noise addition process, gradually adding Gaussian noise to the signal using a predefined noise schedule; constructing a denoising network, including a time step encoding, feature extraction, and signal reconstruction module; using the extracted spatio-temporal features as prior conditions; and integrating feature information into the denoising process through an attention mechanism.
[0063] Furthermore, the training process of the denoising probability diffusion model includes: randomly selecting time steps for training samples based on the extracted spatio-temporal features as conditional information, calculating the corresponding noise level according to the noise schedule; sampling random noise to generate noisy samples; inputting the noisy samples, time steps, and conditional features into the denoising network; calculating the mean squared error loss between the predicted noise and the actual noise; and updating the network parameters using an optimizer.
[0064] Furthermore, the generation process of the denoising probability diffusion model includes: using Gaussian white noise as the starting signal; using the spatio-temporal features of the target SSVEP as conditional information; iterating over each time step: inputting the current signal, time step, and conditional features, predicting the noise through the trained denoising network, and updating the signal according to the predicted noise to obtain the finally generated multi-channel SSVEP signal.
[0065] Furthermore, it also includes the step of evaluating the generated signal: evaluating the frequency characteristics of the generated signal; evaluating the time-domain waveform of the generated signal; and evaluating the spatial distribution of the generated signal.
[0066] As Figure 1 shown, this embodiment discloses a multi-channel SSVEP generation method integrating a neural population model and a diffusion model, including:
[0067] Obtaining preliminary multi-channel SSVEP signals and real SSVEP signals, performing data preprocessing on the real SSVEP signals, and extracting spatial features and time-frequency features;
[0068] Training a spatio-temporal denoising probability diffusion model using the multi-channel SSVEP signals, and optimizing the model parameters in combination with a loss function to generate a trained spatio-temporal denoising probability diffusion model; the spatio-temporal denoising probability diffusion model includes: a forward noise addition sub-model and a backward denoising sub-model;
[0069] During the training of the spatio-temporal denoising probability diffusion model, the multi-channel SSVEP signal is input into the forward noise-adding sub-model, Gaussian noise is added to the multi-channel SSVEP signal, and the corresponding noise level is calculated using randomly sampled time steps to generate a noisy sample. The noisy sample is input into the backward denoising sub-model, and the spatial and time-frequency features of the real SSVEP signal are extracted based on the graph convolutional network and the long short-term memory network to form spatio-temporal features. Using the spatio-temporal features as prior conditions, a gating mechanism is used to fuse the noisy sample, the spatio-temporal features, and the time step, and the fused result is denoised;
[0070] Obtain the multi-channel SSVEP signal to be measured, and input the multi-channel SSVEP signal to be measured into the trained spatio-temporal denoising probability diffusion model to generate a multi-channel SSVEP signal.
[0071] Specifically: Taking the multi-channel coupled neuron population model as the basis of the electrophysiological source model, combined with the spatio-temporal feature extraction model, to guide the denoising probability diffusion model to generate SSVEP signals. The neuron population model can simulate the oscillatory activities of the cerebral cortex, especially the frequency-specific responses in SSVEP signals. The spatio-temporal feature extraction model analyzes the multi-channel electroencephalogram signals to obtain the time-frequency features and the functional connectivity characteristics between regions of the signals. These features are used as prior conditions to guide the generation process of the diffusion model. The diffusion model starts from the output of the neuron population model and gradually generates the target SSVEP signal through multi-step iterative denoising. This method can generate artificial SSVEP signals that are highly similar to real signals in terms of frequency characteristics, time-domain waveforms, and spatial distributions.
[0072] Furthermore, obtaining the preliminary multi-channel SSVEP signal includes:
[0073] Obtain the neuron population model, parallelly weight-control different cell subgroups in the neuron population model with PSP blocks having dynamic characteristics to form a multi-dynamic neuron population model, couple multiple multi-dynamic neuron population models, obtain a multi-channel coupled neuron population model, and use the multi-channel coupled neuron population model to simulate the activities of the cerebral cortex region to generate a preliminary multi-channel SSVEP signal.
[0074] Specifically: Construction of a multi-channel coupled neuron population model, construction of a dynamic spatio-temporal feature extraction model, construction of a spatio-temporal fusion model, construction of a spatio-temporal denoising probability model.
[0075] Construction of a multi-channel coupled neuron population model:
[0076] The Neuron Group Model (NMM) is a macroscopic modeling method that mainly studies the electrical activity of the brain by describing the overall characteristics of specific types of neuron populations, without the need to model individual neurons in the network structure. The core idea of NMM is "average area approximation", that is, the average behavior of the entire cell population in the neural network is represented by concentrated state variables. This method not only simplifies the complexity of the neuron system, but also has strong physiological significance, because it constructs a model of electroencephalogram signals from the perspective of the "organizational structure" of the nervous system.
[0077] Different from traditional single-neuron modeling methods, NMM reflects the interactions between neuron groups by modeling the group behavior. In NMM, the behavior of the neuron population is regarded as a whole, and the state of the population is described by a set of lumped variables, such as the average membrane potential and the average firing rate. This method can capture the macroscopic characteristics of brain activity and reflect the cooperative effects of neuron populations at a larger scale. In particular, the coupled NMM can simulate the interconnections between neuron populations, thus simulating large-scale neural network interactions at the macroscopic level.
[0078] In this embodiment, the Jansen-Rit basic neuron group model adopted is exactly based on this idea, which simplifies the complex neural network structure of the brain by describing the collective behavior of neuron populations. As Figure 2 shown, the basic model consists of three types of neuron sub-populations: excitatory interneuron population, pyramidal cell population, and inhibitory interneuron population, and they are mutually coupled through feedback loops.
[0079] In the NMM model, the three cell sub-populations that make up the neuron group are all composed of two modules. The first module is the activation function, which describes how to convert the presynaptic average membrane potential V(t) into the average impulse density S(t) of the population action potential, that is, the average firing rate. The second module is the synaptic response, which describes how to convert the average impulse density S(t) of the action potential into the average postsynaptic membrane voltage y(t), which can be excitatory or inhibitory. This module is called the postsynaptic potential (PSP) module.
[0080] The activation function is a Sigmoid function, specifically as follows:
[0081]
[0082] where V(t) is the presynaptic average membrane potential, S(t) is the average impulse density of the action potential, 2e 0 is the maximum firing rate, V 0 is relative to the ignition rate e 0The postsynaptic potential, where r represents the degree of curvature of the Sigmoid function. The electrical signal output by the Sigmoid function is an S-shaped curve, which changes fastest when V(t) is close to V 0 and changes slowest when V(t) is far from V 0 . The shape of this curve is determined by e 0 , r, and V 0 , reflecting the firing characteristics of neurons.
[0083] Each PSP block represents a linear transformation, divided into an excitatory transformation PSP e and an inhibitory transformation PSP i , and its impulse response is given by the following second-order constant coefficient linear differential equation:
[0084]
[0085] where τ is the time constant that summarizes the passive dendritic cable delay and neurotransmitter dynamics in the synapse, representing τ under excitatory transformation e or τ under inhibitory transformation i ; H is the average synaptic gain, used to adjust the maximum value of the postsynaptic membrane voltage, representing H under excitatory transformation e or H under inhibitory transformation i ; V(t) is the average membrane potential, representing the output signal of the neuron population; x(t) is the input signal of the neuron population, that is, the average impulse density of the action potential entering the neuron subpopulation, which is determined by the input of other subpopulations and external stimuli.
[0086] As Figure 2 shown, the excitatory and inhibitory interneuron populations receive excitatory feedback from the pyramidal cell population as input. The pyramidal cell population receives excitatory and inhibitory feedback from the excitatory and inhibitory interneuron populations, as well as input signals from other regions as input. The input signals from other regions are usually represented by Gaussian white noise N(μ, σ 2 ) passing through an excitatory transformation PSP e , where μ and σ 2 represent the mean and variance of the Gaussian white noise respectively. The parameters C1, C2, C3, and C4 represent the average number of synaptic connections between the subgroups in the neuron population model. The presynaptic average membrane potential of the pyramidal cell population is used as the output signal of the neuron population model, thus representing the neural oscillation activity.
[0087] In this embodiment, the parameters that determine the generation of the basic single-channel EEG signal are: synaptic gains H e and H i , time constants τ e and τ i, the number of synaptic connections C1, C2, C3, C4, and the mean μ and variance σ of Gaussian white noise 2 . These parameters jointly determine the dynamic behavior and output of the single-channel NMM model.
[0088] The basic NMM model can only reflect the dynamic characteristics of a relatively single cell population and generate narrowband signals of different frequencies. However, the spectrum of actual EEG signals is relatively wide, and signals of different rhythms reflect different behaviors of the brain. Therefore, a model that can generate a richer oscillator frequency is more reasonable. Therefore, in this embodiment, a multi-dynamic NMM model as shown in Figure 3 is constructed. Each cell subgroup is jointly composed of multiple excitatory or inhibitory cell subgroups, and different cell subgroups may have different dynamic characteristics, that is, the pyramidal cell subgroup, the excitatory and inhibitory interneuron subgroups are all composed of N PSP blocks PSP with different dynamic characteristics 1 ,...PSP N under parallel weighted control, and the weight coefficient W = {ω ij} ∈ [0, 1], and satisfies By adjusting the weight W and setting the excitatory and inhibitory parameters of N groups of different linear transformation blocks, the relative proportion of signals in cortical regions with different dynamic characteristics can be adjusted, thereby generating signals with richer frequencies and wider frequency bands. The connection parameters C1, C2, C3, C4 remain unchanged, that is, it is assumed that the connection structure of the pyramidal cell population and the interneuron population remains unchanged.
[0089] By coupling multiple dynamic NMM models, a multi-channel coupled neuron population model as shown in Figure 4 can be constructed, which can generate results similar to multi-channel EEG signals. Figure 4 The coupling signal of the j-th channel to the k-th channel in
[0090]
[0091] where v jk represents the coupling coefficient of the j-th channel to the k-th channel, RM(x) = x - mean(x) is the de-mean function, S is the non-linear Sigmoid function, δ is the delay time, and i = 1, 2,..., N represents that there are N subgroups with different dynamic characteristics in parallel in each channel, represents the first state variable of the i-th oscillator in the j-th channel, that is, the postsynaptic membrane voltage of the excitatory interneuron population, is the second state variable of the i-th oscillator in the j-th channel, that is, the postsynaptic membrane voltage of the inhibitory interneuron population. By setting the coupling coefficient matrix V = {v jk}, j = 1, ..., M, k = 1, ..., M, where M represents the number of channels of the generated EEG signals, and different intensities of coupling between different channels can be achieved.
[0092] After the construction of the multi-channel coupled neural population model, it is first necessary to determine the initial parameters. Generally, there are the following several methods for constructing the initial parameters:
[0093] Method based on experimental data: Use real EEG signals to estimate the parameters.
[0094] Method based on stochastic process: Generate the initial parameters by using a stochastic process, and then use them as the initial state of the neural population model.
[0095] Method based on optimization: Use an optimization algorithm to find the optimal initial parameters and use them as the initial parameters of the neural population model.
[0096] In this embodiment, the method based on experimental data is used to estimate the internal parameters of the neural population model. The main reason is that the internal parameters of the neural population model are described by complex non-linear differential equations. Using a stochastic process to generate parameters may lead to problems such as parameter mismatch, instability, and uncontrollability. Although the optimization algorithm can provide a better solution, its calculation process is usually time-consuming. Therefore, the method based on experimental data becomes a more suitable choice to ensure the accuracy and stability of the parameters.
[0097] After the parameters are determined, the constructed neural population model can be used for the synthesis of EEG signals. However, the generated multi-channel time series signals only reflect a single dynamic behavior, do not consider the coupling effect between different channels, and are also difficult to fully reflect the spatial characteristics of EEG signals. Some head models, such as the actual anatomical model constructed based on individual magnetic resonance imaging (MRI) data, can accurately reflect the geometric shape and conductivity distribution of the head, thereby improving the accuracy of the generated data. However, the actual anatomical model depends on magnetic resonance imaging, which not only has a high acquisition cost and large calculation amount, but also has more strict requirements for the electrode position and direction.
[0098] Therefore, in order to overcome the problem of large calculation amount of the actual anatomical model, a spatio-temporal denoising probability diffusion model is adopted. By introducing this model, not only can the spatio-temporal characteristics of the signal be optimized and denoised, but also the coupling effect and signal propagation process between different regions of the brain can be accurately simulated while ensuring the calculation efficiency. Finally, this method helps to generate more real and reliable multi-channel SSVEP signals.
[0099] Furthermore, the data preprocessing of the real SSVEP signal includes: normalizing the real SSVEP signal and suppressing the noise according to the characteristics of the SSVEP signal.
[0100] Furthermore, the extraction of spatial features includes:
[0101] Input the preprocessed real SSVEP signal into the graph convolutional network model, calculate the spectral information of each channel, construct an adjacency matrix based on the cosine similarity between channels, and update the features of each channel by aggregating the information of neighboring nodes to generate spatial features:
[0102] Furthermore, the extraction of time-frequency features includes: Input the spatial features at each time step into the long short-term memory network model to process time dependence and generate time-frequency features.
[0103] Specifically: As Figure 5 shown, the spatio-temporal feature extraction module:
[0104] The spatio-temporal feature extraction module is mainly responsible for processing real multi-channel SSVEP signals. First, it is necessary to perform data preprocessing on it. It is necessary to standardize the input signal to eliminate the amplitude difference between different channels, and perform band-pass filtering for noise suppression according to the characteristics of the SSVEP signal.
[0105] In this embodiment, spatial information is extracted from real SSVEP data through a graph convolutional network (GCN). First, for each time step t, calculate the spectral information of each channel, and construct an adjacency matrix based on the cosine similarity between channels:
[0106]
[0107] where, X i (t), X i (t) represent the frequency features of channels i and j at time step t respectively. The GCN layer updates the features of each channel by aggregating the information of neighboring nodes, representing the features after the spatial convolution operation of each channel. The output H t of GCN is calculated by the following formula:
[0108]
[0109] where, is the normalized adjacency matrix, W 0 and W 1 are trainable weight matrices, σ is the Sigmoid activation function, and Relu is the activation function.
[0110] In this embodiment, a long short-term memory network (LSTM) is used to process time-series data and extract the time dependence features of SSVEP. In LSTM, the reset gate, update gate, and candidate state are the key parts to control the information flow. The spatial features H tAs the input of the LSTM and processed through the gating mechanism of the LSTM to handle temporal dependencies:
[0111] r t = σ(W r [H t , h t-1 + b r ) (6)
[0112] z t = σ(W z [H t , h t-1 + b z ) (7)
[0113]
[0114]
[0115] where r t is the reset gate, which determines how to utilize the memory from the previous moment; z t is the update gate, which controls the weighted combination of the memory from the previous moment and the candidate state at the current moment; is the candidate state, which is determined by the current input information and the memory information from the previous moment; h t is the hidden state at the current moment, which is the weighted combination of the memory from the previous moment and the current candidate state.
[0116] As Figure 6 shown, the spatio-temporal denoising probability diffusion model is constructed:
[0117] In this embodiment, the spatio-temporal denoising probability diffusion model is the core component for generating high-quality SSVEP signals. During the training and generation processes of the model, the forward noise addition process is first executed. This process controls the noise injection intensity through a predefined noise schedule and gradually adds Gaussian noise to the input signal. Specifically, for each training sample, the corresponding noise level is calculated using randomly sampled time steps, and then the noisy samples are generated. The design of the forward process ensures the gradual accumulation of noise, providing a basis for the subsequent denoising process.
[0118] The backward denoising process is the core part of the model and is responsible for gradually converting the noisy signal into the target SSVEP signal. This process first receives the signal containing noise, and at the same time inputs the corresponding time step information and spatio-temporal conditional features. Through the trained denoising network, the model can predict the noise component in the current signal. At each time step, the network improves the signal quality by removing the predicted noise. This process is iterated repeatedly until a clear target signal is obtained.
[0119] In this embodiment, two Transformer encoders based on multi-head attention are used to aggregate temporal and spatial information respectively. Considering the input x t , the following equation is obtained according to the multi-head attention mechanism:
[0120] Q (l) = x t W q , K (l) = x t W k , V (l) = x t W v (10)
[0121]
[0122] where Q (l) , K (l) , V (l) are the query vector, key vector, and value vector of the l-th head respectively, W q , W k , W v are learnable parameters, α(Q (l) (i), K (l) (j)) represents the similarity between the query vector and the key vector, Q (l) (i) represents the i-th query vector of the l-th head, K (l) (j) represents the j-th key vector of the l-th head, Q (l) (i) T K (l) (j) represents the dot product of the query vector and the key vector (i.e., the similarity metric), and n represents the number of query vectors. The output is:
[0123]
[0124] where O (l) (i) is the i-th output vector of the l-th head, and M is the total number of key-value pairs.
[0125] For L heads, the average value is obtained by averaging:
[0126]
[0127] where L represents the number of multi-head attentions. In the temporal Transformer encoding, the above formula can be gradually applied in the temporal dimension to generate O′ at the next time step.
[0128] In this embodiment, TE(·) is used as the temporal encoding function to add temporal information to the multi-channel node representation:
[0129]
[0130] TE(Δt) = cos(Δt / scale) (15)
[0131] Where Δt is the time step, and scale represents a hyperparameter used to control the frequency and amplitude of encoding.
[0132] The conditional generation mechanism is the key to ensuring the characteristics of the generated signal. This mechanism first injects the extracted spatio-temporal feature vector into the denoising network processing flow. This design enables the denoising process to always meet the constraint conditions of the spatio-temporal characteristics of the real SSVEP signal, ensuring that the generated signal has the required frequency and spatial characteristics. In this embodiment, a gate mechanism is used to fuse the noise input, time step, and learned spatio-temporal features, and the fused information is used as the input of the denoising process to more accurately predict the noise:
[0133]
[0134] x′ t = W P x t ⊙σ(W c c)+W b c (17)
[0135] Where PE(·) is the position embedding layer, t is the diffusion step, W P , W c , W b are learnable parameters, c is the extracted spatio-temporal feature vector, and σ(W c c) is the gating mechanism that controls which channels, time steps, and frequency bands of the SSVEP signal are more important for the output.
[0136] In this embodiment, two multi-head attention-based Transformer encoders are used to process the fused feature x′ t , and output the predicted noise ∈ θ , that is:
[0137] ∈ θ = Transformer(x t ′) (18)
[0138] The denoising network makes the predicted ∈ θ as close as possible to the real noise ∈ added in the forward process by learning the parameter θ.
[0139] In this embodiment, the model is trained using randomly sampled time steps, and the loss function L diff is constructed by calculating the difference between the predicted noise and the real noise, and the network parameters are updated using an optimizer to minimize this loss:
[0140]
[0141] wherein, denotes the calculation of the expected value with respect to noise ∈, the original sample x 0 and the time step t. ∈ represents the noise sampled from a standard Gaussian distribution during the training process, the model needs to learn to recover the real sample from this noise, x 0 represents the original sample data, i.e., the output signal of the NMM model, α t represents the time step parameter in the diffusion process, which controls the intensity of adding noise at each step, ∈ θ represents the output of the denoising model, i.e., the predicted noise, represents the remaining part of the original sample at the current step, represents the noise part at the current time step, ||·|| 2 represents the square of the L2 norm.
[0142] In the signal generation stage, the model starts from the simulated multi-channel SSVEP data output by the NMM, and iteratively executes the denoising step under the guidance of conditional features to predict the noise ∈ θ . Update the signal according to the predicted noise:
[0143]
[0144] wherein, the final denoising result x 0 is the generated multi-channel SSVEP signal with target features. This design not only ensures the quality of the generated signal, but also guarantees the stability and controllability of the generation process.
[0145] In this embodiment, a multi-channel coupled neural population model is established, the initial values of the parameters of the neural population model are determined using experimental data to generate initial macroscopic multi-channel data; a spatio-temporal feature extraction model is established to extract the time-domain and space-domain features of the real SSVEP signal; the output of the NMM model is used as the input, and the extracted spatio-temporal features are used as prior conditions to guide the reverse generation process of the denoising probability diffusion model. Through multi-step iterative denoising, the output of the neural population model is gradually converted into the finally generated SSVEP signal.
[0146] This embodiment mainly runs on a computer device, can generate high-quality electroencephalogram signals urgently needed in the current brain-computer interface field, and improves the difficult situation of electroencephalogram signal acquisition and the small number of high-quality electroencephalogram signals in the traditional method. In the foreseeable future, this embodiment can be widely applied in the field of brain-computer interface.
[0147] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multi-channel SSVEP generation method integrating a neuron group model and a diffusion model, characterized in that: include: Acquiring a preliminary multi-channel SSVEP signal and a real SSVEP signal, performing data preprocessing on the real SSVEP signal, and extracting spatial features and time-frequency features; The multi-channel SSVEP signal is used to train a spatiotemporal denoising probability diffusion model, and the model parameters are optimized in combination with a loss function to generate a trained spatiotemporal denoising probability diffusion model; the spatiotemporal denoising probability diffusion model includes: a forward denoising sub-model and a backward denoising sub-model; In the process of training the spatiotemporal denoising probability diffusion model, the multi-channel SSVEP signal is input into the forward denoising sub-model, Gaussian noise is added to the multi-channel SSVEP signal, and the corresponding noise level is calculated using a random sampling time step to generate a noisy sample, the noisy sample is input into the backward denoising sub-model, and the spatial features and time-frequency features of the real SSVEP signal are extracted based on a graph convolutional network and a long short-term memory network to form a spatiotemporal feature, the spatiotemporal feature is used as a priori condition, a gate mechanism is used to fuse the noisy sample, the spatiotemporal feature and the time step, and the fusion result is denoised; A multi-channel SSVEP signal to be tested is obtained, and the multi-channel SSVEP signal to be tested is input into a trained spatiotemporal denoising probability diffusion model to generate a multi-channel SSVEP signal.
2. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: Acquiring the preliminary multi-channel SSVEP signal comprises: A neuron group model is obtained, and different cell subgroups in the neuron group model are weighted and controlled in parallel with a PSP block with dynamic characteristics to form a multi-dynamic neuron group model, and a plurality of multi-dynamic neuron group models are coupled to obtain a multi-channel coupled neuron group model, and the multi-channel coupled neuron group model is used to simulate the activity of the cerebral cortex area to generate a preliminary multi-channel SSVEP signal.
3. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: The data preprocessing of the real SSVEP signal includes: standardizing the real SSVEP signal and performing noise suppression on the characteristics of the SSVEP signal.
4. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: Extracting the spatial features includes: The preprocessed real SSVEP signal is input into the graph convolutional network model, the spectrum information of each channel is calculated, and the adjacency matrix is constructed based on the cosine similarity between channels: Among them, X i (t),X i (t) represents the frequency characteristics of channel i, j at time step t; The features of each channel are updated using the information of aggregated neighbor nodes to generate the spatial features: in, represents the normalized adjacency matrix, W0 and W1 represent the trainable weight matrices, σ represents the Sigmoid activation function, Relu represents the activation function, H t is the spatial feature, X t is the node feature matrix, which represents the matrix composed of the frequency features of all channels at time step t.
5. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: Extracting the time-frequency features includes: Inputting the spatial features at each time step into the long short-term memory network model, processing the time dependency, and generating the time-frequency features; Processing the time dependency includes: r t =σ(W r [H t ,h t-1 ]+b r ) z t =σ(W z [H t ,h t-1 ]+b z ) Among them, r t Represents the reset gate, which determines how to use the memory of the previous moment; z t represents the update gate, which controls the weighted combination of the memory of the previous moment and the candidate state of the current moment; represents the candidate state, which is determined by the current input information and the memory information of the previous moment; h t represents the hidden state at the current moment, which is a weighted combination of the memory at the previous moment and the current candidate state, b r represents the bias term of the reset gate, b z represents the bias term of the update gate, b h represents the bias term of the candidate state, W r represents the trainable weight matrix of the reset gate, W z represents the trainable weight matrix of the update gate, W h A trainable weight matrix representing the candidate states.
6. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: Forming the spatiotemporal features includes: TE(Δt)=cos(Δt / scale) Among them, O′ is the average value of the output of the multi-head attention mechanism, is the average value after adding time encoding, TE(·) is the time encoding function, Δt is the time step, and scale represents a hyperparameter.
7. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 6 is characterized in that: Obtaining the average value of the multi-head attention mechanism output includes: Among them, L is L heads, α(Q (l) (i),K (l) (j)) represents the similarity between the query vector and the key vector, O (l) (i) is the i-th query vector of the l-th head, O (l) (i) is the i-th query vector of the l-th head, V (l) (j) is the j-th value vector of the l-th head.
8. The multi-channel SSVEP generation method of the fusion neuron group model and the diffusion model according to claim 1 is characterized in that: Using a gate mechanism to fuse the noise-added sample, the spatiotemporal features, and the time step includes: x t ′=W P x t ⊙σ(W c c)+W b c Where PE(·) is the position embedding layer, t is the number of diffusion steps, and W P ,W c ,W b is a learnable parameter, c is the extracted spatiotemporal feature vector, σ(W c c) is the gating mechanism, x t is the noise sample of the tth step in the diffusion process, that is, the noisy signal of the current time step, x t ′ represents the updated state after the fusion of spatiotemporal features, that is, the new sample that combines the spatiotemporal feature conditions in the denoising process.