An EEG Semantic Visualization Method Based on the Diffusion Model Framework
Through the EEG semantic visualization method of the diffusion model framework, the EVR-Net and DDPM models are used to solve the individual differences in EEG semantic feature extraction and visualization, achieving high-quality semantic visualization effects, and enhancing the application potential of EEG control technology.
Patent Information
- Application Number
- CN202210788432.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-06
AI Technical Summary
The existing methods of semantic feature extraction and visualization of brain electroencephalograms have poor performance due to individual differences, making it difficult to effectively visualize semantic information in EEG data.
Using a diffusion model framework approach, EEG semantic features are extracted through EEG semantic feature extraction network model (EVR-Net), and visual images are generated using denoising diffusion probability model (DDPM), and data processing and image generation are combined with deep learning technology.
It realizes the effective correlation between EEG semantic features and visual images, has good generalization performance and high-quality visualization effects, and enhances the practical application potential of EEG control technology.
Smart Images

Figure CN115470810B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioelectroencephalogram signal, and specifically refers to an electroencephalogram semantic visualization method based on a diffusion model framework. Background Art
[0002] Electroencephalogram (EEG) is a brain response signal generated by humans under external stimuli. A large amount of semantic information is contained in the electroencephalogram, and this information can be used to visualize it into images with the same semantics. Usually, after semantic feature extraction, EEG data is used in various models for classification of states, emotions, image stimuli, etc. However, there is little work on visualizing the semantic feature information extracted from EEG data. And the visualization technology of EEG semantics has important prospects, such as solving communication barriers, multimedia applications, and grading discrimination of Alzheimer's disease. However, the individual differences of EEG lead to poor performance of current methods. Therefore, how to design effective EEG semantic feature extraction methods and EEG semantic feature visualization methods is very important. Summary of the Invention
[0003] In view of the deficiencies of the prior art, the present invention proposes an electroencephalogram semantic visualization method based on a diffusion model framework. An electroencephalogram semantic feature extraction network model (EVR-Net) based on deep learning is used to extract the electroencephalogram semantic features contained in EEG data, and the extracted semantic features are used as a guide to generate a visualization image with the same semantics as the EEG using a denoising diffusion probability model (DDPM).
[0004] To solve the above technical problems, the technical solution of the present invention is as follows:
[0005] An electroencephalogram semantic visualization method based on a diffusion model framework, comprising the following steps:
[0006] S1. Original data acquisition, collecting EEG data through an EEG acquisition device
[0007] S2. Data processing
[0008] S2-1. Data preprocessing, importing the collected EEG data into a python script and eliminating noise and interference;
[0009] S2-2. Time slice division, dividing the preprocessed EEG signal into several sequences with equal time lengths and storing them in a file with the format of npy;
[0010] S3. Obtaining a semantic feature vector e through the EVR-Net model
[0011] S3-1. Construct an initial EVR-Net model, where the initial EVR-Net model includes eight layers of networks, namely five convolutional layers, two residual blocks, and one average pooling layer;
[0012] S3-2. Train the initial EVR-Net model to obtain the optimal EEG semantic feature extractor EVR-Net parameters;
[0013] S3-3. Import the optimal EEG semantic feature extractor EVR-Net parameters into the initial EVR-Net model to obtain the final EVR-Net model;
[0014] S3-4. Input the processed EEG signal into the trained final EVR-Net model to obtain the semantic feature vector e;
[0015] S4. Obtain the visualization image x0 through the DDPM model
[0016] S4-1. Construct an initial DDPM model, where the initial DDPM model includes nine layers of networks, namely one fully connected layer, four convolutional layers, and four deconvolutional layers;
[0017] S4-2. Train the initial DDPM model to obtain the optimal semantic visualizer DDPM parameters;
[0018] S4-3. Import the optimal semantic visualizer DDPM parameters into the initial DDPM model to obtain the final DDPM model;
[0019] S4-4. Input the semantic feature vector e into the trained final DDPM model to obtain the visualization image x0.
[0020] Preferably, in step S2-1, the method for eliminating noise and interference is: calling the built-in band-pass filtering algorithm in the python library to extract EEG signals of 1-50 Hz.
[0021] Preferably, in step S3-2, before training the initial EVR-Net model, the parameters μ EVR-Net of the initial EVR-Net model need to be initialized:
[0022] The activation function of the initial EVR-Net model after the average pooling layer is the Sigmoid function, and its formula is as follows:
[0023]
[0024] Among them, x in the Sigmoid function represents the sample, and the activation function between other layers or blocks of the network is the ReLU function, and its formula is as follows:
[0025] ReLU(x) = max(0, x)
[0026] When the initial EVR-Net model is trained, the CrossEntropyLoss cross-entropy function is used as the loss function Lc, and its formula is as follows:
[0027]
[0028] In the cross-entropy function, i represents the classification situation under the current label state, K represents the total number of classifications, y represents the value of the label, and p represents the probability of semantic prediction under the current model parameters.
[0029] Preferably, in the step S3-2, μ in the initial EVR-Net model EVR-Net After initialization, the following training is carried out:
[0030] The data enters the first convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 3, and the stride is set to 2. The number of output channels is 32. After convolution, the data is activated by the activation function ReLU;
[0031] The data enters the second convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 3, and the stride is set to 2. The number of output channels is 64; after convolution, the data is activated by the activation function ReLU;
[0032] The data enters the third residual block, and the residual block with a convolutional kernel size of 3 is used to process the data. After processing, the data is activated by the activation function ReLU;
[0033] The data enters the fourth convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 3, and the stride is set to 2. The number of output channels is 128. After convolution, the data is activated by the activation function ReLU;
[0034] The data enters the fifth convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 3, and the stride is set to 2. The number of output channels is 256; after convolution, the data is activated by the activation function ReLU;
[0035] The data enters the sixth residual block, and the residual block with a convolutional kernel size of 3 is used to process the data. After processing, the data is activated by the activation function ReLU;
[0036] The data enters the seventh convolutional layer, and two-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 3, and the stride is set to 2. The number of output channels is 512. After convolution, the data is activated by the activation function ReLU;
[0037] The data enters the average pooling layer of the eighth layer. After passing through the average pooling layer, the data is converted into a tensor matrix of size 512, and the data is activated using the Sigmoid function;
[0038] The loss function Lc used in the training process is the CrossEntropyLoss cross-entropy function. Through the Lc loop training, the optimal EVR-Net parameters are obtained and saved to a local file.
[0039] Preferably, in step S4-2, before training the initial DDPM model, the parameter μ of the initial DDPM model DDPM is initialized as follows:
[0040] Define the standard deviation schedule β t = 1 - α t , and initialize it according to the linear formula. The formula is as follows:
[0041]
[0042] where T is the total number of initialization steps, t is a certain step number, both t and T values are integers. For the convenience of training the initial DDPM model, define as the product of α1...α t , and its formula is as follows:
[0043]
[0044] Define the value of x t to be obtained from the at the t-th step and the Gaussian noise ∈. The formula is as follows:
[0045]
[0046] The loss function L used during the training of the initial DDPM model DDPM has the following formula:
[0047]
[0048] where ∈ θ is the diffusion model function, x0 is the image, and e is the EEG semantic feature.
[0049] Preferably, in step S4-2, after the parameter μ of the initial DDPM model DDPM is initialized, the following training is performed:
[0050] The semantic feature vector e passes through a fully connected layer, and is vector concatenated with the mapped feature obtained by passing the current t value through a fully connected layer to obtain the EEG semantic fusion feature at the t-th step;
[0051] The xt enters the first convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 64. After convolution, the EEG semantic fusion feature position is embedded into the data, and the data is activated with the activation function ReLU;
[0052] The data enters the second convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 128. After convolution, the EEG semantic fusion feature position is embedded into the data, and the data is activated with the activation function ReLU;
[0053] The data enters the third convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 256. After convolution, the data is activated with the activation function ReLU;
[0054] The data enters the fourth convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 512. After convolution, the data is activated with the activation function ReLU;
[0055] The data enters the fifth deconvolutional layer, and two-dimensional deconvolution is used to perform deconvolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 256; after convolution, the data is activated with the activation function ReLU; then it is concatenated with the data obtained from the processing of the third convolutional layer, and the EEG semantic fusion feature position is embedded into the data;
[0056] The data enters the sixth deconvolutional layer, and two-dimensional deconvolution is used to perform deconvolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 128; after convolution, the data is activated with the activation function ReLU; then it is concatenated with the data obtained from the processing of the second convolutional layer, and the EEG semantic fusion feature position is embedded into the data;
[0057] The data enters the seventh deconvolutional layer, and two-dimensional deconvolution is used to perform deconvolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 64; after convolution, the data is activated with the activation function ReLU; then it is concatenated with the data obtained from the processing of the third convolutional layer;
[0058] The data enters the eighth deconvolutional layer, and two-dimensional deconvolution is used to perform deconvolution on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 3; after convolution, the data is activated with the activation function ReLU;
[0059] The loss function L used during training DDPM , through L DDPMThe optimal semantic visualizer DDPM parameters are obtained through iterative training and saved to a local file.
[0060] Preferably, it further includes
[0061] Step S5, testing and result evaluation,
[0062] The testing method is as follows: Input the EEG signal and Gaussian noise x T into the trained final EVR-Net model to obtain the semantic feature vector e; input the semantic feature vector e into the final DDPM model, and according to the Markov chain, let t iterate from T to 0 to obtain the visualization image x0. The Markov chain formula is as follows:
[0063]
[0064] where z is Gaussian noise, and δ t is the standard deviation of x T .
[0065] Preferably, the evaluation method in step S5 is as follows:
[0066] For the visualization image x0 generated in the testing stage, two metrics are used for evaluation:
[0067] The semantic accuracy of the generated visualization image x0, that is, whether the semantics of the generated visualization image x0 are the same as those of the EEG;
[0068] Use the image quality metric Inception Score for evaluation. The higher the value of the Inception Score, the higher the quality of the generated image;
[0069] Implementation of the evaluation metrics:
[0070] For the evaluation of semantic accuracy, count the number of correspondences between the semantics of the visualization image x0 and the EEG semantics;
[0071] For the image quality evaluation, input the visualization image x0 into a Python script. The Python script calls the Inception Score calculation algorithm built into the library to obtain the value of the Inception Score.
[0072] The present invention has the following features and beneficial effects:
[0073] With the above technical solution, first, the EEG semantic feature extractor EVR-Net is used to extract semantic features from EEG data, and then, guided by the EEG semantic features, the EEG is visualized as an image with the same semantics through DDPM. This solves the one-sided problem in conventional machine learning and deep learning methods that only focus on semantic recognition or image visualization, and associates the EEG semantics with the visualized image. Moreover, the present invention can achieve stable and high-quality results on different data sets and has good generalization performance.
[0074] In summary, the present invention has good effects in the semantic visualization of EEG, significantly improves the brain control technology in practical applications, and at the same time lays a foundation for future practical brain-computer interaction applications, having a broader application prospect. Brief Description of the Drawings
[0075] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts.
[0076] Figure 1 It is the flowchart of the method according to the embodiment of the present invention.
[0077] Figure 2 It is the overall network framework diagram of the training stage of the embodiment of the present invention.
[0078] Figure 3 It is the model diagram of the training stage of the EEG semantic feature extractor network model of the present invention.
[0079] Figure 4 It is the model diagram of the training stage of the EEG semantic visualizer DDPM of the present invention.
[0080] Figure 5 It is the overall network framework diagram of the testing stage of the embodiment of the present invention. Detailed Description of the Embodiments
[0081] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0082] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plural" is two or more.
[0083] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "installed", "connected", "connected to" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0084] The present invention provides an electroencephalogram semantic visualization method based on a diffusion model framework, as Figure 1 and Figure 2 shown, and includes the following steps:
[0085] S1. Acquisition of raw data: Collect EEG data through an EEG acquisition device.
[0086] Specifically, the acquisition of raw data is to build an image stimulation system through the electroencephalogram paradigm design platform Openvibe and the C# programming language, and design an image stimulation paradigm in the system. The data of this simulation system is recorded by the data receiving software EASYCAP corresponding to the acquisition device and forwarded to Openvibe, and Openvibe saves the data in CSV format. The sampling frequency of the EEG acquisition device is 128 Hz, there are 32 EEG electrode positions in total, both are referenced and grounded to the two earlobes of the FCz channel, and the impedance is kept below 10 KΩ.
[0087] Based on this system, a number of healthy young people with normal vision or corrected to normal vision were selected as subjects to collect EEG signals. The task was arranged in a quiet place without external interference. The subjects sat on a chair in a comfortable position, kept a distance of 60 cm from the computer display screen, and then put on the EEG data acquisition device. After the start instruction of the experiment was issued, they watched the paradigm images that appeared on the computer display screen. During the experiment, the paradigm images on the computer display screen were switched every 0.5 seconds, so as to achieve the effect of giving the subjects an event-related potential stimulus. When the image interface disappeared and the end prompt appeared, the task ended. This experiment included multiple groups of experimental acquisitions (26 groups).
[0088] S2. Data Processing
[0089] S2-1. Data preprocessing: Import the collected EEG data into a Python script and eliminate noise and interference.
[0090] Specifically, the EEG signals in the range of 1-50 Hz are extracted by calling the built-in filtering algorithm in the Python library, so as to eliminate noise and other interferences.
[0091] S2-2. Time slice division: The preprocessed EEG signals are segmented into several sequences with equal time lengths and stored in a file with the format of npy for subsequent model training.
[0092] S3. Obtain the semantic feature vector e through the EVR-Net model, where the EVR-Net model refers to the EEG semantic feature extraction network model, and its specific sub-steps are as follows:
[0093] S3-1. Construct the initial EVR-Net model, and the initial EVR-Net model includes eight layers of networks, namely four convolutional layers, three residual blocks and one average pooling layer.
[0094] S3-2. As Figure 3 shown, train the initial EVR-Net model to obtain the optimal EEG semantic feature extractor EVR-Net parameters.
[0095] It can be understood that the parameters μ of the initial EVR-Net model need to be initialized before training the initial EVR-Net model EVR-Net as follows:
[0096] The activation function of the initial EVR-Net model after passing through the average pooling layer is the Sigmoid function, and its formula is as follows:
[0097]
[0098] Among them, x in the Sigmoid function represents the sample, and the activation function between other layers or blocks of the network is the ReLU function. Its formula is as follows:
[0099] ReLU(x) = max(0, x)
[0100] When training the initial EVR-Net model, the CrossEntropyLoss cross-entropy function is used as the loss function Lc. Its formula is as follows:
[0101]
[0102] In the cross-entropy function, i represents the classification situation under the current label state, K represents the total number of classifications, y represents the value of the label, and p represents the probability of semantic prediction under the current model parameters.
[0103] Furthermore, μ in the initial EVR-Net model EVR-Net After completion of initialization, the following training is performed:
[0104] The data enters the first convolutional layer. One-dimensional convolution is used to perform convolution processing on the data. The convolution kernel size is 3, the stride is set to 2, the number of output channels is 32. After convolution, the data is activated by the activation function ReLU;
[0105] The data enters the second convolutional layer. One-dimensional convolution is used to perform convolution processing on the data. The convolution kernel size is 3, the stride is set to 2, the number of output channels is 64; after convolution, the data is activated by the activation function ReLU;
[0106] The data enters the third residual block. The residual block with a convolution kernel size of 3 is used to process the data. After processing, the data is activated by the activation function ReLU;
[0107] The data enters the fourth convolutional layer. One-dimensional convolution is used to perform convolution processing on the data. The convolution kernel size is 3, the stride is set to 2, the number of output channels is 128. After convolution, the data is activated by the activation function ReLU;
[0108] The data enters the fifth convolutional layer. One-dimensional convolution is used to perform convolution processing on the data. The convolution kernel size is 3, the stride is set to 2, the number of output channels is 256; after convolution, the data is activated by the activation function ReLU;
[0109] The data enters the sixth residual block. The residual block with a convolution kernel size of 3 is used to process the data. After processing, the data is activated by the activation function ReLU;
[0110] The data enters the convolutional layer of the seventh layer, and two-dimensional convolution is used to process the data. The convolution uses a convolutional kernel of size 3, the stride is set to 2, the number of output channels is 512, and after convolution, the ReLU activation function is used to activate the data;
[0111] The data enters the average pooling layer of the eighth layer. After passing through the average pooling layer, the data is converted into a tensor matrix of size 512, and the Sigmoid function is used to activate the data;
[0112] The loss function Lc used in the training process is the CrossEntropyLoss cross-entropy function. Through the loop training of Lc, the optimal EVR-Net parameters, that is, the optimal parameter μ EVR-Net is obtained and saved to a local file.
[0113] It should be noted that the data fed into the EVR-Net model is the EEG data read from the folder in the Python script.
[0114] It can be understood that the data of the first layer of the network is the EEG data read from the folder in the Python script, and the data fed into the second layer of the network is the data output by the first layer of the network, and so on.
[0115] S3-3. Import the optimal EEG semantic feature extractor EVR-Net parameters into the initial EVR-Net model to obtain the final EVR-Net model.
[0116] S3-4. Input the processed EEG signal into the trained final EVR-Net model to obtain the semantic feature vector e;
[0117] S4. Obtain the visualization image x0 through the DDPM model
[0118] S4-1. Construct an initial DDPM model. The initial DDPM model includes nine layers of networks, namely a fully connected layer, four convolutional layers, and four transposed convolutional layers.
[0119] S4-2. As Figure 4 shown, train the initial DDPM model to obtain the optimal semantic visualizer DDPM parameters.
[0120] It can be understood that before training the initial DDPM model, the parameters μ DDPM of the initial DDPM model need to be initialized:
[0121] Define the standard deviation schedule β t = 1 - α t and initialize it according to the linear formula. The formula is as follows:
[0122]
[0123] Where T is the total number of initialization steps, t is a certain step number, and both t and T are integers. For the convenience of training the DDPM model, define as the product of α1…α t The formula is as follows:
[0124]
[0125] Define the value of x t to be obtained from the at the t-th step and Gaussian noise ∈. The formula is as follows:
[0126]
[0127] The loss function L used during the training of the DDPM model DDPM The formula is as follows:
[0128]
[0129] Where ∈ θ is the diffusion model function, x0 is the image, and e is the EEG semantic feature.
[0130] Furthermore, the parameter μ of the DDPM model DDPM After completion of initialization, the following training is carried out:
[0131] The semantic feature vector e passes through a fully connected layer, and is concatenated with the mapped feature obtained by passing the current t value through a fully connected layer to obtain the EEG semantic fusion feature at the t-th step;
[0132] x t Enters the first convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution kernel size is 4, the stride is set to 2, the number of output channels is 64. After convolution, the EEG semantic fusion feature position is embedded into the data, and the data is activated using the activation function ReLU;
[0133] The data enters the second convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution kernel size is 4, the stride is set to 2, the number of output channels is 128. After convolution, the EEG semantic fusion feature position is embedded into the data, and the data is activated using the activation function ReLU;
[0134] The data enters the third convolutional layer, and two-dimensional convolution is used to perform convolution on the data. The convolution kernel size is 4, the stride is set to 2, the number of output channels is 256. After convolution, the data is activated using the activation function ReLU;
[0135] The data enters the fourth convolutional layer, and two-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, the number of output channels is 512, and after convolution, the ReLU activation function is used to activate the data;
[0136] The data enters the fifth transposed convolutional layer, and two-dimensional transposed convolution is used to perform transposed convolution processing on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 256; after convolution, the ReLU activation function is used to activate the data; then it is concatenated with the data obtained from the processing of the third convolutional layer, and the EEG semantic fusion feature positions are embedded into the data;
[0137] The data enters the sixth transposed convolutional layer, and two-dimensional transposed convolution is used to perform transposed convolution processing on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 128; after convolution, the ReLU activation function is used to activate the data; then it is concatenated with the data obtained from the processing of the second convolutional layer, and the EEG semantic fusion feature positions are embedded into the data;
[0138] The data enters the seventh transposed convolutional layer, and two-dimensional transposed convolution is used to perform transposed convolution processing on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 64; after convolution, the ReLU activation function is used to activate the data; then it is concatenated with the data obtained from the processing of the third convolutional layer;
[0139] The data enters the eighth transposed convolutional layer, and two-dimensional transposed convolution is used to perform transposed convolution processing on the data. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels is 3; after convolution, the ReLU activation function is used to activate the data;
[0140] The loss function L used during training DDPM , through L DDPM cyclic training to obtain the optimal semantic visualizer DDPM parameters and save them to a local file.
[0141] S4-3. Import the optimal semantic visualizer DDPM parameters into the initial DDPM model to obtain the final DDPM model;
[0142] S4-4. Input the semantic feature vector e into the trained final DDPM model to obtain the visualization image x0.
[0143] It can be understood that in the above technical solution, it is an artificial intelligence method that maps data to a high-dimensional space to find the correlation between data and extract semantic features. Compared with the artificial search for the correlation between data and extraction of semantic features, the deep learning method can complete the task quickly and accurately.
[0144] The diffusion model is an image generation model based on the likelihood function. This model divides the training process from a random image to noise into T training steps. Through the Markov chain, it learns the transformation process from the image to noise step by step in the reverse step, so that the noise can be transformed into a high-quality image in the forward step.
[0145] A further setting of the present invention is as Figure 5 shown, and further includes step S5, testing and result evaluation,
[0146] The testing method is to input the EEG signal and Gaussian noise x T into the trained EVR-Net model to obtain the semantic feature vector e; input the semantic feature vector e into the DDPM model, and according to the Markov chain, let t iterate from T to 0 to obtain the visual image x0. The Markov chain formula is as follows:
[0147]
[0148] where z is Gaussian noise and δ t is the standard deviation of x T .
[0149] Preferably, the evaluation method in step S5 is as follows:
[0150] For the visual image x0 generated in the test stage, two indicators are used for evaluation:
[0151] The semantic accuracy of the generated visual image x0, that is, whether the semantics of the generated visual image x0 are the same as those of the EEG;
[0152] Use the image quality metric Inception Score for evaluation. The higher the value of the Inception Score, the higher the quality of the generated image;
[0153] Implementation of the evaluation indicators:
[0154] For the evaluation of semantic accuracy, count the corresponding number of the semantics of the visual image x0 and the EEG semantics;
[0155] For the image quality evaluation, input the visual image x0 into the python script, and the python script calls the Inception Score calculation algorithm built into the library to obtain the value of the Inception Score.
[0156] It can be understood that through step S5, it can be ensured that the method provided by the present invention can achieve stable and high-quality effects on different data sets.
[0157] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments including components without departing from the principles and spirit of the present invention, and still fall within the protection scope of the present invention.
Claims
1. An electroencephalogram semantic visualization method based on a diffusion model framework, characterized in that It includes the following steps: S1. Original data acquisition: Collect EEG data through an EEG acquisition device. S2. Data processing S2-1. Data preprocessing: Import the collected EEG data into a Python script and eliminate noise and interference. S2-2. Time slice division: Divide the preprocessed EEG signal into several sequences of equal time length and store them in a file with the.npy format. S3. Obtain the semantic feature vector e through the EVR-Net model S3-1. Construct an initial EVR-Net model, where the initial EVR-Net model includes eight layers of networks, namely five convolutional layers, two residual blocks, and one average pooling layer. S3-2. Train the initial EVR-Net model to obtain the optimal EEG semantic feature extractor EVR-Net parameters. S3-3. Import the optimal EEG semantic feature extractor EVR-Net parameters into the initial EVR-Net model to obtain the final EVR-Net model. S3-4. Input the EEG signal after data processing into the trained final EVR-Net model to obtain the semantic feature vector e. S4. Obtain the visualization image x0 through the DDPM model S4-1. Construct an initial DDPM model, where the initial DDPM model includes nine layers of networks, namely one fully connected layer, four convolutional layers, and four deconvolutional layers. S4-2. Train the initial DDPM model to obtain the optimal semantic visualizer DDPM parameters. S4-3. Import the optimal semantic visualizer DDPM parameters into the initial DDPM model to obtain the final DDPM model. S4-4. Input the semantic feature vector e into the trained final DDPM model to obtain the visualization image x0.
2. The EEG semantic visualization method based on the diffusion model framework according to claim 1, wherein In the step S2-1, the method for eliminating noise and interference is: Call the built-in band-pass filtering algorithm in the Python library to extract EEG signals of 1-50 Hz.
3. The EEG semantic visualization method based on the diffusion model framework according to claim 1, wherein In the step S3-2, the parameter μ of the initial EVR-Net model needs to be initialized before the initial EVR-Net model is trained EVR-Net as follows: The activation function after the average pooling layer of the initial EVR-Net model is the Sigmoid function, and its formula is as follows: Among them, x in the Sigmoid function represents the sample, and the activation function between other layers or blocks of the network is the ReLU function, and its formula is as follows: ReLU(x) = max(0, x) When training the initial EVR-Net model, the CrossEntropyLoss cross-entropy function is used as the loss function Lc, and its formula is as follows: In the cross-entropy function, i represents the classification situation under the current label state, K represents the total number of classifications, y represents the value of the label, and p represents the probability of semantic prediction under the current model parameters.
4. The EEG semantic visualization method based on the diffusion model framework according to claim 3, characterized in that In the step S3-2, μ in the initial EVR-Net model EVR-Net After initialization, the following training is carried out: The data enters the first convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolution kernel of size 3, the stride is set to 2, the number of output channels is 32, and after convolution, the data is activated by the activation function ReLU. The data enters the second convolutional layer, and one-dimensional convolution is used to perform convolution processing on the data. The convolution uses a convolution kernel of size 3, the stride is set to 2, the number of output channels is 64; after convolution, the data is activated by the activation function ReLU; The data enters the residual block of the third layer, and the data is processed by the residual block with a convolution kernel size of 3. After processing, the data is activated by the activation function ReLU; The data enters the convolutional layer of the fourth layer, and the data is convolved by one-dimensional convolution. The convolution uses a convolution kernel of size 3, the stride is set to 2, and the number of output channels is 128. After convolution, the data is activated by the activation function ReLU; The data enters the convolutional layer of the fifth layer, and the data is convolved by one-dimensional convolution. The convolution uses a convolution kernel of size 3, the stride is set to 2, and the number of output channels is 256; after convolution, the data is activated by the activation function ReLU; The data enters the residual block of the sixth layer, and the data is processed by the residual block with a convolution kernel size of 3. After processing, the data is activated by the activation function ReLU; The data enters the convolutional layer of the seventh layer, and the data is convolved by two-dimensional convolution. The convolution uses a convolution kernel of size 3, the stride is set to 2, and the number of output channels is 512. After convolution, the data is activated by the activation function ReLU; The data enters the average pooling layer of the eighth layer. After passing through the average pooling layer, the data is converted into a tensor matrix of size 512, and the data is activated by the Sigmoid function; The loss function Lc used in the training process is the CrossEntropyLoss cross-entropy function. The optimal EVR-Net parameters are obtained through cyclic training with Lc and saved to a local file.
5. The EEG semantic visualization method based on the diffusion model framework according to claim 1, wherein In the step S4-2, before training the initial DDPM model, the parameter μ of the initial DDPM model needs to be initialized: DDPM Initialize: Define the standard variance plan β t = 1 - α t , and initialize it according to the linear formula as follows: where T is the total number of initialization steps, t is a certain step number, both t and T are integers. For the convenience of DDPM model training, define as the product of α1…α t as follows: Define x t The value of which is obtained from at the t-th step and Gaussian noise ∈, and the formula is as follows: The loss function L used during the initial training of the DDPM model DDPM The formula is as follows: where ∈ θ Diffusion model function, x0 is the image, and e is the EEG semantic feature.
6. The EEG semantic visualization method based on the diffusion model framework according to claim 5, characterized in that, In the step S4-2, the parameter μ of the initial DDPM model DDPM After completion of the initialization, the following training is performed: The semantic feature vector e passes through a fully connected layer, and is concatenated with the mapped feature obtained by passing the current t value through the fully connected layer to obtain the EEG semantic fusion feature at the t-th step; x t Enter the first convolutional layer, perform convolutional processing on the data using two-dimensional convolution. The convolution uses a convolutional kernel of size 4, sets the stride to 2, and the number of output channels to 64. After convolution, embed the EEG semantic fusion feature positions into the data, and activate the data using the ReLU activation function; The data enters the convolutional layer of the second layer, and the data is convolved by two-dimensional convolution. The convolution uses a convolution kernel of size 4, the stride is set to 2, and the number of output channels is 128. After convolution, the EEG semantic fusion feature position is embedded into the data, and the data is activated by the activation function ReLU; The data enters the convolutional layer of the third layer, and the data is convolved by two-dimensional convolution. The convolution uses a convolution kernel of size 4, the stride is set to 2, and the number of output channels is 256. After convolution, the data is activated by the activation function ReLU; The data enters the convolutional layer of the fourth layer, and the data is convolved by two-dimensional convolution. The convolution uses a convolution kernel of size 4, the stride is set to 2, and the number of output channels is 512. After convolution, the data is activated by the activation function ReLU; The data enters the transposed convolutional layer of the fifth layer, and the data is deconvolved by two-dimensional transposed convolution. The convolution uses a convolution kernel of size 4, the stride is set to 2, and the number of output channels is 256; after convolution, the data is activated by the activation function ReLU; then it is concatenated with the data processed by the convolutional layer of the third layer, and the EEG semantic fusion feature position is embedded into the data; The data enters the transposed convolutional layer of the sixth layer, and the data is deconvolved by two-dimensional transposed convolution. The convolution uses a convolution kernel of size 4, the stride is set to 2, and the number of output channels is 128; after convolution, the data is activated by the activation function ReLU; then it is concatenated with the data processed by the convolutional layer of the second layer, and the EEG semantic fusion feature position is embedded into the data; The data enters the seventh deconvolution layer, and is deconvolved using two-dimensional deconvolution. The convolution uses a convolution kernel of size 4, sets the step size to 2, and the number of output channels to 64. After the convolution, the data is activated using the activation function ReLU. The data is then concatenated with the data processed by the third convolution layer. The data enters the eighth deconvolution layer, and is deconvolved using two-dimensional deconvolution. The convolution uses a convolution kernel of size 4, sets the step size to 2, and the number of output channels to 3. After the convolution, the activation function ReLU is used to activate the data. The loss function L used during the training DDPM , through L DDPM perform cyclic training to obtain the optimal semantic visualizer DDPM parameters and save them to a local file.
7. The method for electroencephalogram semantic visualization based on the diffusion model framework according to claim 6, characterized in that, Also includes Step S5: testing and result evaluation, The test method is as follows: input the EEG signal and Gaussian noise x T into the finally trained EVR-Net model to obtain the semantic feature vector e; input the semantic feature vector e into the final DDPM model, and according to the Markov chain, let t iterate from T to 0 to obtain the visualized image x0. The Markov chain formula is as follows: where z is Gaussian noise, and δ t is the standard deviation of x T .
8. The method for visualizing EEG semantics based on the diffusion model framework according to claim 7, wherein The evaluation method in step S5 is as follows: For the visualization image x0 generated in the test phase, two indicators are used for evaluation: The semantic accuracy of the generated visualization image x0, that is, whether the semantics of the generated visualization image x0 is the same as that of EEG; The image quality indicator Inception Score is used for evaluation. The higher the Inception Score value, the higher the quality of the generated image. Evaluation indicators implementation: For the evaluation of semantic accuracy, the number of correspondences between the semantics of the visualized image x0 and the EEG semantics was counted; For image quality assessment, the visualized image x0 is input into the Python script, and the Python script calls the Inception Score calculation algorithm in the library to obtain the value of the Inception Score.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Methods and compositions using peptides and proteins with C-terminal elements
CN102869384A