Fine-grained image generation method and system based on multilevel electroencephalogram semantic features

Through the fine-grained image generation method of multi-level electroencephalogram semantic features, the high-level and low-level semantic features in the EEG signal are extracted using a multi-channel convolutional neural network, and the diffusion model is conditionally regulated, which solves the problem of semantic consistency and low fidelity of image generation in the prior art, and achieves high-quality and rich in detail image generation.

CN120125691AInactive Publication Date: 2025-06-10HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510190917.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art ignores the multi-level association between low-level and high-level semantic features in EEG signals in fine-grained image generation, resulting in low semantic consistency and fidelity of the generated images.

Method used

A fine-grained image generation method of multi-level electroencephalogram semantic features is adopted to extract high-level and low-level semantic features from preprocessed electroencephalogram signal data through a multi-channel convolutional neural network, and these features are used to conditionally regulate the diffusion process of the diffusion model to generate images with consistent semantics and rich details.

Benefits of technology

The fine control of the image generation process is realized, and the generated images have high semantic consistency and rich details, which significantly improves the quality and generalization ability of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125691A_ABST
    Figure CN120125691A_ABST
Patent Text Reader

Abstract

The invention provides a fine-grained image generation method and system based on multilevel electroencephalogram semantic features, and relates to the technical field of artificial intelligence, and the method comprises the steps: collecting electroencephalogram signal data; performing band-pass filtering and time slice segmentation on the electroencephalogram signal data to obtain preprocessed signals; a multi-channel convolutional neural network is adopted to extract multi-level semantic features from the preprocessed signals; the multi-level semantic features comprise high-level semantic features and low-level semantic features; and performing condition regulation and control on a diffusion process of a preset diffusion model by utilizing the multi-level semantic features so as to generate a fine-grained image which is consistent with the electroencephalogram signal data semantics. The problem that in the prior art, most methods only focus on high-level category related semantics, and consequently the semantic consistency and fidelity of generated images are low is solved. Experimental results on different data sets show that the method has good performance and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a fine-grained image generation method and system based on multi-level EEG semantic features. Background Art

[0002] Electroencephalogram (EEG) captures neural signals of brain activities in a non-invasive manner, with high temporal resolution, and can reflect the brain state during visual perception and imagination. Rich semantic information, such as category, color, contour, etc., is contained in EEG signals, and these information can provide a basis for decoding visual perception and imagination. However, existing methods mainly focus on decoding high-level semantic information (such as object category), and less on processing low-level detail information (such as color and contour).

[0003] Currently, for EEG signal processing methods based on deep learning, convolutional neural networks (CNNs) or recurrent neural networks (RNNs) are mostly used for classification tasks or generating low-resolution images related to categories. These methods often ignore the multi-level associations between low-level and high-level semantic features in EEG signals, restricting their applications in fine-grained image generation.

[0004] As a new type of image generation framework, diffusion models have received extensive attention in recent years due to their high-quality performance in image generation tasks. Diffusion models can generate high-quality images consistent with the input conditions by gradually generating data from a noise distribution. Therefore, introducing multi-level semantic features in EEG signals into diffusion models to achieve fine-grained image generation has important technical significance. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the purpose of the present invention is to provide a fine-grained image generation method and system based on multi-level EEG semantic features, which realizes fine control of the image generation process, thereby generating high-quality images with consistent semantics and rich details.

[0006] To achieve the above purpose, the present invention provides the following solutions:

[0007] A fine-grained image generation method based on multi-level EEG semantic features, comprising:

[0008] Collecting EEG signal data;

[0009] Performing band-pass filtering and time slicing on the EEG signal data to obtain preprocessed signals;

[0010] Extracting multi-level semantic features from the preprocessed signals by using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features;

[0011] Use the multi-level semantic features to conditionally regulate the diffusion process of a preset diffusion model to generate fine-grained images consistent with the semantics of the electroencephalogram (EEG) signal data.

[0012] Preferably, it further includes:

[0013] Use the Inception score and the Fréchet Inception distance to evaluate the quality and diversity of the fine-grained images.

[0014] Preferably, the high-level semantic features are obtained by using a pre-trained residual network as a teacher model, guiding the training of a multi-channel convolutional neural network through response distillation and feature distillation, and performing feature extraction according to the trained network; the low-level semantic features are obtained by constructing a symmetric image autoencoder and using a joint training strategy to train the multi-channel convolutional neural network to align the EEG wave features with the image features, and performing feature extraction according to the trained network.

[0015] Preferably, collecting the EEG signal data includes:

[0016] Use an EEG paradigm design platform to design a rapid serial visual presentation experimental paradigm for picture stimuli;

[0017] Conduct an experiment based on the rapid serial visual presentation experimental paradigm, and use a computer monitor to switch the paradigm images every 0.5 seconds to stimulate the event-related potential response. After the image interface disappears, the experiment ends, and the EEG signal data is obtained.

[0018] Preferably, performing band-pass filtering and time slice segmentation on the EEG signal data to obtain preprocessed signals includes:

[0019] Use a band-pass filtering algorithm to limit the frequency of the EEG signal data within the range of 1 - 50 Hz to remove noise and other interference signals, thereby retaining effective EEG wave data;

[0020] Slice the EEG wave data into a time slice every 0.5 seconds;

[0021] Obtain the frequency domain data of the EEG wave data in each time slice through Fourier transform;

[0022] Combine the EEG wave data with the frequency domain data to obtain the preprocessed signals.

[0023] Preferably, the multi-channel convolutional neural network includes five layers of neural networks connected in sequence; the input of the first layer of the neural network is the preprocessed signal; the first layer of the neural network is used to perform convolutional operations on the preprocessed signal in two dimensions of time domain and frequency domain, and normalize the output of the first convolutional operation and activate it with the activation function ReLU to obtain the first convolutional result; the input of the second layer of the neural network is the first convolutional result; the second layer of the neural network is used to perform convolutional operations on the first convolutional result in two dimensions of time domain and frequency domain, and normalize the output of the second convolutional operation and activate it with the activation function ReLU to obtain the second convolutional result; the input of the third layer of the neural network is the second convolutional result; the third layer of the neural network is used to perform convolutional operations on the second convolutional result in two dimensions of spatial domain and frequency domain, and normalize the output of the third convolutional operation and activate it with the activation function ReLU to obtain the third convolutional result; the fourth layer of the neural network is used to perform convolutional operations on the third convolutional result in two dimensions of spatial domain and frequency domain, and normalize the output of the fourth convolutional operation and activate it with the activation function ReLU to obtain the fourth convolutional result; the fifth layer of the neural network is used to perform convolutional operations on the fourth convolutional result in three dimensions of time domain, spatial domain and frequency domain, and normalize the output of the fifth convolutional operation and activate it with the activation function ReLU.

[0024] Preferably, the process of extracting the high-level semantic features includes:

[0025] Select the residual network of ResNet50 as the teacher model, and replace the last classifier of the residual network with a classification head adapted to the category;

[0026] Use the sample image data to pre-train the teacher model, and fix the network parameters of the teacher model when training the student model based on the multi-channel convolutional neural network;

[0027] Connect an adaptive average pooling layer at the last output of the multi-channel convolutional neural network to map the extracted brain wave features into the domain image feature space;

[0028] Take the classification probability distribution of the last of the teacher model as the soft label, and use the loss function of KL divergence to let the student model learn the classification knowledge of the teacher model; the formula of the loss function of KL divergence is: where P is the probability distribution output by the teacher model, Q is the probability distribution output by the student model, i is the category, and L KL is the loss function of the KL divergence;

[0029] Align the feature vectors of the last layer of the teacher model and the student model; the feature vectors output by the teacher model contain robust high-level semantic information;

[0030] Using the cosine similarity loss function, align the feature vectors output by the teacher model with the feature vectors output by the student model in the semantic space, so that the student model can learn the high-level semantic knowledge related to categories in the teacher model; the formula of the cosine similarity loss function is: where f teacher and f student are the feature vectors of the teacher model and the student model respectively; L CS is the cosine similarity loss function;

[0031] Using the cross-entropy loss function, enable the student model to classify through the class label as a hard label; the formula of the cross-entropy loss function is as follows: where y is the true class distribution, and L CE is the cross-entropy loss function;

[0032] Use the trained network to extract high-level semantic features.

[0033] Preferably, the extraction process of the low-level semantic features includes:

[0034] Take the multi-channel convolutional neural network as the backbone network, and connect an adaptive average pooling layer at the last output of the multi-channel convolutional neural network, so that the brain wave features are mapped into the image feature space;

[0035] Based on the joint training strategy, align the low-level semantic features output by the backbone network with the low-level image semantic features output by the encoder of the image autoencoder; the alignment loss function uses the cosine similarity loss function and the reconstruction auxiliary loss function; the formula of the reconstruction loss function is: where h and w are the height and width of the image, I O is the original image, and I R is the image reconstructed by the image autoencoder; L rec is the reconstruction loss function;

[0036] Use the trained network to extract low-level semantic features.

[0037] Preferably, use the multi-level semantic features to conditionally regulate the diffusion process of a preset diffusion model to generate fine-grained images consistent with the semantics of the EEG signal data, including:

[0038] Construct a diffusion model; the diffusion model includes a multi-level FiLM injection module and a U-Net network; the U-Net network includes 8 layers of networks; the first four layers of networks are used for downsampling, and the last four layers are used for upsampling. Each layer of the network includes two residual modules and an attention module; the residual module is composed of four layers of convolutional neural networks and is connected by residual connections to ensure the effectiveness of feature information transmission; the attention module is used to capture global dependency relationships; the multi-level FiLM injection module is used to convert the semantic conditional vector into two learnable parameters, and use the two learnable parameters to integrate the semantic conditions into the diffusion model. The formula is: F' = γ(c)·F + β(c); where F is the feature map output by each layer of the U-Net, c is the semantic conditional vector, γ and β are learnable parameters for converting semantic conditions respectively, and F′ is the output feature map calculated by FiLM.

[0039] Adopt a linear strategy for the standard deviation scheduling of the diffusion model, and perform noise addition processing on the original image by adding step Gaussian noise; the formula for the standard deviation scheduling is: β t = 1 - α t ; where β is the standard deviation, α t is a factor for controlling β, α t varies with time, T is the total number of steps, and t is the number of steps at a certain moment; the product formula of α 1 ...α t is:

[0040] Add t-step Gaussian noise to the original image x 0 to obtain the noise-added result at the t-th step; the formula for the noise addition process is: where x t is the noise-added result after t steps, and ε is random Gaussian noise;

[0041] During the training process, use the L 1 loss function to guide the denoising process of the diffusion model and combine multi-level semantic features as conditions for control; the formula for the L 1 loss function is: where, is the diffusion model function, c h and c l are the high-level semantic feature and the low-level semantic feature respectively;

[0042] Denoise the noise-added result through a Markov chain to generate a fine-grained image that is semantically consistent with the electroencephalogram signal data; the formula for the Markov chain is: where δ t is the standard deviation of the Gaussian noise ε, and t iterates from T to 0 to obtain the fine-grained image.

[0043] A fine-grained image generation system based on multi-level EEG semantic features, comprising:

[0044] A data acquisition unit for acquiring EEG signal data;

[0045] A data preprocessing unit for performing band-pass filtering and time-slice segmentation on the EEG signal data to obtain preprocessed signals;

[0046] A feature extraction unit for extracting multi-level semantic features from the preprocessed signals by using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features;

[0047] An image generation unit for conditionally regulating the diffusion process of a preset diffusion model by using the multi-level semantic features to generate fine-grained images with semantics consistent with the EEG signal data.

[0048] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0049] The present invention provides a fine-grained image generation method and system based on multi-level EEG semantic features. The method includes: acquiring EEG signal data; performing band-pass filtering and time-slice segmentation on the EEG signal data to obtain preprocessed signals; extracting multi-level semantic features from the preprocessed signals by using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features; conditionally regulating the diffusion process of a preset diffusion model by using the multi-level semantic features to generate fine-grained images with semantics consistent with the EEG signal data. The present invention solves the problem that most methods in the prior art only focus on high-level category-related semantics, resulting in low semantic consistency and fidelity of the generated images. The experimental results of the present invention on different data sets show that it has good performance and generalization ability. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;

[0052] Figure 2 It is a flowchart of the technical route provided by the embodiment of the present invention;

[0053] Figure 3It is the overall framework diagram of EEG2IM provided by the embodiment of the present invention;

[0054] Figure 4 It is the unified neural network structure diagram for EEG feature extraction provided by the embodiment of the present invention

[0055] Figure 5 It is the structure diagram for high-level semantic feature extraction provided by the embodiment of the present invention

[0056] Figure 6 It is the structure diagram for low-level semantic feature extraction provided by the embodiment of the present invention

[0057] Figure 7 It is the structure diagram of the diffusion model network provided by the embodiment of the present invention. Specific implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] The purpose of the present invention is to provide a fine-grained image generation method and system based on multi-level EEG semantic features to achieve fine control of the image generation process, so as to generate high-quality images with consistent semantics and rich details.

[0060] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0061] Figure 1 It is the method flow chart provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a fine-grained image generation method based on multi-level EEG semantic features, including:

[0062] Step 100: Collect EEG signal data;

[0063] Step 200: Perform band-pass filtering and time-slice segmentation on the EEG signal data to obtain preprocessed signals;

[0064] Step 300: Extract multi-level semantic features from the preprocessed signals by using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features;

[0065] Step 400: Use the multi-level semantic features to conditionally regulate the diffusion process of a preset diffusion model to generate fine-grained images consistent with the semantics of the EEG signal data.

[0066] Specifically, the technical route and overall model framework diagram of this embodiment are as shown in Figure 2 and Figure 3 shown below:

[0067] Step S1. Dataset acquisition:

[0068] First, use an electroencephalogram (EEG) paradigm design platform to design a rapid serial visual presentation (RSVP) experimental paradigm for picture stimulation, and construct an EEG-image dataset. The EEG signal acquisition device includes 32 electrodes, with a sampling frequency of 128 Hz, and the reference electrodes are AFz and FCz. To ensure signal quality, the impedance of all electrodes is maintained below 5000 Ω. The EEG data acquisition subjects are healthy participants with normal vision (or corrected to normal vision) and no history of neurological or mental diseases. The acquisition process is carried out in a quiet isolation room to minimize visual and auditory interference. The participants sit 60 cm away from a 24-inch monitor, with an observation angle of approximately 2.5°, wear the EEG acquisition device and get ready for the experiment. When the experiment starts, the computer monitor switches the paradigm images every 0.5 seconds to stimulate the event-related potential (ERP) response. After the image interface disappears, the experiment ends.

[0069] Step S2. Data preprocessing and time slice division:

[0070] Preprocess the collected raw EEG signals. First, use the band-pass filtering algorithm in the Python library to limit the frequency of the EEG signals to the range of 1 - 50 Hz to remove noise and other interference signals, thereby retaining the effective EEG data. Then, divide the preprocessed EEG data into a time slice every 0.5 seconds, and the dimension of each time slice is (64 * 32). The EEG data of each time slice is transformed into frequency domain data through Fourier transform, and the time domain data and frequency domain data will be combined as independent input units for the training of the subsequent multi-level semantic feature extraction model. The conversion formula from time domain to frequency domain is as follows:

[0071]

[0072] where F represents the frequency domain data and E represents the time domain data. where C represents the number of EEG channels and N represents the number of EEG time slice sampling points.

[0073] Step S3. Define a unified neural network structure for EEG feature extraction:

[0074] As shown in Figure 4As shown, this neural network structure is based on a multi-channel convolutional neural network (CNN) and is used to extract semantic features from EEG data in different dimensions. The network includes five convolutional layers, specifically as follows:

[0075] The input of the first-layer neural network is EEG data, which is composed of time-domain and frequency-domain data combined. Specifically, the time-domain and frequency-domain data with dimensions of (64*32) are each extended by one dimension to become (1*64*32), and then concatenated on the new dimension to obtain the final input EEG data with dimensions of (2*64*32). This input data not only contains time-domain and frequency-domain feature information but also takes into account the spatial correlation between channels. A one-dimensional convolutional neural network is used to perform convolutional operations on the time-domain and frequency-domain dimensions, and the output of the convolutional operation is normalized and activated by the ReLU activation function.

[0076] The input of the second-layer neural network is the output of the first-layer neural network. Then, a one-dimensional convolutional neural network is used to perform convolutional operations on the time-domain and frequency-domain dimensions, and the output of the convolutional operation is normalized and activated by the ReLU activation function.

[0077] The input of the third-layer neural network is the output of the second-layer neural network. A one-dimensional convolutional neural network is used to perform convolutional operations on the spatial-domain and frequency-domain dimensions, and the output of the convolutional operation is normalized and activated by the ReLU activation function.

[0078] The input of the fourth-layer neural network is the output of the third-layer neural network. Then, a one-dimensional convolutional neural network is used to perform convolutional operations on the spatial-domain and frequency-domain dimensions, and the output of the convolutional operation is normalized and activated by the ReLU activation function.

[0079] The input of the fifth-layer neural network is the output of the fourth-layer neural network. A two-dimensional convolutional neural network is used to perform an overall convolutional operation on the time-domain, spatial-domain, and frequency-domain dimensions, and the output of the convolutional operation is normalized and activated by the ReLU activation function.

[0080] In the neural network for EEG feature extraction, the formula for normalization is as follows:

[0081]

[0082] where x i is a sample data of the current batch, μ is the sample mean of the current batch, σ 2 is the sample variance of the current batch, is a sample data after normalization of the current batch, and λ is a constant set to 0.00005 to prevent division by zero.

[0083] The formula for the ReLU activation function is as follows:

[0084] ReLU(x) = max(0, x) (5)

[0085] Where x is the current input data, and the max function takes the larger value between 0 and x.

[0086] Step S4. Multi-level semantic feature extraction:

[0087] S4-1: High-level semantic feature extraction

[0088] As Figure 5 shown, in this embodiment, ResNet50 is selected as the teacher network model, and its final classifier is replaced with a classification head suitable for the category. Then, it is pre-trained using a large-scale image data set (such as the ImageNet data set), and its parameters are fixed when training the EEG high-level semantic encoder student model. Among them, the high-level semantic encoder is the EEG encoder introduced in step S3. In addition, in order to better adapt to the distillation training architecture, a layer of adaptive average pooling layer is connected at the end of the high-level semantic encoder to map the extracted EEG features to the domain image feature space.

[0089] The adopted distillation learning method combines two techniques based on response distillation and feature distillation to guide the training of the EEG high-level semantic encoder, and at the same time, the category label is used as the classification supervision label of the EEG high-level semantic encoder.

[0090] The response-based distillation technique uses the final classification probability distribution of the teacher model as the soft label, and uses the loss function of KL divergence to let the student model learn the classification knowledge of the teacher model. The formula of the KL divergence loss function is as follows:

[0091]

[0092] Where P is the probability distribution output by the teacher model, Q is the probability distribution output by the student model, and i is the category.

[0093] The feature-based distillation technique aligns the feature vectors of the last layer of the teacher model and the student model. The feature vector output by the teacher model contains robust high-level semantic information. Using the cosine similarity loss function, the feature vector of the teacher model is aligned with the feature vector of the student model in the semantic space, so that the student model can learn the high-level semantic knowledge related to the category in the teacher model. The formula of the cosine similarity loss function is as follows:

[0094]

[0095] Where f teacher and f student are the feature vectors of the teacher model and the student model respectively.

[0096] In addition, the student model is classified using category labels as hard labels, and its loss function formula is as follows:

[0097]

[0098] where y is the true category distribution, Q is the probability distribution output by the student model, and i is the category.

[0099] The trained EEG high-level semantic encoder will be used to extract high-level semantic features.

[0100] S4-2: Low-level semantic feature extraction

[0101] As shown in Figure 6 , the EEG low-level semantic encoder uses the encoder in step S3 as the backbone network, which is similar to the structure of the EEG high-level semantic encoder. At the end, an adaptive average pooling layer is connected to map the EEG features into the image feature space. The output low-level semantic features are aligned with the image low-level semantic features output by the encoder of the image autoencoder using a joint training strategy. The alignment loss function still uses the cosine similarity loss function. At the same time, to ensure the correctness of the image low-level semantic features extracted by the image autoencoder, an additional reconstruction auxiliary loss function is added for joint optimization. The formula of the reconstruction loss function is as follows:

[0102]

[0103] where h and w are the height and width of the image, I O is the original image, and I R is the image reconstructed by the autoencoder.

[0104] The image autoencoder is used to extract low-level features such as color and contour, and assist the EEG low-level semantic encoder in extracting low-level features. It consists of an encoder and a decoder.

[0105] The role of the encoder is to compress the original image into the latent space, retaining low-level semantic information such as color and contour. It is composed of a five-layer convolutional neural network. The size of the convolutional kernel in each layer is 3, the stride is 2, and the number of output channels is 32, 64, 128, 256, and 512 in sequence. After each convolutional operation, normalization is performed, and a non-linear transformation is carried out through the ReLU activation function. Since the stride is 2, after each convolutional layer, the width and height of the feature map are reduced to 1 / 2 of the original. Finally, the number of output channels is 512, and the width and height are reduced to 1 / 32 of the original, forming a compact latent space feature representation.

[0106] The role of the decoder is to restore the compressed latent space features to the original image. Its structure is symmetric to that of the encoder and consists of five layers of transposed convolutional neural networks. The convolutional kernel size of each layer is 4, the stride is 2, and the number of output channels is 256, 128, 64, 32, and 3 in sequence. Each transposed convolutional operation doubles the width and height of the feature map, and at the same time undergoes normalization processing and activation by the ReLU activation function. Finally, the decoder maps the latent space features back to the same resolution as the original image, while restoring low-level semantic information such as the color and contour of the image. Transposed convolution is not only used for upsampling, but also can enhance key features through learnable parameters, making the generated image closer to the original input.

[0107] Step S5. Fine-grained image generation:

[0108] S5-1: Diffusion model structure:

[0109] As Figure 7 shown, the Net network is the main body of the diffusion model, and also includes a multi-level FiLM injection module to meet the input of multi-level semantic features. The U-Net network structure is similar to that of the image autoencoder and is also a symmetric structure, mainly including 8 layers of networks. The first four layers of networks are used for downsampling, and the last four layers are used for upsampling. Each layer of the network includes three modules, two residual modules and one attention module. The residual module is composed of four layers of convolutional neural networks and is connected by a residual connection to ensure the effectiveness of feature information transmission. The formula for the residual connection is:

[0110]

[0111] where f is the residual module composed of four layers of convolutional neural networks, and x is the input of the residual module.

[0112] The attention module is used to capture global dependencies. First, the input x is linearly transformed to generate query (Query), key (Key), and value (Value) matrices. The formula is:

[0113] Q = xW Q , K = xW K , V = xW V (11)

[0114] where W Q , W K , W V are all learnable transformation matrices.

[0115] The dimension d of the key vector is used k to scale the dot product of self-attention to ensure numerical stability. The self-attention formula is:

[0116]

[0117] Among them, softmax is a normalization function, whose role is to convert similarity scores into a probability distribution, making the sum of weights at all positions equal to 1, so as to assign the influence degree of different positions on the current query position.

[0118] FiLM converts the semantic conditional vector into two learnable parameters, and these two parameters can be used to incorporate semantic conditions into the diffusion model. Its formula is:

[0119] F' = γ(c)·F + β(c) (13)

[0120] Among them, F is the feature map output by each layer of the U-Net, c is the semantic conditional vector, γ and β are two learned parameters for converting semantic conditions, and F′ is the output feature map calculated by FiLM.

[0121] S5-2: Image generation with multiple semantic conditions:

[0122] The standard variance scheduling of the diffusion model adopts a linear strategy, and the formula is as follows:

[0123] β t = 1 - α t (14)

[0124]

[0125] Among them, α is a factor used to control β, which changes with time. β is the standard variance, and T is the total number of steps, and t is the number of steps at a certain moment. The product formula of α 1 ...α t is as follows:

[0126]

[0127] The original image x 0 After adding Gaussian noise for t steps, the noisy result at the t-th step is obtained. The noise-adding process formula is as follows:

[0128]

[0129] Among them, x t is the noisy result after t steps, and ε is random Gaussian noise.

[0130] The diffusion model uses the L 1 loss function during training, and its formula is as follows:

[0131]

[0132] Among them, φ θ is the diffusion model function, t is the time step, c h and cl They are high-level semantic features and low-level semantic features.

[0133] During the training process, the multi-level semantic features are used as conditions to guide the denoising process of the diffusion model. At the same time, the FiLM technology is used for fine-grained control, enabling the diffusion model to flexibly capture rich information in the semantic conditions.

[0134] During the image generation process, denoising is performed on the noisy image. Through the control of multi-level semantic feature conditions, high-quality images that conform to the EEG semantics are generated. The denoising process uses a Markov chain, and the formula is as follows:

[0135]

[0136] where δ t is the standard deviation of the Gaussian noise ε, and t iterates from T to 0 to obtain the image.

[0137] Step S6. Testing and result evaluation:

[0138] In the testing stage, this embodiment uses two metrics, the Inception Score (IS) and the Frechet Inception Distance (FID), to evaluate the quality and diversity of the generated images. Specifically, a higher IS value indicates higher quality of the generated images, while a lower FID value means a smaller distribution difference between the generated images and the real images. To calculate the IS and FID metrics, this embodiment uses the InceptionScore and FrechetInceptionDistance methods provided in the Python library for evaluation.

[0139] A fine-grained image generation method (EEG2IM) based on multi-level semantic feature extraction and generation in two stages from EEG signals. The method framework consists of a multi-level semantic feature extraction module and a diffusion model generation module, aiming to generate high-quality images with high semantic consistency and rich details from EEG signals. Specifically, in the first stage of the present invention, multi-level semantic features in EEG signals are extracted through the idea of cross-modal knowledge transfer. For high-level semantic features, a knowledge distillation method combining response distillation and feature distillation techniques is used to train a high-level semantic feature extractor, and then abstract semantic information related to categories is extracted. For low-level semantic features, through a joint training strategy, EEG low-level features are aligned with the low-level latent space of images to capture detailed information such as colors and contours, ensuring the diversity and integrity of semantic features. In the second stage, a multi-level FiLM module is designed to inject the extracted multi-level semantic features of EEG into the generation process of the diffusion model. The FiLM module dynamically adjusts the feature maps at each stage of the diffusion model through scaling factors and offsets to achieve fine control of the image generation process, thereby generating high-quality images with consistent semantics and rich details.

[0140] Corresponding to the above method, this embodiment also provides a fine-grained image generation system based on multi-level EEG semantic features, including:

[0141] A data acquisition unit for collecting EEG signal data;

[0142] A data preprocessing unit for performing band-pass filtering and time slicing on the EEG signal data to obtain preprocessed signals;

[0143] A feature extraction unit for extracting multi-level semantic features from the preprocessed signals using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features;

[0144] An image generation unit for conditionally regulating the diffusion process of a preset diffusion model using the multi-level semantic features to generate fine-grained images with semantics consistent with the EEG signal data.

[0145] The beneficial effects of the present invention are as follows:

[0146] The present invention proposes a two-stage feature extraction and generation process to achieve fine-grained image generation based on electroencephalogram (EEG) signals. In the feature extraction stage, a cross-modal knowledge transfer method is adopted, and the knowledge distillation technology is used to extract high-level semantic features in EEG signals, such as categories; low-level semantic features, such as colors and shapes, are extracted through a joint training strategy. In the generation stage, through a multi-level FiLM modulation module, multi-level semantic features are used as conditional inputs to generate high-quality images that conform to the semantics of EEG signals. The EEG2IM method solves the problem that most existing methods only focus on high-level category-related semantics, resulting in low semantic consistency and fidelity of the generated images. The experimental results of the present invention on different datasets show that it has good performance and generalization ability.

[0147] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0148] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A fine-grained image generation method based on multi-level EEG semantic features, characterized in that: include: Collect EEG signal data; Performing bandpass filtering and time slice segmentation on the EEG signal data to obtain a preprocessed signal; A multi-channel convolutional neural network is used to extract multi-level semantic features from the preprocessed signal; the multi-level semantic features include high-level semantic features and low-level semantic features; The multi-level semantic features are used to conditionally regulate the diffusion process of the preset diffusion model to generate a fine-grained image that is semantically consistent with the EEG signal data.

2. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1 is characterized in that: Also includes: The Inception score and Fréchet Inception distance are used to evaluate the quality and diversity of the fine-grained images.

3. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1 is characterized in that: The high-level semantic features are obtained by using a pre-trained residual network as a teacher model, guiding the training of a multi-channel convolutional neural network through response distillation and feature distillation, and extracting features based on the trained network; the low-level semantic features are obtained by constructing a symmetrical image autoencoder, training the multi-channel convolutional neural network with a joint training strategy to align brain wave features with image features, and extracting features based on the trained network.

4. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1 is characterized in that: Collect EEG signal data, including: Use the EEG paradigm design platform to design a rapid serial visual presentation experimental paradigm of picture stimulation; The experiment was conducted based on the rapid serial visual presentation experimental paradigm, using a computer display to switch the paradigm image every 0.5 seconds to stimulate event-related potential responses. After the image interface disappeared, the experiment ended and the EEG signal data was obtained.

5. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1 is characterized in that: The EEG signal data is subjected to bandpass filtering and time slice segmentation to obtain a preprocessed signal, including: Using a bandpass filtering algorithm to limit the frequency of the EEG signal data to a range of 1-50 Hz to remove noise and other interfering signals, thereby retaining valid EEG data; Divide the EEG data into time slices of 0.5 seconds each; The EEG data of each time slice is transformed into frequency domain data through Fourier transformation; The brain wave data is combined with the frequency domain data to obtain the preprocessed signal.

6. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1 is characterized in that: The multi-channel convolutional neural network includes five layers of neural networks connected in sequence; the input of the first layer of the neural network is the preprocessed signal; the first layer of the neural network is used to perform convolution operations on the preprocessed signal in two dimensions of time domain and frequency domain, and the output of the first convolution operation is normalized and activated by the activation function ReLU to obtain the first convolution result; the input of the second layer of the neural network is the first convolution result; the second layer of the neural network is used to perform convolution operations on the first convolution result in two dimensions of time domain and frequency domain, and the output of the second convolution operation is normalized and activated by the activation function ReLU to obtain the second convolution result; the input of the third layer of the neural network is the The second convolution result; the third layer of the neural network is used to perform convolution operations on the second convolution result in two dimensions, spatial and frequency domains, and normalize the output of the third convolution operation and activate the activation function ReLU to obtain the third convolution result; the fourth layer of the neural network is used to perform convolution operations on the third convolution result in two dimensions, spatial and frequency domains, and normalize the output of the fourth convolution operation and activate the activation function ReLU to obtain the fourth convolution result; the fifth layer of the neural network is used to perform convolution operations on the fourth convolution result in three dimensions, namely, time, spatial and frequency domains, and normalize the output of the fifth convolution operation and activate the activation function ReLU.

7. The fine-grained image generation method based on multi-level EEG semantic features according to claim 3 is characterized in that: The extraction process of the high-level semantic features includes: The residual network of ResNet50 is selected as the teacher model, and the last classifier of the residual network is replaced with a classification head of the adapted category; Pre-training the teacher model using sample image data, and fixing network parameters of the teacher model when training a student model based on a multi-channel convolutional neural network; Connecting an adaptive average pooling layer to the last output of the multi-channel convolutional neural network to map the extracted brain wave features into the domain image feature space; The final classification probability distribution of the teacher model is used as a soft label, and the loss function of KL divergence is used to allow the student model to learn the classification knowledge of the teacher model; the formula of the loss function of KL divergence is: Where P is the probability distribution of the teacher model output, Q is the probability distribution of the student model output, i is the category, L KL is the loss function of the KL divergence; Aligning the feature vectors of the last layer of the teacher model and the student model; the feature vector output by the teacher model contains robust high-level semantic information; The cosine similarity loss function is used to align the feature vector output by the teacher model with the feature vector output by the student model in the semantic space, so that the student model can learn the high-level semantic knowledge related to the category in the teacher model. The formula of the cosine similarity loss function is: Among them, f teacher and f student are the feature vectors of the teacher model and the student model respectively; L CS is the cosine similarity loss function; The cross entropy loss function is used to make the student model classify using the category label as a hard label; the formula of the cross entropy loss function is as follows: Among them, y is the true category distribution, L CE is the cross entropy loss function; The trained network is used to extract high-level semantic features.

8. The fine-grained image generation method based on multi-level EEG semantic features according to claim 3 is characterized in that: The extraction process of the low-level semantic features includes: The multi-channel convolutional neural network is used as a backbone network, and an adaptive average pooling layer is connected to the last output of the multi-channel convolutional neural network to map the brain wave features into the image feature space; Based on the joint training strategy, the low-level semantic features output by the backbone network are aligned with the low-level semantic features of the image output by the encoder of the image autoencoder; the alignment loss function adopts the cosine similarity loss function and the reconstruction auxiliary loss function; the formula of the reconstruction loss function is: Where h and w are the height and width of the image, I O is the original image, I R is the image reconstructed by the image autoencoder; L rec is the reconstruction loss function; The trained network is used to extract low-level semantic features.

9. The fine-grained image generation method based on multi-level EEG semantic features according to claim 1, characterized in that: The multi-level semantic features are used to conditionally regulate the diffusion process of the preset diffusion model to generate a fine-grained image that is semantically consistent with the EEG signal data, including: Construct a diffusion model; the diffusion model includes a multi-level FiLM injection module and a U-Net network; the U-Net network includes an 8-layer network; the first four layers of the network are used for downsampling, and the last four layers are used for upsampling, and each layer of the network includes two residual modules and an attention module; the residual module is composed of a four-layer convolutional neural network and is connected through residuals to ensure the effectiveness of feature information transmission; the attention module is used to capture global dependencies; the multi-level FiLM injection module is used to convert the semantic condition vector into two learnable parameters, and the semantic conditions are integrated into the diffusion model using the two learnable parameters, and the formula is: F'=γ(c)·F+β(c); wherein F is the feature map output by each layer of U-Net, c is the semantic condition vector, γ and β are respectively learnable parameters for converting semantic conditions, and F′ is the output feature map calculated by FiLM; A linear strategy is used to perform standard deviation scheduling of the diffusion model, and the original image is denoised by adding step Gaussian noise; the formula for standard variance scheduling is: β t =1-α t ; Where β is the standard deviation, α is t is a factor used to control β, α t Changes over time, T is the total number of steps, t is the moment of a certain step; α1...α t The product formula is: Add t steps of Gaussian noise to the original image x0 to get the t-th step of noise addition result; the formula of the noise addition process is: Among them, x t is the noise result after t steps, ε is random Gaussian noise; During the training process, the L1 loss function is used to guide the denoising process of the diffusion model, and multi-level semantic features are combined as conditions for control; the formula of the L1 loss function is: in, is the diffusion model function, c h and c l They are high-level semantic features and low-level semantic features respectively; The noise-added result is denoised by a Markov chain to generate a fine-grained image that is semantically consistent with the EEG signal data. The formula of the Markov chain is: Among them, δ t is the standard deviation of Gaussian noise ε, and t is iterated from T to 0 to obtain the fine-grained image.

10. A fine-grained image generation system based on multi-level EEG semantic features, characterized in that: include: A data acquisition unit, used for acquiring EEG signal data; A data preprocessing unit, used for performing bandpass filtering and time slice segmentation on the EEG signal data to obtain a preprocessed signal; A feature extraction unit, used for extracting multi-level semantic features from the preprocessed signal using a multi-channel convolutional neural network; the multi-level semantic features include high-level semantic features and low-level semantic features; The image generation unit is used to conditionally regulate the diffusion process of the preset diffusion model using the multi-level semantic features to generate a fine-grained image that is semantically consistent with the EEG signal data.