High-quality video reconstruction method based on electroencephalogram signals

Through the multimodal fusion video reconstruction method based on EEG data, the problems of low time resolution and complex operation in fMRI technology are solved, high-quality video reconstruction is realized, and the authenticity and application value of the video are improved.

CN119987549AActive Publication Date: 2025-05-13南昌理工学院 +1

Patent Information

Application Number
CN202510070661.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In the prior art, when using fMRI technology to reconstruct videos, due to low time resolution, high equipment cost, complex operation and professional skills, the video dynamic information extraction is incomplete, low time consistency, inter-frame jumps and incoherence are caused.

Method used

Using a multimodal fusion video reconstruction method based on EEG data, the time-domain-space-frequency domain features are extracted, and the contour and depth information are combined, and high-quality video is generated using adaptive fusion and contrast learning methods.

Benefits of technology

It realizes high-quality video reconstruction, improves the authenticity and application value of the video, optimizes the extraction of brain signals, effectively captures the time and space information of the video, and generates videos with richer information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987549A_ABST
    Figure CN119987549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and particularly discloses a high-quality video reconstruction method based on electroencephalogram signals, which comprises the following steps: preprocessing electroencephalogram data, down-sampling video frames, and segmenting the electroencephalogram data by using a sliding window technology; selecting a channel with high correlation with video reconstruction; electroencephalogram time features are extracted; electroencephalogram space-time features are extracted; electroencephalogram time-frequency features are extracted; the electroencephalogram time-space-frequency features are fused; extracting the category and semantic information of the electroencephalogram; adopting a comparative learning method to align electroencephalogram-text-picture features; generating a reference video frame by using a stable diffusion model; and generating a high-quality video by using the video diffusion model. According to the method, the information of electroencephalogram features is further deeply extracted, the semantic content of the video is remarkably enhanced, meanwhile, the introduced video diffusion model gradually constructs a clear and natural picture through a step-by-step denoising process, the consistency between video frames is remarkably improved, and therefore the high-quality video rich in details is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a high-quality video reconstruction method based on electroencephalogram signals. Background Art

[0002] The human visual system is an advanced intelligent information processing mechanism. The brain processes a large amount of visual input every day and perceives and understands this information through complex neural processes. The development of deep learning technology has enabled scientists to extract visual information from human brain activity patterns and reconstruct visual stimuli such as videos. This technology has great potential and has an important impact on the development of brain-computer interfaces (BCI). The core goal of BCI technology is to establish a direct bridge between the human brain and external devices, so that people can control computers, robotic arms and other devices only through brain activity or slight muscle movements. This is particularly important for people with disabilities to restore their communication and mobility, and is also of great significance for enhancing the interaction efficiency and experience of normal people.

[0003] In recent years, most studies on reconstructing videos using brain activity are based on functional magnetic resonance imaging (fMRI) technology. However, fMRI technology measures changes in blood dynamics, which is a relatively slow process. As a result, it cannot capture rapid changes in brain activity and has a low temporal resolution. It has a natural disadvantage for tasks that require high temporal resolution, such as reconstructing videos. Secondly, fMRI data acquisition faces high equipment costs, complex operating procedures, and the need for highly specialized skills. These factors make the effective acquisition and analysis of fMRI data a complex and expensive task, which often requires professional technicians to have in-depth domain knowledge and rich practical experience, which is not conducive to the promotion of this research. In addition, traditional research methods often directly extract video information from brain activity, which leads to insufficient semantic information captured, low temporal consistency of the generated video, and a lack of smoothness and artifacts.

[0004] In order to solve these problems, the present invention uses a multimodal fusion video reconstruction method to reconstruct videos based on EEG data. The present invention preprocesses the EEG data; then, the time-domain-spatial-frequency domain features of the EEG are integrated to comprehensively describe the characteristics of brain activity and provide richer information for subsequent video reconstruction; finally, the contour and depth information are used to enhance the temporal consistency of the video and generate high-quality videos. This method optimizes the extraction of brain signals, effectively captures the spatiotemporal information of the video, and generates a video with richer information. Summary of the invention

[0005] In order to solve the problems existing in the prior art, the present invention provides a high-quality video reconstruction method based on EEG signals, which optimizes the extraction of brain signals, effectively captures the spatiotemporal information of the video, generates a video with richer information, can achieve high-quality video reconstruction, and solves the problems mentioned in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a high-quality video reconstruction method based on EEG signals, comprising the following steps:

[0007] S1. Preprocess the EEG data, downsample the video frames, and segment the EEG data using the sliding window technique;

[0008] S2. Select channels with high correlation with video reconstruction: Use the common spatial pattern (CSP) technology to spatially filter the EEG data to extract signal components related to specific tasks or stimuli, and combine it with the attention mechanism to select channels with high correlation with video reconstruction;

[0009] S3, using the masked coding method and the masked coding MAE model to extract EEG temporal features;

[0010] S4. Use the spatiotemporal attention mechanism to capture the correlation and importance between different time points and spatial positions and extract EEG spatiotemporal features;

[0011] S5. Using wavelet transform method, extract the frequency domain information of EEG data, and using convolutional neural network to extract the time domain features of EEG data to generate EEG time-frequency features;

[0012] S6, using an adaptive fusion method to fuse EEG temporal features, spatiotemporal features, and time-frequency features to generate EEG time-space-frequency features;

[0013] S7, using EEG classifier to extract video category information, using BLIP model to extract video semantic information, using semantic fusion network SF-Net to fuse these two kinds of information and generate text features;

[0014] S8, using contrastive learning method to align EEG-text-image features to obtain optimized EEG features;

[0015] S9. Using the generated optimized EEG features as a condition, a reference video frame is generated using a stable diffusion model; using the generated reference video frame as an auxiliary condition, a high-quality video is generated using a video diffusion model.

[0016] Preferably, in step S1, it specifically includes: preprocessing the EEG data, the preprocessing steps include locating channel data, deleting useless electrodes, re-referencing, bandpass filtering, notch filtering and independent principal component analysis, so as to ensure the quality of EEG data; then downsampling the video to 3fps; then using a sliding window of fixed length to segment the EEG data, and each window is processed as a time frame.

[0017] Preferably, in step S2, the common spatial pattern CSP technique is used to perform spatial filtering on the EEG signal, and the attention mechanism is applied to the relevant channels selected by the CSP method, and the features of these channels are weighted by the learned weight vector to extract the signal components related to the specific task or stimulus. The objective function of the CSP spatial filter ω is as follows:

[0018]

[0019] in, represents the EEG data of the ith channel, TS is all time points, AC is the number of all channels; C i is the spatial covariance matrix of the EEG data of the ith channel; ω represents the spatial filter, and T represents the transpose of the matrix.

[0020] Preferably, in step S3, a masked coding method is used to extract the time domain features of the EEG data, the EEG data of the channel selected in step S2 is input into the masked coding MAE model, the EEG data is masked by 75%, and then information is extracted from the unmasked EEG data to reconstruct the EEG data, and finally the predicted EEG data is compared with the original EEG data. The specific calculation formula of the loss function is as follows:

[0021]

[0022] Where N represents the number of samples, y i ' represents the predicted value generated by the MAE model, y i Represents a primitive value.

[0023] Preferably, in step S4, the EEG data of the channel selected in step S2 is input into a spatiotemporal feature extraction network;

[0024] First, the EEG data passes through the temporal convolution layer to capture the dynamic changes of the EEG data in the temporal dimension, thereby extracting the temporal features; then, the deep Transformer model is used to model the EEG data in the spatial dimension, thereby extracting the spatial information of the EEG data, including the associations and interactions between different brain regions; finally, the extracted features are input into the spatial convolution model to further fuse the temporal and spatial information to obtain the spatiotemporal characteristics of the EEG data.

[0025] Preferably, in step S5, the frequency domain features of the EEG data are extracted from the EEG data corresponding to the selected channels with high correlation with the video reconstruction task based on the wavelet transform method, and the time domain features of the EEG data are extracted from the EEG data corresponding to the selected channels with high correlation with the video reconstruction task based on the convolutional neural network, and then the extracted frequency domain features of the EEG data and the time domain features of the EEG data are feature fused to obtain the EEG time-frequency features; the specific formula of the wavelet transform is as follows:

[0026]

[0027] in, represents the data of all channels on the EEG at time t, a represents the scale factor, b represents the translation factor, γ * (t) represents the complex conjugate of the wavelet function, and dt represents the integral with respect to t.

[0028] Preferably, in step S6, it specifically includes: first, using an adaptive feature fusion method to fuse the EEG spatiotemporal features and the EEG time-frequency features to obtain EEG spatiotemporal-time-frequency fusion features; then, using a convolutional neural network automatic learning method to obtain a weight coefficient matrix; then, based on the obtained weight coefficient matrix, the EEG time features and the EEG spatiotemporal-time-frequency fusion features are pooled and fully connected to obtain EEG time-space-frequency features; the specific formula is as follows:

[0029] F=[(m1,m2)·V',m3]

[0030] Among them, F represents the fusion feature of EEG time feature, EEG spatiotemporal feature and EEG time-frequency feature, m1 represents EEG time-frequency feature, m2 represents EEG spatiotemporal feature, m3 represents EEG time feature, and V' represents the weight coefficient matrix.

[0031] Preferably, in step S7, it specifically includes: using an EEG classifier to obtain the category of the EEG corresponding to the video frame; using a guided language-image pre-training BLIP model to extract the semantic information of the video frame in the preprocessed original video; using a semantic fusion network SF-Net to fuse the video frame category information and the video frame semantic information to generate text features.

[0032] Preferably, in step S8, the downsampled video frame obtained in step S1 is input into the CLIP picture encoder to extract picture features; then the text features generated in step S7 are input into the CLIP text encoder to obtain text features; together with the EEG time-space-frequency features generated in step S6, they are input into the CLIP model as features for comparative learning, and finally, the EEG-picture and EEG-text alignment is performed by cosine similarity calculation, and the specific formula is as follows:

[0033]

[0034]

[0035] S clip =S ei +S et

[0036] Among them, l eeg is the EEG feature that has been projected to make it more suitable for video generation, l image , l text are picture features and text features respectively; S ei represents the cosine similarity of EEG-image, S et represents the cosine similarity between EEG and text, S clip It is the overall contrastive learning loss of the time-space-frequency features, image features and text features in CLIP. Finally, the EEG time-space-frequency features are further optimized to obtain the optimized EEG features.

[0037] Preferably, in step S9, it specifically includes: using the generated optimized EEG features as conditions, using the stable diffusion model to generate a reference video frame; using the Canny edge detection algorithm to extract the contour information of the reference video frame; extracting the depth information of the reference video frame based on the Monodepth depth estimation algorithm; using the preset spatiotemporal conditional fusion encoder STCF-net to integrate the contour information and depth information into the conditions to obtain the fusion conditions; generating a high-quality reconstructed video based on the video diffusion model and the fusion conditions.

[0038] The beneficial effects of the present invention are as follows: the multimodal fusion video reconstruction method based on brain activity of the present invention makes up for the problems of incomplete extraction of video dynamic information, low consistency between video frames, frame jumps and incoherence by adopting time window segmentation method, attention mechanism, masking coding method, spatiotemporal attention mechanism method, wavelet transform method, contrastive learning and other methods, and realizes high-quality video reconstruction, thereby greatly improving the authenticity of the video, so as to make the application value higher. The present invention can extract visual information from human brain activity patterns, help develop brain-computer interfaces, provide new ways for disabled people to restore communication and control equipment, and enhance the interaction efficiency and experience of normal people. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A flowchart of video reconstruction using EEG data according to an embodiment of the present invention;

[0040] Figure 2 A sample result diagram of an embodiment of the present invention;

[0041] Figure 3Schematic diagram of comparison between sample results of an embodiment of the present invention and a traditional method. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] This implementation uses the EEG data collected when watching video stimulation as an example. The EEG data is self-collected data. All subjects collected EEG data while watching the video using the Bo Rui Kang 64-wet electrode EEG acquisition device. During the acquisition process, the subjects maintained a comfortable and relaxed state, sitting on a chair about 1 meter away from the screen. The subjects were required to remain still while watching the video and reduce actions such as blinking and swallowing saliva.

[0044] See also Figure 1 The present invention provides a technical solution: a high-quality video reconstruction method based on EEG signals, comprising the following steps:

[0045] Step S1: Preprocess the EEG data by first downsampling the original video to 3fps, and then segmenting the EEG data using a sliding window of fixed length, with each window being processed as a time frame.

[0046] Step S2: Use CSP technology to spatially filter the EEG signal to extract signal components related to specific tasks or stimuli, and combine it with the attention mechanism to select channels with high correlation with video reconstruction.

[0047] Step S3: Adopting the masked coding method and using the MAE model to extract EEG temporal features.

[0048] Step S4: Use the spatiotemporal attention mechanism method to capture the correlation and importance between different time points and spatial positions and extract EEG spatiotemporal features.

[0049] Step S5: Using the wavelet transform method to extract the frequency domain information of the EEG data, and using the convolutional neural network to extract the time characteristics of the EEG data to generate EEG time-frequency characteristics.

[0050] Step S6: Using an adaptive fusion method, the EEG time features, space-time features, and time-frequency features are fused to generate EEG time-space-frequency features.

[0051] Step S7: Use the EEG classifier to extract video category information, use the BLIP model to extract video semantic information, and use the ensemble learning method to fuse these two types of information to generate text features.

[0052] Step S8: Using contrastive learning method, the CLIP model is used to align the EEG features of step S6 with the image features, and the EEG features of step S6 with the text features of step S7 to obtain optimized EEG features.

[0053] Step S9: Using the optimized EEG features generated in step S8 as conditions, using the stable diffusion model, generating a reference video frame; using the generated reference video frame as an auxiliary condition, combining contour and depth information, using the video diffusion model to generate / reconstruct a high-quality video.

[0054] In step S1, the EEG data is preprocessed, and the main steps of the preprocessing include positioning channel data, deleting useless electrodes, re-referencing, bandpass filtering, notch filtering, independent principal component analysis, etc. Eye movement artifacts, electrocardiogram artifacts, electromyography artifacts, etc. in the acquisition process are removed to the greatest extent, ensuring the quality of EEG data.

[0055] For video downsampling, ffmpeg software was used to downsample the video from 30fps to 3fps.

[0056] For the division of the time window, the present invention adopts a time window division method with a length of 500 data points to correspond the EEG data to the video frames one by one.

[0057] In step S2, in order to effectively extract the key information most relevant to video reconstruction and improve the accuracy and quality of the reconstruction results, the present invention uses CSP technology to perform spatial filtering on the EEG signal, and at the same time applies the attention mechanism to the relevant channels selected by the CSP method, and weights the features of these channels through the learned weight vector to extract the signal components related to the specific task or stimulus. The objective function of the CSP spatial filter ω is as follows:

[0058]

[0059] in, represents the EEG data of the i-th channel, TS is all time points, AC is the number of all channels; C i is the spatial covariance matrix of the EEG data of the i-th channel; ω represents the spatial filter, and T represents the transpose of the matrix.

[0060] During the training process, the attention mechanism helps the model selectively focus on certain parts of the input signal, thereby extracting information related to the video reconstruction task in a targeted manner. The attention mechanism can dynamically adjust the weights of channel features, so that the model can effectively extract relevant information in different situations, optimize tasks, and improve performance.

[0061] In step S3, the masked coding method is used to extract the time domain features of the EEG data. The EEG data of the channel selected in step S2 is input into the MAE model, the EEG data is masked by 75%, and then information is extracted from the unmasked data to reconstruct the EEG data. Finally, the predicted data is compared with the original data. The specific calculation formula of the loss function is as follows:

[0062]

[0063] Where N represents the number of samples, y i ' represents the predicted value generated by the MAE model, y i Represents a primitive value.

[0064] In the step S4, the spatiotemporal variation pattern of the EEG signal contains rich cognitive and behavioral information, and extracting these spatiotemporal features can improve the decoding ability of the brain to process and respond to video content. Therefore, in order to accurately reconstruct the details and dynamics in the video, the present invention introduces the spatiotemporal Transformer method to extract spatiotemporal features. Specifically, the EEG data of the channel selected in step S2 is first input into the spatiotemporal feature extraction network. Then, the data passes through the temporal convolution layer to capture the dynamic changes of the EEG data in the time dimension, thereby extracting its temporal features; then, through the deep Transformer model, the EEG data is modeled in the spatial dimension, thereby extracting the spatial information of the data, including the association and interaction between different brain regions; finally, the extracted features are input into the spatial convolution model to further fuse the temporal and spatial information to obtain the spatiotemporal features of the data, and provide a richer and more comprehensive feature representation for subsequent tasks.

[0065] In the step S5, because the activity of EEG signals at different frequencies reflects the state and process of the brain in different cognitive tasks. Extracting time-frequency features can reveal these spectral changes, helping to better understand how the brain processes different levels of video content, thereby improving the accuracy and quality of video reconstruction. Therefore, the present invention introduces a wavelet transform and a convolutional neural network fusion method to extract time-frequency features. Specifically, the wavelet transform method processes the EEG data of the channel selected in step S2. The wavelet transform first converts the EEG data into the wavelet domain, in which the multi-scale decomposition of the EEG data is achieved by analyzing the wavelet basis functions of different scales and frequencies, thereby capturing the characteristics of the EEG data in the frequency domain. Next, the important frequency domain features of the EEG data are extracted using the efficient representation and compression properties of the wavelet transform. The specific formula of the wavelet transform is as follows:

[0066]

[0067] in, represents the data of all channels on EEG at time t, a represents the scale factor, b represents the translation factor, γ * (t) represents the complex conjugate of the wavelet function, dt represents the integral of t, and the wavelet function used in the present invention is the Morlet wavelet. Morlet wavelet is a commonly used complex wavelet function, which has the advantage of considering both time domain and frequency domain characteristics. By applying the Morlet wavelet transform, the EEG data is converted into the wavelet domain, thereby realizing the time-frequency feature analysis of the EEG data. In the wavelet domain, the Morlet wavelet can provide local feature representations of the EEG data at different scales and frequencies to capture the time-frequency information of the EEG data.

[0068] After extracting frequency domain features using wavelet transform, the EEG data selected in step S2 is input into the convolutional neural network to extract time domain features, and finally the two are fused to obtain time-frequency features.

[0069] In the step S6, the time, space and frequency characteristics of the EEG signal provide different levels of understanding of the signal. Traditional feature fusion methods have some specific problems in feature selection and fusion effect. Redundant features may appear in the feature selection process, that is, different features are highly correlated, which increases the complexity of the model but does not necessarily improve the performance, and important features may be omitted. In terms of feature fusion, how to assign appropriate weights to different features is a challenge, and improper weight assignment may lead to insufficient feature differentiation after fusion. Therefore, the present invention introduces an adaptive feature fusion method to dynamically integrate these multi-dimensional features, extract information in an optimal way, and ensure that valuable information obtained from different features is effectively utilized, thereby improving the accuracy and reliability of the final result. Therefore, an adaptive feature fusion method is used to fuse the time domain features obtained in step S3, the spatiotemporal features generated in step S4, and the time-frequency features of step S5 to obtain EEG time-space-frequency fusion features. First, the spatiotemporal features are fused with the time-frequency features using an adaptive fusion method. Next, a convolutional neural network is used to automatically learn and obtain a weight coefficient matrix. Then, the two features are pooled and fully connected. Since the masked coding model masks and reconstructs the input data, it can capture subtle changes and important patterns in the signal. This high-fidelity feature representation can more comprehensively reflect the EEG features. Therefore, the present invention assigns a higher weight to it and directly merges the obtained features with the time domain features of step S3. The specific formula is as follows:

[0070] F=[(m1,m2)·V',m3]

[0071] Among them, F represents the fusion feature of the three, m1 represents the time-frequency feature obtained in step S5, m2 represents the space-time feature generated in step S4, m3 represents the time domain feature in step S3, and V' represents the weight coefficient matrix.

[0072] In step S7, the EEG classifier is first used to obtain the EEG category corresponding to the video frame. Specifically, EEG data and recurrent neural networks are combined to learn the brain's classification activities of visual objects. Next, the present invention introduces an attention mechanism to increase the model's attention to important features. On this basis, a regressor is trained through a convolutional neural network to transfer the learned classification capabilities to the machine, enabling it to automatically classify images, and use the human brain features integrated with the attention mechanism for visual classification to obtain the category corresponding to the EEG data, such as "cat".

[0073] The BLIP model is used to obtain video semantic information. Specifically, the video frame downsampled in step S1 is input into the BLIP model, and the subtitle generator of the BLIP model is used to generate synthetic subtitles to convert visual information into text descriptions. The generated subtitles are filtered using filters to remove possible noise or irrelevant information, extract key information, and finally obtain subtitles corresponding to the video frame, such as "A man in a white tracksuit runs again".

[0074] When reconstructing a video, category information can provide a basic classification of the video content (such as scenes, objects or activities), while semantic information can describe the specific meaning and context of these categories. By fusing these two types of information, the relationship between EEG signals and video content can be more fully understood, ensuring that the reconstruction process not only takes into account the simple features of the signal, but also captures the deep semantics behind it. Therefore, the present invention proposes a semantic fusion network named SF-Net to fuse category information and semantic information. First, the category information and semantic information are used as different input data, and are effectively encoded and represented using unique hot encoding. Then, the convolutional neural network of the SF-Net category branch is used to extract category features, and the recurrent neural network branch is used to extract semantic features. Finally, at the top of the multimodal fusion network, a fusion layer is introduced to fuse features from different branches to integrate category information and semantic information to generate comprehensive text features.

[0075] In step S8, the downsampled video frame obtained in step S1 is input into the CLIP picture encoder to extract picture features, and then the semantic features generated in step S7 are input into the CLIP text encoder to obtain text features, which are input into the CLIP model together with the time-space-frequency features generated in step S6 as features for comparative learning. Finally, the EEG-picture and EEG-text alignment is performed by cosine similarity calculation. The specific formula is as follows:

[0076]

[0077]

[0078] S clip =S ei +S et

[0079] Among them, l eeg is the EEG feature that has been projected to make it more suitable for video generation, l image , l text are the features extracted from pictures and texts respectively. ei represents the cosine similarity of EEG-image, S etrepresents the cosine similarity between EEG and text, S clip It is the overall contrastive learning loss of the three CLIPs, which ultimately further optimizes the EEG features.

[0080] In step S9, a pre-trained stable diffusion model is used to generate video frames, and stable diffusion gradually denoises the learning data distribution. First, the stable diffusion model is fine-tuned by using EEG and video frame data sets, and it is fine-tuned from a text generation model that generates ordinary images to a model that generates images conditionally based on the electroencephalogram. Then, the EEG features optimized in step S8 are input into the stable diffusion model as conditions. The conditional signal is introduced through the cross-attention mechanism in the UNet network, and the cross-attention can merge the conditional information of the EEG data. The specific formula is as follows:

[0081]

[0082] Among them, Q stands for query, and the query vector refers to the reference and dependence of the model on the previous frame and global information when generating the current frame. It is mainly used to obtain the correlation with other vectors. K stands for key, and the key vector refers to the characteristics of the conditional information embedded in the EEG data, which is used to match with Q to calculate the attention weight. V stands for value, and the value vector is the part used to generate the final output after weighting according to the attention weight. It represents the EEG features weighted by the cross-attention mechanism, and this information will be used to guide the image content in the generation process. d represents the column vector dimension of Q and K, and T represents the transpose of the matrix.

[0083] In step S9, the video diffusion model is used to generate high-quality video using the generated video frame as an auxiliary condition. First, the contour of the video frame is extracted using the Canny edge detection algorithm, and then the depth information of the video frame is extracted using the Monodepth depth estimation algorithm. Then, the present invention proposes a spatiotemporal conditional fusion encoder named STCF-net, which integrates the contour and depth information into the conditions, further captures the spatiotemporal information of the video, and improves the video quality.

[0084] For STCF-net, specifically, in the spatial information extraction stage, two two-dimensional convolutional layers and a maximum pooling layer are first used to capture the input data features, and then contour information is introduced as an auxiliary input, and channels are added to the convolutional layer to process the contour information, so as to better express the structure and content of the picture; at the same time, the depth information is combined with the spatial information, and a one-dimensional channel can be added to process the depth information, so that the model can better understand the three-dimensional sense and depth information of the object; in the temporal feature extraction stage, compared with a single contour map, the contour sequence can provide more temporal information, so the temporal transformers module is continued to be used to extract the contour sequence and depth information, so that the model can fully utilize the advantages of multi-source information for video generation, and finally generate high-quality video through the fusion of conditions.

[0085] The present invention first proposes a time-space-frequency EEG feature fusion method to more comprehensively capture brain activity information, so as to more deeply understand and analyze the brain's activity patterns. Then, the present invention proposes SF-Net to fuse category and video semantic information to more comprehensively fuse coarse-grained and fine-grained semantic information. Finally, in the generation stage, the present invention first uses a stable diffusion model to generate reference video frames. At the same time, in order to enhance the consistency between frames, the proposed STCF-NET is used to fuse contour and depth information to generate high-quality video through a video diffusion model. Figure 2 From the above, the present invention reconstructs the video with accurate semantics, motion and scene dynamics. Figure 3 From the above, the method of the present invention is closer to the real video in terms of the action, number of subjects and color of the reconstructed video than the traditional method, and the reconstructed video is more realistic. The present invention has made certain contributions to the development of brain-computer interface and computer vision.

[0086] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0087] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0088] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0089] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0090] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0091] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A high-quality video reconstruction method based on electroencephalogram signals, characterized in that: The steps include: S1. Preprocess the EEG data, downsample the video frames, and segment the EEG data using the sliding window technique; S2. Select channels with high correlation with video reconstruction: Use the common spatial pattern (CSP) technology to spatially filter the EEG data to extract signal components related to specific tasks or stimuli, and combine it with the attention mechanism to select channels with high correlation with video reconstruction; S3, using the masked coding method and the masked coding MAE model to extract EEG temporal features; S4. Use the spatiotemporal attention mechanism to capture the correlation and importance between different time points and spatial positions and extract EEG spatiotemporal features; S5. Using wavelet transform method, extract the frequency domain information of EEG data, and using convolutional neural network to extract the time domain features of EEG data to generate EEG time-frequency features; S6, using an adaptive fusion method to fuse EEG temporal features, spatiotemporal features, and time-frequency features to generate EEG time-space-frequency features; S7, using EEG classifier to extract video category information, using BLIP model to extract video semantic information, using semantic fusion network SF-Net to fuse these two kinds of information and generate text features; S8, using contrastive learning method to align EEG-text-image features to obtain optimized EEG features; S9, generating a reference video frame using a stable diffusion model based on the generated optimized EEG features; The generated reference video frames are used as auxiliary conditions to generate high-quality videos using the video diffusion model.

2. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S1, specifically, it includes: preprocessing the EEG data, the preprocessing steps include locating channel data, deleting useless electrodes, re-referencing, bandpass filtering, notch filtering and independent principal component analysis, so as to ensure the quality of EEG data; then downsampling the video to 3fps; then using a fixed-length sliding window to segment the EEG data, and each window is processed as a time frame.

3. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S2, the common spatial pattern (CSP) technique is used to spatially filter the EEG signal. At the same time, the attention mechanism is applied to the relevant channels selected by the CSP method. The features of these channels are weighted by the learned weight vector to extract the signal components related to the specific task or stimulus. The objective function of the CSP spatial filter ω is as follows: in, represents the EEG data of the ith channel, TS is all time points, AC is the number of all channels; C i is the spatial covariance matrix of the EEG data of the ith channel; ω represents the spatial filter, and T represents the transpose of the matrix.

4. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S3, the masked coding method is used to extract the time domain features of the EEG data. The EEG data of the channel selected in step S2 is input into the masked coding MAE model, the EEG data is masked by 75%, and then information is extracted from the unmasked EEG data to reconstruct the EEG data. Finally, the predicted EEG data is compared with the original EEG data. The specific calculation formula of the loss function is as follows: Where N represents the number of samples, y i ' represents the predicted value generated by the MAE model, y i Represents a primitive value.

5. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S4, the EEG data of the channel selected in step S2 is input into the spatiotemporal feature extraction network; First, the EEG data passes through the temporal convolution layer to capture the dynamic changes of the EEG data in the temporal dimension, thereby extracting the temporal features; then, the deep Transformer model is used to model the EEG data in the spatial dimension, thereby extracting the spatial information of the EEG data, including the associations and interactions between different brain regions; finally, the extracted features are input into the spatial convolution model to further fuse the temporal and spatial information to obtain the spatiotemporal characteristics of the EEG data.

6. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S5, the frequency domain features of the EEG data are extracted from the EEG data corresponding to the selected channels with high correlation with the video reconstruction task based on the wavelet transform method, and the time domain features of the EEG data are extracted from the EEG data corresponding to the selected channels with high correlation with the video reconstruction task based on the convolutional neural network, and then the extracted frequency domain features of the EEG data and the time domain features of the EEG data are feature fused to obtain the EEG time-frequency features; the specific formula of the wavelet transform is as follows: in, represents the data of all channels on the EEG at time t, a represents the scale factor, b represents the translation factor, γ * (t) represents the complex conjugate of the wavelet function, and dt represents the integral with respect to t.

7. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S6, it specifically includes: first, using an adaptive feature fusion method to fuse the EEG spatiotemporal features and the EEG time-frequency features to obtain EEG spatiotemporal-time-frequency fusion features; then, using a convolutional neural network automatic learning method to obtain a weight coefficient matrix; then, based on the obtained weight coefficient matrix, the EEG time features and the EEG spatiotemporal-time-frequency fusion features are pooled and fully connected to obtain EEG time-space-frequency features; the specific formula is as follows: F=[(m1,m2)·V',m3] Among them, F represents the fusion feature of EEG time feature, EEG spatiotemporal feature and EEG time-frequency feature, m1 represents EEG time-frequency feature, m2 represents EEG spatiotemporal feature, m3 represents EEG time feature, and V′ represents the weight coefficient matrix.

8. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S7, it specifically includes: using the EEG classifier to obtain the category of the EEG corresponding to the video frame; using the guided language-image pre-training BLIP model to extract the semantic information of the video frame in the preprocessed original video; using the semantic fusion network SF-Net to fuse the video frame category information and the video frame semantic information to generate text features.

9. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S8, the downsampled video frame obtained in step S1 is input into the CLIP picture encoder to extract picture features; Then, the text features generated in step S7 are input into the text encoder of CLIP to obtain text features; together with the EEG time-space-frequency features generated in step S6, they are input into the CLIP model as features for comparative learning. Finally, the EEG-image and EEG-text alignment are performed through cosine similarity calculation. The specific formula is as follows: S clip =S ei +S et Among them, l eeg is the EEG feature that has been projected to make it more suitable for video generation, l image , l text are picture features and text features respectively; S ei represents the cosine similarity of EEG-image, S et represents the cosine similarity between EEG and text, S clip It is the overall contrastive learning loss of the time-space-frequency features, image features and text features in CLIP. Finally, the EEG time-space-frequency features are further optimized to obtain the optimized EEG features.

10. The high-quality video reconstruction method based on EEG signals according to claim 1, characterized in that: In step S9, it specifically includes: using the generated optimized EEG features as conditions, using the stable diffusion model to generate a reference video frame; using the Canny edge detection algorithm to extract the contour information of the reference video frame; extracting the depth information of the reference video frame based on the Monodepth depth estimation algorithm; using the preset spatiotemporal conditional fusion encoder STCF-net to integrate the contour information and depth information into the conditions to obtain the fusion conditions; generating a high-quality reconstructed video based on the video diffusion model and the fusion conditions.

Citation Information

Patent Citations

  • Electroencephalogram visual classification method based on time-frequency domain fusion Transform

    CN114298216A

  • Electroencephalogram signal feature enhancement method based on reinforcement learning combined denoising and time-space relation modeling

    CN114841192A

  • Visual target semantic positioning map generation method based on electroencephalogram signals

    CN118015264A

  • Multi-scale space-time perception electroencephalogram data amplification method and system based on diffusion enhancement

    CN118940031A

  • Fixing Device For Fire-Fighting Pipe

    KR102195038B1

Cited By

  • Electroencephalogram vision reconstruction method and system for performing multi-mode alignment in combination with voice

    CN120540524A

  • A method and system for brain television reconstruction in conjunction with multi-modal alignment of speech

    CN120540524B