Two-channel electroencephalogram signal compression and reconstruction method based on variational auto-encoder
Through a deep learning method based on a variational autoencoder, combined with residual network and hyper-prior modeling, the problems of high complexity of EEG signal compression reconstruction algorithms and difficult parameter optimization in the existing technology are solved, and efficient and real-time EEG signal compression and reconstruction are achieved, which is suitable for real-time brain-computer interface systems.
Patent Information
- Application Number
- CN202510296262.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has problems such as high algorithm complexity, difficulty in parameter optimization, priority compression capability but high computational complexity in the compression and reconstruction of EEG signals, and cannot be applied to real-time brain-computer interface systems with high real-time performance and high signal quality requirements.
The dual-channel EEG signal compression reconstruction method based on a variational autoencoder is adopted, and the distribution characteristics of the original signal are learned through the deep learning model, and the trade-off between signal encoding code rate and signal reconstruction distortion is used as the objective function to realize adaptive EEG signal compression reconstruction. This method introduces residual network and hyper-priori modeling, which improves the generalization ability and stability of the model.
It realizes efficient EEG signal compression and reconstruction, reduces the risk of overfitting, improves the robustness of the model, and meets the high requirements of real-time brain-computer interface systems for real-time, expands its application scope in the field of brain-computer interfaces.
Smart Images

Figure CN120234560A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a dual-channel electroencephalogram signal compression and reconstruction method based on a variational autoencoder. Background Art
[0002] With the development of neuroscience, materials science, and the progress of artificial intelligence technology, wireless brain-computer interfaces based on electroencephalograms have become a hot research direction. The wireless brain-computer interface technology based on electroencephalograms has the advantages of non-invasiveness and high time resolution, can monitor electroencephalogram activities in real time, and is widely used in multiple application scenarios such as disease monitoring, emotion detection, and motor imagery. Correspondingly, after signal acquisition by a wireless sensor device, the transmitted signal needs to be transmitted through a wireless channel to the receiving end to complete subsequent tasks. However, wireless electroencephalogram acquisition devices generally include several to hundreds of acquisition channels, and the high sampling frequency, high sampling point resolution, and long sampling time all result in a large data volume of electroencephalogram signals before transmission. Therefore, in order to reduce transmission energy consumption and ensure the signal fidelity of downstream tasks after data is restored at the decoding end, it is of great significance to perform efficient adaptive compression on electroencephalogram signals before transmission. At the same time, electroencephalogram signals vary greatly between different people and at different detection times of the same person. For example, there are differences in electroencephalogram changes for the same event and in the energy values of electroencephalograms at different frequencies.
[0003] In brain-computer interface technology based on EEG signals, EEG signals are electrophysiological manifestations of brain activity and can explain an individual's cognitive state. Brain-computer interface (BCI) systems based on EEG signals build a bridge between the brain and external devices by accurately interpreting these signals, achieving efficient information transmission. This system not only allows users to control external actuators in almost real time, but also provides a new way to monitor and interpret cognitive states. In the medical field, EEG-based BCI technology was originally developed to help people with disabilities restore their communication and control abilities and improve their quality of life. However, with the continuous advancement of technology, its application scope has expanded from the medical field to non-medical fields, showing a wider range of application potential. In the non-medical field, EEG-based BCI technology is used to improve the life efficiency of healthy people, promote collaboration, and support personal development, such as optimizing the work and learning environment by monitoring attention, emotions, and fatigue status. In recent years, the development of EEG-based BCI technology has shown a clear trend of shifting from the medical field to the non-medical field. This trend not only reflects the widespread application of EEG-based BCI technology in non-medical fields, but also foreshadows its more profound social and economic impact in the future. As the technology matures further and costs decrease, EEG-based BCI technology is expected to be more widely used in multiple fields such as smart home control, entertainment, education, and security monitoring, bringing more convenience and innovation to people's lives.
[0004] Transformation-based EEG signal compression and reconstruction methods are an important research direction in the field of signal processing. They aim to reduce the amount of EEG signal data through mathematical transformation while retaining the key information in the signal as much as possible for easy storage and transmission. These methods usually use the sparsity or compressibility of EEG signals to achieve efficient data compression and reconstruction through transform domain processing. R. (2015). Fast DCT algorithms for EEG data compression in embedded systems. Comput. Sci. Inf. Syst., 12, 49 - 62.) The discrete cosine transform is used to divide the signal into a set of 8 samples, and each set of samples uses 3 different fast discrete cosine transforms. Finally, the unimportant parts in the coefficients are removed and filled with zeros before the inverse transform. The literature (B. Nguyen, D. Nguyen, W. Ma and D. Tran, "Wavelet transform and adaptive arithmetic coding techniques for EEG lossy compression," 2017 International Joint Conference on Neural Networks (IJCNN), Anchorage, AK, USA, 2017, pp. 3153 - 3160, doi: 10.1109 / IJCNN.2017.7966249.) transforms the original EEG signal from one - dimensional to two - dimensional matrix, and its main steps include discrete wavelet transform, quantization, thresholding, and adaptive arithmetic coding. The literature (L. Lin, Y. Meng, J. Chen, and Z. Li, ‘2015 Multichannel EEG compression based on ICA and SPIHT’, Biomedical Signal Processing and Control, vol. 20, pp. 45–51, Jul. 2015, doi: 10.1016 / j.bspc.2015.04.001 .) Based on the subspace compression method, the EEG signal is pre - processed using independent component analysis and principal component analysis, and then each independent component is arranged in matrix form, and then compressed using the set partitioning in hierarchical trees algorithm. EEG signal compression algorithms based on sparse representation and compressive sensing generally transform the signal from the time domain to the frequency domain for processing, and assume that it has a certain sparsity in the frequency domain (D. Kanemoto and T. Hirose, "EEG Measurements with Compressed Sensing Utilizing EEG Signals as the Basis Matrix," 2023 IEEE International Symposium on Circuits and Systems (ISCAS), Monterey, CA, USA, 2023, pp. 1 - 5, doi: 10.1109 / ISCAS46773.2023.10181710.).
[0005] The transform-based EEG signal compression and reconstruction method transforms the time-domain signal into a sparse signal in other domains through different domain transforms, thereby completing the transformation of the original dense signal. Although the compression method based on DCT transform has low computational complexity, it can only implement data compression and reconstruction schemes with several specific compression ratios, and there are significant differences in the signal fidelity reconstruction of different datasets. The hybrid compression method based on transform and entropy coding needs to consider the wavelet decomposition order, quantization bit number, threshold size, etc., which is cumbersome in parameter setting. The subspace compression method and the hierarchical tree set partitioning algorithm preprocess the original data in a method of information concentration, but it is necessary to complete the trade-off between the number of subspace principal components and signal reconstruction; at the same time, although the hierarchical tree set partitioning algorithm has the performance of gradual approximation, it is highly dependent on the data content and has a high computational complexity. The EEG signal compression algorithm based on compressive sensing assumes that the frequency domain has a certain sparsity. Its algorithm is relatively easy in sparse representation, but the compressive sensing algorithm usually involves the solution of non-convex optimization problems in the process of data reconstruction, and the computational complexity is high. The above-mentioned transform-based EEG signal compression and reconstruction algorithms have problems such as high algorithm complexity, difficult parameter optimization, limited compression ability, and high computational complexity, and are not applicable to real-time brain-computer interface systems with high real-time requirements and high signal quality requirements.
[0006] In recent years, the method of compressing and reconstructing electroencephalogram (EEG) signals based on deep learning has become a research hotspot. Its core advantage lies in the ability to automatically learn the abstract and high-level features in the data, thus eliminating the need to rely on manually designed feature extractors. This automated feature learning process not only improves efficiency but also enables a deeper exploration of the internal structure and patterns of EEG signals, providing a richer information basis for compression and reconstruction. In existing research, the literature (A. Ben Said, A. Mohamed and T. Elfouly, "Deep learning approach for EEG compression in mHealth system," 2017 13th International Wireless Communications and Mobile Computing Conference (IWCMC), Valencia, Spain, 2017, pp. 1508-1512, doi: 10.1109 / IWCMC.2017.7986507.), (Y. Cao, H. Zhang, Y.-B. Choi, H. Wang and S. Xiao, "Hybrid Deep Learning Model Assisted Data Compression and Classification for Efficient Data Delivery in Mobile Health Applications," in IEEE Access, vol. 8, pp. 94757-94766, 2020, doi: 10.1109 / ACCESS.2020.2995442.), (A. Valenti, M. Barsotti, R. Brondi, D. Bacciu and L. Ascari, "ROS-Neuro Integration of Deep Convolutional Autoencoders for EEG Signal Compression in Real-time BCIs," 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Toronto, ON, Canada, 2020, pp. 2019-2024, doi: 10.1109 / SMC42975.2020.9283397.) demonstrates the use of stacked convolutional layers to construct an autoencoder architecture to achieve end-to-end EEG signal compression.This method automatically extracts the features of EEG signals through a multi-layer convolutional network and reconstructs the signals at the decoding end. This end-to-end learning method simplifies the complex preprocessing and feature extraction steps in traditional compression methods, making the entire compression and reconstruction process more efficient and integrated. For example, the model proposed in the literature gradually reduces the spatial dimension of the signal through multi-layer convolution and pooling operations, while extracting higher-level feature representations, and then gradually restores the original dimension of the signal through the decoder, achieving efficient compression and good reconstruction quality. In addition, the literature (Dasan, E., & Gnanaraj, R. (2022). Joint ECG–EMG–EEG signal compression and reconstruction with incremental multimodal autoencoder approach. Circuits, Systems, and Signal Processing, 41, 6152-6181.), (Panneerselvam, Ithaya Rani. "Transfer learning autoencoder used for compressing multimodal biosignal." Multimedia Tools and Applications 81.13 (2022): 17547-17565.) further expands this field and proposes a multi-modal deep denoising convolutional autoencoder architecture for the joint compression and reconstruction of electrocardiogram (ECG), electromyogram (EMG), and electroencephalogram (EEG) signals. This method not only considers the compression of single-modal signals but also explores the correlation and complementarity between multi-modal signals. Through joint compression, the information between different modal signals can be more effectively utilized, improving the compression efficiency and better restoring the details of the signal during decompression. For example, the multi-modal autoencoder architecture in the literature learns the common features of different modal signals by sharing hidden layers while retaining the unique information of each modality, achieving high-quality signal reconstruction at a lower bit rate.
[0007] Currently, the electroencephalogram (EEG) signal compression and reconstruction methods based on deep learning mainly adopt the autoencoder architecture. However, such methods are prone to overfitting on datasets with a small sample size, resulting in the model's inability to fully learn the comprehensive features of the data. In addition, although the existing joint compression and reconstruction methods for multi-modal electrophysiological signals (such as electrocardiogram, electromyogram, and electroencephalogram) are innovative in technology, they have limitations in practical applications. Specifically, these methods assume that multiple types of electrophysiological signals can be obtained simultaneously during the acquisition stage, which is often difficult to achieve in real-world scenarios. Therefore, the applicability of such joint compression and reconstruction methods is relatively weak, especially in the application scenarios of EEG-based brain-computer interfaces, where their generality and practicality are limited. Summary of the Invention
[0008] To solve the problems existing in the prior art, the present invention provides a dual-channel EEG signal compression and reconstruction method based on variational autoencoders.
[0009] The present invention focuses on the source compression of EEG signals in the wireless brain-computer interface scenario, designs a deep learning-based EEG signal compression and reconstruction model, and uses an encoder-decoder model based on the variational encoder framework to learn the original signal distribution characteristics. By using the trade-off between the signal coding rate and the signal reconstruction distortion as the objective function, an adaptive EEG signal compression and reconstruction method based on deep learning is realized.
[0010] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0011] A dual-channel EEG signal compression and reconstruction method based on variational autoencoders provided by the present invention mainly includes the following steps:
[0012] Step 1: Construct a dual-channel EEG signal dataset;
[0013] Step 2: Construct an EEG signal compression and reconstruction model;
[0014] The EEG signal compression and reconstruction model is composed of an encoder module and a decoder module. First, through the theoretical derivation of the transformation distribution, the loss function for model training is obtained. The encoder module is used to extract the features of the dual-channel EEG signals, and the decoder module is used to generate the reconstructed EEG signals.
[0015] Step 3: Use the dual-channel EEG signal dataset to train the EEG signal compression and reconstruction model. By optimizing the network parameters, the EEG signal compression and reconstruction model can learn the compressed representation and reconstruction strategy of the EEG signals to minimize the difference between the original signal and the reconstructed signal.
[0016] Step 4: Input the dual-channel EEG signals to be compressed and reconstructed into the trained EEG signal compression and reconstruction model, and output the compressed representation and the reconstructed EEG signals. Evaluate the EEG signal compression and reconstruction effect by calculating the error index between the reconstructed signal and the original signal.
[0017] Furthermore, the dual-channel EEG signal dataset is obtained from the open-source EEG CHB-MIT dataset, and each EEG signal sample contains the EEG activities of two different channels.
[0018] Furthermore, the specific operation process of Step 4 is as follows:
[0019] S4.1: Estimate the system coding rate and the reconstruction distortion estimation model;
[0020] S4.2: Establish the model loss function;
[0021] S4.3: Evaluate the EEG signal compression and reconstruction index;
[0022] S4.4: Model the system objective.
[0023] Furthermore, in Step S4.1, first transform the original signal \(x\in R\) L into the latent variable \(y = g\) a (x;\(\theta\) g ), \(L\) represents the number of sampling points contained in the intercepted EEG signal time window, and \(g\) a represents the non-linear encoder of the model, and \(\theta\) g represents all the parameter weights of the neurons of the non-linear encoder; then scalar quantize the latent variable \(y\) to the discrete value point closest to it, that is \(Q\) represents the quantization operation, which quantizes the value to the discrete point value closest to it, thus generating a rate-distortion optimization problem. The distortion is the expected difference between the reconstructed vector and the original signal \(x\), and its mathematical expression is:
[0024]
[0025] where represents the error metric used.
[0026] Furthermore, in Step S4.2, assuming that the distributions of the prior \(q(y)\) are consistent, transform the optimization objective of variational inference into minimizing the KL loss between the distribution \(p(x)\) of the known variable \(x\) and the generated data distribution \(q(x)\), and its mathematical expression is:
[0027]
[0028] where the non-linear encoder \(g\) aThe output latent variable y is equivalent to a set of observed values of variable x, and the non-linear encoder g a is equivalent to the inference model p θ (y|x), and the non-linear decoder g s is equivalent to the generative model
[0029] The first term of formula (5) is a constant value and is ignored in the optimization problem; the second term represents the expected code length of the latent variable If the latent variable is modeled as a variable with an independent Gaussian distribution, the objective function is equivalent to the rate-distortion optimization problem, and the distortion metric is the mean square error.
[0030] Furthermore, the non-linear encoder g a includes:[[]]
[0031] (1) The first part: A convolutional layer converts the number of channels of the original input two-channel EEG signal into 64 channels;
[0032] (2) The second part: Four convolutional layers with gradually increasing expansion parameters, with a convolutional kernel size of 9, are used to obtain features with different levels of field of view;
[0033] (3) The third part: A downsampling step with a scaling factor of 2;
[0034] (4) The fourth part: Four convolutional layers with gradually increasing expansion parameters, with a convolutional kernel size of 5, further complete feature extraction;
[0035] (5) The fifth part: Through a channel self-attention layer, explicit modeling of the weights on each channel is completed, and then the original feature map is weighted, so that each channel has different amplitudes of importance, and the redundancy of the correlation between channels is removed;
[0036] (6) The sixth part: A downsampling step with a scaling factor of 2;
[0037] The non-linear decoder g s has the same model hierarchy and parameters as the non-linear encoder g a and forms symmetry.
[0038] Furthermore, by constructing a set of hyper priors z to represent the dependence relationship of the quantized latent variable y, that is, z = h a (y; θ h ), scalar quantization is performed on the hyper prior z The quantized hyper prior z is reconstructed through a non-linear transformation to obtain and formula (5) is transformed into minimizing the KL divergence between the variational probability density and the posterior probability density and its mathematical expression is:
[0039]
[0040] Among them, θ h and respectively represent all the parameter weights of the non-linear transformations h a (·) and h s (·) neurons. h a is the hyperprior encoder, and h s is the hyperprior decoder;
[0041] The first term in formula (6) is a constant value and is ignored in the optimization problem; the second term represents the bit rate of the latent variable under the condition of known side information ; the third term represents the bit rate of the side information ; the fourth term represents the distortion measure.
[0042] Furthermore, the hyperprior encoder h a includes:
[0043] (1) The first part: A convolutional layer converts the output result of the analysis transformation of the original input from the number of y channels to 32 channels;
[0044] (2) The second part: Three convolutional layers with sequentially increasing expansion parameters, with a convolutional kernel size of 9, are used to obtain features with different levels of vision;
[0045] (3) The third part: A downsampling step with a scaling factor of 2;
[0046] (4) The fourth part: Two convolutional layers with sequentially increasing expansion parameters, with a convolutional kernel size of 5, further complete feature extraction;
[0047] (5) The fifth part: A downsampling step with a scaling factor of 2;
[0048] The hyperprior decoder h s has the same model hierarchy and parameters as the hyperprior encoder h a and forms symmetry.
[0049] Furthermore, in step S4.3, the electroencephalogram signal compression reconstruction evaluation indexes include a reconstruction accuracy evaluation index and a compression efficiency evaluation index; the calculation formula of the reconstruction accuracy evaluation index is:
[0050]
[0051] The calculation formula of the compression efficiency evaluation index is:
[0052]
[0053] Further, in step S4.4, the calculation formula of the loss function of the electroencephalogram signal compression and reconstruction model is as follows:
[0054] L = (R y + R z ) + λD MSE (10)
[0055] where R y and R z respectively represent the bit rates of the latent variable and the hyperprior. λ is a positive trade-off parameter. Different values represent different trade-offs between the bit rate and the reconstruction accuracy. D MSE represents the distortion measure between the original signal and the reconstructed signal.
[0056] The beneficial effects of the present invention are as follows:
[0057] (1) Innovative technical architecture
[0058] Application of residual network: The residual network (ResNet) is introduced into the convolutional layers of the encoder and the decoder, effectively solving the degradation problem that occurs as the depth of the deep neural network increases. The residual block makes it easier for the network to learn the identity mapping, maintains the information consistency between the shallow layer and the deep layer, improves the generalization ability and stability of the model, and further enhances the compression and reconstruction performance.
[0059] Hyperprior modeling: A hyperprior is constructed to represent the dependencies between latent variables, and the hyperprior is subjected to non-linear transformation and quantization to more accurately simulate the distribution of latent variables. Using the hyperprior modeling method provides new ideas and means for optimizing the rate-distortion optimization problem, helps to further improve the compression efficiency and reconstruction quality, and reflects the innovation of the present invention in the technical architecture.
[0060] (2) High compression performance
[0061] Adaptive feature extraction: The deep convolutional neural network is used to automatically learn the features of the electroencephalogram signal, which can accurately capture the key information in the signal, so as to remove redundant data during the compression process and achieve efficient adaptive compression. Compared with the traditional transform-based method, there is no need to manually design the feature extractor, reducing human intervention and improving the accuracy and adaptability of feature extraction.
[0062] The quantization of the latent space vector is to quantize it to discrete points, which is convenient for subsequent discrete value arithmetic coding.
[0063] (3) Good generalization ability
[0064] Learning signal distribution: By learning the distribution characteristics of the original EEG signals, the model of the present invention can better adapt to different types of EEG signals, including signal differences between different individuals and within the same individual at different detection times. This enables the model to maintain stable compression and reconstruction performance when facing diverse EEG signal data, and has strong generalization ability.
[0065] Reduced overfitting risk: Compared with traditional autoencoder architectures, variational autoencoders introduce probability distributions and normal distribution priors. The model performs compression and reconstruction tasks by learning the distribution of input signals, and can learn the data distribution of unseen signals, thus alleviating the overfitting phenomenon to a certain extent. When the sample size is large enough, and the optimizer for model training is appropriately selected and the number of training epochs is sufficient for the model to converge to the optimal value, the model can fully learn the data features and improve the robustness of the model.
[0066] (4) Enhanced practicality
[0067] Applicable to real-time systems: The compression and reconstruction method of the present invention uses a deep learning model with short inference time, which can complete the compression and decompression processes of EEG signals in a short time, meeting the high real-time requirements of real-time brain-computer interface systems. This is crucial for brain-computer interface applications that require fast response, such as disease monitoring, motor imagery, etc., and can ensure that the system processes and transmits EEG signals in a timely and accurate manner.
[0068] Strong generality: The dual-channel compression and reconstruction framework for EEG signals is further extended based on single-channel EEG signal compression, removing the redundancy between channels, further improving the signal compression efficiency, and having a wider applicability. The high-efficiency EEG signal compression and reconstruction efficiency and data reconstruction quality expand its application scope in the field of brain-computer interfaces, providing strong support for the popularization and application of brain-computer interface technology. Brief description of the drawings
[0069] Figure 1 It is a flowchart of a dual-channel EEG signal compression and reconstruction method based on variational autoencoder provided by the present invention.
[0070] Figure 2 It is the EEG signal compression and reconstruction model constructed by the present invention.
[0071] Figure 3 It is the variational autoencoder mathematical inference model. Detailed implementation manners
[0072] The present invention will be further described in detail below with reference to the drawings.
[0073] A dual-channel electroencephalogram (EEG) signal compression and reconstruction method based on variational autoencoder provided by the present invention has an innovative technical architecture, efficient EEG signal compression performance, and strong generalization ability.
[0074] First, considering that EEG signals vary greatly among different people and at different detection times for the same person, the present invention designs an encoder-decoder model based on the variational autoencoder framework to learn the distribution characteristics of the original signals. Using the trade-off between the signal coding rate and the signal reconstruction distortion as the objective function, it realizes adaptive EEG signal compression and reconstruction. This design enables the model to better adapt to EEG signals with faster time-domain changes and higher randomness, improving the generalization ability and robustness of the model.
[0075] Second, the present invention realizes adaptive EEG signal compression and reconstruction by designing an encoding layer, a decoding layer, a quantization layer, and an entropy estimation layer based on a convolutional neural network. In terms of the encoder-decoder design, stacked convolutional layers are used to realize automatic EEG signal feature extraction, and the stacked convolutional layer structure uses a residual network (ResNet) to solve the degradation problem caused by the deepening of the deep neural network as the network deepens.
[0076] Finally, the present invention introduces a channel self-attention mechanism in the encoder and decoder parts. By explicitly modeling the weights on each channel through the channel self-attention module and then weighting the original feature map, each channel has different degrees of importance, thereby removing the correlation redundancy between channels and further improving the compression efficiency and reconstruction quality.
[0077] See Figure 1 As shown, a dual-channel EEG signal compression and reconstruction method based on variational autoencoder provided by the present invention has the following specific implementation process:
[0078] Step S1: Construct a dual-channel EEG signal dataset;
[0079] The dual-channel EEG signal dataset is obtained from the open-source EEG CHB-MIT dataset. Each EEG signal sample contains the EEG activities of two different channels, providing a data basis for the subsequent training of the compression and reconstruction model.
[0080] Step S2: Construct an EEG signal compression and reconstruction model;
[0081] The EEG signal compression and reconstruction model constructed by the present invention is a variational autoencoder (VAE) network based on a convolutional neural network (CNN) (CNN-based VAE network), mainly composed of an encoder module and a decoder module. Specifically, the encoder module is used to extract the features of the dual-channel EEG signals, and the decoder module is used to generate the reconstructed EEG signals according to the probability distribution.
[0082] In the present invention, the constructed electroencephalogram (EEG) signal compression and reconstruction model specifically uses a variational autoencoder network, which is a generative model. Its goal is to learn the latent representation and distribution of the data, so that the samples generated by sampling from this distribution can match the original signal. The neural network structure of the model introduces a normal distribution prior in the latent space, making the points in the latent space correspond to continuous changes in the data distribution. This makes the samples generated by sampling from the latent space more continuous and smooth in the data space. At the same time, the variational autoencoder network not only learns how to reconstruct the input data, but also learns how to generate samples that conform to the latent distribution. This feature can help the network better learn the temporal and spatial correlation relationships of EEG signals.
[0083] Step S3: Train the EEG signal compression and reconstruction model;
[0084] Use the constructed dual-channel EEG signal dataset to train the EEG signal compression and reconstruction model. By optimizing the network parameters, the network can learn the compressed representation and reconstruction strategy of the EEG signals to minimize the difference between the original signal and the reconstructed signal.
[0085] Step S4: Evaluate the EEG signal compression and reconstruction effect;
[0086] Input the dual-channel EEG signal to be compressed and reconstructed into the trained EEG signal compression and reconstruction model, and output the compressed representation and the reconstructed EEG signal. Evaluate the EEG signal compression and reconstruction effect by calculating the error index between the reconstructed signal and the original signal, so as to verify the effectiveness and practicality of the model.
[0087] Specifically, as Figure 2 shown, its specific implementation process is as follows:
[0088] S4.1: System coding rate estimation and reconstruction distortion estimation model;
[0089] The original signal, that is, the originally input EEG signal, can be represented as a vector \(x\in\mathbb{R}^L\), where \(L\) represents the number of sampling points included in the intercepted EEG signal time window, and the probability density function of the vector \(x\) can be expressed as \(p(x)\). Through the non-linear transformation based on the artificial neural network coding function, the source vector \(x\in\mathbb{R}^L\) is transformed into a latent variable \(y\), that is, \(y = g(x;\theta)\), where \(g\) represents the non-linear encoder of the model, and \(\theta\) represents all the parameter weights of the neurons in the non-linear encoder. L where \(L\) represents the number of sampling points included in the intercepted EEG signal time window, and the probability density function of the vector \(x\) can be expressed as \(p\) x (x). Through the non-linear transformation based on the artificial neural network coding function, the source vector \(x\in\mathbb{R}^L\) L is transformed into a latent variable \(y\), that is, \(y = g\) a (x;\theta g ), where \(g\) a represents the non-linear encoder of the model, and \(\theta\) g represents all the parameter weights of the neurons in the non-linear encoder.
[0090] Similarly, the latent variable obtained at the decoding end passes through the decoder \(g\) sObtain the reconstructed and restored vector This process can be expressed as where g s represents the non - linear decoder of the model, represents all the parameter weights of the non - linear encoder neurons.
[0091] The non - linear encoder g a outputs the latent variable y which is a continuous - value vector. For compression purposes, the latent variable y will be further scalar - quantized to the discrete - value point closest to it, that is Q represents the quantization operation, which quantizes the value to the nearest discrete - point value. Since quantization introduces errors, a rate - distortion optimization problem is thus generated. If entropy coding is performed on the quantized latent variable the expected code length of the quantized latent variable can be expressed as formula (1), where represents the probability density of the latent variable and represents taking the expectation under the probability distribution of the known source vector x.
[0092]
[0093] The distortion is the expected difference between the reconstructed vector and the original signal x, which can be expressed as formula (2), where represents the error metric used.
[0094]
[0095] S4.2: System Reconstruction Distortion Estimation Model;
[0096] As Figure 3 shown, in variational inference, its optimization goal is, given the distribution p(x) of the known variable x, to make the generated data distribution q(x) as consistent as possible with the distribution of the known variable x. This optimization problem can be transformed into minimizing the Kullback - Leibler Divergence (KL) loss between p(x) and q(x), and its mathematical expression is formula (3).
[0097] min D KL (p(x),q(x))(3)
[0098] The non - linear encoder g a outputs the latent variable y which is equivalent to a set of observed values of the variable x, and the non - linear encoder g a is equivalent to the inference model p θ (y|x), and the non - linear decoder g s is equivalent to the generative model First, assume that the prior q(y) is independent and identically distributed. Therefore, the above optimization problem is transformed into Equation (4), and can be further transformed into Equation (5).
[0099]
[0100] Among them, the first term of Equation (5) is a constant value and can be ignored in the optimization problem; the second term represents the expected code length of the latent variable . If the latent variable is modeled as a variable with an independent Gaussian distribution, the objective function is equivalent to the rate-distortion optimization problem, and the distortion measure is the mean squared error (MSE).
[0101] However, in fact, the elements of the latent variable are not independent of each other. Therefore, the entropy of the latent variable is overestimated, and the distribution of the latent variable cannot be accurately simulated. Thus, on this basis, another set of latent variables z is constructed to represent the dependence relationship of the latent variable , which is called the hyperprior in the present invention. Among them correspondingly, scalar quantization needs to be performed on the hyperprior z At this time, the optimization problem of Equation (5) is transformed into minimizing the KL divergence between the variational probability density and the posterior probability density , as shown in Equation (6). At the same time, the quantized hyperprior z is reconstructed through a non-linear transformation to obtain where, θ h and respectively represent all the parameter weights of the non-linear transformations h a (·) and h s (·) neurons.
[0102]
[0103] Among them, the first term in Equation (6) is a constant value and can be ignored in the optimization problem; the second term represents the bit rate of the latent variable under the condition of known side information ; the third term represents the bit rate of the side information ; the fourth term represents the distortion measure.
[0104] Under the encoding-decoding structure, each convolutional layer is a basic unit of a Residual Neural Network (ResNet). The encoding-decoding structure is stacked by ResNets. The advantage of using a residual neural network is that its unique residual blocks make it easier for the network to learn the identity mapping, which helps to maintain the information consistency between the shallow and deep layers, and helps to improve the generalization ability of the model and reduce overfitting.
[0105] As shown in Table 1, the non-linear encoder g a mainly consists of:
[0106] (1) The first part: A convolutional layer converts the number of channels of the original input dual-channel EEG signal into 64 channels.
[0107] (2) The second part: Through convolutional layers with 4 types of expanding parameters increasing in sequence, the kernel size is 9, which is used to obtain features with different levels of field of view.
[0108] (3) The third part: Through a downsampling step with a scaling factor of 2.
[0109] (4) The fourth part: Through convolutional layers with 4 types of expanding parameters increasing in sequence, the kernel size is 5, which further completes feature extraction.
[0110] (5) The fifth part: Through a channel self-attention layer (Squeeze-and-Excitation Networks, SENet), explicit modeling of the weights on each channel is completed, and then the original feature map is weighted, so that each channel has different degrees of importance, and the redundancy of the correlation between channels is removed.
[0111] (6) The sixth part: Through a downsampling step with a scaling factor of 2.
[0112] Correspondingly, the corresponding non-linear decoder g s has the same model hierarchy and parameters as the non-linear encoder g a and forms a symmetry.
[0113] Table 1
[0114] Module Number of channels Convolution kernel Dilation parameter Scaling Convolution layer Conv 64 9 0 - Dilated convolution layer Conv×4 64 9 1,2,4,8 - Downsampling 64 9 0 2 Dilated convolution layer Conv×4 64 5 1,2,4,8 - Channel self-attention layer 64 Fully connected layer 0 - Downsampling 4 5 0 2
[0115] Similarly, the corresponding model parameters of the hyperprior encoder h a are shown in Table 2 and mainly include 5 parts:
[0116] (1) The first part: A convolutional layer converts the number of channels of the output result y of the original input analysis transformation into 32 channels.
[0117] (2) The second part: Convolutional layers with three expansion parameters increasing sequentially, with a kernel size of 9, used to obtain features with different levels of field of view.
[0118] (3) The third part: A downsampling step with a scaling factor of 2.
[0119] (4) The fourth part: Convolutional layers with two expansion parameters increasing sequentially, with a kernel size of 5, to further complete feature extraction.
[0120] (5) The fifth part: A downsampling step with a scaling factor of 2.
[0121] Correspondingly, the corresponding hyperprior decoder h s and the hyperprior encoder h a have the same model levels and parameters, forming symmetry.
[0122] Table 2
[0123] Module Number of channels Convolution kernel Dilation parameter Scaling Convolution layer Conv 32 9 0 - Dilated convolution layer Conv×3 32 9 1,2,4 - Downsampling 32 9 0 2 Dilated convolution layer Conv×2 32 5 1,2 - Downsampling 2 5 0 2
[0124] S4.3: Evaluation metrics for EEG signal compression and reconstruction;
[0125] For EEG signals, the objective reconstruction accuracy evaluation metric is the percentage root mean square difference (PRD), and its calculation formula is shown in Formula (7). The smaller the percentage root mean square difference, the smaller the error between the original signal x k and the reconstructed signal , and the higher the compression and reconstruction accuracy.
[0126]
[0127] For EEG signals, the evaluation metric for compression efficiency is the compression ratio (CR), and its calculation formula is shown in Formula (8), which is the ratio of the size of the original data (original data size) to the size of the compressed data, and the larger its value, the higher the compression efficiency.
[0128]
[0129] S4.4: System objective modeling;
[0130] According to the variational inference results, the loss function of the model is defined as Formula (9), and the distortion measure between the original signal and the reconstructed signal is the mean square error MSE. Then Formula (9) can be simplified to Formula (10), where R y and R zrespectively represent the bit rates of the latent variable and the hyperprior, and λ is a positive-valued trade-off parameter. Different values represent different trade-offs between the bit rate and the reconstruction accuracy.
[0131]
[0132] L = (R y + R z ) + λD MSE (10)
[0133] A dual-channel electroencephalogram (EEG) signal compression and reconstruction method based on variational autoencoder provided by the present invention. First, in the scenario of a brain-computer interface based on EEG signals, a compression and reconstruction model of EEG signals based on deep learning is proposed. Through innovative technical architectures such as the application of residual networks and hyperprior modeling, this model solves the degradation problem of deep neural networks and more accurately simulates the distribution of latent variables, thereby improving the generalization ability and stability of the model and further enhancing the performance of compression and reconstruction. Second, the present invention focuses on the temporal correlation of single-channel EEG signals and the spatial correlation between channels, and realizes the compression and reconstruction of dual-channel EEG signals. Through adaptive feature extraction and the construction of a continuous latent space, the present invention can efficiently remove redundant data in EEG signals, achieve efficient adaptive compression, and reconstruct the signal more smoothly during decompression, reduce errors, and improve the reconstruction quality. In addition, by learning the distribution characteristics of the original EEG signals, the present invention enhances the adaptability of the model to different types of EEG signals, reduces the risk of overfitting, and improves the robustness of the model. Third, the present invention aims to improve the practicality and applicability of the EEG signal compression and reconstruction method. The compression and reconstruction method of the present invention uses a deep learning model and can complete the compression and decompression processes of EEG signals in a short time, meeting the high requirements for real-time performance of real-time brain-computer interface systems. At the same time, the present invention has a wider range of applicability. The present invention can perform efficient compression and reconstruction, expands its application scope in the field of brain-computer interfaces, and provides strong support for the popularization and application of brain-computer interface technologies.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A dual-channel EEG signal compression reconstruction method based on variational autoencoder, characterized in that: The following steps are involved: Step 1: Construct a dual-channel EEG signal dataset; Step 2: Construct an EEG signal compression reconstruction model; The EEG signal compression reconstruction model is composed of an encoder module and a decoder module. First, the loss function of the model training is obtained through theoretical derivation of the transformation distribution; the encoder module is used to extract the features of the dual-channel EEG signal, and the decoder module is used to generate the reconstructed EEG signal; Step 3: Use the dual-channel EEG signal dataset to train the EEG signal compression reconstruction model. By optimizing the network parameters, the EEG signal compression reconstruction model can learn the compression representation and reconstruction strategy of the EEG signal to minimize the difference between the original signal and the reconstructed signal. Step 4: Input the dual-channel EEG signal to be compressed and reconstructed into the trained EEG signal compression and reconstruction model, output the compressed representation and reconstructed EEG signal, and judge the EEG signal compression and reconstruction effect by calculating the error index between the reconstructed signal and the original signal.
2. According to claim 1, a dual-channel EEG signal compression reconstruction method based on variational autoencoder is characterized in that: The dual-channel EEG signal dataset is obtained from the open source EEG signal CHB-MIT dataset, and each EEG signal sample contains EEG activities of two different channels.
3. According to claim 1, a dual-channel EEG signal compression reconstruction method based on variational autoencoder is characterized in that: The specific operation process of step 4 is as follows: S4.1: System coding rate estimation and reconstruction distortion estimation model; S4.2: Establish model loss function; S4.3: Evaluation index of EEG signal compression reconstruction; S4.4: System goal modeling.
4. According to claim 3, a dual-channel EEG signal compression reconstruction method based on variational autoencoder is characterized in that: In step S4.1, the original signal x∈R L Transformed into latent variable y = g a (x;θ g ), L represents the number of sampling points contained in the intercepted EEG signal time window, g a represents the nonlinear encoder of the model, θ g Indicates that it contains all parameter weights of nonlinear encoder neurons; Then the latent variable y is scalar-quantized to the nearest discrete value point, that is, Q represents the quantization operation, which quantizes the value to the nearest discrete point value, thus generating a rate-distortion optimization problem. The distortion is the reconstruction vector The expected difference from the original signal x is expressed mathematically as: in, Indicates the error metric to use.
5. According to claim 3, a dual-channel EEG signal compression reconstruction method based on variational autoencoder is characterized in that: In step S4.2, assuming that the distribution of the prior q(y) is consistent, the optimization objective of variational inference is transformed into minimizing the KL loss between the distribution p(x) of the known variable x and the generated data distribution q(x), and its mathematical expression is: Among them, the nonlinear encoder g a The output latent variable y is equivalent to a set of observations of the variable x, and the nonlinear encoder g a Equivalent to inferring the model p θ (y|x), nonlinear decoder g s Equivalent to a generative model The first term of formula (5) is a fixed value and is ignored in the optimization problem; the second term represents the hidden variable The expected code length is If the variables are modeled as independent Gaussian distributions, the objective function is equivalent to the rate-distortion optimization problem, and the distortion metric is the mean square error.
6. The dual-channel EEG signal compression reconstruction method based on variational autoencoder according to claim 5 is characterized in that: The nonlinear encoder g a include: (1) Part 1: A convolutional layer converts the original input two-channel EEG signal into 64 channels; (2) The second part: After four convolutional layers with increasing expansion parameters, the convolution kernel size is 9, which is used to obtain features at different levels of vision; (3) The third part: after a downsampling step with a scaling factor of 2; (4) Part 4: After four convolutional layers with increasing expansion parameters, the convolution kernel size is 5, and feature extraction is further completed; (5) The fifth part: After the channel self-attention layer, the weight on each channel is explicitly modeled, and then the original feature map is weighted so that each channel has a different amplitude importance, and the redundancy of the correlation between channels is removed; (6) Part 6: After a downsampling step with a scaling factor of 2; The nonlinear decoder g s With nonlinear encoder g a The model levels and parameters are consistent, forming a symmetrical structure.
7. The dual-channel EEG signal compression reconstruction method based on variational autoencoder according to claim 5 is characterized in that: By constructing a set of hyper-prior z to represent the dependency of the quantized latent variable y, i.e. z = h a (y;θ h ), scalar quantization of the hyper-prior z The quantized super prior z is reconstructed through nonlinear transformation to obtain And transform formula (5) into minimizing the variational probability density and the posterior probability density The KL divergence of is expressed as: Among them, θ h and They represent the nonlinear transformation h a (·) and h s (·) All parameter weights of neurons, h a is the super prior encoder, h s is a super-prior decoder; The first term in formula (6) is a fixed value and is ignored in the optimization problem; the second term represents the Under the condition of hidden variables The third term represents the side information The fourth term represents the distortion measure.
8. The dual-channel EEG signal compression reconstruction method based on variational autoencoder according to claim 7 is characterized in that: The super prior encoder h a include: (1) Part 1: A convolutional layer converts the number of channels of the output result y of the original input analysis transformation to 32 channels; (2) The second part: After three convolutional layers with increasing expansion parameters, the convolution kernel size is 9, which is used to obtain features at different levels of vision; (3) The third part: after a downsampling step with a scaling factor of 2; (4) Part 4: After two convolutional layers with increasing expansion parameters, the convolution kernel size is 5, and feature extraction is further completed; (5) Part 5: After a downsampling step with a scaling factor of 2; The super-a priori decoder h s With the super prior encoder h a The model levels and parameters are consistent, forming a symmetrical structure.
9. The dual-channel EEG signal compression reconstruction method based on variational autoencoder according to claim 3 is characterized in that: In step S4.3, the EEG signal compression reconstruction evaluation index includes a reconstruction accuracy evaluation index and a compression efficiency evaluation index; the calculation formula of the reconstruction accuracy evaluation index is: The calculation formula of the compression efficiency evaluation index is:
10. The dual-channel EEG signal compression reconstruction method based on variational autoencoder according to claim 3, characterized in that: In step S4.4, the loss function of the EEG signal compression reconstruction model is calculated as: L=(R y +R z )+λD MSE (10) Among them, R y and R z denote the latent variable and the hyper-prior bit rate respectively, λ is a positive trade-off parameter, and different values represent different trade-offs between the bit rate and the reconstruction accuracy. MSE Represents the distortion measure between the original signal and the reconstructed signal.