Epilepsy electroencephalogram data expansion method and device based on exponential power diffusion model
Through the epilepsy EEG data expansion method based on the exponential power diffusion model, the shape parameters are adaptively estimated and the forward diffusion module and the reverse denoising module are constructed, the problem of complex characteristics and insufficient data volume of epilepsy EEG data set is solved. The generated data is highly similar to the real data, which improves the performance of the epilepsy prediction model.
Patent Information
- Application Number
- CN202510734935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing epilepsy EEG data sets have the problem of complex feature distribution and insufficient data volume. The existing generative models such as VAE and GAN are difficult to process complex data, and the generated data lacks diversity and is far from the distribution of real data.
Using an exponential power diffusion model, the forward diffusion module and the reverse denoising module are constructed by setting the best shape parameters, and the shape parameters are adaptively estimated using maximum likelihood estimation and quasi-Newtonian optimization algorithm, combined with the U-Net network for data expansion to generate high-quality epilepsy EEG data.
The modeling accuracy and stability of epilepsy EEG data is improved, and the generated data is highly similar to the real data distribution, solving the problem of insufficient data volume and improving the performance of epilepsy prediction model.
Smart Images

Figure CN120260962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of epilepsy detection, and particularly to a method and device for epileptic electroencephalogram data augmentation based on an exponential power diffusion model. Background Art
[0002] Epilepsy is a common chronic neurological disease with an extremely complex seizure mechanism, involving abnormal discharge phenomena of brain neurons. Although electroencephalogram (EEG) has been widely used as an efficient detection tool in epilepsy seizure prediction, the characteristics of EEG signals, such as high-dimensionality, non-linearity, high individual differences, and difficulty in acquisition, make the existing epilepsy datasets generally have problems of complex feature distribution and insufficient data volume, which in turn leads to the difficulty for existing epilepsy prediction methods to meet the expectations of clinical use.
[0003] To break through the data bottleneck in the field of epilepsy prediction, researchers have proposed different generative models for augmenting epileptic EEG data, especially the Variational Autoencoder (VAE) and Generative Adversarial Networks (GAN). However, these generative models are limited by their model architectures and optimization functions, making it difficult to process complex data, and the generated data often lacks diversity. Specifically, VAE learns the latent distribution of data by optimizing the Evidence Lower Bound (ELBO) and samples from the latent space when generating data. Since the latent space is usually low-dimensional and continuous, the model tends to generate smooth but limited-diversity data, and its performance is often limited when facing high-dimensional complex data. On the other hand, GAN generates data through the adversarial training of a generator and a discriminator. The generator tends to generate samples that can deceive the discriminator, which may lead to mode collapse, that is, the generator only generates a few types of samples and lacks diversity. At the same time, GAN also often shows unstable situations during training. In recent years, with the significant achievements of diffusion models in fields such as image generation, speech synthesis, and biosignal processing, diffusion models have become an important direction in generative model research and a main method for data augmentation. However, directly applying diffusion models to the task of epileptic EEG data augmentation still faces many problems. Due to the characteristics of EEG data such as high-dimensionality, complexity, and individual differences, it is often difficult to capture and model its true feature distribution. Traditional diffusion models based on Gaussian distribution assume that data follows a simple probability distribution. However, a simple Gaussian distribution cannot fully model complex EEG signals, resulting in a large gap between the generated data distribution and the true data distribution. Summary of the Invention
[0004] In view of this, the present invention provides a method and device for epileptic EEG data augmentation based on an exponential power diffusion model, so as to solve the problem in the prior art that a simple Gaussian distribution cannot fully model complex EEG signals, resulting in a large gap between the generated data distribution and the real data distribution.
[0005] In a first aspect, the present invention provides a method for epileptic EEG data augmentation based on an exponential power diffusion model, the method comprising: Setting the optimal shape parameter of the exponential power distribution; Constructing a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter; Performing forward diffusion on the forward diffusion module based on the exponential power distribution with the pre-acquired original EEG data to obtain noise data; Training the reverse denoising module based on the exponential power distribution with the noise data obtained by forward diffusion to obtain a trained exponential power diffusion model; Sampling the probability density function of the preset exponential power distribution to generate noise data subject to a Gaussian distribution, and inputting the noise data subject to a Gaussian distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
[0006] The method for epileptic EEG data augmentation based on an exponential power diffusion model provided by the present invention can adaptively fit the complex distribution characteristics of epileptic EEG data by setting the optimal shape parameter of the exponential power distribution. Compared with traditional fixed-parameter models, it can more accurately depict EEG data in different stages such as the interictal, preictal, ictal, and postictal stages of epilepsy, effectively retain the characteristic information of the original data, and improve the quality and effectiveness of data augmentation. Based on the optimal shape parameter, a forward diffusion and a reverse denoising module are constructed. The forward diffusion module can achieve controllable noise addition to the data, and the reverse denoising module can efficiently learn the data distribution pattern and has stronger feature extraction and restoration capabilities when reconstructing the original data, significantly improving the performance of the entire exponential power diffusion model, with higher stability and reliability. Forward diffusion is performed on the original EEG data to obtain noise data, and the reverse denoising module is trained with this. The entire training process makes full use of the information of the original data. The trained exponential power diffusion model can reconstruct and generate a large number of high-quality new EEG data from the probability density function of the preset exponential power distribution, realizing the modeling of complex EEG signals, effectively solving the problem of insufficient sample size of epileptic EEG data. The new EEG data generated by this method is highly similar to the real data distribution, solving the problem in the prior art that a simple Gaussian distribution cannot fully model complex EEG signals, resulting in a large gap between the generated data distribution and the real data distribution.
[0007] In an optional implementation manner, setting the optimal shape parameter of the exponential power distribution includes: Adaptive estimation of the optimal shape parameter of the input data based on the maximum likelihood estimation method and the quasi - Newton optimization algorithm.
[0008] In an alternative embodiment, adaptive estimation of the optimal shape parameter of the input data based on the maximum likelihood estimation method and the quasi - Newton optimization algorithm includes: Set the probability density function of the exponential power distribution; Based on the probability density function and the pre - acquired original EEG data, use the maximum likelihood estimation algorithm to calculate the maximum likelihood function, and calculate the log - likelihood function corresponding to the maximum likelihood function; Equivalent the log - likelihood function to an unconstrained optimization problem, and use the quasi - Newton optimization algorithm to solve the unconstrained optimization problem to obtain the optimal shape parameter of the input data.
[0009] An epilepsy EEG data augmentation method based on the exponential power diffusion model provided by the present invention can fully explore the potential distribution law of epilepsy EEG data by setting the probability density function of the exponential power distribution and combining the maximum likelihood estimation method to calculate the maximum likelihood function and the log - likelihood function. Different from the traditional fixed - parameter setting, this method can dynamically adjust the shape parameter according to the characteristics of the original EEG data, so that the exponential power distribution can better fit the complex distributions of epilepsy EEG data in different stages such as the inter - seizure period and the seizure period, accurately depict the data characteristics, and lay a solid foundation for subsequent data processing and model construction. Transforming the log - likelihood function into an unconstrained optimization problem and using the quasi - Newton optimization algorithm to solve it avoid problems such as high computational complexity and slow convergence speed that may exist in traditional optimization algorithms. The adaptively estimated optimal shape parameter can enhance the expression ability of the exponential power diffusion model for epilepsy EEG data. The forward diffusion and reverse denoising modules constructed based on this parameter can more accurately simulate the real changes of the data during the data diffusion and denoising reconstruction processes, effectively retain the data features, thereby improving the stability and reliability of the exponential power diffusion model, and finally generating high - quality augmented EEG data, providing more powerful data support for the research and clinical diagnosis of epilepsy diseases.
[0010] In an alternative embodiment, construct a forward diffusion module based on the exponential power distribution based on the optimal shape parameter, including: Specify the Markov chain of the forward process, set the time step of the forward diffusion, and construct an input data sample based on the pre - acquired original EEG data; Based on the Markov chain, use a scheduling strategy within the time step to generate a constant sequence corresponding to the time step for controlling the mean and variance of the exponential power distribution; Construct a forward diffusion process with a preset number of steps, and gradually add exponentially powered noise controlled by a constant sequence to the input data sample at each step based on the optimal shape parameter, so that the input data sample is converted into a sample sequence obeying the exponentially powered distribution after the above time steps, and a forward diffusion module based on the exponentially powered distribution is obtained.
[0011] A method for augmenting epileptic EEG data based on an exponentially powered diffusion model provided by the present invention specifies the Markov chain of the forward process and sets the time step, and can decompose the diffusion process of the data into ordered discrete steps, so that the data change in each step follows a specific law and has high predictability. This controllable diffusion process avoids irregular changes in the data during processing, ensures the stability and consistency of the data during the diffusion process, and facilitates the analysis and regulation of the data diffusion process. Based on the optimal shape parameter, combined with the scheduling strategy, the input data sample obeys the exponentially powered distribution after the above time steps. Since the optimal shape parameter is adaptively estimated according to the original EEG data and can accurately adapt to the data distribution, both the input data sample and the added noise can more realistically simulate the changes of epileptic EEG data in practice. Compared with the traditional diffusion method, this method can more accurately capture the data characteristics, make the diffused data more conform to the internal distribution law of the original data, and improve the quality of data augmentation. Constructing a forward diffusion process with a preset number of steps ensures the standardization and consistency of the data processing flow. Adding noise based on the optimal shape parameter makes the forward diffusion module highly compatible with the reverse denoising module in data feature processing. In the subsequent model training and data reconstruction process, this stability helps to reduce the fluctuations in model training, accelerate the model convergence speed, improve the overall performance of the exponentially powered diffusion model, and thus generate more reliable augmented EEG data.
[0012] In an optional implementation manner, constructing a forward diffusion process with a preset number of steps, and gradually adding exponentially powered noise controlled by a constant sequence to the input data sample at each step based on the optimal shape parameter, so that the input data sample is converted into a sample sequence obeying the exponentially powered distribution after the above time steps, and a forward diffusion module based on the exponentially powered distribution is obtained, including: Introduce the sample sequence obtained by gradually adding exponentially powered noise in the forward diffusion process into the reparameterization technique, and reconstruct the sample sequence obeying the exponentially powered distribution into a form conforming to the gamma distribution and the uniform distribution; Convert the exponentially powered distribution into a new representation form after linear scaling and translation of the noise sampled from the uniform distribution; Calculate the mean and variance of the specified sub-term in the new representation form, and calculate the relationship from the start time to any time in the forward diffusion process based on the Lyapunov central limit theorem, and use the relationship as the forward diffusion module based on the exponentially powered distribution.
[0013] An epileptic EEG data augmentation method based on the exponential power diffusion model provided by the present invention. The reparameterization technique reconstructs the exponential power distribution into the form of a gamma distribution and a uniform distribution. This conversion can directly sample from the standard distribution, avoiding complex numerical calculations and greatly improving the sampling efficiency. Linear scaling and translation convert the exponential power distribution into the form of a linearly transformed uniform distribution sampling, further simplifying the calculation process, reducing the computational complexity, and making the implementation of the forward diffusion process more efficient. Through the linear transformation that converts the exponential power distribution into a uniform distribution, the model has stronger adaptability to different data characteristics. When processing data with complex distribution characteristics such as epileptic EEG data, it can more accurately capture the data characteristics and enhance the flexibility and adaptability of the model. The relationship obtained based on the Lyapunov central limit theorem provides a more intuitive basis for adjusting the model parameters. The model parameters can be flexibly adjusted according to actual needs to optimize the model performance. Using the Lyapunov central limit theorem to derive the relationship from the starting moment to any moment provides a solid theoretical support for the forward diffusion process. The obtained relationship clearly expresses the change law of the data in the diffusion process, facilitating the analysis and understanding of the model and providing theoretical guidance for the optimization and improvement of the model.
[0014] In an alternative embodiment, a reverse denoising module based on the exponential power distribution is constructed based on the optimal shape parameter, including: Input the pre-acquired original EEG data into the forward diffusion module based on the exponential power distribution for forward diffusion to obtain noisy data; Construct a U-Net network for encoding time step information based on the forward diffusion module based on the exponential power distribution. Input the noisy data into the U-Net network, learn the distribution pattern of the noisy data through a multi-layer convolution, downsampling, and upsampling structure, add residual connections and attention mechanisms to the U-Net network, and at the same time integrate the time step information into each layer of the U-Net network through the time embedding layer to obtain the output of the U-Net network; Construct a reverse denoising process with a preset number of steps based on the output of the U-Net network, and continuously remove noise from the random noise according to the exponential power distribution at each time step based on the optimal shape parameter to obtain a reverse denoising module based on the exponential power distribution.
[0015] An epilepsy EEG data augmentation method based on the exponential power diffusion model provided by the present invention customizes the exponential power distribution through the optimal shape parameter, making the reverse denoising process more conform to the actual distribution characteristics of epilepsy EEG data, capable of more accurately identifying and removing noise, and retaining the characteristics of real EEG signals. The encoder and decoder architectures formed by the multi-layer convolution and downsampling-uppersampling structures can capture EEG signal characteristics at different scales, and can effectively process from microscopic waveform details to macroscopic time series patterns. The use of residual connections, attention mechanisms, and time embedding layers enhances the feature extraction ability of the U-Net network. The reverse process of the preset steps corresponds to the time steps of the forward diffusion, forming a clear denoising path to ensure that the process of gradually recovering the original signal from the noisy data is controllable and stable. The exponential power distribution constraint removes noise according to the exponential power distribution at each time step, forming a theoretical closed loop with the forward diffusion process, ensuring the mathematical consistency and physical meaning of the denoising process. The noise data generated by the forward diffusion is directly used as the input of the reverse denoising. The two modules are constructed based on the same exponential power distribution framework, seamlessly connected, and reducing the data conversion error. The exponential power distribution can better describe the spike-and-heavy-tail characteristics of EEG signals, especially the abnormal discharge patterns during epileptic seizures, which has more advantages than the traditional Gaussian distribution.
[0016] In an alternative embodiment, a reverse denoising process with a preset number of steps is constructed based on the output of the U-Net network, and noise is continuously removed from the random noise according to the exponential power distribution at each time step based on the optimal shape parameter, obtaining a reverse denoising module based on the exponential power distribution, including: Construct a learning model to learn the first conditional probability of the exponential power distribution, and calculate the second conditional probability using Bayes' formula based on the first conditional probability and the sampled data points at the starting moment; Obtain a preset number of exponential power distributions obtained in the forward diffusion, and substitute the preset number of exponential power distributions into Bayes' formula to obtain a deformed form of the second conditional probability; Based on the deformed form of the second conditional probability and the standard exponential power distribution, obtain the analytical expressions of the corresponding mean and variance, and obtain the expression of the sampled data points at the starting moment based on the derivation process of the forward diffusion process; Substitute the expression of the sampled data points at the starting moment back into the analytical expressions of the mean and variance to obtain a reverse denoising module based on the exponential power distribution.
[0017] An epilepsy EEG data augmentation method based on an exponential power diffusion model provided by the present invention establishes a strict mathematical framework by learning the conditional probability of the exponential power distribution and applying Bayes' formula, making the reverse denoising process have a solid theoretical basis. The analytical expressions of the mean and variance are obtained from the exponential power distribution, avoiding complex numerical calculations, improving the interpretability of the model and the accuracy of derivation. Substituting a preset number of exponential power distributions into Bayes' formula to obtain a deformed form realizes form optimization, simplifies the calculation process, reduces the amount of calculation, and makes the model more efficient in processing large-scale EEG data. Compared with the sampling-based method, the analytical expression directly gives the calculation methods of the mean and variance, without the need for a large number of samplings and iterations, significantly improving the speed of reverse denoising. Customizing the exponential power distribution based on the optimal shape parameter enables the model to better adapt to the non-Gaussian characteristics of epilepsy EEG data, improving the accuracy of noise removal. The expression of the sampling data points at the starting moment is obtained through the derivation of the forward diffusion process and substituted into the mean and variance expressions, ensuring the temporal consistency between the reverse denoising process and the forward diffusion, and making the reconstructed data more in line with physiological laws. Combining the noise data distribution pattern learned by the U-Net network with Bayesian derivation makes full use of the feature extraction ability of the network and improves the recognition and recovery ability of epilepsy-specific patterns.
[0018] In a second aspect, the present invention provides an epilepsy EEG data augmentation device based on an exponential power diffusion model, and the device includes: An optimal shape parameter setting module, configured to set the optimal shape parameter of the exponential power distribution; An exponential power diffusion model construction module, configured to construct a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter; A forward diffusion module, configured to perform forward diffusion on the forward diffusion module based on the exponential power distribution based on the pre-obtained original EEG data to obtain noise data; A reverse denoising module, configured to train the reverse denoising module based on the exponential power distribution based on the noise data of the forward diffusion to obtain a trained exponential power diffusion model; A reconstruction and augmentation module, configured to sample the probability density function of the preset exponential power distribution to generate noise data subject to a Gaussian distribution, and input the noise data subject to a Gaussian distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
[0019] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the epilepsy EEG data augmentation method based on the exponential power diffusion model in the first aspect or any corresponding implementation manner thereof.
[0020] Fourthly, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the epilepsy EEG data augmentation method based on the exponential power diffusion model according to the first aspect or any corresponding embodiment thereof above.
[0021] Fifthly, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the epilepsy EEG data augmentation method based on the exponential power diffusion model according to the first aspect or any corresponding embodiment thereof above. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 is a schematic flowchart of the epilepsy EEG data augmentation method based on the exponential power diffusion model according to an embodiment of the present invention; Figure 2 is a schematic flowchart of another epilepsy EEG data augmentation method based on the exponential power diffusion model according to an embodiment of the present invention; Figure 3 is a schematic flowchart of yet another epilepsy EEG data augmentation method based on the exponential power diffusion model according to an embodiment of the present invention; Figure 4 is a schematic diagram of the forward diffusion and reverse denoising processes according to an embodiment of the present invention; Figure 5 is a schematic diagram of another forward diffusion and reverse denoising processes according to an embodiment of the present invention; Figure 6 is a schematic diagram of the U-Net network structure in the reverse denoising module based on the exponential power diffusion model according to an embodiment of the present invention; Figure 7 is a schematic flowchart of the U-Net network working process according to an embodiment of the present invention; Figure 8 is a schematic diagram of the EEG waveform of epileptic seizures generated by the exponential power diffusion model according to an embodiment of the present invention; Figure 9 is a schematic diagram of the EEG waveform of non-epileptic seizures generated by the exponential power diffusion model according to an embodiment of the present invention; Figure 10It is a structural block diagram of an epileptic EEG data augmentation device based on an exponential power diffusion model according to an embodiment of the present invention; Figure 11 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] First, the principle of the classical diffusion model is briefly introduced here. The classical diffusion model proposed by Jonathan Ho et al. is constructed based on the Gaussian distribution and a series of its mathematical properties. The two core processes of the classical diffusion model are the forward diffusion and the reverse denoising process. In the Markov chain of the forward diffusion process, Jonathan Ho gradually adds Gaussian noise (i.e., Gaussian noise) that conforms to the Gaussian distribution at each time step, so that after T time steps, the original real data approaches a random Gaussian noise, that is, the original data follows the Gaussian distribution after T time steps, and at the same time, the model learns the Gaussian noise added at each time step. In the reverse denoising process, the model predicts the Gaussian noise added in the forward process at this step at each time step according to the learned pattern, and continuously denoises according to the predicted noise of the model through a series of mathematical transformations, and finally restores the original data.
[0026] According to the above process, it can be seen that the Gaussian distribution is the core component of the two processes of the diffusion model. Whether it is forward diffusion or reverse denoising, it is inseparable from the mathematical principle of the Gaussian distribution. And there is an important assumption in the diffusion model, that is, the feature distribution of the original data follows the Gaussian distribution, which is mainly reflected in that data can be continuously denoised and generated from random Gaussian noise. When processing natural image data, the feature distribution of the image is close to the Gaussian distribution, which makes this assumption valid, that is, the diffusion model based on the Gaussian distribution has high performance in processing image tasks. However, when facing some more complex data sets, such as the epileptic EEG data set mentioned in the present invention, the simple Gaussian distribution often no longer has the ability to accurately capture the data features, which makes it necessary to use a distribution pattern with stronger and more flexible modeling capabilities, that is, the core technical innovation point of this invention - the diffusion model based on the exponential power distribution is introduced.
[0027] An embodiment of the present invention provides a method for augmenting epileptic EEG data based on an exponential power diffusion model. The exponential power distribution, also known as the power-exponential distribution. Compared with the Gaussian distribution in the classical diffusion model in the prior art, the exponential power distribution introduces a shape parameter, making it have a more flexible form and stronger modeling ability. By reconstructing the core formulas of its forward diffusion and reverse denoising, the embodiment of the present invention effectively improves the modeling ability of the diffusion model for real complex data, and applies the diffusion model based on the exponential power distribution (abbreviation: exponential power diffusion model) to the task of augmenting epileptic EEG data with complex feature distributions. By constructing a large-scale and high-quality EEG data set, it solves the problem that the simple Gaussian distribution in the prior art cannot fully model complex EEG signals, resulting in a large gap between the generated data distribution and the real data distribution, and can also solve problems such as uneven class distribution, low data quality, and small scale commonly existing in existing data sets.
[0028] According to an embodiment of the present invention, an embodiment of a method for augmenting epileptic EEG data based on an exponential power diffusion model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0029] In this embodiment, a method for augmenting epileptic EEG data based on an exponential power diffusion model is provided, which can be used in the above computer device. Figure 1 is a flowchart of a method for augmenting epileptic EEG data based on an exponential power diffusion model according to an embodiment of the present invention, as Figure 1 shown, this process includes the following steps: Step S101, set the optimal shape parameter of the exponential power distribution.
[0030] Specifically, the shape parameter of the exponential power distribution (Exponential Power distribution, EP) refers to the numerical parameter that affects the shape of the distribution. The optimal shape parameter not only affects the shape of the distribution, but also determines the concentration degree and tail thickness of the exponential power distribution, such as a sharp peak and thick tail or a flat and thin tail. An adaptive parameter method is used to adaptively estimate the optimal shape parameter of the input data , and the optimal shape parameter is closely related to the probability density function of the exponential power distribution, improving the ability of the exponential power diffusion model to depict the input data.
[0031] Step S102, construct a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter.
[0032] Specifically, the forward diffusion module based on the exponential power distribution refers to the process in which data is gradually added with noise during the forward diffusion process and finally becomes Gaussian noise. Specifically, the forward diffusion process adds noise to the data step by step through a series of mathematical models until the data distribution approaches a standard Gaussian distribution.
[0033] The reverse noise module based on the exponential power distribution is the inverse process of the forward diffusion process, aiming to recover data from the noise. The reverse noise module based on the exponential power distribution includes a U-Net network. The result of the forward diffusion is used as the data of the reverse noise module through the U-Net network, and finally the original data is reconstructed from the noisy data.
[0034] Step S103, perform forward diffusion on the forward diffusion module based on the exponential power distribution using the pre-acquired original EEG data to obtain noisy data.
[0035] Specifically, the pre-acquired original EEG data is the EEG data collected and recorded for each patient during the interictal, pre-ictal, ictal, and post-ictal periods. Before inputting the original EEG data into the forward diffusion module based on the exponential power distribution, preprocessing of the original EEG data is required, including: Resample the original EEG data, and the sampling frequency is determined according to needs, such as 512 Hz or 256 Hz. Then, use a sliding window to intercept segments according to a certain time step. The window size is determined according to needs, such as 1 s, 10 s, 30 s, etc., and the intercepted segments are stored by category.
[0036] Construct the input of the exponential power diffusion model (including the forward diffusion module based on the exponential power distribution and the reverse noise module based on the exponential power distribution). According to the period to which the EEG data activity in the preprocessed segment belongs, that is, the interictal, pre-ictal, ictal, and post-ictal periods, select the data type that needs to be augmented, and load the data of this category completed in the preprocessing step. And further perform data processing before inputting into the model: (1) Normalize the data using the Min-Max method; (2) Pad or truncate the data in the channel dimension to ensure the consistency of the data size. Obtain the input data (B, L, C, D) of the exponential power diffusion model, where B is the batch size, L is the segment length, C is the number of channels, and D is the feature dimension (determined by the sampling frequency).
[0037] Input the preprocessed original EEG segments into the forward diffusion module based on the exponential power distribution, and perform training on the forward diffusion of the forward diffusion module based on the exponential power distribution to finally obtain noisy data.
[0038] Step S104, train the reverse denoising module based on the exponential power distribution using the noisy data from the forward diffusion to obtain a trained exponential power diffusion model.
[0039] Specifically, the output result of the forward diffusion module based on the exponential power distribution, i.e., the noise data, is used as the input of the reverse denoising module based on the exponential power distribution to train the reverse denoising module based on the exponential power distribution. That is, the noise data is input into the U-Net network of the reverse denoising module based on the exponential power distribution to train the U-Net network. After T times of denoising iterations, a trained exponential power diffusion model is obtained and new denoised data is reconstructed and generated. .
[0040] Step S105: Sample the probability density function of the preset exponential power distribution to generate noise data that follows the exponential power distribution, and input the noise data that follows the exponential power distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
[0041] Specifically, load the trained exponential power diffusion model, sample the exponential power distribution through the forward diffusion module in the exponential power diffusion model to obtain a large amount of noise data that follows the exponential power distribution, input these noise data into the exponential power diffusion model, and through the sampling module in the reverse denoising module, a large amount of high-quality epileptic EEG data is denoised and reconstructed.
[0042] The method for augmenting epileptic EEG data based on the exponential power diffusion model provided in this embodiment can adaptively fit the complex distribution characteristics of epileptic EEG data by setting the optimal shape parameter of the exponential power distribution. Compared with the traditional fixed-parameter model, it can more accurately depict the EEG data in different stages such as the interictal, pre-ictal, ictal, and post-ictal periods of epilepsy, effectively retain the characteristic information of the original data, and improve the quality and effectiveness of data augmentation. The forward diffusion and reverse denoising modules are constructed based on the optimal shape parameter. The forward diffusion module can achieve controllable noise addition to the data, and the reverse denoising module can efficiently learn the data distribution pattern and has stronger feature extraction and restoration capabilities when reconstructing the original data, which significantly improves the performance of the entire exponential power diffusion model, and has higher stability and reliability. Forward diffusion is performed on the original EEG data to obtain noise data, and the reverse denoising module is trained with this. The entire training process makes full use of the information of the original data. The trained exponential power diffusion model can reconstruct and generate a large amount of high-quality new EEG data from the probability density function of the preset exponential power distribution, realizing the modeling of complex EEG signals and effectively solving the problem of insufficient sample size of epileptic EEG data. The new EEG data generated by this method is highly similar to the real data distribution, solving the problem in the prior art that a simple Gaussian distribution cannot fully model complex EEG signals, resulting in a large gap between the generated data distribution and the real data distribution.
[0043] In this embodiment, a method for augmenting epileptic EEG data based on an exponential power diffusion model is provided, which can be used in the above-mentioned computer device. Figure 2 It is a flowchart of a method for augmenting epileptic EEG data based on an exponential power diffusion model according to an embodiment of the present invention, as Figure 2 shown. The process includes the following steps: Step S201, set the optimal shape parameter of the exponential power distribution.
[0044] Specifically, the above step S201 includes: Step a, adaptively estimate the optimal shape parameter of the input data based on the maximum likelihood estimation method and the quasi-Newton optimization algorithm.
[0045] In some optional embodiments, the above step a includes: Step a1, set the probability density function of the exponential power distribution.
[0046] First, define the probability density function of the exponential power distribution as: (1); where is the gamma function, is the central moment parameter, is the mean parameter, is the shape parameter, represents a random variable, and here only the probability density function of the exponential power distribution is defined first.
[0047] Step a2, calculate the maximum likelihood function based on the probability density function and the pre-acquired original EEG data using the maximum likelihood estimation algorithm, and calculate the logarithmic likelihood function corresponding to the maximum likelihood function.
[0048] Specifically, the pre-acquired original EEG data is processed through data preprocessing and input data processing to obtain a training dataset. The training dataset is , and its likelihood function can be calculated: (2); Then, its logarithmic likelihood function can be obtained: (3); where is the sample data in the training dataset; is the number of sample data.
[0049] Step a3, equivalent the logarithmic likelihood function to an unconstrained optimization problem, and use the quasi-Newton optimization algorithm to solve the unconstrained optimization problem to obtain the optimal shape parameter of the input data.
[0050] Specifically, the log-likelihood function is equivalent to an unconstrained optimization problem, that is, to find the shape parameter that maximizes the value of the log-likelihood function , that is, the optimal shape parameter. This problem can be formulated as: (4); Use the BFGS algorithm (BFGS algorithm, an inverse rank-2 quasi-Newton method) in the Quasi-Newton Method to solve this unconstrained optimization problem and obtain the optimal shape parameter . For the specific solution steps, refer to the related technology and will not be elaborated here
[0051] Step S202: Construct a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter
[0052] Specifically, the specific process of constructing a forward diffusion module based on the exponential power distribution based on the optimal shape parameter is as follows (1) Set the time step T, which is used to set the number of time steps that the model needs to go through during the forward diffusion process
[0053] (2) Select a scheduling strategy, such as: linear scheduling, square scheduling or inverse time scheduling, to generate a data sequence of length T, which is used to control the noise intensity of the exponential power noise at each step during the forward diffusion process
[0054] (3) According to the adaptive parameter method in step S101, set the value of the shape parameter to control the shape of the exponential power distribution, such as a sharp peak and thick tail or a flat and thin tail
[0055] (4) Construct a forward diffusion process with T steps, and add exponential power noise to the data at each step according to the shape parameter . The original data will obtain the diffused noise data after T steps
[0056] The specific process of constructing a reverse denoising module based on the exponential power distribution based on the optimal shape parameter is as follows (1) Construct a U-Net network, which is used to encode the information of the time step , and learn the distribution pattern of the input noise data through a multi-layer convolution and downsampling-upsampling structure, and adopt residual connections and attention mechanisms to enhance the feature extraction ability. At the same time, the time step information is integrated into each layer of the network through a time embedding layer
[0057] (2) Construct a reverse denoising process with steps, and gradually reconstruct the original data from the new denoised data according to the output of the U-Net network. .
[0058] Furthermore, in the above step S202, a forward diffusion module based on the exponential power distribution is constructed based on the optimal shape parameter. Specifically, the derivation and construction of the specific forward diffusion process include: Step S2021: Specify the Markov chain of the forward process, set the time step of the forward diffusion, and construct the input data sample based on the pre-acquired original EEG data.
[0059] Specifically, given a data point sampled from the distribution of real data (original EEG data) , and define a forward process, which is a Markov chain, that is, the current state is only determined by the previous state and has nothing to do with the states at any other time. respectively represent the distribution of real data represents the starting moment value of the sampled data point.
[0060] Step S2022: Based on the Markov chain, use the scheduling strategy within the time step to generate a constant sequence corresponding to the time step for controlling the mean and variance of the exponential power distribution.
[0061] Specifically, in this forward diffusion process, through steps, a series of samples are generated, where the step size is given by the set sequence, and it is assumed that each obtained ( represents the data point obtained after t time steps) satisfies the exponential power distribution, that is, the entire forward process simultaneously satisfies the following formula: (5); where, is the exponential power distribution, is the shape parameter of the exponential power distribution, represents the conditional probability distribution, that is, under the given condition, follows the distribution, represents the identity matrix, the joint variable of the data point , represents that under the given condition, the joint variable follows the distribution, Denotes hyperparameters used to control the mean and variance of the exponential power distribution.
[0062] Step S2023: Construct a forward diffusion process with a preset number of steps. Based on the optimal shape parameter, gradually add exponential power noise controlled by a constant sequence to the input data samples at each step, so that the input data samples are converted into a sample sequence obeying the exponential power distribution after the above time steps, and a forward diffusion module based on the exponential power distribution is obtained.
[0063] In some alternative embodiments, the above step S2023 includes: Step b1: Introduce the reparameterization technique into the sample sequence obtained by gradually adding exponential power noise in the forward diffusion process, and reconstruct the sample sequence obeying the exponential power distribution into a form conforming to the gamma distribution and the uniform distribution.
[0064] Specifically, to ensure the differentiability of the sampling process and enable gradient propagation in subsequent training, the reparameterization trick is introduced to reconstruct the exponential power distribution into the following form: (6); Where, Is the gamma distribution, Is the uniform distribution, , , All represent intermediate variables and have no practical meaning.
[0065] Step b2: Convert the exponential power distribution into a new representation form after linear scaling and translation of the noise sampled from the uniform distribution.
[0066] Specifically, therefore, the exponential power distribution in formula (6) can be re-represented by the noise Sampled from the uniform distribution after linear scaling and translation as the following form: (7); Let , And , , we can get: (8); Where, , Both represent variables obeying the uniform distribution, , Respectively represent the components obtained by linearly translating and converting the corresponding intermediate variables z and Through the reparameterization technique, , , All represent intermediate process quantities, and there are 。
[0067] Step b3: Calculate the mean and variance of the specified sub - terms in the re - representation form, and calculate the relationship formula from the starting moment to any moment in the forward diffusion process based on the Lyapunov central limit theorem. Take the relationship formula as the forward diffusion module based on the exponential - power distribution.
[0068] Specifically, consider the sub - term in the above formula , to obtain and The relationship between them requires calculating the mean and variance of this sub - term: (9); (10); where, represents the mean of the sub - term , represents the variance of the sub - term . k represents the k - th sub - term among all the sub - terms mentioned above, and k takes an integer value from 0 to t - 1.
[0069] For the sub - terms and (11); Get: (12); Immediately afterwards, it is necessary to verify whether it satisfies the Lyapunov condition, that is, whether there exists such that when , it satisfies: (13); Let , , we can get: (14); It can be seen that the Lyapunov condition holds, that is, it satisfies the Lyapunov central limit theorem. Let: (15); Then there is , where is a Gaussian distribution. According to the properties of the Gaussian distribution, the above formula can be written in the following form: (16); where, there is .
[0070] Therefore, substituting the above formula into the chain formula of the forward diffusion process, a more simplified form can be obtained: (17); Finally, in the forward diffusion process, the relationship from the starting moment to any moment is obtained, that is to The relationship, whose exponential power distribution satisfies: (18); It can be verified that when the shape parameter the exponential power distribution is equivalent to the standard Gaussian distribution, and the above formula (18) can be simplified to: (19).
[0071] That is, when the shape parameter of the exponential power distribution while the exponential power distribution degenerates into the standard Gaussian distribution, the forward diffusion formula of the exponential power diffusion model also degenerates into the formula of the forward diffusion process of the traditional diffusion model.
[0072] The above content details the construction of the forward diffusion process of the exponential power diffusion model and the derivation of the core formula in the embodiments of the present invention from a theoretical level. The exponential power diffusion model not only theoretically expands the basis of the traditional diffusion model, but also provides the possibility of modeling more complex noise distributions by introducing the shape parameter 𝑝 of the exponential power distribution.
[0073] Furthermore, in step S202, a reverse denoising module based on the exponential power distribution is constructed based on the optimal shape parameter, that is, the derivation and construction of the specific reverse denoising process include: Step S2024, input the pre-acquired original EEG data into the forward diffusion module based on the exponential power distribution for forward diffusion to obtain noise data.
[0074] Specifically, the schematic diagrams of the forward diffusion and reverse denoising processes are as shown in Figure 4 and Figure 5 The result of the forward diffusion module is used as the input of the reverse denoising module. Figure 4 and Figure 5 The X0, X1, X2,..., X T in all represent sampled data points.
[0075] Step S2025, construct a U-Net network for encoding time step information based on the forward diffusion module of the exponential power distribution, input the noise data into the U-Net network, learn the distribution pattern of the noise data through multi-layer convolution, downsampling, and upsampling structures, add residual connections and attention mechanisms in the U-Net network, and integrate the time step information into each layer of the U-Net network through the time embedding layer to obtain the output of the U-Net network.
[0076] Specifically, construct a U-Net network: The U-Net network structure is as shown in Figure 6 , and the working process of the U-Net network is as shown in Figure 7 . Through its encoder part, multi-level feature extraction is performed on the input noisy data , gradually capturing the global and local information of the data. Each layer of the encoder converts the input data into a series of high-dimensional feature representations through convolutional operations and downsampling operations. These feature representations not only contain the statistical characteristics of the noisy data but also implicitly contain the dynamic information of the time step .
[0077] In the decoder part, the U-Net network gradually reconstructs the features extracted by the encoder into denoised data through upsampling and skip connections. The role of the skip connections is to combine the low-level features of the encoder with the high-level features of the decoder, thereby retaining more detailed information and avoiding the loss of important structural features during the denoising process. To further enhance the model's ability to model the time step t, a time embedding mechanism is added to each layer of the U-Net. The time embedding maps the time step t into a high-dimensional vector and combines it with the input features, enabling the model to dynamically adjust its parameters to adapt to different noise levels. This design enables the model to adaptively learn the distribution pattern of the noisy data according to the change of the time step t during the reverse denoising process, thereby improving the accuracy and stability of denoising. In addition, to enhance the model's expressive ability, an attention mechanism is introduced in the convolutional layer of the U-Net network. The attention mechanism enables the model to better capture the long-range dependencies in the data by calculating the correlations between feature maps, thereby more accurately restoring the global structure of the data during the denoising process.
[0078] Step S2026: Based on the output of the U-Net network, construct a reverse denoising process with a preset number of steps, and continuously remove noise from the random noise according to the exponential power distribution at each time step based on the optimal shape parameter, obtaining a reverse denoising module based on the exponential power distribution.
[0079] In some optional embodiments, the process of constructing the reverse denoising module of the exponential power diffusion model: The reverse denoising process corresponds to the forward diffusion process. Different from the forward process that gradually adds random noise to the data , the reverse process removes noise from the random noise continuously according to the exponential power distribution at each time step. However, the model cannot directly estimate because this requires the entire dataset. The above step S2026 includes: Step c1, construct a learning model to learn the first conditional probability of the exponential power distribution, and calculate the second conditional probability based on the first conditional probability and the sampled data points at the starting moment using the Bayesian formula.
[0080] Specifically, therefore, a model needs to be learned to approximate the conditional probability (the first conditional probability), so as to realize the whole reverse denoising process, that is: (19); Although is unknown, but after adding the condition later, (the second conditional probability) can be calculated by the Bayesian formula: (20).
[0081] Step c2, obtain a preset number of exponential power distributions obtained in the forward diffusion, and substitute the preset number of exponential power distributions into the Bayesian formula to obtain a deformed form of the second conditional probability.
[0082] Specifically, in the derivation process from step S2021 to step S2023, the following three exponential power distributions are obtained: (21); Let: . Substitute the three exponential power distributions of formula (21) into the Bayesian formula (20), and we can get: (22); Among them, is the part that does not involve and can be omitted, is a mathematical symbol, indicating proportional to.
[0083] Step c3, obtain the analytical expressions of the corresponding mean and variance based on the deformed form of the second conditional probability and the standard exponential power distribution, and obtain the expression of the sampled data points at the starting moment based on the derivation process of the forward diffusion process.
[0084] Specifically, continue to further simplify the above formula (22), and following the forward diffusion module, according to the standard exponential power distribution, the analytical expressions of the mean and variance can be given: (23); Among them, , respectively represent the mean and variance of the reverse denoising. Step c4, substitute the expression of the sampled data points at the starting moment back into the analytical expressions of the mean and variance to obtain the reverse denoising module based on the exponential power distribution.
[0085] Specifically, according to the derivation process of forward diffusion in steps S2021 to S2023, it can be obtained that: (24); where is the output of the U-Net network.
[0086] Substitute the expression of into the expression of , and the mean and variance of the data obtained after each denoising step in the reverse denoising process can be obtained. Through its mean and variance, the noisy data can be sampled and reconstructed.
[0087] It can be verified that when the shape parameter of the exponential power distribution, the exponential power distribution is equivalent to the Gaussian distribution, , and The expressions of can be simplified to: (25).
[0088] That is, when the shape parameter of the exponential power distribution, while the exponential power distribution degenerates into the standard Gaussian distribution, in the exponential power diffusion model , and The expressions of also degenerate into the core formula of the reverse denoising process in the traditional diffusion model.
[0089] So far, all the derivations of the core formula in the reverse denoising process of the exponential power diffusion model have been completed. Combining the forward diffusion module and the reverse denoising module, the entire exponential power diffusion model is systematically constructed.
[0090] Step S203, perform forward diffusion on the forward diffusion module based on the exponential power distribution using the pre-acquired original EEG data to obtain noisy data. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.
[0091] Step S204, train the reverse denoising module based on the exponential power distribution using the noisy data from forward diffusion to obtain a trained exponential power diffusion model. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.
[0092] Step S205, sample the probability density function of the preset exponential power distribution to generate noisy data that follows the exponential power distribution, and input the noisy data that follows the exponential power distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data. For details, please refer toFigure 1 Step S105 of the illustrated embodiment will not be elaborated here.
[0093] The method for augmenting epileptic EEG data based on the exponential power diffusion model provided in this embodiment uses the maximum likelihood estimation method and the quasi-Newton optimization algorithm to adaptively estimate the optimal shape parameter . In the exponential power diffusion model, one of the major factors affecting the model performance is the shape parameter . The selection of the shape parameter is one of the core parameters of the exponential power distribution, which directly determines the shape characteristics of the distribution. By means of the adaptive parameter method, the cumbersome process and resource waste caused by manual parameter debugging are avoided, and at the same time, the model performance can be effectively improved. Introducing the exponential power distribution into the traditional diffusion model reconstructs the core formulas of the forward diffusion and reverse denoising of the diffusion model. When facing complex data, the diffusion model based on the simple Gaussian distribution often has difficulty in accurately capturing the true distribution characteristics of the data. This is because the distributions of many actual data do not fully conform to the Gaussian distribution, but may exhibit characteristics of super-Gaussian (peaked and heavy-tailed) or sub-Gaussian (flat and thin-tailed). Therefore, based on the classical diffusion model, the present invention introduces a more flexible exponential power distribution, enabling the model to have the ability to depict super-Gaussian and sub-Gaussian distributions, enhancing the learning and modeling capabilities of the diffusion model, and better completing the task of data augmentation.
[0094] In this embodiment, a method for augmenting epileptic EEG data based on the exponential power diffusion model is provided, which can be used in the above-mentioned computer device. Figure 3 is a flowchart of the method for augmenting epileptic EEG data based on the exponential power diffusion model according to the embodiment of the present invention. As Figure 3 shown, the process includes the following steps: Step S301, set the optimal shape parameter of the exponential power distribution. For details, please refer to Figure 2 Step S201 of the illustrated embodiment will not be elaborated here.
[0095] Step S302, construct a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter. For details, please refer to Figure 2 Step S202 of the illustrated embodiment will not be elaborated here.
[0096] Step S303, perform forward diffusion on the forward diffusion module based on the exponential power distribution based on the pre-acquired original EEG data to obtain noise data.
[0097] Specifically, construct a training module for the exponential power diffusion model: construct a training sub-module for the forward diffusion module, and randomly perform steps of the forward diffusion process on the input data of the exponential power diffusion model to obtain noise data .
[0098] First, select the input data from the dataset, and through the input steps of constructing the exponential power diffusion model, select the data types to be augmented according to the periods to which the EEG activities in the segments belong, namely the interictal period, pre-ictal period, ictal period, and post-ictal period. Load the data of this category that has completed data preprocessing, and perform data processing before inputting into the model: (1) Normalize the data using the Min-Max method; (2) Pad or truncate the data in the channel dimension to ensure the consistency of the data size. Obtain the input data of the exponential power diffusion model , where is the batch size, is the segment length, is the number of channels, is the feature dimension (determined by the sampling frequency).
[0099] Perform a random forward diffusion process for steps on the input data after the input steps of constructing the exponential power diffusion model to obtain the noisy data .
[0100] Step S304: Train the inverse denoising module based on the exponential power distribution using the noisy data from the forward diffusion to obtain the trained exponential power diffusion model.
[0101] Specifically, construct a noise prediction module: Input the time step and the noisy data into the U-Net network to obtain the predicted noise ; Finally, calculate the mean square error between the two as the loss value according to the exponential power noise added to the data at the th time step and the noise predicted by the inverse denoising module, and train the U-Net network.
[0102] Construct a sampling module: First, sample a noise from the exponential power distribution; then, iteratively reconstruct the noisy data . At each iteration, input the current time step and the noisy data into the trained U-Net network to obtain the predicted noise at the current time step . Calculate the corresponding mean , variance through the and obtained during the inverse denoising process, and then reconstruct ; Finally, after After the next denoising iteration, new noise data is generated. .
[0103] Step S305: Sample the probability density function of the preset exponential power distribution to generate noise data that follows the exponential power distribution, and input the noise data that follows the exponential power distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
[0104] Specifically, during the epileptic EEG data augmentation process, first load the U-Net network of the trained exponential power diffusion model, including the noise prediction module and the sampling module, and set the exponential power diffusion model to the evaluation mode. Randomly sample the probability density function of the exponential power distribution to generate a large amount of initial noise data; then input these noise data into the trained exponential power diffusion model, and use the sampling module in the reverse denoising module to perform step-by-step denoising and reconstruction. Finally, a large amount of EEG data with similar time-frequency characteristics to real epileptic EEG signals is generated. This process ensures the high quality and diversity of the generated data by adjusting the diffusion step size and the adaptive shape parameter, thereby effectively augmenting the clinically scarce EEG datasets during epileptic seizures and the pre-epileptic seizure period, providing sufficient data support for the training of subsequent epileptic prediction models.
[0105] The method for augmenting epileptic EEG data based on the exponential power diffusion model provided in this embodiment customizes the exponential power distribution through the optimal shape parameter, making the reverse denoising process more conform to the actual distribution characteristics of epileptic EEG data, and can more accurately identify and remove noise, retaining the characteristics of real EEG signals. The encoder and decoder architectures formed by the multi-layer convolution and downsampling-upsampling structures can capture EEG signal characteristics at different scales, and can effectively process both microscopic waveform details and macroscopic time series patterns. The use of residual connections, attention mechanisms, and time embedding layers enhances the feature extraction ability of the U-Net network. The reverse process of the preset steps corresponds to the time steps of the forward diffusion, forming a clear denoising path to ensure that the process of gradually restoring the original signal from the noise data is controllable and stable. The exponential power distribution constraint removes noise according to the exponential power distribution at each time step, forming a theoretical closed loop with the forward diffusion process, ensuring the mathematical consistency and physical meaning of the denoising process. The noise data generated by the forward diffusion is directly used as the input for the reverse denoising. The two modules are constructed based on the same exponential power distribution framework, seamlessly connected, reducing data conversion errors. The exponential power distribution can better describe the spiky and heavy-tailed characteristics of EEG signals, especially the abnormal discharge patterns during epileptic seizures, and has more advantages than the traditional Gaussian distribution.
[0106] As one or more specific application embodiments of the present invention, in combination with Figure 8 and Figure 9A further detailed description of the epilepsy EEG data augmentation method based on the exponential power diffusion model provided by the present invention is as follows: In this embodiment, the training batch size is set to 64, the Adam optimizer is used (learning rate 2e-4, β1 = 0.9, β2 = 0.999), and a total of 1000 epochs are trained. The diffusion process timesteps are set to 2000, the input data shape is (30, 16, 256), and the linear scheduling strategy is used for the diffusion process. The present invention aims to solve the problems commonly existing in existing epilepsy EEG datasets, such as small scale, low quality, and class imbalance. The present invention includes three modules: a forward diffusion module based on the exponential power distribution, a reverse denoising module based on the exponential power distribution, and a training and sampling module of the exponential power diffusion model. Among them, the forward diffusion module based on the exponential power distribution can diffuse EEG data into exponential power noise, and the reverse denoising module based on the exponential power distribution can reconstruct the original EEG data from the exponential power noise. Through the training and sampling module of the exponential power diffusion model, new high-quality EEG data can be reconstructed from random noise. In two experiments of the present invention, a publicly available EEG dataset CHB-MIT and a publicly available natural image dataset CIFAR-10 are respectively used to conduct qualitative and quantitative experimental verification on the performance of the present invention. Among them, in the dataset CHB-MIT, typical features during epileptic seizures, such as sharp waves and spike waves, can be clearly captured from the epileptic seizure signals generated by the exponential power diffusion model (as shown in Figure 8); normal EEG activity patterns, such as alpha waves and beta waves, can be observed from the generated non-epileptic seizure signals (as shown in Figure 9). In the dataset CIFAR-10, the FID (Fréchet Inception Distance) index of the exponential power diffusion model in the data generation task reaches 3.1889, and the IS (Inception Score) index reaches 9.0846 ± 0.0761. Both indexes are better than those of the classical diffusion model. The results show that the exponential power diffusion model is superior to the classical diffusion model in both data generation quality and data generation diversity. Figure 8 It is 30s epileptic seizure data generated by the exponential power diffusion model trained based on the CHB-MIT dataset, Figure 9 It is 30s non-epileptic seizure data generated by the exponential power diffusion model trained based on the CHB-MIT dataset.
[0107] The epilepsy EEG data augmentation method based on the exponential power diffusion model provided by the present invention shows unique advantages in capturing the characteristics of complex EEG signals and learning data distribution. The present invention can effectively generate large-scale, high-quality, and diverse epilepsy EEG data, break through the data bottleneck in the field of epilepsy prediction, and thus effectively improve the performance of epilepsy prediction methods. This technological breakthrough effectively solves the long-existing data scarcity problem in the field of epilepsy prediction, significantly expands the data scale available for model training, and more importantly, by providing rich and high-quality training samples, can significantly enhance the generalization ability and diagnostic accuracy of epilepsy prediction models. This makes the present invention have higher practical value and clinical application potential.
[0108] In this embodiment, an epilepsy EEG data augmentation device based on the exponential power diffusion model is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0109] This embodiment provides an epilepsy EEG data augmentation device based on the exponential power diffusion model, as Figure 10 shown, including: The optimal shape parameter setting module 1001 is used to set the optimal shape parameter of the exponential power distribution.
[0110] The exponential power diffusion model construction module 1002 is used to construct a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter.
[0111] The forward diffusion module 1003 is used to perform forward diffusion on the forward diffusion module based on the exponential power distribution based on the pre-acquired original EEG data to obtain noise data.
[0112] The reverse denoising module 1004 is used to train the reverse denoising module based on the exponential power distribution with the noise data of the forward diffusion to obtain a trained exponential power diffusion model.
[0113] The reconstruction and augmentation module 1005 is used to sample the probability density function of the preset exponential power distribution to generate noise data that follows a Gaussian distribution, and input the noise data that follows a Gaussian distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
[0114] In some alternative implementation manners, the optimal shape parameter setting module 1001 includes: An adaptive estimation unit for adaptively estimating the optimal shape parameter of input data based on the maximum likelihood estimation method and the quasi-Newton optimization algorithm.
[0115] In some alternative embodiments, the adaptive estimation unit includes: A probability density function setting subunit for setting the probability density function of the exponential power distribution.
[0116] A log-likelihood function calculation subunit for calculating the maximum likelihood function based on the probability density function and the pre-acquired original EEG data using the maximum likelihood estimation algorithm, and calculating the log-likelihood function corresponding to the maximum likelihood function.
[0117] An optimal shape parameter calculation subunit for equating the log-likelihood function to an unconstrained optimization problem and solving the unconstrained optimization problem using the quasi-Newton optimization algorithm to obtain the optimal shape parameter of the input data.
[0118] In some alternative embodiments, the exponential power diffusion model construction module 1002 includes: A time step setting unit for specifying the Markov chain of the forward process, setting the time step of the forward diffusion, and constructing an input data sample based on the pre-acquired original EEG data.
[0119] A sample sequence generation unit for generating a constant sequence corresponding to the time step for controlling the mean and variance of the exponential power distribution within the time step based on the Markov chain using a scheduling strategy.
[0120] A forward diffusion module construction unit for constructing a forward diffusion process with a preset number of steps, and gradually adding exponential power noise controlled by the constant sequence to the input data sample in each step based on the optimal shape parameter, so that the input data sample is converted into a sample sequence subject to the exponential power distribution after the above time step, and obtaining a forward diffusion module based on the exponential power distribution.
[0121] In some alternative embodiments, the forward diffusion module construction unit includes: A reparameterization introduction and reconstruction subunit for introducing the reparameterization technique into the sample sequence obtained by gradually adding exponential power noise in the forward diffusion process, and reconstructing the sample sequence subject to the exponential power distribution into a form conforming to the gamma distribution and the uniform distribution.
[0122] A first derivation subunit for converting the exponential power distribution into a new representation form obtained by linearly scaling and translating the noise sampled from the uniform distribution.
[0123] A second derivation sub-unit, configured to calculate the mean and variance of specified sub-items in the re-representation form, and calculate the relationship from the start time to any time during the forward diffusion process based on the Lyapunov central limit theorem, and use the relationship as a forward diffusion module based on the exponential power distribution.
[0124] In some alternative embodiments, the exponential power diffusion model construction module 1002 further includes: A forward diffusion unit, configured to input pre-acquired original EEG data into the forward diffusion module based on the exponential power distribution for forward diffusion to obtain noise data.
[0125] A U-Net network construction unit, configured to construct a U-Net network for encoding time step information based on the forward diffusion module based on the exponential power distribution, input the noise data into the U-Net network, learn the distribution pattern of the noise data through a multi-layer convolution and downsampling and upsampling structure, add residual connections and attention mechanisms to the U-Net network, and at the same time integrate the time step information into each layer of the U-Net network through a time embedding layer to obtain the output of the U-Net network.
[0126] A reverse denoising module construction unit, configured to construct a reverse denoising process with a preset number of steps based on the output of the U-Net network, and continuously remove noise from random noise according to the exponential power distribution at each time step based on the optimal shape parameter to obtain a reverse denoising module based on the exponential power distribution.
[0127] In some alternative embodiments, the reverse denoising module construction unit includes: A third derivation sub-unit, configured to construct a learning model to learn the first conditional probability of the exponential power distribution, and calculate the second conditional probability based on the first conditional probability and the sampling data points at the starting time using Bayes' formula.
[0128] A fourth derivation sub-unit, configured to obtain a preset number of exponential power distributions obtained in the forward diffusion, and substitute the preset number of exponential power distributions into Bayes' formula to obtain a deformed form of the second conditional probability.
[0129] A fifth derivation sub-unit, configured to obtain the analytical expressions of the corresponding mean and variance based on the deformed form of the second conditional probability based on the standard exponential power distribution, and obtain the expression of the sampling data points at the starting time based on the derivation process of the forward diffusion process.
[0130] A reverse denoising module construction sub-unit, configured to re-substitute the expression of the sampling data points at the starting time into the analytical expressions of the mean and variance to obtain a reverse denoising module based on the exponential power distribution.
[0131] The further functional descriptions of the above-mentioned modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0132] The epileptic EEG data augmentation device based on the exponential power diffusion model in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0133] An embodiment of the present invention also provides a computer device having the above-mentioned Figure 10 epileptic EEG data augmentation device based on the exponential power diffusion model as shown.
[0134] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As shown in Figure 11 , the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 11 In
[0135] Processor 10 can be a central processor, a network processor, or a combination thereof. Among them, processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0136] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0137] The memory 20 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0138] The memory 20 may include volatile memory, such as random access memory. The memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive. The memory 20 may also include a combination of the above types of memory.
[0139] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 11 Taking connection through a bus as an example.
[0140] The input device 30 may receive input digital or character information and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0141] Embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0142] A part of the present invention can be applied as a computer program product, for example, computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present invention can be called or provided. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0143] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An epileptic EEG data augmentation method based on an exponential power diffusion model, characterized in that, The method includes: Setting the optimal shape parameter of the exponential power distribution; Constructing a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter; Performing forward diffusion on the forward diffusion module based on the exponential power distribution using the pre-acquired original EEG data to obtain noise data; Training the reverse denoising module based on the exponential power distribution using the noise data from the forward diffusion to obtain a trained exponential power diffusion model; Sampling the probability density function of the preset exponential power distribution to generate noise data that follows the exponential power distribution, and inputting the noise data that follows the exponential power distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
2. The method according to claim 1, wherein The setting of the optimal shape parameter of the exponential power distribution includes: Adaptive estimation of the optimal shape parameter of the input data based on the maximum likelihood estimation method and the quasi-Newton optimization algorithm.
3. The method according to claim 2, wherein The adaptive estimation of the optimal shape parameter of the input data based on the maximum likelihood estimation method and the quasi-Newton optimization algorithm includes: Setting the probability density function of the exponential power distribution; Calculating the maximum likelihood function using the maximum likelihood estimation algorithm based on the probability density function and the pre-acquired original EEG data, and calculating the log-likelihood function corresponding to the maximum likelihood function; Equating the log-likelihood function to an unconstrained optimization problem, and solving the unconstrained optimization problem using the quasi-Newton optimization algorithm to obtain the optimal shape parameter of the input data.
4. The method according to claim 1, wherein Constructing a forward diffusion module based on the exponential power distribution based on the optimal shape parameter includes: Specifying the Markov chain of the forward process, setting the time step of the forward diffusion, and constructing an input data sample based on the pre-acquired original EEG data; Based on the Markov chain, generating a constant sequence corresponding to the time step for controlling the mean and variance of the exponential power distribution within the time step using a scheduling strategy; Constructing a forward diffusion process with a preset number of steps, and gradually adding exponential power noise controlled by the constant sequence to the input data sample in each step based on the optimal shape parameter, so that the input data sample is converted into a sample sequence that follows the exponential power distribution after the above time step, obtaining a forward diffusion module based on the exponential power distribution.
5. The method according to claim 4, wherein The constructing of a forward diffusion process with a preset number of steps, and gradually adding exponential power noise controlled by the constant sequence to the input data sample in each step based on the optimal shape parameter, so that the input data sample is converted into a sample sequence that follows the exponential power distribution after the above time step, obtaining a forward diffusion module based on the exponential power distribution includes: Introducing the reparameterization technique into the sample sequence obtained by gradually adding exponential power noise in the forward diffusion process, and reconstructing the sample sequence that follows the exponential power distribution obtained after the forward diffusion into a form that conforms to the gamma distribution and the uniform distribution; Converting the exponential power distribution into a new representation form after linear scaling and translation of the noise sampled from the uniform distribution. Calculate the mean and variance of the specified sub-items in the re-representation form, and calculate the relationship from the starting moment to any moment in the forward diffusion process based on the Lyapunov central limit theorem, and use the relationship as the forward diffusion module based on the exponential power distribution.
6. The method according to claim 1, characterized in that, Construct a reverse denoising module based on the exponential power distribution based on the optimal shape parameter, including: Input the pre-acquired original EEG data into the forward diffusion module based on the exponential power distribution for forward diffusion to obtain noise data; Construct a U-Net network for encoding time step information based on the forward diffusion module based on the exponential power distribution. Input the noise data into the U-Net network, learn the distribution pattern of the noise data through multi-layer convolution and downsampling and upsampling structures, and add residual connections and attention mechanisms to the U-Net network. At the same time, integrate the time step information into each layer of the U-Net network through the time embedding layer to obtain the output of the U-Net network; Construct a reverse denoising process with a preset number of steps based on the output of the U-Net network, and continuously remove noise from the random noise according to the exponential power distribution at each time step based on the optimal shape parameter to obtain the reverse denoising module based on the exponential power distribution.
7. The method according to claim 6, wherein The constructing a reverse denoising process with a preset number of steps based on the output of the U-Net network, and continuously removing noise from the random noise according to the exponential power distribution at each time step based on the optimal shape parameter to obtain the reverse denoising module based on the exponential power distribution includes: Construct a learning model to learn the first conditional probability of the exponential power distribution, and calculate the second conditional probability based on the first conditional probability and the sampling data points at the starting moment using Bayes' formula; Obtain a preset number of exponential power distributions obtained in the forward diffusion, and substitute the preset number of exponential power distributions into Bayes' formula to obtain a deformed form of the second conditional probability; Derive the analytical expressions of the corresponding mean and variance based on the deformed form of the second conditional probability and the standard exponential power distribution, and obtain the expression of the sampling data points at the starting moment based on the derivation process of the forward diffusion process; Substitute the expression of the sampling data points at the starting moment back into the analytical expressions of the mean and variance to obtain the reverse denoising module based on the exponential power distribution.
8. An epileptic EEG data augmentation device based on an exponential power diffusion model, characterized in that, The device includes: An optimal shape parameter setting module for setting the optimal shape parameter of the exponential power distribution; An exponential power diffusion model construction module for constructing a forward diffusion module based on the exponential power distribution and a reverse denoising module based on the exponential power distribution based on the optimal shape parameter; A forward diffusion module for performing forward diffusion on the forward diffusion module based on the exponential power distribution based on the pre-acquired original EEG data to obtain noise data; A reverse denoising module for training the reverse denoising module based on the exponential power distribution with the noise data of the forward diffusion process to obtain a trained exponential power diffusion model; The reconstruction and expansion module is used to sample the probability density function of the preset exponential power distribution to obtain noise data obeying the exponential power distribution, and input the noise data obeying the exponential power distribution into the reverse denoising module of the exponential power diffusion model to gradually reconstruct and generate new EEG data.
9. A computer device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the epileptic EEG data expansion method based on the exponential power diffusion model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the epileptic EEG data expansion method based on the exponential power diffusion model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Generation method and system for simulating and collecting electroencephalogram signals
CN115238745A
Unsupervised epilepsy detection system based on de-noising diffusion probability model
CN117338314A
Epileptic seizure prediction method, device and equipment based on complementary frequency conversion perception model
CN118078211A
Method and system for training electroencephalogram signal enhancement model based on diffusion model
CN118606707A
Epilepsy electroencephalogram signal detection system and method based on distraction attention model
CN119279521A