Type 2 diabetes data processing method based on photoelectric volume pulse wave signals
Through the improved CWGAN-GP model generation combined with the self-attention mechanism, the problem of scarcity and imbalance of PPG signal data is solved, high-quality samples are generated, the data set is expanded, and the accuracy and model stability of type 2 diabetes prediction are improved.
Patent Information
- Application Number
- CN202510508513.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
In the intelligent diagnosis of Chinese medicine pulse diagnosis, deep learning algorithms are limited by the scarcity of PPG signal data and the imbalance of data sets, resulting in poor model training results.
The improved conditional Wasserstein generative adversarial network (CWGAN-GP) model is adopted, and the gated recurrent unit (GRU) network and self-attention mechanism is combined. Adversarial training is performed through the gradient punishment mechanism, PPG signal samples of the specified categories are generated, and the balanced data set is constructed mixed with the real samples. The multi-classifier fusion framework is used for prediction.
The data set size has been significantly expanded, the category distribution is balanced, the authenticity and diversity of generated samples have been improved, the training efficiency and generalization ability of the model have been improved, and the accuracy of type 2 diabetes prediction has been enhanced.
Smart Images

Figure CN120372552A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical health monitoring, and particularly relates to a method for processing type 2 diabetes data based on photoplethysmogram (PPG). Background Art
[0002] In recent years, deep learning algorithms have shown excellent performance in the intelligent diagnosis of traditional Chinese medicine pulse diagnosis. However, their demand for large-scale data limits their practical applications. The acquisition of pulse diagnosis data is not only time-consuming and laborious but also restricted by privacy protection regulations, resulting in the available sample size usually being only a few hundred cases. This problem has caused data scarcity and dataset imbalance in the diagnosis of diabetes using PPG signals, thus limiting the training effect and performance improvement of deep learning models. To address this challenge, data augmentation techniques have become an important solution. Generative Adversarial Networks (GANs) are recognized as an efficient data augmentation technique. This technique has not only proven its excellent performance in the fields of image and speech synthesis but has also been further extended to the field of physiological signal generation. However, due to the complex temporal characteristics and important clinical relevance of PPG signals, traditional GANs are difficult to accurately capture their features, resulting in significant differences between the generated samples and the real data. Summary of the Invention
[0003] To solve the above problems, the present invention proposes a method for processing type 2 diabetes data based on photoplethysmogram signals, including the following steps:
[0004] S1. Collect the original photoplethysmogram (PPG) signal data and perform normalization and denoising preprocessing on the data;
[0005] S2. Construct a Conditional Wasserstein Generative Adversarial Network (CWGAN-GP) model, where the generator adopts a dual-channel structure of a gated recurrent unit (GRU) network and a self-attention mechanism, and the discriminator adopts a multi-layer convolutional neural network (CNN);
[0006] S3. Input the diabetes label as a conditional variable into the generator and perform adversarial training by introducing a gradient penalty mechanism;
[0007] S4. Use the trained generator to synthesize PPG signal samples of a specified category;
[0008] S5. Mix the generated samples with the real samples to construct a dataset with a balanced class distribution;
[0009] S6. Adopt a multi-classifier fusion framework to perform type 2 diabetes prediction analysis on the augmented data.
[0010] Further, the generator structure in step S2 includes:
[0011] The input layer receives the concatenated vector of the noise vector and the class label;
[0012] The GRU time-series feature extraction module consists of a 3-layer gated recurrent network with 1024 hidden units;
[0013] The self-attention optimization module calculates the feature weight distribution based on the attention mechanism;
[0014] The Dropout layer prevents overfitting;
[0015] The fully-connected output layer maps the 128-dimensional feature vector to 2100-dimensional PPG waveform data.
[0016] Furthermore, the discriminator in step S2 adopts a 6-layer deep convolutional architecture, with a 1×5 convolutional kernel of 32, 64, 128, 256, and 512 channels configured in each layer in sequence. The LeakyReLU activation function and max pooling for dimensionality reduction are used between layers, and the output end is connected to a fully-connected layer and a class label embedding layer.
[0017] Furthermore, the gradient penalty mechanism in step S3 is specifically as follows:
[0018] The Wasserstein distance is used as the distance metric between the generator and the discriminator to replace the cross-entropy loss function in GAN. The formula for the Wasserstein distance is:
[0019] ; Among them, is the real sample, is the conditional vector of the noise vector, represents the real distribution, represents the generated distribution, represents the expected value, represents and the set of all joint distributions combined by At this time, the loss function is: ; is any random input of the generator, is the output of the generator from any random input, is the result obtained by discriminating the real sample using the Wasserstein distance, is the result obtained by discriminating the fake sample using the Wasserstein distance; The gradient penalty is used to replace the gradient clipping, and the replaced loss function is: ; ; Among them, is the penalty term, is and random interpolation sampling on the connection line, is distribution, is a random number in [0, 1].
[0020] Furthermore, in the multi-classifier fusion framework described in step S6, the classifiers include a convolutional neural network (CNN), a long short-term memory network (LSTM), a two-layer LSTM, and a random convolutional kernel transform (ROCKET), and the output results of each classifier are fused by weighted voting or stacking ensemble methods.
[0021] Furthermore, the similarity between the generated samples and the real samples is evaluated by root mean square error (RMSE), maximum mean discrepancy (MMD), and dynamic time warping (DTW) metrics to ensure the authenticity and diversity of the synthetic data.
[0022] The present invention also provides a type 2 diabetes data processing system based on photoplethysmogram signals, which is characterized by including:
[0023] A PPG signal acquisition module for acquiring original PPG signal data;
[0024] A data preprocessing module for normalizing and denoising the original data;
[0025] A data augmentation module that deploys an improved CWGAN-GP model to generate and expand the PPG data set;
[0026] An intelligent diagnosis module that integrates CNN, LSTM, and ROCKET multi-classifiers to perform type 2 diabetes prediction and analysis on the augmented data;
[0027] A performance evaluation module for multi-dimensional evaluation of the generated data and classification results by RMSE, MMD, DTW, accuracy, recall, precision, and F1 score.
[0028] Furthermore, the generator of the data augmentation module adopts a structure combining a GRU network and a self-attention mechanism, the discriminator adopts a multi-layer convolutional neural network structure, and class label embedding is introduced.
[0029] Advantages of the present invention:
[0030] In view of the problems of limited number of PPG signal data samples and extremely unbalanced class distribution in type 2 diabetes prediction, this invention proposes an improved conditional Wasserstein generative adversarial network with gradient penalty (CWGAN-GP) algorithm. Through high-quality synthetic sample generation and data augmentation, the scale of the dataset is significantly expanded, and the balance of class distribution is achieved, providing a more sufficient and diverse data basis for subsequent model training.
[0031] This invention introduces a gated recurrent unit (GRU) network and a self-attention mechanism into the generator, which can better capture the complex temporal characteristics and global dependencies of PPG signals. The generated synthetic data is highly consistent with the real data in terms of amplitude, phase, frequency domain distribution, and time-frequency characteristics, greatly improving the authenticity and diversity of synthetic samples.
[0032] By introducing a gradient penalty mechanism, this invention effectively constrains the Lipschitz continuity of the discriminator, prevents training instability problems such as gradient vanishing or explosion, and improves the training efficiency and stability of the generative adversarial network. At the same time, the model after data augmentation shows stronger generalization ability when facing new samples. Brief Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0034] Figure 1 Structure diagram of the improved CWGAN-GP model;
[0035] Figure 2 Comparison of the generated pulse waveforms by the improved CWGAN-GP;
[0036] Figure 3 Comparison of the generated pulse spectra by the improved CWGAN-GP;
[0037] Figure 4 Time-frequency diagrams of the original pulse (left) and the pulse generated by the improved CWGAN-GP (right);
[0038] Figure 5 Comparison of the distributions before and after data augmentation. Detailed Embodiments
[0039] The following will describe the present application in conjunction with specific embodiments:
[0040] Embodiment 1:
[0041] This embodiment provides a method for processing type 2 diabetes data based on photoplethysmogram (PPG) signals, including the following steps:
[0042] S1. Collect the original photoplethysmogram (PPG) signal data, and perform normalization and denoising preprocessing on the data;
[0043] S2. Construct a conditional Wasserstein generative adversarial network (CWGAN-GP) model, whose generator adopts a dual-channel structure of a gated recurrent unit (GRU) network and a self-attention mechanism, and the discriminator adopts a multi-layer convolutional neural network (CNN);
[0044] S3. Input the diabetes label as a conditional variable into the generator, and perform adversarial training by introducing a gradient penalty mechanism;
[0045] The generator structure includes: an input layer that receives the concatenated vector of the noise vector and the class label; a GRU temporal feature extraction module, which is a 3-layer gated recurrent network with 1024 hidden units; a self-attention optimization module that calculates the feature weight distribution based on the attention mechanism; a Dropout layer to prevent overfitting; and a fully connected output layer that maps the 128-dimensional feature vector to 2100-dimensional PPG waveform data. The discriminator adopts a 6-layer deep convolutional architecture, with 1×5 convolutional kernels with 32, 64, 128, 256, and 512 channels configured in each layer in turn. The LeakyReLU activation function and max pooling for dimensionality reduction are used between layers, and the output end is connected to a fully connected layer and a class label embedding layer.
[0046] GAN is a deep learning model, which consists of two parts: a generator and a discriminator. During the training process of the generative adversarial network, a fierce competition unfolds between the generator and the discriminator, similar to a duel between two chess players on a chessboard. The generator is committed to learning how to generate increasingly realistic data samples in order to deceive the discriminator; while the discriminator is constantly improving its recognition ability, striving to distinguish between real data and the fake data created by the generator. Mathematically, GAN can be described as a min-max problem. During training, the generator attempts to minimize the loss function, and the discriminator attempts to maximize the loss function. Its loss function is shown in formula (1):
[0047]
[0048] In the above formula, is the real sample, is any random input of the generator, is the output of the generator from any random input. represents the real data distribution, represents the generator data distribution. represents the expected value, Represents probability.
[0049] To enable the GAN to generate samples with specified features, one solution is the CGAN, which adds an additional conditional vector to the input noise vector , which can more precisely control the output generated by the GAN. The loss function is as shown in Equation (2):
[0050]
[0051] where is the conditional vector.
[0052] Since traditional GAN uses the JS divergence as the value to measure the difference between two samples, when the real data distribution and the generated data distribution have no overlap or the overlap degree is small, the loss function of the GAN will approach a constant. The generator and discriminator cannot generate effective gradients, which may lead to the problem of gradient disappearance.
[0053] This embodiment adopts the Wasserstein GAN method. By introducing the Wasserstein distance as the training objective, it effectively overcomes the defect of the JS divergence in traditional GAN training. Even when the overlap between the real data distribution and the generated data distribution is small, the Wasserstein distance can still provide meaningful gradient information, thus optimizing the training process. Its main idea is to use the Wasserstein distance as the distance metric between the generator and the discriminator to replace the cross-entropy loss function in GAN. First, the Wasserstein distance is shown as follows:
[0054]
[0055] where, represents the real distribution, represents the generated distribution, represents and the set of all joint distributions obtained by combining .
[0056] At this time, the formula of the loss function is as follows:
[0057]
[0058] is the result obtained by discriminating the real samples using the Wasserstein distance, is the result obtained by discriminating the fake samples using the Wasserstein distance.
[0059] WGAN-GP overcomes the problem of unstable training in WGAN, replaces gradient clipping with gradient penalty, and its loss function is:
[0060]
[0061]
[0062] where, is the penalty term, is and random interpolation sampling on the line connecting them, is 's distribution, is a random number in [0, 1].
[0063] The self-attention mechanism is a type of attention mechanism. The attention mechanism learns the important information in the input information and calculates the weights of these important information and all the input information. The self-attention mechanism calculates a weight parameter for each element of the input, pays more attention to the parts similar to the input element according to the weight parameter. The calculation process remains the same, but the calculation object has changed. It can calculate the relationship between each element in the input sequence and other elements and use these relationships to better represent the input sequence. The self-attention mechanism allows the model to dynamically allocate different attention weights at each time step to focus on the part of the sequence that is most relevant to the current processing unit.
[0064] For the input sequence , through three weight matrices , and are linearly transformed into the query vector Q, the key vector K, and the value vector V respectively, as shown in formulas (7)-(9):
[0065]
[0066]
[0067]
[0068] For the three inputs, first calculate the similarity between Q and K through the scoring function, usually obtaining the attention score by dot product. Then, scale the attention score and normalize it through Softmax to obtain the attention weights. Finally, perform weighted summation on V according to the attention weights to obtain the final output of self-attention. The formula is as follows:
[0069]
[0070] In the formula, The dimension of K, which is usually the same as the dimensions of the query and value vectors.
[0071] S4. Use the trained generator to synthesize PPG signal samples of a specified category.
[0072] S5. Mix the generated samples with the real samples to construct a dataset with balanced class distribution. The similarity between the generated samples and the real samples is evaluated by metrics such as Root Mean Square Error (RMSE), Maximum Mean Discrepancy (MMD), and Dynamic Time Warping (DTW) to ensure the authenticity and diversity of the synthetic data.
[0073] S6. Adopt a multi-classifier fusion framework to perform type 2 diabetes prediction analysis on the enhanced data. Here, the classifiers include Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), Double-Layer LSTM, and Random Convolutional Kernel Transform (ROCKET), and the output results of each classifier are fused by weighted voting or stacking ensemble methods.
[0074] Example 2:
[0075] This example provides a type 2 diabetes data processing system based on photoplethysmogram (PPG) signals, including:
[0076] A PPG signal acquisition module for collecting raw PPG signal data.
[0077] A data preprocessing module for normalizing and denoising the raw data.
[0078] A data augmentation module that deploys an improved CWGAN-GP model to generate and expand the PPG dataset; in the CWGAN-GP model, the generator adopts a structure that combines a GRU network and a self-attention mechanism, the discriminator adopts a multi-layer convolutional neural network structure, and class label embedding is introduced.
[0079] An intelligent diagnosis module that integrates CNN, LSTM, and ROCKET multi-classifiers to perform type 2 diabetes prediction analysis on the enhanced data.
[0080] A performance evaluation module for multi-dimensional evaluation of the generated data and classification results in terms of RMSE, MMD, DTW, accuracy, recall, precision, and F1 score.
[0081] Example 3:
[0082] For the type 2 diabetes data processing method and system provided in Examples 1 and 2, this example gives a specific implementation method as follows:
[0083] The structure of the improved CWGAN-GP involved in this example is as Figure 1As shown, as the design core of the model, the generator of CWGAN-GP consists of two parts: the self-attention mechanism and the GRU network. The discriminator uses CNN to ensure that it can effectively extract local features during the process of discriminating between generated signals and real signals, thus providing more powerful feedback to the generator.
[0084] The generator provided in this embodiment is designed to generate simulated data with the same dimension as the real pulse wave data. The input consists of a random noise vector and a class label, and data generation is achieved through a multi-layer network. First, the random noise vector and the class label are concatenated to form an input tensor, and then the time series features are extracted through the GRU network. The hidden layer size of the GRU network is set to 1024, and the parameters are initialized using a normal distribution with a mean of 0 and a standard deviation of 0.01 to ensure the stability of training. Subsequently, the self-attention mechanism is used to optimize the output features of the GRU. The self-attention module generates query, key, and value vectors through linear transformation, calculates the attention scores, and uses the Softmax function to determine the weights. These weights are used to perform weighted summation on the value vectors to form a context feature vector, thereby capturing global dependencies. In addition, the feature vector passes through a Dropout layer, and some neurons are randomly deactivated with a probability of 0.5 to prevent overfitting and ensure the generalization ability of the model. Finally, the feature vector is mapped to the target dimension (i.e., the length of the pulse wave data 2100) through a fully connected layer to generate an output tensor consistent with the real data.
[0085] The discriminator consists of CNN and contains six convolutional layers. Each layer uses 5 convolutional kernels of size 1, with a stride of 1 and a padding of 2. The number of channels increases layer by layer from 32 to 512, gradually enhancing the feature expression ability of the model. After each convolutional layer, a LeakyReLU activation function is connected to introduce non-linearity to improve the model's learning ability for complex features and alleviate the vanishing gradient problem. Then, further dimensionality reduction is performed through a max-pooling layer with a window size and stride of 2 to extract important features and reduce the computational complexity. After the convolutional and pooling operations, the feature map is flattened and passed to the fully connected layer. The fully connected layer maps the high-dimensional features to an output value, which is used to represent the probability that the input sample is real data. In addition, the model introduces an embedding layer to map the class label to a vector of dimension 1 and concatenate it with the pulse wave signal in the feature dimension, enabling the discriminator to use both the pulse wave data and the class information for discrimination. This embodiment provides an optimized set of hyperparameter configurations, as shown in Table 1.
[0086] Table 1 Model Parameter Settings
[0087]
[0088] For the data generated by the above model, performance evaluation is required. When evaluating the performance of time series data, a series of quantitative indicators are usually relied on to measure the accuracy and similarity of the prediction model. In this paper, the DTW value, RMSE, and MMD will be used as indicators to measure the sequence similarity:
[0089] Root Mean Square Error (RMSE) is an indicator to measure the overall error between the generated signal and the original signal in the time domain. RMSE measures the magnitude of the error by calculating the average of the sum of the squares of the prediction errors and then taking the square root. The lower the RMSE value, the smaller the difference between the generated signal and the original signal. The RMSE is defined as shown in the following formula:
[0090]
[0091] where is the value of the original signal, is the value of the generated signal, is the total number of data points.
[0092] Maximum Mean Discrepancy (MMD) is a statistical measure based on kernel methods for evaluating the difference between two time series distributions. MMD measures their similarity by comparing the mean differences between the two distributions. The lower the MMD value, the closer the two distributions are, that is, the more similar the distribution of the predicted sequence generated by the model is to the actual sequence. The calculation formula of MMD is as follows:
[0093]
[0094] where is the Gaussian kernel function.
[0095] Dynamic Time Warping (DTW) is an algorithm for measuring the shape similarity of two time series. It can capture the non-linear relationship between time series and the changes on the time axis. DTW finds an optimal matching method by dynamically adjusting the correspondence between points in the time series, thereby minimizing the distance between the two sequences. The smaller the value of DTW, the more similar the two time series are in shape. The calculation of DTW involves constructing a cost matrix and finding a path that minimizes the cumulative distance between the sequences. The specific formula is as follows:
[0096]
[0097] where and are two time series, is the distance metric, is an alignment function that allows for changes along the timeline.
[0098] To verify the effectiveness of the above method, this study used a publicly available dataset from a certain hospital as the source of pulse wave data. This dataset included 219 adult subjects aged between 21 and 86 years old. Each signal segment had a duration of 2.1 seconds and was recorded at a sampling frequency of 1KHz. In addition, this dataset also detailed the basic physiological characteristics, disease status, blood pressure values, heart rate data, and three waveform materials of each subject.
[0099] As Figure 2 shown, the generated pulse signal showed good synchronization with the original signal along the timeline, indicating that the improved generation algorithm can effectively capture and simulate the periodic characteristics of the original pulse signal. Further observing the amplitude changes, it was found that the two curves maintained a high degree of consistency in the overall trend. The generated signal not only resembled the original signal in amplitude but also showed good correspondence in the occurrence time points of the peaks and valleys, further confirming the simulation effect of the generated signal in amplitude and phase.
[0100] As Figure 3 shown, in the frequency range of 0 to 20 Hz, the amplitude curves of the pulse data generated by CWGAN-GP and the real pulse data were very close, which is also the main frequency range of the pulse signal.
[0101] As Figure 4 shown, the main frequency band energy distributions (intensity of time-frequency) were highly consistent, indicating that the generation model successfully learned the core characteristics of the pulse signal. The trends of energy changes updated over time were similar, and the generated signal simulated the real signal well in terms of time resolution. It can be seen from the time-frequency diagram that the pulse data generated by CWGAN-GP achieved very similar time-frequency characteristics to the original pulse data. It performed well in terms of frequency domain distribution, consistency, and noise suppression, indicating that the model has strong capabilities in generating high-quality pulse signals.
[0102] In addition to the icon comparison, this embodiment also verified the effectiveness of this method through quantitative analysis. To comprehensively evaluate the performance of the proposed method, Table 2 summarizes the RMSE, RMSE, and DTW scores calculated between the real data and the samples generated by different data augmentation methods:
[0103] Table 2 Comparison of evaluation indicators for generated signals
[0104]
[0105] It can be seen from the data in the table that although the traditional SMOTE method is simple and easy to use, its performance in the three indicators is not ideal. Compared with SMOTE, CGAN has certain improvements, but there are still limitations in the distribution consistency of the generated signals and the preservation of time series features. On this basis, CWGAN significantly reduces MMD and DTW by introducing the Wasserstein distance, and both the distribution consistency of the generated signals and the time series features are improved. Further, CWAN-GP effectively solves the limitation of weight clipping on the basis of CWGAN by introducing a gradient penalty mechanism, and all three indicators are significantly optimized. However, the method in this paper is further optimized on the basis of CWAN-GP. The structural design and algorithm optimization of this method make it reach the current optimal level in terms of the amplitude accuracy, distribution consistency, and time series feature fidelity of pulse signals.
[0106] Facing the problem of serious imbalance in the number of samples of diabetic patients and normal people in the PPG dataset, where there are only 27 samples of diabetic patients and as many as 177 samples of normal people. To improve the generalization ability and prediction accuracy of the model, the data distribution after expanding the dataset using the improved CWAN-GP is as Figure 5 shown.
[0107] To verify the effectiveness and superiority of the proposed algorithm, various data augmentation methods are used to expand the PPG dataset and apply it to the diabetes prediction task. Specifically, four classifiers are used in the experiment, including CNN, Long Short-Term Memory (LSTM), double-layer LSTM, and Random Convolutional Kernel Transform (ROCKET), to classify the expanded dataset. To comprehensively evaluate the performance of the classifiers, accuracy, recall, precision, and F1 score are selected as performance indicators, and the performance of the datasets generated by different data augmentation methods in the classification task is compared. In the experiment, the training set and the test set are divided at a ratio of 80% and 20% to ensure the rationality of training and validation. The performance results of each classifier on the datasets generated by different data augmentation methods according to the above three performance indicators are shown in Table 3.
[0108] Table 3 Prediction performance results of different methods
[0109]
[0110] Without sampling, the overall performance of the classifier is weak. In particular, LSTM, double-layer LSTM, and ROCKET perform poorly, and only CNN is slightly better. This indicates that it is difficult to achieve effective improvement in classification performance by directly using the original data under the problem of class imbalance. After introducing SMOTE, the classification performance is significantly improved. The F1 score of CNN increases from 0.6 to 0.8564, and that of LSTM increases from 0.4286 to 0.8246. However, the improvement of double-layer LSTM and ROCKET is relatively limited, probably because the distribution difference between the SMOTE synthetic samples and the real data is large, resulting in model overfitting. CGAN further improves the classification performance, especially for double-layer LSTM and ROCKET, which indicates that CGAN can better solve the problem of data distribution deviation compared with SMOTE. However, the improvement of CGAN for CNN is limited. CWGAN overcomes the problems of instability and mode collapse in the training process of the CGAN model by introducing the WGAN mechanism, and the classification effect is enhanced. On this basis, CWGAN-GP further optimizes the classification performance by introducing gradient penalty, indicating that gradient penalty effectively improves the stability of the generator and the quality of samples.
[0111] By comparing the methods of without sampling, SMOTE, CGAN, CWGAN, and CWAN-GP, it can be seen that the method in this paper is more effective in dealing with the problem of class imbalance. While optimizing the accuracy, it maintains a high recall rate. Its core advantage lies in that for long time series data, it can generate high-quality samples by capturing global time dependencies. These samples not only have stronger time consistency but also are more balanced in class distribution, thus greatly reducing the negative impact of class imbalance on classification performance. From the experimental data, the accuracy of all classifiers under the method in this paper is higher than that of other methods. For example, the accuracy of the CNN classifier reaches 89.50%, which is significantly better than other methods. And for other classifiers such as LSTM, double-layer LSTM, and ROCKET, obvious performance improvements are also achieved, fully verifying the effectiveness of the method in this paper in practical applications. Thanks to the input of high-quality samples, the performance of the classifier on the imbalanced dataset is more stable, and performance improvements are achieved under various architectures. This indicates that the combination of GRU and the self-attention mechanism can more effectively extract global and local features in time series, significantly improve the quality of generated samples, and thus provide more reliable support for classification tasks.
[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0113] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention (such as quantity, shape, position, etc.), and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A method for processing type 2 diabetes data based on photoplethysmogram signals, characterized in that, It includes the following steps: S1. Collect the original photoplethysmogram (PPG) signal data, and perform normalization and denoising preprocessing on the data; S2. Construct a conditional Wasserstein generative adversarial network (CWGAN-GP) model. Its generator adopts a dual-channel structure of a gated recurrent unit (GRU) network and a self-attention mechanism, and the discriminator adopts a multi-layer convolutional neural network (CNN); S3. Input the diabetes label as a conditional variable into the generator, and perform adversarial training by introducing a gradient penalty mechanism; S4. Use the trained generator to synthesize PPG signal samples of a specified category; S5. Mix the generated samples with the real samples to construct a dataset with a balanced class distribution; S6. Adopt a multi-classifier fusion framework to perform type 2 diabetes prediction analysis on the enhanced data.
2. The method for processing type 2 diabetes data based on photoplethysmogram signals according to claim 1, wherein The generator structure described in step S2 includes: An input layer that receives the concatenated vector of the noise vector and the class label; A GRU time series feature extraction module, which is a 3-layer gated recurrent network with 1024 hidden units; A self-attention optimization module that calculates the feature weight distribution based on the attention mechanism; A Dropout layer to prevent overfitting; A fully connected output layer that maps the 128-dimensional feature vector to 2100-dimensional PPG waveform data.
3. The method for processing type 2 diabetes data based on photoplethysmogram signals according to claim 1, wherein: The discriminator described in step S2 adopts a 6-layer deep convolutional architecture. Each layer is sequentially configured with 1×5 convolutional kernels with 32, 64, 128, 256, and 512 channels. The LeakyReLU activation function and max pooling for dimensionality reduction are used between layers, and the output end is connected to a fully connected layer and a class label embedding layer.
4. The method for processing type 2 diabetes data based on photoplethysmogram signals according to claim 1, wherein The gradient penalty mechanism in step S3 is specifically: Use the Wasserstein distance as the distance metric between the generator and the discriminator to replace the cross-entropy loss function in GAN. The Wasserstein distance formula is: ; Among them, is the real sample, is the conditional vector of the noise vector, represents the real distribution, represents the generated distribution, represents the expected value, represents and all the joint distributions obtained by combination of the set, and the loss function at this time is: ; is any random input to the generator, is the output of the generator from any random input, is the result obtained by discriminating real samples using the Wasserstein distance, is the result obtained by discriminating fake samples using the Wasserstein distance; Use gradient penalty to replace gradient clipping. The replaced loss function is: ; ; Among them, is the penalty term, is and random interpolation sampling on the connection line of is distribution of is a random number in [0, 1].
5. The method for processing type 2 diabetes data based on photoplethysmogram signals according to claim 1, wherein: In the multi-classifier fusion framework described in step S6, the classifiers include a convolutional neural network (CNN), a long short-term memory network (LSTM), a double-layer LSTM, and a random convolutional kernel transform (ROCKET), and the output results of each classifier are fused by means of weighted voting or stacking ensemble.
6. The method for processing type 2 diabetes data based on photoplethysmogram signals according to any one of claims 1-5, characterized in that: The similarity between the generated samples and the real samples is evaluated by indicators such as root mean square error (RMSE), maximum mean discrepancy (MMD), and dynamic time warping (DTW) to ensure the authenticity and diversity of the synthetic data.
7. A type 2 diabetes data processing system based on photoplethysmogram signals, characterized in that, It includes: A PPG signal acquisition module for collecting the original PPG signal data; A data preprocessing module for normalizing and denoising the original data; A data augmentation module that deploys an improved CWGAN-GP model to generate and expand the PPG dataset; An intelligent diagnosis module that integrates CNN, LSTM, and ROCKET multi-classifiers to perform type 2 diabetes prediction analysis on the enhanced data; A performance evaluation module for multi-dimensional evaluation of the generated data and classification results in terms of RMSE, MMD, DTW, accuracy, recall, precision, and F1 score.
8. The system according to claim 7, characterized in that: The generator of the data augmentation module adopts a structure that combines a GRU network and a self-attention mechanism. The discriminator adopts a multi-layer convolutional neural network structure and introduces class label embedding.