Health data enhancement and anomaly detection method based on GAN (Generative Adversarial Network)

By generating synthetic data that conforms to medical priors through the generation of adversarial networks, combining multimodal feature fusion and privacy protection, the problems of scarcity and high false alarm rates of medical data are solved, high-precision abnormality detection and cross-organization collaboration are achieved, and it is suitable for real-time deployment of edge devices.

CN120277590AActive Publication Date: 2025-07-08ZHEJIANG YISHAN SMART MEDICAL RES CO LTD

Patent Information

Application Number
CN202510766324.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Medical data is scarce, heterogeneity and high false alarm rates. The images generated by traditional data augmentation technology do not conform to anatomical rationality, and privacy protection is difficult to achieve cross-organizational collaboration.

Method used

A health data augmentation system based on a generative adversarial network is adopted, including a generator module, a discriminator module, a federated learning protocol module and a physiological signal generation module. Combined with anatomical constraints and privacy protection, synthetic data that conforms to medical priors is generated, and abnormal detection is carried out through a hierarchical attention mechanism.

Benefits of technology

The generated synthetic data complies with medical priors, realizes high-precision abnormality detection, reduces false alarm rates, supports cross-hospital collaborative training, has a high success rate of privacy protection, and is suitable for real-time deployment of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_75
    Figure SMS_75
  • Figure SMS_77
    Figure SMS_77
  • Figure SMS_141
    Figure SMS_141
Patent Text Reader

Abstract

The invention discloses a GAN (Generative Adversarial Network)-based health data enhancement and anomaly detection method, which realizes data system enhancement through arrangement of a generator module, a discriminator module, a federal learning protocol module and a physiological signal generation module and adjustment of multi-modal features, hierarchical attention and threshold values, and realizes health data enhancement and anomaly detection through training of the system. The problems of scarcity, heterogeneity and high false alarm rate of medical data are solved, and improvement of medical diagnosis accuracy is facilitated; the medical level is improved, and benefits are brought to patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for health data augmentation and anomaly detection based on the Generative Adversarial Network (GAN). Background Art

[0002] The rapid development of medical artificial intelligence (AI) relies on high-quality and diverse medical data. However, the acquisition of medical data faces multiple challenges. According to the World Health Organization's 2023 report, the global annual volume of medical data generated exceeds 2.5 zettabytes, but only 0.7% of the data is effectively annotated and used for AI model training. In the field of rare diseases, such as amyotrophic lateral sclerosis, the average number of effective samples accumulated by a single medical institution is less than 50 cases. In addition, strict privacy protection regulations for medical data severely restrict the sharing and circulation of data, resulting in the widespread phenomenon of "data silos". Traditional data augmentation techniques (such as rotation, translation, noise injection) can partially alleviate the problem of data insufficiency, but the generated images often do not conform to anatomical rationality, such as the displacement of the heart position or the breakage of blood vessel structures. Therefore, there is a need to develop an emerging method for health data augmentation and anomaly detection. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a data augmentation system that supports privacy protection and anatomical constraints, as well as a real-time anomaly detection method for multi-modal fusion, to solve the problems of scarce, heterogeneous medical data and high false alarm rates, and a method for health data augmentation and anomaly detection based on the Generative Adversarial Network (GAN) with practicality and wide application.

[0004] To solve the above problems, the present invention adopts the following technical solutions: A method for health data augmentation based on the Generative Adversarial Network, which is implemented by a health data augmentation system based on the Generative Adversarial Network. The system includes the following components: a generator module, a discriminator module, a federated learning protocol module, and a physiological signal generation module; The generator module: includes a parallel image generation branch and a U-Net segmentation branch; The image generation branch consists of a 5-layer transposed convolution network, and each layer is in turn: a transposed convolution layer, a batch normalization layer, and a LeakyReLU activation function, and outputs a synthetic medical image with a resolution of 512×512; The U-Net segmentation branch includes 4 layers of encoders and 4 layers of decoders, and outputs a segmentation mask of the target organ; the loss function of the segmentation branch is the Dice coefficient loss, with a weight ratio of 30%; The federated learning protocol module supports cross - institutional collaborative training. When participating parties conduct local training, differential privacy stochastic gradient descent is adopted, with the gradient clipping threshold C = 0.5 and the Gaussian noise scale σ = 1.2. At the same time, global model aggregation is performed every 10 rounds of communication, and the aggregation method is weighted average, with weights allocated according to the data volume of each participating party. At the same time, the privacy budget calculation satisfies ε = 0.8, meeting the HIPAA privacy protection standard. The physiological signal generation module uses a temporal - constraint WGAN - GP to generate ECG signals, and the constraint conditions include: QRS wave width of 80 - 120 ms and heart rate of 60 - 100 BPM. The time - continuity loss function of the physiological signal generation module is defined as the L2 - norm difference between adjacent signal points: , where, is the smoothing loss value, which reflects the smoothness of the entire sequence data; is the number of elements in the sequence minus 1, that is, the length of the sequence minus 1; represents the -th element in the sequence; represents the +1 - th element; is the L2 - norm for calculating the difference between two adjacent elements.

[0005] Preferably, the loss function of the generator is: ; where, is the Wasserstein distance adversarial loss, and the gradient penalty coefficient λ = 10; is the Dice coefficient loss of the segmentation branch, which constrains the positions of key organs such as the heart and lungs; is the federated aggregation loss, which weighted - averages the model parameters of each participating party; The communication data encryption of the federated learning protocol module uses a homomorphic encryption algorithm with a key length of 2048 bits.

[0006] Preferably, the physiological signal generation module further includes: a QRS wave width detection algorithm and abnormal signal filtering. The QRS wave width detection algorithm is based on dynamic time warping to match the standard QRS template, with an error tolerance of ±5 ms. Abnormal signal filtering: If the heart rate or QRS width of the generated signal exceeds the preset range, it is automatically discarded and regenerated.

[0007] Preferably, a health data enhancement method based on a generative adversarial network includes the following steps: 1) Data preparation and pre - processing; It includes data collection and cleaning, collecting medical data, removing noise, and setting a standardized data format; Including data annotation, where doctors mark key areas and perform privacy data desensitization processing; 2) Select enhancement methods, adopting traditional methods and deep generation methods; 3) Implement enhancement processing; 4) Evaluate the generated data: conduct quality evaluation and privacy protection evaluation; 5) Deploy the enhanced data.

[0008] A method for detecting anomalies in health data based on a generative adversarial network includes the following steps: 1) Perform multi-modal feature extraction; specifically: perform image feature extraction: use a pre-trained ResNet-50 model, remove the fully connected layer and replace it with an adaptive average pooling layer, and output a 256-dimensional feature vector; perform temporal signal feature extraction: a 1D-CNN network extracts local features, and sinusoidal positional encoding is superimposed to inject temporal information; perform text feature extraction: the BioBERT model tokenizes and performs entity recognition on the medical record text, and outputs a 768-dimensional clinical semantic embedding; 2) Perform hierarchical attention fusion: Specifically: The first-level attention: Using the image features as the Query, the temporal signal features as the Key and Value, calculate the attention features from the image to the signal: ; Among them, is the query vector obtained by linearly transforming the image features; is the key vector obtained by linearly transforming the signal features; is the key vector of the dimension; is to apply the Softmax function to the similarity score matrix obtained above and convert it into a probability distribution form; is the value vector obtained by linearly transforming the signal features; The second-level attention: Using the output of the first level as the Query, the text features as the Key and Value, calculate the final fusion features: ; Among them, is the attention feature from the image to the signal; is the attention feature from the fusion feature to the text; is the layer normalization operation, which normalizes the input feature vector in the dimension of the layer to make the input data of each layer have a similar distribution; 3) Perform dynamic threshold adjustment and anomaly determination, specifically: update the threshold based on exponential weighted moving average, set the smoothing factor β = 0.9, and the calculation formula is: ; Among them, is the time step The mean estimated value; is the time step The mean estimated value; is the time step The observed value; is the time step The variance estimated value; is the time step The variance estimated value; is the current observed value and the current mean estimated value The square of the difference; is the time step The standard deviation estimated value; Its anomaly determination condition: If the anomaly probability score is continuous for 3 times , trigger a warning signal, where represents the anomaly probability of the detected object; represents time The normal data mean under; Corresponding to time The normal data standard deviation under, when is greater than the mean plus 3 times the standard deviation it is determined to be abnormal.

[0009] Preferably, the layer structure of the 1D-CNN network is: Conv1D(64)-ReLU-MaxPool1D(2)-Conv1D(128)-ReLU-MaxPool1D(2)-Flatten; The position encoding formula is: ; where represents the position encoding, d = 256 is the feature dimension, is the time sequence position, is the dimension index in the position encoding vector, starting from 0.

[0010] Preferably, the following optimizations are carried out when deployed to edge devices, specifically including: model lightweighting and real-time interface optimization; Model lightweighting is to use TensorRT to perform 8-bit integer quantization on the model, reducing the memory occupancy by 75%; merging the convolutional layer and the batch normalization layer to reduce the inference latency; Real-time interface optimization supports the HL7 FHIR standard, receiving DICOM images and JSON format physiological signals through the RESTful API; the output result is encapsulated as an HL7 warning message and pushed to the hospital information system.

[0011] Preferably, the update trigger conditions for the dynamic threshold adjustment are: time trigger, data volume trigger, and emergency trigger; specifically: Time trigger: Force an update of the threshold every 24 hours; Data volume trigger: Update the threshold after accumulating 100 new samples; Emergency trigger: If it is detected that the abnormal probability exceeds the current threshold for 5 consecutive times, recalculate the threshold immediately.

[0012] Preferably, a method for detecting abnormal health data based on a generative adversarial network includes the following steps: 1) Conduct data preparation and feature analysis; Perform data partitioning, dividing the training set, validation set, and test set according to time or patient ID; Perform data extraction. For image data: Use a pre-trained model to extract deep features; for time-series data: Extract time-domain and frequency-domain features; for text data: Extract semantic embeddings of medical records through BERT; 2) Select a detection model; For image anomaly detection: Adopt an autoencoder + reconstruction error threshold method. The specific steps are as follows: A. Train the AE using only normal images; B. Calculate the reconstruction error of the test samples; C. Set the dynamic threshold: Mean of normal sample errors + 3 × standard deviation; For time-series anomaly detection, adopt an LSTM-VAE hybrid model. The specific steps are as follows: A. The encoder compresses the ECG signal into a latent vector; B. The decoder reconstructs the signal and calculates the reconstruction error; C. Jointly optimize the ELBO and the anomaly score; For tabular data anomalies, adopt the Isolation Forest method. The specific steps are as follows: A. Randomly select features and split values to construct a tree structure; B. Calculate the path length of the sample; C. Sort according to the path length to determine the anomaly level; 3) Conduct model training and tuning; 4) Conduct model evaluation, specifically: Index selection and interpretability analysis; 5) Conduct deployment and monitoring; Deploy to edge devices, convert the model to the TFLite format, and deploy it to intelligent wearable devices; Conduct continuous learning, and regularly fine-tune the model with new data to prevent distribution drift.

[0013] The beneficial effects of the present invention are: 1. Generate synthetic data that conforms to medical prior knowledge under the premise of protecting privacy through an anatomically constrained federated generative adversarial network.

[0014] 2. Adopt a hierarchical attention mechanism to fuse image, temporal signal, and text features, and combine with a dynamic threshold to achieve high-precision anomaly detection.

[0015] 3. The accuracy of this system in the pneumonia classification task has been increased to 92.3%, the false alarm rate of ICU sepsis early warning is as low as 1.1%, and it supports real-time deployment on edge devices.

[0016] 4. Generate synthetic data that conforms to medical priors.

[0017] 5. Achieve federated training for cross-hospital collaboration, and the success rate of privacy attack defense is ≥97%.

[0018] 6. Fuse image, temporal signal, and text features, and the false alarm rate of anomaly detection is ≤1.5%.

[0019] 7. Support real-time inference on edge devices, with a latency of ≤100ms. Detailed implementation

[0020] A health data augmentation method based on a generative adversarial network, which is implemented by a health data augmentation system based on a generative adversarial network. The system includes the following components: a generator module, a discriminator module, a federated learning protocol module, and a physiological signal generation module; The generator module: includes a parallel image generation branch and a U-Net segmentation branch; The image generation branch consists of a 5-layer transposed convolution network. Each layer is in turn: a transposed convolution layer, a batch normalization layer, and a LeakyReLU activation function, and outputs a synthetic medical image with a resolution of 512×512; The first layer: input noise vector ( represents a variable, represents a 100-dimensional real vector space), the transposed convolution kernel is 4×4, the stride is 1, the padding is 1, and the output channels are 512; The second to fourth layers: the transposed convolution kernel is 4×4, the stride is 2, and the output channels are 256, 128, and 64 in turn; The fifth layer: the transposed convolution kernel is 4×4, the stride is 2, the output channels are 1 (single-channel medical image), and the activation function is Tanh.

[0021] The U-Net segmentation branch includes 4 layers of encoders and 4 layers of decoders, and outputs a segmentation mask of the target organ; the loss function of the segmentation branch is the Dice coefficient loss, and the weight ratio is 30%; The segmentation targets include key organs such as the heart, lungs, and liver, and the loss function is extended to a multi-organ Dice loss: ; where is the number of organ categories, is the real segmentation mask, is the predicted mask.

[0022] During the implementation process, the weights of the loss function are dynamically adjusted: the loss weights are dynamically adjusted according to the training stage, specifically: First, adaptive weight allocation; Initial stage (the first 50 rounds): The adversarial loss weight is increased to 0.8 to accelerate the convergence of the generator; Middle stage (50 - 150 rounds): The anatomical constraint loss weight is increased to 0.4 to strengthen anatomical rationality; Late stage (after 150 rounds): The federated aggregation loss weight is increased to 0.2 to improve the generalization of the model; Adjustment formula: ; where t is the training round.

[0023] Second, gradient penalty optimization is performed, and the gradient penalty coefficient λ of the Wasserstein adversarial loss changes with the training round: ; where, is the total number of training rounds; The said federated learning protocol module supports cross - institutional collaborative training. When participating parties perform local training, differential privacy stochastic gradient descent is adopted, the gradient clipping threshold C = 0.5, and the Gaussian noise scale σ = 1.2; at the same time, global model aggregation is performed every 10 rounds of communication, and the aggregation method is weighted average, and the weights are allocated according to the data volume of each participating party; at the same time, the privacy budget calculation satisfies ε = 0.8, meeting the HIPAA privacy protection standard; During the implementation process, the federated learning protocol is enhanced: Dynamic participant selection: Dynamically adjust the weights of participants according to data quality (such as annotation consistency, signal - to - noise ratio); Gradient compression: Adopt Top - K gradient sparsification (retain the top 10% gradient values) to reduce the communication overhead by 40%; Heterogeneous model support: Allow the generator structures of each participating party to be different (such as the number of layers, the number of channels), and align the feature spaces through knowledge distillation.

[0024] The said physiological signal generation module uses a temporal - constraint WGAN - GP to generate ECG signals, and the constraint conditions include: QRS wave width 80 - 120ms, heart rate 60 - 100 BPM; The time - continuity loss function of the said physiological signal generation module is defined as the L2 - norm difference between adjacent signal points: ; where, is the smoothing loss value, which reflects the smoothness of the entire sequence data; is the number of elements in the sequence minus 1, that is, the length of the sequence minus 1; represents the element; indicating the +(1)th element; is to calculate the L2 norm of the difference between two adjacent elements.

[0025] Furthermore, the loss function of the generator is: ; where is the Wasserstein adversarial loss, and the gradient penalty coefficient λ = 10; is the Dice loss; is the federated average loss.

[0026] Furthermore, the communication data encryption of the federated learning protocol module adopts the homomorphic encryption algorithm with a key length of 2048 bits. At the same time, it has multi-algorithm compatibility: it supports Paillier, RSA-OAEP, and LWE homomorphic encryption algorithms, and the key length is optional (2048 / 4096 bits).

[0027] Furthermore, the physiological signal generation module further includes: QRS wave width detection algorithm and abnormal signal filtering; the QRS wave width detection algorithm is based on the dynamic time warping to match the standard QRS template with an error tolerance of ±5 ms; abnormal signal filtering: if the heart rate or QRS width of the generated signal exceeds the preset range, it will be automatically discarded and regenerated; if the generated signal exceeds the tolerance range, the reinforcement learning mechanism will be triggered and the generator will be re-optimized.

[0028] Furthermore, the physiological signal generation module supports multi-modal signals: including ECG, EEG, PPG, and the constraint conditions are shown in Table 1 below;

[0029] During the implementation process, the enhancement of QRS wave detection is carried out as follows: Multi-template matching: 10 standard QRS waveform templates (normal, left bundle branch block, right bundle branch block, etc.) are pre-stored, and the similarity is calculated by dynamic time warping (DTW); Adaptive tolerance: Adjust the error tolerance according to the patient's historical data: ; where Tolerance is the tolerance value, 5 ms is a fixed time threshold, and AvgQRSWidth represents the average value of the QRS complex width; Real-time correction: If the generated signal exceeds the limit for 3 consecutive times, the generator will be fine-tuned, and the learning rate will be reduced to 1 / 10 of the initial value.

[0030] At the same time, the expansion of abnormal signal filtering is carried out: a multi-level filtering mechanism is adopted; as shown in Table 2 below;

[0031] Perform signal quality scoring: Based on a weighted scoring model, the formula is: ; Only retain signals with a score ≥ 0.8; Among them, Score is the comprehensive score; QRS_Acc is the characteristic index of the QRS complex. A higher QRS_Acc means more reliable detection of the QRS complex; HR_Stability: Heart rate stability index; This index measures the degree of heart rate fluctuation within a certain period of time, such as obtained by calculating the standard deviation of the heart rate, etc.; ST_Slope is the ST segment slope index. The ST segment is a segment from the end of the QRS complex to the start of the T wave in the electrocardiogram, and its slope change can reflect physiological states such as myocardial blood supply.

[0032] Furthermore, a health data enhancement method based on a generative adversarial network includes the following steps: 1) Data preparation and preprocessing; Including data collection and cleaning, collecting medical data, removing noise, and setting a standardized data format; Including data annotation, where doctors perform key area annotation and privacy data desensitization processing; 2) Select an enhancement method, using traditional methods and deep generation methods; 3) Implement enhancement processing; 4) Evaluate the generated data, perform quality evaluation and privacy protection evaluation; 5) Deploy the enhanced data.

[0033] A health data anomaly detection method based on a generative adversarial network includes the following steps: 1) Perform multi-modal feature extraction; Specifically: Perform image feature extraction: Use a pre-trained ResNet-50 model, remove the fully connected layer and replace it with an adaptive average pooling layer, and output a 256-dimensional feature vector; Perform time series signal feature extraction: A 1D-CNN network extracts local features, and superimposes sinusoidal positional encoding to inject time series information; Perform text feature extraction: The BioBERT model performs word segmentation and entity recognition on the medical record text, and outputs a 768-dimensional clinical semantic embedding; Image features: Support 3D medical images (such as CT, MRI volume data), and use 3D-ResNet to extract spatio-temporal features; Time series signals: Adaptive sampling: Dynamically adjust the sampling rate according to the signal frequency (such as ECG 250Hz, EEG 512Hz); Text features: Entity relationship graph construction: Identify the "drug - symptom - disease" association and generate graph neural network (GNN) embeddings.

[0034] 2) Perform hierarchical attention fusion: Specifically: First-level attention: Using the image features as the Query, the temporal signal features as the Key and Value, calculate the attention features from the image to the signal: ; is the query vector obtained by linearly transforming the image features; is the key vector obtained by linearly transforming the signal features; is the dimension of the key vector ; is to apply the Softmax function to the similarity score matrix obtained above to convert it into a probability distribution form; is the value vector obtained by linearly transforming the signal features.

[0035] Second-level attention: Using the output of the first level as the Query, the text features as the Key and Value, calculate the final fusion features: ; where is the attention feature from the image to the signal; is the attention feature from the fusion feature to the text; is the layer normalization operation, which normalizes the input feature vector in the dimension of the layer to make the input data of each layer have a similar distribution.

[0036] The enhanced content of hierarchical attention in the implementation process: 1. Each level of attention is expanded to 8 heads to enhance the feature interaction ability; 2. Introduce cross-modal attention between modalities (such as text → image), and the formula is: ; where, is the query vector obtained by linearly transforming the text features; is the key vector obtained by linearly transforming the image features; is the value vector obtained by linearly transforming the image features; is the key vector 's dimension, and CrossAttn represents the cross-modal attention calculation function.

[0037] 3. The output of each level of attention is superimposed on the original input to avoid gradient disappearance.

[0038] 3) Perform dynamic threshold adjustment and anomaly determination. Specifically: Update the threshold based on the exponential weighted moving average, set the smoothing factor β = 0.9, and the calculation formula is: ; where, is the mean estimated value at time step ; is the time step The mean estimated value; is the time step of the observed value; is the time step of the variance estimated value; is the time step of the variance estimated value; is the current observed value and the current mean estimated value the square of the difference; is the time step of the standard deviation estimated value.

[0039] Its abnormal determination condition: If the abnormal probability score for 3 consecutive times triggers a warning signal, where represents the abnormal probability of the object being detected; represents time the mean of the normal data under; corresponding to time the standard deviation of the normal data under, when is greater than the mean plus 3 times the standard deviation it is determined as abnormal.

[0040] During the implementation process, adjust the smoothing factor β according to the abnormal probability distribution, ; where represents the time series from −100 to the variance within these 100 time steps; represents the time series the maximum variance in the entire historical record.

[0041] At the same time, combine the patient's historical data (such as underlying diseases, age) and adjust the threshold offset Δ personalized: ; where represents the threshold at time ; represents the time step of the mean estimate; represents based on the time step the standard deviation estimated value calculated; represents the patient-specific adjustment term.

[0042] Furthermore, the layer structure of the 1D-CNN network is: Conv1D(64)-ReLU-MaxPool1D(2)-Conv1D(128)-ReLU-MaxPool1D(2)-Flatten; The position encoding formula is: ; Among them, represents the position encoding, d = 256 is the feature dimension, is the temporal position, is the dimension index in the position encoding vector, starting from 0.

[0043] Furthermore, the following optimizations are carried out when deploying to edge devices, specifically including: model lightweighting and real-time interface optimization. Model lightweighting is to use TensorRT to perform 8-bit integer quantization on the model, reducing the memory occupancy by 75%; merging the convolutional layer and the batch normalization layer to reduce the inference latency. During implementation, multi-level optimization of model lightweighting is carried out, as shown in Table 3 below;

[0044] At the same time, a layer fusion strategy is carried out: merging adjacent Conv1D and BatchNorm layers to reduce the number of memory accesses; fusing activation functions (such as ReLU) into the computational graph of the previous layer to save instruction cycles.

[0045] Carry out the expansion optimization of the real-time structure and support multiple protocols, as shown in Table 4 below;

[0046] Carry out the expansion optimization of the real-time structure and adopt a fault tolerance mechanism, specifically: Data verification: Use CRC32 checksum to verify the data integrity and discard data packets with an error rate > 1%; Retransmission mechanism: If the transmission fails continuously for 3 times, switch to the backup communication channel (such as Bluetooth → Wi-Fi); For example: Deployed to a smart home gateway, integrating a smart watch (PPG), a weighing scale (impedance data), and an environmental sensor (temperature and humidity).

[0047] The real-time interface optimization supports the HL7 FHIR standard, receives DICOM images and physiological signals in JSON format through the RESTful API; the output result is encapsulated as an HL7 warning message and pushed to the hospital information system.

[0048] Furthermore, the update trigger conditions for the dynamic threshold adjustment are: time trigger, data volume trigger, and emergency trigger; specifically: Time trigger: Force an update of the threshold every 24 hours; Data volume trigger: Update the threshold after accumulating 100 new samples; Emergency trigger: If it is detected that the abnormal probability exceeds the current threshold continuously for 5 times, recalculate the threshold immediately.

[0049] When calculating the threshold, offset, temperature compensation, and motion state compensation need to be considered.

[0050] A method for abnormal detection of health data based on a generative adversarial network, including the following steps: 1) Conduct data preparation and feature analysis; Perform data partitioning, dividing the training set, validation set, and test set by time or patient ID; Perform data extraction. For image data: Use a pre-trained model (such as ResNet-50) to extract deep features; for time series data: Extract time domain (mean, variance) and frequency domain (FFT) features; for text data: Extract semantic embeddings of medical record texts through BERT; 2) Select a detection model; For image anomaly detection: Adopt an autoencoder (AE) + reconstruction error threshold method. The specific steps are as follows: A. Train the AE using only normal images; B. Calculate the reconstruction error (MSE or SSIM) of the test samples; C. Set a dynamic threshold: Mean of normal sample errors + 3 × standard deviation; For time series anomaly detection, adopt LSTM-VAE (variational autoencoder). The specific steps are as follows: A. The encoder compresses the ECG signal into a latent vector; B. The decoder reconstructs the signal and calculates the reconstruction error; C. Jointly optimize the ELBO (evidence lower bound) and the anomaly score; For tabular data anomalies, adopt the isolation forest method. The specific steps are as follows: A. Randomly select features and split values to construct a tree structure; B. Calculate the path length of the samples (the path of abnormal samples is shorter); C. Sort according to the path length to determine the anomaly level; 3) Conduct model training and tuning; 4) Conduct model evaluation, specifically: Index selection and interpretability analysis; 5) Conduct deployment and monitoring; Deploy to edge devices, convert the model to the TFLite format, and deploy it to intelligent wearable devices; Conduct continuous learning, and regularly fine-tune the model with new data to prevent distribution drift.

[0051] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be thought of without creative work should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope defined by the claims.

Claims

1. A health data augmentation method based on a generative adversarial network, which is implemented by a health data augmentation system based on a generative adversarial network, characterized in that, The system includes the following components: a generator module, a discriminator module, a federated learning protocol module, and a physiological signal generation module; The generator module: includes a parallel image generation branch and a U-Net segmentation branch; The image generation branch consists of a 5-layer transposed convolution network, with each layer being, in sequence: a transposed convolution layer, a batch normalization layer, and a LeakyReLU activation function, and outputs a synthetic medical image with a resolution of 512×512; The U-Net segmentation branch includes 4 layers of encoders and 4 layers of decoders, and outputs a segmentation mask of the target organ; the loss function of the segmentation branch is the Dice coefficient loss, with a weight ratio of 30%; The federated learning protocol module supports cross-institutional collaborative training. When participating parties perform local training, they use differential privacy stochastic gradient descent, with a gradient clipping threshold C = 0.5 and a Gaussian noise scale σ = 1.2; at the same time, global model aggregation is performed every 10 rounds of communication, and the aggregation method is weighted average, with weights allocated according to the data volume of each participating party; at the same time, the privacy budget calculation satisfies ε = 0.8, meeting the HIPAA privacy protection standard; The physiological signal generation module uses a temporal-constrained WGAN-GP to generate ECG signals, and the constraint conditions include: QRS wave width of 80 - 120 ms and heart rate of 60 - 100 BPM; The time continuity loss function of the physiological signal generation module is defined as the L2 norm difference between adjacent signal points: ; where is the smoothing loss value, which reflects the smoothness of the entire sequence data; is the number of elements in the sequence minus 1, that is, the length of the sequence minus 1; represents the -th element in the sequence; represents the +1-th element; is the L2 norm for calculating the difference between two adjacent elements.

2. The health data enhancement method based on a generative adversarial network according to claim 1, wherein: The loss function of the generator is as follows: ; where is the Wasserstein distance adversarial loss, and the gradient penalty coefficient λ = 10; is the Dice coefficient loss of the segmentation branch, which constrains the positions of key organs such as the heart and lungs; is the federated aggregation loss, which weighted-averages the model parameters of each participating party; The federated learning protocol module uses a homomorphic encryption algorithm for communication data encryption, with a key length of 2048 bits.

3. The health data enhancement method based on a generative adversarial network according to claim 1, wherein: The physiological signal generation module further includes: a QRS wave width detection algorithm and abnormal signal filtering; the QRS wave width detection algorithm is based on the dynamic time warping matching of a standard QRS template, with an error tolerance of ±5 ms; Abnormal signal filtering: If the heart rate or QRS width of the generated signal exceeds the preset range, it is automatically discarded and regenerated.

4. A health data enhancement method based on a generative adversarial network according to claim 1, characterized in that: Including the following steps: 1) Data preparation and preprocessing; Including data collection and cleaning, collecting medical data, removing noise, and setting a standardized data format; Including data annotation, where doctors perform key area annotation and perform privacy data desensitization processing; 2) Select enhancement methods, using traditional methods and deep generation methods; 3) Implement enhancement processing; 4) Evaluate the generated data: perform quality evaluation and privacy protection evaluation; 5) Deploy the enhanced data.

5. A method for detecting abnormal health data based on a generative adversarial network, characterized in that, Including the following steps: 1) Perform multi-modal feature extraction; Specifically: perform image feature extraction: use a pre-trained ResNet-50 model, remove the fully connected layer and replace it with an adaptive average pooling layer, and output a 256-dimensional feature vector; perform temporal signal feature extraction: a 1D-CNN network extracts local features, and superimposes sinusoidal positional encoding to inject temporal information; perform text feature extraction: the BioBERT model tokenizes and performs entity recognition on the medical record text, and outputs a 768-dimensional clinical semantic embedding; 2) Perform hierarchical attention fusion: Specifically: First-level attention: Using the image features as Query, the temporal signal features as Key and Value, calculate the attention features from the image to the signal: ; Among them, is the query vector obtained by linearly transforming the image features; is the key vector obtained by linearly transforming the signal features; is the key vector of dimension; applies the Softmax function to the similarity score matrix obtained above to convert it into a probability distribution form; is the value vector obtained by linearly transforming the signal features; Second-level attention: Using the first-level output as the Query, the text features as the Key and Value, calculate the final fused feature: ; Among them, is the attention feature from image to signal; is the attention feature from fused feature to text; is the layer normalization operation, which normalizes the input feature vector in the dimension of the layer to make the input data of each layer have similar distributions; 3) Perform dynamic threshold adjustment and anomaly determination, specifically: update the threshold based on the exponentially weighted moving average. Let the smoothing factor β = 0.9, and the calculation formula is: ; Among them, is the mean estimate of time step ; is the mean estimate of time step ; is the observed value of time step ; is the variance estimate of time step ; is the variance estimate of time step ; is the square of the difference between the current observed value and the current mean estimate ; is the standard deviation estimate of time step ; Its abnormal determination condition: If the abnormal probability score is , a warning signal is triggered, where represents the abnormal probability of the object under detection; represents time the mean value of normal data under; corresponding to time the standard deviation of normal data under, when is greater than the mean value plus 3 times the standard deviation , it is determined as abnormal.

6. The method for abnormal detection of health data based on a generative adversarial network according to claim 5, wherein: The layer structure of the 1D-CNN network is: Conv1D(64)-ReLU-MaxPool1D(2)-Conv1D(128)-ReLU-MaxPool1D(2)-Flatten; the position encoding formula is: ; where represents the position encoding, d = 256 is the feature dimension, is the temporal position, is the dimension index in the position encoding vector, starting from 0.

7. A method for detecting abnormal health data based on a generative adversarial network according to claim 5, characterized in that: When deploying to edge devices, the following optimizations are performed, specifically including: model lightweighting and real-time interface optimization; Model lightweighting is to use TensorRT to perform 8-bit integer quantization on the model, reducing the memory footprint by 75%; merge the convolutional layer and the batch normalization layer to reduce the inference latency; The real-time interface optimization supports the HL7 FHIR standard, receives DICOM images and physiological signals in JSON format through RESTful APIs; the output results are encapsulated as HL7 warning messages and pushed to the hospital information system.

8. The method for detecting abnormal health data based on a generative adversarial network according to claim 5, characterized in that: The update trigger conditions for the dynamic threshold adjustment are: time trigger, data volume trigger, and emergency trigger; specifically: Time trigger: The threshold is forcibly updated every 24 hours; Data volume trigger: The threshold is updated after receiving 100 new samples cumulatively; Emergency trigger: If it is detected that the abnormal probability exceeds the current threshold continuously for 5 times, the threshold is recalculated immediately.

9. The method for detecting abnormal health data based on a generative adversarial network according to claim 5, characterized in that: It includes the following steps: 1) Perform data preparation and feature analysis; Perform data partitioning, partitioning the training set, validation set, and test set by time or patient ID; Perform data extraction. For image data: Use a pre-trained model to extract deep features; for Time series data: Extract time domain and frequency domain features; for text data: Extract semantic embeddings of medical records through BERT; 2) Select a detection model; For image anomaly detection: Adopt the autoencoder + reconstruction error threshold method. The specific steps are: A. Train the AE using only normal images; B. Calculate the reconstruction error of the test samples; C. Set the dynamic threshold: the mean of the normal sample errors + 3 × standard deviation; For time series anomaly detection, adopt the LSTM-VAE hybrid model. The specific steps are: A. The encoder compresses the ECG signal into a latent vector; B. The decoder reconstructs the signal and calculates the reconstruction error; C. Jointly optimize the ELBO and the anomaly score; For tabular data anomalies, adopt the isolation forest method. The specific steps are: A. Randomly select features and split values to construct a tree structure; B. Calculate the path length of the samples; C. Determine the anomaly level according to the path length ranking; 3) Perform model training and tuning; 4) Perform model evaluation, specifically: metric selection and interpretability analysis; 5) Perform deployment and monitoring; Perform edge device deployment, convert the model to the TFLite format, and deploy it to smart wearable devices; perform continuous learning, and fine-tune the model regularly with new data to prevent distribution drift.

Citation Information

Patent Citations

  • Brain tumor segmentation data enhancement method based on generative adversarial network

    CN111833359A

  • Real image generation method based on annotation images under unsupervised training and storage medium

    CN111899203A

  • Pseudo-CT image generation system based on generative adversarial network

    CN113674330A

  • Federal learning fairness improvement method and system for medical images

    CN117115619A

  • Medical image segmentation method based on SwinUnet and generative adversarial network

    CN117974676A

Cited By

  • Electric power marketing business online auditing all-in-one machine digital auditing system and method

    CN120449070A

  • Cross-mechanism federal collaborative modeling method and system for heterogeneous medical data

    CN121391055A