Data privacy protection method and system

Through the combination method of dual variational autoencoder, coupled generation adversarial network and discriminator, the problem of continuous and discrete data association in medical data synthesis is solved, and high-quality synthetic data generation and privacy protection is achieved, which is suitable for medical data processing.

CN119989420AActive Publication Date: 2025-05-13BEIHANG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510473970.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing medical data synthesis technologies are difficult to capture the complex relationship between continuous and discrete data simultaneously, resulting in poor quality and privacy protection of synthetic data.

Method used

The combination method of dual variational autoencoder, coupled generative adversarial network and discriminator is adopted to encode continuous and discrete data into the same hidden space to generate high-quality synthetic data, and enhance privacy protection capabilities through differential privacy protection mechanisms and dynamic privacy budget allocation strategies.

Benefits of technology

It improves the timing and accuracy of synthetic data, enhances the privacy protection capabilities of data, ensures that the quality of synthetic data is similar to the original data, and is suitable for data processing in the medical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989420A_ABST
    Figure CN119989420A_ABST
Patent Text Reader

Abstract

The invention discloses a data privacy protection method and system. The system comprises a dual variational auto-encoder, a coupling generative adversarial network and a discriminator, the dual variational auto-encoder is used for encoding original data of different types and different feature dimensions into features of a hidden space with the same feature dimension; the coupling generative adversarial network is used for generating hidden space features corresponding to the original data; the coupling generative adversarial network comprises a discrete data generator, a continuous data generator, a time sequence full-connection layer and a full-connection layer, high-order features of hidden spaces output by the discrete data generator and the continuous data generator are input into the time sequence connection layer, output of the time sequence full-connection layer is input into the full-connection layer, and output of the full-connection layer is input into the full-connection layer. The output of the full connection layer is input into a comparator; the output of the comparator is decoded and input to the discriminator, so as to output a label for distinguishing the synthetic data from the real data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a data privacy protection method for medical data and also to a corresponding data privacy protection system, belonging to the technical field of medical care informatics. Background Art

[0002] Data synthesis technology is a technology that generates synthetic data by simulating real-world data. It preserves the utility of data while protecting privacy. In the field of medicine, this technology is particularly important. It can generate synthetic data based on medical imaging statistics, which not only protects patient privacy but also promotes data sharing. In addition, these synthetic data can also be used to train algorithms to improve the efficiency of medical services.

[0003] Data synthesis techniques rely on a variety of models, including variational autoencoders, generative adversarial networks, and diffusion models, with the goal of improving the availability of synthetic data. A variety of models have been proposed and verified for the synthesis of continuous data (such as vital signs) and discrete data (such as medical interventions). However, the quality of medical data synthesis is challenging because there are complex associations between these two types of data, making it difficult for existing methods to simultaneously capture their distributions and their relationships.

[0004] Although the main advantage of synthetic data is that it avoids personal privacy leakage, the synthesis method based on raw data still has the risk of privacy leakage when facing attacks such as reverse engineering and re-identification. The privacy protection issues of synthetic data are similar to the privacy attacks faced by machine learning models, including data theft, attribute inference, and membership inference attacks. Deep generative models are usually vulnerable to these attacks, and the use of privacy protection measures such as differential privacy will sacrifice model accuracy, resulting in a decrease in the quality and availability of synthetic data compared to raw data. Therefore, how to improve the quality and availability of synthetic data while protecting privacy is the main challenge facing current synthetic data technology. Summary of the invention

[0005] The primary technical problem to be solved by the present invention is to provide a data privacy protection system for medical data.

[0006] Another technical problem to be solved by the present invention is to provide a data privacy protection method for medical data.

[0007] In order to achieve the above technical objectives, the present invention adopts the following technical solutions: According to a first aspect of an embodiment of the present invention, a data privacy protection system is provided, comprising a dual variational autoencoder, a coupled generative adversarial network and a discriminator; wherein, The dual variational autoencoder is used to encode raw data of different types and different feature dimensions in the medical field into features of a latent space with the same feature dimension; the raw data includes discrete data and continuous data in a medical scenario, wherein the continuous data is patient vital signs data, and the discrete data is binary identification data of medical intervention measures; the dual variational autoencoder includes an encoder for continuous data, a decoder for continuous data, an encoder for discrete data, a decoder for discrete data, a continuous data sampling module located between the encoder and the decoder for continuous data, a discrete data sampling module located between the encoder and the decoder for discrete data, and a comparator; The coupled generative adversarial network is used to generate latent space features corresponding to the original data; wherein the coupled generative adversarial network includes a generator of discrete data and a generator of continuous data, as well as a temporal fully connected layer and a fully connected layer, the high-order features of the latent space output by the generator of discrete data and the generator of continuous data are input to the temporal connected layer, the output of the temporal fully connected layer is input to the fully connected layer, and the output of the fully connected layer is input to the comparator; The discriminator includes a continuous data discriminator and a discrete data discriminator. The output of the comparator is decoded and input into the discriminator to output a label for distinguishing synthetic data from real data.

[0008] Preferably, the coupled generative adversarial network includes a BLSTM network for continuous data and a BLSTM network for discrete data that are coupled to each other, and also includes a fully connected layer for continuous data and a fully connected layer for discrete data. The BLSTM network of continuous data is used as a generator of continuous data, and continuous random noise is input to obtain its The hidden layer output at the moment is input to the fully connected layer of continuous data; The BLSTM network of discrete data is used as a generator of discrete data, and discrete random noise is input to obtain its The hidden layer output at the moment is input to the fully connected layer of the discrete data; exist At time t, the inputs of the BLSTM network for continuous data and the BLSTM network for discrete data both include: The hidden layer output of the BLSTM network for continuous data at time t, and The hidden layer output of the BLSTM network of discrete data at the moment is obtained The hidden layer output at time t.

[0009] Preferably, the inputs of the encoder for continuous data and the encoder for discrete data are respectively The real data at that time, and the time The corresponding decoder output reconstructed data and hidden state tensor when ; The outputs of the continuous data encoder and the discrete data encoder are sampled respectively to obtain the time The latent space high-order features are input into the corresponding decoders respectively; The decoders are respectively based on the time The hidden state tensor of The latent space high-order features, decoding generation time The reconstruction data, The comparator is based on the time Real data and time The reconstructed data is used for comparative learning.

[0010] Preferably, the tensor of continuous data and the tensor of discrete data output by the time series fully connected layer are input into the fully connected layer; the fully connected layer outputs the final latent space high-order feature representation.

[0011] Preferably, the patient's vital signs data include at least one of heart rate, blood pressure, blood oxygen saturation, and respiratory rate; and the medical intervention measures include binary status identifiers of ventilator use, antihypertensive drug use, and pressor drug use.

[0012] According to a second aspect of an embodiment of the present invention, a data privacy protection method is provided, wherein a continuous random noise and vital sign data at a historical moment are input to a continuous data generator to obtain The hidden layer output at time ; Input discrete random noise and the medical intervention status at historical moments into the discrete data generator to obtain The hidden layer output at time ; Will Input the continuous fully connected layer to get the continuous real data The corresponding high-order features of the latent space ;Will Input the discrete fully connected layer to get the discrete real data The corresponding high-order features of the latent space ; The high-order features of the latent space and Input to the time series fully connected layer, The output at each moment is a tensor group , split to get the continuous data tensor and discrete data tensors ,in Indicates from Until the moment The time step of the moment; Will , Input to the fully connected layer to output the final latent space high-order feature representation and to the comparator; The encoder for continuous data and the encoder for discrete data complete the original data respectively. The encoding is sampled by the continuous data sampling module and the discrete data sampling module respectively, and the moment with the same feature dimension is obtained The latent space high-order features are decoded to obtain the reconstructed data, which is input into the comparator; After contrast learning by the comparator, they are encoded separately and input into the continuous data discriminator and the discrete data discriminator to output labels for distinguishing synthetic data from real data.

[0013] Compared with the prior art, the present invention combines dual variational autoencoders, coupled generative adversarial networks and discriminators to generate synthetic data that is highly similar to the original data while protecting personal privacy, and is particularly suitable for the processing of continuous and discrete data in the medical field. The system not only improves the timing and accuracy of synthetic data, but also further enhances the quality of synthetic data through adversarial training and discriminator evaluation. In addition, the present invention also introduces a differential privacy protection mechanism and a dynamic privacy budget allocation strategy to balance privacy protection and model performance, ensuring that the quality of synthetic data will not decrease as training proceeds. Experimental results confirm that the present invention can generate synthetic data with similar distribution, strong correlation and high availability while protecting privacy, showing significant technical effects and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A schematic diagram of the overall architecture of the CoGAN model provided in the first embodiment of the present invention; Figure 2 for Figure 1 Schematic diagram of the architecture of the dual variational autoencoder in ; Figure 3 for Figure 1 Schematic diagram of the architecture of the coupled recurrent network in ; Figure 4 for Figure 1 Schematic diagram of the architecture of the coupled generative adversarial network in ; Figure 5 Schematic diagram of the architecture of the continuous data discriminator; Figure 6 This is a schematic diagram of the structure of a DP-CoGAN model for implementing dynamic privacy budget allocation in an embodiment of the present invention; Fig. 7A This is the Pearson correlation coefficient effect diagram of the real data for 24 hours; Figure 7B For Fig. 7A The corresponding Pearson correlation coefficient effect diagram of the synthetic data; Figure 8 The privacy protection effects of the three models under different privacy budgets. DETAILED DESCRIPTION

[0015] The technical content of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] For the convenience of describing the technical solution of the present invention, the following description assumes Medical data of each patient For the patients, whose medical data are represented as ,in, represents the length of the time series, From 1 o'clock to All medical data of the patient at the time, including continuous data such as blood pressure and heart rate monitored by instruments ; It also includes discrete data indicating whether there are medical intervention measures, whose value is 1 (intervention occurs) or 0 (no intervention), that is, Specifically, for ,patient The continuous data is represented as ,in Represents the dimension of continuous features, It means that the patient At the moment The collected Similarly, the patient The discrete data is represented as ,in Represents the dimension of continuous features, It means that the patient At the moment The collected The value of the dimension data.

[0017] First embodiment Medical data can be mainly divided into two categories: (1) continuous data describing patients’ vital signs and laboratory test data; and (2) discrete data indicating whether medical intervention was adopted.

[0018] Table 1 schematically shows several common types of medical data.

[0019] Table 1 Classification of medical data types It is not difficult to find that there is a correlation between different data types. For example, the use of antihypertensive drugs will affect the change of blood pressure, and the value of blood pressure at a certain moment will also affect the use of antihypertensive drugs. Therefore, when completing the task of generating medical data, the two types of models cannot be trained completely independently and the generation task cannot be completed independently. However, due to the different formats of continuous data and discrete data, the two feature tensors cannot be simply spliced ​​together to train the model. In addition, the patient's vital signs will not change suddenly. At a certain moment, the patient's vital signs are affected by the vital signs of the previous moment on the one hand, and on the other hand, they are also affected by whether there is medical intervention. Therefore, the influence of these two factors should be considered simultaneously when training the generation model.

[0020] like Figure 1 As shown, the data privacy protection system (referred to as CoGAN model) provided by the first embodiment of the present invention includes a dual variational autoencoder ( Figure 1 The dotted line in the upper middle) and the coupled generative adversarial network ( Figure 1 The part outside the dotted line) and the discriminator. Among them, the dual variational autoencoder is used to encode the original data of different types and different feature dimensions in the medical field into features of the latent space with the same feature dimension; the original data includes discrete data and continuous data in the medical scenario, among which the continuous data is the patient's vital signs data, and the discrete data is the binary identification data of the medical intervention measures. The patient's vital signs data here include at least one of heart rate, blood pressure, blood oxygen saturation, and respiratory rate; the medical intervention measures include binary state identifications of ventilator use, antihypertensive drug use, and pressor drug use. The coupled generative adversarial network is used to generate latent space features corresponding to the original data. In Figure 1 In the figure, the variational autoencoder and coupled generative adversarial network in the dotted box on the left are used to generate continuous data; the variational autoencoder and coupled generative adversarial network in the dotted box on the right are used to generate discrete data. When training the dual variational autoencoder, contrastive learning is performed; when training the coupled generative adversarial network, the features of different types of data are combined at the same time, thereby improving the similarity between the synthetic data and the original data and retaining the correlation between the continuous data and the discrete data.

[0021] The dual variational autoencoder (DVAE) is used to encode continuous data (such as vital signs) and discrete data (such as medical interventions) in medical data into the same latent space to solve the problem of heterogeneity of data types. Through joint encoding of the shared latent space, the complex associations between the two types of data (such as the causal relationship between blood pressure changes and drug interventions) can be captured, breaking through the limitations of existing technologies that independently process different types of data.

[0022] Specifically, the dual variational autoencoder converts continuous data and discrete data Encoded as high-level features of the same latent space . and The task of the generative model is to generate raw data Transformed into high-order features of the generated latent space , thereby solving the problem of inconsistent data types. In addition, when training the dual variational autoencoder, in addition to considering the variational lower bound loss, feature matching loss and contrastive learning loss are also considered to strengthen the feature The connection between.

[0023] like Figure 2 As shown, the dual variational autoencoder in the embodiment of the present invention includes: an encoder for continuous data , decoder for continuous data , encoder for discrete data , a decoder for discrete data , a continuous data sampling module located between the continuous data encoder and decoder, a discrete data sampling module located between the discrete data encoder and decoder, and a comparator. The inputs of the two encoders are the time Real data at the time (continuous real data) or (discrete real data), and the moment The corresponding decoder outputs the reconstructed data and hidden state tensor when ; two encoders (continuous encoders and discrete encoders ) are sampled and the time is obtained. The latent space high-order features ( and ), respectively input into the corresponding decoder (continuous decoder and discrete decoders ). Each decoder is based on the time The hidden state tensor and time The latent space high-order features, decoding generation time Reconstruction data (continuous reconstruction data and discrete reconstruction data ). Based on the time Real data and time The reconstructed data, the contrastor contrasts the learning loss, measures the similarities and differences between the real data and the reconstructed data, and achieves the goal of contrastive learning.

[0024] and The original data will be completed separately The encoding is sampled by the continuous data sampling module and the discrete data sampling module respectively to obtain the encoded data with the same feature dimension. . Then the decoder Complete the encoding data to produce output similar to the original data , (Reconstructed data). The comparator measures the similarity between the reconstructed data and the positive sample, and the difference between the reconstructed data and the negative sample, so that the model can distinguish different categories of data and extract information that can represent the characteristics of the data.

[0025] Figure 2 In Represents a continuous loss, which is an unsupervised generative model here, so the label is the sample itself. It is calculated from the continuous sample itself and the reconstructed sample. Similarly, the discrete loss It is calculated from the discrete samples themselves and the reconstructed samples. This figure shows that after pre-training, the encoding results are input into GAN for training. Here Z is the representation of the latent space, X is the representation of the original data space, and GAN learns the representation of the latent space. Contrastive learning is used in pre-training, and the loss is calculated using the end-to-end reconstructed data of the original data space.

[0026] It can be seen that the decoder converts the high-order features of the latent space back to the original data space, learns effective representations of the data, and is able to reconstruct the data from these representations. Through the decoder, the latent space is explored and new data samples are generated. It helps to learn the potential features of the data, which can be used for downstream tasks. Moreover, the output of the decoder can be used to evaluate the performance of the model. By comparing the original data and the reconstructed data through the comparator, information that can represent the data characteristics can be extracted and the model's ability to capture data characteristics can be quantified.

[0027] Specifically, the time Encoder for continuous data The input consists of three parts: 1) time The real data , 2) Moment Decoder for continuous data Output (reconstructed data) ,3) Encoder for continuous data Moment The hidden state tensor of . Encoder for continuous data The output is the corresponding time The latent space high-order feature representation That is, at time When the continuous data encoder The output can be expressed as follows.

[0028] in, The input is the real data at the corresponding time. In addition, it also includes the decoder output of the previous moment , sig represents the Sigmoid function. Decoder for time-continuous data The input includes: time The latent space high-order feature representation ,time The hidden state tensor of , so the decoder for continuous data The output can be expressed as follows.

[0029] Similarly, for discrete data, the encoder of discrete data The output of can also be expressed as formula 1; the decoder of discrete data The output can also be expressed by formula 2.

[0030] The loss function of the dual variational autoencoder is in addition to the variational lower bound loss In addition, it also includes feature matching loss and contrastive learning loss . Variational Lower Bound Loss It includes the sum of reconstruction loss and KL divergence, where the reconstruction loss is usually calculated using mean square error (MSE), while the KL divergence measures the difference between the latent variable distribution output by the encoder and the standard normal distribution. Feature matching loss The loss functions include cross entropy loss (CE), L1 distance or mean absolute error (MAE), etc. Contrastive learning loss functions can use triple loss, InfoNCE loss, etc. In the following description, examples are given to illustrate , , However, this does not constitute a limitation to the present invention, and other loss functions may also be used according to actual needs.

[0031] In one embodiment of the present invention, feature matching loss is used To measure the distance of features in the latent space. Specifically, the encoder and The sample Encoded as a set of feature tensors in the latent space .because Indicates the characteristics of the same patient, which are semantically similar, so The distance in the latent space should also be closer. Therefore, in the embodiment of the present invention, the Euclidean distance between the two tensors is used as the feature matching loss to measure The distance between.

[0032] In one embodiment of the present invention, the two decoders output to the comparator for contrastive learning. Contrastive learning loss (Contra) is used Indicates that, in order to learn the correlation between groups in the training phase, sufficient constraints are added to the dual variational autoencoder. In the embodiment of the present invention, NT-Xent loss (Normalized Temperature-scaled Cross Entropy Loss) is used, but other losses may also be used.

[0033] Here, the features from the same sample in the latent space are Defined as a pair of positive samples, and features from different samples It is defined as a negative sample. Specifically, assuming there are N samples, after the dual variational autoencoder, N sets of feature tensors in the latent space can be obtained. Since the feature dimensions of the latent space features of continuous data and the latent space features of discrete data are the same, all the features of the latent space can be regarded as a set , .use and Represents different types of features (continuous or discrete), that is .in, Represents the characteristics of continuous data Or the characteristics of discrete data ,in ; accordingly, Can be expressed with Corresponding or , In particular, you can use Represents the characteristics of any continuous data and the characteristics of discrete data , So, for patients In terms of is a positive sample pair, and the rest The group sample pair is a negative sample pair, and its contrastive learning loss function is: .

[0034] in, express The cosine similarity between ; is the temperature normalization factor, ; Indicator vector , if and only if The value is 1. The overall contrast loss function can be expressed as: In summary, the loss function of the dual variational autoencoder can be expressed as: in is the weight of the corresponding loss function term, represents the variational lower bound loss, represents the feature matching loss, Denotes contrastive learning loss. The encoder of continuous data can capture the causal relationship between blood pressure changes and drug intervention through variational lower bound loss and contrastive learning loss.

[0035] Dual learning is achieved through the dual variational autoencoder mentioned above. In the dual variational autoencoder, the two variational autoencoders process continuous data and discrete data respectively, so the model can learn the correlation and transformation between the two types of data. It also helps the model to better complete and predict data when facing incomplete or missing data, thereby improving the generalization ability of the model and generating complete and accurate medical data. This ability is particularly important in medical data analysis and prediction, and can help doctors and researchers better understand and use complex medical data.

[0036] Next, the coupled generative adversarial network in the CoGAN model is further explained.

[0037] As mentioned above, when processing medical data, the temporal characteristics of the medical data should be fully considered, and the correlation between continuous data and discrete data should also be considered. Specifically, the embodiment of the present invention improves the accuracy of generated data by: (1) Medical data includes continuous data describing the patient’s vital signs and discrete data describing whether medical intervention has occurred. The latter is a floating point number with a value of The original data is obtained through the aforementioned dual variational autoencoder The latent space representation of The coupled generative adversarial network is used as a generator, whose generation goal is to generate raw data Generate corresponding The latent space features (2) Medical data has a strong temporal correlation. The value at a certain moment is often affected by the value at the previous moment. In order to enable the generative model to learn this temporal correlation and improve the quality of synthetic data, a temporal fully connected layer is added to fully consider Data generated at all times Received Time data impact.

[0038] The difficulty of generating medical data is continuous data With discrete data Interrelated. Tensor Group It can be encoded into a high-order feature representation in the same latent space through a dual variational autoencoder Then the generator (coupled generative adversarial network) receives a random noise to generate a high-order feature representation tensor group in the latent space .

[0039] In generating high-level feature representation tensor groups When With tensor In one embodiment of the present invention, a coupled recurrent neural network (CRN) is used as part of the generator to learn the association between different data; two RNN networks are coupled together to construct a CRN network, which supports learning features from two sets of data in a cyclic manner.

[0040] Since ordinary LSTM networks can learn long-term dependent information but cannot learn the features of two types of data at the same time, the bidirectional long short-term memory network (BLSTM network for short) is optimized in the embodiment of the present invention. By adding hidden states, the BLSTM network can learn the features of different tensor groups, and the BLSTM network is used to construct a coupled recurrent network.

[0041] Compared with the LSTM network, the BLSTM network in the embodiment of the present invention adds a new hidden state ( or ) input, and through the weight matrix , when the update of the gate control unit is completed.

[0042] At the moment , continuous data can be expressed as: Discrete data can be represented as: , in and are two sets of BLSTM networks coupled to each other, and It is used to generate continuous data. Used to generate discrete data. Figure 3 As shown, and The hidden layer output and It not only serves as the input of the same group of BLSTM networks at the next moment, but also as the input of another group of BLSTM networks.

[0043] Specifically, combined Figure 1 and Figure 3 As shown, the BLSTM network for continuous data ( Figure 3 The midpoint line shows the Receive random noise , the last moment network The hidden state , the network of discrete data at the last moment ( Figure 3 The hidden state of . Indicates the cell state at the current time step. BLSTM network for continuous data At the moment The state update formula is as follows.

[0044] in, Indicates the input gate at time step The activation value of Represents the forget gate at time step The activation value of Indicates the output gate at time step The activation value of Represents the candidate cell state at time step Candidate values ​​of Represents the cell state at time step The value of represents the hidden state at time step The value of Indicates the current time step The input vector of the LSTM network of continuous data (random noise); Represents the time step The hidden state value of the BLSTM network for discrete data; and represents the weight matrix, and different superscripts and subscripts represent different gates and different input sources; represents the bias vector; represents the sigmoid activation function, which is used to compress values ​​between 0 and 1; tanh represents the hyperbolic tangent activation function, which is used to compress values ​​between -1 and 1; ⊙ represents element-wise multiplication (Hadamard product), which is used for the gating mechanism. In addition, all are BLSTM networks for continuous data The value of .

[0045] Similarly, for the BLSTM network of discrete data , at the time Receive random noise , the last moment network The hidden state , BLSTM network of continuous data at the previous moment The hidden state . BLSTM network for discrete data At the moment The state update formula is as follows: in, Indicates that the input gate is at time step The activation value of Indicates that the forget gate is at time step The activation value of Indicates that the output gate is at time step The activation value of Represents the unit state at time step Candidate values ​​of Represents the unit state at time step The value of represents the hidden state at time step The value of Indicates the current time step The input vector of the LSTM network of continuous data (random noise); Represents the time step The value of the hidden state of the BLSTM network for continuous data; and represents the weight matrix, and different superscripts and subscripts represent different gates and different input sources; represents the bias vector; represents the sigmoid activation function, which is used to compress values ​​between 0 and 1; tanh represents the hyperbolic tangent activation function; ⊙ represents element-wise multiplication.

[0046] It can be seen that the BLSTM network in the embodiment of the present invention includes a set of LSTM networks for continuous data and a set of LSTM networks for discrete data to generate two types of data respectively. Among them, in the LSTM network for continuous data provided by the embodiment of the present invention, the activation values ​​of the forget gate, input gate and output gate are all obtained by using two hidden states. and and random noise The result is that the hidden state of the LSTM network for continuous data and the hidden state of the LSTM network for discrete data can be used to associate the two types of data; the bidirectional hidden state can also be calculated by running the LSTM layer forward and backward along the time axis. In each group of BLSTM networks, past and future information are processed simultaneously, so as to better capture the contextual relationship in the sequence data. Therefore, the BLSTM network integrates the temporal correlation of continuous data and discrete data (such as the vital signs of the previous moment affecting the current intervention decision) through bidirectional hidden state transmission, and the generated data is more in line with the real medical logic.

[0047] like Figure 4 As shown, the generator architecture based on the aforementioned BLSTM network includes a generator for continuous data and discrete data generators .

[0048] Generator for continuous data BLSTM network including continuous data And the fully connected layer for continuous data .exist At time t, BLSTM network for continuous data The input includes: continuous random noise , time The hidden layer output as well as BLSTM network for discrete time data The hidden layer output , get its The hidden layer output at time .Will Input continuous fully connected layer ,get Generator for time-continuous data Output , that is, with continuous real data The corresponding high-order features of the latent space (See the formula below).

[0049] Discrete data generator It consists of two components: BLSTM network for discrete data And the fully connected layer for discrete data .exist At time instant, the BLSTM network of discrete data The input includes: discrete random noise , time The hidden layer output as well as time The hidden layer output , get its The hidden layer output at time .Will Input discrete fully connected layer ,get Moment Generator Output (See the formula below).

[0050] Generator and High-order features of the output latent space and , is input into the temporal fully connected layer .exist Moment, Generator and generator Output Tensor Group , Until the moment Moment Generator and generator Output Tensor Group . The tensor With tensor Concatenate to get tensor Input to , exist The output at each moment is a tensor group , split to get the continuous data tensor and discrete data tensors . Among them, the dimensions of all tensor groups are the same.

[0051] The output y of the time series fully connected layer is similar to the output of the regular fully connected layer, satisfying: y=Wx+b, where Wx is matrix multiplication, representing the dot product of the weight matrix W and the input vector x; b is the bias vector, which is added to the output of the weight matrix to adjust the output offset. However, unlike the regular fully connected layer, the input vector is two tensors of the same dimension. After being spliced ​​and input to the fully connected layer, the output vector is split into two output vectors of equal length.

[0052] Output of the temporal fully connected layer , is input to the fully connected layer . exist A tensor that receives continuous data at all times and discrete data tensors As input, and output the final latent space high-order feature representation and , see the following formula: Through the above measures, the temporal fully connected layer of the coupled generative adversarial network can dynamically associate the vital signs data at the current moment with the medical intervention status at the previous moment.

[0053] For continuous data and discrete data, train continuous data discriminators separately and discrete data discriminator The output of the comparator is decoded by the decoder and then input into the discriminator.

[0054] like Figure 5 As shown, the discriminator It consists of two parts: a recurrent neural network composed of LSTM and a fully connected layer. The input is real data Or generate data , the output is the label , so the discriminator can be expressed as : .

[0055] Discriminator Structure and discriminator Similar, I will not go into details here.

[0056] It should be noted that the purpose of the discriminator is to output labels. These labels are the result of comparing and identifying the generated data (synthetic data) with the real data. The discriminator evaluates whether the generated data is realistic enough through comparative learning, that is, whether it can imitate the characteristics and distribution of the real data with a high degree of similarity. If the discriminator can accurately output labels and distinguish between real data and synthetic data, it means that the synthetic data successfully retains the key features of the original data without leaking information sufficient to identify the original data. Therefore, the ability to output labels is an important indicator to measure the effectiveness of a data privacy protection system.

[0057] In addition, the generator , Generator High-level feature representation of the latent space of each output , , cannot be used directly as a discriminator , Instead of inputting to the decoder , , and get the final generated data , , and then input to the corresponding discriminator , .

[0058] The loss function of the generator is: in, represents the loss function of the generator, Represents the distribution Random noise sampled in The expected value obtained by averaging is and represent continuous or discrete random noise, respectively. and represent the generator of continuous or discrete data respectively, and represent continuous or discrete data discriminators, respectively.

[0059] Loss function of the discriminator Satisfy (including ): During the training phase, the generator and the discriminator are updated alternately. The generator continuously learns how to generate more realistic samples, while the discriminator continuously learns how to better distinguish between real samples and generated samples. This adversarial process drives the improvement of the quality of samples generated by the generator.

[0060] Second embodiment The second embodiment of the present invention provides a data privacy protection method, which uses a dual variational autoencoder, a fully connected layer, and a generator and a discriminator for two data types respectively to improve the correlation between continuous data and discrete data. In addition, a temporal connection layer is used to improve the temporal nature of generated data, and random noise is added to the generator to improve privacy security.

[0061] Specifically, the data privacy protection method includes the following steps: The continuous data generator is fed with continuous random noise and the vital signs characteristics of historical moments, specifically, continuous random noise , time The hidden layer output as well as BLSTM network for discrete time data The hidden layer output , get The hidden layer output at time ; Input discrete random noise and medical intervention status characteristics at historical moments into the discrete data generator, specifically discrete random noise , time The hidden layer output as well as time The hidden layer output , get The hidden layer output at time ; Here, the temporal correlation between vital sign characteristics and medical intervention status characteristics is jointly modeled by coupling a bidirectional long short-term memory network, which facilitates the generation of synthetic data that conforms to medical logic.

[0062] Will Input the continuous fully connected layer to get the continuous real data The corresponding high-order features of the latent space ;Will Input the discrete fully connected layer to get the discrete real data The corresponding high-order features of the latent space ; The high-order features of the latent space and Input to the time series fully connected layer, The output at each moment is a tensor group , split to get the continuous data tensor and discrete data tensors ,in Indicates from Until the moment The time step of the moment; Will , Input to the fully connected layer to output the final latent space high-order feature representation and to the comparator; Use the encoder for continuous data and the encoder for discrete data to complete the original data respectively. The encoding is sampled by the continuous data sampling module and the discrete data sampling module to obtain the moment with the same feature dimension The latent space high-order features are decoded to obtain the reconstructed data, which is input into the comparator; After comparative learning by the comparator, they are encoded separately and input into the continuous data discriminator and the discrete data discriminator to output labels for distinguishing synthetic data from real data.

[0063] The implementation process of the data privacy protection method includes a pre-training stage and a training stage. In the pre-training stage, the dual variational autoencoder is first trained, and the loss function is the loss function of the dual variational autoencoder. Then, in the training stage, for the identification of real data, the real data is first trained. Encoding to obtain high-order feature representation of latent space , then Input decoder, get , and respectively Input Discriminator ; In adversarial training, through the generator Generating high-level feature representations ,Will Input Decoder , , get the generated data , and generate data respectively Input Discriminator , and obtain the identification result.

[0064] Based on the CoGAN model provided in the first embodiment of the present invention, the present invention further proposes an improved DP-CoGAN model. Figure 6 As shown in Figure 1, the DP-CoGAN model adds a differential privacy protection mechanism to the CoGAN model. It achieves differential privacy by adding random noise in the model training phase, thereby protecting the individual privacy in the training data. Compared with the CoGAN model, the DP-CoGAN model uses differential privacy to train the generative adversarial network, so that the data generated by the generative network also meets the differential privacy specification.

[0065] Specifically, given the DP-CoGAN model Training set and training process satisfy − Differential privacy, so Generate synthetic data for differential privacy generation models It is also differentially private. Intuitively, differentially private generative models should not directly access original sensitive data during training, and at the same time, the generated samples should be usable in specific downstream tasks. Currently, the mainstream solution is to use the differentially private stochastic gradient descent algorithm to complete the model training.

[0066] In the DP-CoGAN model, the degree of data privacy protection is determined by the privacy budget Decide, The smaller the value, the higher the degree of data privacy protection. When performing the differential privacy stochastic gradient descent algorithm, each round of iteration needs to be based on the noise scale. determines the size of the Gaussian noise added, and Determined by the privacy budget allocated to that iteration. At a certain point, as the number of algorithm iterations increases, the total amount of noise added to the DP-CoGAN model also increases, and the model convergence difficulty also increases. Therefore, it is necessary to dynamically adjust the amount of noise added to avoid reducing the quality of synthetic data as the number of training rounds increases. Starting from the characteristics of multiple iterations of deep learning networks, the embodiment of the present invention balances the size of the privacy budget allocated to each round of iteration, reduces the algorithm overhead, and proposes a dynamic privacy budget allocation strategy.

[0067] In order to avoid adding excessive noise to the model as the number of iterations increases, the embodiment of the present invention gradually increases the privacy budget during the training phase and reduces the noise scale of each iteration. Below, two noise attenuation algorithms are proposed: , or, .

[0068] in, represents the initial noise size, represents the expected noise size at the end of training, represents the training round, Represents the decay rate.

[0069] In the above noise attenuation algorithm, the noise As the training rounds As the noise decreases, the privacy budget will gradually increase, but the noise cannot be reduced indefinitely. In exponential noise decay, the noise will not be less than ; In the power function noise attenuation, the parameter Control the attenuation range, when the number of training rounds When the noise from Decay to ;when When the noise is kept constant.

[0070] In addition, the initial noise size , attenuation rate All parameters need to be determined at the beginning of training. For data with high privacy protection strength, a larger initial noise and a smaller decay rate should be selected; similarly, for applications with high requirements for synthetic data quality, a smaller initial noise and a larger decay rate should be selected.

[0071] In one embodiment of the present invention, an adaptive clipping threshold selection algorithm (ACTS algorithm for short) based on Gaussian mechanism is further proposed. In this algorithm, the samples are divided into multiple batches for training. First, is equal to the maximum gradient norm in the current batch of samples, that is, Then, according to the input parameters , the algorithm converts the interval Divide equally, obtain subintervals; again, for each interval , calculate the number of samples of the gradient norm in this interval, and use Indicates that ; Finally, for any interval and ,Towards Add obedience Gaussian noise of the distribution, and the gradient clipping threshold is determined as the maximum frequency of the interval .

[0072] The ACTS algorithm can solve abnormal samples The impact on the overall clipping threshold can also make it comply with the differential privacy specification. On the one hand, abnormal samples The frequency of at most two intervals can be changed. On the other hand, by adding Gaussian noise , differential privacy protection can be achieved. At the same time, due to the pruning threshold determined by the ACTS algorithm , so it can be considered that the optimal threshold is selected, which makes the algorithm converge more easily. Therefore, the clipping threshold calculates the average gradient size of all samples after processing. The threshold is neither too small, resulting in excessive clipping and information loss, nor too large, resulting in insufficient privacy protection. Therefore, it can better balance privacy protection and model performance.

[0073] Next, the technical effect of the data privacy protection method provided by the embodiment of the present invention is further introduced through experimental data.

[0074] In the verification experiment of the present invention, the MIMIC-III dataset is used. The dataset is constructed by the Computational Physiology Laboratory of the Massachusetts Institute of Technology. The MIMIC-III dataset contains ICU patient data from Beth Israel Deaconess Medical Center from 2001 to 2012. The dataset mainly includes demographic data, patient vital sign measurement data, laboratory test results, treatment methods, hospitalization time and other information.

[0075] In data preprocessing, the preprocessing framework proposed by the MIMIC-EXTRACT program is implemented. As shown in Table 2, the MIMIC-EXTRACT program can extract four tables from the original data: Table 2 Dataset table structure After obtaining the data in the four tables, data processing is required to extract features, handle outliers and normalize the data to obtain data for model training. The data screening conditions include: (1) Patient age ≥ 16 years (2) 24 hours ≤ patient stay ≤ 240 hours After inputting the original data into the MIMIC-EXTRACT program, four tables will be obtained. Among them, the table patients is indexed by patient number (subj_id), hospitalization number (hadm_id), and ICU number (icustay_id), while the remaining three tables are indexed by subj_id, hadm_id, icustay_id, and hours_in. According to the data screening conditions, the embodiment of the present invention obtains 14 discrete features and 114 continuous features. Since not every patient has all the features, it is necessary to complete the processing of abnormal values ​​such as null values, and select the records of the last 24 hours of the patient; then, the embodiment of the present invention completes the aggregation of continuous and discrete features in units of hours; finally, complete the normalization of the data and save the maximum and minimum values ​​of each feature to complete the reverse mapping of the generated data.

[0076] After the above processing, the data to be analyzed that can be used for model training obtained in the embodiment of the present invention is shown in Table 3.

[0077] Table 3 Statistics of data preprocessing results In order to accurately evaluate the quality of synthetic data, the embodiments of the present invention mainly conduct experiments on three aspects: the similarity in distribution between synthetic data and original data, the correlation between continuous data (such as vital signs data) and discrete data (such as medical intervention measures), and the usability of synthetic data in practical applications.

[0078] First, the authenticity of the data, that is, the similarity in distribution between the synthetic data and the original data, the embodiment of the present invention compares the overall differences between the two types of data on the one hand; on the other hand, it also makes a more detailed and intuitive comparison on specific indicators. In the comparison of the overall distribution difference, the embodiment of the present invention uses the maximum mean difference (MMD) and the discriminative score to evaluate the performance of continuous data and discrete data respectively. Among them, the maximum mean difference is mainly evaluated from the difference in distribution between the synthetic data and the original data, while the discriminative score evaluates the authenticity of the synthetic data from the difficulty of distinguishing the synthetic data from the original data.

[0079] MMD is often used to measure the consistency of two sets of distributions. It measures the distance between two distributions in Hilbert space. The smaller the maximum mean difference, the smaller the distance between the two distributions and the closer the distributions are; the larger the maximum mean difference, the larger the distance between the two distributions. The calculation process can be expressed by the following formula: in, and Represents two probability distributions. The kernel function expressed by the following formula can embed the probability distribution into different regenerated Hilbert spaces.

[0080] The core idea of ​​the discriminant score is to train a binary discriminator separately from the generative model, and divide a part of the data into a test set during training. After the generative model is trained, the binary discriminant model is trained using synthetic data (labeled 0) and test set data (labeled 1). The synthetic data is discriminated using the binary discriminator. The accuracy of the binary discriminator is the discrimination score. Therefore, the optimal result of the discrimination score is 0.5.

[0081] Secondly, the embodiment of the present invention uses the Pearson pairwise correlation (PCC) to evaluate the correlation between vital sign data and medical intervention data. The Pearson correlation coefficient is widely used to measure the degree of correlation between two variables. The value of the Pearson correlation coefficient is between -1 and 1. The closer the PCC is to 1, the more positively correlated the feature is; the closer it is to -1, the more negatively correlated the feature is.

[0082] Again, the embodiment of the present invention trains a set of machine learning models on real data and generated data, including logistic regression (LR), decision tree (DT), random forest (RF) and multi-layer perceptron (MLP) models. And calculate the true positive (True Positive, referred to as TP), false positive (False Positive, referred to as FP), true negative (True Negative, referred to as TN), false negative (False Negative, referred to as FN) of each model, and thus calculate the sensitivity of the model , specificity , positive predictive value and negative predictive value Etc. In terms of indicator selection, the AUC score and PR-AUC score were calculated by ROC curve and Precision-Recall curve respectively as model evaluation indicators.

[0083] In the selection of comparison models, since most models cannot generate continuous data and discrete data at the same time, for continuous data, the embodiment of the present invention selects three models such as R(C)GAN, TimeGAN and DP-CGAN as comparison, which are typical GAN ​​models for generating time series data, and DP-CGAN is also a medical data generation model; for discrete data, the comparison models select medGAN, seqGAN and DP-CGAN, which perform better in generating discrete data.

[0084] In the design of the ablation experiment, the embodiment of the present invention mainly verifies the influence of the coupled recurrent network based on the BLSTM network on the generated data. Therefore, the ablation experiment uses the following model: VAE-GAN model: It consists of two parts: a dual variational autoencoder and a generative adversarial network. The dual variational autoencoder is consistent with the previous introduction, and the generator part of the generative adversarial network uses a common LSTM network, which only considers the characteristics of a single type of data.

[0085] CoGAN model: The complete model proposed in the embodiment of the present invention uses a BLSTM network to build a coupled recurrent network as a generator compared to the VAE-GAN model.

[0086] The experimental results are as follows: (1) Authenticity evaluation of continuous synthetic data As shown in the table, the MMD scores of the continuous synthetic data generated by the CoGAN model proposed in the embodiment of the present invention are compared with the MMD values ​​of the synthetic data generated by the existing algorithms such as R(C)GAN, Time-GAN and DP-CGAN.

[0087] Table 4 Comparison of MMD values ​​of continuous synthetic data From the comparative experiments, it can be seen that the CoGAN model proposed in the embodiment of the present invention is better than other models in generating continuous time series data. From the ablation experiments, it can be seen that the generator based on the coupled recurrent network performs better when facing the generation task of different types of data.

[0088] As shown in Table 5, the embodiment of the present invention also uses a discriminant score to evaluate the generator, wherein the discriminant score value of the CoGAN model is 0.755, which is better than the other three models and also better than the VAE-GAN model.

[0089] Table 5 Discrimination scores of continuous synthetic data (2) Authenticity Assessment of Discrete Synthetic Data In the evaluation of discrete synthetic data, since the discrete data takes the value of 0 or 1, the maximum mean difference cannot be used to evaluate the synthetic data. Therefore, the embodiment of the present invention uses the discriminant score to evaluate the discrete value. As shown in Table 6, the CoGAN model proposed in the embodiment of the present invention is also better than the DP-CGAN model in the performance of discrete synthetic data. Through the ablation experiment, it can be found that the data generated by the coupled recurrent network is better than the data generated by the ordinary LSTM.

[0090] Table 6 Discriminant scores of discrete data The embodiment of the present invention also calculates the mean, variance and KS-Statts representing the difference between distributions of each distribution to reflect the performance of the synthetic data on specific indicators, as shown in Table 7.

[0091] Table 7 Comparison of mean and standard deviation of real data and generated data Among them, KS-Stats represents KS statistic (Kolmogorov-Smirnov statistic), which is a statistic used to compare the difference between two distributions. It can be seen that in the comparative experiments with R(C)GAN, Time-GAN, and DP-CGAN, the CoGAN model performs best in all six indicators.

[0092] (3) Relevance assessment In order to evaluate whether there is a correlation between continuous data and discrete data in the synthetic data, the embodiment of the present invention conducts a correlation experiment. Specifically, the embodiment of the present invention uses the Pearson correlation coefficient (PCC) to evaluate the correlation. The embodiment of the present invention selects 24 hours of real data ( Fig. 7A ) and synthetic data ( Figure 7B ) to calculate the Pearson correlation coefficient. As can be seen from the figure, whether in real data or synthetic data, there is a correlation between diastolic blood pressure and pressor drugs, and the correlation is negative. The reason is that a decrease in diastolic blood pressure will lead to medical intervention, that is, the lower the diastolic blood pressure, the higher the possibility of using pressor drugs. In addition, it can be seen that the heart rate and respiratory rate are positively correlated. The reason is that when the heart rate increases, the oxygen consumption is regulated by the nervous system, and the respiratory rate will increase accordingly.

[0093] (4) Usability evaluation Table 8 shows the usability evaluation results using original data, CoGAN generated data, DP-CGANS generated data, TimeGAN generated data, and R(C)GAN generated data. The logistic regression (LR), decision tree (DT), random forest (RF), and multi-layer perceptron (MLP) models were trained and the AUC and PR-AUC scores were calculated. In the tests of the Adult dataset, Census dataset, and Mimic-III dataset, the CoGAN model scored higher than the scores of other generation models.

[0094] Table 8 Detailed results of usability evaluation Since there is still a correlation between synthetic data and real data, there may be some potential privacy leakage risks, including membership inference attacks, re-identification attacks, and attribute inference attacks. For membership inference attacks, the embodiment of the present invention divides the original data into a training set (50%) and a test set (50%), and then only uses the training set to train the model. After generating synthetic data, the k nearest neighbor (kNN) model is trained on the synthetic data. Assume that the attacker has all the real data including training data and test data. Using the kNN model, the nearest neighbor is identified for each sample in the real data set. Use a minimum threshold (e.g., minimum Hamming distance) to predict whether a real data sample belongs to the training data. The experimental results are shown in Table 9. It can be seen that no matter which distance metric is used, the average gap between the DP-CoGAN model provided by the embodiment of the present invention and the optimal score does not exceed 5%.

[0095] Table 9 Privacy attack experiment results In addition, the embodiment of the present invention also studies the impact of different privacy budgets on privacy protection. Given a model and a record, when the attacker knows that the record is used for model training, it is considered information leakage. The embodiment of the present invention selects the CoGAN model without differential privacy, the DP-CoGAN model with differential privacy, and the DP-CGAN model in the prior art for experimentation. When implementing the member reasoning attack, a model with the same structure as the discriminator is used as a shadow model, and then a fully connected network is constructed as an attack model to evaluate the privacy protection effects of the three models under different privacy budgets. The results are shown in Figure 2. Figure 8 As shown. Figure 8 In the figure, the horizontal axis represents the privacy budget , the vertical axis represents the accuracy of member reasoning attacks, where the closer it is to 0.5, the better the model is at preventing member reasoning attacks and the higher the degree of protection of individual privacy. Figure 8 It can be seen that even when When , the accuracy of the DP-CoGAN model with differential privacy never exceeded 0.6, while the accuracy of the CoGAN model without differential privacy reached above 0.7, which shows that the privacy protection effect is obvious.

[0096] It should be noted that the above embodiments are only examples, and the technical solutions of the various embodiments can be combined, and the order of the steps can be changed, all within the protection scope of the present invention.

[0097] The above is a detailed description of the data privacy protection method and system provided by the present invention. For those skilled in the art, any obvious changes made to it without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.

Claims

1. A data privacy protection system, characterized in that It includes dual variational autoencoder, coupled generative adversarial network and discriminator; among them, The dual variational autoencoder is used to encode raw data of different types and different feature dimensions in the medical field into features of a latent space with the same feature dimension; the raw data includes discrete data and continuous data in a medical scenario, wherein the continuous data is patient vital signs data, and the discrete data is binary identification data of medical intervention measures; the dual variational autoencoder includes an encoder for continuous data, a decoder for continuous data, an encoder for discrete data, a decoder for discrete data, a continuous data sampling module located between the encoder and the decoder for continuous data, a discrete data sampling module located between the encoder and the decoder for discrete data, and a comparator; The coupled generative adversarial network is used to generate latent space features corresponding to the original data; wherein the coupled generative adversarial network includes a generator of discrete data and a generator of continuous data, as well as a temporal fully connected layer and a fully connected layer, the high-order features of the latent space output by the generator of discrete data and the generator of continuous data are input to the temporal connected layer, the output of the temporal fully connected layer is input to the fully connected layer, and the output of the fully connected layer is input to the comparator; The discriminator includes a continuous data discriminator and a discrete data discriminator. The output of the comparator is decoded and input into the discriminator to output a label for distinguishing synthetic data from real data.

2. The data privacy protection system according to claim 1, characterized in that: The coupled generative adversarial network includes a BLSTM network for continuous data and a BLSTM network for discrete data that are coupled to each other, and also includes a fully connected layer for continuous data and a fully connected layer for discrete data; wherein, The BLSTM network of continuous data is used as a generator of continuous data, and continuous random noise is input to obtain The hidden layer output at the moment is input to the fully connected layer of continuous data; The BLSTM network of discrete data is used as a generator of discrete data, and discrete random noise is input to obtain The hidden layer output at the moment is input to the fully connected layer of the discrete data; exist At time t, the inputs of the BLSTM network for continuous data and the BLSTM network for discrete data both include: The hidden layer output of the BLSTM network for continuous data at time t, and The hidden layer output of the BLSTM network of discrete data at the moment is obtained The hidden layer output at time t.

3. The data privacy protection system according to claim 2, characterized in that: The inputs of the continuous data encoder and the discrete data encoder are respectively the time The real data at that time, and the time The reconstructed data and hidden state tensor output by the decoder corresponding to ; The outputs of the encoder for continuous data and the encoder for discrete data are sampled respectively to obtain the time The latent space high-order features are input into the corresponding decoders respectively; The decoders are respectively based on the time The hidden state tensor of The latent space high-order features, decoding generation time The reconstruction data, The comparator is based on the time Real data and time The reconstructed data is used for comparative learning.

4. The data privacy protection system according to claim 3, characterized in that: The temporal fully connected layer outputs a tensor of continuous data and a tensor of discrete data, which are input into the fully connected layer; the fully connected layer outputs the final latent space high-order feature representation.

5. The data privacy protection system according to claim 4, characterized in that: The input vector of the time series fully connected layer is two tensors of the same dimension, one continuous and one discrete. After being spliced ​​and input to the fully connected layer, the output vector is split into a tensor of continuous data and a tensor of discrete data of equal length.

6. The data privacy protection system according to claim 1, characterized in that: The patient's vital signs data include at least one of heart rate, blood pressure, blood oxygen saturation, and respiratory rate; the medical intervention measures include binary status identifiers of ventilator use, antihypertensive drug use, and pressor drug use.

7. A data privacy protection method, implemented based on the data privacy protection system according to any one of claims 1 to 6, characterized in that The steps include: Input continuous random noise and historical vital signs data into the continuous data generator to obtain The hidden layer output at time ; Input discrete random noise and the medical intervention status at historical moments into the discrete data generator to obtain The hidden layer output at time ; Will Input the continuous fully connected layer to get the continuous real data The corresponding high-order features of the latent space ;Will Input the discrete fully connected layer to get the discrete real data The corresponding high-order features of the latent space ; The high-order features of the latent space and Input to the time series fully connected layer, The output at each moment is a tensor group , split to get the continuous data tensor and discrete data tensors ,in Indicates from Until the moment The time step of the moment; Will , Input to the fully connected layer to output the final latent space high-order feature representation and to the comparator; Use the encoder for continuous data and the encoder for discrete data to complete the original data respectively. The encoding is sampled by the continuous data sampling module and the discrete data sampling module to obtain the moment with the same feature dimension The latent space high-order features are decoded to obtain the reconstructed data, which is input into the comparator; After comparative learning by the comparator, they are encoded separately and input into the continuous data discriminator and the discrete data discriminator to output labels for distinguishing synthetic data from real data.

8. The data privacy protection method according to claim 7, characterized in that: By coupling a bidirectional long short-term memory network, the temporal correlation between vital sign characteristics and medical intervention status characteristics is jointly modeled to generate synthetic data that conforms to medical logic.

9. The data privacy protection method according to claim 7, characterized in that: The encoder of the continuous data captures the causal relationship between blood pressure changes and drug intervention through variational lower bound loss and contrastive learning loss; the temporal fully connected layer of the coupled generative adversarial network dynamically associates the vital sign data at the current moment with the medical intervention status at the previous moment.

10. The data privacy protection method according to claim 7, characterized in that: The encoder of the continuous data and the encoder of the discrete data are constructed into a dual variational autoencoder, and the loss function of the dual variational autoencoder is: variational lower bound loss , feature matching loss Or contrastive learning loss .

Citation Information

Patent Citations

  • Privacy protection data generation method based on generative adversarial network

    CN115936107A

  • Video prediction method and system based on space-time decoupling and self-attention difference LSTM

    CN116524419A

  • Industrial process time series data generation method and system

    CN116894186A

  • Medical data synthesis method and device

    CN117077641A

  • Electronic health record generation method based on diffusion model and convolution auto-encoder

    CN117789904A