Data generation method and model based on generative adversarial network and generative model

Through the data generation method of the generative network and the generation model, the problem of limited access to medical data is solved, high-quality and high-diversity data generation is achieved, model performance is improved, and data privacy protection is ensured.

CN120199510APending Publication Date: 2025-06-24SECOND AFFILIATED HOSPITAL OF COLLEGE OF MEDICINEOF XIAN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100486.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The acquisition of medical data is limited by patient privacy protection, which leads to the lack of data in large model training, which in turn affects the development of machine learning in biomedical research.

Method used

By introducing data generation methods based on adversarial generation networks and generation models, a unique network architecture is used to achieve the comprehensive generation of continuous and discrete data, combining multiple loss functions such as matching loss, contrast loss and semantic loss, ensuring the high quality and diversity of generated data, and protecting data privacy through differential privacy and other measures.

Benefits of technology

It significantly improves the quality and diversity of data, solves the problem of data lack, improves the performance and generalization capabilities of the model, and ensures the security of data. It is suitable for many fields such as medical diagnosis and disease prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199510A_ABST
    Figure CN120199510A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and model based on an adversarial generative network and a generative model, continuous data and discrete data are jointly mapped to a shared potential space through a generative model network, and the generative model network is pre-trained to learn the shared potential space. A synthetic potential embedding representation is generated using a conditional random network (CRN)-based sequential coupling generator, decoded into synthetic patient trajectory data in an observation space by a decoder. And training is continued based on the classifier, and data evaluation is performed to improve the quality of the synthetic data. Feature information is constructed through feature information extracted from the synthesized patient trajectory data, a loss function of the feature information is fed back and acts on the data generation network, and weight parameters in the network are optimized. The method particularly emphasizes the modeling of uncertainty, and improves the robustness and credibility of generated data through the combined modeling of conditional Gaussian distribution and variational Dropout technology. The method solves the problem of data shortage in a large model training process, and is suitable for sensitive fields such as medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of machine learning, and in particular, to a data generation method and model based on a generative adversarial network and a generative model. Background Art

[0002] In the past decade, due to the explosive growth of medical data such as electronic health records (EHRs), breakthroughs have been made in the field of computational health. The secondary use of EHRs has sparked a variety of research, especially machine learning (ML)-based digital health solutions for improving care services.

[0003] However, in practice, the benefits of data-driven research are limited to healthcare organizations (HCOs) that have the data. Due to concerns about patient privacy, HCO stakeholders are reluctant to share patient data. The access to clinical data is usually restricted or too costly, resulting in the problem of data shortage in the large model training process, further causing machine learning in biomedical research to lag behind other fields of artificial intelligence. Summary of the Invention

[0004] The present invention realizes the comprehensive generation of continuous and discrete data through a unique network architecture, significantly improving the quality and diversity of the data. This method not only ensures a high degree of similarity between the generated data and the real data in multiple dimensions, but also further guarantees the high quality and diversity of the data by introducing various loss functions, such as matching loss, contrast loss, and semantic loss. In addition, this method also performs well in privacy protection, ensuring the security of the data through measures such as differential privacy, which is particularly important for sensitive fields such as healthcare.

[0005] The technical innovation of this method lies in the introduction of a sequential coupling generator based on a conditional random network (CRN) and a loss function feedback mechanism for feature information. These technologies not only improve the quality of data generation, but also enable the model to handle more complex data types and relationships. The generated synthetic patient trajectory data has broad application prospects and can be applied to multiple aspects such as medical diagnosis, disease prediction, and patient management, providing more accurate and comprehensive diagnostic basis for doctors. At the same time, it can also be used to train large machine learning models, solve the problem of data shortage, and improve the performance and generalization ability of the models.

[0006] In contrast, although other data generation methods have made contributions in specific fields, they have limitations in the comprehensive performance and application breadth of data generation. For example, "Differentially private synthetic medical data generation using convolutional GANs" focuses on privacy protection but is somewhat lacking in data quality and diversity; "Brain Imaging Generation with Latent Diffusion Models" mainly focuses on the generation of high-resolution 3D brain images in the field of medical imaging, and its application scope is relatively limited; "Synthetic Subject Generation with Coupled Coherent Time Series Data" focuses on the coherent and consistent generation of time series data and does not involve the comprehensive generation of multiple types of data; "Biomedical Data-to-Text Generation via Fine-Tuning Transformers" mainly focuses on text generation and does not involve privacy protection measures.

[0007] Generally speaking, the present invention demonstrates significant advantages and broad application prospects in the field of data generation. It not only improves the quality and diversity of data generation but also ensures the security of data through advanced privacy protection measures. At the same time, this method also has high technological innovation, bringing new development opportunities to the field of data generation. With the continuous popularization and application of this method, it is believed that it will provide higher-quality data support and services for more industries.

[0008] According to the first aspect of the embodiments of the present invention, a data generation method based on an adversarial generative network and a generative model is provided, including the following steps: S1. Jointly map continuous data and discrete data to a shared latent space through a generative model network; S2. Pre-train the generative model network to learn the shared latent space; S3. Generate a synthetic latent embedding representation of the continuous data and the discrete data through an order-coupled generator constructed based on a conditional random network CRN; S4. Input the synthetic latent embedding representation into a pre-trained decoder, and decode it into synthetic patient trajectory data in the observation space through the decoder; S5. Separate the uncertainty of the synthetic patient trajectory data, and continue to train and evaluate the data based on a classifier to obtain high-quality synthetic patient trajectory data; S6. Construct feature information from the feature information extracted from the synthetic patient trajectory data, and feedback the loss function of the constructed feature information to act on the data generation network, and optimize the weight parameters in the data generation network to improve the quality of the synthetic patient trajectory data.

[0009] In one implementation, before step S1, it further includes: processing heterogeneous patient data into the continuous data and the discrete data.

[0010] In another implementation, step S1 specifically includes: starting from two encoders that train the generative model network, that is Its embedding function is:

[0011]

[0012] After passing the data of x C and x D through two encoders, a pair of embedding vectors is obtained in the shared latent space H S Then the decoders of the two domains and and further reconstruct features based on the latent embedding using the mapping function in the opposite direction.

[0013] In another implementation, step S2 specifically includes:

[0014] Using the continuous data and the discrete data to pre-train the generative model network to learn the total objective function for utilizing the shared latent space with d ∈ {C, D}:

[0015]

[0016] wherein, the generative model network consists of a pair of encoders (flow C , flow D ) and decoders (Inverse C , Inverse D );

[0017] Using the generative model network to learn the shared latent space, where the gap between the embedding representations of the two domains is minimized, l d is the total objective function, is the loss function in the generative model of domain d ∈ {C, D}, l Match represents the matching loss function, ensuring that the low-dimensional latent space can be shared among heterogeneous features, and combining the pairwise matching loss to motivate the encoder to minimize the distance within the corresponding representation pairs, l Contra represents the contrastive loss function, and the corresponding auxiliary task for calculating the contrastive loss is applied to the scenario of learning representations from mixed-type time series, to determine whether a set of representations transformed from the observation space belongs to the same patient. For the patient data of N records, N pairs of latent representations are obtained from the encoders in the generative model network. For the patient with index i, h i Cand respectively represent the embeddings from the continuous - value and discrete - value observation spaces. Due to the symmetric architecture of the generative model network, different domains are denoted by d and d’, and the positive - sample pair of patient i is denoted as i d , and the other 2(N - 1) samples are used as negative - sample pairs. Then, the contrastive loss of the positive - sample pair can be defined as The final contrastive loss function l Contra is the mean of denotes the semantic loss function, l certainty , l uncertainty denote the deterministic and uncertainty loss functions respectively.

[0018] In another implementation, step S3 specifically includes: The sequential - coupling generator constructed based on the conditional random network CRN takes the random - noise vector as input and iteratively passes through the internal transformation function across time steps t ∈ {1, 2,..., T} to obtain the synthetic latent - embedding representations of the continuous data and the discrete data

[0019] In another implementation, step S6 specifically includes: Processing the input features through convolution operations and pooling operations to obtain the output feature A of the atrous spatial pyramid pooling (ASPP) i ; Processing the output feature A after average pooling and max pooling through convolution operations using the channel - attention factor s1 i ; Processing the original output feature A through a depth - separable convolution using the spatial - attention factor s2 i ; Using the feedback - feature selection operation S to multiply the output feature A i by the channel - attention factor s1 and the spatial - attention factor s2 to generate the weighted feature R i to perform weighted selection on the channel and spatial features of the output feature A i .

[0020] According to the second aspect of the embodiments of the present invention, there is provided a model for implementing the data generation method based on the adversarial generative network and the generative model as described in the first aspect above, including: a mapping module for jointly mapping continuous data and discrete data to a shared latent space through a generative model network; a learning module for pre-training the generative model network to learn the shared latent space; a generation module for generating a synthetic latent embedding representation of continuous data and discrete data through a sequential coupling generator constructed based on a conditional random network (CRN); a decoding module for inputting the synthetic latent embedding representation into a pre-trained decoder and decoding it into synthetic patient trajectory data in the observation space through the decoder; an evaluation module for separating the uncertainty of the synthetic patient trajectory data and continuously training and evaluating the data based on a classifier to obtain high-quality synthetic patient trajectory data; and a feedback module for constructing feature information from the feature information extracted from the synthetic patient trajectory data, and feeding back the loss function of the constructed feature information to act on the data generation network to optimize the weight parameters in the data generation network to improve the quality of the synthetic patient trajectory data.

[0021] In summary, the solution of the present invention realizes the generation of data with high quality and high availability, solves the problem of data shortage in the large model training process, and can capture the temporal dynamics and correlations between various features. It particularly emphasizes the modeling of uncertainties, including stochastic uncertainty and parameter uncertainty, and jointly models them through conditional Gaussian distribution and variational Dropout technology, improving the robustness and credibility of the generated data. Moreover, it can support downstream tasks such as association analysis and result prediction by accurately reconstructing the interdependencies between features and complex clinical relationships. In addition, by introducing multiple loss functions (such as matching loss, contrast loss, and semantic loss), the high quality and diversity of the data are further ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0023] Figure 1 is a flowchart of the steps of the data generation method based on the adversarial generative network and the generative model according to the embodiments of the present invention;

[0024] Figure 2 For Figure 1 the corresponding data generation framework diagram of the data generation method based on the adversarial generative network and the generative model;

[0025] Figure 3Model framework diagram of the data generation method based on the adversarial generative network and the generative model according to the embodiments of the present invention;

[0026] Figure 4 Feedback mechanism diagram of the data generation method based on the adversarial generative network and the generative model according to the embodiments of the present invention;

[0027] Figure 5 Dual diffusion generative model framework diagram of the data generation method based on the adversarial generative network and the generative model according to the embodiments of the present invention. Detailed implementation manners

[0028] For a clearer understanding of the technical features, objectives, and effects of the embodiments of the present invention, the specific implementation manners of the embodiments of the present invention will now be described with reference to the accompanying drawings.

[0029] In this document, "exemplarily" means "serving as an instance, example, or illustration", and any illustration or implementation manner described as "schematic" in this document should not be interpreted as a more preferred or more advantageous technical solution.

[0030] To make the drawings concise, only the parts related to the present invention are schematically shown in each drawing, and they do not represent the actual structure of the product. In addition, for the sake of simplicity and clarity of the drawings for easy understanding, in some drawings, components with the same structure or function are only schematically shown for one or more of them, or only one or more of them are labeled.

[0031] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention shall fall within the scope of protection of the embodiments of the present invention.

[0032] The following further illustrates the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.

[0033] See Figure 1 Step flowchart of the data generation method based on the adversarial generative network and the generative model according to the embodiments of the present invention, Figure 2 For Figure 1 The corresponding data generation framework diagram of the data generation method based on the adversarial generative network and the generative model, in combination with Figure 1 、 Figure 2 , the embodiments of the present invention include the following steps:

[0034] Step S1, jointly map continuous data and discrete data to a shared latent space through a generative model network;

[0035] Step S2: Pre-train the generative model network to learn the shared latent space;

[0036] Step S3: Generate synthetic latent embedding representations of continuous and discrete data through a sequential coupling generator constructed based on the conditional random network CRN;

[0037] Step S4: Input the synthetic latent embedding representation into the pre-trained decoder, and decode it into synthetic patient trajectory data in the observation space through the decoder;

[0038] Step S5: Separate the uncertainty of the synthetic patient trajectory data, and continue to train and evaluate the data based on the classifier to obtain high-quality synthetic patient trajectory data.

[0039] Step S6: Construct feature information from the feature information extracted from the synthetic patient trajectory data, and feedback the loss function of the constructed feature information to the data generation network to optimize the weight parameters in the data generation network to improve the quality of the synthetic patient trajectory data.

[0040] It should be understood that the generative model network in the present invention, i.e., the generative model, is not limited to a specific type, and can be a series of generative models such as a double VAE model, a double diffusion model, etc.

[0041] For the sake of convenience of description, in the subsequent steps of this embodiment, the double diffusion generative model is taken as an example for data generation. See Figure 5 Fig. is the framework diagram of the double diffusion generative model of the data generation method based on the adversarial generative network and the generative model according to the embodiment of the present invention.

[0042] It should also be understood that different feedback mechanisms can be adopted for different generated data, such as: spatial attention mechanism, Laplacian pyramid pooling, etc., which are not specifically limited herein.

[0043] The generated data targeted by the embodiment of the present invention is synthetic patient trajectory data, and the subsequent steps will be described by taking the atrous spatial pyramid pooling (ASPP) as an example.

[0044] In summary, the solution of the present invention realizes the generation of data with high quality and high availability, solves the problem of data shortage in the process of large model training, and can capture the temporal dynamics and correlations between various features. It particularly emphasizes the modeling of uncertainties, including stochastic uncertainty and parameter uncertainty, and jointly models them through conditional Gaussian distribution and variational Dropout techniques, improving the robustness and credibility of the generated data. Moreover, it can support downstream tasks such as association analysis and result prediction by accurately reconstructing the interdependencies and complex clinical relationships between features. In addition, by introducing multiple loss functions (such as matching loss, contrast loss, and semantic loss), the high quality and diversity of the data are further ensured.

[0045] Further, before step S1, it further includes: processing heterogeneous patient data into continuous data and discrete data.

[0046] It should be understood that a prerequisite for successfully training the EHR-M adversarial generation network to generate reversible latent codes is to satisfy the following assumption: for the same patient with index i, and can both be encoded into the same latent space where |S| represents its spatial dimension.

[0047] It should also be understood that the synthetic patient trajectory data is represented as

[0048] To achieve this goal, the present invention intends to use a dual diffusion generation model framework, which uses two dual diffusion generation model networks to encode continuous and discrete multivariate time series into a dense representation based on multiple constraints in H S within.

[0049] Taking the dual diffusion generation model as the generation model of this embodiment for exemplary description, step S1 specifically includes: starting from the two encoders of the trained dual diffusion generation model network, that is whose embedding function is:

[0050]

[0051] After passing the data of x C and x D through the two encoders, a pair of embedding vectors S are obtained in the shared latent space H Then the decoders of the two domains and use the mapping functions in the opposite direction to further reconstruct the features based on the latent embeddings:

[0052]

[0053] Further, step S2 specifically includes:

[0054] Pre - train the double - diffusion generative model network using continuous data and discrete data to learn the total objective function for the shared latent space using d ∈ {C, D}:

[0055]

[0056] Among them, the double - diffusion generative model network consists of a pair of encoders (flow C , flow D ) and decoders (Inverse C , Inverse D );

[0057] Use the double - diffusion generative model network to learn the shared latent space where the gap between the embedding representations of the two domains is minimized, l is the total objective function, d is the loss function in the double - diffusion generative model of domain and its expression is: Among them, p

[0058]

[0059] (x B ) is the trajectory distribution, N is Gaussian noise; is Gaussian noise;

[0060]

[0061] Among them, f is the drift term and g is the diffusion coefficient;

[0062] l Match represents the matching loss function, ensuring that the low - dimensional latent space can be shared among heterogeneous features, and combining the pairwise matching loss to encourage the encoder to minimize the distance within the corresponding representation pairs:

[0063]

[0064] Among them, when l Match is zero, the pairwise matching loss reaches the best;

[0065] l Contra represents the contrastive loss function. The corresponding auxiliary task for calculating the contrastive loss is applied to the scenario of learning representations from mixed - type time series. The goal of this task is to determine whether a set of representations transformed from the observation space belongs to the same patient, which results in corresponding positive pairs (true) and negative pairs (false). For the patient data of N records, N pairs of latent representations are obtained from the encoder in the double - diffusion generative model network. For the patient with index i, h iC and respectively represent the embeddings from the continuous - valued and discrete - valued observation spaces. Due to the symmetric architecture of the double - diffusion generative model network, d and d’ are used to represent different domains. The positive - sample pair of patient i is denoted as i d , and the other 2(N - 1) samples are used as negative - sample pairs. Then, the contrastive loss of the positive - sample pair can be defined as

[0066]

[0067] where sim(u, v)=u T ·v / ||u||·||v|| represents the cosine similarity between two vectors, and τ > 0 represents the temperature hyper - parameter, represents 1 if and only if i dd‘ ≠i d and i dd' ∈{1, 2, …, 2N} represents the index of the latent embeddings of the two data types. The final contrastive loss function l Contra is the mean of l Contra , and the expression of l

[0068]

[0069] represents the semantic loss function. Using logistic regression as a linear classifier, the cross - entropy is calculated as the semantic loss for the two domains. For d ∈ {C, d}, given the latent embedding vector z d and the conditional information vector y:

[0070]

[0071] where represents the linear classifier for the corresponding domain.

[0072] It should be understood that the semantic loss is imposed to better arrange patients with the same label (condition) into the same latent - space cluster. For the two domains, an additional linear classifier is trained to classify the latent embeddings according to their corresponding conditions in the observation space.

[0073]

[0074] where is the output of the linear classifier, and y j is the true label of class j in the conditional vector y.

[0075] l certainty l uncertaintyThey represent the certainty and uncertainty of data respectively. In the process of medical data synthesis, uncertainty is crucial for the quality and reliability of the generated data. This embodiment proposes a joint modeling method for aleatoric uncertainty and epistemic uncertainty, so as to more accurately capture the complexity and potential risks in the data. This method can effectively improve the robustness of the generation model and the credibility of the generated data.

[0076] Modeling Aleatoric Uncertainty and Heteroscedastic Noise

[0077] Aleatoric Uncertainty refers to the irreducible ambiguity caused by the noise of the data itself or measurement errors. To model aleatoric uncertainty, the present invention assumes that the generated patient trajectory data follows a conditional Gaussian distribution, and the variance is related to the input data:

[0078]

[0079] where μ(x; ω1) is the predicted mean of the input data x, and Σ(x; ω2) is the conditional covariance matrix, which reflects the input-dependent noise level. By minimizing the negative log-likelihood loss, the predicted mean and variance of the model can be optimized simultaneously:

[0080]

[0081] Modeling Epistemic Uncertainty and Variational Dropout

[0082] Epistemic Uncertainty reflects the ambiguity that the model may generate when facing limited data. The present invention models epistemic uncertainty by introducing the variational dropout technique and uses variational inference to approximate the true posterior distribution of the parameters. The specific method is to apply Dropout in each layer of the model to generate multiple network samples, thereby obtaining the prediction distributions under different parameter configurations.

[0083] The optimization objective of variational dropout is to minimize the KL divergence:

[0084] KL[q(ω)||p(ω|D)]

[0085] At each forward pass, the parameter ω is regarded as a random variable sampled from the approximate posterior q(ω), and different prediction results are generated accordingly.

[0086] Joint Modeling of Aleatoric Uncertainty and Epistemic Uncertainty

[0087] To capture both aleatoric uncertainty and epistemic uncertainty simultaneously, the present invention combines variational dropout and the heteroscedastic Gaussian model in the generation model. The prediction distribution of the output y of the generation model is:

[0088] p(y|x,D) = ∫p(y|x; ω)p(ω|D)d

[0089] Estimate this distribution through Monte Carlo sampling:

[0090]

[0091] where ω (t) is the network parameter sampled from the posterior distribution. This method accurately reflects the uncertainty in the generated data through the means and variances predicted by multiple models.

[0092] Uncertainty Decomposition and Propagation

[0093] When analyzing the output of the generative model, the total uncertainty can be decomposed into two parts: **parameter uncertainty and stochastic uncertainty**:

[0094] V[y] = V 参数不确定性 [y] + V 随机不确定性 [y]

[0095] Parameter uncertainty reflects the impact of changes in model parameters on the output, while stochastic uncertainty reflects the inherent noise in the data itself. By calculating the prediction variances under different parameters, these two types of uncertainties can be estimated separately, thus better understanding and optimizing the prediction behavior of the model.

[0096] The deterministic part represents the core prediction result of the generative model for the input data. To make the mean of the generated data as consistent as possible with the real data, the Mean Squared Error (MSE) is used as the loss function to optimize the mean prediction μ(x) of the model.

[0097] The optimization objective is as follows:

[0098]

[0099] where y i is the real data, representing the actual trajectory of the patient, and μ(x i ) is the mean prediction of the generative model for the input x i .

[0100] By minimizing this mean loss, the model can more accurately capture the core structural features of the data, ensuring that the generated data is consistent with the real data in key aspects.

[0101] Optimization of the Uncertainty Part

[0102] The uncertainty part reflects the model's estimation of noise or measurement errors in the data. To optimize the model's performance in high-noise regions, the Negative Log-Likelihood (NLL) loss is introduced, which can dynamically adjust the model's ability to handle uncertain regions.

[0103] The uncertainty optimization objective is expressed as:

[0104]

[0105] where is the variance estimate of the model for the input x i corresponding to the uncertainty of the model for this data point, y i is the real data, and μ(x i ) is the mean prediction of the model.

[0106] By minimizing the uncertainty loss, the model can better handle high-noise or highly uncertain regions, enabling the generated data to maintain high credibility in these regions.

[0107] β0, β1, β2, β3, β4, β5 are scalar loss weights used to balance the loss terms.

[0108] In another implementation, step S3 specifically includes: The sequential coupling generator based on the Conditional Random Network (CRN) takes a random noise vector as input and iteratively passes through the internal transformation function across time steps t ∈ {1, 2,..., T} to obtain the synthetic latent embedding representations of continuous and discrete data

[0109] Furthermore, step S6 specifically includes: Processing the input features through convolution operations and pooling operations to obtain the output feature A of the Atrous Spatial Pyramid Pooling (ASPP) i ; Processing the output feature A after average pooling and max pooling through convolution operations using the channel attention factor s1 i ; Processing the original output feature A through a depthwise separable convolution using the spatial attention factor s2 i ; Using the feedback feature selection operation S to multiply the output feature A i by the two attention factors, the channel attention factor s1 and the spatial attention factor s2, to generate the weighted feature R i for weighted selection of the channel and spatial features of the output feature A i .

[0110] In summary, the solution of the present invention realizes the generation of data with high quality and high availability, solves the problem of data shortage in the process of large model training, and can capture the temporal dynamics and correlations between various features. It particularly emphasizes the modeling of uncertainty, including stochastic uncertainty and parameter uncertainty, and jointly models through conditional Gaussian distribution and variational Dropout technology, improving the robustness and credibility of the generated data. Moreover, it can support downstream tasks such as association analysis and result prediction by accurately reconstructing the interdependencies and complex clinical relationships between features. In addition, by introducing multiple loss functions (such as matching loss, contrast loss, and semantic loss), the high quality and diversity of the data are further ensured.

[0111] It should be noted that step S5 is to evaluate the quality of synthetic data generation to obtain high-quality synthetic patient trajectory data, and data utility is used for evaluation here. The data utility metric measures the extent to which the statistical properties of the real (private) data are captured and transferred to the synthetic dataset. The classification metric adopted is CrCl-RS, which involves training on real data and testing on the held-out data from both real and synthetic datasets. This metric is particularly useful for evaluating whether the statistical properties of the actual data are similar to those of the synthetic data. The available real data is divided into a training set and a test set.

[0112] The classifier is trained on the training set (real) and applied to the test set (held real) and synthetic data. The classification performance metric is calculated on both sets. CrCl-RS is defined as the ratio between the performance of the synthetic data and the held real data. The classification performance depends on the selected classifier. Here, due to the discrete nature of the dataset, a decision tree is regarded as the classifier. To perform classification, one of the variables is used as the target, and the remaining variables are used as predictors. This process is repeated for each variable as the target, and the average value is reported. Generally, for these two cross-classification metrics, a value close to 1 is ideal.

[0113] It should also be noted that step S6 is used to extract feature information from the synthetic patient trajectory data generated by the classifier, construct feature information based on the feature information, and feed back the constructed feature information to the data generation network to optimize the weight parameters in the generation network and improve the quality of the generated data, i.e., the synthetic patient trajectory data.

[0114] Exemplarily, referring to Figure 4 the problem of data quality in the feedback mechanism diagram shown, the present invention adopts a feedback method to improve the quality of the generated data (synthetic patient trajectory data). The feedback mechanism in this embodiment takes ASPP as an example.

[0115] Since ASPP uses global information and large receptive field information to assist in describing local semantic information, high-semantic lesion information with a larger feature area is concerned, while low-semantic local noise is suppressed. The channel attention factor and the spatial attention factor respectively use global average pooling and large convolutional kernels to suppress noise, capture long-distance spatial dependencies, and generate selection weights for each channel and each position. The calculations of ASPP, the channel attention factor, and the spatial attention factor are as follows:

[0116]

[0117] s1 = σ(Conv(AvgPool(A i ) + MaxPool(A i )))

[0118] s2 = σ(Conv(A i , k = 7))

[0119] Among them, represents the input feature, A i represents the output feature of ASPP, Conv represents the convolution function with a convolution kernel of 1, Conv(r = 3) represents the dilated convolution with a dilation rate of 3, Conv(r = 6) represents the dilated convolution with a dilation rate of 6, AvgPool represents average pooling, MaxPool represents max pooling, Conv(k = 7) represents the depthwise separable convolution with a convolution kernel of 7, and σ represents the Sigmoid function.

[0120] The feedback feature selection operation S generates the weighted feature R i The calculation process is as follows:

[0121] R i = A i × s1 × s2

[0122] Among them, A i represents the output feature of ASPP, s1 represents the channel attention factor, s2 represents the spatial attention factor, and R i represents the weighted feature.

[0123] It should be understood that the feedback feature selection operation uses the channel attention factor and the spatial attention factor to perform weighted selection on the channel and spatial features of the multi-scale feature A i . The attention factor is learned through the backpropagation method to reduce the classification loss and the regression loss, and its size depends on the effectiveness of the feature for the detection task.

[0124] According to another aspect of the embodiments of the present invention, there is provided a model for implementing the above data generation method based on a generative adversarial network and a generative model, including: a mapping module for jointly mapping continuous data and discrete data to a shared latent space through a generative model network; a learning module for pre-training the generative model network to learn the shared latent space; a generation module for generating a synthetic latent embedding representation of continuous data and discrete data through a sequential coupling generator constructed based on a conditional random network (CRN); a decoding module for inputting the synthetic latent embedding representation into a pre-trained decoder and decoding it into synthetic patient trajectory data in the observation space through the decoder; an evaluation module for separating the uncertainty of the synthetic patient trajectory data and continuing to train and perform data evaluation based on a classifier to obtain high-quality synthetic patient trajectory data; and a feedback module for constructing feature information from the feature information extracted from the synthetic patient trajectory data, feeding back the loss function of the constructed feature information to the data generation network, and optimizing the weight parameters in the data generation network to improve the quality of the synthetic patient trajectory data.

[0125] For the specific model framework diagram of the model in this embodiment, refer to Figure 3 The model in this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0126] In addition, the implementation of the functions of each module in the model of this embodiment can be referred to the description of the corresponding parts in the foregoing method embodiments, which will not be elaborated here either.

[0127] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0128] It should be understood that although this specification is described according to each embodiment, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.

[0129] The above embodiments are preferred implementation manners of the present invention, but the implementation manners of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention shall be defined by the claims.

Claims

1. A data generation method based on a generative adversarial network and a generative model, characterized in that: The following steps are involved: S1, jointly mapping continuous and discrete data into a shared latent space through a generative model network; S2. pre-training the generative model network to learn the shared latent space; S3, generating a synthetic potential embedding representation of the continuous data and the discrete data through a sequential coupling generator constructed based on a conditional random network CRN; S4, inputting the synthesized latent embedding representation into a pre-trained decoder, and decoding it into synthesized patient trajectory data in observation space by the decoder; S5, separating the uncertainty of the synthetic patient trajectory data, and continuing training and performing data evaluation based on the classifier to obtain high-quality synthetic patient trajectory data; S6. Constructing feature information by extracting the characterization information from the synthesized patient trajectory data, feeding back the loss function of the constructed feature information to a data generation network, and optimizing weight parameters in the data generation network to improve the quality of the synthesized patient trajectory data.

2. The method according to claim 1, characterized in that Before step S1, the following steps are also included: The heterogeneous patient data is processed into the continuous data and the discrete data.

3. The method according to claim 1, characterized in that The step S1 specifically includes: We start by training the two encoders of the generative model network, namely Its embedding function is: x C and x D After the data is passed through the two encoders, in the shared latent space H S We get a pair of embedding vectors Then the two domains and The decoder further reconstructs features based on the latent embedding using a mapping function in the opposite direction.

4. The method according to claim 1, characterized in that The step S2 specifically includes: The generative model network is pre-trained using the continuous data and the discrete data to learn the overall objective function of the shared latent space utilization d∈{C,D}: The generative model network consists of a pair of encoders (flow C ,flow D ) and decoder (Inverse C ,Inverse D )composition; The generative model network is used to learn a shared latent space where the embedding representations of the two domains The gap between ld is the overall objective function, is the loss function in the generative model of the domain d∈{C,D}, l Match The representation matching loss function ensures that the low-dimensional latent space can be shared between heterogeneous features, and is combined with the pairwise matching loss to encourage the encoder to minimize the distance within the corresponding representation pair, l Contra Representation contrast loss function, the corresponding auxiliary task for calculating contrast loss is applied to the scenario of learning representations from mixed-type time series, determining whether a set of representations transformed from the observation space belong to the same patient, for N records of patient data, obtain N pairs of potential representations from the encoder in the generative model network, for the patient with index i, h i C and They represent the embedding from the continuous value and discrete value observation space respectively. Due to the symmetric architecture of the generative model network, d and d' are used to represent different domains. The positive sample pair of patient i is represented as i d , the other 2(N-1) samples are used as negative sample pairs, then the contrast loss of the positive sample pair can be defined as The final contrast loss function l Contra for The mean of represents the semantic loss function, l certainty , l uncertainty They represent deterministic and uncertain loss functions respectively.

5. The method according to claim 1, characterized in that The step S3 specifically includes: The sequential coupling generator based on the conditional random network CRN is constructed with random noise vector As input, and iteratively passes through the internal transformation function across time steps t∈{1,2,...,T} to obtain the synthetic latent embedding representation of the continuous data and the discrete data 6. The method according to claim 1, characterized in that: The step S6 specifically includes: Process input features through convolution and pooling operations Get the output feature A of the hollow spatial convolution pooling pyramid ASPP i ; The channel attention factor s1 is used to process the output feature A after average pooling and maximum pooling through convolution operation. i ; The original output feature A is processed by a depthwise separable convolution using the spatial attention factor s2. i ; The feedback feature selection operation S is used to select the output feature A i Multiply it with the channel attention factor s1 and the spatial attention factor s2 to generate the weighted feature R i , to output feature A i The channel and spatial features are weighted for selection.

7. A model for implementing the data generation method based on a generative adversarial network and a generative model as described in any one of claims 1 to 6, characterized in that: include: A mapping module for jointly mapping continuous and discrete data to a shared latent space through a generative model network; A learning module, configured to pre-train the generative model network to learn the shared latent space; A generation module, for generating a synthetic latent embedding representation of the continuous data and the discrete data through a sequentially coupled generator constructed based on a conditional random network (CRN); A decoding module, configured to input the synthesized latent embedding representation into a pre-trained decoder, and decode the synthesized latent embedding representation into synthesized patient trajectory data in an observation space by the decoder; An evaluation module, used to separate the uncertainty of the synthetic patient trajectory data, and continue training and data evaluation based on the classifier to obtain high-quality synthetic patient trajectory data; A feedback module is used to construct feature information by extracting characterization information from the synthetic patient trajectory data, and to feed back the loss function of the constructed feature information to a data generation network to optimize weight parameters in the data generation network to improve the quality of the synthetic patient trajectory data.