Controllable Conditional Clinical Data Synthesis Method and Device Supporting Logical Relationship Perception
By constructing a logical relationship perception model and a conditional diffusion model in electronic clinical data synthesis, the problems of logical inconsistency and lack of controllability in synthetic data are solved, and high-quality electronic clinical data generation and multi-condition control are achieved.
Patent Information
- Application Number
- CN202510322101.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing electronic clinical data synthesis methods are difficult to effectively capture the inherent semantic logical relationships in the data, resulting in logical inconsistency in the synthetic data, and lack effective control of the data generation process, so it is impossible to flexibly synthesize data that meets the combination of multiple user-defined conditions.
A controllable conditional electronic clinical data synthesis method based on diffusion model is proposed. By constructing a logical relationship perception model and a conditional diffusion model, the preprocessed data extracts potential variables and realizes conditional generation in the latent space to ensure the logical consistency and conditional correlation of the synthetic data.
It effectively solves the problem of logical inconsistency in synthetic data, enhances the controllability of the data synthesis process, can efficiently capture and learn the inherent semantic logical relationships in the data, generate high-quality electronic clinical data, and meets multiple user-defined conditions combinations.
Smart Images

Figure CN119851845B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of clinical data synthesis, and particularly relates to a controllable conditional electronic clinical data synthesis method and device based on a diffusion model with logical relationship perception ability. Background Art
[0002] Electronic clinical data is an important carrier of patient health information, usually presented in a structured table form, consisting of columns and rows. Columns represent various attributes of the domain, such as the patient's gender, age, medication records, and various laboratory test indicators, etc., while rows represent specific instances of these attributes. For example, data of patients in the intensive care unit (MIMIC-III), human cancer clinical data (TCGA), and clinical data related to prediabetes and obesity, etc. This digital clinical data can be used to train machine learning models, promoting the application of artificial intelligence technology in fields such as precision medicine, personalized risk prediction, and health trajectory analysis.
[0003] Electronic clinical data synthesis technology is crucial for the medical field. Firstly, most clinical data contains sensitive information and cannot be made public due to privacy protection and social ethics principles. Electronic clinical data synthesis can solve the problem of privacy protection by generating high-quality virtual data without sensitive information to replace the real data that cannot be made public. Secondly, due to the difficulty of the clinical data collection process, there are usually missing values and noises in electronic clinical data, and this technology can be used to complete the missing data and improve the quality of the data set. In addition, in some diseases or specific treatments, the available real electronic clinical data may be very limited. Synthetic data can expand these scarce data sets and improve the effect and generalization ability of model training. Therefore, the current medical field has an increasingly urgent need for high-quality electronic clinical data synthesis technology, and its development is of great significance for promoting scientific and technological progress in the field of medical health.
[0004] In recent years, applying deep learning models to the task of electronic clinical data generation has significantly improved the quality of synthetic data, and the electronic clinical data synthesis scheme based on deep generative models has gradually become the mainstream scheme in this field. However, the difficulty in using deep generative models is that electronic clinical data usually contains mixed heterogeneous data types, including continuous variables (such as numerical data) and discrete variables (such as categorical data), which makes it difficult for deep learning models to learn the joint probability between columns. The standard input format of most deep learning models is continuous data, which poses a major challenge to processing discrete data in clinical data. For the above problems, the commonly used solution is to use digital encoding technology to convert discrete data into digital form to solve this problem, such as one-hot encoding or analog bit encoding. There are also some methods that use different models to process numerical features and categorical features separately, or try to embed the original data into a unified high-dimensional space.
[0005] However, the above solutions may lead to logical inconsistency problems. The logical inconsistency problem refers to the situation where the synthesized clinical data contains data that is logically inconsistent with the real data. For example, using the above solutions may synthesize a sample with a marital status of "never married" but the marital relationship is wrongly marked as "husband". Such a contradictory relationship does not exist in the original dataset, and this phenomenon is called logical inconsistency. In the medical field, this problem is particularly important. In addition to the marital status, there are high dependencies and logical relationships among various vital indicators in clinical data. Even a very small logical inconsistency may lead to very serious consequences. Existing synthesis methods often require a large number of preprocessing operations on the original data, resulting in serious information loss, which limits the model's ability to capture underlying semantic relationships and makes it difficult for the model to learn the correlation relationships between columns. Moreover, current research on clinical data synthesis mainly focuses on improving the performance of downstream classification models or enhancing the effect of privacy protection by synthesizing data, but ignores the importance of maintaining the consistency between data features.
[0006] In addition, most current clinical data synthesis techniques do not support conditional generation, or only use a single feature as a simple conditional input, lacking effective control over the data generation process. However, in reality, it is often necessary to flexibly synthesize data that meets various combinations of user-defined conditions. For example, when solving the problems of data imbalance and data missing, only data that meets the specified conditions needs to be synthesized. Therefore, the function of supporting controllable conditional synthesis is crucial. Summary of the Invention
[0007] To solve the existing technical problems, the present invention proposes a controllable conditional electronic clinical data synthesis method and device based on a diffusion model with logical relationship perception ability, to solve the problem of logical inconsistency in synthesized data and enhance the conditional controllability of the data synthesis process.
[0008] In the first aspect of the present invention, a controllable conditional clinical data synthesis method supporting logical relationship perception is provided, including the following steps:
[0009] Step S1, preprocess numerical variable data and categorical variable data;
[0010] Step S2, construct a logical relationship perception model, and use the preprocessed data to train the logical relationship perception model, and extract latent variables in the data using the trained model;
[0011] The processing process of the logical relationship perception model is as follows: First, use the tokenization module to encode the semantic features in the data into the original D-dimensional semantic vector representation. Then, use the variational autoencoder structure in the logical perception module to obtain the latent variables rich in semantic logical relationships. Subsequently, decode the latent variables to generate the reconstructed D-dimensional vector, and minimize the difference between the reconstructed D-dimensional vector and the original D-dimensional semantic vector. Finally, use the detokenization module to convert the D-dimensional reconstructed vector back to the form of the data; where D is a hyperparameter;
[0012] Step S3, construct a conditional diffusion model, and use the latent variables extracted in step S2 to train the conditional diffusion model;
[0013] Step S4, use the trained conditional diffusion model, given any set of conditional inputs, starting from randomly sampled Gaussian noise, synthesize data that meets the specified conditions;
[0014] Step S5, decode the synthesized data from the latent space back to the real space, convert the synthesized electronic clinical data into a readable and usable form, and complete the entire synthesis process.
[0015] Furthermore, the data preprocessing in step S1 is specifically as follows: For numerical variable data, use a non-linear quantile transformation function to achieve data normalization; for categorical variable data, use textification operations to convert it from the original categorical form to text format.
[0016] Furthermore, in step S2, a differential processing strategy is adopted for the preprocessed numerical variable data and categorical variable data to adapt to the characteristics of different data types. The processing process of the tokenization module is represented by the following formula:
[0017] ,
[0018] where, represents the linear mapping layer. For each textified categorical variable data , use the pre-trained BERT model to extract the semantic features of the text, and finally use a linear structure to encode the semantic features into a D-dimensional vector; for the normalized numerical variable data , use two trainable parameters and to encode the normalized features into a unified D-dimensional vector space, and are the D-dimensional categorical and numerical semantic vectors after tokenization, and are the indices of the numerical variable data and categorical variable data respectively.
[0019] Further, the detokenization module forms a corresponding relationship with the tokenization module, and the processing process is represented by the following formula:
[0020] ,
[0021] wherein, , , and are all learnable parameters, and are respectively a D-dimensional reconstructed numerical vector and a categorical vector, and are respectively the restored numerical vector and the categorical vector; for numerical variable data, a linear decoder is used to restore the vector to ; for categorical variable data, after a linear transformation, a Softmax structure is used to obtain the probabilities of all candidate values, and finally the candidate value with the highest probability is selected.
[0022] Further, the variational autoencoder VAE is a -VAE structure based on Transformer, which consists of an encoder and a decoder. The input received by the encoder is a unified original D-dimensional semantic vector representation, and the output is two vectors: the mean and the variance , and these two output vectors jointly define a Gaussian distribution , wherein represents the latent variable, and this Gaussian distribution represents the probability distribution of the latent variable when the input data is given; the decoder samples from this probability distribution to generate data;
[0023] The encoder includes two groups of Transformer structures. One group is used to generate the mean , and the other group is responsible for generating the log variance . Subsequently, through the reparameterization trick, a latent variable is sampled from the Gaussian distribution. By introducing a random variable of a standard normal distribution, the process of sampling to obtain is realized, that is ; each Transformer structure contains a self-attention layer and a feed-forward layer. The feed-forward layer consists of two linear structures and a ReLU activation layer. The decoder is also an N-layer Transformer, which matches the structure of the encoder, and its task is to process the latent variable Map back to the original data space from the latent space to generate a reconstructed D-dimensional vector.
[0024] Furthermore, the overall loss function of the logical relationship perception model is:
[0025]
[0026] Among them, the specific form of the KL divergence is:
[0027]
[0028] And the reconstruction loss Consists of the following two parts:
[0029]
[0030] Reconstruction loss Quantifies the difference between the reconstructed data And the original data For categorical variable data, the cross-entropy loss function Is used to measure the difference between the predicted value and the actual value, where Represents the original categorical variable data, Represents the reconstructed categorical variable data; for numerical variable data, the mean squared error loss function Is used to evaluate the deviation between the predicted value and the actual value, Represents the original numerical variable data, Represents the reconstructed numerical variable data; The divergence represents the KL divergence between the posterior distribution Of the latent variable And its prior distribution The parameter Is used to balance the relative importance between the reconstruction loss and The divergence.
[0031] Furthermore, the training process of the conditional diffusion model in step S3 includes a forward noise-adding process and a backward denoising process. The forward noise-adding process is to gradually add predefined Gaussian noise to the initial data, that is, the latent variable obtained in step 2, at each time step And gradually transform the data into noise; in the backward denoising process, the original data is recovered from the noise, thus realizing the generation of data.
[0032] Furthermore, the backward denoising process in the conditional diffusion model is implemented through a noise prediction network. The inputs include The noise data at time , the current time step , and the control condition , the output is the predicted noise , the first layer of the noise prediction network is an additive embedding layer, whose function is to embed and integrate the time step information into the input data using the sine position encoding technique; following the additive embedding layer is the multi-head cross-attention layer, and the specific structure of the cross-attention layer is as follows:
[0033] , where
[0034] where represents the output of the cross-attention layer, 、 、 represent query, key, and value respectively; is responsible for embedding conditional semantic information , and receive the intermediate data output by the embedding layer , do not participate in the introduction of conditional information, but achieve the fusion of conditional information through interaction with ; 、 、 are learnable weight matrices used to map the conditional information and the intermediate data to the corresponding 、 、 space, is 、 、 's dimension, is the function used to calculate the attention weights;
[0035] controls the condition is calculated by the feature combination module, and the function of this feature combination module is to extract the semantic features of any conditional combination defined by the user, and the calculation formula is as follows:
[0036]
[0037] where is the D-dimensional semantic vector obtained by the categorical variable through the tokenization module, is the index of the categorical variable, represents 's arbitrary subset, is the number of categorical variables. From this, it can be seen that represents 's sum of elements of any subset, that is, the combination of semantic features of any number of categorical variables.
[0038] Further, the training objective function of the noise prediction network is as follows:
[0039]
[0040] Among them, represents the standard normal distribution, represents the expectation. This objective function will minimize the difference between the actual noise and the predicted noise to achieve the goal of noise prediction.
[0041] The present invention also provides a controllable conditional clinical data synthesis device supporting logical relationship perception, including the following units:
[0042] A data preprocessing unit for preprocessing numerical variable data and categorical variable data;
[0043] A logical relationship perception model construction and training unit for constructing a logical relationship perception model, training the logical relationship perception model using the preprocessed data, and extracting latent variables in the data using the trained model;
[0044] The processing process of the logical relationship perception model is as follows: First, encode the semantic features in the data into the original D-dimensional semantic vector representation using the tokenization module, then obtain the latent variables rich in semantic logical relationships using the variational autoencoder structure in the logical perception module, then decode the latent variables to generate the reconstructed D-dimensional vector, and minimize the difference between the reconstructed D-dimensional vector and the original D-dimensional semantic vector. Finally, use the detokenization module to convert the D-dimensional reconstructed vector back to the form of the data; where D is a hyperparameter;
[0045] A conditional diffusion model construction and training unit for constructing a conditional diffusion model and training the conditional diffusion model using the extracted latent variables;
[0046] A controllable data synthesis unit for using the trained conditional diffusion model, given any set of conditional inputs, starting from randomly sampled Gaussian noise, to synthesize data that meets the specified conditions;
[0047] A data decoding and restoration unit for decoding the synthesized data from the latent space back to the real space, converting the synthesized electronic clinical data into a readable and usable form, and completing the entire synthesis process.
[0048] The beneficial effects that can be brought about by the implementation of the technical solution provided by the present invention are:
[0049] 1. The present invention proposes a logical relationship perception model based on a pre-trained BERT model and a VAE structure, which can not only solve the problem of lossy preprocessing in previous solutions, but also efficiently capture and learn the inherent semantic logical relationships in the original data, overcome the problem of logical inconsistency in electronic clinical data synthesis technology, thereby greatly improving the quality and usability of synthetic data, and providing a set of high-quality data synthesis solutions for the medical field.
[0050] 2. The present invention proposes a controllable conditional generation mechanism based on a diffusion model designed specifically for clinical data, which can skillfully integrate the embeddings of multiple semantic conditions and flexibly generate data that meets various user-defined condition combinations. This mechanism can be widely applied to specified class generation tasks and missing data completion tasks, overcoming the problem of lack of controllability in previous data synthesis methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0052] Figure 1 It is a flowchart of a controllable conditional electronic clinical data synthesis method with logical relationship perception ability proposed by the present invention.
[0053] Figure 2 It is a schematic structural diagram of the logical relationship perception model in the embodiment of the present invention;
[0054] Figure 3 It is a schematic structural diagram of the variational autoencoder model in the embodiment of the present invention;
[0055] Figure 4 It is a flowchart of the working process of the diffusion model in the embodiment of the present invention;
[0056] Figure 5 It is a schematic structural diagram of the noise prediction network model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] In order to elaborate in detail the technical solutions, model construction, implementation objectives and implementation effects of the present invention, the following will describe the implementation methods in detail with reference to the drawings. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0058] The controllable conditional medical electronic clinical data synthesis method based on the diffusion model with logical relationship perception ability provided by the embodiments of the present invention includes two key stages. The logical relationship perception stage: This stage is responsible for extracting semantic feature vectors from the original data and mapping them into a high-dimensional latent space to capture and utilize the inherent semantic information in the original data, enabling the model to learn the logical relationships between columns. The conditional diffusion model generation stage: This stage uses the diffusion model to fit the distribution characteristics of the latent variables extracted in the previous stage and cleverly integrates conditional semantic embeddings, aiming to achieve flexible multi-conditional control generation tasks. Through the mutual cooperation of these two stages, the difficult electronic clinical data generation task is decomposed into two relatively simple problems, efficiently solving the problems of logical inconsistency and lack of controllability. Among them, the logical relationship perception stage mainly includes three core modules, namely the tokenization module, the detokenization module, and the logical perception module.
[0059] As Figure 1 shown, the method of the present invention specifically includes the following steps:
[0060] Step S1, data preprocessing. First, select the original training data from the target dataset and perform preprocessing on it. Specifically, perform normalization operations on the numerical variables in the dataset to eliminate the influence of dimensions and make them comparable on the same scale; perform textification processing on the categorical variables, that is, convert the categorical labels into text forms so that the subsequent model can process them. After these processing steps, the obtained data can be directly used for model training.
[0061] Step S2, construction and training of the logical relationship perception model. Then, construct the logical relationship perception model and use the preprocessed data to train the model. This process aims to enable the model to learn and capture the inherent semantic logical relationships in the data. Subsequently, use the trained model to extract the semantic logical relationships in the data and map these relationships into the latent space.
[0062] Step S3, construction and training of the conditional diffusion model. Further, construct the conditional diffusion model and use the latent variables and control condition information extracted in S2 to train the model. The goal of this step is to enable the model to fully learn the distribution characteristics of the latent variables and the corresponding relationship between the samples and the conditions, thereby ensuring the logical consistency and conditional relevance of the synthetic data.
[0063] Step S4, conditional control data synthesis. Use the trained conditional diffusion model, given any set of conditional inputs, starting from randomly sampled Gaussian noise, to controllably synthesize data that meets the specified conditions. This step reflects the controllable conditional synthesis ability of the present invention, allowing users to generate data according to specific conditions.
[0064] Step S5, data decoding and de-tokenization. Finally, using the decoder of the trained variational autoencoder (VAE) and the de-tokenization module, the generated data is decoded from the latent space back to the real space. This process converts the synthetic electronic clinical data into a readable and usable form, completing the entire synthesis process.
[0065] Specifically, for the preprocessing operation in step S1, since the distribution types of numerical variables and categorical variables are different, different methods are needed to process these two types of data separately. An original data sample can be represented as . Among them represents numerical variables, mostly continuous numbers, such as numerical indicators like age, weight, blood pressure, etc. represents categorical variables, mostly finite discrete data, such as gender, marital status, blood type, etc.
[0066] For numerical features, the present invention adopts a non-linear quantile transformation function to perform normalization processing on the data. The specific steps are as follows: First, by estimating the cumulative distribution function (CDF) of the feature, the original data values are mapped to a uniform distribution. This step can convert the original distribution of the data into a uniform distribution between 0 and 1, thereby eliminating the skewness of the original data distribution. Subsequently, using the quantile function of the inverse normal distribution, the values of the uniform distribution are further mapped to a normal distribution. This process makes the distribution of the data closer to the standard normal distribution, that is, the distribution with a mean of 0 and a standard deviation of 1. For the feature values of new data or unobserved data that exceed the fitting range, they are mapped to the boundary values of the output distribution. This processing method ensures the consistency of data preprocessing and prevents the influence of extreme values on the model. Through this data preprocessing method, the influence of skewness and outliers in the data on the analysis results can be effectively reduced, making the distribution of the data more reasonable, and thus more accurately reflecting the true distribution of the data. This method is widely regarded as an important step in improving the performance and stability of models in the fields of statistical analysis and machine learning.
[0067] For categorical features, a textification operation is required to convert them from the original categorical form to a text format. This textification operation comprehensively considers the header information and context information, converting a single feature into a text representation rich in semantics, which is crucial for subsequent in-depth semantic analysis and feature extraction of the data. The core objective of this step is to ensure that the feature can comprehensively capture and contain the rich semantic information contained in the dataset, laying a foundation for the semantic feature extraction work of the subsequent pre-trained language model.
[0068] Specifically, step S2 is used to construct and train a logical relationship awareness model, and its specific model structure is as follows Figure 2 shown, which mainly includes three core modules, namely the tokenization module, the detokenization module, and the logic awareness module.
[0069] In the process of semantic feature extraction, the tokenization module plays a crucial role. Its main function is to encode the semantic features in the original data into a unified D-dimensional semantic vector representation. Here, D is the dimension of the semantic vector, which can be used as a hyperparameter of the model, and different datasets will have different optimal choices. This module adopts a differential processing strategy for the preprocessed numerical variables and categorical variables to adapt to the characteristics of different data types.
[0070] For the preprocessed numerical variables , the tokenization module maps it to a D-dimensional vector through a linear transformation . As a one-dimensional scalar, the original data is encoded into a high-dimensional space, and its semantic expression ability is significantly enhanced. The specific encoding process follows the following formula:
[0071]
[0072] where and are two learnable parameters. During the model training process, these two parameters will be continuously updated through optimization algorithms such as gradient descent in order to reach the optimal solution of the model objective function. Through this encoding mechanism, the tokenization module can not only effectively convert the original information of numerical variables into high-dimensional semantic vectors, but also gradually improve the model's ability to capture data semantic features by adjusting the learned parameters.
[0073] For the processing of categorical variables, the tokenization module adopts a semantic feature extractor based on the pre-trained BERT model, aiming to extract deep semantic features from the texturized categorical variables and map them into a D-dimensional semantic vector space. The pre-trained BERT model is well-known for its excellent semantic extraction ability and has shown excellent performance in natural language processing fields such as sentiment analysis. The following is the specific structure description of this semantic feature extraction module:
[0074] Let represent the text variable of the categorical variable after the texturization operation . Input into the BERT model, and extract the [CLS] token of this text, which contains the global semantic features of the entire text. Subsequently, map this feature to a unified D-dimensional vector through a linear layer . The specific calculation process can be expressed as:
[0075]
[0076] Among them, represents the textification operation. represents the semantic feature extraction process of the pre-trained BERT model, represents the [CLS] token extracted by BERT. represents the linear mapping layer, which is responsible for compressing the high-dimensional semantic feature vector output by the BERT model into a D-dimensional space. Through this structure, the tokenization module can effectively extract rich semantic information from categorical variables, providing strong support for the model to deeply understand the semantic structure of the data.
[0077] After obtaining the unified dimensional vector representation of all data, the logical perception module uses the variational autoencoder (VAE) structure to capture the complex logical relationships between data columns. The specific model architecture of the VAE is as Figure 3 shown, mainly consisting of two major parts: an encoder and a decoder. The output of the encoder is a probability distribution (usually a Gaussian distribution), and the decoder samples from this distribution to generate data.
[0078] The function of the encoder is to map the input data into the latent space. Specifically, the input received by the encoder is the unified D-dimensional semantic vector and . Its output is two vectors: the mean and the variance . These two output vectors jointly define a Gaussian distribution , which represents the probability distribution of the latent variable given the input data . Here, represents the latent variable, represents all the D-dimensional semantic vectors, represents a Gaussian distribution with a mean of and a variance of .
[0079] The encoder proposed in the present invention consists of N Transformer structures. Each Transformer structure contains a self-attention layer and a feed-forward layer. The feed-forward layer consists of two linear structures with a dimension of 128 and a ReLU activation layer. There are two groups of Transformers in the encoder, one group is used to generate the mean , and the other group is responsible for generating the log variance .
[0080] Subsequently, through the reparameterization trick, from the above Gaussian distribution Sample a latent variable . Specifically, first introduce a random variable from a standard normal distribution , and then calculate , thus realizing the process of sampling from to obtain . is the latent space variable, which contains deep logical relationships. Mapping the original data to the latent space variable can enhance the semantic expression ability of the model and enable the model to fully capture the semantic logical relationships between features.
[0081] The task of the decoder is to map the latent variable back to the original data space to generate the reconstructed D-dimensional vectors and . That is, the input of the decoder is , and the output is and . The goal of the decoder is to minimize the difference between the reconstructed D-dimensional vectors and and the original D-dimensional vectors and . The decoder proposed in the present invention is also an N-layer Transformer, which is matched with the structure of the encoder. Through this structure, the VAE model can not only learn the distribution characteristics of the data, but also capture and learn the complex logical relationships between data columns, providing strong support for subsequent data analysis and applications.
[0082] Further, the role of this detokenization module is to convert the D-dimensional reconstructed vectors and back to the form of the original data. This module forms a corresponding relationship with the tokenization module and adopts different processing methods for numerical variables and categorical variables respectively.
[0083] For numerical variables, the detokenization module will use a linear transformation to restore it from the D-dimensional reconstructed vector to a one-dimensional scalar. The specific decoding formula is as follows:
[0084]
[0085] Wherein, and are both learnable parameters of the linear module, and these parameters are adjusted through optimization algorithms during the model training process to ensure the accuracy of reconstruction.
[0086] For categorical variables, the tokenization module also uses a linear encoding to decode it from the D-dimensional reconstructed vector to a -dimensional vector represents the original categorical variable The number of candidate values. Then, this dimensional vector is input into the Softmax activation function to obtain this predicted probability distribution of the candidate values. Finally, the candidate value with the highest probability is selected to restore the original data form. The specific decoding function is as follows:
[0087]
[0088] where and are both learnable parameters. These parameters are also optimized during the model training process to improve the classification accuracy.
[0089] After the detailed elaboration, we have fully described the specific structure of the logical relationship perception model. After constructing the required model, the next step is to use the preprocessed data obtained in step S1 to train the model. Based on this, the overall loss function in the logical relationship perception stage is defined as follows:
[0090] ,
[0091] where the reconstruction loss consists of the following two parts:
[0092]
[0093] The reconstruction loss quantifies the difference between the reconstructed data and the original data For categorical features, the cross-entropy loss function is used to measure the difference between the predicted value and the actual value; for numerical features, the mean squared error loss function is used to evaluate the deviation between the predicted value and the actual value.
[0094] The specific form of the KL divergence is:
[0095]
[0096] The divergence represents the KL divergence between the posterior distribution of the latent variable and its prior distribution The parameter is used to balance the relative importance between the reconstruction loss and the divergence, so as to balance between the model's fitting degree and the continuity of the latent space.
[0097] After sufficient training following the minimization principle of the loss function definition, the logical perception model will be able to effectively learn and capture the inherent logical relationships between features. Through this training process, the model can not only identify explicit associations in the data but also infer implicit and deep semantic connections. Using the trained model, we encode the original data into latent space variables containing rich semantic logical relationships. This encoding process maps the data from its original high-dimensional space to a low-dimensional, dense latent representation space, where each latent variable contains the complex semantic information and logical structure of the data.
[0098] Specifically, step S3 is used to construct and train a conditional diffusion model. This model is responsible for fitting the latent variables obtained from the previous stage, while fully learning the distribution characteristics of the latent variables and the correspondence between samples and conditions, and realizing conditional generation in the latent space. The training process of the diffusion model includes a forward noise addition process and a reverse denoising generation process, as Figure 4 shown.
[0099] First, we denote the latent space variables obtained in step S2 as , where the subscript of represents the initial time step , that is, represents the initial latent variable. Assume .
[0100] The forward noise addition process involves gradually introducing predefined Gaussian noise to the initial data at each time step . This process starts from the initial data without noise , and then gradually adds noise to it, successively obtaining and other intermediate data, where represents the pure noise data obtained at the final time step . This continuous noise addition process simulates the evolution of data from a deterministic state to a random state, providing the necessary dynamic structure for the reverse denoising process of the diffusion model, thus allowing the model to gradually recover the structure and features of the original data from pure noise. The noise addition formula from time step to time step is as follows:
[0101]
[0102] where is a predefined hyperparameter, represents Gaussian noise sampled from the standard normal distribution . Its function is to adjust the proportional relationship between Gaussian noise and the data of the previous state. By adjusting , the model can finely control the rate of noise introduction. Write the forward noise addition formula in the form of a probability transition kernel:
[0103]
[0104] This formula means that given the previous state , then the current state follows a normal distribution with a mean of and a variance of . Among them, represents the identity matrix.
[0105] In the reverse process of the diffusion model, the model tries to reverse this noise addition process, gradually denoising from a completely noisy state to recover the original data and achieve data generation. The reverse denoising process involves training a Gaussian transition kernel to achieve the goal of gradual denoising:
[0106]
[0107] This formula means that given the previous state , then the current state follows a normal distribution with a mean of and a variance of . Specifically:
[0108] ,
[0109] Among them, , . The parameter is the parameter of the noise prediction network of the diffusion model. is the output of this noise prediction network, that is, the predicted noise added to the data of the previous state. The input of this noise prediction network includes the noise data at time , the current time step , and the control condition , and the output is the predicted noise . By training the noise prediction network and obtaining its predicted output , we can calculate and . Furthermore, according to the reverse transition kernel formula , the model can gradually remove noise from pure Gaussian noise and achieve the data generation task.
[0110] The training objective function of this noise prediction network is as follows:
[0111]
[0112] This objective function minimizes the difference between the actual noise and the predicted noise to achieve the goal of noise prediction.
[0113] This noise prediction network is responsible for predicting the noise added to the data of the previous state , and its model structure is the core part of the diffusion model and the key to supporting conditional generation. Its specific model structure is as Figure 5 shown. The present invention proposes an innovative noise prediction network structure based on the multi-head cross-attention mechanism, which ingeniously integrates conditional semantic embeddings.
[0114] The first layer of the noise prediction network model is an additive embedding layer, whose function is to embed and integrate the time step information into the input data using the sine positional encoding technique. Specifically, this layer encodes the time step into an embedding vector and adds it to the corresponding dimension of the input data, thereby providing the time step information for the model. The specific calculation formula of this embedding layer is as follows:
[0115]
[0116] where represents the sine positional encoding, is responsible for mapping the sine-encoded data to the specified dimension. Specifically, the sine positional encoding is achieved through the following formula:
[0117] ,
[0118] where represents the time step, is the index of the embedding dimension, and is the dimension of the embedding vector. This encoding method can not only capture the periodic changes of the time step but also provide multi-scale feature representations, helping the model make more accurate denoising decisions during the generation process.
[0119] Following the additive embedding layer, a multi-head cross-attention layer is introduced into the model architecture. This layer is the core component for effectively fusing conditional information. The multi-head cross-attention mechanism allows the model to exhibit higher flexibility and efficiency when processing conditional information. In this structure, the query sequence can effectively incorporate the semantic conditions Embed it into the model to ensure that the model can deeply understand and utilize the input conditional information, and achieve feature alignment between the conditional information and the data samples. The specific structure is as follows:
[0120] , where
[0121] where represents the output of the cross-attention structure, and and represent Query, Key, and Value respectively. is responsible for embedding conditional semantic information . and receive the intermediate data of the model , do not participate in the introduction of conditional information, but achieve the fusion of conditional information through interaction with . and and are learnable weight matrices for mapping conditional information and intermediate data to the corresponding and and space. is and and 's dimension, is the function used to calculate the attention weights. The design of the multi-head cross-attention layer enables the model to flexibly incorporate conditional control during generation, provides the model with the ability to capture conditional information at different scales, enhances the model's sensitivity to conditional changes, and improves the accuracy and applicability of data synthesis.
[0122] Specifically, the conditional semantic information is calculated by the feature combination module. The function of this module is to extract the semantic features of any conditional combination defined by the user. The calculation formula of this module is as follows:
[0123]
[0124] where is the D-dimensional semantic vector obtained by the categorical variable through the tokenization module, represents 's arbitrary subset. It can be seen from this that represents The sum of the elements of any subset, that is, the combination of the semantic features of any number of categorical variables. Through this method, any combination of categorical variables can be specified as a condition to guide and control the synthesis process, so as to generate data containing the specified features.
[0125] The above outlines the model structure and training method proposed by the present invention, which includes two core training stages, namely the logical relationship perception stage and the conditional diffusion model generation stage. There is an objective function to be trained in both stages, which are responsible for different functions respectively. After the model training of the above two stages is completed, the clinical data generation task can be carried out.
[0126] Further, step S4 is responsible for data generation. In this step, the user can set a specific combination of conditions and use the conditional diffusion model trained in step S3, starting from the Gaussian noise sampled randomly as the starting point, following the trained denoising transfer kernel , and following the direction of conditional control, gradually denoise , and finally specifically synthesize data that meets the preset conditions .
[0127] Finally, step S5 uses the decoder of the trained variational autoencoder (VAE) and the de-tokenization module to decode the generated data from the latent space back to the real space data and . This process converts the synthesized electronic clinical data into a readable and usable form, completing the entire synthesis process. Thus, we have successfully achieved the task of generating conditional electronic clinical data, ensuring the consistency of the synthesized data set in structure and semantics with real-world data, and providing high-quality data resources for subsequent data analysis and application research.
[0128] To comprehensively evaluate the effectiveness and practicality of the electronic clinical data synthesis method proposed by the present invention, a series of experimental tests are carried out below, and these tests include two levels: qualitative analysis and quantitative evaluation.
[0129] Qualitative test results: The experimental verification of the electronic clinical data synthesis method proposed by the present invention reveals positive results. The experimental results show that the proposed model performs excellently in capturing the logical dependencies between data columns, can significantly reduce the problem of logical inconsistency in the synthesized data, thus ensuring the reliability and effectiveness of the synthesized data. In addition, the experimental results further confirm that the conditional control mechanism introduced in the present invention exhibits excellent matching accuracy and can efficiently produce data that meets specific condition combinations. These findings fully prove the effectiveness and practicality of this method in synthesizing electronic clinical data that meets specific conditions.
[0130] Quantitative test analysis: To evaluate the performance of the electronic clinical data synthesis method proposed in this invention, quantitative tests were conducted on the clinical obesity analysis dataset and comparative analysis was carried out with three most widely used current methods. These three methods are respectively based on different generative model frameworks: the data synthesis scheme TVAE based on variational autoencoder (VAE), the data synthesis scheme CTGAN based on generative adversarial network (GAN), and the data synthesis scheme TabDDPM based on diffusion model. Two key evaluation metrics were adopted in the experiment: F1 Score is used to quantitatively evaluate the classification performance of the classification model on the real test dataset after training on the synthetic dataset. The higher the value of this metric, the better the performance. The Correlation Coefficient Difference (CCD) measures the consistency of the logical relationship between the synthetic data and the real data. The lower the value of this metric, the better the consistency. The data in Table 1 clearly show that the method proposed in this study performs excellently in improving the classification effect of the downstream classification model, confirming its remarkable practicality. At the same time, it also shows excellent performance in maintaining the consistency of the logical relationship between the synthetic data and the real data, confirming its high effectiveness. These quantitative results provide strong empirical support for the effectiveness and practicality of this method.
[0131] Table 1 Synthesis effects of multiple synthesis schemes on the clinical obesity analysis dataset
[0132]
[0133] This invention also provides a controllable conditional electronic clinical data synthesis device based on diffusion model with logical relationship perception ability, including the following units:
[0134] Data preprocessing unit, used to preprocess numerical variable data and categorical variable data;
[0135] Logical relationship perception model construction and training unit, used to construct a logical relationship perception model, train the logical relationship perception model with the preprocessed data, and extract latent variables in the data using the trained model;
[0136] The processing process of the logical relationship perception model is as follows: First, use the tokenization module to encode the semantic features in the data into the original D-dimensional semantic vector representation, then use the variational autoencoder structure in the logical perception module to obtain latent variables rich in semantic logical relationships, then decode the latent variables to generate a reconstructed D-dimensional vector, and minimize the difference between the reconstructed D-dimensional vector and the original D-dimensional semantic vector. Finally, use the detokenization module to convert the D-dimensional reconstructed vector back to the form of data; where D is a hyperparameter;
[0137] A conditional diffusion model construction and training unit, which is used to construct a conditional diffusion model and train the conditional diffusion model using the extracted latent variables;
[0138] A controllable data synthesis unit, which is used to utilize the trained conditional diffusion model, given any set of conditional inputs, and synthesize data that meets the specified conditions starting from randomly sampled Gaussian noise;
[0139] A data decoding and restoration unit, which is used to decode the synthesized data from the latent space back to the real space, convert the synthesized electronic clinical data into a readable and usable form, and complete the entire synthesis process.
[0140] The specific implementation methods of each module are the same as the above steps and will not be elaborated here.
[0141] The above has elaborated in detail different embodiments of the present invention. It should be clear that these descriptions are exemplary in nature, not exhaustive, and at the same time, they should not be construed as the only limitations to which the present invention may be applied. Obviously, those skilled in the art of this technology can make various modifications and adjustments to the present invention without departing from the core scope and spirit of the present invention. If these modifications and adjustments fall within the scope of the claims of the present invention and their equivalent technologies, they should also be regarded as part of what the present invention intends to protect.
Claims
1. A controllable conditional clinical data synthesis method supporting logical relationship perception, characterized in that: The steps include: Step S1, preprocessing the digital variable data and the categorical variable data; Step S2, constructing a logical relationship perception model, and using the preprocessed data to train the logical relationship perception model, and using the trained model to extract potential variables in the data; The processing process of the logical relationship perception model is as follows: first, the semantic features in the data are encoded into the original D-dimensional semantic vector representation using the tokenization module, then the latent variables rich in semantic logical relations are obtained using the variational autoencoder structure in the logical perception module, then the latent variables are decoded to generate a reconstructed D-dimensional vector, and the difference between the reconstructed D-dimensional vector and the original D-dimensional semantic vector is minimized, and finally the D-dimensional reconstructed vector is converted back to the form of data using the de-tokenization module; where D is a hyperparameter; Step S3, constructing a conditional diffusion model, and using the latent variables extracted in step S2 to train the conditional diffusion model; Step S4, using the trained conditional diffusion model, given any set of conditional inputs, starting from randomly sampled Gaussian noise, synthesizing data that meets the specified conditions; In step S5, the synthesized data is decoded from the latent space back to the real space, and the synthesized electronic clinical data is converted into a readable and usable form, completing the entire synthesis process.
2. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 1, characterized in that: The data preprocessing in step S1 is specifically as follows: for numerical variable data, a nonlinear quantile transformation function is used to achieve data normalization; for categorical variable data, a text operation is used to convert it from the original classification form to a text format.
3. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 1, characterized in that: In step S2, a differentiated processing strategy is adopted for the preprocessed digital variable data and categorical variable data to adapt to the characteristics of different data types. The processing process of the tokenization module is expressed by the following formula: , in, Represents the linear mapping layer, for each text-based categorical variable data , use the pre-trained BERT model to extract the semantic features of the text, and finally use a linear structure to encode the semantic features into a D-dimensional vector; for the normalized digital variable data , using two trainable parameters and Encode the normalized features into a unified D-dimensional vector space. and That is, the D-dimensional classification type and numerical semantic vector after tokenization. and They are the indexes of numeric variable data and categorical variable data respectively.
4. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 3, characterized in that: The de-tokenization module forms a corresponding relationship with the tokenization module, and the processing process is expressed by the following formula: , in, , , and are all learnable parameters. and They are D-dimensional reconstructed numerical vectors and categorical vectors, and are the restored numerical vectors and categorical vectors respectively; for digital variable data, a linear decoder is used to convert the vector Restore to ; For categorical variable data, after using a linear transformation, a Softmax structure is used to obtain the probabilities of all candidate values, and finally the candidate value with the largest probability is selected.
5. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 1, characterized in that: Variational Autoencoder VAE is based on Transformer -VAE structure, consisting of an encoder and a decoder. The encoder accepts a unified original D-dimensional semantic vector representation as input and outputs two vectors: mean and variance , these two output vectors together define a Gaussian distribution ,in represents the latent variable, and the Gaussian distribution represents the given input data When the latent variable The probability distribution of The decoder samples and generates data from this probability distribution; The encoder consists of two sets of Transformer structures, one for generating the mean , and the other group is responsible for generating the log variance , then, by the reparameterization technique, a latent variable is sampled from the Gaussian distribution , by introducing a standard normally distributed random variable , to achieve Sampling in process, that is ; Each Transformer structure contains a self-attention layer and a feedforward layer. The feedforward layer consists of two linear layers and a ReLU activation layer. The decoder is also an N-layer Transformer, matching the structure of the encoder. Its task is to transform the latent variables Mapping from the latent space back to the original data space to produce a reconstructed D-dimensional vector.
6. The controllable conditional clinical data synthesis method supporting logical relationship perception as claimed in claim 5, characterized in that: The overall loss function of the logical relationship perception model is: Among them, the specific form of KL divergence is: The reconstruction loss It consists of the following two parts: Reconstruction loss Quantified reconstruction data and the original data For categorical variable data, the cross entropy loss function is used to measure the difference between the predicted value and the actual value, where Represents the original categorical variable data, Represents the reconstructed categorical variable data; for numerical variable data, the mean square error loss function is used To evaluate the deviation between the predicted value and the actual value, Represents the original numeric variable data, Represents the reconstructed numeric variable data; Divergence indicates latent variables The posterior distribution of With its prior distribution The KL divergence between Used to balance the reconstruction loss and The relative importance of divergences.
7. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 1, characterized in that: The training process of the conditional diffusion model in step S3 includes a forward denoising process and a reverse denoising process. The forward denoising process is performed at each time step. In the process of de-noising, predefined Gaussian noise is gradually added to the initial data, i.e., the latent variables obtained in step 2, gradually converting the data into noise; in the reverse denoising process, the original data is restored from the noise, thereby realizing data generation.
8. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 1, characterized in that: The inverse denoising process in the conditional diffusion model is implemented through a noise prediction network, and the input includes Noise data at time , the current time step , and control conditions , the output is the predicted noise The first layer of the noise prediction network is an additive embedding layer, whose function is to embed the time step information into the input data using the sinusoidal position encoding technique; following the additive embedding layer is a multi-head cross attention layer, and the specific structure of the cross attention layer is as follows: , in, in, represents the output of the cross attention layer, , , Represents query, key and value respectively; Responsible for embedding conditional semantic information , and Receive the intermediate data output by the embedding layer , does not participate in the introduction of conditional information, but through The interaction of can realize the fusion of conditional information; , , is a learnable weight matrix used to transform conditional information and intermediate data Map to the corresponding , , space, yes , , The dimension of is the function used to calculate the attention weight; Control conditions It is calculated by the feature combination module, which is used to extract the semantic features of any combination of conditions defined by the user. The calculation formula is as follows: in, is the D-dimensional semantic vector obtained by the tokenization module for the categorical variable, is the index of the categorical variable, express Any subset of is the number of categorical variables, from which we can see that express The sum of the elements of any subset of , that is, the combination of the semantic features of any number of categorical variables.
9. The controllable conditional clinical data synthesis method supporting logical relationship perception according to claim 8, characterized in that: The training objective function of the noise prediction network is as follows: in, represents the standard normal distribution, represents the expectation, and the objective function minimizes the actual noise and prediction noise The difference between them can achieve the goal of noise prediction.
10. A controllable conditional clinical data synthesis device supporting logical relationship perception, characterized in that: The following units are included: A data preprocessing unit, used for preprocessing digital variable data and categorical variable data; A logical relationship perception model construction and training unit is used to construct a logical relationship perception model, train the logical relationship perception model using preprocessed data, and extract potential variables in the data using the trained model; The processing process of the logical relationship perception model is as follows: first, the semantic features in the data are encoded into the original D-dimensional semantic vector representation using the tokenization module, then the latent variables rich in semantic logical relations are obtained using the variational autoencoder structure in the logical perception module, then the latent variables are decoded to generate a reconstructed D-dimensional vector, and the difference between the reconstructed D-dimensional vector and the original D-dimensional semantic vector is minimized, and finally the D-dimensional reconstructed vector is converted back to the form of data using the de-tokenization module; where D is a hyperparameter; A conditional diffusion model construction and training unit, used to construct a conditional diffusion model and train the conditional diffusion model using the extracted latent variables; A controllable data synthesis unit is used to use the trained conditional diffusion model to synthesize data that meets the specified conditions from randomly sampled Gaussian noise given any set of conditional inputs; The data decoding and restoration unit is used to decode the synthesized data from the latent space back to the real space, convert the synthesized electronic clinical data into a readable and usable form, and complete the entire synthesis process.
Citation Information
Patent Citations
Method and device for generating synthetic data
CN117493325A
Eye fundus image quality enhancement method and system based on conditional diffusion model
CN117830131A