Digital twin drug data analysis system

By employing a digital twin-based drug data analysis system with adaptive calibration and lightweight generation fusion mechanisms, the system addresses the challenges of integrating individual differences and heterogeneous multi-source data, enabling personalized and scientific prediction of drug responses and providing individualized regulatory strategies.

CN121031367AActive Publication Date: 2025-11-28SHULIJU (HAINAN) TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511530886.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-11-28
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing drug response prediction technologies lack the ability to dynamically calibrate and predict individual differences, are difficult to integrate multi-source heterogeneous drug data, and are difficult to systematically integrate knowledge of drug response mechanisms, resulting in a lack of interpretability and scientific validity in the prediction results.

Method used

A drug data analysis system employing digital twins achieves dynamic adjustment and updating of individual-level drug response trends through adaptive calibration and lightweight generation fusion mechanisms, constructs a unified drug data foundation, and integrates pharmacological and clinical prior knowledge into a multi-layered digital twin model for causal analysis.

Benefits of technology

This has improved the accuracy and interpretability of individualized drug response prediction, enhanced the utilization efficiency of multi-source heterogeneous data and the scientific validity of prediction results, and provided a basis for individualized regulatory strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121031367A_ABST
    Figure CN121031367A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence medical treatment, in particular to a digital twinning drug data analysis system which comprises a data fusion module, a multi-layer digital twinning modeling module, a cross-domain causal and mechanism fusion module and an optimization decision module. The dynamic adjustment and updating of the drug reaction trend of the individual layer are realized, so that the drug reaction prediction of each individual is more accurate and personalized; a data fusion module is adopted to carry out standardization, semantic mapping and feature alignment on data of a molecular layer, an individual layer, a group layer and a production quality layer, a unified drug data base is constructed, the utilization efficiency of multi-source heterogeneous data is improved, and a reliable basis is provided for digital twin modeling; through a cross-domain causal and mechanism fusion module, pharmacology and clinical priori knowledge is integrated into a multi-layer digital twinborn model, drug-target-pathway-clinical outcome causal analysis is realized, and the scientificity and interpretability of a prediction result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence medical technology, specifically to a digital twin drug data analysis system. Background Technology

[0002] With the rapid development of artificial intelligence in the medical field, a large amount of multi-level, multi-source heterogeneous data has been accumulated in the process of drug development and clinical use, including drug molecule information, individual clinical data, population trial data, and production quality data. However, existing drug response prediction technologies still have the following problems:

[0003] 1. Individual differences lead to unpredictable drug response: Traditional methods mainly rely on population average data to predict drug response, which lacks the ability to dynamically calibrate and predict individual differences, making it difficult to accurately reflect the drug response trend of each individual.

[0004] 2. Difficulty in integrating multi-source heterogeneous drug data: Existing technologies lack a unified data standardization and feature alignment mechanism when processing data at the molecular, individual, population, and production quality levels, resulting in low data utilization efficiency and affecting prediction accuracy;

[0005] 3. Difficulty in systematically integrating knowledge of drug response mechanisms: Existing methods struggle to effectively integrate pharmacological knowledge with prior clinical information, resulting in a lack of interpretability and scientific validity in drug response predictions, which limits their application in personalized medicine. Summary of the Invention

[0006] To address the above issues and overcome the shortcomings of existing technologies, this invention provides a digital twin drug data analysis system. Addressing the problem of unpredictable drug responses due to individual differences, it introduces an adaptive calibration mechanism and a lightweight generation fusion mechanism to dynamically adjust and update individual-level drug response trends, making drug response predictions more accurate and personalized for each individual. To address the difficulty of integrating multi-source heterogeneous drug data, a data fusion module standardizes, semantically maps, and aligns features of data at the molecular, individual, population, and production quality levels, constructing a unified drug data foundation and improving the utilization efficiency of multi-source heterogeneous data, providing a reliable basis for digital twin modeling. Furthermore, to address the difficulty of systematically integrating knowledge of drug response mechanisms, a cross-domain causal and mechanism fusion module integrates pharmacological and clinical prior knowledge into a multi-layered digital twin model, enabling causal analysis of drug-target-pathway-clinical outcome, improving the scientific validity and interpretability of prediction results, and providing a basis for personalized regulatory strategies.

[0007] The technical solution adopted in this invention is as follows: This invention provides a digital twin drug data analysis system, including a data fusion module, a multi-layer digital twin modeling module, a cross-domain causal and mechanism fusion module, and an optimization decision-making module, specifically including the following:

[0008] The data fusion module collects drug data, including drug-level data, individual-level data, group-level data, and production quality-level data. It processes the drug data using semantic mapping and feature alignment techniques to obtain a unified drug data foundation.

[0009] The multi-layer digital twin modeling module constructs multi-layer digital twins based on a unified drug data foundation, including molecular-level twins, individual-level twins, and population-level twins.

[0010] The cross-domain causality and mechanism fusion module introduces pharmacological and clinical prior knowledge into a multi-level digital twin to construct a causal structure of drug-target-pathway-clinical outcome. It then uses causal inference and counterfactual analysis methods to analyze the causal inference results.

[0011] The optimization decision-making module, based on multi-level digital twins and causal inference results, introduces a multi-objective method and outputs individualized control strategies.

[0012] Furthermore, in the data fusion module, the drug data includes drug-level data, individual-level data, population-level data, and production quality-level data, specifically including the following:

[0013] Drug-level data: drug molecular structure, physicochemical properties, pharmacokinetic parameters, and pharmacodynamic parameters;

[0014] Individual-level data: genomics, proteomics, clinical testing data, medical imaging data, pathological status, genetic background, and environmental factors;

[0015] Population-level data: laboratory testing data, clinical trial data, and clinical drug use data;

[0016] Production quality data: pharmaceutical manufacturing process and quality control data.

[0017] Furthermore, the multi-layer digital twin modeling module constructs a multi-layer digital twin based on a unified drug data foundation, specifically including the following steps:

[0018] Step S1: Molecular-level feature extraction. Processed drug-level data is retrieved from a unified drug data base. Combining molecular dynamics simulation, graph neural networks, and feature encoding methods, key molecular-level action features are extracted, including binding affinity, metabolic pathway parameters, and toxicological indicators.

[0019] Step S2: Molecular-level twin modeling. Based on the key role characteristics of the molecular level, and combining quantum chemical calculations and deep learning hybrid models, a molecular-level digital twin is constructed to simulate the drug-target interaction mechanism.

[0020] Step S3: Individual data representation. Individual-level data is retrieved from a unified drug data foundation. Multimodal representation learning and mixed-effects models are used to model individual differences and obtain individual-level drug response characteristics.

[0021] Step S4: Individual-level twin modeling. Based on individual-level drug response characteristics, construct individual-level digital twins, introduce adaptive calibration and lightweight generation fusion mechanisms, and obtain dynamically updated drug response predictions.

[0022] Step S5: Population data aggregation. Population-level data is retrieved from a unified drug data base. Hierarchical Bayesian modeling and multi-density clustering methods are used to identify differences in drug use and efficacy patterns among populations.

[0023] Step S6: Population-level twin modeling. Based on the individual-level twins, construct a population-level digital twin to enable long-term drug efficacy assessment, epidemiological trend prediction, and public health strategy support.

[0024] Furthermore, step S4 specifically includes the following steps:

[0025] Step S41: Call up the individual-level drug response features, standardize and screen the individual-level drug response features using multidimensional features, remove redundant features and enhance key effect factors to obtain the processed drug response features;

[0026] Step S42: Preliminary digital twin generation. Based on the processed drug response characteristics, a preliminary digital twin is constructed. Preliminary predictions are made for each feature dimension to generate an individual-level drug response trend sequence.

[0027] Step S43: Introduce an adaptive calibration mechanism. Based on the initial digital twin, introduce a recurrent convolutional long short-term memory network to construct an adaptive calibration module. Gradually correct the predicted drug response trend sequence to obtain the calibrated drug response characteristics. The calibrated twin model is obtained through the adaptive calibration mechanism and used as the adaptive calibration result.

[0028] Step S44: Introduce a lightweight generation fusion mechanism. In the calibrated twin model, construct a lightweight generation fusion module and introduce a dual-branch architecture of candidate response construction unit and response effect evaluation unit. The candidate response construction unit generates candidate response sequences based on the calibrated drug response characteristics, and the response effect evaluation unit performs discriminative feedback to obtain the lightweight generation fusion result.

[0029] Step S45: Digital twin optimization, fusing adaptive calibration results with lightweight generation fusion results, performing final optimization on the calibrated twin model, and outputting dynamically updated drug response predictions, including the ADME process of the drug in vivo and adverse reaction dynamics.

[0030] Furthermore, step S43 specifically includes the following steps:

[0031] Step S431: Sequence difference calculation. Collect historical individual-level drug response trend sequences, combine them with the individual-level drug response trend sequences, construct a difference function, and calculate the residual values. The formula used is as follows:

[0032] ;

[0033] in, Represents the residual value. This represents a sequence of individual-level drug response trends. This represents a historical individual-level trend sequence of drug response. Indicates a time index;

[0034] Step S432: Adaptive weighted fusion. Adaptive weights are introduced to fuse the historical individual-level drug response trend sequence and the individual-level drug response trend sequence to obtain the weighted individual-level drug response trend sequence. The formula used is as follows:

[0035] ;

[0036] in, This represents the obtained weighted individual-level drug response trend sequence. Indicates adaptive weights;

[0037] Step S433: Gated calibration update. Based on the residual values ​​calculated in step S431 and the weighted individual-level drug response trend sequence obtained by fusion in step S432, a gated calibration unit is constructed. The gate parameters are dynamically adjusted in conjunction with the residual values ​​to update the weighted individual-level drug response trend sequence. The formula used is as follows:

[0038] ;

[0039] in, This represents the calibrated individual-level drug response trend sequence. Indicates the gating parameters;

[0040] Step S434: Calibrate the twin output, set the residual threshold, gradually converge the initial digital twin, repeat step S433 iteratively until the residual value converges to the residual threshold, output the calibrated individual-level drug response trend sequence, map to generate the calibrated drug response features, and output the calibrated twin model obtained through the adaptive calibration mechanism.

[0041] Furthermore, step S44 specifically includes the following steps:

[0042] Step S441: Candidate response sequence construction. Based on the calibrated drug response features output in step S434, candidate response sequences are constructed. A lightweight generation network is used to generate the sequences, and the formula used is as follows:

[0043] ;

[0044] in, Candidate reaction sequences, To create lightweight generative networks, The calibrated drug response characteristics;

[0045] Step S442: Response effect assessment, constructing a response effect assessment unit, discriminating and feeding back candidate response sequences, and obtaining sequence discrimination feedback results;

[0046] Step S443: Lightweight generation fusion, combining the candidate reaction sequence and the sequence discrimination feedback result to obtain the candidate reaction sequence after lightweight generation fusion, which is used as the lightweight generation fusion result.

[0047] The beneficial effects achieved by the present invention using the above solution are as follows:

[0048] (1) To address the problem of unpredictable drug response caused by individual differences, an adaptive calibration mechanism and a lightweight generation fusion mechanism are introduced to achieve dynamic adjustment and updating of drug response trends at the individual level, making drug response predictions for each individual more accurate and personalized.

[0049] (2) To address the difficulty of integrating multi-source heterogeneous drug data, a data fusion module is used to standardize, semantically map, and align features of data at the molecular, individual, population, and production quality levels, thereby constructing a unified drug data foundation, improving the utilization efficiency of multi-source heterogeneous data, and providing a reliable foundation for digital twin modeling.

[0050] (3) To address the problem of the difficulty in systematically integrating knowledge of drug response mechanisms, a cross-domain causal and mechanism fusion module is used to integrate pharmacological and clinical prior knowledge into a multi-layer digital twin model, thereby realizing causal analysis of drug-target-pathway-clinical outcome, improving the scientificity and interpretability of prediction results, and providing a basis for individualized regulation strategies. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of a digital twin drug data analysis system proposed in this invention.

[0052] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0053] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0054] Example 1, see Figure 1 The present invention provides a digital twin drug data analysis system, comprising a data fusion module, a multi-layer digital twin modeling module, a cross-domain causal and mechanism fusion module, and an optimization decision-making module, specifically including the following:

[0055] The data fusion module collects drug data, including drug-level data, individual-level data, group-level data, and production quality-level data. It processes the drug data using semantic mapping and feature alignment techniques to obtain a unified drug data foundation.

[0056] The multi-layer digital twin modeling module constructs multi-layer digital twins based on a unified drug data foundation, including molecular-level twins, individual-level twins, and population-level twins.

[0057] The cross-domain causality and mechanism fusion module introduces pharmacological and clinical prior knowledge into a multi-level digital twin to construct a causal structure of drug-target-pathway-clinical outcome. It then uses causal inference and counterfactual analysis methods to analyze the causal inference results.

[0058] The optimization decision-making module, based on multi-level digital twins and causal inference results, introduces a multi-objective method and outputs individualized control strategies.

[0059] Example 2, based on the above examples, describes a data fusion module where drug data includes drug-level data, individual-level data, population-level data, and production quality-level data, specifically including the following:

[0060] Drug-level data: drug molecular structure, physicochemical properties, pharmacokinetic parameters, and pharmacodynamic parameters;

[0061] Individual-level data: genomics, proteomics, clinical testing data, medical imaging data, pathological status, genetic background, and environmental factors;

[0062] Population-level data: laboratory testing data, clinical trial data, and clinical drug use data;

[0063] Production quality data: pharmaceutical manufacturing process and quality control data.

[0064] Example 3, based on the above examples, describes a multi-layered digital twin modeling module that constructs a multi-layered digital twin based on a unified drug data foundation. Specifically, it includes the following steps:

[0065] Step S1: Molecular-level feature extraction. Processed drug-level data is retrieved from a unified drug data base. Combining molecular dynamics simulation, graph neural networks, and feature encoding methods, key molecular-level action features are extracted, including binding affinity, metabolic pathway parameters, and toxicological indicators.

[0066] Step S2: Molecular-level twin modeling. Based on the key role characteristics of the molecular level, and combining quantum chemical calculations and deep learning hybrid models, a molecular-level digital twin is constructed to simulate the drug-target interaction mechanism.

[0067] Step S3: Individual data representation. Individual-level data is retrieved from a unified drug data foundation. Multimodal representation learning and mixed-effects models are used to model individual differences and obtain individual-level drug response characteristics.

[0068] Step S4: Individual-level twin modeling. Based on individual-level drug response characteristics, construct individual-level digital twins, introduce adaptive calibration and lightweight generation fusion mechanisms, and obtain dynamically updated drug response predictions.

[0069] Step S5: Population data aggregation. Population-level data is retrieved from a unified drug data base. Hierarchical Bayesian modeling and multi-density clustering methods are used to identify differences in drug use and efficacy patterns among populations.

[0070] Step S6: Population-level twin modeling. Based on the individual-level twins, construct a population-level digital twin to enable long-term drug efficacy assessment, epidemiological trend prediction, and public health strategy support.

[0071] In this embodiment, taking the antihypertensive drug amlodipine as an example, a multi-level digital twin is constructed based on a unified drug data foundation to achieve drug response prediction from the molecular level, individual level to the population level.

[0072] First, during the molecular-level modeling process, the molecular structure, physicochemical properties, and pharmacokinetic parameters of amlodipine are retrieved from the drug data base. Combined with molecular dynamics simulation, the binding process between amlodipine and calcium ion channel target proteins is dynamically simulated. At the same time, a graph neural network is used to encode the molecular graph and extract key action features, including binding affinity, metabolic pathway parameters, and potential toxicological indicators. Based on these features, a molecular-level digital twin is constructed through a hybrid model of quantum chemical calculations and deep learning, thereby reproducing the interaction mechanism between amlodipine and calcium ion channels in a virtual environment.

[0073] In the individual-level modeling process, genomic data, clinical test data, medical imaging, and lifestyle information of patient A are retrieved from the drug data foundation. These data are fused using a multimodal representation learning method and combined with a mixed-effects model to obtain the individual-level drug response characteristics of patient A, thus constructing an individual-level digital twin of patient A. Based on the initially predicted blood pressure response trend sequence, an adaptive calibration mechanism is introduced to iteratively correct the prediction results. When there is a residual between the predicted blood pressure decrease and the actual monitored value, the predicted sequence is dynamically updated through a gating calibration unit. Simultaneously, a lightweight generation and fusion mechanism is introduced to generate possible blood pressure change candidate sequences, which are then screened and fed back through an effect assessment module, thereby obtaining a more realistic individual response prediction.

[0074] In the population-level modeling process, long-term follow-up data of clinical trial populations and multi-center clinical drug use data are retrieved from the drug data base. Through hierarchical Bayesian modeling and multi-density clustering methods, differences in drug response among populations with different genetic backgrounds and lifestyles are identified. Based on the modeling results of individual-level twins, a population-level digital twin is constructed to simulate the long-term efficacy distribution of amlodipine in the population and predict its antihypertensive effect and potential adverse reaction trends in different populations, thereby providing a basis for public health policy formulation and clinical drug use guidance.

[0075] Example 4, based on the above examples, specifically includes the following steps in step S4:

[0076] Step S41: Call up the individual-level drug response features, standardize and screen the individual-level drug response features using multidimensional features, remove redundant features and enhance key effect factors to obtain the processed drug response features;

[0077] Step S42: Preliminary digital twin generation. Based on the processed drug response characteristics, a preliminary digital twin is constructed. Preliminary predictions are made for each feature dimension to generate an individual-level drug response trend sequence.

[0078] Step S43: Introduce an adaptive calibration mechanism. Based on the initial digital twin, introduce a recurrent convolutional long short-term memory network to construct an adaptive calibration module. Gradually correct the predicted drug response trend sequence to obtain the calibrated drug response characteristics. The calibrated twin model is obtained through the adaptive calibration mechanism and used as the adaptive calibration result.

[0079] Step S44: Introduce a lightweight generation fusion mechanism. In the calibrated twin model, construct a lightweight generation fusion module and introduce a dual-branch architecture of candidate response construction unit and response effect evaluation unit. The candidate response construction unit generates candidate response sequences based on the calibrated drug response characteristics, and the response effect evaluation unit performs discriminative feedback to obtain the lightweight generation fusion result.

[0080] Step S45: Digital twin optimization, fusing adaptive calibration results with lightweight generation fusion results, performing final optimization on the calibrated twin model, and outputting dynamically updated drug response predictions, including the ADME process of the drug in vivo and adverse reaction dynamics.

[0081] In this embodiment, in a clinical study of a novel antidepressant, the system selects multi-source data of patient B as input, including gene sequencing results, basic liver and kidney function tests, previous drug use history, blood drug concentration monitoring data, and clinical adverse reaction scores. First, these data are standardized and feature-screened to remove redundant information and retain key factors closely related to drug metabolism, including drug-metabolizing enzyme activity, plasma half-life parameters, and adverse reaction sensitivity indicators.

[0082] Based on this, the system generated a preliminary digital twin of patient B to predict the metabolic trend of the drug in his body. The preliminary results showed that the blood drug concentration of patient B would remain in the effective range within 12 hours after drug administration, but would drop below the efficacy threshold after 24 hours. An adaptive calibration mechanism was introduced to correct the prediction results through a recurrent convolutional long short-term memory network. Combining the patient's previous history of antidepressant use and current liver function indicators, the model dynamically adjusted the metabolic rate prediction. The calibrated results showed that the concentration of patient B began to decrease at 16 hours after drug administration, and the probability of mild gastrointestinal adverse reactions increased.

[0083] A lightweight generation fusion mechanism is introduced into the calibrated twin model. The candidate response building unit generates multiple possible drug response trajectories, including "normal metabolism", "delayed metabolism with mild adverse reactions" and "rapid metabolism with insufficient efficacy". The response effect assessment unit discriminates and screens them, and finally the fusion outputs the two most likely situations: one is that the drug concentration decreases after 16 hours, resulting in insufficient efficacy, and the other is that the patient has a 30% probability of experiencing mild gastrointestinal discomfort in the first 24 hours.

[0084] The adaptive calibration results and lightweight fusion results are comprehensively optimized to output a dynamically updated individualized drug response prediction model. Through the optimized digital twin, doctors can intuitively obtain the drug concentration change trajectory, ADME process and potential adverse reaction risks of patient B at different time points, and adjust the dosing regimen accordingly, including shortening the dosing interval and reducing the single dose, to balance efficacy and safety.

[0085] Example 5, based on the above examples, specifically includes the following steps in step S43:

[0086] Step S431: Sequence difference calculation. Collect historical individual-level drug response trend sequences, combine them with the individual-level drug response trend sequences, construct a difference function, and calculate the residual values. The formula used is as follows:

[0087] ;

[0088] in, Represents the residual value. This represents a sequence of individual-level drug response trends. This represents a historical individual-level trend sequence of drug response. Indicates a time index;

[0089] Step S432: Adaptive weighted fusion. Adaptive weights are introduced to fuse the historical individual-level drug response trend sequence and the individual-level drug response trend sequence to obtain the weighted individual-level drug response trend sequence. The formula used is as follows:

[0090] ;

[0091] in, This represents the obtained weighted individual-level drug response trend sequence. Indicates adaptive weights;

[0092] Step S433: Gated calibration update. Based on the residual values ​​calculated in step S431 and the weighted individual-level drug response trend sequence obtained by fusion in step S432, a gated calibration unit is constructed. The gate parameters are dynamically adjusted in conjunction with the residual values ​​to update the weighted individual-level drug response trend sequence. The formula used is as follows:

[0093] ;

[0094] in, This represents the calibrated individual-level drug response trend sequence. Indicates the gating parameters;

[0095] Step S434: Calibrate the twin output, set the residual threshold, gradually converge the initial digital twin, repeat step S433 iteratively until the residual value converges to the residual threshold, output the calibrated individual-level drug response trend sequence, map to generate the calibrated drug response features, and output the calibrated twin model obtained through the adaptive calibration mechanism.

[0096] In this embodiment, the code used is as follows:

[0097] enter:

[0098] H_t: Historical individual-level drug response trend sequence

[0099] S_t: Current individual observed drug response trend sequence

[0100] max_iter: Maximum number of iterations

[0101] residual_threshold: Residual convergence threshold

[0102] Output:

[0103] S_calibrated: Calibrated individual-level drug response trend sequence

[0104] features: Drug response features generated by mapping

[0105] step:

[0106] 1. Initialization:

[0107] S_calibrated = S_t

[0108] residual = ∞

[0109] iter = 0

[0110] 2. Iterate repeatedly until residual converges or the maximum number of iterations is reached:

[0111] while residual > residual_threshold and iter < max_iter:

[0112] / / Step S431: Sequence Differential Calculation

[0113] R_t = S_calibrated - H_t

[0114] / / Step S432: Adaptive Weight Fusion

[0115] w_t = compute_adaptive_weight(R_t) / / Dynamically calculate weights based on residuals

[0116] F_t = w_t * H_t + (1 - w_t) * S_calibrated

[0117] / / Step S433: Gating Calibration Update

[0118] G_t = gate_update(R_t) / / Gating parameter, which can be adjusted based on sigmoid or residual.

[0119] S_calibrated = G_t * S_calibrated + (1 - G_t) * F_t

[0120] / / Update residuals and record iterations

[0121] residual = compute_residual(S_calibrated, H_t)

[0122] iter += 1

[0123] 3. Step S434: Calibrate the twin output

[0124] features = map_to_features(S_calibrated) / / Map the calibration sequence to generate drug response features

[0125] return:

[0126] S_calibrated, features.

[0127] Example 6, based on the above examples, specifically includes the following steps in step S44:

[0128] Step S441: Candidate response sequence construction. Based on the calibrated drug response features output in step S434, candidate response sequences are constructed. A lightweight generation network is used to generate the sequences, and the formula used is as follows:

[0129] ;

[0130] in, Candidate reaction sequences, To create lightweight generative networks, The calibrated drug response characteristics;

[0131] Step S442: Response effect assessment, constructing a response effect assessment unit, discriminating and feeding back candidate response sequences, and obtaining sequence discrimination feedback results;

[0132] Step S443: Lightweight generation fusion, combining the candidate reaction sequence and the sequence discrimination feedback result to obtain the candidate reaction sequence after lightweight generation fusion, which is used as the lightweight generation fusion result.

[0133] In this embodiment, the code used is as follows:

[0134] enter:

[0135] X_hat_t: The calibrated drug response characteristics output in step S434

[0136] G_theta: Lightweight generative network model

[0137] max_iter: Maximum number of iterations for discrimination

[0138] Output:

[0139] C_fused: Lightweight generation of candidate reaction sequences after fusion

[0140] step:

[0141] 1. Step S441: Construction of candidate reaction sequences

[0142] C_t = G_theta(X_hat_t) / / Generate candidate reaction sequences using a lightweight generative network

[0143] 2. Step S442: Evaluation of Response Effect

[0144] for iter in 1 to max_iter:

[0145] feedback = evaluate_effect(C_t) / / Construct a response effect evaluation unit to discriminate and provide feedback on the sequence.

[0146] C_t = update_sequence(C_t, feedback) / / Adjust the candidate sequence based on feedback

[0147] 3. Step S443: Lightweight generation and fusion

[0148] C_fused = fuse_sequence(C_t, feedback) / / Fuse the candidate sequence and the discrimination feedback result to obtain the final sequence.

[0149] return:

[0150] C_fused.

[0151] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0152] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0153] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A digital twin-based drug data analysis system, characterized in that: It includes a data fusion module, a multi-layer digital twin modeling module, a cross-domain causal and mechanism fusion module, and an optimization decision-making module, specifically including the following: The data fusion module collects drug data, including drug-level data, individual-level data, group-level data, and production quality-level data. It processes the drug data using semantic mapping and feature alignment techniques to obtain a unified drug data foundation. The multi-layer digital twin modeling module constructs multi-layer digital twins based on a unified drug data foundation, including molecular-level twins, individual-level twins, and population-level twins. The cross-domain causality and mechanism fusion module introduces pharmacological and clinical prior knowledge into a multi-level digital twin to construct a causal structure of drug-target-pathway-clinical outcome. It then uses causal inference and counterfactual analysis methods to analyze the causal inference results. The optimization decision-making module is based on multi-level digital twins and causal inference results, introduces a multi-objective method, and outputs individualized control strategies. The multi-layer digital twin modeling module constructs a multi-layer digital twin based on a unified drug data foundation, specifically including the following steps: Step S1: Molecular-level feature extraction. Processed drug-level data is retrieved from a unified drug data base. Combining molecular dynamics simulation, graph neural networks, and feature encoding methods, key molecular-level action features are extracted, including binding affinity, metabolic pathway parameters, and toxicological indicators. Step S2: Molecular-level twin modeling. Based on the key role characteristics of the molecular level, and combining quantum chemical calculations and deep learning hybrid models, a molecular-level digital twin is constructed to simulate the drug-target interaction mechanism. Step S3: Individual data representation. Individual-level data is retrieved from a unified drug data foundation. Multimodal representation learning and mixed-effects models are used to model individual differences and obtain individual-level drug response characteristics. Step S4: Individual-level twin modeling. Based on individual-level drug response characteristics, construct individual-level digital twins, introduce adaptive calibration and lightweight generation fusion mechanisms, and obtain dynamically updated drug response predictions. Step S5: Population data aggregation. Population-level data is retrieved from a unified drug data base. Hierarchical Bayesian modeling and multi-density clustering methods are used to identify differences in drug use and efficacy patterns among populations. Step S6: Population-level twin modeling. Based on the individual-level twins, construct a population-level digital twin to enable long-term drug efficacy assessment, epidemiological trend prediction, and public health strategy support.

2. The digital twin drug data analysis system according to claim 1, characterized in that: The data fusion module includes drug-level data, individual-level data, population-level data, and production quality-level data, specifically including the following: Drug-level data: drug molecular structure, physicochemical properties, pharmacokinetic parameters, and pharmacodynamic parameters; Individual-level data: genomics, proteomics, clinical testing data, medical imaging data, pathological status, genetic background, and environmental factors; Population-level data: laboratory testing data, clinical trial data, and clinical drug use data; Production quality data: pharmaceutical manufacturing process and quality control data.

3. The digital twin drug data analysis system according to claim 1, characterized in that: Step S4 specifically includes the following steps: Step S41: Call up the individual-level drug response features, standardize and screen the individual-level drug response features using multidimensional features, remove redundant features and enhance key effect factors to obtain the processed drug response features; Step S42: Preliminary digital twin generation. Based on the processed drug response characteristics, a preliminary digital twin is constructed. Preliminary predictions are made for each feature dimension to generate an individual-level drug response trend sequence. Step S43: Introduce an adaptive calibration mechanism. Based on the initial digital twin, introduce a recurrent convolutional long short-term memory network to construct an adaptive calibration module. Gradually correct the predicted drug response trend sequence to obtain the calibrated drug response characteristics. The calibrated twin model is obtained through the adaptive calibration mechanism and used as the adaptive calibration result. Step S44: Introduce a lightweight generation fusion mechanism. In the calibrated twin model, construct a lightweight generation fusion module and introduce a dual-branch architecture of candidate response construction unit and response effect evaluation unit. The candidate response construction unit generates candidate response sequences based on the calibrated drug response characteristics, and the response effect evaluation unit performs discriminative feedback to obtain the lightweight generation fusion result. Step S45: Digital twin optimization, fusing adaptive calibration results with lightweight generation fusion results, performing final optimization on the calibrated twin model, and outputting dynamically updated drug response predictions, including the ADME process of the drug in vivo and adverse reaction dynamics.

4. The digital twin drug data analysis system according to claim 3, characterized in that: Step S43 specifically includes the following steps: Step S431: Sequence difference calculation. Collect historical individual-level drug response trend sequences, combine them with the individual-level drug response trend sequences, construct a difference function, and calculate the residual values. The formula used is as follows: ; in, Represents the residual value. This represents a sequence of individual-level drug response trends. This represents a historical individual-level trend sequence of drug response. Indicates a time index; Step S432: Adaptive weighted fusion. Adaptive weights are introduced to fuse the historical individual-level drug response trend sequence and the individual-level drug response trend sequence to obtain the weighted individual-level drug response trend sequence. The formula used is as follows: ; in, This represents the obtained weighted individual-level drug response trend sequence. Indicates adaptive weights; Step S433: Gated calibration update. Based on the residual values ​​calculated in step S431 and the weighted individual-level drug response trend sequence obtained by fusion in step S432, a gated calibration unit is constructed. The gate parameters are dynamically adjusted in conjunction with the residual values ​​to update the weighted individual-level drug response trend sequence. The formula used is as follows: ; in, This represents the calibrated individual-level drug response trend sequence. Indicates the gating parameters; Step S434: Calibrate the twin output, set the residual threshold, gradually converge the initial digital twin, repeat step S433 iteratively until the residual value converges to the residual threshold, output the calibrated individual-level drug response trend sequence, map to generate the calibrated drug response features, and output the calibrated twin model obtained through the adaptive calibration mechanism.

5. A digital twin drug data analysis system according to claim 3, characterized in that: Step S44 specifically includes the following steps: Step S441: Candidate response sequence construction. Based on the calibrated drug response features output in step S434, candidate response sequences are constructed. A lightweight generation network is used to generate the sequences, and the formula used is as follows: ; in, Candidate reaction sequences, To create lightweight generative networks, The calibrated drug response characteristics; Step S442: Response effect assessment, constructing a response effect assessment unit, discriminating and feeding back candidate response sequences, and obtaining sequence discrimination feedback results; Step S443: Lightweight generation fusion, combining the candidate reaction sequence and the sequence discrimination feedback result to obtain the candidate reaction sequence after lightweight generation fusion, which is used as the lightweight generation fusion result.

Citation Information

Patent Citations

  • Medical industry digital twin platform based on supercomputing

    CN117831640A

  • Artificial intelligence-based pharmaceutical knowledge graph construction method and system

    CN120375910A

  • Digital twin workshop management and control method and system based on data multi-layer fusion

    CN120542835A

  • System and methods for ai-enhanced cellular modeling and simulation

    US20250259715A1

  • Method, apparatus, and computer-readable medium for generating predictions with a digital twin architecture

    WO2025212686A1