Clinical trial simulation methods, equipment, media, and products based on molecular perturbation spectra
By using a causal inference system based on molecular perturbation spectra, the problem of dependence on real-world drug data in existing clinical trial simulation technologies has been solved, enabling reliable efficacy evaluation in innovative drug development and precision medicine, and overcoming simulation obstacles caused by incomplete data and patient heterogeneity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-30
AI Technical Summary
Existing clinical trial simulation technologies rely on high-quality real-world drug use data and cannot overcome simulation barriers caused by a lack of drug use records, incomplete data, or patient molecular heterogeneity, which severely restricts their application in the early stages of innovative drug development and precision medicine.
By establishing a causal inference system based on molecular perturbation spectra, using molecular perturbation spectra as a direct characterization of drug biological effects, combining genomics and high-dimensional omics prediction models, optimizing individual allocation models, and constructing simulated clinical trials, a reliable estimation of drug intervention effects can be achieved.
In the absence of real-world drug use data, it can accurately assess the effects of drug interventions, overcome simulation barriers caused by incomplete data and patient heterogeneity, and support the development of innovative drugs and the application of precision medicine.
Smart Images

Figure CN122314461A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a clinical trial simulation method, device, medium and product based on molecular perturbation spectrum. Background Technology
[0002] Clinical trials are the gold standard for verifying efficacy and safety in drug development, but they are inherently limited by their long duration and high cost. Clinical trial simulation technology, through computer simulation, predicts clinical trial results, which can significantly shorten the time and reduce costs, thereby addressing the shortcomings of traditional clinical trials—long duration and high cost—and improving the efficiency of new drug development.
[0003] However, the inventors have discovered that existing clinical trial simulation technologies have at least the following technical problems:
[0004] Current mainstream methods rely solely on existing clinical data. For example, when studying the efficacy of a drug for lung cancer, it's necessary to screen patients who "used the drug" and those who "did not use the drug" from electronic medical records or health insurance data, and then compare the survival outcomes of these two groups to draw conclusions. Clearly, the core premise of this method is the existence of a large number of clinical records showing actual patient use of the drug. Therefore, if a drug is not yet on the market or has never been used in a particular disease population, the lack of corresponding clinical data makes it impossible to conduct effective clinical trial simulations. This limits the application of this technology to the early stages of innovative drug development and to expanding drug indications in new disease areas.
[0005] Furthermore, current technologies heavily rely on the quality and completeness of clinical diagnostic data. For drug types where medication information is difficult to record and statistically analyze (such as over-the-counter drugs and traditional Chinese medicine preparations), effective clinical trial simulations are often difficult to conduct due to the inability to guarantee data accuracy. Simultaneously, existing technological solutions typically fail to adequately consider the heterogeneity of patients at the molecular level. That is, for patient groups with the same disease but different molecular characteristics, current methods struggle to effectively distinguish them during simulation, a limitation that may affect the accuracy and reliability of simulation results.
[0006] In summary, existing clinical trial simulation technologies rely on high-quality real-world drug use data and cannot overcome simulation barriers caused by a lack of drug use records, incomplete data, or patient molecular heterogeneity, which severely restricts their application in the early stages of innovative drug development and precision medicine. Summary of the Invention
[0007] One objective of this application is to provide a clinical trial simulation method, device, medium, and product based on molecular perturbation spectrum, at least to address the technical problem that existing clinical trial simulation technologies must rely on high-quality real-world medication data and cannot overcome simulation obstacles caused by lack of medication records, incomplete data, or patient molecular heterogeneity, which severely restricts their application in early-stage innovative drug development and precision medicine.
[0008] To achieve the above objectives, some embodiments of this application provide the following aspects:
[0009] In a first aspect, some embodiments of this application also provide a clinical trial simulation method based on molecular perturbation spectra. The method includes: determining the molecular perturbation spectrum of a target drug; determining a predictive model from genomics to high-dimensional omics based on genomics data and high-dimensional omics data in a first population dataset; determining the predictive omics data of an individual based on the predictive model and the genomics data of an individual in a second population dataset; wherein the second population dataset is also associated with clinical outcome data; determining a grouping scheme for the simulated clinical trial by optimizing an individual allocation model based on the molecular perturbation spectrum and the predictive omics data of all individuals, so as to maximize the similarity between the predicted omics difference and the molecular perturbation spectrum between the simulated experimental group and the simulated control group; constructing a simulated clinical trial based on the grouping scheme; and determining an estimated value of the intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group in the simulated clinical trial.
[0010] Secondly, some embodiments of this application also provide an electronic device, the electronic device comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.
[0011] Thirdly, some embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the method described above.
[0012] Fourthly, some embodiments of this application also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.
[0013] Compared with related technologies, the solution provided in this application, by establishing a complete causal inference system, can effectively solve the core technical problems faced by existing clinical trial simulation technologies: First, this application cleverly eliminates the dependence on real-world drug use data by introducing molecular perturbation spectra as a direct characterization of drug biological effects. Even for unmarketed drugs or drug types lacking drug use records (such as traditional Chinese medicine preparations), subsequent simulation analysis can be carried out simply by obtaining their molecular perturbation characteristics through experiments, which breaks through the application limitations of traditional methods in the early-stage development of innovative drugs. Second, a genomics and high-dimensional omics prediction model is established based on the first population dataset and applied to the second population dataset to obtain predictive omics data. This step innovatively uses genetic variation as an instrumental variable, which can effectively correct for various biases, including unobserved confounding. This method does not rely on the completeness of clinical data and can still ensure the reliability of effect estimation even in scenarios with poor data quality, such as using over-the-counter drugs, overcoming the reliability problem of traditional methods in the case of incomplete data. Third, by optimizing the individual allocation model to determine the grouping scheme, the baseline comparability between the experimental group and the control group is ensured at the molecular level. This approach fully considers the molecular heterogeneity of patients. When constructing simulated clinical trials, it ensures baseline equilibrium between the experimental and control groups for patients with different molecular characteristics, solving the problem of population heterogeneity that traditional simulation techniques cannot handle. Finally, based on the constructed simulated clinical trials, the differences in clinical outcomes between the two groups are compared to obtain reliable estimates of the intervention effect. The entire process establishes a complete causal chain of "drug intervention → molecular perturbation → clinical outcome," enabling accurate efficacy assessment even in the absence of real-world drug use data.
[0014] In summary, this application innovatively solves the problem of excessive reliance on real-world drug data in existing simulation technologies by establishing a complete causal chain of "drug intervention → molecular perturbation → clinical outcome". At the same time, it overcomes the simulation obstacles caused by incomplete data and patient heterogeneity, providing effective technical support for innovative drug development and precision medicine practice. Attached Figure Description
[0015] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0016] Figure 1 An exemplary flowchart of a clinical trial simulation method based on molecular perturbation spectrum provided for embodiments of this application;
[0017] Figure 2An exemplary flowchart illustrating an application example of the clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0018] Figure 3 An exemplary flowchart of step S104 in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0019] Figure 4 An exemplary schematic diagram of GPCs that are significantly associated with the risk of lung cancer in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0020] Figure 5 This is an exemplary schematic diagram illustrating the difference in GPCs between two groups in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment, which is consistent with the perturbation effect of a certain drug A on GPCs.
[0021] Figure 6 An exemplary schematic diagram illustrating the grouping and weighting of individuals in a simulated trial dataset in a clinical trial simulation method based on molecular perturbation spectra provided in this application embodiment;
[0022] Figure 7 An exemplary schematic diagram regarding the cumulative event incidence rate in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0023] Figure 8 An exemplary schematic diagram illustrating the estimation of the preventive effect of a certain drug A on lung cancer in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0024] Figure 9 An exemplary schematic diagram of negative control results in a clinical trial simulation method based on molecular perturbation spectrum provided in this application embodiment;
[0025] Figure 10 This is an exemplary flowchart of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] The following terms are used in this document.
[0028] 1. Clinical trial: A scientific research method that randomly assigns research subjects (such as patients) to experimental and control groups to evaluate the safety and efficacy of a specific drug or treatment.
[0029] 2. Molecular perturbation spectrum: refers to the systemic changes in the molecular composition (such as gene expression, protein abundance, metabolite levels, etc.) of a biological system (such as cells, tissues, or organisms) under the influence of specific stimuli (such as drug treatment, gene knockout, environmental changes, etc.). In this application, this term specifically refers to conducting drug intervention trials on cell samples, animal models, or a small number of human subjects, using high-throughput omics technologies (such as transcriptomics, proteomics, etc.) to detect and compare omics changes before and after drug action, thereby characterizing the global impact of the intervention perturbation on the biological system.
[0030] 3. GWAS population data: This data type is based on Genome-Wide Association Studies (GWAS). Its core characteristics require the study population to possess two pieces of information: firstly, completed genomic testing; and secondly, follow-up records showing disease incidence, outcomes, and other phenotypes. Currently, several public GWAS population databases are available, such as those from the UK Biobank.
[0031] 4. Clinical trial simulation: This is a research method that uses existing real-world database data or simulated data to construct a virtual cohort of subjects, and then reproduces the existing clinical trial process and results in the cohort, or predicts the potential clinical trial process and results.
[0032] 5. Instrumental variables: These are statistical variables used in causal inference to address endogeneity bias (such as confounding factors or reverse causality). Such variables must meet three core conditions: first, they must be strongly correlated with the exposure variable; second, they must only indirectly affect the outcome variable through the exposure variable, not directly affect the outcome; and third, they must be independent of confounding factors that may cause endogeneity bias. In this application, genetic variation is specifically used as the instrumental variable for causal inference.
[0033] First Embodiment
[0034] The first embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. For example... Figure 1 and Figure 2 As shown, the method may include the following steps:
[0035] Step S101: Determine the molecular perturbation spectrum of the target drug;
[0036] Step S102: Based on the genomics data and high-dimensional omics data in the first population dataset, determine the prediction model from genomics to high-dimensional omics.
[0037] Step S103: Determine the predictive omics data of the individual based on the prediction model and the genomic data of the individual in the second population dataset; wherein, the second population dataset is also associated with clinical outcome data;
[0038] Step S104: Based on the molecular perturbation spectrum and the predicted omics data of all individuals, determine the grouping scheme for the simulated clinical trial by optimizing the individual allocation model, so as to maximize the similarity between the predicted omics differences and the molecular perturbation spectrum between the simulated experimental group and the simulated control group.
[0039] Step S105: Construct a simulated clinical trial according to the grouping scheme, and determine the estimated value of the intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group in the simulated clinical trial.
[0040] This application aims to estimate the intervention effect of target drug T on clinical outcome O in a population using computer simulation methods. The following sections provide detailed explanations of each step described above.
[0041] For step S101, for example, the target drug may include, but is not limited to, small molecule chemical drugs, biological agents, natural products, or effective components of traditional Chinese medicine. The molecular perturbation spectrum can be obtained through small-scale experiments, thereby significantly reducing early-stage research and development costs.
[0042] In some cases, molecular perturbation spectra can be obtained through standardized experimental procedures: First, drug treatment is performed, exposing cells or tissues to the target drug and a solvent control group, respectively, and independent replicate experiments are established. Then, samples such as treated RNA or proteins are extracted. Next, high-throughput sequencing is performed on the aforementioned samples, and bioinformatics analysis is used to quantify expression levels, identify differentially expressed molecules between the drug intervention group and the control group, and generate molecular perturbation spectra quantified by log2 fold change. For example, the molecular perturbation spectrum can be denoted as an eigenvector { , ,…, }
[0043] Optionally, in some cases, sequencing data can be obtained from public databases (such as GEO, CMap) or randomized controlled clinical trial data can be used as supplementary data sources to ensure the reliability and reproducibility of molecular perturbation spectra.
[0044] For step S102, exemplarily, the first population dataset is used to characterize a dataset consisting of genomic data G and high-dimensional omics data X collected simultaneously in a specific population. The genomic data may include, but is not limited to, genetic polymorphisms such as single nucleotide polymorphisms. The high-dimensional omics data may include, but is not limited to, transcriptomic, proteomic, or metabolomic data.
[0045] Based on this dataset, a mapping relationship, i.e., a prediction model F, can be established from genomic data G to high-dimensional omics data X using statistical learning or machine learning methods. The prediction model F establishes a quantitative association between genomic data and high-dimensional omics data. For any given individual, the predicted value X' of the corresponding high-dimensional omics data X can be calculated based on the individual's genomic data G through the functional relationship X'=F(G).
[0046] For step S103, exemplarily, the second population dataset is a dataset characterizing genomic data G collected simultaneously in a specific population and clinical outcome data O of research interest. The genomic data in this dataset may include genetic variation information relevant to GWAS population data. By cleverly utilizing such population datasets containing genetic information, the efficacy of drugs can be evaluated without actual drug administration.
[0047] In practical applications, the genomic data G of each individual in the second population dataset can be used as input and substituted into the constructed prediction model F. The predicted value X' of the high-dimensional omics data X corresponding to that individual can be calculated through the functional relationship X'=F(G). This calculation process enables the indirect inference of the intrinsic molecular phenotype in a population with only genomic data and clinical outcome data, laying the foundation for establishing the association between omics characteristics and clinical outcomes, and thus supporting the conduct of simulated drug intervention trials in this population.
[0048] For step S104, for example, a machine learning model can be trained to assign individuals in the second population dataset to the simulated experimental group and the simulated control group. This process can determine the optimal grouping scheme based on the molecular perturbation spectrum ΔX of the target drug and the predicted omics data X' of all individuals by optimizing the individual allocation model (essentially a machine learning model with individual allocation capabilities).
[0049] During the optimization process, the parameters of the individual allocation model can be continuously adjusted based on the molecular perturbation spectrum ΔX of the target drug, so that the predicted omics difference ΔX' between the simulated experimental group and the control group is as close as possible to the molecular perturbation spectrum ΔX, i.e., maximizing their similarity. Through this optimization mechanism, the molecular effects of drug intervention can be simulated at the population level, while ensuring that the simulated experimental group and the control group reach baseline equilibrium at the molecular level, providing a reliable basis for subsequent causal inference.
[0050] The grouping scheme output by the final determined individual allocation model is the optimal scheme, which can be used to construct a virtual queue with good comparability to support subsequent causal effect assessment.
[0051] For step S105, for example, a simulated clinical trial (Sim-Trial) can be constructed based on a determined grouping scheme. In this simulated clinical trial, various statistical analysis methods applicable to clinical trials can be used to compare the differences in clinical outcomes between the simulated experimental group and the simulated control group, thereby obtaining an estimate of the effect of the target drug on the clinical outcome.
[0052] Optionally, to achieve a quantitative assessment of drug efficacy, the estimated effect value can be further converted into an estimated drug intervention effect value at a specific clinical dose based on the conversion relationship between the experimental dose and the clinical target dose, thereby achieving an accurate quantitative analysis of drug efficacy.
[0053] See Figure 2 As shown, Figure 2 This is an exemplary flowchart illustrating an application example of the clinical trial simulation method based on molecular perturbation spectra provided in this application. Specifically, differentially expressed molecules are identified by comparing the drug intervention group and the solvent control group, thus constructing a drug-specific molecular perturbation spectrum. Using nearly randomly assigned genetic variations at birth (such as single nucleotide polymorphisms) as a tool, the expression levels of specific molecules (such as genes or proteins) in an individual are predicted; this process simulates the drug's intervention on molecules. Subsequently, genetic tools are generated using dimensionality reduction methods such as principal component analysis, and a machine learning model is constructed to calculate an allocation score for each individual. Based on this, individuals in a large-scale population dataset are divided into a simulated intervention group and a simulated control group, thereby constructing a simulated clinical trial. By statistically analyzing the differences in clinical outcomes between the two groups in this simulated clinical trial, the potential effects of the drug in the real population can be inferred. Therefore, this application integrates in vitro discovered drug perturbation mechanisms with in vivo human genetic evidence, providing a reliable bridge for causal inference from cell phenotype to human efficacy.
[0054] It is not difficult to see that, compared with related technologies, the solution provided in this application, by establishing a complete causal inference system, can effectively solve the core technical problems faced by existing clinical trial simulation technologies: First, this application cleverly eliminates the dependence on real-world drug use data by introducing molecular perturbation spectra as a direct characterization of drug biological effects. Even for unmarketed drugs or drug types lacking drug use records (such as traditional Chinese medicine preparations), subsequent simulation analysis can be carried out simply by obtaining their molecular perturbation characteristics through experiments, which breaks through the application limitations of traditional methods in the early development of innovative drugs. Second, a genomics and high-dimensional omics prediction model is established based on the first population dataset and applied to the second population dataset to obtain predictive omics data. This step innovatively uses genetic variation as an instrumental variable, which can effectively correct for various biases, including unobserved confounding. This method does not rely on the completeness of clinical data and can still ensure the reliability of effect estimation even in scenarios with poor data quality, such as using over-the-counter drugs, overcoming the reliability problem of traditional methods in the case of incomplete data. Third, by optimizing the individual allocation model to determine the grouping scheme, the baseline comparability between the experimental group and the control group is ensured at the molecular level. This approach fully considers the molecular heterogeneity of patients. When constructing simulated clinical trials, it ensures baseline equilibrium between the experimental and control groups for patients with different molecular characteristics, solving the problem of population heterogeneity that traditional simulation techniques cannot handle. Finally, based on the constructed simulated clinical trials, the differences in clinical outcomes between the two groups are compared to obtain reliable estimates of the intervention effect. The entire process establishes a complete causal chain of "drug intervention → molecular perturbation → clinical outcome," enabling accurate efficacy assessment even in the absence of real-world drug use data.
[0055] In summary, this application innovatively solves the problem of excessive reliance on real-world drug data in existing simulation technologies by establishing a complete causal chain of "drug intervention → molecular perturbation → clinical outcome". At the same time, it overcomes the simulation obstacles caused by incomplete data and patient heterogeneity, providing effective technical support for innovative drug development and precision medicine practice.
[0056] Second Embodiment
[0057] The second embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. The second embodiment is an improvement upon the first embodiment, specifically in that it provides a concrete method for determining the molecular perturbation spectrum of a target drug.
[0058] Specifically, determining the molecular perturbation spectrum of the target drug may include:
[0059] Step S1011: Perform differential analysis on the high-dimensional omics data before and after drug intervention, and determine the results of the differential analysis;
[0060] Step S1012: Based on the difference analysis results, determine the molecular perturbation feature vector composed of differentially expressed molecules and their change amplitudes.
[0061] For step S1011, for example, analysis can be performed based on high-dimensional omics data obtained from the experiment. This data can be derived from omics testing of the experimental subjects before and after intervention with the target drug.
[0062] The experimental subjects may include, but are not limited to, cells, animal models, organoids, or human subjects. The high-dimensional omics data may include, but are not limited to, transcriptomics, proteomics, metabolomics, and other types.
[0063] For step S1012, for example, molecular features that show significant differences before and after drug intervention can be extracted based on the results of differential analysis. These features include, but are not limited to, differentially expressed genes, proteins, or metabolites. By quantifying the magnitude of these differentially expressed molecules (e.g., expressed as log2 values), they can be integrated to form a complete molecular perturbation feature vector, thereby obtaining a molecular perturbation spectrum ΔX that can characterize the biological effects of the target drug.
[0064] It is readily apparent that, in this embodiment, by performing differential analysis on high-dimensional omics data before and after drug intervention, molecular-level changes caused by drug action can be accurately identified, establishing a direct chain of evidence for drug-biological system interactions. Then, the molecular perturbation feature vector constructed based on the differential analysis results can quantify the specific biological effects of the drug, providing a reliable molecular characteristic benchmark for subsequent simulated clinical trials. This molecular perturbation spectrum-based method overcomes the dependence of traditional clinical trial simulations on real-world drug use data, enabling the establishment of a complete drug effect assessment system through experimental data, even for unmarketed drugs or drug types lacking complete usage records.
[0065] Third Embodiment
[0066] The third embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. The third embodiment is an improvement on the first embodiment, specifically in that: in this embodiment, a method is provided that, based on both genomic data and high-dimensional omics data in a first population dataset, the specific implementation of the predictive model from genomics to high-dimensional omics is determined.
[0067] Optionally, in some embodiments, step S102, which involves determining a predictive model from genomics to high-dimensional omics based on genomics data and high-dimensional omics data from the first population dataset, may include:
[0068] Step S1021: Determine the quantitative relationship between genotype and gene expression level based on the principle of allele fold change;
[0069] Step S1022: Construct the prediction model based on the quantitative relationship.
[0070] In some cases, the gene / protein expression profile of an individual can be predicted by linking genetic variation with molecular expression. Specifically, a quantitative relationship between genotype and gene expression level can be established based on the allele fold change (aFC) principle. In practical applications, three types of data are required: genomic data of cohort individuals, cis-QTL data, and allele fold change (aFC) data from publicly available aggregated statistical data. The core computational logic is based on the following assumption: the total gene expression level (e) of a genotype is defined as the sum of the expression levels of the two haplotypes, specifically in three cases:
[0071] Reference homozygote:
[0072] Heterozygote:
[0073] Alternative homozygotes:
[0074] in, For reference haplotype expression levels, To replace the expression level of haplotypes.
[0075] aFC is expressed on a log2 scale: ,in aFC value This is used to express the fold of a haplotype relative to a reference haplotype.
[0076] Furthermore, a complete prediction model can be constructed based on the aforementioned quantitative relationship. Specifically, given the known expression levels of the reference homozygote... with aFC value In this case, the total expression level of each genotype can be simplified as:
[0077] Reference homozygote:
[0078] Heterozygote:
[0079] Alternative homozygotes:
[0080] In actual calculations, we set By predicting relative expression levels and performing log2 transformation, the molecular expression matrix of an individual's genes / proteins can be output, which is the relative expression level predicted based on genetic variation. This output result can directly constitute the core component of the prediction model, completing the construction of a prediction model from genomic data to omics data.
[0081] Optionally, in some embodiments, step S102, which involves determining a predictive model from genomics to high-dimensional omics based on genomics data and high-dimensional omics data from the first population dataset, may include:
[0082] Step S1021': Determine the association weights between genetic variation and omics characteristics based on the aggregated statistical data of tissue-specific quantitative trait loci;
[0083] Step S1022': Construct the prediction model based on the association weights.
[0084] In some cases, the association weights between genetic variations and omics traits can be determined based on aggregated statistical data of quantitative trait loci (QTLs) from specific tissues (such as liver tissue, brain tissue, etc.). In practical applications, aggregated statistical data of QTLs from target tissues can be used to determine the association weights between each genetic variation locus (such as SNP) and target omics traits (such as gene expression levels) through statistical analysis. These association weights can reflect the magnitude and direction of the effect of a specific genetic variation on omics traits.
[0085] Furthermore, a predictive model can be constructed based on the determined association weights. In specific implementation, any of the following modeling strategies can be adopted: deriving the expression level through linear weighted combination based on the linear association coefficient of QTL; or using machine learning algorithms such as PrediXcan and JTI-prediXcan to directly complete the end-to-end prediction from genetic variation to omics features.
[0086] It should be noted that this embodiment can also be an improvement based on the second embodiment.
[0087] It is readily apparent that this application provides two optional implementation methods for constructing genomic prediction models: a prediction model based on the principle of allele fold change can establish a quantitative causal relationship between genotype and gene expression levels, accurately reflecting the biological effects of genetic variation on molecular phenotypes; while a linear prediction model based on tissue-specific quantitative trait loci establishes a causal prediction path from genomics to high-dimensional omics through the association weights between genetic variation and omics characteristics. Both methods utilize genetic variation as an instrumental variable, effectively controlling the interference of confounding variables such as environmental factors and lifestyle on the prediction results, providing a reliable technical foundation for achieving accurate omics status prediction in subsequent second population datasets.
[0088] Fourth embodiment
[0089] The fourth embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectrum. The fourth embodiment is an improvement upon the first embodiment, specifically in that it introduces principal component analysis as an optional data dimensionality reduction process.
[0090] Optionally, in some embodiments, after determining the predictive omics data of an individual based on the predictive model and the genomic data of an individual in the second population dataset, the method may further include:
[0091] Step S201: Based on the predicted omics data, determine a set of low-dimensional principal component features through principal component analysis;
[0092] Step S202: Determine the projection of the molecular perturbation spectrum onto the low-dimensional principal component feature space, and use it as the drug perturbation principal component vector.
[0093] For example, principal component analysis can be performed on the relative expression levels of predicted genetic variations to convert high-dimensional predictive omics data into low-dimensional principal component features, thereby outputting principal component GPCs for predicted genetic variations of individuals in the cohort.
[0094] Furthermore, the molecular perturbation spectral vector can be transformed to the corresponding principal component dimension to obtain the perturbation effect vector ∆X_GPC of the drug on GPCs. This perturbation effect vector can be used to quantify the magnitude of the effect produced by drug intervention in each principal component dimension.
[0095] Optionally, in some embodiments, after determining the projection of the molecular perturbation spectrum onto the low-dimensional principal component feature space as the drug perturbation principal component vector based on the molecular perturbation spectrum (i.e., step S202), the method may further include step S203: determining a subset of features from the low-dimensional principal component features that meet a preset significance level with respect to the clinical outcome and / or the drug based on the statistical test results; wherein, in the step of determining the grouping scheme for the simulated clinical trial by optimizing the individual allocation model based on the molecular perturbation spectrum and the predictive omics data of all individuals, only the subset of features is used.
[0096] In the implementation process, statistical tests can be used to analyze the association between each GPC and the clinical outcome variable, identifying GPCs with significant associations. Simultaneously, statistical tests can be used to assess the impact of drug intervention on each GPC, identifying GPCs that show significant differences before and after drug intervention. Based on these analytical results, GPCs with significant associations to drug intervention and / or clinical outcomes can be retained as a feature subset, which will be used only for optimization analysis when determining the grouping scheme for the simulated clinical trial.
[0097] It should be noted that this embodiment may also be an improvement based on the second embodiment and / or the third embodiment.
[0098] It is readily apparent that this application's embodiments incorporate principal component analysis (PCA), an optional data dimensionality reduction process. This process first reduces the dimensionality of the output of the prediction model F, transforming high-dimensional omics data into low-dimensional principal component features with clear biological significance. Subsequently, by analyzing the relationship between each low-dimensional component and drug intervention and / or clinical outcomes, the subset of features most relevant to drug effects and disease outcomes is selected and retained. This approach not only improves the reliability of the overall computational process through dimensionality reduction but also establishes a clearer causal path by focusing on key feature dimensions: drug intervention first affects core molecular principal component features, and then ultimately influences clinical outcomes through these key mediating variables. This effectively eliminates the interference of irrelevant variables, enhancing the inference accuracy and biological interpretability of the causal chain of "drug intervention - molecular change - clinical outcome."
[0099] Fifth Embodiment
[0100] The fifth embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. The fifth embodiment is an improvement upon the first embodiment, specifically in that it provides a method for determining the grouping scheme of the simulated clinical trial by optimizing an individual allocation model based on the molecular perturbation spectra and predictive omics data of all individuals.
[0101] Specifically, see Figure 3 As shown, step S104, which involves determining the grouping scheme for the simulated clinical trial by optimizing the individual allocation model based on the molecular perturbation spectrum and the predictive omics data of all individuals, may include:
[0102] Step S1041: Based on the genomic data of individuals in the second population dataset, determine the allocation score for each individual using an individual allocation model;
[0103] Step S1042: Based on the allocation score, individuals are assigned to the simulated trial group and the simulated control group to construct the initial simulated clinical trial;
[0104] Step S1043: Based on the initial simulated clinical trial, determine the grouping scheme for the simulated clinical trial by optimizing the individual allocation model.
[0105] For example, a machine learning model M, i.e., an individual allocation model, is used, with the genomic data G of each individual in the second population dataset (corresponding to...) Figure 3 Using the genetic mutation in the genome as input, the individual's assignment score S is calculated through the functional relationship S=M(G). This assignment score is a continuous value between 0 and 1, representing the probability that the individual will be assigned to the simulated experimental group based on their genomic characteristics.
[0106] Optionally, in some embodiments, the process of assigning individuals to the simulated experimental group and the simulated control group based on the allocation score can be carried out by random allocation or weighted allocation, in which individuals in the second population dataset are assigned to the experimental group T and the control group C, respectively.
[0107] Specifically, in the random allocation method, each individual in the population is assigned to the simulated experimental group with a probability of allocation score S and to the simulated control group with a probability of 1-S. The simulated random allocation process is realized through a pseudo-random number generator, thereby constructing the initial simulated clinical trial.
[0108] Specifically, in the weighted allocation method, by constructing a pseudo-population, two observation records are constructed for each individual in the queue. One record assigns the individual to the simulated intervention group (A=1) and assigns a machine learning score S, while the other record assigns the individual to the control group (A=0) and assigns a weight 1-S. Through this weighted cloning method, a simulated intervention effect simulation dataset is constructed, which includes the individual's grouping variable A and the corresponding allocation weight.
[0109] Both methods described above are based on machine learning to assign scores, achieving effective allocation of individuals to experimental and control groups. This provides a reliable foundation of simulated clinical trial data for subsequent intervention effect evaluation. The weighted allocation method, by constructing a pseudo-population, can better address the bias issues arising from non-random allocation.
[0110] Optionally, in some embodiments, the step of determining the grouping scheme for the simulated clinical trial based on the initial simulated clinical trial by optimizing the individual allocation model, i.e., step S1043, may include:
[0111] Step S10431: In the initial simulated clinical trial, the omics difference characteristics are determined based on the predicted omics data of the simulated experimental group and the control group;
[0112] Step S10432: Calculate the similarity between the omics differential features and the molecular perturbation spectrum;
[0113] Step S10433: Optimize the individual allocation model with the goal of maximizing the similarity to obtain the optimal individual allocation model.
[0114] Specifically, in the initial Sim-Trial, the predicted omics data X' of the experimental group and the control group can be compared to extract omics difference features, i.e., simulated perturbation spectrum △X'. Then, the molecular perturbation spectrum △X is compared with the simulated perturbation spectrum △X', and the similarity (similarity(△X, △X')) between the two is calculated. The individual allocation model is then continuously trained with the goal of maximizing this similarity to obtain the optimized individual allocation model. The allocation scheme determined based on the optimized individual allocation model is the final grouping scheme for the simulated clinical trial.
[0115] It should be noted that this embodiment may also be an improvement based on any one or more of the second to fourth embodiments.
[0116] It is readily apparent that in this embodiment, individual allocation scores are determined based on genomic data, establishing a quantitative association between genetic characteristics and experimental grouping. An initial simulated clinical trial is constructed using these allocation scores, achieving an effect of near-randomized grouping in observational data. Further optimization of the individual allocation model ensures that the predicted omics differences between the experimental and control groups approximate the drug perturbation spectrum, essentially simulating the molecular effects of drug intervention. This hierarchical optimization mechanism constructs a complete causal path from "genetic background → molecular characteristics → clinical outcome," utilizing genetic variation as an instrumental variable to control confounding bias while ensuring accurate estimation of causal effects through precise molecular-level matching.
[0117] Sixth Embodiment
[0118] The sixth embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectrum. The sixth embodiment is an improvement upon the fifth embodiment, specifically in that it provides a concrete implementation of the individual allocation model.
[0119] Optionally, in some embodiments, the individual allocation model is specifically a machine learning function whose output is the probability weight of an individual being assigned to the simulated experimental group, expressed as:
[0120]
[0121] Where Gi represents the genomic data of the i-th individual or the principal component features of the genomic data after dimensionality reduction; The probability weights are defined as follows: A represents the intervention allocation, where A=1 represents allocation to the simulated trial group, and A=0 represents allocation to the simulated control group. In this embodiment, this probabilistic allocation method allows the individual allocation model to flexibly construct simulated clinical trials, rather than simply rigidly dividing the population. The probability weights are... It can be equal to the machine learning score S.
[0122] Optionally, in some embodiments, the individual allocation model is specifically implemented using a composite model of linear combination and the Sigmoid function, expressed as:
[0123]
[0124] in, GPC2,…,GPCn are the first to nth principal component features obtained after dimensionality reduction by principal component analysis of the genomic data; α1,α2,…,α n The parameter weights that the model needs to learn are assigned to the individual, reflecting the relative importance of each principal component feature in the assignment decision. This is the Sigmoid function.
[0125] For example, the Sigmoid function This is responsible for mapping the result of a linear combination to a probability space between 0 and 1, ensuring that the weights assigned to the output have a clear probabilistic interpretation, where z represents the input value of the sigmoid function. This represents the output value of the Sigmoid function, where e is a mathematical constant and the base of the natural logarithm.
[0126] It should be noted that this embodiment may also be an improvement based on any one or more of the first to fourth embodiments.
[0127] It is readily apparent that, in this embodiment, by introducing a machine learning model based on genomic data or its principal component features, the probability weights for assigning individuals to simulated experimental or control groups are precisely quantified. This approach eliminates the reliance on randomization or subjective judgment in intervention allocation (A=1 or A=0), instead basing it on causal inference based on individual genetic characteristics. This enhances the scientific rigor and interpretability of the experimental grouping, facilitating more reliable estimation of the causal effects of interventions in observational studies or real-world data. This method provides a computationally achievable and reproducible modeling path for causal analysis in the context of high-dimensional genomic data.
[0128] Seventh Embodiment
[0129] The seventh embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. The seventh embodiment is an improvement upon the fifth embodiment, specifically in that it provides two concrete implementation methods for calculating the similarity between the omics differential features and the molecular perturbation spectra.
[0130] Optionally, in some embodiments, the calculation of the similarity between the omics differential features and the molecular perturbation spectrum is specifically achieved through any of the following methods:
[0131] Based on the squared correlation coefficient between the omics differential features and the molecular perturbation spectrum; or based on the Euclidean distance between the omics differential features and the scaled molecular perturbation spectrum.
[0132] Optionally, in some embodiments, the squared correlation coefficient between the omics differential features and the molecular perturbation spectrum is specifically implemented using the following formula:
[0133]
[0134] in, This represents the vector of perturbation effects of drugs on genomic principal components (GPCs). This represents the difference vector between the simulated intervention group and the simulated control group in GPCs during the simulation experiment. This is the squared value of the correlation coefficient between two vectors. It can be understood that this method primarily captures the consistency in the changing patterns of two vectors; the closer the squared correlation coefficient is to 1, the more similar the changing trends of the two vectors.
[0135] Optionally, in some embodiments, the Euclidean distance between the omics differential features and the scaled molecular perturbation spectrum is specifically achieved by the following formula:
[0136]
[0137] in, This represents the vector of perturbation effects of drugs on genomic principal components (GPCs). This represents the difference vector between the simulated intervention group and the simulated control group in GPCs during the simulation experiment. It is an adaptive scaling coefficient. This method measures similarity by calculating the direct distance between two vectors in numerical space; the smaller the distance value, the higher the similarity. It is worth mentioning that this embodiment introduces an adaptive scaling coefficient. This allows the method to automatically adjust the magnitude ratio of the two vectors, thereby enhancing its adaptability to data of different scales.
[0138] It is worth emphasizing that during model training, the individual assignment model is optimized by minimizing the objective function. The objective function can employ any of the similarity calculation methods mentioned above; the core objective is to minimize the difference in GPCs between the two groups in the simulation experiment. ) and the perturbation effect of drugs on GPCs ( To achieve maximum consistency, the drug's effects can be accurately simulated in the population, and machine learning assignment scores can be output for individuals in the output queue to complete the model optimization process.
[0139] It should be noted that this embodiment may also be an improvement based on any one or more of the first to fourth and sixth embodiments.
[0140] It is readily apparent that, in the embodiments of this application, by calculating the similarity between omics differential features and known molecular perturbation spectra (using the squared correlation coefficient or scaled Euclidean distance), the complex assessment of causal effects is transformed into a quantifiable and robust similarity comparison problem. Specifically, if the omics differential features generated by drug intervention are highly similar (high correlation or short distance) to the spectrum of a specific molecular perturbation (such as gene knockout), it strongly suggests that the drug may exert its effects by perturbing the same biological pathways or targets, thus providing a mechanistic explanation and evidence for the causal chain between "drug-phenotype" at the molecular level. Compared to traditional methods, this similarity-based causal inference approach lowers the threshold for interpreting high-dimensional omics data and enhances the objectivity and reproducibility of inferring drug action mechanisms.
[0141] Eighth embodiment
[0142] The eighth embodiment of this application relates to a clinical trial simulation method based on molecular perturbation spectra. The eighth embodiment is an improvement upon the first embodiment, specifically in that it provides a method for determining the estimated intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group.
[0143] Optionally, in some embodiments, determining the estimated intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group in the simulated clinical trial, i.e., step S105, may include:
[0144] Step S1051: Using weighted regression analysis, with the assigned weights as the weighting values, calculate the difference in clinical outcomes between the simulated experimental group and the simulated control group;
[0145] Step S1052: Based on the proportional relationship between the molecular perturbation spectrum and the simulated omics differences, the differences in clinical outcomes are converted into estimated values of drug intervention effects at specific doses.
[0146] Optionally, in some embodiments, the weighted regression analysis method may include, but is not limited to, weighted linear regression, weighted logistic regression, or weighted Cox regression.
[0147] In practical applications, based on the basic information and follow-up outcome data, grouping variables and allocation weights of each individual in the simulation test dataset, the weighted regression analysis method is used to estimate the difference in clinical outcomes between the simulated experimental group and the simulated control group by using the allocation score output by the individual allocation model as the weighting value, and to obtain a preliminary estimate of the drug intervention effect E.
[0148] Furthermore, the drug dose used when obtaining the molecular perturbation spectrum ΔX in the experimental stage can be converted into the corresponding human equivalent dose d, and then corrected by multiplying by the fitness coefficient k. Based on the assumption of linear dose-response relationship, the drug intervention effect E obtained in step S1051 is combined with the target drug dose d0 to be used in clinical practice. The final estimated value of the drug intervention effect at the target dose d0 is calculated by the formula (E×d0) / (k×d), thereby completing the quantitative transformation from simulated effect to real clinical effect.
[0149] It should be noted that this embodiment may also be an improvement based on any one or more of the second to seventh embodiments.
[0150] It is readily apparent that in this embodiment, weighted regression analysis is first employed to correct for confounding biases in non-randomized studies by assigning weights. This allows for the estimation of the pure clinical outcome differences resulting from the intervention in a genomically defined "simulated trial," ensuring the internal validity of the causal effect estimation. Subsequently, an innovative dose conversion mechanism based on molecular perturbation spectra is introduced to proportionally convert the effect estimates under the simulated environment into the actual drug intervention effect at a specific target dose. This causal analysis approach not only answers the causal question of "whether the drug is effective" but also, through a molecular-level scale, quantitatively answers the more clinically valuable question of "at what dose will it be effective," significantly enhancing the translatability and practicality of high-throughput omics data into actual clinical medication decisions.
[0151] Ninth Embodiment
[0152] The ninth embodiment of this application provides a clinical trial simulation method based on molecular perturbation spectra. The ninth embodiment is an improvement upon the first embodiment, specifically in that it introduces a permutation test to verify the reliability of the statistical significance of the intervention effect, and / or uses subgroup analysis to identify the target population most likely to benefit. This enhances the reliability and practical value of the entire causal inference conclusion from both the robustness of statistical inference and the accuracy of clinical application.
[0153] Optionally, in some embodiments, after determining the estimated intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group, the method further includes step S301:
[0154] Step S3011: Perform multiple random permutations on the clinical outcome data in the second population dataset, keeping the covariates unchanged after each permutation, and repeat the simulated clinical trial construction and effect quantification process.
[0155] Step S3012: Based on the effect value distribution obtained from multiple repeated calculations, construct the expected distribution of the effect index under the null hypothesis;
[0156] Step S3013: The reliability of statistical significance is assessed by comparing the difference between the original intervention effect estimate and the expected distribution.
[0157] For example, the clinical outcome data in the second population dataset can be permuted multiple times to generate a simulated negative control dataset, keeping all covariates constant throughout the process. Subsequently, for each permuted dataset, the entire process from constructing the simulated clinical trial to quantifying the effects is completely repeated.
[0158] Furthermore, based on a large number of effect estimates obtained through repeated calculations (usually no less than 100 times), a projected distribution of the effect index under the null hypothesis (i.e., the drug has no real effect) can be constructed. This distribution reflects the range of variation in effect estimates due to random factors in the absence of a real effect.
[0159] Furthermore, the reliability of statistical significance is assessed by comparing the original intervention effect estimates with the expected distribution of this null hypothesis. Specifically, if the original effect estimates deviate significantly from the null hypothesis distribution (e.g., fall outside its 95th percentile), it indicates that the observed effect is unlikely to be caused by random factors alone, thus enhancing the credibility of the results.
[0160] It should be noted that this embodiment may also be an improvement based on any one or more of the second to eighth embodiments.
[0161] It is readily apparent that in this embodiment, the permutation test constructs an empirical distribution of effect sizes when the null hypothesis (i.e., drug ineffectiveness) holds by randomly shuffling the association between clinical outcomes and intervention allocations and repeatedly simulating the experiment. By comparing the actual observed intervention effect estimates with this distribution, the obtained p-value can more reliably assess whether the effect is truly caused by the intervention, thus effectively eliminating the possibility of false positives caused by random fluctuations in the data or unknown confounding factors, ensuring the robustness of causal inference. Subgroup analysis, on the other hand, acknowledges and actively explores the heterogeneity of intervention effects in the population. By independently performing causal effect estimation in subgroups with different pre-defined characteristics, it identifies which specific populations have the strongest causal effect on the drug intervention. This not only answers the average causal question of "whether the drug is effective," but also more precisely reveals the heterogeneous causal effect of "which group is more effective," providing a direct decision-making basis for achieving precision medicine and selecting advantageous populations.
[0162] Tenth Embodiment
[0163] The tenth embodiment of this application provides an application example of a clinical trial simulation method based on molecular perturbation spectrum, which is implemented based on any one or more embodiments of the first to ninth embodiments described above.
[0164] In this application example, the complete workflow of this application is demonstrated by taking the discovery of the potential preventive effect of a certain drug A on lung cancer.
[0165] (I) Research Background and Objectives: Currently, the number of cancer prevention drugs approved by the U.S. Food and Drug Administration (FDA) is very limited, and there are no clear preventive drugs for lung cancer, which has the leading incidence and mortality rate. This study aims to estimate the combined causal effect of drug A on the incidence of lung cancer in a specific population using the methods provided in this application, and to discover drugs that can be used for the prevention of lung cancer.
[0166] (ii) Data sources and methods: The analysis was conducted using publicly available data from public databases, employing the complete technical process provided in this application. Specific data sources include: CMap database (drug perturbation spectrum), GTEx database (lung-specific eQTL data), and the UK Biobank (UKB) (genomic and clinical data).
[0167] (III) Specific Implementation Steps:
[0168] 1. Obtain the molecular perturbation spectrum of drug A.
[0169] Transcriptome data were obtained from the CMap database: Normal lung epithelial cell line NL20 was treated with drug A and DMSO solvent as a control group, and transcriptome data were obtained through sequencing. Differential expression analysis was used to identify differentially expressed molecules between the drug A intervention group and the control group, generating molecular perturbation spectra ∆X quantified by log2 fold change, as shown in Table 1.
[0170] Table 1. Molecular perturbation spectrum of drug A (partial)
[0171] Gene logFC AveExpr P.Value adj.P.Val ENSG00000072736 4.855541992 7.743549392 5.40E-26 5.48E-22 ENSG00000100906 -2.151128912 10.517644 3.27E-18 1.66E-14 ENSG00000170312 -2.832261014 12.03079169 1.72E-15 5.83E-12 ENSG00000175063 -2.579712796 13.09326955 5.08E-15 1.29E-11 ENSG00000131747 -2.304195333 13.58507195 1.61E-14 3.26E-11 ENSG00000175592 1.640500879 10.12071843 1.44E-10 1.82E-07 … … … … …
[0172] 2. Predicting individual gene expression profiles
[0173] Based on individual genomic data from the UKB cohort and lung-specific cis-eQTL and aFC data from the GTEx database, the gene expression profiles of individuals were predicted using an allelic fold change (aFC) model. The gene expression matrix of individuals in the UKB cohort was obtained by summing haplotype expression levels, as shown in Table 2.
[0174] Table 2. Gene expression matrix of individuals in the UKB cohort (partial)
[0175] eid ENSG00000072736 ENSG00000100906 ENSG00000170312 … 1000020 -0.87179 1.411863 0.849657 … 1000036 5.306195 0.751035 0.953541 … 1000048 -0.17786 -1.25741 0.450765 … 1000052 0.718018 0.544179 1.277431 … 1000077 -0.58065 -0.08686 -0.57378 … 1000084 -0.1657 -1.59182 0.348368 … … … … … …
[0176] 3. Component analysis and feature screening
[0177] Principal component analysis was performed on the predicted gene expression matrix to generate principal components (GPCs) of genetic variation prediction. The molecular perturbation spectrum ∆X was transformed to the principal component dimension to obtain the perturbation effect vector ∆X_GPC of the drug on GPCs, as shown in Tables 3 and 4. Statistical tests were used to screen GPCs that were significantly associated with the risk of lung cancer for subsequent analysis. See [link to table]. Figure 4 , Figure 4 The variables in the analysis refer to the various principal genetic components included in the analysis, such as GPC7, GPC18, etc.; the hazard ratio represents the fold change in the risk of clinical outcomes (such as disease occurrence or death) for each unit increase in the variable; the upper and lower limits specifically refer to the lower and upper limits of the 95% confidence interval for the hazard ratio; the p-value is used to test whether the hazard ratio is statistically significant. This figure shows the results of association analysis between a series of principal genetic components and a certain clinical outcome. For example, for the variable GPC7, its hazard ratio is 0.968, and the 95% confidence interval is 0.941-0.996, excluding 1, with a p-value of 0.023. This indicates that GPC7 is a protective factor, and the higher its expression level, the significantly lower the patient's disease risk. Conversely, GPC18 is a risk factor, and the higher its level, the significantly increased the risk.
[0178] Table 3. Principal components (GPCs) predicting genetic variation in individuals from the UKB cohort (partial list)
[0179] eid GPC1 GPC2 GPC3 GPC4 GPC5 GPC6 GPC7 … 1000020 -0.87179 1.411863 0.849657 0.165649 0.469572 -2.05527 -0.33739 … 1000036 5.306195 0.751035 0.953541 -0.33285 0.542753 0.327199 -0.31403 … 1000048 -0.17786 -1.25741 0.450765 0.16193 0.266246 -0.73705 0.091383 … 1000052 0.718018 0.544179 1.277431 -0.79907 -0.37686 -0.57483 -0.74573 … 1000077 -0.58065 -0.08686 -0.57378 -0.34743 -0.7534 -0.40501 -0.10033 … 1000084 -0.1657 -1.59182 0.348368 -1.33336 0.338119 0.777596 0.930816 … … … … … … … … … …
[0180] Table 4. Perturbation effects of drug A on various GPCs (partial)
[0181] GPC1 GPC2 GPC3 GPC4 GPC5 GPC6 GPC7 … -1.4712 -1.96704 -3.11844 -2.43985 0.294056 0.990269 0.832836 …
[0182] 4. Constructing a simulated clinical trial
[0183] Based on the individual GPCs and drug perturbation effect vectors of UKK individuals, an individual allocation model is constructed, such as... For details, please refer to the technical solution described in the sixth embodiment. By minimizing the objective function (based on the square of the correlation coefficient or Euclidean distance), the difference between the two groups in GPCs in the simulation experiment is made consistent with the drug perturbation effect. See [link to relevant documentation]. Figure 5 In the figure, most data points (such as GPC269 and GPC218) fall in the lower right quadrant (RNA expression is downregulated by the drug, and this downregulation shows a protective effect in the simulation) and the upper left quadrant (RNA expression is upregulated by the drug, and this upregulation shows a risk effect in the simulation). This pattern suggests that the direction of molecular changes induced by the drug in cells is consistent with the direction of the causal effect of that molecule on clinical outcomes inferred from genetic data.
[0184] Furthermore, a pseudo-population was constructed using a weighted method, creating two weighted observation records for each individual, forming a simulated clinical trial dataset (see Table 5); See also... Figure 6 This visually illustrates the probability distribution of all individuals in the second target population being assigned to the simulated trial group based on their genetic characteristics. Each point represents an individual. The position of the Y-axis directly reflects the individual's propensity to be assigned to the simulated trial group. A higher weight (higher point position) indicates a closer match between the individual's genetic profile and the drug's expected mechanism of action, thus increasing the likelihood of being assigned to the simulated intervention group in the simulated trial.
[0185] Table 5. Machine learning-assigned scores for individuals in the UKB cohort (partial)
[0186] eid ML-allocation score 1000020 0.937488 1000036 0.328944 1000048 0.270261 1000052 0.542009 1000077 0.102387 1000084 0.075815 1000095 0.000885 1000106 0.886869 … …
[0187] 5. Effectiveness Evaluation and Verification
[0188] The weighted Cox proportional hazards model was used to analyze simulated clinical trial data. The results showed that drug A significantly reduced the risk of lung cancer. (See [link to relevant documentation]). Figure 7The figure clearly shows two survival curves: green represents the simulated intervention group of drug A in the simulation experiment, and red represents the simulated control group. Throughout the time axis, the survival curve of the simulated experimental group is consistently above that of the control group. This indicates that throughout the follow-up period, individuals assigned (or weighted) to the experimental group consistently had a higher probability of survival (or event-free survival) than those in the control group.
[0189] Further, see Figure 8 It clearly presents the estimated intervention effect of the target drug in simulated clinical trials and its statistical reliability.
[0190] Furthermore, to verify the reliability of the results, a negative control analysis can be performed: by randomly permuting the clinical outcome data (at least 10 simulations), an effect distribution under the null hypothesis is constructed to assess the statistical significance of the original effect. See details below. Figure 9 As shown, in the 10 negative control analyses, the hazard ratio confidence intervals for the vast majority of replacement groups (Permut_1–Permut_9) spanned 1, with corresponding p-values greater than 0.05, indicating that these simulated effects were not statistically significant; only Permut_10 had a p-value of 0.042 (close to 0.05). In contrast, the hazard ratio for the original variable, Evodiamine, was 0.917 (95% CI: 0.870–0.967), with a p-value of 0.001, showing a significant statistical difference. This result further supports the reliability of the original analysis findings, suggesting that the observed effects are unlikely to be caused by random factors.
[0191] (iv) Research Results
[0192] Through the complete technical process provided in this application, it was discovered that drug A has a potential effect in preventing lung cancer, providing a new candidate drug for the drug prevention of lung cancer, and at the same time verifying the practical value of this application in the field of drug discovery.
[0193] As can be seen from the above embodiments, the core advantage of this application lies in: replacing real drug use records with molecular perturbation spectra, correcting data defects through genetic instrumental variables, and solving heterogeneity problems through molecular-level grouping optimization, thus providing a complete technical solution for early-stage innovative drug development and precision medicine practices. Specifically:
[0194] 1. Existing simulation technologies rely on actual efficacy data of new drugs in populations, while this application only needs to simulate clinical trial results based on molecular perturbation spectra obtained from basic experimental stages of drugs (such as cells, animals, or a small number of subjects) and combined with public genomic databases (such as GWAS), making it easier to implement. Based on this, the method can be applied to scenarios that are difficult to cover by traditional simulation technologies, such as unmarketed new drugs and traditional Chinese medicine compound prescriptions. Therefore, its application scope is wider than that of existing technical solutions.
[0195] 2. Existing simulation methods often employ PSM matching or IPTW weighting to balance known confounding factors, but they cannot correct for unobserved confounding. This application introduces the concept of instrumental variables, using genetic variations (such as SNPs) as instrumental variables to achieve simulation randomization. This effectively controls various biases, including unobserved confounding, and improves the reliability of causal effect estimation.
[0196] 4. This application ingeniously combines the genetic instrumental variable technique in "Mendelian randomization" with clinical trial simulation, combining the advantages of both methods: it retains the intuitiveness and interpretability of clinical trial simulation, while possessing the rigor and robustness of genetic methods in causal inference.
[0197] 5. This application supports the flexible definition of subgroup populations based on multiple dimensions such as gender, age, genotype, and disease stage, enabling heterogeneity analysis of drug effects. This function helps to accurately identify populations with significant benefits, optimize population-drug matching strategies, and provide a basis for personalized treatment and precise positioning of new drugs.
[0198] In summary, this application demonstrates significant advancements in data requirements, method design, and application flexibility, providing a more efficient, robust, and scalable new clinical trial simulation pathway for drug development.
[0199] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0200] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0201] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 10An exemplary structural diagram of the electronic device is disclosed. The electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0202] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103 and output device 1104 may be connected by a bus or other means, as shown in the figure, which is connected by a bus.
[0203] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0204] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback); and input from the user can be received in any form (e.g., voice input or tactile input).
[0205] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.
[0206] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.
[0207] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0208] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0209] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, read-only optical discs, digital versatile optical discs or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0210] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0211] In the above embodiments, all or part of the implementation can be achieved through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0212] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0213] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0214] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.
[0215] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A clinical trial simulation method based on molecular perturbation spectrum, characterized in that, The method includes: Determine the molecular perturbation spectrum of the target drug; Based on the genomics and high-dimensional omics data in the first population dataset, a predictive model from genomics to high-dimensional omics was determined. Based on the prediction model and the genomic data of individuals in the second population dataset, the predictive omics data of the individual are determined; wherein, the second population dataset is also associated with clinical outcome data; Based on the molecular perturbation spectrum and the predictive omics data of all individuals, the grouping scheme for the simulated clinical trial is determined by optimizing the individual allocation model to maximize the similarity between the predicted omics differences and the molecular perturbation spectrum between the simulated experimental group and the simulated control group; and the simulated clinical trial is constructed according to the grouping scheme. A simulated clinical trial was constructed based on the grouping scheme, and the estimated intervention effect of the target drug was determined based on the difference in clinical outcomes between the simulated experimental group and the simulated control group in the simulated clinical trial.
2. The method according to claim 1, characterized in that, The determination of the molecular perturbation spectrum of the target drug includes: Perform differential analysis on high-dimensional omics data before and after drug intervention to determine the results of the differential analysis; Based on the differential analysis results, a molecular perturbation feature vector consisting of differentially expressed molecules and their variation amplitudes is determined.
3. The method according to claim 1, characterized in that, The step of determining a predictive model from genomics to high-dimensional omics based on genomics and high-dimensional omics data from the first population dataset includes: Based on the principle of allele fold change, a quantitative relationship between genotype and gene expression level is determined; and a prediction model is constructed based on the quantitative relationship. or, Based on the aggregated statistical data of tissue-specific quantitative trait loci, the association weights between genetic variations and omics characteristics are determined; a linear prediction model is then constructed based on these association weights.
4. The method according to claim 1, characterized in that, After determining the predictive omics data of an individual based on the predictive model and the genomic data of an individual in the second population dataset, the method further includes: Based on the predicted omics data, a set of low-dimensional principal component features were determined through principal component analysis. The projection of the molecular perturbation spectrum onto the low-dimensional principal component feature space is determined and used as the drug perturbation principal component vector.
5. The method according to claim 4, characterized in that, After determining the projection of the molecular perturbation spectrum onto the low-dimensional principal component feature space as the drug perturbation principal component vector, the method further includes: Based on the statistical test results, a subset of features that meet the preset significance level with respect to clinical outcomes and / or drugs is determined from the low-dimensional principal component features; In the step of determining the grouping scheme for the simulated clinical trial by optimizing the individual allocation model based on the molecular perturbation spectrum and the predictive omics data of all individuals, only the feature subset is used.
6. The method according to claim 1, characterized in that, The step of determining the grouping scheme for the simulated clinical trial by optimizing the individual allocation model based on the molecular perturbation spectrum and the predictive omics data of all individuals includes: Based on the genomic data of individuals in the second population dataset, the allocation score for each individual is determined using an individual allocation model; Based on the assigned scores, individuals are assigned to a simulated trial group and a simulated control group to construct an initial simulated clinical trial; Based on the initial simulated clinical trial, the grouping scheme for the simulated clinical trial is determined by optimizing the individual allocation model.
7. The method according to claim 6, characterized in that, The process of determining the grouping scheme for the simulated clinical trial by optimizing the individual allocation model based on the initial simulated clinical trial includes: In the initial simulated clinical trial, omics differential characteristics were determined based on the predicted omics data of the simulated experimental group and the control group; Calculate the similarity between the omics differential features and the molecular perturbation spectrum; To maximize the similarity, the individual allocation model is optimized to obtain the optimal individual allocation model.
8. The method according to claim 6, characterized in that, The individual allocation model is specifically a machine learning function whose output is the probability weight of an individual being assigned to the simulated experimental group, expressed as: Where Gi represents the genomic data of the i-th individual or the principal component features of the genomic data after dimensionality reduction; is the probability weight; A represents the intervention allocation, A=1 means allocation to the simulated experimental group, and A=0 means allocation to the simulated control group.
9. The method according to claim 8, characterized in that, The individual allocation model is specifically implemented using a composite model of linear combination and the Sigmoid function, expressed as: ;in, ,GPC2,…,GPC n The first to nth principal component features obtained after dimensionality reduction by principal component analysis of the genomic data; α1, α2, ..., α n Assign the parameter weights that the model needs to learn to the individual; This is the Sigmoid function.
10. The method according to claim 7, characterized in that, The calculation of the similarity between the omics differential features and the molecular perturbation spectrum is specifically achieved through any of the following methods: Based on the squared correlation coefficient between the omics differential features and the molecular perturbation spectrum; or, based on the Euclidean distance between the omics differential features and the scaled molecular perturbation spectrum; Specifically, the squared correlation coefficient between the omics differential characteristics and the molecular perturbation spectrum is achieved through the following formula: Specifically, the Euclidean distance between the omics differential features and the scaled molecular perturbation spectrum is achieved through the following formula: , This represents the vector of perturbation effects of drugs on the principal components of the genome. This represents the difference vector of principal components in the genome between the simulated intervention group and the simulated control group in the simulation experiment. This is the square of the correlation coefficient between the two vectors; It is an adaptive scaling factor.
11. The method according to claim 1, characterized in that, The step of determining the estimated intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group in the simulated clinical trial includes: We used weighted regression analysis, with the assigned weights as the weighting values, to calculate the differences in clinical outcomes between the simulated experimental group and the simulated control group. Based on the proportional relationship between the molecular perturbation spectrum and the simulated omics differences, the differences in clinical outcomes are converted into estimated values of drug intervention effects at specific doses.
12. The method according to any one of claims 1 to 11, characterized in that, After determining the estimated intervention effect of the target drug based on the difference in clinical outcomes between the simulated experimental group and the simulated control group, the method further includes: The clinical outcome data in the second population dataset are subjected to multiple random permutations, with the covariates remaining unchanged after each permutation, and the simulated clinical trial construction and effect quantification process is repeated. Based on the effect value distribution obtained from multiple repeated calculations, the expected distribution of the effect index under the null hypothesis is constructed. The reliability of statistical significance is assessed by comparing the original intervention effect estimates with the expected distribution.
13. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as described in any one of claims 1 to 12.
14. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 12.