A method, apparatus, medium, and program product of constructing a digital organoid
By constructing a GDnet digital organoid model and integrating transcriptome and drug information, the problems of flexibility and accuracy in predicting neoadjuvant therapy for breast cancer in existing technologies have been solved, enabling more precise treatment selection and drug response prediction, and improving the treatment effect of breast cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
- Filing Date
- 2025-01-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for predicting neoadjuvant therapy in breast cancer are based on small sample sizes and single treatment strategies, making it difficult to effectively guide patients in choosing treatment options in real-world scenarios. Furthermore, they lack the flexibility to simulate drug-tumor interactions, resulting in poor predictive performance.
We constructed a deep learning-based digital organoid model, GDnet, and integrated transcriptome data and drug information. By using a bioinformatics neural network to simulate the interaction between drugs and tumors, we established a general predictive model for accurately predicting the response to neoadjuvant therapy in breast cancer.
It improves the predictive accuracy and clinical decision optimization capabilities of neoadjuvant therapy for breast cancer, can be widely applied in different treatment environments, significantly improves the pCR rate, and supports treatment options for precision medicine.
Smart Images

Figure CN120809051B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent healthcare, and more specifically, to a method, apparatus, medium, and program product for constructing digital organoids. Background Technology
[0002] Existing methods for predicting neoadjuvant therapy in breast cancer patients are based on models constructed using genomics, transcriptomics, radiomics, pathology, or even multiple omics. These models are developed based on small sample sizes and single treatment strategy cohorts, and cannot be well applied to real-world scenarios to directly guide patients' treatment options.
[0003] In theory, predicting the efficacy of cancer treatments requires addressing two aspects. First, a more comprehensive characterization of tumor markers is crucial for better assessing their functional status. This task necessitates integrating as many omics data types as possible, including transcriptomics, genomics, proteomics, and pathological information, to achieve a more comprehensive tumor characterization and improve the predictive performance of the model. However, as the number of features increases, the amount of samples required for training also grows significantly. Currently, the availability of patients with multi-omics biological information, apart from transcriptomics data, is limited, and existing large-sample-based models are primarily based on transcriptomics data. Therefore, a comprehensive characterization of transcriptomics status is essential. Due to the flexibility of neural networks, deep learning methods allow the use of as many genes as possible to optimally characterize the tumor state, capturing not only tumor heterogeneity but also information about the tumor microenvironment. However, deep learning methods require a large number of samples, and their results are more difficult to interpret. Integrating prior biological knowledge can facilitate the construction of sparse, high-performance, and interpretable deep learning models.
[0004] The second aspect involves the appropriate characterization of drugs to accurately simulate the drug-tumor interaction, which is crucial for predicting tumor response to treatment. To develop a model that can better guide clinical drug decisions in precision medicine, it is necessary to combine as many different treatment types as possible to learn complex drug-tumor interaction patterns. To incorporate therapeutic drugs into the model, three important factors need to be considered: (1) The types of anticancer drugs used in clinical practice are relatively limited (less than one hundred) and structurally diverse (e.g., small molecules, antibodies). Given the limited number of types of breast cancer therapeutic drugs, it is impossible to understand the structural characterization of compounds, and it is difficult to integrate compound structural data with transcriptomic status. The previous strategy of directly characterizing compounds based on their structures is not suitable for this task. (2) Many chemotherapeutic drugs lack specific targets, and the mechanisms of action of targeted therapies often go beyond the primary target (e.g., antibody-dependent cell-mediated cytotoxicity-related genes in anti-HER2 therapy). Therefore, previous drug characterization based on specific gene targets cannot fully represent all types of drugs or comprehensively characterize their mechanisms. (3) Clinically, the combined use of drugs leads to complex synergistic or antagonistic interactions, requiring flexible model structures to simulate these interactions. Therefore, existing frameworks based solely on late-stage fusion integration cannot achieve optimal efficacy prediction. Generally, a unified genome-wide approach is needed to characterize different types of drugs and explain the complex interactions in combination therapies. Summary of the Invention
[0005] This invention aims to at least address one of the technical problems existing in the prior art. To this end, this invention provides a method, apparatus, medium, and program product for constructing digital organoids, and also develops methods, apparatus, medium, and program products for predicting treatment regimens and / or therapeutic drug sensitivity based on this model. The method of this invention establishes a general deep learning model based on treatment regimens and transcriptome data to accurately predict neoadjuvant therapy responses in various treatment regimens for breast cancer and optimize clinical decision-making.
[0006] The first aspect of this application discloses a method for constructing digital organoids, the method comprising: 101. Obtain the optional treatment regimens, real treatment regimens, and transcriptome data of the training set samples; the treatment regimens include regimens consisting of at least one therapeutic drug; 102. Input the optional treatment options and transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after each optional treatment option. 103. The optional treatment options with a predicted pCR probability greater than the first threshold are assigned to the best group, and the optional treatment options with a predicted pCR probability less than the first threshold are assigned to the worst group. 104. Based on whether the real treatment plan belongs to the best group or the worst group, the training set samples are classified. If the real treatment plan is in the best group, the sample is assigned to the optimization group; if the real treatment plan is in the worst group, the sample is assigned to the control group, thus obtaining digital organoids including the optimization group and the control group.
[0007] In some embodiments, the method for constructing the GDnet model includes: Obtain whole-genome data and real treatment regimens from training set samples that have received neoadjuvant therapy; the therapeutic drugs are represented by Gcd, a gene sensitivity correlation of pan-cancer cell lines, in a cancer drug sensitivity genomics database; The whole-genome data and real treatment protocols were input into a bioinformatics neural network model for training, resulting in the GDnet model; this is a typical binary classification model. The neural network, after a sigmoid transformation, outputs the probability of a pCR (progression response). This can be compared to logistic regression, which is also a binary classification model.
[0008] A second aspect of this application discloses a method for predicting the sensitivity of treatment regimens, the method comprising: 201. Obtain transcriptome data and candidate treatments from the subjects; 202. Input transcriptome data and candidate treatment plans into the digital organoids described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained.
[0009] A third aspect of this application discloses a method for predicting the sensitivity of treatment regimens, the method comprising: 301. Obtain transcriptome data and candidate treatment regimens for the subjects; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 302. Input transcriptome data and candidate treatment plans into the digital organoids described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained. Optionally, the drug types in the candidate treatment regimens include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, and ENG. Optionally, the method further includes: assisting in the selection of a treatment plan based on the classification result; if the classification result indicates that the subject belongs to the optimized group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is suitable; if the classification result indicates that the subject belongs to the control group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is unsuitable. Optionally, if the candidate treatment regimens include any one or more of paclitaxel|doxorubicin|cyclophosphamide|olaparib|PDCD1, 5-fluorouracil|epirarubicin|cyclophosphamide|lapatinib|ERBB2, 5-fluorouracil|epirarubicin|cyclophosphamide|ERBB2 and paclitaxel|ERBB2ERBB2, paclitaxel|doxorubicin|cyclophosphamide|ERBB2MK-2206, paclitaxel|doxorubicin|cyclophosphamide|PDCD1 and paclitaxel|doxorubicin|cyclophosphamide|ENG|ERBB2, an auxiliary prediction result with a high probability of the subject belonging to the optimized group is obtained; Optionally, if the candidate treatment regimen includes any one or more of docetaxel, 5-fluorouracil, epirubicin, cyclophosphamide, docetaxel, and doxorubicin, an auxiliary predictive result indicating the high probability that the subject belongs to the control group is given.
[0010] A fourth aspect of this application discloses a method for predicting the sensitivity of therapeutic drugs, the method comprising: 401. Obtain transcriptome data and candidate therapeutics from the subjects; 402, Input transcriptome data and candidate therapeutic drugs into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained. This application discloses a fifth aspect of a method for predicting the sensitivity of a therapeutic drug, the method comprising: 501. Obtain transcriptome data of the subjects and candidate therapeutic drugs; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 502, Input transcriptome data and candidate therapeutic drugs into the digital organoids described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained.
[0011] This application has the following beneficial effects: 1. This application innovatively discloses a method for constructing digital organoids. It utilizes existing and real-world treatment options, correlates them with transcriptomic data, calculates predicted pCR probabilities, and groups the available treatment options based on these probabilities. The training set samples are then classified based on the grouping results to obtain digital organoids. In routine clinical practice, the model can help select the most suitable treatment from candidate options; in clinical trials, the model can help select patients who may respond to new testing protocols or even previously abandoned protocols, and the model can still predict drug response even if the treatment includes new drugs not included in the training set.
[0012] 2. This application innovatively discloses the construction process of the GDnet model. By fully integrating drug and transcriptome data, drugs are mapped to a series of gene-value pairs, and different methods are used to characterize small molecule drugs and antibody drugs respectively. Based on this, whole-genome data of breast cancer patients who have received neoadjuvant therapy and real treatment regimens are input into the bioinformatics neural network model for training, establishing a general deep learning model as a digital organoid that can simulate drug-tumor interactions to accurately predict neoadjuvant therapy responses in various treatment regimens and optimize clinical decision-making. In this approach, as many transcriptome samples and diverse treatment data as possible are collected, and as many genes as possible are used to characterize tumor status. Various methods for integrating drug and transcriptome data to simulate drug-tumor interactions are explored. Bioinformatics neural networks (a cutting-edge deep learning technology) are applied to achieve optimal predictive performance with a limited sample size, and the model's ability to support clinical decision optimization is explored. This research will significantly advance precision neoadjuvant therapy selection and improve the treatment outcomes of breast cancer.
[0013] The application of GDnet in two series of simulated clinical trials demonstrates that GDnet can serve as a digital organoid to optimize treatment options for breast cancer patients. Compared to previous models, GDnet offers the following advantages: (1) By integrating drug data into the model, GDnet is able to compare all available treatment options before treatment begins, thereby promoting precision medicine specifically for breast cancer patients, something that previous machine learning models could not achieve correctly. In other words, we have built a model that considers organoid function, can test the sensitivity of drugs and treatment options, and inform patients of appropriate drugs. The results of computer simulations based on the ISPY-2 trial and external validation set in our study show that GDnet can be used to optimize the treatment decision-making process for breast cancer patients and significantly improve the pCR rate for breast cancer patients. For each optimization threshold, GDnet maintains its ability to improve the pCR rate, thus confirming the results. In addition, there is a linear correlation between the pCR rate and the optimization intensity, which strongly supports our conclusion that GDnet can optimize neoadjuvant therapy for breast cancer and has the potential to change relevant treatment guidelines. Our model can play an important role in many potential scenarios. In routine clinical practice, our model can help select the most suitable option from candidate options. In clinical trials, our model can help select patients who may respond to new test options or even previously abandoned options. Even when treatment regimens include novel drugs not included in the GDnet training set, our model is still able to predict drug responses because it performs well on a validation set containing many drugs, such as durvalumab, bevacizumab, cisplatin, olaparib, and gemcitabine, which are not present in the training set.
[0014] (2) Most previous machine learning models have faced problems such as limited sample size and lack of sufficient independent external validation leading to unstable performance. This application collected 31 datasets containing 4371 breast cancer samples, which included available pre-treatment transcriptome and neoadjuvant therapy information. This is the largest dataset to date, used to robustly train and validate GDnet; (3) Data were collected from all types of RNA sequencing methods by applying a ranking normalization strategy to each sample. This approach provides robustness against technical artifacts that could otherwise introduce systematic bias into absolute transcript counts, while keeping the overall relative ranking of genes within each cell at a more stable level. Previous studies have been limited to a limited number of RNA sequencing methods due to their ease of normalization and reduction of batch effects; however, these methods tend to reduce the available sample size, making it difficult to generalize to a wider range of scenarios. While ranking-based encoding has limitations, including underutilization of the precise gene expression measurements provided in transcript counts, it makes the model more suitable for realistic conditions with a variety of unexpected variables. Furthermore, unlike previous studies that normalized between patients in different cohorts or testing methods (e.g., z-score transformation) or even between datasets (e.g., Combat algorithm), this approach only includes standardization for each patient. This design choice stems from the understanding that normalization parameters trained across patients or datasets cannot be effectively applied to patients in another dataset with different transcriptome testing protocols, leading to overfitting and poor generalization. Therefore, our model can be directly generalized to new cases tested using any RNA sequencing technology in a wider range of scenarios and complex realistic conditions; (4) Previous models, such as the 21-gene recurrence score, MammaPrint, and Adjutorium, were mainly based on traditional statistical or machine learning models. In contrast, GDnet, based on a deep learning framework, includes a larger number of genes to maximize model capacity. Given that achieving pCR means eliminating all tumor cells, it is crucial to maintain sufficient flexibility in the model to account for tumor heterogeneity and microenvironmental changes, a challenge that traditional machine learning models cannot address. Deep learning frameworks, with their large number of parameters and flexible architecture, promise to solve these challenges. A previous study showed that deep learning methods can reflect or predict tumor heterogeneity and the tumor environment. Furthermore, by integrating prior biological knowledge, GDnet significantly reduces the number of learning parameters, resulting in more stable and better performance than dense models, a point also demonstrated in previous studies. The visualization of the importance and properties of multi-level genes and biological pathways in GDnet enables a multi-level view of the model's interpretation, which can guide researchers to propose hypotheses about the underlying biological processes involved in drug resistance and translate these findings into therapeutic opportunities. Specifically, genes identified by GDnet, such as SRC, CCND1, MCF2L, RPS6KA1, and PSMB7, may play an important role in breast cancer resistance.
[0015] (5) GDnet innovatively integrates treatment options for breast cancer, capturing the interaction patterns between drugs and tumors, thereby improving model performance. Although a few neoadjuvant therapy models for breast cancer consider specific drug types, such as chemotherapy and anti-HER2 therapy, these models often lack the flexibility required to simulate tumor-drug interactions. Previous studies have explored representative drugs based on drug structures; however, these approaches have limitations because they lack sufficient drug diversity for comprehensive training of breast cancer neoadjuvant tasks to represent antibody drugs, or target-based drugs, but target-based drug representation lacks the flexibility to simulate drug-tumor interactions. To address these shortcomings, a simple approach using gene expression correlations to represent drugs is introduced, which helps to incorporate external knowledge related to drug resistance. This approach provides a drug representation structure similar to gene expression in tumors, which helps to fuse and create the potential to simulate drug-transcriptome interactions. We also demonstrate that a better approach to integrating drug representation and transcriptome involves feeding data from both modalities into a neural network in parallel, rather than simply fusing them through specially designed rules or aggregating predictions from drug-specific models; In summary, GDnet is a bioinformatics deep learning approach that integrates drug characterization and transcriptomics to achieve more precise oncology and enhance the treatment decision-making process for neoadjuvant therapy in breast cancer. Integrating drug characterization and bioomics data represents a novel approach to simulating drug-tumor interactions, providing a digital organoid platform for precision oncology that is broadly applicable to diverse treatment environments and various cancer types. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the method flow provided in the first aspect of the present invention; Figure 2 This is a schematic diagram of the method flow provided in the second aspect of the present invention; Figure 3 This is a schematic diagram of the method flow provided in the third aspect of the present invention; Figure 4 This is a schematic diagram of the method flow provided in the fourth aspect of the present invention; Figure 5 This is a schematic diagram of the method flow provided in the fifth aspect of the present invention; Figure 6This is a schematic diagram of a computer device provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the architecture of an exemplary computing device provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the storage medium provided in an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating the construction process of digital organoids provided in the embodiments of the present invention, and the effect of computer clinical trial simulation in optimizing neoadjuvant treatment decisions for breast cancer. Figure 9 a. Schematic diagram of a computer-based clinical trial. Patient transcriptomes and multiple alternative treatments are input into the model to predict the pCR probability for each option. The top N% and bottom N% regions with the highest and lowest pCR probabilities are selected as the best and worst options, respectively. If a patient's treatment option (indicated by a white star "*") is rated as the best, the patient is assigned to the optimization group; if ranked as the worst, they are assigned to the control group. N represents the optimization intensity of the simulated trial. pCR rates of the optimization and control groups in a simulated trial based on the I-SPY2 trial after propensity score matching (PSM) at different optimization intensity levels. Figure 9 b) and simulation experiments based on external validation datasets after PSM ( Figure 9 d). Differences in pCR rates were tested using Fisher's exact test, with p-values marked on the solid green line in the optimized group. The odds ratio of pCR in the optimized group compared to the control group in the I-SPY2-based simulation after PSM ( Figure 9 c) and simulation experiments based on external validation datasets after PSM ( Figure 9 e). The dashed line represents the fitted linear regression function between the y-axis and the optimization threshold; R² and p-values are shown at the top. Opt: Optimization group; Ctl: Control group; Figure 10 This is a schematic diagram illustrating the fusion of drug representation and transcriptome and protocol representation provided in an embodiment of the present invention. Figure 10 'a' represents the strategy of representing drugs through gene relevance. Small molecules represent significant genome-wide relevance with the half-maximal inhibitory concentration (IC50) value, while antibody drugs represent significant genome-wide relevance with the target gene. Figure 10 b represents the GDnet framework, a biological deep learning model that integrates transcriptome and protocol representations. Transcriptome data is sorted and normalized, and drug representations in the treatment protocol are summed genetically to be input into the neural network. Binary masks containing information about the dependencies between genes and / or pathways are incorporated into the network to impose constraints on nodes and edges; IC50: half-maximal inhibitory concentration; Gcd: correlation between gene and drug sensitivity; Ge: gene expression; Figure 11This is a heatmap visualization of the gene correlation of each drug provided in this embodiment of the invention, with correlation values represented by color coding. Rows and columns are clustered using consensus clustering, and drugs are annotated by type. PDCD1 represents anti-PD-1 therapy (pembrolizumab and durvalumab), ANG represents ANG peptide inhibitor (AMG-386), VEGFA represents anti-VEGFA antibody (bevacizumab), IGF1R represents anti-IGF1R antibody (ganitumab), capecitabine is represented by 5-fluorouracil, Ganetspib is represented by luminespib, and ixaspirone is represented by epoch-yam B. Figure 12 This report compares the computational performance of GDnet, provided in this embodiment of the invention, with other models that have external validation sets. The performance metrics reported include the area under the ROC curve (AUC). Figure 12 a) and the area under the precision-recall curve (AUPRC) Figure 12 b). In terms of both metrics, GDnet's median significantly outperforms other models. Figure 12 a- Figure 12 The data in b are represented as box plots, where the middle line is the median, the lower and upper hinges correspond to the 1st and 3rd quartiles, respectively, and the dashed lines correspond to the minimum or maximum values within 1.5 × IQR (where IQR is the quartile range) of the hinge. Any data points outside the dashed lines are considered outliers. For all comparisons between GDNet and other models, and between GDNNet and GANNNet, p < 0.05. Figure 12 c represents the ROC curve of GDnet compared to other models (n=912). Figure 12 The bar chart for d shows the AUC of GDNet and other models integrating biological knowledge on each individual GEO dataset in the test set. The datasets are sorted by sample size, with the largest on the left and the smallest on the right. All models are ( Figure 12 c) and ( Figure 12 d) 10 repeated integrations; Figure 13 This diagram shows the proportions of the optimized group (middle) and the control group (right) provided in this embodiment of the invention, the difference in proportions between the two groups (left), and the external validation set results of simulation experiments based on different optimization thresholds. The optimization thresholds are color-coded. Opt: Optimized group; Ctl: Control group; Figure 14 This is an overview of the model provided in this embodiment of the invention for comparison with GDnet. Figure 14a. The GAnet framework integrates drug representation and transcriptome data. First, transcriptome data and protocol representation are aggregated using a function, where Gatt represents the attention value of gene expression upon treatment. Then, the Gatt values are input into... Figure 10 In the biological information neural network described in b. Figure 14 b. Gnetens framework. A separate transcriptome-only model was constructed for each drug type, and the predictions of all transcriptome-only models corresponding to drugs administered to all patients were integrated by averaging to produce the final prediction. Figure 14 c. The Gnet framework is only for transcriptome models. Transcriptome data is directly input into... Figure 10 In the bioinformatics neural network described in section b, Gcd represents the correlation between genes and drug sensitivity; Ge represents gene expression; and Gatt represents the gene expression attention value. Figure 15 This is a heatmap showing all patient treatment drug combinations provided in this embodiment of the invention. Applied drugs are marked with 1 and color-coded in red, while unapplied drugs are marked with 0 and color-coded in blue. ER status, HER2 status, and pCR results are annotated at the top, with white representing missing values. Figure 16 This embodiment of the invention compares the computational performance of GDNet with other models using 10x cross-validation on the training set. Metrics include the area under the ROC curve (AUC). Figure 16 a) Area under the precision-recall curve (AUPRC) Figure 16 b). Figure 16 a and Figure 16 The data in b is represented as a box plot, where the middle line is the median, the lower and upper hinges correspond to the first and third quartiles respectively, and the whisker line corresponds to the distance from the upper... Figure 1 The minimum or maximum value hinge within the 0.5 × IQR range (where IQR is the interquartile range). Data points outside the whisker are considered outliers. No statistical significance was observed in any combination of comparisons; Figure 17 This is a comparison of the computational performance of GDnet provided in this embodiment of the invention with other models that have external validation sets. Figure 17 a. The bar chart shows the AUC of all models for each individual GEO dataset in the test set. Figure 17 b. The bar chart shows the AUPRC for all models in each individual GEO dataset in the test set. The datasets are sorted by sample size. Figure 17 a- Figure 17 In b), the left side has the largest value, and the right side has the smallest. Both models are ensembles with 10 repetitions. Figure 18This is a computer-based clinical trial simulation based on the ISPY-2 trial provided in this embodiment of the invention. In propensity score matching (PSM)... Figure 18 a) before and PSM ( Figure 18 e) Subsequently, the sample size for each group varied with the optimization intensity (N%). Baseline tumor transcriptome status (inferred from Gnet prediction scores) Figure 18 b) Within the optimization group and control group at different optimization thresholds before and after PSM ( Figure 18 f), the difference in Gnet scores between the two groups was tested using the p-value marked on the green solid line of the optimized group, by means of the Wilcoxon grading test. Figure 18 c. pCR rates of the optimization and control groups under different optimization intensity levels. Differences in pCR rates were tested using Fisher's exact test, with p-values marked on the solid green line in the optimization group. (The pCR rate difference was determined using PSM...) Figure 18 d) and treatment-weighted inverse probability (IPTW) Figure 18 g) Calculate the pCR ratio between the optimized group and the control group. The dashed line represents the fitted linear regression function between the y-axis and the optimization threshold, R0. 2 The p-value is displayed at the top.
[0018] Figure 19 This is a computer-based clinical trial simulation based on an external validation dataset provided in this embodiment of the invention. In propensity score matching (PSM)... Figure 19 a) before and PSM ( Figure 19 e) Subsequently, the sample size for each group varied with the optimization intensity (N%). Baseline tumor transcriptome status (inferred from Gnet prediction scores) Figure 19 b) Within the optimization group and control group at different optimization thresholds before and after PSM ( Figure 19 f), the difference in Gnet scores between the two groups was tested using the p-value marked on the green solid line of the optimized group, by means of the Wilcoxon grading test. Figure 19 c. pCR rates of the optimization and control groups under different optimization intensity levels. Differences in pCR rates were tested using Fisher's exact test, with p-values marked on the solid green line in the optimization group. (The pCR rate difference was determined using PSM...) Figure 19 d) and treatment-weighted inverse probability (IPTW) Figure 19 g) Calculate the pCR ratio between the optimized group and the control group. The dashed line represents the fitted linear regression function between the y-axis and the optimization threshold, R0. 2 The p-value is displayed at the top. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0020] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Figure 1 This is a schematic flowchart of a method for constructing digital organoids according to an embodiment of the present invention. Specifically, the method includes the following steps: 101: Obtain optional treatment regimens, real treatment regimens, and transcriptome data from the training set samples; the treatment regimens include regimens consisting of at least one therapeutic drug; optionally, the therapeutic drugs include any one or more of the following: small molecule drugs, antibody drugs; the small molecule drugs are represented by calculating the IC50 value and the correlation between genes significantly associated with the small molecule drugs in the whole genome data, denoted as Gcd; specifically, the representation of the small molecule drugs includes: calculating the IC50 value of the small molecule drug (used to represent drug sensitivity, where a high IC50 corresponds to low cell-killing ability, and a low IC50 corresponds to high cell-killing ability; the lower the IC50 value, the higher the sensitivity); determining genes significantly associated with the small molecule drugs in the whole genome data based on the IC50 value, denoted as Ge; calculating the correlation between the IC50 value and the significantly associated genes, denoted as gene-drug sensitivity correlation Gcd; Gcd is used to characterize the small molecule drugs; Optionally, if the genomics database lacks a first small molecule drug, the most relevant drug within the database that has a similar mechanism of action and is of the same type of antitumor drug is used to represent the first small molecule drug; Optionally, if the genomics database includes at least two databases, and the second small molecule drug exists in at least two databases simultaneously, the database with higher recognition is preferred for calculating the representation of the second small molecule drug; for example, when the database includes GDSC and CTRP datasets, the GDSC dataset is used. Optionally, the small molecule drug includes any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, and ENG; optionally, the drug does not include drugs for which there is no corresponding drug sensitivity data in the GDSC or CTRP database, such as T-DM1 and endocrine therapy drugs.
[0023] Optionally, the cancer drug sensitivity genomics database includes the GDSC and CTRP datasets; Optionally, the antibody drug is represented by the correlation between the target gene and genes significantly related to the target gene in the whole genome data, denoted as Gcd; specifically, the representation of the antibody drug includes: identifying the target gene of the antibody drug; identifying genes significantly related to the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the significantly related genes, denoted as gene-drug sensitivity correlation Gcd; Gcd is used to characterize the antibody drug; Preferably, the antibody drug is represented by the negative value of the correlation between the target gene and genes that are significantly related to the target gene in the whole genome data, and a negative sign is added before the expression level of the target gene of the antibody drug; optionally, the antibody drug includes: HER2 drug; Optionally, the correlation is calculated using common statistical methods, such as the Pearson correlation coefficient. The correlation value ranges from -1 to 1. When the treatment regimen includes at least two therapeutic drugs, the sum of the drug sensitivity-related Gcd values of the same gene in each drug representation is used to represent the treatment regimen. 102. Input the optional treatment options and transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after each optional treatment option. In some embodiments, the method for constructing the GDnet model includes: acquiring whole-genome data and real treatment regimens of a training set sample of breast cancer patients who have received neoadjuvant therapy; representing the therapeutic drugs using gene sensitivity correlation (Gcd) of pan-cancer cell lines in a cancer drug sensitivity genomics database; inputting the whole-genome data and real treatment regimens into a bioinformatics neural network model for training to obtain the GDnet model; optionally, the bioinformatics neural network model includes at least an input layer, a gene layer, at least one biological process layer, and an output layer; optionally, the bioinformatics neural network model includes the GDnet framework.
[0024] In some embodiments, the significantly related gene expression is denoted as Ge, and the input method for inputting whole-genome data and real treatment plans into the bioinformatics neural network model for training includes: inputting Ge and Gcd into the neural network model in parallel, or aggregating Ge and Gcd to obtain the attention value Gat of gene expression at the time of treatment, and inputting the attention value into the neural network model; optionally, the attention value Gat is calculated as follows: Gat = Gex * (Exp - Gcd).
[0025] 103. The optional treatment options with a predicted pCR probability greater than the first threshold (N%) are assigned to the best group, and the optional treatment options with a predicted pCR probability less than the first threshold are assigned to the worst group. 104. Based on whether the real treatment plan belongs to the best group or the worst group, the training set samples are classified. If the real treatment plan is in the best group, the sample is assigned to the optimization group; if the real treatment plan is in the worst group, the sample is assigned to the control group, thus obtaining digital organoids including the optimization group and the control group.
[0026] like Figure 2 As shown, a second aspect of this application discloses a method for predicting the sensitivity of a treatment regimen, the method comprising: 201. Obtain the transcriptome data and candidate treatment plans of the subjects; 202. Input the transcriptome data and candidate treatment plans into the digital organoids described in the first aspect of this application, and output the classification result of whether the subjects belong to the optimized group or the control group; if the classification result of the subjects belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subjects belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained. Optionally, the method further includes: assisting in the selection of a treatment plan based on the classification result; if the classification result indicates that the subject belongs to the optimized group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is suitable; if the classification result indicates that the subject belongs to the control group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is unsuitable.
[0027] like Figure 3 As shown, a third aspect of this application discloses a method for predicting the sensitivity of a treatment regimen, the method comprising: 301. Obtain transcriptome data and candidate treatment regimens for the subjects; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 302. Input transcriptome data and candidate treatment plans into the digital organoids described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained. Optionally, the drug types in the candidate treatment regimens include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, and ENG. Optionally, the method further includes: assisting in the selection of a treatment plan based on the classification result; if the classification result indicates that the subject belongs to the optimized group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is suitable; if the classification result indicates that the subject belongs to the control group, obtaining an auxiliary prediction result indicating that the candidate treatment plan is unsuitable. Optionally, if the candidate treatment regimens include any one or more of paclitaxel|doxorubicin|cyclophosphamide|olaparib|PDCD1, 5-fluorouracil|epirarubicin|cyclophosphamide|lapatinib|ERBB2, 5-fluorouracil|epirarubicin|cyclophosphamide|ERBB2 and paclitaxel|ERBB2ERBB2, paclitaxel|doxorubicin|cyclophosphamide|ERBB2MK-2206, paclitaxel|doxorubicin|cyclophosphamide|PDCD1 and paclitaxel|doxorubicin|cyclophosphamide|ENG|ERBB2, an auxiliary predictive result indicating a higher probability that the subject belongs to the optimized group is obtained; if the candidate treatment regimens include any one or more of docetaxel|5-fluorouracil|5-fluorouracil|epirarubicin|cyclophosphamide, docetaxel and doxorubicin, an auxiliary predictive result indicating a higher probability that the subject belongs to the control group is given.
[0028] like Figure 4As shown, the fourth aspect of this application discloses a method for predicting the sensitivity of a therapeutic drug, the method comprising: 401. Obtain the transcriptome data and candidate therapeutic drugs of the test subjects; 402. Input the transcriptome data and candidate therapeutic drugs into the digital organoids described in the first aspect of this application, and output the classification result of whether the test subjects belong to the optimized group or the control group; if the classification result of the test subjects belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drugs is obtained; if the classification result of the test subjects belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drugs is obtained; optionally, the candidate drugs also include new drugs not included in the training set.
[0029] like Figure 5 As shown, the fifth aspect of this application discloses a method for predicting the sensitivity of a therapeutic drug, the method comprising: 501. Obtain transcriptome data of the subjects and candidate therapeutic drugs; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 502, Input transcriptome data and candidate therapeutic drugs into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained. Optionally, the candidate therapeutic drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, and ENG; optionally, the candidate therapeutic drugs also include new drugs not included in the training set. Optionally, if the candidate treatment includes any one or more of PDCD1, ERBB2, olaparib, lapatinib, and MK-2206, an auxiliary prediction result with a higher probability of predicting pCR can be given.
[0030] In some embodiments, the terms “subject,” “subject,” “test subject,” or “sample to be tested” as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., which will become the recipient of a particular treatment. Generally, the terms “subject” and “patient” are used interchangeably herein when referring to human subjects. Preferably, the subject is a human.
[0031] In some embodiments, the auxiliary prediction results include, but are not limited to, paper or electronic reports. These results are obtained by the intelligent machine based on the subject's relevant data and are intended only as a reference for medical personnel, not as the subject's final diagnosis. In some embodiments, the first threshold is obtained through training on training set samples. It can be a specific threshold or a range, and its specific form is not specifically limited in this embodiment.
[0032] Figure 6 This is a schematic diagram of a computer device provided in an embodiment of the present invention, such as... Figure 6 As shown, the device may include: one or more processors and one or more memories; wherein the memories store computer-readable code that, when run by the one or more processors, can perform the methods described above.
[0033] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 or ARM architecture.
[0034] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0035] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 7 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 7As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 7 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 7 One or more components in the computing device shown.
[0036] This invention also includes a computer-readable storage medium, such as... Figure 8 The diagram illustrates a storage medium provided in an embodiment of the present invention. The computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the method described above according to embodiments of the present disclosure can be performed. The computer-readable storage medium in the embodiments of the present disclosure can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchronous Link Dynamic Random Access Memory (SLDRAM), and Direct Memory Bus Random Access Memory (DRRAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0037] This disclosure also provides a computer program product or system, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0038] In some embodiments, this embodiment also discloses a system for constructing digital organoids, the system comprising: The first acquisition module 601 is used or configured to acquire optional treatment plans, real treatment plans, and transcriptome data of the training set samples; the treatment plan includes a plan consisting of at least one therapeutic drug; pCR probability calculation module 602 is used or configured to input optional treatment options and transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after treatment with each optional treatment option. The optional treatment plan allocation module 603 is used or configured to allocate optional treatment plans with a predicted pCR probability greater than a first threshold to the best group, and allocate optional treatment plans with a predicted pCR probability less than the first threshold to the worst group. The model training module 604 is used or configured to classify the training set samples based on whether the real treatment plan belongs to the best group or the worst group. If the real treatment plan is in the best group, the sample is assigned to the optimal group; if the real treatment plan is in the worst group, the sample is assigned to the control group, thus obtaining digital organoids including the optimal group and the control group.
[0039] In some embodiments, this embodiment also discloses a system for predicting the sensitivity of treatment regimens, the system comprising: The second acquisition module 701 is used or configured to acquire the transcriptome data of the subjects and candidate treatment regimens; The first result output module 702 is used or configured to input transcriptome data and candidate treatment plans into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained. Optionally, the system further includes a second result output module 703: used or configured to assist in selecting a treatment plan based on the classification result; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of the candidate treatment plan being suitable is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of the candidate treatment plan being unsuitable is obtained.
[0040] In some embodiments, this embodiment also discloses a system for predicting the sensitivity of treatment regimens, the system comprising: The third acquisition module 801 is used or configured to acquire the transcriptome data and candidate treatment regimens of the subject; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; The third result output module 802 is used or configured to input transcriptome data and candidate treatment plans into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained. Optionally, the system further includes a fourth result output module 803: assisting in the selection of treatment plans based on the classification results; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of the candidate treatment plan being suitable is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of the candidate treatment plan being unsuitable is obtained.
[0041] In some embodiments, this embodiment also discloses a system for predicting therapeutic drug sensitivity, the system comprising: The fourth acquisition module 901 is used or configured to acquire the transcriptome data of the subject and candidate therapeutic drugs; optionally, the candidate drugs may also include new drugs not included in the training set.
[0042] The fifth result output module 902 is used or configured to input transcriptome data and candidate therapeutic drugs into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained. In some embodiments, this embodiment also discloses a system for predicting therapeutic drug sensitivity, the system comprising: The fifth acquisition module 1001 is used or configured to acquire transcriptome data and candidate therapeutic drugs of the subject; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; optionally, the candidate therapeutic drugs also include new drugs not included in the training set; The sixth result output module 1002 is used or configured to input transcriptome data and candidate therapeutic drugs into the digital organoid described in the first aspect of this application, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained. Specific implementation examples: Results: GDnet Workflow Diagram: To fully integrate drug and transcriptome data, mapping drugs to a series of gene-value pairs is one of the most suitable methods. The selection of gene-value pairs is a key consideration. The GDSC and CTRP datasets, containing transcriptome states of over 800 cell lines and their drug sensitivities to over 500 molecular drugs, are excellent reference datasets. The heterogeneous responses of different cancer cell lines to drugs in GDSC or CTRP can be attributed to the heterogeneity of tumor transcriptome states; therefore, genes significantly associated with drug sensitivity are more likely to be genes associated with response or resistance, and thus more likely to be genes that are interfered with when tumors are exposed to a specific drug. Therefore, we used the genome-wide expression correlation of the half-maximal inhibitory concentration (IC50) values (where lower values indicate higher sensitivity) of each drug in different cell lines to represent the expression of each drug. Figure 10 a). It is worth noting that antibody or peptide drugs (e.g., anti-HER2 drugs) are not included in these drug sensitivity datasets, and IC50 values are not traditionally used to measure drug sensitivity through complex mechanisms. Given the strong positive correlation between the sensitivity of these targeted drugs and the expression of the target gene, we use the negative correlation between genome-wide expression and drug target expression as a representative value for antibody drugs. Figure 10 a). For treatment regimens involving multiple drugs, the gene summation is performed on the representative gene for each drug in the treatment regimen ( Figure 10 b).
[0044] To better simulate the interaction between drugs and the transcriptome, we input rank-normalized transcriptome and scheme representations in parallel into the gene-level layer of a subsequent neural network, which we call GDnet. Figure 10 b). We further explored three additional model frameworks: (1) GAnet, which utilizes a mimicking key-value attention mechanism ( Figure 14 (a) Methods that integrate drugs and transcriptomes; (2) Gnetens, which uses ensemble learning to integrate predictions for patients ( Figure 14 b) All drugs; (3) Relying solely on transcriptome information ( Figure 14 c) Gnet. Next, the gene layer is advanced to the bioinformatics neural network proposed by HA et al., which incorporates gene hierarchy (child-parent) relationships, fine pathways, and more complex pathways based on the Reactome pathway dataset (version 2022). Figure 10 b). The final predictive model was designed to accurately predict the response of breast cancer patients to neoadjuvant therapy and to assist in the selection of multiple candidate treatments.
[0045] Drug Representation Based on GDSC and CTRP Datasets: To generate gene-value pairs for drug characterization, we calculated the correlation between IC50 values of various drugs and gene expression in different cell lines, as well as their correlation with the expression levels of antibody or peptide drug targets. This calculation was performed for all combinations of drugs and genes. For drugs appearing in both datasets, GDSC correlation was prioritized. Most drugs administered to patients have corresponding representations in either the GDSC or CTRP datasets. However, some small molecule drugs are missing in both datasets, which we replaced with the most relevant drugs with similar mechanisms and the same antitumor drug type in the GDSC / CTRP datasets, respectively. Representations of all drugs used in existing neoadjuvant breast cancer therapy datasets are summarized as follows ( Figure 11 Methotrexate had the most gene-value pairs, totaling 5378. Among the 22 drug characteristics, 17 had gene-value pairs greater than 1000, and 5 (ANG, lapatinib, carboplatin, epothilone B, and doxorubicin) had gene-value pairs less than 1000. Consistent clustering of the drugs grouped these 5 drugs together with fewer pairs. Figure 11 For drugs with a large number of pairs, three main clusters were generated. The first cluster included anti-angiogenic therapies, anti-HER2 therapies, and anti-IGF1R therapies; the second cluster included chemotherapy therapies, AKT inhibitors, and PARP inhibitors; and the last cluster included anti-PD-1 therapies and HSP90 inhibitors. Figure 11 These drug characterizations were then combined with the patient transcriptome to predict neoadjuvant therapy response.
[0046] GDNet demonstrated superiority in predicting neoadjuvant therapy outcomes: We collected 9550 breast cancer samples from 68 independent public datasets. After removing duplicates and filtering out samples with non-preprocessed states or missing key information (outcomes and treatment), 31 datasets containing 4371 samples and 16133 genes were filtered out. The model was trained using 17 datasets with the most cases (a total of 3459 samples). Hyperparameter tuning and internal validation were performed using 10x cross-validation, with a final learning rate of 0.005. Due to the sparsity of the model, L2 regulation was not applied (set to 0). The remaining 14 datasets with the fewest cases, totaling 912 samples, were used as external validation sets. After cross-validation with genes involved in the biological network from the Reactome dataset, a total of 8377 genes were ultimately retained.
[0047] For benchmark comparisons, we also trained dense deep neural network (DNN) variants of GDnet, GAnet, Gnetens, and Gnet (without biological knowledge), named GDNNet, GANNet, GNNetens, and GNNet, respectively. In Gnetens and GNNetens, only drugs appearing in more than 10% of the samples were used to build individual models. Five drug types were selected (anthracyclines, microtubule inhibitors, cyclophosphamide, pyrimidine analogs, and anti-HER2 antibodies) Figure 15 Then, tenfold cross-validation was applied using the training set to compare model performance. The results showed that GDnet's area under the curve (AUC) and area under the curve for precision and recall (AUPRC) (median AUC = 0.808, median AUPRC = 0.739) were higher than all other models (for example, Gnet's median AUC = 0.799, median AUPRC = 0.731) in interval validation, although there was no significant difference. Figure 16 a and Figure 16 (b) These models were trained 10 times using all the training sets and tested individually using an external validation set, facilitating robust comparisons between different models. In external validation, GDNet significantly outperformed all other models (median AUC = 0.712, median AUPRC = 0.691) (for example, GNet's median AUC = 0.676; median AUPRC = 0.632). Figure 12 ab). It's noteworthy that GDNNet, with the most parameters (AUC=0.678, AUPRC=0.661), performed worse than GANNet (AUC=0.695, AUPRC=0.668), demonstrating the importance of prior knowledge in guiding training and improving performance. To robustly compare the ROC curves between models, we created an ensemble by averaging the predictions of models trained for 10 epochs for the final prediction. Again, GDNet (AUC=0.725) achieved the highest AUC ( Figure 12 c), whose ROC curve is significantly different from Gnet (AUC=0.683) (Delong's test, p=0.037). We then tested GDnet's performance on 13 separate validation datasets with 10 replicates of the ensemble model to determine if GDnet could outperform Gnet in limited protocol settings. The GSE191127 dataset was not included in the individual tests because it only includes patients with residual disease (RD). The results showed that GDnet achieved a higher AUC than Gnet on 8 of the 13 datasets (c). Figure 12d). Although GDNet exhibits a lower AUC on the left 5 datasets, all differences are less than 0.03. Furthermore, GDNet achieves a higher AUC than GAnet and Gnetens on 7 out of 13 datasets. Figure 12 d). GDNet also shows a higher AUC than all dense models across various datasets. Figure 17 a). Similar results were observed for AUPRC values ( Figure 17 b). These results indicate that GDnet successfully captured the interaction patterns between drugs and the transcriptome, rather than simply overfitting the drug distribution biases between datasets.
[0048] GDnet can optimize regimen selection for better outcomes: To assess GDnet's ability to guide clinical regimen selection, we subsequently conducted a computer-simulated clinical trial to examine whether GDnet can improve neoadjuvant therapy outcomes in breast cancer. Figure 9 a). Given multiple possible treatment options and patient transcriptomes, we fed them into the GDnet model to calculate the predicted probability of pathological complete response (pCR) for each case. We selected the top N% of treatment options with the highest predicted pCR probability, which were likely to benefit the patient, while the bottom N% of treatment options with the lowest predicted pCR probability were not suitable for the patient. If the treatment actually used by the patient was rated as one of the best options, we assigned the patient to the optimal group; if it was rated as one of the worst options, we assigned the patient to the control group. To examine the baseline characteristics of the two groups, we also calculated prediction scores using a transcriptome-only model to assess tumor malignancy between them, prioritizing Gnet because it is structurally similar to GDnet and is more robust than GNNet (AUC Levene test, p=0.026, AUPRC Levene test, p=0.017). Figure 12 ab).
[0049] We first conducted a series of computer-based experiments based on the ISPY-2 trial (one of the datasets used to train our model). A total of 653 patients receiving one of the 11 ISPY-2 regimens participated in our simulation study, where these regimens were set as optional. We set different optimization thresholds, starting at 50% and decreasing in 5% increments (e.g., 50%, 45%, 40%, 35%), until any group size decreased to below 30 ( Figure 18 a, Figure 18 e), Figure 19 a, Figure 19 e) categorizes the best and worst solutions to ensure robustness of the results. A lower threshold indicates higher optimization rigor. Results show that among all optimization thresholds, the pCR rate of the optimization group ( Figure 18 c) were all significantly higher than the control group. Furthermore, with increased optimization (lower optimization threshold), the odds ratio (OR) of pCR increased linearly (R²=0.72, p=0.008). Figure 18 d). However, higher Gnet prediction scores were observed in the optimization group ( Figure 18 b) indicates that the tumors in this group have a lower degree of malignancy. This difference may affect the performance of our model. Therefore, we performed propensity score matching (PSM) to predict scores based on Gnet ( Figure 18 f) to balance transcriptome profiles. Similarly, the pCR rate in the optimized group was significantly higher than that in the control group ( Figure 9 (b) As optimization intensifies, the OR increases linearly (R² = 0.87, p < 0.001). Figure 9 c). Treatment weighted inverse probability (IPTW) analysis also confirms this (R²=0.46, p=0.043). Figure 18 g), indicating that GDnet has the ability to optimize treatment options.
[0050] We further conducted a series of simulations using external validation datasets. Given the limited variety of treatment options in each validation dataset, we combined all available validation aversions, resulting in 912 patients and 18 treatment options. Similar to what was observed in the ISPY-2 simulations above, the results from the external validation set showed that, across all optimization thresholds, the pCR rate of the optimized group ( Figure 19 c) were all significantly higher than the control group, with the OR increasing linearly with increasing optimization (R²=0.68, p=0.003). Figure 19 d), and even lower Gnet prediction scores were observed in the optimization group ( Figure 19 b). PSM analysis based on Gnet predicted scores ( Figure 19 f) also showed that the pCR rate in the optimized group was significantly higher than that in the control group ( Figure 9 d) OR increases linearly (through PSM, R²=0.94, p<0.001; through IPTW, R²=0.46, p=0.03) as optimization intensifies ( Figure 9 e, Figure 19 (g), further demonstrating GDnet's excellent ability to guide the selection of neoadjuvant therapy for breast cancer.
[0051] The interpretation of GDNet reveals key genes and drugs involved in the response: To understand the importance of features contributing to the model and their interactions, we visualized the layers of GDNet (figures not shown). A comprehensive gradient attribution method was applied to obtain the importance score of each node in each layer for each patient in the external validation set. The average importance score for all patients was used as the final score for each node, shown as the node length in the Sankey diagram. From 10-replication GDNet models, we selected GDNet models with similar AUC values (0.727) to the ensemble GDNet of 10-replication for interpretation. Regarding transcriptomics and treatment regimens, transcriptomics information contained more information than drug information; however, drug treatment regimens provided additional information for more accurate prediction of response. The lack of diversity (homogeneity) in regimens used in neoadjuvant therapy for breast cancer in the dataset may have resulted in a relatively low contribution of regimen information to GDNet. High-scoring genes included SRC, CCND1, MCF2L, RPS6KA1, and PSMB7. Pathways at various levels, such as the ER-Phagosome pathway at level 1, reportedly involved in the innate immune system and phagocytosis; the RHOGTPasc cycle at level 2, reportedly involved in tumorigenesis, invasion, and metastasis; and the immune system at level 5, have been identified, highlighting their crucial role in GDnet performance. Further experimental investigations are needed to elucidate the more detailed mechanisms of neoadjuvant therapy for breast cancer.
[0052] We then investigated which types of treatment contributed to the response. (Targeting...) Figure 9 As shown in Figure a, regarding the treatment optimization strategy, we further examined the distribution of regimens and drugs used in each group under different optimization thresholds in the PSM analysis. The differences in the proportion of regimens and drugs in different groups can indicate which regimens and drugs contribute more to the model. The results showed that the regimens paclitaxel|doxorubicin|cyclophosphamide|olaparib|PDCD1 and paclitaxel|doxorubicin|cyclophosphamide|ERBB2|ERBB2 appeared much more frequently in the optimized group than in the control group, and increased with the degree of optimization. Figure 13 Docetaxel | cyclophosphamide, doxorubicin and 5-fluorouracil | epirubicin | cyclophosphamide | lapatinib are the most common, or even only, regimens in the control group, with a trend towards increasing use as optimization intensifies. Figure 13 Specifically, drugs such as PDCD1 (PD-1 inhibitor), olaparib, and ERBB2 (anti-HER2 antibody) may be effective in (…). Figure 10GDnet plays a crucial role in improving pCR rates. We further examined the regimens and drug distributions in the ISPY-2 simulation trial. Similar to the results of external validation, paclitaxel|doxorubicin|cyclophosphamide|ERBB2|MK-2206, paclitaxel|doxorubicin|cyclophosphamide|PDCD1, and paclitaxel|doxorubicin|cyclophosphamide|ANG|ERBB2 were key regimens in the optimized group, with PDCD1, ERBB2, and MK-2206 playing significant roles in improving pCR rates. By integrating drug data into the model, GDnet can compare all available treatment options before treatment begins, thus promoting precision medicine specifically for breast cancer patients—something previous machine learning models could not achieve correctly. In other words, we built a model that considers organoid function, allowing us to test the sensitivity of drugs and treatment regimens and inform patients of appropriate medications. Our computer simulation results based on the ISPY-2 trial and external validation set demonstrate that GDnet can be used to optimize the treatment decision-making process for breast cancer patients and significantly improve pCR rates. For each optimization threshold, GDnet maintained its ability to improve the pCR rate, thus confirming the findings. Furthermore, a linear correlation existed between the pCR rate and optimization intensity, strongly supporting our conclusion that GDnet can optimize neoadjuvant therapy for breast cancer and has the potential to change relevant treatment guidelines. Our model can play a significant role in many potential scenarios. In routine clinical practice, our model can help select the most suitable regimen from candidate options. In clinical trials, our model can help select patients who may respond to new testing regimens or even previously abandoned regimens. Even when treatment regimens include novel drugs not included in the GDnet training set, our model is still able to predict drug response because it performs well on a validation set containing many drugs, such as durvalumab, bevacizumab, cisplatin, olaparib, and gemcitabine, which are not present in the training set.
[0053] Methods: Public dataset processing: Transcriptome data related to neoadjuvant therapy for breast cancer were retrieved from PubMed and the Gene Expression Organization (GEO) database system on February 15, 2024. The search terms used were (“neoadjuvant therapy” [MeSH term] OR neoadjuvant) AND (“breast tumor” [MeSH term] OR ((breast or mammary gland) AND (cancer* OR tumor* OR tumor* OR cancer* OR tumor* or oncology or malignancy*))). High-throughput sequencing such as next-generation sequencing (NGS) and microarrays were included, and only datasets containing at least 20 samples were selected. A total of 9550 breast cancer samples from 68 independent public datasets were accessed. Gene annotation of the microarray probe set was performed using reference tables provided in GEO. For gene and sample quality control, 16133 genes appearing in more than 50% of the samples were retained, and 4580 samples with less than 50% missing genes were retained. These datasets come from 31 datasets (GSE194040, GSE25066, GSE16716, GSE41998, GSE180962, GSE20271, GSE34138, GSE50948, GSE149322, GSE22226, GSE32603, GSE22358, GSE32646, GSE16446, GSE130788, GSE231629, GSE1 Data including GSE23845, GSE4779, GSE173839, GSE22093, GSE21997, GSE42822, GSE66399, GSE181574, GSE23988, GSE41656, GSE8465, GSE122630, GSE21974, GSE207248, and GSE191127, covering complete pCR, treatment protocols, and sufficient transcriptomic information, were used for model construction and validation.
[0054] Transcriptome ranking values are encoded: GEOs are collected from RNA-seq counts or patient-level normalized data (e.g., kilobase million transcripts (TPM) or kilobase million fragments (FPKM)), microarray patient-level normalized data (e.g., robust multi-chip average (RMA) or other methods). Access to multi-level data allows for the maintenance of a larger upstream level. To facilitate the translation between gene expression assay technology platforms, we applied a ranking transformation (ranking value encoding) to normalize the data. This involves converting gene expression values from microarray intensity or RNA-seq counts or their normalized formats to their respective ranks. This transformation is performed gene-wise. Therefore, all gene expression values for each gene are ranked from lowest to highest. The ranking is then scaled to quantiles (ranging from 0 to 1). Missing values are filled with 0. This ranking-based approach may be more robust to technical artifacts that can systematically bias absolute transcript count values, while maintaining a more stable overall relative ranking of genes within each cell.
[0055] Drug and Regimen Representation: To facilitate the fusion of drug therapy and transcriptomics, our goal is to represent drugs by combining relevant genes and correlation values. To construct a sufficiently large pharmacogenomics dataset for drug characterization, raw drug sensitivity data of cell lines to molecular drugs were retrieved from the Drug Sensitivity Genomics database v2, along with RMA-normalized / TPM-normalized and log-transformed transcriptome (GDSC) data of cell lines and the Cancer Therapy Response Portal v2 (CTRP). We used the IC50 metric to represent drug sensitivity, where a high IC50 value corresponds to low cell-killing ability, and a low IC50 value corresponds to high cell-killing ability. For each drug, we calculated the correlation between gene expression levels across all cell lines and log2(IC50+1) of the cell line's response to each drug. Considering the need to test nearly 5*10... 4 For each gene, the significance threshold for correlation was set to 1*10⁻⁶ based on Bonferroni correction. -6 To reduce the false discovery rate, genes identified as significant are more likely to play a crucial role in predicting drug responses to cancer treatments. Correlation values, being prior knowledge, indicate the importance of a gene to a drug response, thus reflecting the perturbation of the gene by the drug. Genes not present in the retained patient transcriptome set of 16,133 genes are removed. Therefore, each drug is ultimately represented by its significant genes and correlation values (ranging from -1 to 1). If a drug appears in both the GDSC and CTRP datasets, the GDSC dataset is considered first.
[0056] There are several special cases in drug representation. Certain types of drugs are not found in the GSDC or CTRP datasets. Therefore, we select the most relevant drug within the same anti-tumor drug class with a similar mechanism of action to represent it. Examples include capecitabine, represented by 5-fluorouracil; ganestepib, represented by luminespib; and some drug descriptions in the GEO dataset, such as taxanes or anthracyclines, lack sufficient drug detail; therefore, we select the average representative of taxanes or anthracyclines. Certain types of drugs are antibody- or peptide-based drugs, which are not included in either dataset and are generally not expressed using IC50 to measure sensitivity to drugs with complex mechanisms. For these drugs, we use the negative correlation between their target genes and genes in other cell lines, as responses to these drugs are closely related to target expression. Examples include PDCD1 therapies (pembrolizumab and durvalumab), VEGFA (bevacizumab), and IGF1R (ganitumab), an anti-IGF1R antibody. If a drug has multiple targets, we sum the genetic correlations between all targets as the final representation, including ANG, which comprises the ANGPT1 and ANGPT2 genes, representing the ANG peptide inhibitor (AMG-386). Furthermore, drugs such as T-DM1 (an antibody-drug conjugate) and endocrine therapy lacked corresponding drug sensitivity data in GDSC or CTRP. Given the small number of patients receiving these drugs in our dataset, patients receiving regimens such as T-DM1 or endocrine therapy were excluded. For patients receiving multiple drugs in a single regimen, we sum the genetic correlations for each drug to represent that regimen.
[0057] Fusion of drug characterization and transcriptomics: Three approaches integrating therapeutic regimens and transcriptomic information were tested. The first approach involved using gene expression (denoted as Ge) and gene sensitivity correlation (denoted as Gcd) as inputs, which we termed GDnet (GDnet). Figure 10 b). The second method uses As the final input, a method similar to key-value attention is used ( Figure 14 a), and named the model GAnet. In this mechanism, for each query (Ge), we find the corresponding key (gene) in the drug representation, and the attention (Gatt) is calculated as follows. , where this value represents the Gcd represented by the drug. Attention calculation is set to e. valueThis ensures that Gatt can mimic the tumor's response to drug exposure conditions; in other words, when the drug is absent, Gatt is expected to reduce or increase genes that are responsive (Gcd<0) or resistant (Gcd>0) to the drug. For example, ERBB2 is a gene with poor prognosis, but its expression is expected to decrease when exposed to an anti-HER2 drug that responds to ERBB2 (Gcd<0). Furthermore, we explored a third approach involving the use of ensemble learning. This approach focuses on constructing separate transcriptome-only models for each drug and combining the final predictions by averaging the predictions of all transcriptome-only models corresponding to the drug applied to the patient. This model is called Gnetens (Gnetens). Figure 14 b). A pure transcriptome model called Gnet was also explored ( Figure 14 c).
[0058] Design of the bioinformatics model: We incorporate prior knowledge into the training process, drawing inspiration from the P-NET framework developed by Elmarakeby HA et al., to accelerate model training and improve performance and interpretability. The complete Reactome dataset was downloaded and processed into a hierarchical network. This network consists of one layer of features, one layer of genes, and five layers of pathways. The membership relationships between adjacent layer nodes (representing genes or pathways) are represented by six binary matrices. The constructed hierarchical network originates from genes available in both our dataset and the Reactome dataset (the first layer). Pathways directly related to these genes form the second layer. Subsequently, paths directly related to the pathways in the second layer constitute the third layer, and so on, until a sixth layer with 18 pathways is constructed.
[0059] The deep learning model is implemented in PyTorch. The number of layers and nodes in the prior knowledge network defines the structure of the basic feedforward dense neural network model. Six binary matrices containing the dependencies between genes or pathways in adjacent layers are input into the neural network model as masks between adjacent layers. During weight updates, after each epoch, this mask is multiplied by the weights, thus imposing constraints on nodes and edges. This configuration produces a sparse neural network with six layers, including 8377 nodes (gene layers), 846 (layer 1 paths), 218 (layer 2 paths), 108 (layer 3 paths), 51 (layer 4 paths), and 18 nodes (layer 5 paths) in each layer. Specifically, for GDnet and GDNNet, an additional gene and drug layer with 19300 nodes (half gene expression and half drug representation) is added before the 8377 gene layers. Due to the sufficiently sparse model, L2 modulation is set to 0. Each node encodes a biological entity (e.g., a gene or pathway), while each edge represents a known relationship between the corresponding entities. Node constraints provide a better understanding of the states of different biological components. Edge constraints result in fewer parameters compared to fully connected networks with the same number of nodes, potentially leading to less computation. Due to the imbalanced dataset, we apply different weights to the classes to reduce the network's bias towards a particular class based on the bias in the training set. The model was trained for 1000 epochs using the Adam optimizer with default parameters to reduce the imbalanced binary cross-entropy loss function. To make each layer useful on its own, we added a prediction layer with sigmoid activation after each hidden layer. `gene`, `Layer1`, `Layer2`, `Layer3`, `Layer4`, and `Layer5` are all hidden layers, ultimately outputting the new auxiliary pCR probability. Since fitting the data to the later layers with fewer weights is more challenging, we applied higher loss weights to the later layer results (weights of 1, 1.2, 1.4, 1.6, 1.8, 2, and 2.5 from the `gene` and `drug` layers to later layers) during optimization. The network's final prediction is calculated by averaging the results across all layers.
[0060] Model Training and Validation: This analysis included datasets conforming to specific criteria, providing a total of 4371 available patients. The 17 largest datasets, comprising approximately 3459 patients (79% of the total), were designated as the training set. The remaining 14 datasets, comprising approximately 912 patients (21% of the total), were used as external validation sets, with no subjective assignment during allocation. The model was trained using the training set, and hyperparameters such as the learning rate were optimized through a tenfold validation pass. The trained model was then validated using an external test set. The predictive performance of different models was compared using the validation set by calculating the median AUC and AUPRC over 10 epochs of replicate training using the roc_auc_score and average_precision_score implemented in the Python sklearn.metrics library, respectively. We then averaged the results across 10 replicates of the model to derive the final predictions for stability and robustness. Implementations of the proposed system and reproducible results are available upon request and will be publicly available on GitHub upon posting acceptance.
[0061] Model Explanation: We applied IntegratedGradients, a gradient-based attribution method implemented in the tcaptum.attr library, to compute the sample-level importance of all nodes across all layers. The absolute importance score (always positive) of each layer's nodes was normalized using the following formula to promote comparability between layers: Where Ij is the normalized absolute importance score of each node in the layer, Aj is the absolute importance score, and n is the total number of nodes in the layer.
[0062] To calculate the overall node importance, we aggregated the sample-level importance scores of all samples in the validation set (dataset-level interpretation: the average of all patient-level weights in the validation set). The importance score of each node is represented as the node length in the Sankey plot. The importance score of the link between the source node (previous layer) and the target node (next layer) is represented by the normalized absolute weight (within the absolute weights of all source nodes pointing to the same target node), as shown in the following equation: Ls-t j is the importance score of the link from the j-th source node (s) to the target node (t), is the normalized absolute importance score of the target node, Wsj is the weight of the j-th source node pointing to the target node in the deep learning model, and n is the number of source nodes pointing to the target node.
[0063] Statistical Analysis: The DeLong test was used to examine the variation of the area under the ROC curve between GDnet and other models. The Wilcoxon rank test was used to compare the AUCs of GDnet and other models based on 10-fold cross-validation and 10-fold repeated training and testing. The Levene test was used to compare the AUCs and AUPRC variances between Gnet and GNNet to assess performance robustness. Odds ratios (ORs) were calculated to determine the preference for pathological complete response (pCR) in the simulation trial, and differences were compared using Fisher's exact test. A linear regression model was fitted to examine the relationship between optimization and OR. The significance of the regression coefficients was assessed using a t-test, and the R-squared value was used to quantify the proportion of variance explained by the model. To address potential confounding effects from the imbalance in transcriptomic profiles (inferred from Gnet predicted scores) between the optimization and control groups, propensity score matching (PSM) and treatment-weighted inverse probability (IPTW) were employed. PSM was performed by estimating propensity scores using logistic regression, including relevant covariates of Gnet scores. Nearest neighbor matching (1:1 ratio) and a caliper of 0.1 were used for matching without replacement to minimize imbalance. Furthermore, IPTW is applied by weighting each individual with the inverse probability of the observed assignments, thereby creating a balanced pseudopopulation. Stable weights are used to maintain accurate estimates and ensure robust results. All tests are two-sided. Unless otherwise stated, a p-value < 0.05 is considered statistically significant. All statistical analyses were performed using R statistical software version 4.3.
[0064] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of this disclosure, and such modifications should fall within the scope of this disclosure.
Claims
1. A method for constructing digital organoids, characterized in that, The method includes:
101. Obtain the optional treatment regimens, real treatment regimens, and transcriptome data of the training set samples; the treatment regimens include regimens consisting of at least one therapeutic drug; 102. Input the optional treatment options and transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after each optional treatment option.
103. The optional treatment options with a predicted pCR probability greater than the first threshold are assigned to the best group, and the optional treatment options with a predicted pCR probability less than the first threshold are assigned to the worst group.
104. Based on whether the real treatment plan belongs to the best group or the worst group, the training set samples are classified. If the real treatment plan is in the best group, the sample is assigned to the optimization group; if the real treatment plan is in the worst group, the sample is assigned to the control group, thus obtaining digital organoids including the optimization group and the control group. The method for constructing the GDnet model includes: The study acquires whole-genome data and real treatment regimens from training set samples that have received neoadjuvant therapy; the therapeutic drugs are represented using gene sensitivity correlations in pan-cancer cell lines from a cancer drug sensitivity genomics database; wherein, the therapeutic drugs include small molecule drugs and / or antibody drugs, small molecule drugs are represented by calculating IC50 values and the correlation between genes significantly associated with small molecule drugs in whole-genome data, and antibody drugs are represented by calculating the correlation between target genes and genes significantly associated with target genes in whole-genome data, this correlation is denoted as Gcd; the expression of the significantly associated genes is denoted as Ge; The attention value Gat is calculated by inputting Ge and Gcd into a bioinformatics neural network model in parallel, or by aggregating Ge and Gcd to obtain the gene expression attention value Gat during treatment. This attention value is then input into the bioinformatics neural network model for training, resulting in the GDnet model. The methods for aggregating Ge and Gcd to calculate the attention value Gat include: Gat... x = Ge x × (Exp -Gcdx ).
2. The method for constructing digital organoids according to claim 1, characterized in that, The bioinformatics neural network model includes at least an input layer, a gene layer, at least one biological process layer, and an output layer; the input layer is a Gene+drug layer, and gene, Layer1, Layer2, Layer3, Layer4, and Layer5 are all hidden layers, and finally, the output of the new auxiliary pCR probability is given.
3. The method for constructing digital organoids according to claim 1, characterized in that, The representation of the small molecule drug includes: calculating the IC50 value of the small molecule drug; identifying genes that are significantly associated with the small molecule drug in whole-genome data based on the IC50 value; and calculating the correlation between the IC50 value and the significantly associated genes, denoted as the gene-drug sensitivity correlation.
4. The method for constructing digital organoids according to claim 3, characterized in that, If the first small molecule drug is missing from the genomics database, the most relevant drug within the same antitumor drug type and with a similar mechanism of action in the database shall be used to represent the first small molecule drug.
5. The method for constructing digital organoids according to claim 3, characterized in that, If the genomics database includes at least two databases, and the second small molecule drug exists in at least two databases simultaneously, the database with higher acceptance shall be used preferentially to calculate the representation of the second small molecule drug.
6. The method for constructing digital organoids according to claim 3, characterized in that, The small molecule drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, 5-fluorouracil, epirubicin, lapatinib, MK-2206, and ENG.
7. The method for constructing digital organoids according to claim 1, characterized in that, The cancer drug sensitivity genomics database includes the GDSC and CTRP datasets.
8. The method for constructing digital organoids according to claim 1, characterized in that, The representation of the antibody drug includes: identifying the target gene of the antibody drug; identifying genes that are significantly related to the target gene in whole genome data; and calculating the correlation between the expression level of the target gene and the significantly related gene, denoted as gene-drug sensitivity correlation.
9. The method for constructing digital organoids according to claim 8, characterized in that, The antibody drugs are represented by negative values of the correlation between the target gene and genes that are significantly related to the target gene in the whole genome data.
10. The method for constructing digital organoids according to claim 9, characterized in that, When the treatment regimen includes at least two therapeutic drugs, the sum of the drug sensitivity-related Gcd values for the same gene in each drug representation is used to represent the treatment regimen.
11. The method for constructing digital organoids according to claim 1, characterized in that, The bioinformatics neural network model includes the GDnet framework.
12. A method for predicting the sensitivity of a treatment regimen, characterized in that, The method includes:
201. Obtain transcriptome data and candidate treatments from the subjects; 202. Input transcriptome data and candidate treatment plans into the digital organoid according to any one of claims 1-11, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained.
13. The method for predicting treatment sensitivity according to claim 12, characterized in that, The method further includes: assisting in the selection of treatment plans based on the classification results; if the classification result of the subject belonging to the optimization group is obtained, obtaining an auxiliary prediction result that the candidate treatment plan is suitable; if the classification result of the subject belonging to the control group is obtained, obtaining an auxiliary prediction result that the candidate treatment plan is unsuitable.
14. A method for predicting the sensitivity of a treatment regimen, characterized in that, The method includes:
301. Obtain transcriptome data and candidate treatment regimens for the subjects; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 302, Input transcriptome data and candidate treatment plans into the digital organoid according to any one of claims 1-11, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result with high sensitivity of the candidate treatment plan is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result with low sensitivity of the candidate treatment plan is obtained.
15. The method for predicting treatment sensitivity according to claim 14, characterized in that, The candidate treatment options include any one or more of the following drugs: paclitaxel, doxorubicin, cyclophosphamide, olaparib, 5-fluorouracil, epirubicin, lapatinib, MK-2206, and ENG.
16. The method for predicting treatment sensitivity according to claim 14, characterized in that, The method further includes: assisting in the selection of treatment plans based on the classification results; if the classification result of the subject belonging to the optimization group is obtained, obtaining an auxiliary prediction result that the candidate treatment plan is suitable; if the classification result of the subject belonging to the control group is obtained, obtaining an auxiliary prediction result that the candidate treatment plan is unsuitable.
17. A method for predicting the sensitivity of a therapeutic drug, characterized in that, The method includes:
401. Obtain transcriptome data and candidate therapeutics from the subjects; 402, input transcriptome data and candidate therapeutic drugs into the digital organoid according to any one of claims 1-11, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained.
18. The method for predicting therapeutic drug sensitivity according to claim 17, characterized in that, The candidate therapeutics also include new drugs not included in the training set.
19. A method for predicting the sensitivity of a therapeutic drug, characterized in that, The method includes:
501. Obtain transcriptome data of the subjects and candidate therapeutic drugs; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 502, input transcriptome data and candidate therapeutic drugs into the digital organoid according to any one of claims 1-11, and output the classification result of whether the subject belongs to the optimized group or the control group; if the classification result of the subject belonging to the optimized group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained.
20. The method for predicting therapeutic drug sensitivity according to claim 19, characterized in that, The candidate therapeutic agents include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, 5-fluorouracil, epirubicin, lapatinib, MK-2206, and ENG.
21. The method for predicting therapeutic drug sensitivity according to claim 19, characterized in that, The candidate therapeutics also include new drugs not included in the training set.
22. The method for predicting therapeutic drug sensitivity according to claim 19, characterized in that, If the candidate treatment includes one or more of olaparib, lapatinib, and MK-2206, provide an auxiliary prediction result that indicates a higher probability of pCR.
23. A computer device, characterized in that, The device includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1-22.
24. A computer-readable storage medium, characterized in that, It stores a computer program thereon, which, when executed by a processor, implements the steps of the method as described in any one of claims 1-22.
25. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-22.