Digital organoid construction method and device, medium and program product

By constructing a GDnet digital organoid model and integrating transcriptome data and multiple treatment options, the limitations of flexibility and interpretability in existing neoadjuvant therapy models for breast cancer have been addressed. This has enabled more accurate treatment prediction and decision optimization, thereby improving the treatment outcomes for breast cancer patients.

CN120809051AActive Publication Date: 2025-10-17CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510122427.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-10-17
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

Existing methods for predicting neoadjuvant therapy in breast cancer are based on small sample sizes and single treatment strategies, making it difficult to effectively guide patients in choosing treatment options in real-world scenarios. Furthermore, existing models lack flexibility and interpretability, and cannot comprehensively characterize tumor status and drug-tumor interactions.

Method used

We constructed a deep learning-based digital organoid model, GDnet, which integrates transcriptome data and multiple treatment options. By using bioinformatics neural networks to simulate drug-tumor interactions, we established a general predictive model to accurately predict responses to neoadjuvant therapy in breast cancer and optimize clinical decision-making.

Benefits of technology

It improves the predictive accuracy of neoadjuvant therapy for breast cancer and optimizes clinical decision-making, achieving optimal predictive performance with limited sample size, and provides interpretable treatment options, significantly improving the complete clinical response rate for breast cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809051A_ABST
    Figure CN120809051A_ABST
Patent Text Reader

Abstract

The invention provides a digital organoid construction method, equipment, a medium and a program product, and further provides a method, equipment, a medium and a program product for developing and predicting therapeutic schedule and / or therapeutic drug sensitivity based on a model, and relates to the field of intelligent medical treatment. The constructed digital organoid is widely applied to daily clinical practice, and the model can help to select the most suitable scheme from candidate schemes; in clinical trials, the model can help select patients who may respond to new test regimens, even previously abandoned regimens, and the model can predict drug response even if the treatment regimens contain new drugs not contained in the training set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent medical treatment, and more particularly, to a method, device, medium and program product for constructing a digital organoid. BACKGROUND

[0002] The existing method for predicting neoadjuvant therapy of breast cancer patients is a model constructed based on genomics, transcriptomics, radiomics, pathology, or even multiomics. The model is developed based on small sample size and single treatment strategy queue, and cannot be well applied to the real scene to directly guide the selection of treatment plan for patients.

[0003] In theory, predicting the efficacy of cancer treatment needs to address two aspects. First, more comprehensive characterization of tumor markers is essential for better assessment of their functional status. This task requires integrating as many omics data types as possible, including transcriptome, genome, proteome, and pathological information, to achieve more comprehensive tumor characterization and improve the predictive performance of the model. However, as the number of features increases, the required sample size for training also increases significantly. Currently, in addition to transcriptome data, the availability of patients with biological multiomics information is limited, and the existing large-sample-based model is mainly based on transcriptome data. Therefore, comprehensive characterization of the transcriptome state is crucial. Due to the flexibility of neural networks, deep learning methods allow the use of as many genes as possible to best characterize the state of the tumor, capturing not only tumor heterogeneity but also tumor microenvironment information. However, deep learning methods require a large number of samples, and the results are more difficult to interpret. Integrating prior biological knowledge can facilitate the construction of sparse, high-performance, and interpretable deep learning models.

[0004] The second aspect relates to proper characterization of drugs to accurately simulate the interaction between drugs and tumors, which is crucial for predicting the response of tumors to treatment. In order to develop a model that can better guide clinical drug decisions in precision medicine, it is necessary to learn the complex drug-tumor interaction patterns by combining as many different types of treatments as possible. In order to incorporate therapeutic drugs into the model, three important factors need to be considered: (1) The types of anticancer drugs used in clinical practice are relatively limited (less than one hundred) and structurally diverse (e.g. small molecules, antibodies). Given the insufficient number of breast cancer treatment drug types, it is difficult to understand the compound structure characterization and integrate compound structure data with transcriptome status, and previous strategies of directly characterizing compounds according to their structures are not suitable for this task.(2) Many chemotherapy drugs lack specific targets, and the mechanism of action of targeted therapy often goes beyond the primary target (e.g. antibody-dependent cell-mediated cytotoxicity-related genes in anti-HER2 therapy). Therefore, previous drug characterization based on specific gene targets cannot fully represent all types of drugs or fully characterize their mechanisms.(3) In clinical practice, the combination of drugs leads to complex synergistic or antagonistic interactions, requiring a flexible model structure to simulate these interactions. Therefore, the existing framework based on late fusion integration cannot achieve the best efficacy prediction. In general, a unified genome-wide approach is needed to characterize different types of drugs and explain the complex interactions in combination therapy. SUMMARY

[0005] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a method, device, medium and program product for constructing a digital organoid, and a method, device, medium and program product for developing a prediction of a treatment plan and / or a sensitivity of a therapeutic drug based on the model; the method of the present application establishes a general deep learning model based on a treatment plan and transcriptome data for accurately predicting the response of breast cancer to neoadjuvant therapy in various treatment plans and optimizing clinical decisions.

[0006] The first aspect of the present application discloses a method for constructing a digital organoid, the method comprising:

[0007] 101. Obtain the optional treatment plan, the real treatment plan and the transcriptome data of the training set samples; the treatment plan comprises a plan composed of at least one therapeutic drug;

[0008] 102. Input the optional treatment plan and the transcriptome data into the GDnet model, and calculate the predicted pCR probability of different samples after treatment with each optional treatment plan;

[0009] 103. Assign the optional treatment plan with a predicted pCR probability greater than a first threshold value to the best group, and assign the optional treatment plan with a predicted pCR probability less than the first threshold value to the worst group;

[0010] 104, based on the real treatment scheme belongs to the best group or worst group of the training set sample classification, if the real treatment scheme is in the best group, the sample is assigned to the optimization group; if the real treatment scheme is in the worst group, the sample is assigned to the control group, to obtain a digital organoid including the optimization group and the control group.

[0011] In some embodiments, the method for constructing the GDnet model comprises:

[0012] Obtaining the whole genome data and the real treatment scheme of the training set sample which has accepted neoadjuvant therapy; the treatment drug is represented by the gene sensitivity related Gcd of the pan-cancer cell line in the cancer drug sensitivity genomics database;

[0013] Inputting the whole genome data and the real treatment scheme into a biological information neural network model for training to obtain the GDnet model; the GDnet model is a typical binary classification model. The neural network is transformed by sigmoid, and the output is the probability of pCR. It can be referred to as logistics regression, which belongs to a binary classification model.

[0014] The second aspect of the present application discloses a method for predicting the sensitivity of a treatment scheme, comprising:

[0015] 201, obtaining the transcriptome data of a subject and a candidate treatment scheme;

[0016] 202, inputting the transcriptome data and the candidate treatment scheme into the digital organoid of the first aspect of the present application, and outputting the classification result of the subject belonging to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, the auxiliary prediction result of the high sensitivity of the candidate treatment scheme is obtained; if the classification result of the subject belonging to the control group is obtained, the auxiliary prediction result of the low sensitivity of the candidate treatment scheme is obtained.

[0017] The third aspect of the present application discloses a method for predicting the sensitivity of a treatment scheme, comprising:

[0018] 301, obtaining the transcriptome data of a subject and a candidate treatment scheme; the genes in the transcriptome data include any one or several of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1;

[0019] 302, inputting the transcriptome data and the candidate treatment scheme into the digital organoid of the first aspect of the application, and outputting a classification result of the subject belonging to an optimization group or a control group; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of high sensitivity of the candidate treatment scheme is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate treatment scheme is obtained;

[0020] Optionally, the drug species in the candidate treatment scheme includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG;

[0021] Optionally, the method further comprises: assisting in selecting a treatment scheme based on the classification result; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of the candidate treatment scheme being suitable is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of the candidate treatment scheme being unsuitable is obtained;

[0022] Optionally, if the candidate treatment scheme includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, 5-fluorouracil, epirubicin, cyclophosphamide, ERBB2, paclitaxel, ERBB2ERBB2, paclitaxel, doxorubicin, cyclophosphamide, ERBB2MK-2206, paclitaxel, doxorubicin, cyclophosphamide, PDCD1, and paclitaxel, doxorubicin, cyclophosphamide

[0023] |ENG|ERBB2, an auxiliary prediction result of the subject belonging to the optimization group with a high probability is obtained;

[0024] Optionally, if the candidate treatment scheme includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, 5-fluorouracil, epirubicin, cyclophosphamide, ERBB2, paclitaxel, ERBB2ERBB2, paclitaxel, doxorubicin, cyclophosphamide, ERBB2MK-2206, paclitaxel, doxorubicin, cyclophosphamide, PDCD1, and paclitaxel, doxorubicin, cyclophosphamide

[0025] The fourth aspect of the application discloses a method for predicting the sensitivity of a treatment drug, which comprises:

[0026] 401, obtaining the transcriptome data and the candidate treatment drug of the subject;

[0027] 402, input the transcriptome data and the candidate therapeutic drug into the digital organoid of the first aspect of the application, output the classification result of the subject belonging to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, the auxiliary prediction result of the candidate therapeutic drug with high sensitivity is obtained; if the classification result of the subject belonging to the control group is obtained, the auxiliary prediction result of the candidate therapeutic drug with low sensitivity is obtained;

[0028] The fifth aspect of the application discloses a method for predicting the sensitivity of a therapeutic drug, the method comprising:

[0029] 501, obtaining the transcriptome data of the subject and the candidate therapeutic drug; the genes in the transcriptome data include any one or several of PSMC5, CCND1, PSMB3, PSMD14 and CUL1;

[0030] 502, input the transcriptome data and the candidate therapeutic drug into the digital organoid of the first aspect of the application, output the classification result of the subject belonging to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, the auxiliary prediction result of the candidate therapeutic drug with high sensitivity is obtained; if the classification result of the subject belonging to the control group is obtained, the auxiliary prediction result of the candidate therapeutic drug with low sensitivity is obtained.

[0031] The application has the following beneficial effects:

[0032] 1, the application innovatively discloses a method for constructing a digital organoid, which utilizes existing optional treatment schemes and real treatment schemes, associates with transcriptome data, calculates the predicted pCR probability and groups the optional treatment schemes based on the predicted pCR probability, and classifies the training set samples based on the grouping result to obtain the digital organoid. In daily clinical practice, the model can help to select the most suitable scheme from the candidate schemes; in clinical trials, the model can help to select patients who may respond to new test schemes or even previously abandoned schemes, even if the treatment scheme contains new drugs not included in the training set, the model can still predict drug response.

[0033] 2, The application innovatively discloses the construction process of the GDnet model, maps the drug to a series of gene and value pairs by fully integrating the drug and transcriptome data, and uses different methods for small molecule drugs and antibody drugs respectively for characterization; on this basis, the whole genome data and the real treatment scheme of the breast cancer training set samples that have received neoadjuvant therapy are input into the biological information neural network model for training, a general deep learning model is established, which can simulate the interaction between drugs and tumors as a digital organoid, to accurately predict the neoadjuvant therapy response of breast cancer in various treatment schemes and optimize the clinical decision. In this scheme, as many transcriptome samples and diversified treatment data as possible are collected, as many genes as possible are used to represent the tumor state; various methods of integrating drug and transcriptome data to simulate the interaction between drug and tumor are explored; the biological information neural network (a leading deep learning technology) is applied to achieve the best prediction performance with limited sample size, and the ability of the model to support the optimization of clinical decision is explored. This study will significantly promote the precise neoadjuvant therapy selection and improve the treatment effect of breast cancer.

[0034] The application of GDnet in two series of simulated clinical trials shows that GDnet can be used as a digital organoid to optimize the treatment selection of breast cancer patients. Compared with previous models, GDnet has the following advantages:

[0035] (1) By integrating drug data into the model, GDnet can compare all optional treatment options before starting treatment, thus promoting precision medicine specifically for breast cancer patients, which has not been properly implemented by all previous machine learning models. In other words, we have established a model that takes into account organoid function, which can test drug and treatment sensitivity and inform patients of appropriate drugs. The results of our computer simulation tests based on the ISPY-2 trial and external validation set in our study show that GDnet can be used to optimize the treatment decision-making process for breast cancer patients and significantly improve the pCR rate for breast cancer patients. For each optimization threshold, GDnet maintains the ability to improve the pCR rate, thus confirming the research results. In addition, there is a linear correlation between the pCR rate and the optimization intensity, which strongly supports our conclusion that GDnet can optimize neoadjuvant therapy for breast cancer and has the potential to change the relevant treatment guidelines. Our model can play an important role in many potential scenarios. In daily clinical practice, our model can help select the most suitable candidate from the candidate scheme. In clinical trials, our model can help select patients who may respond to new test schemes or even previously abandoned schemes. Even if the treatment scheme contains new drugs not included in the GDnet training set, our model can still predict drug response because it performs well on the validation set containing many drugs, such as durvalumab, bevacizumab, cisplatin, olaparib, and gemcitabine, which are not present in the training set.

[0036] (2) Most previous machine learning models face the problem of limited sample size, lack of sufficient independent external validation leading to unstable performance, etc. This application collects 31 data sets containing 4371 breast cancer samples containing available pre-treatment transcriptome and neoadjuvant therapy information, which is the largest number of data sets so far to robustly train and validate GDnet;

[0037] (3) Data were collected from all types of RNA sequencing methods by applying a rank-normalization strategy to each sample. This approach provides robustness against technical artifacts that can otherwise introduce systematic bias in absolute transcript counts while maintaining the overall relative ranking of each gene within a cell at a more stable level. Previous studies were limited to a limited number of RNA sequencing methods due to ease of standardization and reduction of batch effects; however, these approaches often reduce the amount of available samples, making it difficult to generalize to a wider range of scenarios. Although rank-based encoding has limitations, including underutilization of the precise gene expression measurements provided in transcript counts, it makes the model more applicable to real conditions with various unexpected variables. In addition, unlike previous studies, which normalized between patients in different cohorts or test methods (e.g., z-score transformation) or even between datasets (e.g., Combat algorithm), this method only includes standardization for each patient. This design choice stems from the understanding that normalization parameters trained across patients or datasets cannot be effectively applied to patients in another dataset with different transcriptome test protocols, resulting in overfitting and poor generalizability. Therefore, our model can be directly generalized to new cases tested using any RNA sequencing technology in a wider range of scenarios and complex real conditions;

[0038] (4) Previous models, such as the 21-gene recurrence score, MammaPrint, Adjutorium, are mainly based on traditional statistical models or machine learning models. In contrast, GDnet is based on a deep learning framework and contains a larger number of genes to maximize model capacity. Considering that achieving pCR means eliminating all tumor cells, it is important to maintain sufficient flexibility in the model to consider tumor heterogeneity and microenvironment changes, a challenge that traditional machine models cannot address. The deep learning framework, with its large number of parameters and flexible architecture, is expected to address these challenges. A previous study showed that deep learning methods can reflect or predict tumor heterogeneity and tumor environment. In addition, by integrating prior biological knowledge, GDnet significantly reduces the number of learning parameters, resulting in more stable and better performance than dense models, as demonstrated in previous studies. The visualization of the importance and properties of multi-level genes and biological pathways in GDnet enables a multi-level view of model interpretation, which can guide researchers to make hypotheses about the potential biological processes involved in drug resistance and translate these findings into therapeutic opportunities. Specifically, GDnet-identified genes such as SRC, CCND1, MCF2L, RPS6KA1, and PSMB7 may play an important role in breast cancer resistance.

[0039] (5) GDnet innovatively integrates breast cancer treatment regimens, capturing the interaction patterns between drugs and tumors, thereby improving model performance. Although a few breast cancer neoadjuvant therapy models consider specific drug types, such as chemotherapy and anti-HER2 treatment, these models often lack the flexibility needed to simulate the interaction between tumors and drugs. Previous studies have explored drug representations based on drug structures, however, these methods have limitations because they lack sufficient drug diversity for comprehensive training on the breast cancer neoadjuvant task and cannot represent antibody drugs, or use targets to drugs, but target representations lack flexibility to simulate drug-tumor interactions. To address these shortcomings, a simple method is introduced that utilizes gene expression correlations to represent drugs, which helps to introduce external knowledge related to drug resistance. This method provides a drug representation structure similar to gene expression in tumors, which helps to integrate and create the potential to simulate drug and transcriptome interactions. We also demonstrate that a better approach to integrating drug representations and transcriptomes involves parallel inputting data from both modalities into a neural network, rather than simply fusing them through specially designed rules or aggregating predictions from specific drug models;

[0040] In summary, GDnet is a bioinformatics deep learning method that integrates drug representations and transcriptomes to achieve more accurate oncology and enhance the treatment decision-making process for breast cancer neoadjuvant therapy. Integrating drug representations and omics data represents a new approach to simulating drug-tumor interactions, providing a digital organoid platform for precise oncology that can be widely applied to different therapeutic environments and various cancer types. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0042] Figure 1 is a method flowchart provided by the first aspect of the embodiment of the present application;

[0043] Figure 2 is a method flowchart provided by the second aspect of the embodiment of the present application;

[0044] Figure 3 is a method flowchart provided by the third aspect of the embodiment of the present application;

[0045] Figure 4 is a method flowchart provided by the fourth aspect of the embodiment of the present application;

[0046] Figure 5 is a schematic diagram of the method provided by the fifth aspect of the embodiments of the present application;

[0047] Figure 6 is a schematic diagram of a computer device provided by the embodiments of the present application;

[0048] Figure 7 is a schematic diagram of the architecture of an exemplary computer device provided by the embodiments of the present application;

[0049] Figure 8 is a schematic diagram of a storage medium provided by the embodiments of the present application;

[0050] Figure 9 is a schematic diagram of the construction process of a digital organoid and the effect of computer clinical trial simulation on optimizing breast cancer neoadjuvant therapy decision-making provided by the embodiments of the present application; Figure 9 a, computer clinical trial simulation. Patient transcriptome and multiple options are input into the model to predict the pCR probability of each option. The top N% of regions with the highest pCR probability and the bottom N% of regions with the lowest pCR probability are selected as the best option and the worst option. If the patient's treatment option (represented by a white star "*") is ranked as the best treatment option, the patient is assigned to the optimization group; if ranked as the worst, they are assigned to the control group. N represents the optimization intensity of the simulation trial. pCR rates of the optimization group and the control group of the simulation trial based on the I-SPY2 trial after propensity score matching (PSM) at different optimization intensity levels Figure 9 b) and the simulation trial based on the external validation dataset after PSM Figure 9 d). The difference in pCR rate is tested by Fisher's exact test, and the p value is marked on the green solid line of the optimization group. The odds ratio of pCR in the optimization group Figure 9 c) and the simulation trial based on the external validation dataset after PSM Figure 9 e). The dashed line is a fitted linear regression function between the y-axis and the optimization threshold, and R2 and p value are displayed at the top. Opt: optimization group; Ctl: control group;

[0051] Figure 10 is a schematic diagram of the fusion of drug representation and transcriptome and option representation provided by the embodiments of the present application. Wherein Figure 10 a represents the strategy of representing drugs through gene correlation. Small molecule drugs represent significant whole genome correlation with half maximal inhibitory concentration (IC50) values, while antibody drugs represent significant whole genome correlation with target genes. Figure 10b represents the GDnet framework, a biological deep learning model that fuses transcriptomic and regimen representations. Transcriptomic data is rank-normalized and drug representations in the treatment regimen are summed gene-wise to be input into a neural network. Binary masks containing information about dependencies between genes and / or pathways are incorporated into the network to impose constraints on nodes and edges; IC50: half maximal inhibitory concentration; Gcd: gene-drug sensitivity correlation; Ge: gene expression;

[0052] Figure 11 is a heatmap visualization of each drug-gene correlation provided by embodiments of the present invention, with correlation values represented by color coding. Rows and columns are clustered using consensus clustering, and drugs are annotated by type. PDCD1 stands for anti-PD-1 therapy (pembrolizumab and durvalumab), ANG stands for ANG peptide inhibitors (AMG-386), VEGFA stands for anti-VEGFA antibodies (bevacizumab), IGF1R stands for anti-IGF1R antibodies (ganitumab), capecitabine stands for 5-fluorouracil, ganetespib stands for luminespib, and ixabepilone stands for epothilone B;

[0053] Figure 12 is a comparison of the computational performance of GDnet with other models with an external validation set provided by embodiments of the present invention. The reported performance metrics include the area under the ROC curve (AUC) ( Figure 12 a) and the area under the precision-recall curve (AUPRC) ( Figure 12 b). In terms of these two metrics, the median of GDnet is significantly better than other models. Figure 12 a- Figure 12 Data in b is represented as a box plot, where the middle line is the median, and the lower and upper hinges correspond to the 1st and 3rd quartiles, respectively. Dotted lines correspond to the minimum or maximum values that are not more than 1.5xIQR away from the hinges (where IQR is the interquartile range). Any data that falls outside the dotted lines is considered an outlier. For all comparisons between GDnet and other models, and between GDNNet and GANNnet, p<0.05. Figure 12 c represents the ROC curve of GDnet compared to other models (n=912). Figure 12 The bar plot in d shows the AUC of GDnet and other models in integrating biological knowledge in each individual GEO dataset in the test set. Datasets are ordered by sample size, with the largest on the left and the smallest on the right. All models are trained on the training set and tested on the test set. Figure 12 c) and ( Figure 12 d) with 10 repeats of integration;

[0054] Figure 13is the GDnet check and interpretation provided by embodiments of the present application. Figure 13 a, visualization of the inner layers of the GDnet, showing the estimated relative importance of different nodes in each layer. The nodes of the leftmost layer represent input types, the nodes of the second layer represent genes, the following layers represent higher-level biological entities, and the last layer represents the model outcome. The longer the node, the more important it is. The contribution of a source node to a target node, such as the importance of each input type (drug or transcriptome) to each gene, the contribution of each gene to each pathway, and the contribution of each pathway to each broader pathway, is described by a Sankey diagram, for example, the importance of the PSMC5 gene is mainly driven by the transcriptome. Figure 13 b, the proportion of regimens used by the optimization group (middle) and the control group (right) and the difference in the proportion between the two groups (left), as well as the external validation set of simulation tests based on different optimization thresholds. The optimized thresholds are color-coded. Opt: optimization group Ctl: control group;

[0055] Figure 14 is the model overview provided by embodiments of the present application for comparison with GDnet. Figure 14 a, GAnet framework, which integrates drug representation and transcriptome data. First, the transcriptome data and regimen representation are aggregated using a function, where Gatt represents the attention value of gene expression when treated. Then the Gatt value is input into the Figure 10 b, the bioinformatics neural network described in b. Figure 14 b, the Gnetens framework. Different transcriptome-only models are constructed for each drug type, and the predictions of all transcriptome-only models corresponding to all drugs administered to all patients are integrated by averaging to produce the final prediction. Figure 14 c, Gnet framework transcriptome-only model. The transcriptome data is directly input into the Figure 10 b, the bioinformatics neural network described in b. Gcd: gene-drug sensitivity correlation; Ge: gene expression; Gatt: gene expression attention value;

[0056] Figure 15 is the heat map provided by embodiments of the present application showing the combination of all patient treatment drugs. The applied drugs are marked as 1, color-coded as red, and the unapplied drugs are marked as 0, color-coded as blue. The ER status, HER2 status, and pCR result are annotated at the top, where white represents missing values;

[0057] Figure 16 is the performance comparison of GDnet and other models using 10-fold cross-validation on the training set provided by embodiments of the present application. The indicators include the area under the ROC curve (AUC) (a), the area under the precision-recall curve (AUPRC) (b), and the area under the precision-recall curve (AUPRC) (b). Figure 16 Figure 16 ​a and Figure 16 Data in b are represented as box plots with the middle line as the median, lower and upper hinges corresponding to the first and third quartiles, and whiskers corresponding to the minimum or maximum hinges within 5xIQR range (where IQR is the interquartile range). Data points beyond the whiskers are considered outliers. No statistical significance was observed in all comparison combinations; Figure 1

[0058] Figure 17 Figure 17 is a computer clinical trial simulation based on ISPY-2 trial provided by embodiments of the present application. The sample size of each group varies with the optimization strength (N%) before propensity score matching (PSM) Figure 17 a, bar plots showing AUC of all models for each individual GEO dataset in the test set. Figure 17 b, bar plots showing AUPRC of all models for each individual GEO dataset in the test set. Datasets are ordered by sample size, Figure 17 a- Figure 17 Left max and right min in b. Both models are ensemble of 10 repeats;

[0059] Figure 18 Figure 17 is a computer clinical trial simulation based on ISPY-2 trial provided by embodiments of the present application. The sample size of each group varies with the optimization strength (N%) before propensity score matching (PSM) Figure 18 a) and after PSM Figure 18 e). Baseline tumor transcriptomic state (inferred by Gnet prediction score) Figure 18 b) within the optimization and control groups at different optimization thresholds before and after PSM Figure 18 f), the difference in Gnet scores between the two groups was tested by Wilcoxon ranksum test using the p-value labeled on the green solid line of the optimization group. Figure 18 c, pCR rates of the optimization and control groups at different optimization strength levels. The difference in pCR rates was tested by Fisher’s exact test with p-value labeled on the green solid line of the optimization group. The sample size of the control group is Figure 18 d) and inverse probability of treatment weighted (IPTW) Figure 18 g) to calculate the pCR odds ratio of the optimization group versus the control group. The dashed line is the fitted linear regression function between the y-axis and the optimization threshold, R 2 and p-value are shown on top.

[0060] Figure 19 Figure 17 is a computer clinical trial simulation based on ISPY-2 trial provided by embodiments of the present application. The sample size of each group varies with the optimization strength (N%) before propensity score matching (PSM) Figure 19 a) and after PSM Figure 19 ​e) After that, the sample size of each group varies with the change of the optimization intensity (N%). Baseline tumor transcriptome status (inferred by Gnet prediction score) Figure 19 b) Within the optimization group and the control group at different optimization thresholds before and after PSM Figure 19 f) Difference of Gnet scores between the two groups was tested by Wilcoxon rank sum test. The p-value is labeled on the green line of the optimization group. Figure 19 c) pCR rates of the optimization group and the control group at different optimization intensity levels. The difference of pCR rates was tested by Fisher's exact test. The p-value is labeled on the green line of the optimization group. By PSM Figure 19 d) and inverse probability of treatment weighting (IPTW) Figure 19 g) Calculate the pCR odds ratio of the optimization group and the control group. The dotted line is the fitted linear regression function between the y-axis and the optimization threshold, R 2 and p-value is shown at the top. DETAILED DESCRIPTION

[0061] In order to make the person skilled in the art better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application.

[0062] In some of the descriptions in the specification and claims of the present application and the above-mentioned drawings, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed or in parallel without the order appearing in the text. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second" and the like described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. Also, "first" and "second" are different types.

[0063] The technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0064] Figure 1is a construction method flowchart of a digital organoid provided by an embodiment of the present application. Specifically, the method comprises the following steps: 101: obtaining optional treatment schemes, real treatment schemes and transcriptome data of training set samples; the treatment scheme comprises a scheme composed of at least one treatment drug; optionally, the treatment drug comprises any one or more of the following: small molecule drugs, antibody drugs; the small molecule drug is represented by calculating the IC50 value and the correlation between the genes significantly related to the small molecule drug in the whole genome data, denoted as Gcd; specifically, the representation of the small molecule drug comprises: calculating the IC50 value of the small molecule drug (used to represent drug sensitivity, wherein high IC50 corresponds to low cell killing ability, and low IC50 corresponds to high cell killing ability; the lower the IC50 value, the higher the sensitivity); determining the genes significantly related to the small molecule drug in the whole genome data based on the IC50 value, denoted as Ge; calculating the correlation between the IC50 value and the significantly related genes, denoted as gene-drug sensitivity correlation Gcd; Gcd is used to represent the small molecule drug;

[0065] Optionally, if the first small molecule drug is missing in the genomics database, the most relevant drug in the database with similar mechanism of action and the same antitumor drug type is used to represent the first small molecule drug;

[0066] Optionally, if the genomics database comprises at least two, and the second small molecule drug exists in at least two databases at the same time, the database with high recognition degree is preferentially used to calculate the representation of the second small molecule drug; for example, when the database comprises GDSC and CTRP data sets, the GDSC data set is used.

[0067] Optionally, the small molecule drug comprises any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; optionally, the drug does not include drugs without corresponding drug sensitivity data in the GDSC or CTRP database, such as: T-DM1, endocrine therapy drugs.

[0068] Optionally, the cancer drug sensitivity genomics database comprises GDSC and CTRP data sets.

[0069] Optionally, the antibody drug is represented by calculating the correlation between the target gene and the gene significantly related to the target gene in the whole genome data, denoted as Gcd; specifically, the representation of the antibody drug includes: determining the target gene of the antibody drug; determining the gene significantly related to the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the significantly related gene, denoted as gene-drug sensitivity correlation Gcd; Gcd is used to characterize the antibody drug;

[0070] Preferably, the antibody drug is represented by calculating the negative value of the correlation between the target gene and the gene significantly related to the target gene in the whole genome data, and a negative sign is added before the expression level of the target gene of the antibody drug; optionally, the antibody drug includes: HER2 drug;

[0071] Optionally, the calculation method of the correlation is a commonly used statistical method, such as Pearson correlation coefficient, etc. The value range of the correlation is greater than or equal to -1 and less than or equal to 1;

[0072] When the treatment scheme includes at least 2 treatment drugs, the sum of the drug sensitivity correlations Gcd of the same gene in the representation of each drug is used to represent the treatment scheme;

[0073] 102, inputting the optional treatment scheme and the transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after treatment of each optional treatment scheme;

[0074] In some embodiments, the construction method of the GDnet model includes: obtaining the whole genome data and the true treatment scheme of the breast cancer training set samples that have received neoadjuvant therapy; the treatment drug is represented by the gene sensitivity correlation Gcd of the pan-cancer cell line in the cancer drug sensitivity genomics database; inputting the whole genome data and the true treatment scheme into a biological information neural network model for training to obtain the GDnet model; optionally, the biological information neural network model includes at least an input layer, a gene layer, at least one biological process layer, and an output layer; optionally, the biological information neural network model includes: a GDnet framework.

[0075] In some embodiments, the significantly related gene expression is denoted as Ge, and the input method of inputting the whole genome data and the true treatment scheme into the biological information neural network model for training includes: inputting Ge and Gcd into the neural network model in parallel, or obtaining the attention value Gatt of the gene expression when receiving treatment by aggregating Ge and Gcd, and inputting the attention value into the neural network model; optionally, the calculation method of the attention value Gatt includes: Gatt=Gex*(Exp-Gcd).

[0076] 103, assigning the optional treatment regimen with the predicted pCR probability greater than a first threshold (N%) to a best group, and assigning the optional treatment regimen with the predicted pCR probability less than the first threshold to a worst group;

[0077] 104, classifying the training set sample based on whether the real treatment regimen belongs to the best group or the worst group, assigning the sample to an optimization group if the real treatment regimen is in the best group, and assigning the sample to a control group if the real treatment regimen is in the worst group, to obtain a digital organoid including the optimization group and the control group.

[0078] The second aspect of the present application discloses a method for predicting sensitivity of a treatment regimen, the method comprising:

[0079] 201, obtaining transcriptome data of a subject and a candidate treatment regimen; 202, inputting the transcriptome data and the candidate treatment regimen into the digital organoid of the first aspect of the present application, and outputting a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of high sensitivity of the candidate treatment regimen is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate treatment regimen is obtained;

[0080] Optionally, the method further comprises: assisting in selecting a treatment regimen based on the classification result; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of the candidate treatment regimen being suitable is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of the candidate treatment regimen being unsuitable is obtained.

[0081] The third aspect of the present application discloses a method for predicting sensitivity of a treatment regimen, the method comprising:

[0082] 301, obtaining transcriptome data of a subject and a candidate treatment regimen; the genes in the transcriptome data include any one or several of PSMC5, CCND1, PSMB3, PSMD14, and CUL1;

[0083] 302, inputting the transcriptome data and the candidate treatment regimen into the digital organoid of the first aspect of the present application, and outputting a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of high sensitivity of the candidate treatment regimen is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate treatment regimen is obtained;

[0084] Optionally, the drug species in the candidate treatment regimen includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG;

[0085] Optionally, the method further comprises: assisting in selecting a treatment regimen based on the classification result; if the classification result obtained is that the subject belongs to the optimization group, an assisted prediction result that the candidate treatment regimen is suitable is obtained; if the classification result obtained is that the subject belongs to the control group, an assisted prediction result that the candidate treatment regimen is unsuitable is obtained.

[0086] Optionally, if the candidate treatment regimen includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, 5-fluorouracil, epirubicin, cyclophosphamide, ERBB2, and paclitaxel, ERBB2, ERBB2MK-2206, paclitaxel, doxorubicin, cyclophosphamide, PDCD1, and paclitaxel, doxorubicin, cyclophosphamide

[0087] | ENG | ERBB2, an assisted prediction result that the subject has a high probability of belonging to the optimization group is obtained; if the candidate treatment regimen includes any one or several of the following: docetaxel, 5-fluorouracil, 5-fluorouracil, epirubicin, cyclophosphamide, docetaxel, and doxorubicin, an assisted prediction result that the subject has a high probability of belonging to the control group is given.

[0088] The fourth aspect of the present application discloses a method for predicting the sensitivity of a therapeutic drug, the method comprising:

[0089] 401, obtaining the transcriptome data of a subject and a candidate therapeutic drug; 402, inputting the transcriptome data and the candidate therapeutic drug into the digital organoid of the first aspect of the present application, and outputting a classification result that the subject belongs to an optimization group or a control group; if the classification result obtained is that the subject belongs to the optimization group, an assisted prediction result that the candidate therapeutic drug has high sensitivity is obtained; if the classification result obtained is that the subject belongs to the control group, an assisted prediction result that the candidate therapeutic drug has low sensitivity is obtained; optionally, the candidate drug further includes a new drug not included in the training set.

[0090] The fifth aspect of the present application discloses a method for predicting the sensitivity of a therapeutic drug, the method comprising:

[0091] 501, obtaining the transcriptome data of a subject and a candidate therapeutic drug; the genes in the transcriptome data include any one or several of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1;

[0092] 502, inputting the transcriptome data and the candidate therapeutic drug into the digital organoid of the first aspect of the application, and outputting a classification result of the subject belonging to an optimization group or a control group; if the classification result of the subject belonging to the optimization group is obtained, an auxiliary prediction result of high sensitivity of the candidate therapeutic drug is obtained; if the classification result of the subject belonging to the control group is obtained, an auxiliary prediction result of low sensitivity of the candidate therapeutic drug is obtained;

[0093] Optionally, the candidate therapeutic drug species includes any one or several of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; optionally, the candidate therapeutic drug further includes a new drug not included in the training set;

[0094] Optionally, if the candidate therapeutic drug includes any one or several of PDCD1, ERBB2, olaparib, lapatinib, and MK-2206, a high-probability auxiliary prediction result of high probability of predicting pCR is given.

[0095] In some embodiments, the term "subject" or "patient" or "testee" or "test sample" used herein refers to any animal (e.g., a mammal), including but not limited to a human, a non-human primate, a rodent, etc., who will be the recipient of a particular treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to a human subject. Preferably, the subject is a human.

[0096] In some embodiments, the auxiliary prediction result includes but is not limited to a paper or electronic report form, which is only an analysis result of an intelligent machine based on relevant data of the subject, and is only a reference for medical personnel, and is not the final diagnosis result of the subject. In some embodiments, the first threshold is obtained by training the training set samples, which can be a specific threshold or an interval range, and the specific form is not limited in the present embodiment.

[0097] Figure 6 is a schematic diagram of a computer device provided by an embodiment of the application, as shown in Figure 6 the device can include one or more processors and one or more memories; wherein the memory has computer readable code stored therein, which, when executed by the one or more processors, can perform the method as described above.

[0098] The processor in this embodiment can be an integrated circuit chip with signal processing capability. The processor can be a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. The methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be X86 architecture or ARM architecture.

[0099] Generally, various example embodiments of the present disclosure can be implemented in hardware or special-purpose circuitry, software, firmware, logic, or any combination thereof. Certain aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device. When aspects of the present disclosure are illustrated or described as a block diagram, flow chart, or using some other pictorial representation, it will be understood that the blocks, apparatus, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special-purpose circuitry, general purpose hardware or controller or other computing device, or some combination thereof.

[0100] For example, the methods or apparatuses according to embodiments of the present disclosure can also be implemented by means of Figure 7 The architecture of the computing device 3000 shown can be implemented. As Figure 7 shown, the computing device 3000 can include a bus 3010, one or more CPUs 3020, read-only memory (ROM) 3030, random access memory (RAM) 3040, a communication port connected to a network 3050, input / output components 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as ROM 3030 or hard disk 3070, can store various data or files used in processing and / or communication of the methods provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 can also include a user interface 3080. Of course, Figure 7 The architecture shown is only exemplary, and when implementing different devices, one or more components in the computing device shown can be omitted Figure 7 in the computing device.

[0101] The embodiments of the present disclosure are also a computer readable storage medium, such as Figure 8As shown, the schematic diagram of the storage medium provided by the embodiment of the present application, the computer storage medium 4020 stores computer readable instructions 4010. When the computer readable instructions 4010 are run by the processor, the method according to the embodiment of the present application can be executed, which is described with reference to the above drawings. The computer readable storage medium in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example and not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRAM). It should be noted that the memory of the method described herein is intended to include but not limited to these and any other suitable types of memory. It should be noted that the memory of the method described herein is intended to include but not limited to these and any other suitable types of memory.

[0102] The embodiment of the present application also provides a computer program product or system, comprising a computer program which is executed by a processor to realize the steps of the above method.

[0103] In some embodiments, the embodiment also discloses a digital organoid construction system, the system comprising:

[0104] The first acquisition module 601 is configured to acquire the optional treatment scheme, the real treatment scheme and the transcriptome data of the training set samples; the treatment scheme comprises a scheme composed of at least one treatment drug;

[0105] The pCR probability calculation module 602 is configured to input the optional treatment scheme and the transcriptome data into the GDnet model, and calculate the predicted pCR probability of different samples after treatment of each optional treatment scheme;

[0106] The optional treatment scheme distribution module 603 is configured to distribute the optional treatment scheme with the predicted pCR probability greater than the first threshold to the best group, and distribute the optional treatment scheme with the predicted pCR probability less than the first threshold to the worst group;

[0107] The model training module 604 is configured to classify the training set samples based on whether the real treatment scheme belongs to the best group or the worst group, and assign the sample to the optimization group if the real treatment scheme is in the best group, or assign the sample to the control group if the real treatment scheme is in the worst group, to obtain a digital organoid including the optimization group and the control group.

[0108] In some embodiments, the present embodiment also discloses a system for predicting the sensitivity of a treatment scheme, which comprises:

[0109] The second acquisition module 701 is configured to acquire the transcriptome data of a subject and a candidate treatment scheme.

[0110] The first result output module 702 is configured to input the transcriptome data and the candidate treatment scheme into the digital organoid of the first aspect of the present application, and output a classification result of whether the subject belongs to the optimization group or the control group.

[0111] Optionally, the system further comprises a second result output module 703 configured to assist in selecting a treatment scheme based on the classification result.

[0112] In some embodiments, the present embodiment also discloses a system for predicting the sensitivity of a treatment scheme, which comprises:

[0113] The third acquisition module 801 is configured to acquire the transcriptome data of a subject and a candidate treatment scheme.

[0114] The third result output module 802 is configured to input the transcriptome data and the candidate treatment scheme into the digital organoid of the first aspect of the present application, and output a classification result of whether the subject belongs to the optimization group or the control group.

[0115] Optionally, the system further comprises a fourth result output module 803: based on the classification result, assisting in selecting a treatment scheme; if the classification result shows that the subject belongs to the optimization group, obtaining an auxiliary prediction result that the candidate treatment scheme is suitable; if the classification result shows that the subject belongs to the control group, obtaining an auxiliary prediction result that the candidate treatment scheme is unsuitable.

[0116] In some embodiments, the present embodiment also discloses a system for predicting the sensitivity of a treatment drug, comprising:

[0117] A fourth acquisition module 901 for or configured to acquire transcriptome data of a subject and a candidate treatment drug; optionally, the candidate drug further comprises a new drug not included in the training set.

[0118] A fifth result output module 902 for or configured to input the transcriptome data and the candidate treatment drug into the digital organoid of the first aspect of the present application, and output a classification result that the subject belongs to an optimization group or a control group; if the classification result shows that the subject belongs to the optimization group, an auxiliary prediction result that the candidate treatment drug has high sensitivity is obtained; if the classification result shows that the subject belongs to the control group, an auxiliary prediction result that the candidate treatment drug has low sensitivity is obtained.

[0119] In some embodiments, the present embodiment also discloses a system for predicting the sensitivity of a treatment drug, comprising:

[0120] A fifth acquisition module 1001 for or configured to acquire transcriptome data of a subject and a candidate treatment drug; the genes in the transcriptome data include any one or several of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; optionally, the candidate treatment drug further comprises a new drug not included in the training set;

[0121] A sixth result output module 1002 for or configured to input the transcriptome data and the candidate treatment drug into the digital organoid of the first aspect of the present application, and output a classification result that the subject belongs to an optimization group or a control group; if the classification result shows that the subject belongs to the optimization group, an auxiliary prediction result that the candidate treatment drug has high sensitivity is obtained; if the classification result shows that the subject belongs to the control group, an auxiliary prediction result that the candidate treatment drug has low sensitivity is obtained. Specific embodiments:

[0123] Results: GDnet workflow schematic: To fully integrate drug and transcriptomic data, mapping drugs to a series of gene and value pairs is one of the most suitable approaches. The selection of gene and value pairs is a key consideration. The GDSC and CTRP datasets, which contain transcriptomic states of over 800 cell lines and their drug sensitivity to over 500 molecular drugs, are good reference datasets. The heterogeneous responses of different cancer cell lines to drugs in GDSC or CTRP can be attributed to the heterogeneity of tumor transcriptomic states; therefore, genes that are significantly associated with drug sensitivity are more likely to be genes associated with response or resistance, and thus more likely to be genes that are perturbed when a tumor is exposed to a particular drug. Therefore, we used the genome-wide expression correlation of the half maximal inhibitory concentration (IC50) values (where lower values indicate higher sensitivity) of each drug in different cell lines to represent each drug Figure 10 a). It is worth noting that antibody or peptide drugs (e.g. anti-HER2 drugs) are not included in these drug sensitivity datasets, and IC50 values are not traditionally used to measure drug sensitivity through complex mechanisms. Given the strong positive correlation between the sensitivity of these targeted drugs and the expression of the targeted genes, we used the negative values of the genome-wide expression correlation with the expression of the drug target as a representation of antibody drugs Figure 10 a). For a treatment regimen of multiple drugs, we performed gene summation on the representation of each drug in the treatment regimen Figure 10 b).

[0124] To better simulate the interaction between drugs and transcriptomes, we input the rank-normalized transcriptome and regimen representation into the gene-level layer of the subsequent neural network in parallel, which we call GDnet Figure 10 b). We further explored 3 additional model frameworks: (1) GAnet, which integrates drugs and transcriptomes using a method that mimics the key-value attention mechanism Figure 14 a); (2) Gnetens, which uses an ensemble learning approach to integrate all drugs predicted for a patient Figure 14 b); (3) Gnet, which relies only on transcriptomic information Figure 14 c). Next, the gene layer is pushed forward to the bioinformatic neural network proposed by H. A. et al., which incorporates the dependency (child-parent) relationships between genes, fine pathways, and more complex pathways based on the Reactome Pathways dataset (version 2022) Figure 10 b). The final constructed prediction model aims to accurately predict the response of breast cancer patients to neoadjuvant therapy, assisting in the selection of multiple candidate regimens.

[0125] Drug representations based on GDSC and CTRP datasets: To generate gene and value pairs representing drugs, we computed the correlation of IC50 values of various drugs in different cell lines with gene expression, and with the expression level of antibody or polypeptide drug targets. This computation was done for all combinations of drugs and genes. For drugs that appear in both datasets, the GDSC correlation was prioritized. Most drugs given to patients have their corresponding representation in either the GDSC or CTRP dataset. However, a few small molecule drugs are missing in both datasets, we replaced them with the most relevant drugs in the GDSC / CTRP dataset with similar mechanism and same antitumor drug class, respectively. The representation of all drugs used in the existing breast cancer neoadjuvant dataset was summarized as Figure 11 ). Methotrexate has the most number of gene-value pairs, with 5378 pairs. Among the 22 drug representations, 17 have more than 1000 gene-value pairs, and 5 (ANG, lapatinib, carboplatin, epothilone B, doxorubicin) have less than 1000. Consistent clustering of drugs groups the 5 drugs with fewer number of pairs together Figure 11 ). For drugs with larger number of pairs, three main clusters emerged, the first cluster contains anti-angiogenic therapies, anti-HER2 therapies, and anti-IGF1R therapies, the second cluster contains chemotherapy therapies, AKT inhibitors, and PARP inhibitors, and the last cluster contains anti-PD-1 therapies and HSP90 inhibitors Figure 11 ). These drug representations were then combined with patient transcriptomes to predict neoadjuvant therapy response.

[0126] GDnet shows superiority in predicting neoadjuvant therapy outcome: We collected 9550 breast cancer samples from 68 independent public datasets. After removing duplicates and filtering out samples with non-preprocessed state or missing critical information (outcome and therapy), 31 datasets containing 4371 samples and 16133 genes were filtered out. The 17 datasets with the most number of cases (totaling 3459 samples) were used to train the model. Ten-fold cross-validation was adopted for hyperparameter tuning and internal validation, with a final learning rate of 0.005. Due to the sparse nature of the model, L2 regularization was not applied (set to 0). The remaining 14 datasets with the least number of cases, totaling 912 samples, were used as an external validation set. After intersecting with genes involved in biological networks from the Reactome dataset, a total of 8377 genes were finally retained.

[0127] For benchmarking purposes, we also trained dense deep neural network (DNN) variants of GDnet, GAnet, Gnetens, and Gnet (without biological knowledge), named GDNNet, GANNet, GNNetens, and GNNet, respectively. In Gnetens and GNNetens, separate models were constructed using only drugs that occurred in more than 10% of samples. Five drug types were selected (anthracyclines, microtubule inhibitors, cyclophosphamide, pyrimidine analogs, and anti-HER2 antibodies) Figure 15 ). Then, ten-fold cross-validation was applied using the training set to compare model performance. Results showed that the area under the curve (AUC) and the area under the precision recall curve (AUPRC) of GDnet (median AUC = 0.808, median AUPRC = 0.739) were higher than all other models (median AUC = 0.799, median AUPRC = 0.731 for Gnet) in the interval validation, although not significantly Figure 16 ). These models were trained 10 times using all training sets and tested separately using the external validation set, facilitating a robust comparison between different models. In the external validation, GDnet performed significantly better than all other models (median AUC = 0.712, median AUPRC = 0.691) (median AUC = 0.676; median AUPRC = 0.632 for Gnet) Figure 12 a-b). Notably, the performance of GDNNet, which has the largest number of parameters (AUC = 0.678, AUPRC = 0.661), was lower than GANNet (AUC = 0.695, AUPRC = 0.668), demonstrating the importance of prior knowledge in guiding training and improving performance. To robustly compare the ROC curves between models, we created an ensemble by averaging the predictions of the models trained 10 times to make a final prediction. Again, GDnet (AUC = 0.725) achieved the highest AUC Figure 12 c), whose ROC curve was significantly different from that of Gnet (AUC = 0.683) (Delong’s test, p = 0.037). Then, we tested the performance of GDnet in 13 separate validation datasets by 10 repeated ensemble models to determine whether GDnet could perform well in limited protocol settings compared to Gnet. The GSE191127 dataset was not included in the individual testing because it contained only residual disease (RD) patients. Results showed that GDnet achieved higher AUC than Gnet on 8 out of 13 datasets Figure 12d). Although GDnet shows lower AUC in the left 5 datasets, all differences are less than 0.03. Moreover, GDnet achieves higher AUC than GAnet and Gnetens on 7 out of 13 datasets Figure 12 d). GDnet also shows higher AUC than all dense models in each dataset Figure 17 a). Similar results are observed for AUPRC values Figure 17 b). These results show that GDnet successfully captures the interaction patterns between drugs and transcriptomes, rather than simply overfitting the drug distribution bias between datasets.

[0128] GDnet can optimize regimen selection for better outcomes: To evaluate the ability of GDnet to guide clinical regimen selection, we subsequently conducted a computer simulated clinical trial to examine whether GDnet can improve neoadjuvant therapy outcomes for breast cancer Figure 9 a). Given multiple possible treatment regimens and patient transcriptomes, we input them into the GDnet model to compute the predicted probability of pathological complete response (pCR) for each case. We select the top N% of best regimens that are likely to benefit the patient, and the bottom N% of worst regimens that are not suitable for the patient, based on the highest and lowest predicted pCR probabilities, respectively. If the regimen actually used by the patient is among the best ones, we assign the patient to the optimization group, and if it is among the worst ones, we assign the patient to the control group. To check the baseline characteristics of the two groups, we also computed the predicted scores by the transcriptome-only model to assess the tumor malignancy between them, with preference given to Gnet, as it is similar in structure to GDnet and more robust in performance than GNNet (AUC Levene’s test, p = 0.026, AUPRC Levene’s test, p = 0.017) Figure 12 a-b).

[0129] We first performed a series of computer trials based on the ISPY-2 trial (one of the datasets used to train our model). A total of 653 patients who received one of the 11 regimens in ISPY-2 participated in our simulation study, with these regimens being set as optional choices. We set different optimization thresholds, starting from 50%, with a 5% decrement (e.g., 50%, 45%, 40%, 35%), until the size of either group dropped below 30 Figure 18 a, Figure 18 e), Figure 19 a, Figure 19 e), to classify the best and worst regimens, ensuring the robustness of the results. The lower the threshold, the higher the optimization stringency. The results show that the pCR rate of the optimization group is higher than that of the control group in all optimization thresholds Figure 18c) were significantly higher than those in the control group. In addition, with the strengthening of optimization (lowering of optimization threshold), the odds ratio (OR) of pCR increased linearly (R2=0.72, p=0.008) ( Figure 18 d). However, higher Gnet prediction scores were observed in the optimized group ( Figure 18 b), indicating that the tumors in this group were less malignant. This difference may affect the performance of our model. Therefore, we performed propensity score matching (PSM) to compare the Gnet prediction score ( Figure 18 f) to balance the transcriptome profile. Similarly, the pCR rate of the optimized group was significantly higher than that of the control group ( Figure 9 b), with the strengthening of optimization, the OR increased linearly (R2=0.87, p<0.001) ( Figure 9 c). This was also confirmed by the inverse probability of treatment weight (IPTW) analysis (R2=0.46, p=0.043) ( Figure 18 g), indicating that GDnet has the ability to optimize treatment options.

[0130] We further conducted a series of simulations using the external validation dataset. Given the limited variety of treatments in each validation dataset, we combined all available validation sets, resulting in 912 patients and 18 treatments. Similar to what was observed in the ISPY-2 simulations above, the results from the external validation set showed that the pCR rate ( Figure 19 c) were significantly higher than those in the control group, among which the OR increased linearly with the strengthening of optimization (R2=0.68, p=0.003) ( Figure 19 d), even lower Gnet prediction scores were observed in the optimized group ( Figure 19 b). PSM analysis based on Gnet prediction score ( Figure 19 f) also showed that the pCR rate in the optimization group was significantly higher than that in the control group ( Figure 9 d), OR increased linearly (by PSM, R2 = 0.94, p < 0.001, by IPTW, R2 = 0.46, p = 0.03) with the strengthening of optimization ( Figure 9 e, Figure 19 g), further demonstrating the excellent ability of GDnet in guiding the selection of neoadjuvant therapy for breast cancer.

[0131] GDnet interpretation reveals important genes and drugs involved in the response: To understand the importance of features contributing to the model and the interactions between them, we visualized the layers of GDnet ( Figure 13a). The ensemble gradient attribution method was applied to obtain the importance score of each node in each layer for each patient in the external validation set. The average importance score of all patients was used as the final score for each node, which was shown as the node length in the Sankey diagram. From the 10 repeated GDnet models, we selected the GDnet model with similar AUC value (0.727) to the 10 repeated ensemble GDnet for interpretation. In terms of transcriptome and treatment regimen, the transcriptome contains more information than the drug information; however, the drug treatment regimen provides additional information for more accurate prediction of response Figure 13 a). The regimen applied in the neoadjuvant treatment of breast cancer lacks diversity (homogeneity) in the dataset, which can lead to a relatively low contribution of regimen information in the GDnet. High-score genes include SRC, CCND1, MCF2L, RPS6KA1, and PSMB7. Pathways at each level, such as the ER-Phagosome pathway in the first layer, which is reported to be associated with the innate immune system and phagocytosis, the RHOGTPasc cycle in the 2nd layer, which is reported to be associated with tumorigenesis, invasion, and metastasis, and the immune system in the 5th layer have been identified, emphasizing their key role in the performance of the GDnet. Further experiments are needed for more detailed mechanisms of neoadjuvant treatment of breast cancer.

[0132] Then, we investigated which treatment types contributed to the response. We focused on Figure 9 a. To further investigate the distribution of regimens and drugs used in each group in the PSM analysis under different optimization thresholds, the difference in the proportion of regimens and drugs in different groups can inform which regimens and drugs contribute more to the model. The results show that the paclitaxel | doxorubicin | cyclophosphamide | olaparib | PDCD1 and paclitaxel | doxorubicin | cyclophosphamide | ERBB2 | ERBB2 regimens appear much more frequently in the optimization group than in the control group, and the frequency increases with the degree of optimization Figure 13 b). Docetaxel | cyclophosphamide, doxorubicin, and 5-fluorouracil | epirubicin | cyclophosphamide | lapatinib are regimens that occur in most or even only in the control group, with a trend of increasing with the degree of optimization Figure 13 b). Specifically, drugs such as PDCD1 (PD-1 inhibitor), olaparib, ERBB2 (anti-HER2 antibody), etc. may be more suitable for the optimization of the GDnet model Figure 10) play an important role in improving pCR rate. We further investigated the regimen and drug distribution in ISPY-2 simulation, and the results of external validation were similar. Paclitaxel | doxorubicin | cyclophosphamide | ERBB2 | MK-2206, paclitaxel | doxorubicin | cyclophosphamide | PDCD1, and paclitaxel | doxorubicin | cyclophosphamide | ANG | ERBB2 were important regimens in the optimization group, in which PDCD1, ERBB2, and MK-2206 played an important role in improving pCR rate. By integrating drug data into the model, GDnet can compare all available treatment regimens before starting treatment, thus promoting precision medicine specifically for breast cancer patients, which has never been properly achieved by all previous machine learning models. In other words, we established a model that considers organoid function, which can test drug and treatment regimen sensitivity and inform patients of appropriate drugs. The results of our computer simulation tests based on the ISPY-2 trial and the external validation set in our study showed that GDnet can be used to optimize the treatment decision-making process for breast cancer patients and significantly improve the pCR rate of breast cancer patients. For each optimization threshold, GDnet maintained the ability to improve pCR rate, thus confirming the research results. In addition, there is a linear correlation between pCR rate and optimization strength, which strongly supports our conclusion that GDnet can optimize breast cancer neoadjuvant therapy and has the potential to change the relevant treatment guidelines. Our model can play an important role in many potential scenarios. In daily clinical practice, our model can help select the most suitable regimen from candidate regimens. In clinical trials, our model can help select patients who are likely to respond to new test regimens or even previously abandoned regimens. Even if the treatment regimen contains new drugs not included in the GDnet training set, our model can still predict drug response because it performs well on the validation set containing many drugs, such as durvalumab, bevacizumab, cisplatin, olaparib, and gemcitabine, which are not present in the training set.

[0133] Methods: Public data collection processing: On February 15, 2024, transcriptomic data related to neoadjuvant therapy in breast cancer were retrieved from PubMed and Gene Expression Omnibus (GEO) database system. The search terms used were (“neoadjuvant therapy” [MeSH term] OR neoadjuvant) AND (“breast neoplasms” [MeSH term] OR ((breast OR breast) AND (cancer* OR tumor* OR tumor* OR cancer* OR tumor* or oncology or malignant tumor*))). High-throughput sequencing such as next-generation sequencing (NGS) and microarray were included and only datasets containing not less than 20 samples were selected. A total of 9550 breast cancer samples from 68 independent public datasets were accessed. Microarray probesets were gene annotated against the reference table provided in GEO. For gene and sample quality control, 16133 genes present in more than 50% of samples were retained, and 4580 samples with missing genes less than 50% were retained. Among these datasets, data from 31 datasets (GSE194040, GSE25066, GSE16716, GSE41998, GSE180962, GSE20271, GSE34138, GSE50948, GSE149322, GSE22226, GSE32603, GSE22358, GSE32646, GSE16446, GSE130788, GSE231629, GSE123845, GSE4779, GSE173839, GSE22093, GSE21997, GSE42822, GSE66399, GSE181574, GSE23988, GSE41656, GSE8465, GSE122630, GSE21974, GSE207248, GSE191127) encompassing complete pCR, treatment regimen, and sufficient transcriptomic information were used for model construction and validation.

[0134] Ranking value encoding of transcriptomes: RNA-seq counts or patient-level normalized data (e.g., transcripts per kilobase million (TPM) or fragments per kilobase million (FPKM)), microarray patient-level normalized data (e.g., robust multi-chip average (RMA) or other methods) were collected from GEO. If multiple-level data were accessed, a larger upstream level could be maintained. To facilitate the conversion between gene expression testing technology platforms, we applied a ranking transformation (ranking value encoding) to normalize the data. For this, gene expression values were transformed from microarray intensities or RNAseq counts or their normalized formats to respective ranks. This transformation was done on a gene-wise basis. Thus, all gene expression values for each gene were ranked according to the order from lowest to highest. The ranks were then scaled to quantiles (ranging from 0 to 1). Missing values were padded with 0. This ranking-based approach can be more robust to technical artifacts that can systematically bias absolute transcript count values, while the overall relative ranking of a gene within each cell remains more stable.

[0135] Therapeutic drugs and profiles represent: To facilitate the fusion of drug treatments and transcriptomes, our goal was to represent drugs by combining relevant genes and relevant values. To construct a large enough drug genome dataset for drug characterization, we retrieved the original drug sensitivity data of cell lines to molecular drugs as well as the RMA normalized / TPM normalized and log-transformed transcriptomes of the cell lines from the Genomics of Drug Sensitivity in Cancer database v2 (GDSC) and the Cancer Therapeutic Response Portal v2 (CTRP). We used the IC50 indicator to represent drug sensitivity, where a high IC50 value corresponds to a low cell-killing ability and a low IC50 value corresponds to a high cell-killing ability. For each drug, we computed the correlation between the gene expression levels of all cell lines and the log2(IC50+1) of the cell lines’ response to each drug. Considering that nearly 5*10 4 genes were to be tested, the significance cutoff of the correlation was set to 1*10 -6 to reduce the false discovery rate. Genes identified as significant are more likely to play an important role in predicting the drug response of cancer treatments, and the correlation value is prior knowledge that can inform the importance of a gene to the drug response, which can reflect the perturbation of the drug to the gene. If a relevant gene does not exist in the retained 16133 genes of the patient transcriptome, it is deleted. Thus, each drug is finally represented by its significant genes and correlation values (ranging from -1 to 1). If a drug appears in both the GDSC and CTRP datasets, the GDSC dataset is considered first.

[0136] There are several special cases in drug representation. Some types of drugs were not found in either the GSDC or CTRP datasets. Therefore, we selected the most relevant drugs within the same anti-tumor drug class with similar mechanism of action to represent it. Examples include capecitabine, which is represented by 5-fluorouracil; ganetespib, which is represented by luminespib; some drugs in the GEO dataset, such as taxanes or anthracyclines, lack sufficient drug details; therefore, we selected the average representation of taxanes or anthracyclines. Some types of drugs are antibody- or peptide-based drugs, which are not included in these two datasets and are usually not measured with IC50 for drug sensitivity with complex mechanisms. For these drugs, we used their target genes with the negative correlation of other cell line genes in the cell, as the reactions of these drugs are closely related to the expression of the target, including PDCD1 therapy (pembrolizumab and durvalumab), VEGFA (bevacizumab), and IGF1R (ganitumab), which is an anti-IGF1R antibody. If a drug has multiple targets, we sum the genes between all targets as the final representation, which includes ANG, containing ANGPT1 and ANGPT2 genes, representing ANG peptide inhibitors (AMG-386). In addition, T-DM1 (an antibody drug conjugate) and drugs such as endocrine therapy do not have corresponding drug sensitivity data in GDSC or CTRP. Considering the small number of patients in our dataset who received these drugs, patients who received T-DM1 or endocrine therapy were excluded. For patients who received multiple drugs in one regimen, we summed the gene correlation representation of each drug to represent the regimen.

[0137] Drug representation and transcriptome fusion: Three methods of integrating treatment regimens and transcriptome information were tested. The first method involves using gene expression (represented as Ge) and gene sensitivity correlation (represented as Gcd) as inputs, which we call GDnet( Figure 10 b). The second method uses Gatt= (Ge)(e Gcd ) as the final input, using a method similar to the key-value attention mechanism ( Figure 14 a), and the model is named GAnet. In this mechanism, for each query (Ge), we find the corresponding key (gene) in the drug representation, and the attention (Gatt) is calculated as attention= (query)(e value ), where the value represents the Gcd of the drug representation. The attention calculation is set to e valueto ensure that Gatt mimics the response of the tumor under drug exposure conditions; in other words, when the drug is not present, Ge is expected to decrease or increase relative to drug response (Gcd < 0) or resistance (Gcd > 0) genes. For example, ERBB2 is a poor prognosis gene, but when exposed to an anti-HER2 drug that targets ERBB2 as a response gene (Gcd < 0), ERBB2 expression is expected to decrease. In addition, we explored a third approach that involved the use of an ensemble learning method. This approach focuses on building a different transcriptome-only model for each drug and combining the final predictions by averaging the predictions of all transcriptome-only models corresponding to the drug applied to the patient. This model is called Gnetens ( Figure 14 b). A pure transcriptome model named Gnet was also explored ( Figure 14 c).

[0138] Design of the bioinformatics model: We introduced prior knowledge into the training process, drawing inspiration from the P-NET framework constructed by ElmarakebyHA et al. to accelerate model training and improve performance and interpretability. The complete Reactome dataset was downloaded and processed into a hierarchical network. The network consists of a layer of features, a layer of genes, and five layers of pathways. The affiliation between adjacent layer nodes (representing genes or pathways) is represented by 6 binary matrices. The constructed hierarchical network is derived from genes available in both our dataset and the Reactome dataset (first layer). The pathways directly related to these genes form the second layer. Subsequently, the paths directly related to the paths in the second layer constitute the third layer, and so on, until the sixth layer with 18 paths is constructed.

[0139] The deep learning model was implemented in PyTorch, with the number of layers and nodes in the prior knowledge network defining the structure of the basic feed-forward dense neural network model. Six binary matrices containing the dependencies between genes or pathways in adjacent layers were input into the neural network model as masks between adjacent layers. During weight updates, the mask was multiplied by the weights after each epoch, thereby imposing constraints on the nodes and edges. This configuration resulted in a sparse neural network with six layers, including 8377 nodes (gene layer), 846 (Layer 1 pathway), 218 (Layer 2 pathway), 108 (Layer 3 pathway), 51 (Layer 4 pathway), and 18 nodes (Layer 5 pathway) in each layer. Specifically, for GDnet and GDNNet, an additional gene and drug layer of 19300 nodes (half gene expression and half drug representation) was added before the 8377 gene layer. Since the model was sufficiently sparse, the L2 regularization was set to 0. Each node encodes a biological entity (e.g., a gene or a pathway), while each edge represents a known relationship between the respective entities. Node constraints can better understand the status of different biological components. Compared with a fully connected network with the same number of nodes, edge constraints result in fewer parameters, and thus less computation. Due to the imbalance of the dataset, we applied different weights to the classes to reduce the bias of the network towards a certain class according to the bias in the training set. The model was trained for 1000 rounds using the Adam optimizer with default parameters to reduce the unbalanced binary cross-entropy loss function. In order to make each layer useful in itself, we added a prediction layer with sigmoid activation after each hidden layer. gene, Layerl, Layer2, Layer3, Layer4, and Layer5 are all hidden layers, and the output gives the new adjuvant pCR probability. Since it is more challenging to fit the data using a smaller number of weights in later layers, we used higher loss weights for the results of later layers (from the gene and drug layer to the later layers, the weights are 1, 1.2, 1.4, 1.6, 1.8, 2, and 2.5) during the optimization process. The final prediction of the network is calculated by averaging all layer results.

[0140] Model training and validation: This analysis included datasets that met specific criteria, providing a total of 4371 patients available. The 17 largest datasets, including approximately 3,459 patients (79% of the total), were designated as the training set. The remaining 14 datasets, approximately 912 (21% of the total), were used as an external validation set, assigned without subjective assignment. The model was trained using the training set and optimized learning rate and other hyperparameters through ten-fold validation. The trained model was finally validated through the external test set. The validation set was used to compare the predictive performance of different models by calculating the median AUC and AUPRC of 10 replicates of training using roc auc score and average precision score implemented in the sklearn.metrics library of Python, respectively. Then, we averaged the 10 repetitions of the model to obtain the final prediction of stability and robustness. The implementation of the proposed system and reproducible results can be provided upon request and will be publicly available on GitHub upon acceptance for publication.

[0141] Model interpretation: We applied Integrated Gradients, a gradient-based attribution method implemented in the tcaptum.attr library, to calculate sample-level importance for all nodes in all layers. The absolute importance score for each layer node (always positive) was normalized to facilitate comparability between layers by the following equation:

[0142]

[0143] where Ij is the normalized absolute importance score for each node in the layer, Aj is the absolute importance score, and n is the total number of nodes in the layer.

[0144] To calculate the total node-level importance, we aggregated the sample-level importance scores for all samples in the validation set (dataset-level interpretation represented as the average of all patient-level weights in the validation set). The importance score for each node was represented as the node length in the Sankey diagram. The importance score of the link between the source node (previous layer) and the target node (next layer) was represented by multiplying the normalized (within the absolute weight of all source nodes pointing to the same target node) absolute weight of the source node pointing to the target node in the model, as shown in the following equation: Ls-t j is the importance score of the link from the jth source node (s) to the target node (t), is the normalized absolute importance score of the target node, Wsj is the weight of the jth source node pointing to the target node in the deep learning model, and n is the number of source nodes pointing to the target node.

[0145] Statistical analysis: The change in area under the ROC curve of GDnet and other models was tested by the DeLong test. The AUCs from ten-fold cross-validation and ten-fold repeated training and testing of GDnet and other models were compared by Wilcoxon rank-sum test. The variance of AUC and AUPRC between Gnet and GNNet was compared by Levene’s test to assess the robustness of performance. Odds ratio (OR) was calculated to determine the preference of pathological complete response (pCR) in simulation trials, and the difference was compared by Fisher’s exact test. Linear regression models were fitted to examine the relationship between the degree of optimization and OR. The significance of regression coefficients was assessed by t-test, and the R-squared value was used to quantify the proportion of variance explained by the model. To address potential confounding caused by imbalance in transcriptomic profiles (inferred by Gnet prediction scores) between the optimization and control groups, we employed propensity score matching (PSM) and inverse probability of treatment weighting (IPTW). PSM was performed by estimating propensity scores using logistic regression, including relevant covariates of Gnet scores. Matching was performed using nearest neighbor matching (1 : 1 ratio) and a caliper of 0.1 without replacement to minimize imbalance. In addition, the application of IPTW was by weighting each individual by the inverse probability of receiving the observed assignment, thereby creating a balanced pseudo-population. Stable weights were used to maintain accurate estimates and ensure robust results. All tests were two-sided. Unless otherwise specified, P values < 0.05 were considered statistically significant. All statistical analyses were performed by R statistical software version 4.3.

[0146] The example embodiments of the present disclosure described in detail above are merely illustrative, rather than restrictive. It should be understood by those skilled in the art that various modifications and combinations of these embodiments or features thereof can be made without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for constructing a digital organoid, characterized in that: The method comprises:

101. Obtain optional treatment regimens, actual treatment regimens, and transcriptome data for training set samples; the treatment regimen includes a regimen consisting of at least one therapeutic drug; 102. Input the optional treatment regimens and transcriptome data into the GDnet model to calculate the predicted pCR probability of different samples after treatment with each optional treatment regimen; 103, allocating the optional treatment regimen with the predicted pCR probability greater than a first threshold to the best group, and allocating the optional treatment regimen with the predicted pCR probability less than the first threshold to the worst group; 104. Classify the training set samples based on whether the true treatment regimen belongs to the best group or the worst group. If the true treatment regimen is in the best group, the sample is assigned to the optimization group; if the true treatment regimen is in the worst group, the sample is assigned to the control group, thereby obtaining digital organoids including the optimization group and the control group.

2. The method for constructing a digital organoid according to claim 1, wherein: The method for constructing the GDnet model includes: obtaining whole genome data and actual treatment plans of training set samples that have received neoadjuvant therapy; the therapeutic drug is represented by the gene sensitivity correlation Gcd of the pan-cancer cell line in the cancer drug sensitivity genomics database; Inputting the whole genome data and the actual treatment plan into a bioinformatics neural network model for training to obtain the GDnet model; Optionally, the bioinformatics neural network model includes at least an input layer, a gene layer, at least one biological process layer, and an output layer; the input layer is a Gene+drug layer, and gene, Layer 1, Layer 2, Layer 3, Layer 4, and Layer 5 are all hidden layers, and finally an output of the neoadjuvant pCR probability is given; Optionally, the therapeutic drug includes any one or more of the following: small molecule drugs, antibody drugs; the small molecule drug is represented by calculating the IC50 value and the correlation between genes significantly correlated with the small molecule drug in the whole genome data, recorded as Gcd; the representation of the small molecule drug includes: calculating the IC50 value of the small molecule drug; determining the genes significantly correlated with the small molecule drug in the whole genome data based on the IC50 value, recorded as Ge; calculating the correlation between the IC50 value and the significantly correlated genes, recorded as gene-drug sensitivity correlation Gcd; Gcd is used to characterize the small molecule drug; Optionally, if the first small molecule drug is missing from the genomic database, the most relevant drug in the database with a similar mechanism of action and the same anti-tumor drug type is used to represent the first small molecule drug; Optionally, if the genomic databases include at least two, and the second small molecule drug exists in at least two databases, the database with the higher recognition is preferentially used to calculate the representation of the second small molecule drug; Optionally, the small molecule drug includes any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; Optionally, the cancer drug sensitivity genomics database includes GDSC and CTRP datasets; Optionally, the antibody drug is represented by calculating the correlation between the target gene and the gene significantly correlated with the target gene in the whole genome data, which is recorded as Gcd; specifically, the representation of the antibody drug includes: determining the target gene of the antibody drug; determining the gene significantly correlated with the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the significantly correlated gene, which is recorded as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize the antibody drug; Preferably, the antibody drug is represented by calculating the negative value of the correlation between the target gene and the gene significantly correlated with the target gene in the whole genome data; Optionally, when the treatment regimen includes at least two therapeutic drugs, the sum of the drug sensitivity correlation Gcd of the same gene in each drug representation is calculated to represent the treatment regimen.

3. The method for constructing a digital organoid according to claim 1, wherein: The significantly correlated gene expression is denoted as Ge, and the input method of inputting the whole genome data and the actual treatment plan into the bioinformatics neural network model for training includes: inputting Ge and Gcd into the neural network model in parallel, or aggregating Ge and Gcd to obtain the attention value Gatt of the gene expression when receiving treatment, and inputting the attention value into the neural network model; Optionally, the calculation method of the attention value Gatt includes: Gatt = Ge x *(Exp -Gcd ); Optionally, the bioinformatics neural network model includes: a GDnet framework.

4. A method for predicting sensitivity to a treatment regimen, characterized in that The method comprises: 201, obtain transcriptome data and candidate treatment options for the subjects; 202. Input the transcriptome data and the candidate treatment regimen into the digital organoid according to any one of claims 1-3, and output a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, obtain an auxiliary prediction result of high sensitivity of the candidate treatment regimen; if the classification result of the subject belonging to the control group is obtained, obtain an auxiliary prediction result of low sensitivity of the candidate treatment regimen; Optionally, the method further includes: assisting in selecting a treatment plan based on the classification result; if the classification result indicates that the subject belongs to the optimization group, obtaining an auxiliary prediction result that the candidate treatment plan is suitable; if the classification result indicates that the subject belongs to the control group, obtaining an auxiliary prediction result that the candidate treatment plan is unsuitable.

5. A method for predicting sensitivity to a treatment regimen, characterized in that The method comprises: 301, obtaining transcriptome data and candidate treatment options for a subject; wherein the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, and CUL1; 302. Input the transcriptome data and candidate treatment regimens into the digital organoid according to any one of claims 1-3, and output a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, obtain an auxiliary prediction result of high sensitivity of the candidate treatment regimen; if the classification result of the subject belonging to the control group is obtained, obtain an auxiliary prediction result of low sensitivity of the candidate treatment regimen; optionally, the types of drugs in the candidate treatment regimen include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; optionally, the method further comprises: assisting in selecting a treatment regimen based on the classification result; if the classification result of the subject belonging to the optimization group is obtained, obtain an auxiliary prediction result of appropriateness of the candidate treatment regimen; if the classification result of the subject belonging to the control group is obtained, obtain an auxiliary prediction result of inappropriateness of the candidate treatment regimen; Optionally, if the candidate treatment regimens include any one or more of paclitaxel|doxorubicin|cyclophosphamide|olaparib|PDCD1, 5-fluorouracil|epirubicin|cyclophosphamide|lapatinib|ERBB2, 5-fluorouracil|epirubicin|cyclophosphamide|ERBB2 and paclitaxel|ERBB2ERBB2, paclitaxel|doxorubicin|cyclophosphamide|ERBB2MK-2206, paclitaxel|doxorubicin|cyclophosphamide|PDCD1 and paclitaxel|doxorubicin|cyclophosphamide|ENG|ERBB2, an auxiliary prediction result with a high probability that the subject belongs to the optimized group is obtained; Optionally, if the candidate treatment options include any one or more of docetaxel|5-fluorouracil|5-fluorouracil|epirubicin|cyclophosphamide, docetaxel, and doxorubicin, an auxiliary prediction result with a high probability of the subject belonging to the control group is given.

6. A method for predicting sensitivity to therapeutic drugs, characterized in that The method comprises: 401, obtaining transcriptome data and candidate therapeutic drugs of the subjects; 402. Input the transcriptome data and the candidate therapeutic drug into the digital organoid according to any one of claims 1-3, and output a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, obtain an auxiliary prediction result of high sensitivity to the candidate therapeutic drug; if the classification result of the subject belonging to the control group is obtained, obtain an auxiliary prediction result of low sensitivity to the candidate therapeutic drug; optionally, the candidate drug also includes new drugs not included in the training set.

7. A method for predicting sensitivity to therapeutic drugs, characterized in that: The method comprises: 501, obtaining transcriptome data of the subject and candidate therapeutic drugs; the genes in the transcriptome data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 502. Input transcriptome data and candidate therapeutic drugs into the digital organoid according to any one of claims 1-3, and output a classification result of whether the subject belongs to the optimization group or the control group; if the classification result of the subject belonging to the optimization group is obtained, obtain an auxiliary prediction result of high sensitivity to the candidate therapeutic drug; if the classification result of the subject belonging to the control group is obtained, obtain an auxiliary prediction result of low sensitivity to the candidate therapeutic drug; optionally, the candidate therapeutic drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; optionally, the candidate therapeutic drugs also include new drugs not included in the training set; Optionally, if the candidate therapeutic drugs include any one or more of PDCD1, ERBB2, olaparib, lapatinib and MK-2206, an auxiliary prediction result with a high probability of predicting a high probability of pCR is given.

8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • PCR prediction method based on multi-type medical data

    CN110991535A

  • Drug IC50 deep learning model prediction method based on molecular structure and gene expression

    CN114373550A

  • Prediction model-based prediction method and device for breast cancer drug regimen

    CN115376706A

  • Tumor drug sensitivity prediction method, system and equipment and storage medium

    CN115966316A

  • Deep learning prediction method for tumor medication scheme based on gene detection

    CN117079716A