Construction method and device for predicting new adjuvant therapy result model and medium

By constructing a neoadjuvant treatment outcome model combining drug characterization and biomics data based on genome-wide drug characterization, the shortcomings in predicting neoadjuvant treatment outcomes in the prior art are solved, and more accurate breast cancer treatment options and higher pathological complete remission rates are achieved.

CN120220952APending Publication Date: 2025-06-27CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510122431.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the results of neoadjuvant therapy, especially in patients with breast cancer. The existing model data type is single and the sample size is small, and it is impossible to fully characterize the tumor status, and it is difficult to integrate compound structure data and transcriptome status.

Method used

A new adjuvant treatment outcome model was constructed by calculating the IC50 value of small molecule drugs and the correlation with significantly related gene expression. The model combines drug characterization and biomics data to optimize treatment selection using a bioinformatic deep learning approach.

Benefits of technology

It significantly improves the performance of predicting neoadjuvant treatment results for breast cancer, can more accurately simulate the interaction between drugs and tumors, optimize treatment options, and improve the rate of complete pathological remission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220952A_ABST
    Figure CN120220952A_ABST
Patent Text Reader

Abstract

The invention provides a method, equipment, a medium and a program product for characterizing drugs based on a whole genome. A new adjuvant therapy result model is constructed on the basis of a method for representing drugs by using gene expression correlation, and a construction method, equipment, a medium and a program product for predicting the new adjuvant therapy result model are provided. The invention provides a method, equipment, a medium and a program product for predicting a new adjuvant therapy result in an application stage, and relates to the field of intelligent medical treatment. The new adjuvant therapy result model serves as a biological information deep learning method, drug characterization and transcriptome are fused, more accurate oncology can be achieved, and the treatment decision process of breast cancer new adjuvant therapy is enhanced. Drug characterization and biomics data are integrated to represent a novel method for simulating interaction of drugs and tumors, a digital organ platform is provided for accurate oncology, and the method can be widely applied to different treatment environments and various cancer types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for constructing a model for predicting the results of neoadjuvant therapy. Background Art

[0002] Neoadjuvant therapy for tumors refers to a treatment before the main treatment means (surgery), including chemotherapy, radiotherapy, endocrine therapy, targeted drugs, immune drugs, etc. The purpose of neoadjuvant therapy is to shrink the tumor, kill invisible metastatic cells, and improve the overall treatment effect. Pathological complete remission (pCR) means that after receiving neoadjuvant therapy, the tumor specimen is taken for pathological microscopic examination, and all malignant tumor cells disappear. Generally speaking, the higher the pCR, the better the effect of neoadjuvant therapy.

[0003] Literature reports that approximately 50% of breast cancer patients receiving neoadjuvant therapy do not achieve pathological complete remission, about 21.4% of patients will have local recurrence within 15 years, and 38.2% of patients will have distant recurrence. Previous models based on genomics, transcriptomics, radiomics, pathology, or even multi-omics were developed based on small sample sizes and single-treatment-strategy cohorts. Although the above models can achieve high prediction performance in internal validation, they cannot be well generalized to more complex real-world application scenarios. The reasons are as follows: (1) The data types used in existing models are single and the sample sizes are small, and they cannot comprehensively characterize the tumor state to improve the prediction performance of the models; (2) Since the types of anti-cancer drugs are relatively limited and their structures are diverse, it is difficult to understand the compound structure characterization and integrate the compound structure data with the transcriptome state. Therefore, directly characterizing compounds with compound structures in the past is not applicable; since many chemotherapy drugs lack specific targets and the mechanism of action of targeted therapy often goes beyond the main targets, drug characterization based on specific gene targets cannot fully represent all types of drugs or comprehensively characterize their mechanisms; the combined use of drugs in clinical practice will lead to complex synergistic or antagonistic interactions, and these interactions need to be simulated, but existing models cannot meet the above requirements to appropriately characterize drugs to accurately simulate the interaction between drugs and tumors, and the interaction between drugs and tumors is crucial for predicting the response of tumors to treatment. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for characterizing drugs based on whole-genome; constructs a neoadjuvant treatment outcome model based on a method for characterizing drugs using gene expression correlation, and provides a method, device, medium and program product for predicting the construction of the neoadjuvant treatment outcome model; in the application stage, provides a method, device, medium and program product for predicting the neoadjuvant treatment outcome, and biological deep learning combining drug characterization and biomics profile can be used as a digital organoid to optimize neoadjuvant treatment options for diseases such as breast cancer.

[0005] The first aspect of the present application discloses a method for characterizing drugs based on whole-genome, the method comprising:

[0006] 101. Obtain the whole-genome data and drugs of pan-cancer cell lines in the Cancer Drug Sensitivity Genomics database; the drugs include small molecule drugs;

[0007] 102. Calculate the IC50 value of the small molecule drug;

[0008] 103. Determine the gene expression significantly related to the small molecule drug in the whole-genome data and denote it as Ge;

[0009] 104. Calculate the correlation between the IC50 value and the significantly related gene and denote it as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize the small molecule drug.

[0010] In some embodiments, if the first small molecule drug is missing in the genomics database, the most relevant drug within the same anti-tumor drug type with a similar mechanism of action in the database is used to represent the first small molecule drug;

[0011] Optionally, if the genomics database includes at least 2, and the second small molecule drug exists in at least 2 databases at the same time, the database with a higher recognition rate is preferentially used to calculate and represent the second small molecule drug; for example, when the database includes GDSC and CTRP data sets, the GDSC data set is adopted;

[0012] Optionally, the small molecule drug includes any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2 MK-2206, ENG;

[0013] Optionally, the value range of the correlation is greater than or equal to -1 and less than or equal to 1;

[0014] Optionally, the Cancer Drug Sensitivity Genomics database includes GDSC and CTRP data sets;

[0015] Optionally, the drug does not include drugs without corresponding drug sensitivity data in the GDSC or CTRP databases, such as: T-DM1, endocrine therapy drugs.

[0016] In some embodiments, the small molecule drug can be replaced by an antibody drug; or, the drug further includes an antibody drug;

[0017] Optionally, the characterization method of the antibody drug includes: determining the target gene of the antibody drug; determining the genes significantly related to the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the expression Ge of the significantly related genes, denoted as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize the antibody drug.

[0018] Preferably, the negative value of the correlation between the expression level of the target gene and the expression Ge of the significantly related genes characterizes the antibody drug.

[0019] Optionally, the antibody drug includes: HER2 drugs.

[0020] The second aspect of the present application discloses a method for constructing a model for predicting the outcome of neoadjuvant therapy, the method comprising:

[0021] 201. Obtain the whole genome data and treatment plan of the training set samples that have received neoadjuvant therapy; the treatment plan is a plan composed of at least 1 drug; the drug is characterized by the gene-drug sensitivity correlation Gcd described in the first aspect of the present application.

[0022] 202. Input the whole genome data and treatment plan of the training set samples into a bioinformatics neural network model including an input layer, a gene layer, at least 1 biological process layer and an output layer for training to obtain the neoadjuvant therapy outcome model.

[0023] In some embodiments, when the treatment plan includes at least 2 drugs, calculate the sum of the drug sensitivity correlations Gcd of the same genes in the characterization methods of each drug to represent the treatment plan.

[0024] Optionally, the method further includes performing rank normalization and standardization processing on the whole genome data of each sample in 201.

[0025] Optionally, the way of inputting the whole genome data and treatment plan in 202 into the neural network model includes: parallel input of Ge and Gcd into the neural network model, or aggregating Ge and Gcd to obtain the attention value Gatt of gene expression when receiving treatment, and inputting the attention value into the neural network model.

[0026] Optionally, the calculation method of the attention value Gatt includes: Gatt = Gex * (Exp - Gcd);

[0027] Optionally, the types of inputs input by the input layer include whole-genome data and drugs;

[0028] Optionally, the biological process layer includes any one or more of the following: a fine pathway layer, a more complex biological pathway, and a biological process;

[0029] Optionally, the biological information neural network model includes: a GDnet framework;

[0030] Optionally, a prediction layer with sigmoid activation is added after each hidden layer;

[0031] Optionally, the loss weights of the results of the input layer, the gene layer, and the biological process layer increase step by step;

[0032] Optionally, the method further includes: calculating the sample-level importance scores of all nodes in all layers of the model;

[0033] Optionally, the training set samples are samples after removing duplicates and filtering out samples in a non-preprocessed state or missing key information (results and treatments).

[0034] The third aspect of the present application discloses a method for predicting the results of neoadjuvant therapy, and the method includes:

[0035] 301. Obtain the genomic data of the subject and the types of therapeutic drugs;

[0036] 302. Input the genomic data and the types of therapeutic drugs into the neoadjuvant therapy result model described in the second aspect of the present application, and output a reaction prediction value after receiving neoadjuvant therapy with the types of therapeutic drugs;

[0037] Optionally, the method further includes 303: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the reaction prediction value; when the reaction prediction value is greater than the first threshold, output an auxiliary prediction result that the subject has a good therapeutic effect using the types of therapeutic drugs; when the prediction score is less than the first threshold, output an auxiliary prediction result that the subject has a poor therapeutic effect using the types of therapeutic drugs.

[0038] The fourth aspect of the present application is a method for predicting the results of neoadjuvant therapy, and the method includes:

[0039] 401. Obtain the genomic data of the subject and the types of therapeutic drugs; the genes in the genomic data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1;

[0040] 402. Input the genomic data and types of therapeutic drugs into the neoadjuvant treatment outcome model described in the second aspect of this application, and output the response prediction value after neoadjuvant treatment with the types of therapeutic drugs.

[0041] Optionally, the types of therapeutic drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2 MK-2206, ENG.

[0042] Optionally, the method further includes 403: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the response prediction value; when the response prediction value is greater than the first threshold, output an auxiliary prediction result that the test subject has a good therapeutic effect using the type of therapeutic drug; when the prediction score is less than the first threshold, output an auxiliary prediction result that the test subject has a poor therapeutic effect using the type of therapeutic drug. The fifth aspect of this application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method described in the first aspect and / or the second aspect and / or the third aspect of this application.

[0043] The sixth aspect of this application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect and / or the second aspect and / or the third aspect of this application.

[0044] The seventh aspect of this application discloses a computer program product, including a computer program, which when executed by a processor, implements the steps of the method described in the first aspect and / or the second aspect and / or the third aspect of this application.

[0045] This application has the following beneficial effects:

[0046] 1. This application innovatively discloses a simple method of representing drugs using gene expression correlation, fully integrating drugs with transcriptome data, mapping drugs into a series of gene and value pairs, and using different methods to characterize different drugs (small molecule drugs and antibody drugs) respectively, more comprehensively characterizing tumor markers, and appropriately characterizing drugs to accurately simulate the interaction between drugs and tumors; 2. This application innovatively constructs a GDnet model based on the interaction between the transcriptome and treatment regimens. Compared with the transcriptome-only model, the GDnet model has achieved significantly higher performance in predicting the results of neoadjuvant treatment for breast cancer. The application of GDnet in two series of simulated clinical trials shows that GDnet can be used as a digital organoid to optimize the treatment options for breast cancer patients. Compared with previous models, GDnet has the following advantages:

[0047] (1) Most previous machine learning models have faced problems such as limited sample size and unstable performance due to the lack of sufficient independent external validation. This application collected 31 datasets, including 4371 breast cancer samples, which contain available pre-treatment transcriptomes and neoadjuvant treatment information, the largest dataset to date, to robustly train and validate GDnet;

[0048] (2) Data was collected from all types of RNA sequencing methods by applying a rank normalization strategy to each sample. This approach provides robustness against technical artifacts that may otherwise introduce systematic biases in absolute transcript counts while keeping the overall relative ranking of genes within each cell at a more stable level. Due to ease of normalization and reduction of batch effects, previous studies were limited to a limited number of RNA sequencing methods; however, these methods tend to reduce the available sample size, making it difficult to generalize to broader scenarios. Although rank-based encoding has limitations, including not fully utilizing the precise gene expression measurements provided in transcript counts, it makes the model more applicable to real-world conditions with various confounding variables. Additionally, different from previous studies that normalized between patients (e.g., z-score transformation) or even datasets (e.g., Combat algorithm) in different cohorts or testing methods, this method only includes normalization for each patient. This design choice stems from the understanding that normalization parameters trained across patients or datasets cannot be effectively applied to patients in another dataset with a different transcriptome testing protocol, resulting in overfitting and poor generalization. Therefore, our model can be directly generalized to new cases tested with any RNA sequencing technology under broader scenarios and complex real-world conditions;

[0049] (3)Previous models, such as the 21-gene recurrence score, MammaPrint, and Adjutorium, were mainly based on traditional statistical models or machine learning models. In contrast, GDnet is based on a deep learning framework and includes a larger number of genes to maximize model capacity. Given that achieving pCR means eliminating all tumor cells, it is important to maintain sufficient flexibility in the model to account for tumor heterogeneity and microenvironmental changes, which are challenges that traditional machine models cannot address. The deep learning framework, with its large number of parameters and flexible architecture, is expected to solve these challenges. A previous study showed that deep learning methods are capable of reflecting or predicting tumor heterogeneity and the tumor environment. Additionally, by integrating prior biological knowledge, GDnet significantly reduces the number of learning parameters, resulting in a more stable and better-performing model than dense models, as demonstrated in previous studies. The visualization of the importance and properties of multi-level genes and biological pathways in GDnet enables a multi-level view of model interpretation, which can guide researchers in formulating hypotheses about the potential biological processes involved in drug resistance and translating these findings into therapeutic opportunities. Specifically, genes identified by GDnet, such as SRC, CCND1, MCF2L, RPS6KA1, and PSMB7, may play important roles in breast cancer resistance.

[0050] (4)GDnet innovatively integrates breast cancer treatment regimens, capturing the interaction patterns between drugs and tumors, thereby improving model performance. Although a few neoadjuvant breast cancer treatment models consider specific drug types, such as chemotherapy and anti-HER2 treatment, these models often lack the flexibility required to simulate the interaction between tumors and drugs. Previous studies have explored drug representatives based on drug structures; however, these methods have limitations because they lack sufficient drug diversity for comprehensive training in neoadjuvant breast cancer tasks and cannot represent antibody drugs, or use targets to represent drugs, but target-representative drugs lack the flexibility to simulate the interaction between drugs and tumors. To address these drawbacks, a simple method using gene expression correlations to represent drugs was introduced, which helps incorporate external knowledge related to drug resistance. This method provides a drug representation structure similar to gene expression in tumors, facilitating fusion and creating the potential to simulate the interaction between drugs and the transcriptome. We also demonstrated that a better approach to integrating drug representation and the transcriptome involves feeding the data of these two modalities into the neural network in parallel, rather than simply fusing them through specially designed rules or aggregating the predictions of specific drug models;

[0051] In summary, GDnet is a deep learning method for bioinformatics that integrates drug representation and transcriptomics, enabling more precise oncology and enhancing the treatment decision-making process for neoadjuvant breast cancer treatment. Integrating drug representation and biomic data represents a new approach to simulating the interaction between drugs and tumors, providing a digital organoid platform for precision oncology that can be widely applied to different treatment settings and various cancer types. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0053] Figure 1 is a schematic flowchart of the method provided in the first aspect of the embodiment of the present invention;

[0054] Figure 2 is a schematic flowchart of the method provided in the second aspect of the embodiment of the present invention;

[0055] Figure 3 is a schematic flowchart of the method provided in the third aspect of the embodiment of the present invention;

[0056] Figure 4 is a schematic flowchart of the method provided in the fourth aspect of the embodiment of the present invention;

[0057] Figure 5 is a schematic diagram of the computer device provided in the embodiment of the present invention;

[0058] Figure 6 is a schematic diagram of the architecture of the exemplary computing device provided in the embodiment of the present invention;

[0059] Figure 7 is a schematic diagram of the storage medium provided in the embodiment of the present invention;

[0060] Figure 8 is a schematic diagram of the fusion of drug representation, transcriptomics, and protocol representation provided in the embodiment of the present invention. Among them Figure 8 a represents the strategy of representing drugs through gene correlation. Small molecules represent significant genome-wide correlations with the half-maximal inhibitory concentration (IC50) value, while antibody drugs represent significant genome-wide correlations with target genes. Figure 8b represents the GDnet framework, a biological deep learning model that integrates transcriptomics and regimen representation. Transcriptomic data is sorted and normalized, and the drug representations in the treatment regimens are summed gene-wise and input into the neural network. A binary mask containing information about the dependencies between genes and / or pathways is incorporated into the network to impose constraints on the nodes and edges; IC50: half-maximal inhibitory concentration; Gcd: gene-drug sensitivity correlation; Ge: gene expression; Figure 9 This is the heatmap visualization of the gene-drug correlation for each drug provided by the embodiments of the present invention, and the correlation values are represented by color coding. The rows and columns are clustered using consensus clustering, and the drugs are annotated by type. PDCD1 represents anti-PD-1 therapy (pembrolizumab and durvalumab), ANG represents ANG peptide inhibitor (AMG-386), VEGFA represents anti-VEGFA antibody (bevacizumab), IGF1R represents anti-IGF1R antibody (ganitumab), capecitabine is represented by 5-fluorouracil, Ganetespib is represented by luminespib, and ixabepilone is represented by epothilone B;

[0061] Figure 10 This is the comparison of the computational performance of GDnet with other models having an external validation set provided by the embodiments of the present invention. The reported performance metrics include the area under the ROC curve (AUC) ( Figure 10 a) and the area under the precision-recall curve (AUPRC) ( Figure 10 b). In terms of these two metrics, the median of GDnet is significantly better than other models. Figure 10 a - Figure 10 The data in a - b are represented as box plots, where the median line is the median, the lower hinge and the upper hinge correspond to the 1st and 3rd quartiles respectively, and the dashed lines correspond to the minimum or maximum values that are no more than 1.5×IQR away from the hinges (where IQR is the interquartile range). Any data beyond the dashed lines are considered outliers. For all comparisons between GDnet and other models and between GDNNet and GANNnet, p < 0.05. Figure 10 c represents the ROC curve of GDnet compared with other models (n = 912). Figure 10 The bar chart in d shows the AUC of integrating biological knowledge in each individual GEO dataset in the test set for GDnet and other models. The datasets are sorted by sample size, with the largest on the left and the smallest on the right. All models are the integrations of 10 repetitions in ( Figure 10 c) and ( Figure 10 d);

[0062] Figure 11 This is the inspection and interpretation of GDnet provided by the embodiments of the present invention. Figure 11a, Visualization of the inner layer of GDnet, showing the estimated relative importance of different nodes in each layer. Nodes in the leftmost layer represent input types, nodes in the second layer represent genes, the following layers represent higher-level biological entities, and the last layer represents the model results. Nodes with longer lengths are more important. The contribution of a certain source node to a target node, such as the importance of each input type (drug or transcriptome) to each gene, the contribution of each gene to each pathway, and the contribution of each pathway to each broader pathway, is described by a Sankey diagram. For example, the importance of the PSMC5 gene is mainly driven by the transcriptome. Figure 11 b, The proportion of regimens used in the optimization group (middle) and the control group (right), the proportion difference between the two groups (left), and the external validation set of simulation experiments based on different optimization thresholds. The optimization thresholds are color-coded. Opt: Optimization group; Ctl: Control group;

[0063] Figure 12 It is an overview of the model provided in the embodiments of the present invention for comparison with GDnet. Figure 12 a, The GAnet framework, which integrates drug representation and transcriptome data. First, a function is used to aggregate transcriptome data and regimen representation, where Gatt represents the attention value of gene expression when receiving treatment. Then the Gatt value is input into the Figure 8 biological information neural network described in Figure 12 b, The Gnetens framework. Different transcriptome-only models are constructed for each drug type, and the predictions of all transcriptome-only models corresponding to the drugs administered to all patients are integrated by averaging to generate the final prediction. Figure 12 c, The Gnet framework transcriptome-only model. The transcriptome data is directly input into the Figure 8 biological information neural network described in

[0064] Figure 13 It is a heatmap provided in the embodiments of the present invention showing the treatment drug combinations of all patients. Applied drugs are marked as 1 and color-coded red, and unapplied drugs are marked as 0 and color-coded blue. The ER status, HER2 status, and pCR results are annotated at the top, where white represents missing values;

[0065] Figure 14 It is a comparison of the computational performance of GDnet and other models using the training set for 10-fold cross-validation provided in the embodiments of the present invention. The metrics include the area under the ROC curve (AUC) (a), the area under the precision-recall curve (AUPRC) ( Figure 14 b). Figure 14 a and Figure 14The data in b is represented as a box plot, where the median line is the median, the lower and upper hinges correspond to the first and third quartiles respectively, and the whiskers correspond to the minimum or maximum value within 1.5×IQR from the hinges (where IQR is the interquartile range). Data points outside the whiskers are considered outliers. No statistical significance was observed in all comparison combinations; Figure 1 .5×IQR range from the hinges (where IQR is the interquartile range). Data points outside the whiskers are considered outliers. No statistical significance was observed in all comparison combinations;

[0066] Figure 15 is the comparison of the computational performance of GDnet provided by the embodiments of the present invention with other models having an external validation set. Figure 15 a, The bar chart shows the AUC of all models for each individual GEO dataset in the test set. Figure 15 b, The bar chart shows the AUPRC of all models for each individual GEO dataset in the test set. The datasets are sorted by sample size, ( Figure 15 a - Figure 15 b) is the largest on the left and the smallest on the right. Both models are ensembles of 10 repetitions. Detailed implementation manners

[0067] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0068] In some processes described in the specification and claims of the present invention and the above - mentioned accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0070] Figure 1 is a schematic flowchart of a method for characterizing drugs based on the whole - genome provided by the embodiments of the present invention. Specifically, the method includes the following steps:

[0071] 101: Obtain the whole-genome data and drugs of pan-cancer cell lines in the Cancer Drug Sensitivity Genomics Database; the drugs include small molecule drugs;

[0072] In some embodiments, if the first small molecule drug is missing in the genomics database, the most relevant drug within the same type of anti-tumor drug with a similar mechanism of action in the database is used to represent the first small molecule drug;

[0073] In some embodiments, if the genomics database includes at least two, and the second small molecule drug exists in at least two databases simultaneously, the database with a higher recognition rate is preferentially used to calculate and represent the second small molecule drug; for example, when the database includes GDSC and CTRP datasets, the GDSC dataset is adopted;

[0074] In some embodiments, the small molecule drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG;

[0075] Optionally, the Cancer Drug Sensitivity Genomics Database includes GDSC and CTRP datasets, and this dataset includes the transcriptome status of more than 800 cell lines and the drug sensitivity of more than 500 molecular drugs;

[0076] Optionally, the drugs do not include drugs without corresponding drug sensitivity data in the GDSC or CTRP databases, such as: T-DM1, endocrine therapy drugs.

[0077] 102: Calculate the IC50 value of the small molecule drug, where the IC50 value is used to represent drug sensitivity, where a high IC50 corresponds to low cell killing ability, and a low IC50 corresponds to high cell killing ability; the lower the IC50 value, the higher the sensitivity;

[0078] IC50 (halfmaximal inhibitory concentration) refers to the half-inhibitory concentration of the measured antagonist. It can indicate the half amount of a certain drug or substance (inhibitor) in inhibiting certain biological procedures (or certain substances included in this procedure, such as enzymes, cell receptors or microorganisms). In terms of apoptosis, it can be understood that a certain concentration of a certain drug induces apoptosis of 50% of tumor cells, and this concentration is called the 50% inhibitory concentration, that is, the concentration corresponding to when the ratio of apoptotic cells to the total number of cells is equal to 50%. The IC50 value can be used to measure the ability of a drug to induce apoptosis, that is, the stronger the induction ability, the lower this value. Of course, it can also inversely illustrate the tolerance degree of a certain cell to the drug.

[0079] 103. Determine the gene expression significantly related to the small molecule drug in the whole genome data based on the IC50 value and denote it as Ge. 104. Calculate the correlation between the IC50 value and the significantly related gene, denoted as the gene-drug sensitivity correlation Gcd. Gcd is used to characterize the small molecule drug.

[0080] In some embodiments, the value range of the correlation is greater than or equal to -1 and less than or equal to 1. The method for calculating the correlation is log2(IC50 + 1). As prior knowledge, it informs the importance of the gene's response to the drug, thereby reflecting the perturbation of the drug on the gene.

[0081] In some embodiments, the small molecule drug can be replaced by an antibody drug; or, the drug also includes an antibody drug.

[0082] Optionally, the characterization method of the antibody drug includes: determining the target gene of the antibody drug; determining the gene significantly related to the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the gene expression Ge of the significantly related gene, denoted as the gene-drug sensitivity correlation Gcd. Gcd is used to characterize the antibody drug.

[0083] Preferably, the negative value of the correlation between the expression level of the target gene and the gene expression Ge of the significantly related gene characterizes the antibody drug. The reason for adding a negative sign before the expression level of the antibody drug target gene is that the higher the expression level of the target gene, the more sensitive it is to the drug, and the lower the IC50. There is an inverse trend between the expression level of the target gene and the IC50.

[0084] Optionally, the antibody drug includes: HER2 drug.

[0085] Optionally, the calculation method of the correlation is a common statistical method, such as the Pearson correlation coefficient, etc.

[0086] Optionally, if the significant gene is not in the whole genome dataset of the training set samples, directly delete the gene.

[0087] As Figure 2 shown, the second aspect of the present application discloses a method for constructing a model for predicting the outcome of neoadjuvant therapy. The method includes:

[0088] 201. Obtain the whole genome data and treatment plan of the training set samples that have received neoadjuvant therapy. The treatment plan is a plan composed of at least one drug. The drug is characterized by the gene-drug sensitivity correlation Gcd described in the first aspect of the present application.

[0089] In some embodiments, when the treatment regimen includes at least two drugs, the sum of the drug sensitivity correlations Gcd of the same gene in each drug characterization method is calculated to represent the treatment regimen.

[0090] Optionally, the method further includes performing rank normalization and standardization processing on the whole genome data of each sample in 201. The reason for the standardization processing is that in datasets with different transcriptome test protocols, the normalization parameters trained across patients or datasets cannot be effectively applied to patients, resulting in overfitting and reduced generalization. Therefore, our model can be directly generalized to new cases tested by any RNA sequencing technology under a wider range of scenarios and complex real-world conditions.

[0091] 202. Input the whole genome data and treatment regimen of the training set samples into a bioinformatics neural network model including an input layer, a gene layer, at least one biological process layer, and an output layer for training to obtain the neoadjuvant treatment outcome model. The GDnet model is the neoadjuvant treatment outcome model. It belongs to a typical binary classification model. The neural network result undergoes a sigmoid transformation, and the output is the probability of a pCR efficacy. It can refer to logistic regression, which belongs to a binary classification model.

[0092] In some embodiments, the way of inputting the whole genome data and treatment regimen in 202 into the neural network model includes: parallel input of Ge and Gcd into the neural network model, or obtaining the attention value Gatt of gene expression at the time of receiving treatment after aggregating Ge and Gcd, and inputting the attention value into the neural network model; optionally, the calculation method of the attention value Gatt includes: Gatt = Gex * (Exp - Gcd); optionally, the input types of the input layer include whole genome data and drugs;

[0093] Optionally, the biological process layer includes any one or several of the following: a fine pathway layer, a more complex biological pathway, and a biological process;

[0094] Optionally, the bioinformatics neural network model includes: a GDnet framework; optionally, a prediction layer with sigmoid activation is added after each hidden layer; the input layer is a Gene + drug layer, gene, Layer1, Layer2, Layer3, Layer4, and Layer5 are all hidden layers, and finally the output of the neoadjuvant pCR probability is given.

[0095] Optionally, the loss weights of the results of the input layer, gene layer, and biological process layer increase gradually;

[0096] Optionally, the method further includes: calculating the sample-level importance scores of all nodes in all layers of the model;

[0097] Optionally, the training set samples are samples after removing duplicates and filtering out samples in a non-preprocessed state or lacking key information (results and treatments).

[0098] As Figure 2 shown, a method for predicting the outcome of neoadjuvant therapy is disclosed in the third aspect of the present application. The method includes:

[0099] 301. Obtain the genomic data of the subject and the types of therapeutic drugs;

[0100] 302. Input the genomic data and the types of therapeutic drugs into the neoadjuvant therapy outcome model described in the second aspect of the present application, and output a response prediction value after neoadjuvant therapy with the types of therapeutic drugs;

[0101] Optionally, the method further includes 303: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the response prediction value; when the response prediction value is greater than the first threshold, output an auxiliary prediction result indicating that the subject has a good therapeutic effect with the types of therapeutic drugs; when the prediction score is less than the first threshold, output an auxiliary prediction result indicating that the subject has a poor therapeutic effect with the types of therapeutic drugs.

[0102] As Figure 4 shown, a method for predicting the outcome of neoadjuvant therapy is provided in the fourth aspect of the present application. The method includes:

[0103] 401. Obtain the genomic data of the subject and the types of therapeutic drugs; the genes in the genomic data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1;

[0104] 402. Input the genomic data and the types of therapeutic drugs into the neoadjuvant therapy outcome model described in the second aspect of the present application, and output a response prediction value after neoadjuvant therapy with the types of therapeutic drugs;

[0105] Optionally, the types of therapeutic drugs include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2 MK-2206, ENG;

[0106] Optionally, the method further includes 403: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the response prediction value; when the response prediction value is greater than the first threshold, output an auxiliary prediction result indicating that the subject has a good therapeutic effect with the types of therapeutic drugs; when the prediction score is less than the first threshold, output an auxiliary prediction result indicating that the subject has a poor therapeutic effect with the types of therapeutic drugs.

[0107] In some embodiments, the terms "subject", "test subject", "sample to be tested", or "sample to be measured" as used herein refer to any animal (e.g., a mammal), including but not limited to humans, non-human primates, rodents, etc., that will be the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human. In some embodiments, the sample to be tested is a patient clinically used for prognostic evaluation.

[0108] In some embodiments, the auxiliary prediction results include, but are not limited to, paper or electronic report forms. These results are only obtained by the intelligent machine through analysis of the relevant data of the subject and are only for reference by medical staff, not as the final diagnosis result of the subject.

[0109] In some embodiments, the threshold is obtained by training with a training set of samples. It can be a specific threshold or an interval range, and the specific form is not specifically limited in this embodiment.

[0110] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, as Figure 5 shown, the device may include: one or more processors, and one or more memories; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the above-described method can be executed.

[0111] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0112] Generally speaking, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0113] For example, the method or apparatus according to an embodiment of the present disclosure can also be implemented by means of Figure 6 the architecture of the computing device 3000 shown. As Figure 6 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, can store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 the architecture shown is only exemplary, and when implementing different devices, one or more components in the Figure 6 shown computing device may be omitted according to actual needs.

[0114] An embodiment of the present invention also provides a computer-readable storage medium. As Figure 7 shown, it is a schematic diagram of the storage medium provided by an embodiment of the present invention. A computer-readable instruction 4010 is stored on the computer storage medium 4020. When the computer-readable instruction 4010 runs on a processor, it can execute the method according to an embodiment of the present disclosure described with reference to the above drawings. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0115] Embodiments of the present disclosure also provide a computer program product or system, including a computer program which, when executed by a processor, implements the steps of the above method.

[0116] In some embodiments, this embodiment also discloses a system for characterizing drugs based on whole-genome, the system includes:

[0117] A first acquisition module 501, configured to acquire the whole-genome data and drugs of pan-cancer cell lines in a cancer drug sensitivity genomics database; the drugs include small molecule drugs;

[0118] An IC50 value calculation module 502, configured to calculate the IC50 value of the small molecule drug;

[0119] A significant gene determination module 503, configured to determine the gene expression significantly related to the small molecule drug in the whole-genome data based on the IC50 value, denoted as Ge;

[0120] A small molecule drug characterization module 504, configured to calculate the correlation between the IC50 value and the significantly related genes, denoted as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize the small molecule drug.

[0121] In some embodiments, this embodiment also discloses a system for constructing a neoadjuvant treatment outcome prediction model, the system includes:

[0122] A second acquisition module 601, configured to acquire the whole-genome data and treatment regimens of a training set sample that has received neoadjuvant treatment; the treatment regimen is a regimen composed of at least one drug; the drug is characterized by the gene-drug sensitivity correlation Gcd described in the first aspect of the present application;

[0123] A model training module 602, configured to input the whole-genome data and treatment regimens of the training set sample into a bioinformatics neural network model including an input layer, a gene layer, at least one biological process layer, and an output layer for training, to obtain the neoadjuvant treatment outcome prediction model.

[0124] In some embodiments, this embodiment also discloses a system for predicting the outcome of neoadjuvant treatment, the system includes:

[0125] A third acquisition module 701, configured to acquire the genomic data of the subject and the types of treatment drugs;

[0126] A first response prediction value calculation module 702, configured to input the genomic data and the types of treatment drugs into the neoadjuvant treatment outcome prediction model described in the second aspect of the present application, and output a response prediction value after receiving neoadjuvant treatment with the types of treatment drugs;

[0127] In some embodiments, the system further includes a first prediction result output module 703: which is used for or configured to obtain an auxiliary prediction result of the treatment effect of the drug type based on the reaction prediction value; when the reaction prediction value is greater than the first threshold, output an auxiliary prediction result that the treatment effect of the test subject using the treatment drug type is good; when the prediction score is less than the first threshold, output an auxiliary prediction result that the treatment effect of the test subject using the treatment drug type is poor.

[0128] In some embodiments, this embodiment also discloses a system for constructing a neoadjuvant treatment result prediction model, the system includes:

[0129] A fourth acquisition module 801, which is used for or configured to acquire the genomic data and treatment drug type of the test subject; the genes in the genomic data include any one or several of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1;

[0130] A second reaction prediction value calculation module 802, which is used for or configured to input the genomic data and treatment drug type into the neoadjuvant treatment result model described in the second aspect of the present application, and output a reaction prediction value after neoadjuvant treatment with the treatment drug type;

[0131] In some embodiments, the system further includes a second prediction result output module 803: which is used for or configured to obtain an auxiliary prediction result of the treatment effect of the drug type based on the reaction prediction value; when the reaction prediction value is greater than the first threshold, output an auxiliary prediction result that the treatment effect of the test subject using the treatment drug type is good; when the prediction score is less than the first threshold, output an auxiliary prediction result that the treatment effect of the test subject using the treatment drug type is poor. Specific embodiments:

[0133] Results: Schematic diagram of the GDnet workflow: To fully integrate drugs with transcriptome data, mapping drugs to a series of gene and value pairs is one of the most appropriate methods. The selection of gene and value pairs is a key consideration. The GDSC and CTRP datasets contain the transcriptome status of more than 800 cell lines and their drug sensitivities to more than 500 molecular drugs, and are excellent reference datasets. The heterogeneous responses of different cancer cell lines to drugs in GDSC or CTRP can be attributed to the heterogeneity of the tumor transcriptome status; therefore, genes significantly correlated with drug sensitivity are more likely to be genes related to response or drug resistance, and thus more likely to be genes disrupted when the tumor is exposed to a specific drug. Therefore, we use the genome-wide expression correlation of the half-maximal inhibitory concentration (IC50) value of each drug in different cell lines (where a lower value indicates higher sensitivity) to represent each drug ( Figure 8a). It should be noted that antibody or peptide drugs (such as anti-HER2 drugs) are not included in these drug sensitivity datasets, and IC50 values are not traditionally used to measure drug sensitivity through complex mechanisms. Given the strong positive correlation between the sensitivity of these targeted drugs and the expression of the target genes, we used the negative value of the correlation between the genome-wide expression and the expression of the drug target as a proxy for antibody drugs ( Figure 8 a). For treatment regimens with multiple drugs, the genes representing each drug in the treatment regimen are summed up ( Figure 8 b).

[0134] To better simulate the interaction between drugs and the transcriptome, we input the rank-normalized transcriptome and regimen representation in parallel into the gene-level layer of the subsequent neural network, which we call GDnet ( Figure 8 b). We further explored three additional model frameworks: (1) GAnet, which integrates drugs and the transcriptome using a method that mimics the key-value attention mechanism ( Figure 12 a); (2) Gnetens, which uses an ensemble learning method to integrate all drugs predicted for a patient ( Figure 12 b); (3) Gnet, which relies only on transcriptome information ( Figure 12 c). Next, the gene layer is advanced to the bioinformatics neural network proposed by H.A. et al., which incorporates the subordinate (child-parent) relationships, fine-grained pathways, and more complex pathways between genes based on the Reactome pathway dataset (version 2022) ( Figure 8 b). The finally constructed prediction model aims to accurately predict the response of breast cancer patients to neoadjuvant therapy and assist in the selection of multiple candidate regimens.

[0135] Drug representation based on the GDSC and CTRP datasets: To generate gene and value pairs for drug representation, we calculated the correlation between the IC50 values of various drugs in different cell lines and gene expression, as well as the correlation with the expression level of antibody or polypeptide drug targets. This calculation was performed for all combinations of drugs and genes. For drugs that appear in both datasets, the GDSC correlation was given priority. Most drugs administered to patients have their corresponding representations in the GDSC or CTRP datasets. However, some small molecule drugs are missing in both datasets, and we replaced them with the most relevant drugs in the GDSC / CTRP datasets that have similar mechanisms and the same type of anti-tumor drugs. The representations of all drugs used in the existing breast cancer neoadjuvant therapy datasets were summarized as ( Figure 9)。Methotrexate represents the largest number of gene-value pairs, with a total of 5378 pairs. Among the 22 drug characterizations, 17 gene-value pairs are greater than 1000, and 5 (ANG, Lapatinib, Carboplatin, Epothilone B, Doxorubicin) gene-value pairs are less than 1000. The consensus clustering of drugs groups these 5 drugs together with fewer paired numbers ( Figure 9 )。For drugs with a larger number of pairs, three main clusters are generated. The first cluster includes anti-angiogenic therapies, anti-HER2 therapies, and anti-IGF1R therapies. The second cluster includes chemotherapy therapies, AKT inhibitors, and PARP inhibitors. The last cluster includes anti-PD-1 therapies and HSP90 inhibitors ( Figure 9 )。These drug characterizations are then combined with the patient transcriptome to predict neoadjuvant treatment response.

[0136] GDnet demonstrated superiority in predicting neoadjuvant treatment outcomes: We collected 9550 breast cancer samples from 68 independent public datasets. After removing duplicates and filtering out samples in non-pretreated states or lacking key information (outcomes and treatments), 31 datasets containing 4371 samples and 16133 genes were filtered out. Seventeen datasets with the largest number of cases (a total of 3459 samples) were used to train the model. Ten-fold cross-validation was employed for hyperparameter tuning and internal validation, with a final learning rate of 0.005. Due to the sparse nature of the model, L2 regularization was not applied (set to 0). The remaining 14 datasets with the fewest cases, a total of 912 samples, were used as an external validation set. After crossing with genes involved in the biological network from the Reactome dataset, a total of 8377 genes were finally retained.

[0137] For benchmark comparison, we also trained dense deep neural network (DNN) variants of GDnet, GAnet, Gnetens, and Gnet (without biological knowledge), named GDNNet, GANNet, GNNetens, and GNNet respectively. In Gnetens and GNNetens, only drugs that appeared in more than 10% of the samples were used to construct separate models. Five drug types (anthracyclines, microtubule inhibitors, cyclophosphamide, pyrimidine analogs, and anti-HER2 antibodies) were selected ( Figure 13 )。Then, ten-fold cross-validation was applied using the training set to compare the model performance. The results showed that the area under the curve (AUC) and the area under the precision-recall curve (AUPRC) of GDnet (median AUC = 0.808, median AUPRC = 0.739) were higher than those of all other models (taking Gnet as an example, median AUC = 0.799, median AUPRC = 0.731). In the interval validation, although not significant ( Figure 14)。These models were trained 10 times using the entire training set and were separately tested using an external validation set, which facilitated a robust comparison between different models. In the external validation, GDnet performed significantly better than all other models (median AUC = 0.712, median AUPRC = 0.691) (taking Gnet as an example, median AUC = 0.676; median AUPRC = 0.632)( Figure 10 a-b). Notably, the performance of GDNNet, which has the largest number of parameters (AUC = 0.678, AUPRC = 0.661), was lower than that of GANNet (AUC = 0.695, AUPRC = 0.668), which demonstrated the importance of prior knowledge in guiding training and improving performance. To robustly compare the ROC curves between models, we created an ensemble by averaging the predictions of the models trained 10 times for the final prediction. Similarly, GDnet (AUC = 0.725) obtained the highest AUC( Figure 10 c), and its ROC curve was significantly different from that of Gnet (AUC = 0.683) (Delong's test, p = 0.037). Then, we tested the performance of GDnet in 13 separate validation datasets using the ensemble model with 10 repetitions to determine whether GDnet could perform well in limited scenario settings compared with Gnet. The GSE191127 dataset was not included in the individual tests because it only contained patients with residual disease (RD). The results showed that GDnet achieved a higher AUC than Gnet in 8 out of 13 datasets( Figure 10 d). Although GDnet showed a lower AUC in the 5 datasets on the left, all differences were less than 0.03. In addition, GDnet achieved a higher AUC than GAnet and Gnetens in 7 out of 13 datasets( Figure 10 d). GDnet also showed a higher AUC than all dense models in each dataset( Figure 15 a). Similar results were observed for the AUPRC values( Figure 15 b). These results indicated that GDnet successfully captured the interaction pattern between drugs and transcriptomes, rather than simply overfitting the drug distribution bias between datasets.

[0138] GDnet can optimize the treatment plan selection to obtain better results: To evaluate the ability of GDnet to guide the clinical treatment plan selection, we then conducted a computer-simulated clinical trial to examine whether GDnet could improve the results of neoadjuvant treatment for breast cancer. Given multiple possible treatment plans and patient transcriptomes, we input them into the GDnet model to calculate the probability of pathological complete remission (pCR) for each case. We selected the top N% of the best plans with the highest predicted pCR probability, which might benefit the patients, and the bottom N% of the worst plans with the lowest predicted pCR probability, which were not suitable for the patients. If the plan actually used by the patient was rated as one of the best plans, we assigned the patient to the optimization group; if it was rated as one of the worst plans, we assigned the patient to the control group. To examine the baseline characteristics of the two groups, we also calculated the predicted scores through the transcriptome-only model to evaluate the malignancy of the tumors between them, giving priority to Gnet because it was similar in structure to GDnet and more robust in performance than GNNet (AUCLevene test, p = 0.026; AUPRC Levene test, p = 0.017). Figure 10 a-b).

[0139] We first conducted a series of computer trials based on the ISPY-2 trial (one of the datasets used to train our model). A total of 653 patients who received one of the 11 plans in ISPY-2 participated in our simulation study, and these plans were set as optional choices. We set different optimization thresholds, starting from 50% and decreasing in 5% increments (e.g., 50%, 45%, 40%, 35%) until the size of either group dropped below 30, classifying the best and worst plans to ensure the robustness of the results. The lower the threshold, the higher the optimization strictness. The results showed that at all optimization thresholds, the pCR rate in the optimization group was significantly higher than that in the control group. In addition, as the optimization was strengthened (the optimization threshold decreased), the odds ratio (OR) of pCR increased linearly (R2 = 0.72, p = 0.008). However, a higher Gnet predicted score was observed in the optimization group, indicating a lower malignancy of the tumors in this group. This difference might affect the effectiveness of our model. Therefore, we performed propensity score matching (PSM) to balance the transcriptome profiles according to the Gnet predicted scores. Similarly, the pCR rate in the optimization group was significantly higher than that in the control group, and as the optimization was strengthened, the OR increased linearly (R2 = 0.87, p < 0.001). The inverse probability of treatment weights (IPTW) analysis also demonstrated this (R2 = 0.46, p = 0.043), indicating that GDnet has the ability to optimize the treatment plan selection.

[0140] We further conducted a series of simulation experiments using an external validation dataset. Given the limited variety of treatment regimens in each validation dataset, we combined all available validation aversions to generate 912 patients and 18 treatment regimens. Similar to what was observed in the above ISPY-2 simulations, the results of the external validation set showed that across all optimization thresholds, the pCR rate in the optimization group was significantly higher than that in the control group, where the OR linearly increased with the intensification of optimization (R2 = 0.68, p = 0.003), and even a lower Gnet prediction score was observed in the optimization group. The PSM analysis based on the Gnet prediction score also showed that the pCR rate in the optimization group was significantly higher than that in the control group, with a linear increase in OR (by PSM, R2 = 0.94, p < 0.001; by IPTW, R2 = 0.46, p = 0.03) with the intensification of optimization, further demonstrating the excellent ability of GDnet to guide the selection of neoadjuvant breast cancer treatment.

[0141] Interpretation of GDnet revealed important genes and drugs involved in the response: To understand the importance of the features contributing to the model and their interactions, we visualized the layers of GDnet ( Figure 11 a). The integrated gradient attribution method was applied to obtain the importance scores of each node in each layer for each patient in the external validation set. The average importance score of all patients was used as the final score for each node, which was shown as the node length in the Sankey diagram. From the 10 repeated GDnet models, we selected the GDnet model with an AUC value (0.727) similar to that of the integrated GDnet with 10 repeats for interpretation. In terms of transcriptome and treatment regimens, the transcriptome contains more information than drug information; however, the drug treatment regimen provides additional information for more precisely predicting the response ( Figure 11 a). The regimens applied in neoadjuvant breast cancer treatment lack diversity (homogeneity) in the dataset, which may result in a relatively low contribution of regimen information in GDnet. High-score genes include SRC, CCND1, MCF2L, RPS6KA1, and PSMB7. Pathways at various levels, such as the ER-Phagosome pathway in the first layer, which is reported to be related to the innate immune system and phagocytosis, the RHOGTPasc cycle in the second layer, which is reported to be related to tumorigenesis, invasion, and metastasis, and the immune system in the fifth layer have been identified, highlighting their key role in the performance of GDnet. More detailed mechanisms of neoadjuvant breast cancer treatment require further experiments.

[0142] Next, we investigated which treatment types contributed to the response. For the treatment optimization strategy, we further examined the distribution of the regimens and drugs used in each group in the PSM analysis at different optimization thresholds. The differences in the proportions of regimens and drugs in different groups can inform the regimens and drugs that contribute more to the model. The results showed that the frequencies of paclitaxel | doxorubicin | cyclophosphamide | olaparib | PDCD1 and paclitaxel | doxorubicin | cyclophosphamide | ERBB2 | ERBB2 regimens in the optimization group were much higher than those in the control group and increased with the increasing degree of optimization ( Figure 11 b). Docetaxel | cyclophosphamide, doxorubicin, and 5-fluorouracil | epirubicin | cyclophosphamide | lapatinib were regimens that occurred mostly or only in the control group and tended to increase with the intensification of optimization ( Figure 11b). Specifically, drugs such as PDCD1 (PD-1 inhibitor), olaparib, and ERBB2 (anti-HER2 antibody) may play important roles in increasing the pCR rate. We further examined the regimens and drug distributions in the ISPY-2 simulation trial. Similar to the results of external validation, paclitaxel | doxorubicin | cyclophosphamide | ERBB2 | MK-2206, paclitaxel | doxorubicin | cyclophosphamide | PDCD1, and paclitaxel | doxorubicin | cyclophosphamide | ANG | ERBB2 were important regimens in the optimization group, and among them, PDCD1, ERBB2, and MK-2206 played important roles in increasing the pCR rate. By integrating drug data into the model, GDnet was able to compare all optional treatment regimens before the start of treatment, thus promoting precision medicine specifically for breast cancer patients, which could not be correctly achieved by all previous machine learning models. In other words, we established a model considering the function of organoids that could test the sensitivity of drugs and treatment regimens and inform patients about suitable drugs. The results of the in-silico trial based on the ISPY-2 trial and external validation set in our study showed that GDnet could be used to optimize the treatment decision-making process for breast cancer patients and significantly increase the pCR rate of breast cancer patients. For each optimization threshold, GDnet maintained the ability to increase the pCR rate, thus confirming the research results. In addition, there was a linear correlation between the pCR rate and the optimization intensity, strongly supporting our conclusion that GDnet could optimize neoadjuvant treatment for breast cancer and potentially change the relevant treatment guidelines. Our model could play important roles in many potential scenarios. In daily clinical practice, our model could help select the most suitable regimen from candidate regimens. In clinical trials, our model could help select patients who might respond to new test regimens or even previously abandoned regimens. Even if the treatment regimen included new drugs not included in the GDnet training set, our model was still able to predict drug responses because it performed well on a validation set containing many drugs, such as durvalumab, bevacizumab, cisplatin, olaparib, and gemcitabine, which were not present in the training set.

[0143] Method: Public data set processing: Transcriptome data related to neoadjuvant treatment of breast cancer were retrieved from the PubMed and Gene Expression Omnibus (GEO) database systems on February 15, 2024. The search terms used were ("Neoadjuvant Therapy" [MeSH Terms] OR neoadjuvant) AND ("Breast Neoplasms" [MeSH Terms] OR ((breast or mammary gland) AND (cancer* OR tumor* OR neoplasm* OR carcinoma* OR tumour* or oncology or malignancy*))). High-throughput sequencing such as next-generation sequencing (NGS) and microarray was included, and only data sets containing no less than 20 samples were selected. A total of 9550 breast cancer samples from 68 independent public data sets were accessed. Gene annotation of microarray probe sets was performed with reference to the reference table provided in GEO. For quality control of genes and samples, 16,133 genes that appeared in more than 50% of the samples were retained, and 4580 samples with less than 50% missing genes were retained. Among these data sets, data from 31 data sets (GSE194040, GSE25066, GSE16716, GSE41998, GSE180962, GSE20271, GSE34138, GSE50948, GSE149322, GSE22226, GSE32603, GSE22358, GSE32646, GSE16446, GSE130788, GSE231629, GSE123845, GSE4779, GSE173839, GSE22093, GSE21997, GSE42822, GSE66399, GSE181574, GSE23988, GSE41656, GSE8465, GSE122630, GSE21974, GSE207248, GSE191127), covering complete pCR, treatment regimens, and sufficient transcriptome information, were used for model construction and validation.

[0144] Ranking value encoding of transcriptome: RNA-seq counts or patient-level normalized data (e.g., transcripts per million kilobases (TPM) or fragments per million kilobases (FPKM)), microarray patient-level normalized data (e.g., robust multi-chip average (RMA) or other methods) were collected from GEO. If multi-level data were accessed, larger upstream levels could be maintained. To facilitate the conversion between gene expression testing technology platforms, we applied a ranking transformation (ranking value encoding) to normalize the data. To this end, gene expression values were converted from microarray intensities or RNAseq counts or their normalized formats to their respective ranks. This transformation was performed gene-wise. Thus, all gene expression values for each gene were ranked according to the order from lowest to highest. Then the ranks were scaled to quantiles (ranging from 0 to 1). Missing values were filled with 0. This rank-based method may be more robust to technical artifacts that may systematically bias absolute transcript count values, while the overall relative rank of genes within each cell remains more stable.

[0145] Therapeutic drug and protocol representation: To facilitate the integration of drug therapy and transcriptome, our goal was to represent drugs by combining relevant genes and relevant values. To construct a large enough pharmacogenomic dataset for drug characterization, raw drug sensitivity data of cell lines to molecular drugs, as well as the RMA-normalized / TPM-normalized and log-transformed transcriptomes of cell lines from Genomics of Drug Sensitivity in Cancer database v2 (GDSC) and Cancer Therapeutics Response Portal v2 (CTRP) were retrieved. We used the IC50 metric to represent drug sensitivity, where a high IC50 value corresponded to low cell killing ability and a low IC50 value corresponded to high cell killing ability. For each drug, we calculated the correlation between the gene expression levels of all cell lines and log2(IC50 + 1) of the response of cell lines to each drug. Considering that nearly 5*10 4 genes were to be tested, the significance cutoff of the correlation was set to 1*10 -6 according to the Bonferroni correction to reduce the false discovery rate. Genes identified as significant were more likely to play an important role in predicting drug responses in cancer treatment, and the relevant values were prior knowledge that could inform the importance of genes for drug responses, thus reflecting the perturbation of drugs on genes. If the relevant genes were not present in the 16,133 gene set of the retained patient transcriptome, they were removed. Thus, each drug was ultimately represented by its significant genes and relevant values (ranging from -1 to 1). If a drug appeared in both the GDSC and CTRP datasets, the GDSC dataset was considered first.

[0146] There are several special cases in drug agency. Certain types of drugs were not found in the GSDC or CTRP datasets. Therefore, we selected the most relevant drugs within the same type of anti-tumor drugs with similar mechanisms of action to represent it. Examples include capecitabine, which is represented by 5-fluorouracil; ganetespib, which is represented by luminespib; some drug descriptions in the GEO dataset, such as taxanes or anthracyclines, lack sufficient drug details; therefore, we selected the average representatives of taxanes or anthracyclines. Certain types of drugs are antibody- or peptide-based drugs, which are not included in these two datasets and are usually not measured by IC50 to represent drug sensitivity with complex mechanisms. For these drugs, we used the negative correlation between their target genes and other cell line genes in the cell because the response of these drugs is closely related to target expression, including PDCD1 therapies (pembrolizumab and durvalumab), VEGFA (bevacizumab), and IGF1R (ganitumab), which is an anti-IGF1R antibody. If a drug has multiple targets, we perform a gene sum of the correlations between all targets as the final representation, including ANGPT, which contains ANGPT1 and ANGPT2 genes and represents an ANGPT peptide inhibitor (AMG-386). In addition, drugs such as T-DM1 (an antibody-drug conjugate) and endocrine therapy do not have corresponding drug sensitivity data in GDSC or CTRP. Considering the small number of patients receiving these drugs in our dataset, patients receiving regimens such as T-DM1 or endocrine therapy were excluded. For patients receiving multiple drugs in one regimen, we perform a gene sum of the gene-related representations of each drug to represent the regimen.

[0147] Fusion of drug characterization and transcriptome: Three methods of integrating treatment regimens and transcriptome information were tested. The first method involved using gene expression (denoted as Ge) and gene sensitivity correlation (denoted as Gcd) as inputs, which we referred to as GDnet( Figure 8 b). The second method used Gatt=(Ge)(e Gcd ) as the final input, using a method similar to the key-value attention mechanism( Figure 12 a), and named the model GAnet. In this mechanism, for each query (Ge), we find the corresponding key (gene) in the drug representation, and the attention (Gatt) is calculated as attention=(query)(e value ), where the value represents the Gcd of the drug representation. The attention calculation is set as e value, to ensure that Gatt can simulate the tumor's response under drug exposure conditions; in other words, when there is no drug, the expression of genes expected to respond (Gcd < 0) or be resistant (Gcd > 0) to the drug is expected to decrease or increase. For example, ERBB2 is a gene with poor prognosis, but when exposed to an anti-HER2 drug with ERBB2 as the response gene (Gcd < 0), the expression of ERBB2 is expected to decrease. In addition, we also explored a third method, which involves using an ensemble learning approach. The focus of this method is to construct different transcriptome-only models for each drug and combine the final predictions by averaging the predictions of all transcriptome-only models corresponding to the drugs applied to the patients. This model is called Gnetens( Figure 12 b). A pure transcriptome model called Gnet was also explored( Figure 12 c).

[0148] Design of the bioinformatics model: We introduced prior knowledge into the training process, drawing inspiration from the P-NET framework constructed by Elmarakeby HA et al. to accelerate model training and improve performance and interpretability. The complete Reactome dataset was downloaded and processed into a hierarchical network. The network includes one layer of features, one layer of genes, and five layers of pathways. The membership relationships between adjacent layer nodes (representing genes or pathways) are represented by six binary matrices. The constructed hierarchical network is derived from genes (the first layer) that are available in both our dataset and the Reactome dataset. The pathways directly related to these genes form the second layer. Subsequently, the pathways directly related to the pathways in the second layer constitute the third layer, and so on, until the sixth layer with 18 pathways is constructed.

[0149] The deep learning model is implemented in PyTorch. The number of layers and nodes in the prior knowledge network defines the structure of the basic feed-forward dense neural network model. Six binary matrices containing the dependencies between genes or pathways in adjacent layers are input into the neural network model as masks between adjacent layers. During weight updates, after each epoch, this mask is multiplied by the weights, imposing constraints on the nodes and edges. This configuration results in a sparse neural network with six layers, including 8377 nodes (gene layer), 846 (layer 1 pathways), 218 (layer 2 pathways), 108 (layer 3 pathways), 51 (layer 4 pathways), and 18 nodes (layer 5 pathways) in each layer. Specifically, for GDnet and GDNNet, an additional gene and drug layer with 19300 nodes (half gene expression and half drug representation) is added before the 8377-gene layer. Since the model is sparse enough, the L2 regularization is set to 0. Each node encodes a biological entity (e.g., a gene or a pathway), and each edge represents a known relationship between the corresponding entities. Node constraints can better understand the states of different biological components. Compared with a fully connected network with the same number of nodes, edge constraints result in fewer parameters and thus less computational effort. Due to the dataset imbalance, we weighted the classes differently to reduce the network's bias towards a certain class according to the bias in the training set. The model was trained for 1000 epochs using the default parameters of the Adam optimizer to minimize the imbalanced binary cross-entropy loss function. To make each layer useful on its own, we added a prediction layer with sigmoid activation after each hidden layer. Since it is more challenging to fit the data with fewer weights in the later layers, we used higher loss weights for the results of the later layers (weights of 1, 1.2, 1.4, 1.6, 1.8, 2, and 2.5 from the gene and drug layer to the later layers) during the optimization process. The final prediction of the network is calculated by averaging the results of all layers.

[0150] Model Training and Validation: The analysis included a dataset that met specific criteria, providing a total of 4,371 available patients. The 17 largest datasets, including approximately 3,459 patients (79% of the total), were designated as the training set. The remaining 14 datasets, approximately 912 (21% of the total), were used as an external validation set, assigned without subjective allocation. The model was trained using the training set, and hyperparameters such as the learning rate were optimized through ten-fold validation. The trained model was finally validated using an external test set. The prediction performance of different models was compared using the validation set by calculating the median AUC and AUPRC of 10 rounds of replicated training using roc_auc_score and average_precision_score implemented separately in the Python sklearn.metrics library. Then, we averaged the 10 replicates of the model to obtain the final prediction of stability and robustness. The implementation of the proposed system and the reproducible results are available upon request and will be publicly available on GitHub after publication acceptance.

[0151] Model Interpretation: We applied IntegratedGradients, a gradient-based attribution method implemented in the tcaptum.attr library, to calculate the sample-level importance of all nodes in all layers. The absolute importance scores (always positive) of the nodes in each layer were normalized using the following formula to facilitate comparability between layers:

[0152]

[0153] where Ij is the normalized absolute importance score of each node in the layer, Aj is the absolute importance score, and n is the total number of nodes in the layer.

[0154] To calculate the total node-level importance, we aggregated the sample-level importance scores of all samples in the validation set (dataset-level interpretation is represented as the average of all patient-level weights in the validation set). The importance score of each node is represented as the node length in the Sankey diagram ( Figure 11 a). The importance score of the link between the source node (previous layer) and the target node (next layer) is represented by multiplying the absolute weight normalized within the absolute weights of all source nodes pointing to the same target node) of the source nodes pointing to the target node through the importance score of the target node in the model, as shown in the following formula:

[0155]

[0156] Ls-t jis the importance score of the link from the j-th source node(s) to the target node (t), is the normalized absolute importance score of the target node, Wsj is the weight of the j-th source node pointing to the target node in the deep learning model, and n is the number of source nodes pointing to the target node.

[0157] Statistical analysis: The change in the area under the ROC curve between GDnet and other models was tested by the DeLong test. The AUCs obtained from ten-fold cross-validation and ten-fold repeated training and testing of GDnet and other models were compared by the Wilcoxon rank test. The variances of AUC and AUPRC between Gnet and GNNet were compared by the Levene test to evaluate the robustness of performance. The odds ratio (OR) was calculated to determine the preference for pathologic complete remission (pCR) in the simulation trials, and the differences were compared by Fisher's exact test. A linear regression model was fitted to examine the relationship between the degree of optimization and the OR. The significance of the regression coefficients was evaluated by the t-test, and the R-squared value was used to quantify the proportion of variance explained by the model. To address the potential confounding caused by the imbalance in the transcriptome profiles (inferred from the Gnet prediction scores) between the optimization group and the control group, we employed propensity score matching (PSM) and inverse probability of treatment weighting (IPTW). PSM was performed by estimating the propensity scores using logistic regression, including the relevant covariates of the Gnet scores. Matching was performed using nearest neighbor matching (1:1 ratio) and a caliper of 0.1 without replacement to minimize the imbalance. Additionally, the application of IPTW was to weight each individual by the inverse probability of receiving the observed assignment, thus creating a balanced pseudo-population. Stable weights were used to maintain accurate estimates and ensure robust results. All tests were two-sided. Unless otherwise stated, a P-value < 0.05 was considered statistically significant. All statistical analyses were performed using R statistical software version 4.3.

[0158] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0159] In general, the various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general hardware or a controller or other computing device, or some combination thereof.

[0160] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0162] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0163] The example embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for characterizing drugs based on the whole genome, characterized in that: The method comprises:

101. Obtain the whole genome data and drugs of pan-cancer cell lines in the cancer drug sensitivity genomics database; the drugs include small molecule drugs; 102, calculating the IC50 value of the small molecule drug; 103, the gene expression significantly associated with small molecule drugs in the whole genome data determined based on IC50 values ​​is denoted as Ge; 104. Calculate the correlation between the IC50 value and the significantly correlated gene and record it as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize small molecule drugs.

2. The method for drug characterization based on whole genome according to claim 1, characterized in that: If the first small molecule drug is missing from the genomic database, the most relevant drug in the database with a similar mechanism of action and the same anti-tumor drug type is used to represent the first small molecule drug; Optionally, if the genomic databases include at least 2, and the second small molecule drug exists in at least 2 databases at the same time, the second small molecule drug is preferentially calculated using the database with a high degree of recognition; Optionally, the small molecule drug includes any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; Optionally, the value range of the correlation is greater than or equal to -1 and less than or equal to 1; Optionally, the cancer drug sensitivity genomics database includes GDSC and CTRP datasets.

3. The method for drug characterization based on whole genome according to claim 1, characterized in that: The small molecule drug can be replaced by an antibody drug; or, the drug also includes an antibody drug; Optionally, the characterization method of the antibody drug includes: determining the target gene of the antibody drug; determining the gene significantly correlated with the target gene in the whole genome data; calculating the correlation between the expression level of the target gene and the expression level of the significantly correlated gene Ge, recorded as the gene-drug sensitivity correlation Gcd; Gcd is used to characterize the antibody drug; Preferably, a negative value of the correlation between the expression level of the target gene and the significantly correlated gene expression Ge represents an antibody drug; optionally, the antibody drug includes: HER2 drug.

4. A method for constructing a model for predicting neoadjuvant therapy outcomes, characterized in that: The method comprises:

201. Obtaining whole genome data and treatment regimen of training set samples that have received neoadjuvant therapy; the treatment regimen is a regimen comprising at least one drug; the drug is characterized by the gene-drug sensitivity correlation Gcd of any one of claims 1 to 3; 202. Input the whole genome data and treatment regimen of the training set samples into a bioinformatics neural network model including an input layer, a gene layer, at least one biological process layer and an output layer for training to obtain the neoadjuvant therapy outcome model.

5. The method for constructing a model for predicting neoadjuvant therapy outcomes according to claim 4, characterized in that: When the treatment regimen includes at least 2 drugs, the sum of the drug sensitivity correlation Gcd of the same gene in each drug characterization method is calculated to represent the treatment regimen; optionally, the method further includes performing rank normalization processing on the whole genome data of each sample in 201; Optionally, the whole genome data and treatment plan in 202 are input into the neural network model in a manner including: inputting Ge and Gcd into the neural network model in parallel, or aggregating Ge and Gcd to obtain an attention value Gatt of gene expression during treatment, and inputting the attention value into the neural network model; Optionally, the calculation method of the attention value Gatt includes: Gatt = Ge x *(Exp -Gcd ); Optionally, the biological process layer includes any one or more of the following: a refined pathway layer, a more complex biological pathway, and a biological process; Optionally, the bioinformatics neural network model includes: a GDnet framework; Optionally, the loss weights of the input layer, gene layer, and biological process layer results are increased step by step; Optionally, the method further includes: calculating sample-level importance scores of all nodes in all layers in the model.

6. A method for predicting the outcome of neoadjuvant therapy, characterized in that: The method comprises: 301, obtaining the genomic data and types of therapeutic drugs tested; 302, inputting the genomic data and the type of therapeutic drug into the neoadjuvant therapy outcome model according to any one of claims 4-5, and outputting a predicted response value after receiving the neoadjuvant therapy with the type of therapeutic drug; Optionally, the method also includes 303: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the reaction prediction value; when the reaction prediction value is greater than a first threshold, outputting an auxiliary prediction result that the subject has a good therapeutic effect using the therapeutic drug type; when the prediction score is less than the first threshold, outputting an auxiliary prediction result that the subject has a poor therapeutic effect using the therapeutic drug type.

7. A method for predicting the outcome of neoadjuvant therapy, characterized in that: The method comprises: 401, obtaining the genomic data of the subject and the type of therapeutic drug; the genes in the genomic data include any one or more of the following: PSMC5, CCND1, PSMB3, PSMD14, CUL1; 402, inputting the genomic data and the type of therapeutic drug into the neoadjuvant therapy outcome model according to any one of claims 4-5, and outputting a predicted response value after receiving the neoadjuvant therapy with the type of therapeutic drug; Optionally, the therapeutic drug types include any one or more of the following: paclitaxel, doxorubicin, cyclophosphamide, olaparib, PDCD1, 5-fluorouracil, epirubicin, cyclophosphamide, lapatinib, ERBB2, ERBB2MK-2206, ENG; Optionally, the method also includes 403: obtaining an auxiliary prediction effect of the therapeutic effect of the drug type based on the reaction prediction value; when the reaction prediction value is greater than a first threshold, outputting an auxiliary prediction result that the subject has a good therapeutic effect using the therapeutic drug type; when the prediction score is less than the first threshold, outputting an auxiliary prediction result that the subject has a poor therapeutic effect using the therapeutic drug type.

8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Collaborative anti-tumor multi-drug combination effect prediction method based on deep learning

    CN111223577A

  • Drug sensitivity prediction method, electronic equipment and computer readable storage medium

    CN112951327A

  • Esophageal squamous carcinoma typing and key gene screening method and system

    CN118609652A

  • Method for constructing prognosis model related to neoadjuvant chemotherapy pyroptosis of breast cancer

    CN119170195A

  • Deep learning model prediction method of drug IC50 based on molecular structure and gene expression

    US20230298720A1