Therapeutic effect prediction method based on multi-organ metastasis genome data
By integrating genomic data from metastatic lesions in multiple organs, a treatment efficacy prediction model based on machine learning and deep learning was constructed. This solved the problem of insufficient reliance on primary lesion data and nonlinear interaction between the model and existing technologies, and enabled high-precision individualized treatment plans to be guided for metastatic breast cancer.
Patent Information
- Application Number
- CN202510835116.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies rely on primary lesion data to predict the efficacy of treatment for metastatic breast cancer, ignoring clonal evolution during metastasis and genomic heterogeneity of metastatic lesions in multiple organs. They cannot reflect the true biological characteristics, and most models use simple linear regression or risk scoring systems, which are difficult to capture the complex nonlinear interactions between genes and lack the ability to comprehensively evaluate multiple treatment options.
By integrating multi-dimensional genomic data such as somatic mutations, germline mutations, and copy number variations from multiple organ metastases, and combining regularization algorithms and nonlinear models for feature selection, a machine learning and deep learning framework is constructed. A hierarchical randomization method is used to divide the training set and the test set, and an attention-based efficacy prediction model is built. The consistency of the model's recommendations is verified through perturbation testing.
It improves the accuracy of predicting the efficacy of treatment for metastatic breast cancer, provides guidance for multiple individualized treatment options, and solves the problems of existing technologies that rely on primary lesion data, ignore multi-organ heterogeneity, and have insufficient nonlinear modeling capabilities, thus achieving high-precision individualized treatment decisions.
Smart Images

Figure CN120977597A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical models, and in particular to a method for predicting therapeutic efficacy based on multi-organ metastasis genomic data. BACKGROUND
[0002] Metastatic breast cancer as its late stage, due to tumor cell spread to multiple organs and high heterogeneity of genomic characteristics, leading to complex treatment options and difficult efficacy prediction. Epidemiological data shows that about 20-30% of early breast cancer patients will progress to metastatic disease, and the 5-year survival rate after metastasis is only about 25%. With the development of large-scale genomic sequencing technology, the study of metastasis genomic characteristics has gradually deepened, and significant genomic differences have been found between primary and metastatic lesions, and between different organ metastases. These differences directly affect drug target expression and signal pathway activity, and further lead to the diversity of treatment response. How to integrate multi-dimensional genomic data to accurately predict treatment efficacy has become a key challenge in the field of precision medicine for breast cancer.
[0003] Currently, the widely used breast cancer treatment decision tools (such as Oncotype DX, MammaPrint, etc. multi-gene expression profiling, and NGS-based molecular typing system) are mainly designed based on primary tumor samples. Although they play an important role in the treatment of early breast cancer, they have limitations in the context of metastatic breast cancer. Existing technologies can detect mutations in key genes such as PIK3CA, TP53, ESR1, and tumor mutation burden (TMB), mutation signatures, and other indicators to predict the efficacy of endocrine therapy, targeted therapy or immunotherapy, and provide some reference for clinical practice.
[0004] However, the existing technology still has the following shortcomings: first, it mainly relies on primary tumor data, ignoring the clonal evolution and genomic heterogeneity of multi-organ metastases during metastasis, and cannot reflect the true biological characteristics of metastatic breast cancer; second, most models use simple linear regression or risk scoring systems, which cannot capture the complex nonlinear interactions between genes, resulting in limited prediction accuracy; in addition, most existing tools provide predictions for single treatment methods, lack comprehensive evaluation capabilities for multiple treatment options such as chemotherapy, targeted therapy, and immunotherapy, and do not fully integrate multi-dimensional genomic data such as somatic mutations, germline mutations, and copy number variations, which cannot provide accurate guidance for individualized treatment of different organ metastases. SUMMARY
[0005] The purpose of the present application is to provide a method for predicting therapeutic efficacy based on multi-organ metastasis genomic data, which solves the following technical problems:
[0006] The prior art relies on primary tumor data in metastatic breast cancer efficacy prediction, ignores clonal evolution and multi-organ metastasis genomic heterogeneity in the metastasis process, and cannot reflect the true biological characteristics; most models use simple linear regression or risk scoring systems, which are difficult to capture complex nonlinear interactions between genes, resulting in limited prediction accuracy; most of them are for single treatment mode prediction, lack of comprehensive evaluation ability for multiple treatment schemes, and do not fully integrate multi-dimensional genomic data, which cannot provide individualized precise treatment guidance for different organ metastases.
[0007] The purpose of the application can be achieved by the following technical solutions:
[0008] The efficacy prediction method based on multi-organ metastasis genomic data comprises the following steps:
[0009] S1, collecting the clinical pathological characteristics, multi-organ metastasis genomic data and treatment scheme information of a plurality of breast cancer patients; the genomic data includes somatic mutations, germline pathogenic mutations, copy number variations and tumor mutation load; the metastasis covers multiple organs, and the specific site and metastasis load state of each metastasis event are recorded;
[0010] S2, the high-dimensional genomic data is preliminarily reduced by a regularization algorithm, the importance of the characteristics is evaluated in combination with a nonlinear model, and the clinical pathological factors, treatment scheme characteristics and genomic characteristics related to treatment response are screened out;
[0011] S3, inputting the screened characteristics into a machine learning and deep learning framework, dividing the training set and the test set by using a hierarchical randomization method, optimizing the model hyperparameters by cross-validation, and respectively constructing a set model based on a classical machine learning algorithm and an efficacy prediction model based on an attention mechanism;
[0012] S4, by disturbing the treatment scheme data in the test queue, the consistency of the model recommended treatment scheme and the actual treatment scheme is evaluated, and the prediction ability of the efficacy prediction model for treatment response and the clinical practicability are verified.
[0013] As a further scheme of the application, the S1 specifically comprises:
[0014] The clinical pathological characteristics include patient age, hormone receptor status, human epidermal growth factor receptor-2 status, tumor stage and number of metastatic sites; the hormone receptor status is detected by immunohistochemistry to detect the expression level of estrogen receptor and progesterone receptor; the number of metastatic sites is recorded by imaging reports or biopsy results;
[0015] The genomic data is sequenced from tumor samples and paired blood samples by targeted sequencing technology, the targeted sequencing technology covers the coding region and regulatory region of breast cancer related genes, the sequencing process includes DNA fragmentation, library construction, primer capture and high-throughput sequencing, and bioinformatics analysis is performed on the sequencing data, the bioinformatics analysis includes sequence alignment, somatic mutation detection, germline pathogenic mutation identification and copy number variation analysis, the copy number variation analysis is determined by calculating the sequencing depth ratio of tumor samples and normal samples;
[0016] The treatment scheme covers the specific types of endocrine therapy drugs, chemotherapy drugs, targeted therapy drugs and immunotherapy drugs and combined drug regimens, and the targeted therapy drugs include anti-human epidermal growth factor receptor-2 monoclonal antibodies and cyclin-dependent kinase inhibitors.
[0017] The metastasis load state is divided into high metastasis load and low metastasis load according to the number of metastatic sites, the high metastasis load is defined as the number of metastatic sites being greater than or equal to three, and the low metastasis load is defined as the number of metastatic sites being one to two.
[0018] As a further scheme of the application, the S3 specifically includes:
[0019] Linear dimension reduction is performed on high-dimensional genomic data and clinical pathological characteristics by a regularization algorithm, the regularization algorithm eliminates variables that have no significant correlation with treatment response by constraining feature coefficients;
[0020] The importance score of the remaining features is calculated by using a nonlinear model, the nonlinear model evaluates the contribution of features to the prediction of treatment response based on a decision tree structure; the importance score is calculated by combining the number of splits and information gain of features in the model;
[0021] The finally retained features include hormone receptor status, human epidermal growth factor receptor-2 expression level, metastatic site distribution, targeted drug type, combined drug regimen, organ-specific mutation and copy number variation, the organ-specific mutation is defined as a gene mutation that frequently occurs in a specific metastatic site, and the copy number variation is defined as a gene amplification or deletion related to the drug resistance or sensitivity of metastatic foci.
[0022] As a further scheme of the application, the S3 specifically includes:
[0023] The data set is proportionally divided into a training set and a test set, and the stratified randomization method ensures that the distribution of metastasis load state and treatment scheme in the training set and the test set is consistent with the original data set;
[0024] The hyperparameters of the machine learning models are optimized using cross-validation in the training set. The machine learning models include logistic regression, support vector machine, random forest, and gradient boosting models. The hyperparameter optimization is achieved through grid search or random search methods.
[0025] A deep learning model based on an attention mechanism is constructed. The deep learning model captures the interaction between genomic features and clinical features through a bidirectional attention layer. The bidirectional attention layer calculates the weight matrix between features and extracts key patterns related to treatment response.
[0026] The input layer of the deep learning model receives standardized clinicopathological features, genomic features, and treatment plan data; the standardization includes mean and variance normalization of continuous features and one-hot encoding of discrete features; the output layer of the model converts the attention output into treatment response probabilities through a fully connected layer.
[0027] As a further aspect of the present invention: S4 specifically includes:
[0028] Drug regimen perturbation data is generated in the test set. The perturbation data is generated by modifying the drug types or combinations in the original treatment regimen while keeping the patient's clinicopathological characteristics and genomic data unchanged.
[0029] The perturbation data is input into the efficacy prediction model to obtain the predicted response probability under different treatment options; the predicted response probability is generated based on the model's ranking of the efficacy of the treatment options.
[0030] The prediction results are divided into a matching group and a non-matching group according to a preset threshold. The matching group consists of cases in which the model's recommended treatment plan is consistent with the actual treatment plan. The preset threshold is determined by the best classification performance of the validation set.
[0031] The predictive accuracy of the model was verified by statistically analyzing the objective response rate of patients in the matched group; the objective response rate was quantified by imaging assessment or pathological examination results; the differences in treatment response between the matched group and the unmatched group were analyzed for different breast cancer subtypes, including luminal breast cancer, human epidermal growth factor receptor-2 positive breast cancer, and triple-negative breast cancer.
[0032] As a further aspect of the present invention, the process of validating the efficacy prediction model also includes:
[0033] For patients with luminal breast cancer and triple-negative breast cancer, the actual efficacy of the model-recommended treatment regimen was compared with that of the routine treatment regimen based on international clinical guidelines.
[0034] The stability of the efficacy prediction model under different data distributions is evaluated by multiple random partitioning of the test set and the training set, and the stability is quantified by the standard deviation or the coefficient of variation of the prediction performance;
[0035] A test cohort is generated by integrating multiple clinical trial data to verify the recommendation ability of the model for potential treatment strategies by artificially changing the drug treatment regimen; the efficacy data of the test cohort is compared with the model prediction results for consistency.
[0036] As a further scheme of the present application: in the S4, the output result of the efficacy prediction model is the response probability of the treatment regimen, specifically comprising:
[0037] A combination drug regimen is recommended for patients with high metastasis load, and the combination drug regimen preferentially selects a combination of targeted drugs with a predicted response probability higher than a threshold;
[0038] The local treatment strategy is adjusted according to the organ distribution characteristics of metastases, and the adjustment includes selecting a radiotherapy or interventional therapy regimen matched with a specific organ genomic variation pattern;
[0039] During the treatment process, the model input data is updated in combination with new metastasis events or genomic detection results, and the response probability of the treatment regimen is recalculated to optimize the prediction results.
[0040] As a further scheme of the present application: the process of updating the model input data is:
[0041] New metastatic lesion samples of the patient are regularly collected for genomic sequencing, and the sequencing covers somatic mutations, germline mutations and copy number variations;
[0042] The updated genomic data is integrated with the original clinicopathological features, and the model is input to regenerate the treatment response probability;
[0043] The drug regimen is adjusted according to the latest prediction results, and the adjustment includes replacing drugs with a decreased response probability or introducing newly approved targeted therapeutic drugs.
[0044] The beneficial effects of the present application are:
[0045] The present application integrates somatic mutations, germline mutations, copy number variations, tumor mutation load and other comprehensive genomic data of multiple organ metastases, combines regular algorithm and nonlinear model for feature screening, constructs a comprehensive data set containing clinical pathological factors, treatment plan and genomic characteristics, and divides the training set and test set by using hierarchical randomization method, uses the classical machine learning algorithm to construct the set model, and based on the deep learning model of the Transformer architecture, captures the long distance dependence relationship and complex interaction mode of gene variation, improves the prediction accuracy, verifies the model recommendation consistency by disturbing the treatment plan, finds the improvement trend of objective remission rate of the matched group of patients, solves the defects of the prior art such as relying on primary tumor data, ignoring multiple organ heterogeneity, insufficient nonlinear modeling ability of the model and single treatment evaluation, provides a high-precision individualized efficacy prediction system integrating multi-dimensional data and covering multiple treatment plans for metastatic breast cancer, and improves the accuracy of clinical treatment decision. BRIEF DESCRIPTION OF DRAWINGS
[0046] The present application will be further described below in conjunction with the accompanying drawings.
[0047] Figure 1 is a flowchart of the present application;
[0048] Figure 2 is a flowchart of machine learning of the present application;
[0049] Figure 3 is a flowchart of deep learning of the present application;
[0050] Figure 4 is a ROC curve diagram of the machine learning model of the present application;
[0051] Figure 5 is a ROC curve diagram of the deep learning model of the present application;
[0052] Figure 6 is a comparison diagram of the prediction performance of the classical machine learning and deep learning model of the present application;
[0053] Figure 7 is a verification diagram of the test set model efficacy prediction result of the present application;
[0054] Figure 8 is a verification result diagram of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0056] Referring to Figures 1-8 As shown in the accompanying drawings, the present application is a method for predicting efficacy based on multi-organ metastasis genomic data, comprising the following steps:
[0057] 1) Efficacy prediction model input data generation
[0058] Prospectively collected data of 711 patients with newly diagnosed stage IV or recurrent / metastatic breast cancer who visited the Breast Surgery Department of Fudan University Shanghai Cancer Center from April 2018 to August 2023. The baseline characteristics of the patients (such as age at diagnosis, menstrual status, tumor pathological stage, family history, etc.), multi-line treatment regimens and efficacy data were comprehensively collected and organized. Among them, the tumor samples were evaluated in the FUSCC pathology department to determine the estrogen receptor (ER), progesterone receptor (PR) status and human epidermal growth factor receptor-2 (HER2) immunohistochemistry and / or FISH detection results. The suspicious lesions found in the imaging reports or biopsy were listed as metastatic lesions, and the specific sites of metastasis events during the follow-up period of each patient were reviewed and recorded one by one, and the metastasis load status of each patient was defined: high metastasis load (metastatic sites ≥3) and low metastasis load (1-2 metastatic sites). In terms of genomic data collection, 711 pairs of tumor and blood samples were sent to the Chinese National Human Genome Center (CHGC) for high-depth sequencing. The sequencing process includes DNA fragmentation, library construction, primer capture and quantification, high-throughput sequencing and other processes. BWA-MEM algorithm was used for bioinformatics analysis of sequencing data, and Illumina sequencing platform sequences meeting the quality control standards were aligned with the hg19 version of the human genome reference sequence (GRCh37). GATKMutect2 and GATKHaplotypeCaller software were used to identify somatic mutations and germline pathogenic mutations in targeted sequencing data. Gene-level amplification and deletion detection was completed using the FACETS algorithm.
[0059] 2) Efficacy prediction model construction and verification
[0060] 1. High-dimensional data was initially filtered using least absolute shrinkage and selection operator (LASSO) algorithm before building the prediction model. Unimportant feature coefficients were shrunk to zero through L1 regularization, thus retaining a subset of features that had significant linear relationship with the target variable. Subsequently, based on the results of LASSO filtering, the importance of features was further assessed using random forest (RF) model, which captured the complex relationship between features and target variable with its nonlinear modeling capability, and selected features with higher importance scores. Finally, 13 clinicopathological factors, 15 treatment regimens, and 35 genomic features were selected for model training. Clinicopathological factors included diagnosis age, HR and HER2 status, Ki-67 index, metastasis-related variables such as sample site (breast, bone, lung, liver, distant lymph node, and chest wall) and metastasis burden (divided into low and high). Treatment features included endocrine therapy regimen, conventional chemotherapy regimen (gemcitabine, vinorelbine, paclitaxel, platinum, and capecitabine), targeted therapy (anti-HER2 therapy, CDK4 / 6 inhibitor, everolimus, and tyrosine kinase inhibitor, including bevacizumab, famitinib, saracatinib, and pyrotinib), and anti-PD1 immunotherapy. Genomic features included alterations related to metastasis and organ-specific metastasis, as well as genomic alterations observed in at least 10 cases in our study.
[0061] 2. The cohort was split into training and test sets in a 7:3 ratio by stratified randomization. A three-step machine learning pipeline was adopted for feature selection and model training. First, features with mutual correlation higher than 0.9 were removed, and the top correlated features with binary response were retained. Second, the top k features were selected based on ANOVA-F value, and were standardized using z-score. Finally, seven machine learning algorithms were independently trained and validated: logistic regression (LR), support vector classifier (SVC), random forest (RF), eXtreme gradient boosting (XGBoost), light gradient boosting machine (LGBM), K-nearest neighbors (KNN), and Bernoulli naive Bayes (BNB). All hyperparameters were optimized by 1000-step cross-validation to maximize the area under the receiver operating characteristic curve (AUROC). Hyperparameters of the machine learning algorithms were optimized in the training set using five-fold cross-validation Figure 2 ).
[0062] 3. Deep learning implementation adopted tabular prior-data fitted networks (TabPFN) based on context learning and bidirectional attention mechanism, which was selected due to its superior performance in computational efficiency and performance over traditional machine learning algorithms. Similarly, the input cohort was split into training and test sets in a 7:3 ratio by stratified randomization. The analysis used the AutoTabPFNClassifier function to generate optimized prediction results, which automatically combined selected TabPFN models through post-hoc ensemble techniques and optimized their parameters by randomly sampling from a defined search space Figure 3 ).
[0063] 4. Two different algorithms were used to build predictive models of metastatic lesion efficacy in the training set, with the area under the curve (AUC) of the receiver operating characteristic (ROC) curve as the main indicator of the effectiveness of the predictive model. In the validation set data, the average ensemble model of the classic machine learning algorithm had an AUC value of 0.765( Figure 4 ); the TabPFN model of the deep learning algorithm had an AUC value of 0.805( Figure 5 ). Further comparison of the AUC values under random conditions distinguished the algorithm type with better prediction performance, and the deep learning TabPFN model had better prediction performance than the classic machine learning algorithm model in 5 different random number conditions (P = 0.0250) Figure 6 ).
[0064] 5. To evaluate the clinical practicability of the TabPFN deep learning model, data from multiple clinical trials were integrated to generate a test cohort for model validation. The data was disturbed by artificially changing the drug treatment regimen of each test patient (i.e., changing the specific drug used), while keeping other clinical and genomic features unchanged. Based on the efficacy prediction scores generated by these drug regimen disturbances, the treatment regimen recommended by the deep learning model was evaluated and determined. With a prediction probability threshold of 60%, the patients were divided into a matched group and a non-matched group by comparing the consistency of the treatment regimen recommended by the model with the actual treatment regimen received Figure 7 ). The luminal and triple-negative breast cancer patients in the matched group showed a trend of improvement in the objective response rate (ORR) Figure 8 .
[0065] 6. In summary, we found that the clinical guidance provided by the TabPFN deep learning model can accurately identify new potential treatment regimens that improve patient outcomes.
[0066] The above describes one embodiment of the present application in detail, but the content described is only the preferred embodiment of the present application and cannot be considered as limiting the scope of the present application. Any equivalent changes and improvements made in accordance with the scope of the present application should still be within the scope of the present application.
Claims
1. A method for predicting treatment efficacy based on genomic data of multi-organ metastases, characterized in that, Includes the following steps: S1. Collect clinicopathological characteristics, genomic data of multi-organ metastases, and treatment plans of several breast cancer patients; the genomic data includes somatic mutations, germline pathogenic mutations, copy number variations, and tumor mutational burden; the metastases cover multiple organs, and the specific location and metastatic burden status of each metastatic event are recorded; S2. The high-dimensional genomic data is initially reduced in dimensionality using a regularization algorithm. The importance of features is evaluated by combining a nonlinear model, and clinicopathological factors, treatment regimen features, and genomic features related to treatment response are screened out. S3. Input the selected features into the machine learning and deep learning framework, use hierarchical randomization to divide the training set and test set, optimize the model hyperparameters through cross-validation, and build an ensemble model based on classic machine learning algorithms and an efficacy prediction model based on attention mechanism, respectively. S4. By using treatment protocol data in the perturbation test queue, evaluate the consistency between the model-recommended treatment protocols and the actual treatment protocols, and verify the predictive ability and clinical applicability of the efficacy prediction model for treatment response.
2. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 1, characterized in that, S1 specifically includes: The clinicopathological features include patient age, hormone receptor status, human epidermal growth factor receptor-2 status, tumor stage, and number of metastatic sites; the hormone receptor status is determined by immunohistochemistry to detect the expression levels of estrogen receptor and progesterone receptor; the number of metastatic sites is clearly recorded through imaging reports or biopsy results. The genomic data were sequenced from tumor samples and paired blood samples using targeted sequencing technology. This targeted sequencing technology covers the coding and regulatory regions of breast cancer-related genes. The sequencing process includes DNA fragmentation, library construction, primer capture, and high-throughput sequencing. Bioinformatics analysis was performed on the sequencing data, including sequence alignment, somatic mutation detection, germline pathogenic mutation identification, and copy number variation analysis. The copy number variation analysis was determined by calculating the sequencing depth ratio between tumor and normal samples. The treatment plan covers the specific types and combination regimens of endocrine therapy drugs, chemotherapy drugs, targeted therapy drugs, and immunotherapy drugs; the targeted therapy drugs include anti-human epidermal growth factor receptor-2 monoclonal antibodies and cyclin-dependent kinase inhibitors. The load transfer status is divided into two categories based on the number of transfer points: high load transfer and low load transfer. High load transfer is defined as having three or more transfer points, while low load transfer is defined as having one or two transfer points.
3. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 1, characterized in that, S3 specifically includes: Linear dimensionality reduction of high-dimensional genomic data and clinicopathological features is performed using a regularization algorithm. The regularization algorithm eliminates variables that are not significantly correlated with treatment response by constraining feature coefficients. The remaining features are scored for importance using a nonlinear model, which evaluates the contribution of features to treatment response prediction based on a decision tree structure; the importance score is calculated by combining the number of splits of the feature in the model and the amount of information gain. The final retained features include hormone receptor status, human epidermal growth factor receptor-2 expression level, distribution of metastatic sites, type of targeted drug, combination therapy regimen, organ-specific mutations, and copy number variations; organ-specific mutations are defined as gene mutations that occur frequently in specific metastatic sites, and copy number variations are defined as gene amplifications or deletions associated with drug resistance or sensitivity of metastatic lesions.
4. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 1, characterized in that, S3 specifically includes: The dataset is divided into training and testing sets proportionally, and the hierarchical randomization method ensures that the distribution of load states and treatment plans in the training and testing sets is consistent with the original dataset. The hyperparameters of the machine learning models are optimized using cross-validation in the training set. The machine learning models include logistic regression, support vector machine, random forest, and gradient boosting models. The hyperparameter optimization is achieved through grid search or random search methods. A deep learning model based on an attention mechanism is constructed. The deep learning model captures the interaction between genomic features and clinical features through a bidirectional attention layer. The bidirectional attention layer calculates the weight matrix between features and extracts key patterns related to treatment response. The input layer of the deep learning model receives standardized clinicopathological features, genomic features, and treatment plan data; the standardization includes mean and variance normalization of continuous features and one-hot encoding of discrete features; the output layer of the model converts the attention output into treatment response probabilities through a fully connected layer.
5. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 1, characterized in that, S4 specifically includes: Drug regimen perturbation data is generated in the test set. The perturbation data is generated by modifying the drug types or combinations in the original treatment regimen while keeping the patient's clinicopathological characteristics and genomic data unchanged. The perturbation data is input into the efficacy prediction model to obtain the predicted response probability under different treatment options; the predicted response probability is generated based on the model's ranking of the efficacy of the treatment options. The prediction results are divided into a matching group and a non-matching group according to a preset threshold. The matching group consists of cases in which the model's recommended treatment plan is consistent with the actual treatment plan. The preset threshold is determined by the best classification performance of the validation set. The predictive accuracy of the model was verified by statistically analyzing the objective response rate of patients in the matched group; the objective response rate was quantified by imaging assessment or pathological examination results; the differences in treatment response between the matched group and the unmatched group were analyzed for different breast cancer subtypes, including luminal breast cancer, human epidermal growth factor receptor-2 positive breast cancer, and triple-negative breast cancer.
6. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 5, characterized in that, The process of validating the efficacy prediction model also includes: For patients with luminal breast cancer and triple-negative breast cancer, the actual efficacy of the model-recommended treatment regimen was compared with that of the routine treatment regimen based on international clinical guidelines. The stability of the efficacy prediction model under different data distributions was evaluated by randomly dividing the test set and training set multiple times; the stability was quantified by the standard deviation or coefficient of variation of the prediction performance. A test cohort is generated by integrating data from multiple clinical trials. The model's ability to recommend potential treatment strategies is verified by artificially altering drug treatment regimens. The efficacy data of the test cohort is then compared with the model's prediction results for consistency.
7. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 1, characterized in that, In step S4, the output of the efficacy prediction model is the response probability of the treatment plan, specifically including: For patients with high metastatic burden, combination therapy regimens are recommended, with priority given to combinations of targeted drugs whose predicted response probability is higher than a threshold. The local treatment strategy is adjusted according to the organ distribution characteristics of the metastatic lesions, and the adjustment includes selecting a radiotherapy or interventional treatment plan that matches the genomic variation pattern of a specific organ; During treatment, the model input data is updated by incorporating new metastatic events or genomic testing results, and the response probability of the treatment plan is recalculated to optimize the prediction results.
8. The efficacy prediction method based on multi-organ metastatic lesion genomic data according to claim 7, characterized in that, The process of updating the model input data is as follows: Regularly collect new metastatic lesion samples from patients for genome sequencing, which covers somatic mutations, germline mutations, and copy number variations. The updated genomic data was integrated with the original clinicopathological features and input into the model to regenerate the treatment response probability; The medication regimen is adjusted based on the latest forecast results, including replacing drugs with decreased response probability or introducing newly approved targeted therapies.