Esophageal cancer neoadjuvant chemotherapy and immunization effect prediction method and system

Through the deep coupling of multi-source data intelligent governance and dynamic weight decision engine, the problems of data heterogeneity, feature weight solidification and insufficient privacy protection in the prediction of esophageal cancer efficacy have been solved, and efficient and safe prediction of the efficacy of neoadjuvant chemotherapy and immunotherapy for esophageal cancer has been achieved, which enhances the generalization ability of the model and the trust of doctors in the prediction results.

CN120600344APending Publication Date: 2025-09-05ZHEJIANG CANCER HOSPITAL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511108091.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing technology for predicting the efficacy of esophageal cancer faces problems such as high heterogeneity of multi-source data, static feature weights, and insufficient privacy protection, which leads to serious data silos, limited model generalization ability, large spatiotemporal alignment errors between imaging data and clinical data, and insufficient standardization of genomic data.

Method used

Through the deep coupling of multi-source data intelligent governance and dynamic weight decision engine, standardized integration and dynamic weight adjustment of multi-source heterogeneous data are achieved. The XGBoost algorithm is used to train the federated prediction model, and a privacy protection framework and closed-loop optimization strategy are introduced to improve prediction accuracy and adaptability.

Benefits of technology

It improves the accuracy and adaptability of predicting the efficacy of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, enhances the model's generalization ability for complex cases, ensures data security, and enhances doctors' trust in prediction results through personalized modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600344A_ABST
    Figure CN120600344A_ABST
Patent Text Reader

Abstract

The invention provides an esophageal cancer neoadjuvant chemotherapy and immunization effect prediction method and system, and relates to the technical field of medical big data analysis. Aiming at the problems of data islands, feature weight solidification and insufficient privacy protection in a traditional method, the system realizes standardized integration of cross-modal data, and adapts to feature importance differences of different patient groups through a dynamic weight adjustment mechanism, so that the generalization ability of a model to complex cases is enhanced. Meanwhile, a privacy protection framework and a closed-loop optimization strategy are introduced, and on the premise that data security is guaranteed, sustainable evolution of a prediction model and transparent support of clinical decisions are achieved. Finally, the system not only improves the accuracy of curative effect prediction, but also enhances the credibility and the adoption rate of a doctor to a prediction result through personalized modeling and interpretability analysis, and provides an efficient and safe solution for precise medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical big data analysis, and in particular to a method and system for predicting the effects of neoadjuvant chemotherapy and immunization for esophageal cancer. Background Art

[0002] Current esophageal cancer treatment efficacy prediction faces core challenges, including high heterogeneity in multi-source data, static feature weighting, and insufficient privacy protection. Traditional methods rely on a single data source, making it difficult to fully characterize patient status. Feature weights are often assigned in a fixed ratio, ignoring individual patient differences. Furthermore, cross-institutional model training lacks privacy protection mechanisms, leading to severe data silos. Existing technologies suffer from significant spatiotemporal alignment errors between imaging and clinical data, and insufficient standardization of genomic data further limits model generalization. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer. Through the deep coupling of multi-source data intelligent management and a dynamic weight decision engine, the accuracy and adaptability of efficacy prediction are significantly improved.

[0004] This application proposes a method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, which includes: S1: collecting at least one first multi-source heterogeneous data and performing a first normalization process to obtain at least one target multi-source heterogeneous data, and performing a first clustering process on the at least one target multi-source heterogeneous data to obtain a first patient portrait; S2: performing a first dynamic weight adjustment process according to the first patient portrait and at least one target multi-source heterogeneous data to determine at least one first dynamic weight value; S3: training at least one first federated prediction model based on at least one target multi-source heterogeneous data and at least one first dynamic weight value, and determining and distributing a first target federated prediction model based on the first patient profile; S4: Obtain first prediction information according to the first target federated prediction model, perform a first difference analysis process on the first prediction information to obtain a first dynamic weight adjustment rule, and return to S2.

[0005] Preferably, the first multi-source heterogeneous data includes imaging data, clinical test data and genomic test data.

[0006] Preferably, the S1 includes: S11: Collect at least one first multi-source heterogeneous data, and perform a first preprocessing according to the type of the first multi-source heterogeneous data to obtain at least one second multi-source heterogeneous data; S12: performing a first spatiotemporal alignment process on at least one of the second multi-source heterogeneous data to obtain at least one third multi-source heterogeneous data; S13: performing a first semantic consistency process on at least one of the third multi-source heterogeneous data to obtain at least one target multi-source heterogeneous data; S14: Inputting at least one target multi-source heterogeneous data into a patient portrait construction model to output a first patient portrait.

[0007] Preferably, the first spatiotemporal alignment process includes a time axis alignment process and a spatial registration process.

[0008] Preferably, the S2 includes: S21: Constructing a first weight rule base for different patient types; S22: Determine a first weight rule item from the first weight rule library according to the first patient portrait, and perform a first dynamic weight adjustment process on at least one of the target multi-source heterogeneous data based on the first weight rule item to obtain at least one first initial dynamic weight value.

[0009] Preferably, after S22, the method further includes: S23: Perform weight normalization processing on at least one of the first initial dynamic weight values ​​to determine at least one first dynamic weight value.

[0010] Preferably, S3 includes the following sub-steps: S31: Obtaining at least one first federated prediction model based on XGBoost algorithm training according to at least one target multi-source heterogeneous data and at least one first dynamic weight value; S32: Performing a first security aggregation process on at least one of the first federated prediction models to obtain at least one target federated prediction model; S33: Determine a first target federated prediction model from at least one of the target federated prediction models according to the first patient portrait, and perform a first verification and distribution process on the first target federated prediction model.

[0011] Preferably, the S4 includes the following sub-steps: S41: Obtaining a first confidence level from the first prediction information output by the first target federated prediction model, and determining a first human-machine collaborative decision based on the first confidence level; S42: Performing a first difference analysis process on the first prediction information to obtain at least one first dispute feature and at least one first weight adjustment suggestion; S43: Determine at least one first optimization parameter based on at least one first dispute feature and at least one first weight adjustment suggestion and feed it back to S2.

[0012] Preferably, the first patient profile is one of metabolism-dominant, immune-sensitive and gene-driven.

[0013] The present application also proposes a system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, which is used to implement the above-mentioned method for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer.

[0014] The present application proposes a method and system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, which relates to the field of medical big data analysis technology. In response to the problems of data silos, feature weight solidification, and insufficient privacy protection in traditional methods, the system realizes the standardized integration of cross-modal data, and adapts to the differences in feature importance of different patient groups through a dynamic weight adjustment mechanism, thereby enhancing the model's generalization ability for complex cases. At the same time, the introduction of a privacy protection framework and a closed-loop optimization strategy realizes the continuous evolution of the prediction model and transparent support for clinical decision-making while ensuring data security. Ultimately, the system not only improves the accuracy of efficacy prediction, but also enhances doctors' trust and adoption rate of prediction results through personalized modeling and explainable analysis, providing an efficient and safe solution for precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.

[0016] Figure 1 This is a flowchart of an execution of a method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer in the present invention.

[0017] Figure 2 It is a schematic diagram of the execution of the overall optimization and adjustment process in the technical solution of this application. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The exemplary embodiments and descriptions are only used to explain the present invention but are not intended to limit the present invention.

[0020] The following is a detailed description of a method and system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to the present invention.

[0021] This embodiment proposes a method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer. The specific process is as follows: Figure 1 shown.

[0022] S1: Collect at least one first multi-source heterogeneous data and perform a first normalization process to obtain at least one target multi-source heterogeneous data, and perform a first clustering process on the at least one target multi-source heterogeneous data to obtain a first patient portrait.

[0023] In this step, it is necessary to unify multi-source heterogeneous data into standardized feature vectors, eliminate spatiotemporal and semantic differences, and provide an analytical basis for subsequent dynamic weight decisions.

[0024] The types of the first multi-source heterogeneous data include imaging data, clinical test data, and genomic test data. The first normalization process includes a first preprocessing operation and a first spatiotemporal alignment process.

[0025] The S1 includes the following sub-steps: S11: Collect at least one first multi-source heterogeneous data, and perform a first preprocessing according to the type of the first multi-source heterogeneous data to obtain at least one second multi-source heterogeneous data.

[0026] Different types of the first multi-source heterogeneous data correspond to different first preprocessing operation modes, and each second multi-source heterogeneous data is in the form of a feature vector. 1. Image data: Input: CT / MRI images in DICOM format.

[0027] First preprocessing: VTK tools were used to extract tumor location, morphology (such as volume and major diameter), and texture features (such as the entropy value of the gray-level co-occurrence matrix).

[0028] Generate a 3D reconstructed model and quantify the changes in tumor morphology (such as the percentage of volume reduction). This needs to be obtained by comparative analysis with the imaging data of the previous stage.

[0029] Output a normalized image feature vector (eg, volume_change, entropy), in which each dimension corresponds to the data extracted after the first preprocessing operation.

[0030] 2. Clinical test data: Input: Electronic medical records in HL7 format.

[0031] First preprocessing: Laboratory test results (such as HbA1c and albumin levels), pathology reports (such as PD-L1 expression), and treatment records (such as the number of neoadjuvant chemotherapy cycles) were extracted.

[0032] The missing values ​​were filled using K-nearest neighbor interpolation method, and outliers were removed.

[0033] A standardized clinical feature vector is output, in which each dimension corresponds to the data extracted after the first preprocessing operation.

[0034] 3. Genomic testing data: Input: FASTQ files (DNA sequencing data).

[0035] Second preprocessing: The BWA-MEM tool was used to align the genome sequences and extract mutation information of key genes such as TP53 and ERBB2.

[0036] Low-frequency mutations (e.g., frequency < 0.5%) were filtered out, and high-frequency mutations (e.g., TP53 exon 6 mutations) were retained.

[0037] Output a standardized genomic feature vector (e.g., TP53_mutant, ctDNA_concentration), in which each dimension corresponds to the data extracted after the first preprocessing operation.

[0038] S12: Perform a first spatiotemporal alignment process on at least one of the second multi-source heterogeneous data to obtain at least one third multi-source heterogeneous data.

[0039] In S11, feature extraction of different types of data has been completed. For at least one feature vector obtained by extraction, spatiotemporal alignment processing is required in this step to facilitate the subsequent combination of multiple types of data to accurately portray the patient. The first spatiotemporal alignment process may specifically include: 1. Timeline alignment: Input: data of all non-consecutive time points in at least one of the second multi-source heterogeneous data (such as HbA1c at T0+10d and T0+45d).

[0040] Processing process: Taking the date of diagnosis (T0) as the benchmark, the time axis is divided into standard nodes (T0+7d, T0+30d, T0+60d).

[0041] Cubic spline interpolation was used for data with intervals >7 days to ensure temporal continuity.

[0042] Output standardized time series features (such as HbA1c_T0+30d, which indicates the change of HbA1c features within 30 days).

[0043] 2. Spatial registration: Input: all data with spatial attributes in at least one of the second multi-source heterogeneous data, such as images taken by different devices (such as CT and MRI).

[0044] Processing process: The image data were aligned using the B-spline non-rigid registration algorithm, and the error was controlled within 1.5 mm.

[0045] Generate a heat map of tumor volume changes.

[0046] Outputs normalized spatial features (such as tumor_volume_change).

[0047] The standardized time series features and the standardized spatial features are combined to obtain at least one third multi-source heterogeneous data.

[0048] S13: Perform a first semantic consistency process on at least one of the third multi-source heterogeneous data to obtain at least one target multi-source heterogeneous data.

[0049] After S11 and S12, the multi-category and multi-source data corresponding to the patient are unified in time and space. However, in the subsequent portrait clustering processing, all the features to be clustered are required to have high semantic consistency, such as classifying some discrete values ​​into multiple diagnostic categories, unifying the English expressions detected by the instrument into Chinese expressions, and encoding some features with complex expressions.

[0050] Therefore, in this step, it is necessary to perform a first semantic consistency process on at least one of the third multi-source heterogeneous data outputted by S12.

[0051] Preferably, before the first semantic consistency processing, a process of constructing a medical ontology library may be included: The medical ontology library specifies semantic extension rules for each category of features. For example, for features such as adenocarcinoma classification, unified coding can be achieved by comparing RadLex annotation data sources. Example construction rules are as follows:

[0052] The first semantic consistency processing may specifically include the following feature conversion steps. Different categories of features correspond to different conversion methods. The conversion rules may refer to the pre-established medical ontology library. Exemplary conversion rules are as follows: Numerical features: Unit standardization (e.g. kPa → Pa) and normalization to the [0, 1] interval.

[0053] Category features: Hierarchical one-hot encoding (e.g., diabetes classification → "Diabetes-Type 1", "Diabetes-Type 2").

[0054] Image features: Naming standardization (e.g. “texture_entropy” instead of “GLCM_Entropy”).

[0055] After the first semantic consistency processing, at least one target multi-source heterogeneous data can be determined.

[0056] S14: Inputting at least one target multi-source heterogeneous data into a patient portrait construction model to output a first patient portrait.

[0057] The patient portrait construction model is obtained by clustering analysis and training of historical big data. Specifically, it includes: 1. Obtain historical big data for characterizing patient condition samples, and organize each piece of the historical big data into the form of the target multi-source heterogeneous data in accordance with steps S11-S13 to obtain multiple pieces of target multi-source heterogeneous sample data.

[0058] 2. Use the K-means algorithm to cluster multiple target multi-source heterogeneous sample data (for example, set the number of clusters to 5) and select the k-means++ initialization strategy.

[0059] During the clustering process, the silhouette coefficient is calculated in real time to evaluate the clustering effect and ensure that the separation between clusters is high enough.

[0060] An exemplary portrait tag generation process may include:

[0061] In the above exemplary patient portrait labels, through cluster analysis of different category features, multiple types of patient portraits such as metabolic-dominant, immune-sensitive and gene-driven types can be obtained.

[0062] After determining the patient portrait construction model, at least one of the target multi-source heterogeneous data can be input into the patient portrait construction model, so as to determine the first patient portrait corresponding to the patient based on the cluster center with the highest similarity.

[0063] S2: Perform a first dynamic weight adjustment process based on the first patient portrait and at least one target multi-source heterogeneous data to determine at least one first dynamic weight value.

[0064] Both doctors and patients typically have high accuracy requirements for disease diagnosis predictions. Therefore, they need to compare the patient's current diagnosis with previous results in real time to adjust the weights of various indicators and parameters. Furthermore, the weight adjustment strategy varies depending on the patient type.

[0065] Based on the above considerations, in this step, it is necessary to dynamically adjust the weights of all data features in at least one of the target multi-source heterogeneous data according to the first patient portrait.

[0066] The S2 specifically includes the following sub-steps: S21: Construct a first weight rule library for different patient types.

[0067] The first weight rule library specifies the initial weight adjustment rules, including the triggering conditions for weight adjustment, specific adjustment rules, and medical constraints, that is, the adjustment must meet the specified constraints.

[0068] Still taking the three types of patients, namely metabolic-dominant, immune-sensitive and gene-driven, as an example, the weight adjustment trigger conditions, adjustment rules and medical constraints corresponding to each type of patient can be set as follows.

[0069]

[0070] S22: Determine a first weight rule item from the first weight rule library according to the first patient portrait, and perform a first dynamic weight adjustment process on at least one of the target multi-source heterogeneous data based on the first weight rule item to obtain at least one first initial dynamic weight value.

[0071] In this step, it is first necessary to determine the first weight rule item corresponding to the first patient portrait from the first weight rule library, and the first weight rule item records the weight dynamic adjustment rule corresponding to the first patient portrait.

[0072] Dynamic weight adjustment should be performed on all features in at least one target multi-source heterogeneous data. An exemplary execution process of weight adjustment is as follows: According to the first patient portrait (such as "metabolism-dominant type"), the trigger conditions in the matching rule library (such as "metabolic type + history of diabetes >5 years") are determined.

[0073] Adjustment rules were applied to the matched rules (e.g., ×1.6 for metabolic features and ×0.7 for imaging features).

[0074] The boundary protection mechanism is used to limit the amplitude of a single adjustment, and the upper limit of the immune feature weight is ensured to be 0.25, and the lower limit of the gene feature weight is ensured to be 0.1.

[0075] In order to ensure uniformity in the adjustment of weights of various parameters, after S22, the following steps may be further included: S23: Perform weight normalization processing on at least one of the first initial dynamic weight values ​​to determine at least one first dynamic weight value.

[0076] In this step, it is necessary to perform normalization processing on at least one of the first initial dynamic weight values ​​to determine at least one first dynamic weight value.

[0077] The weight normalization process may specifically include: Use the Softmax function to normalize the adjusted weights: 1. Input the adjusted weight vector (e.g. [1.6, 0.7, 1.8, 0.6]).

[0078] 2. Calculate the exponential value of each weight and sum it. Then divide each weight by the sum to get the normalized weight (such as [0.45, 0.12, 0.48, 0.05]).

[0079] 3. Output the normalized weights to the federated learning optimization step in step S3 as custom parameters for local model training.

[0080] S3: Based on at least one of the target multi-source heterogeneous data and at least one of the first dynamic weight values, at least one first federated prediction model is trained and obtained, and based on the first patient portrait, a first target federated prediction model is determined and distributed.

[0081] In this step, under the premise of privacy protection, at least one of the target multi-source heterogeneous data of S1 and at least one first dynamic weight value of S2 are aggregated to improve the generalization ability, and the first weight rule library is fed back to optimize.

[0082] S31: According to at least one of the target multi-source heterogeneous data and at least one of the first dynamic weight values, at least one first federated prediction model is obtained based on XGBoost algorithm training.

[0083] Among them, each of the first federal prediction models corresponds to a patient portrait type.

[0084] The following is an explanation of the specific training process: 1. Model selection: using XGBoost algorithm: The input data is at least one of the target multi-source heterogeneous data output by S1, which are all standardized data (such as HbA1c, TP53 mutation status) and at least one of the first dynamic weight values ​​output by S2 (such as metabolic=1.6).

[0085] 2. Model parameter settings: The following are the example settings of the XGBoost algorithm parameters: max_depth=6, learning_rate=0.1, tree_method = 'gpu_hist'.

[0086] 3. Output results The output data of the model can specifically include the model prediction probability (such as pCR probability = 76%) and confidence level (calculated using the uncertainty estimate built into the model, such as std_dev = 0.05).

[0087] Since the data generated by different hospitals during the detection and diagnosis stages may be somewhat different, in order to ensure the consistency of the models of different hospitals, S32 may be further included after S31.

[0088] S32: Perform a first security aggregation process on at least one of the first federated prediction models to obtain at least one target federated prediction model.

[0089] The first security aggregation process may specifically include: 1. During the model parameter aggregation phase, Gaussian noise (standard deviation σ = 0.7) is added to the model parameters of each hospital to satisfy the differential privacy constraint.

[0090] 2. The aggregation strategy uses the weighted average method, and the weight is determined by the sample size of each hospital (for example, if the sample size of a hospital accounts for 30%, its parameter weight is 0.3).

[0091] 3. Synchronize model parameters every 24 hours, and distribute the updated global model to the local models of each hospital.

[0092] S33: Determine a first target federated prediction model from at least one of the target federated prediction models according to the first patient portrait, and perform a first verification and distribution process on the first target federated prediction model.

[0093] In this step, it is necessary to determine the corresponding first target federated prediction model based on the first patient portrait determined in S1, perform verification processing before distribution, and further improve the prediction accuracy of the model by introducing a feedback mechanism.

[0094] The first verification distribution process may specifically include: 1. Distribute dedicated models (such as "Metabolic Dedicated Model v2.1") according to the patient profile type constructed in S1 (such as "Metabolic Dominant").

[0095] 2. Generate an MD5 checksum before distributing the model (e.g., model_v2.1.md5=abc123...). The recipient verifies the integrity of the model by comparing the checksum.

[0096] 3. Feedback mechanism: The distributed model was used for clinical prediction of S4 (e.g., predicted pCR probability = 76%) and triggered discrepancy analysis between physician decisions and system predictions.

[0097] If the prediction result is significantly different from the doctor's decision (e.g., pCR is predicted but not actually achieved), the rule optimization process of S4 is triggered.

[0098] S4: Obtain first prediction information according to the first target federated prediction model, perform a first difference analysis process on the first prediction information to obtain a first dynamic weight adjustment rule, and return to S2.

[0099] In this step, a first difference analysis and processing optimization is performed on the first prediction information to form a closed loop, and the dynamic weight decision step in S2 is fed back to make the weight adjustment rule in S2 more in line with actual needs.

[0100] The S4 includes the following sub-steps: S41: Obtain a first confidence level from the first prediction information output by the first target federated prediction model, and determine a first human-machine collaborative decision based on the first confidence level.

[0101] Among them, different numerical ranges of the first confidence level correspond to different processing methods. The specific confidence level grading mechanism is as follows:

[0102] S42: Perform a first difference analysis process on the first prediction information to obtain at least one first dispute feature and at least one first weight adjustment suggestion.

[0103] The first difference analysis process includes SHAP value analysis, which is as follows: Use the SHAP algorithm to interpret the prediction results of the XGBoost model in S3 and quantify the contribution of each feature to the prediction (for example, the SHAP value of HbA1c is 0.25).

[0104] Identify at least one first controversial feature that has a significant impact on the prediction results (e.g., an overly high weight for HbA1c leads to prediction bias).

[0105] For the first dispute feature, the first weight adjustment suggestion generated may be: "The weight of HbA1c in metabolic patients is too high, and it is recommended to reduce it from 1.6 to 1.4."

[0106] The quantity threshold of the first dispute feature can be set according to specific needs. If the prediction accuracy requirement is high, the quantity threshold should be set higher, otherwise it should be set lower.

[0107] For all the determined first dispute features, the first weight suggestions need to be generated one by one.

[0108] S43: Determine at least one first optimization parameter based on at least one first dispute feature and at least one first weight adjustment suggestion and feed it back to S2.

[0109] In this step, it is necessary to perform rule adjustment on the corresponding at least one first dispute feature according to the at least one first weight adjustment suggestion determined in S42.

[0110] Preferably, a Bayesian optimization method may be used, and the specific method is as follows: 1. Define the optimization parameter range (e.g. metabolic factor range 0.5-2.0, immune threshold range 30-70).

[0111] 2. Use the Bayesian search algorithm to iteratively optimize parameters to maximize model performance (such as AUC value).

[0112] The optimized parameters are fed back to the rule base in step 2 (e.g., the adjustment rule of the metabolic factor from 1.6 to 1.4).

[0113] In the attached Figure 2 The overall optimization and adjustment process of the technical solution of this application is explained in detail. The clinical feedback suggestions in S4 above can realize dynamic adjustment of weights and real-time updating of the federated learning model.

[0114] The present application also proposes a system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, which is used to implement the above-mentioned method for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer.

[0115] The present application proposes a method and system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, which relates to the field of medical big data analysis technology. In response to the problems of data silos, feature weight solidification, and insufficient privacy protection in traditional methods, the system realizes the standardized integration of cross-modal data, and adapts to the differences in feature importance of different patient groups through a dynamic weight adjustment mechanism, thereby enhancing the model's generalization ability for complex cases. At the same time, the introduction of a privacy protection framework and a closed-loop optimization strategy realizes the continuous evolution of the prediction model and transparent support for clinical decision-making while ensuring data security. Ultimately, the system not only improves the accuracy of efficacy prediction, but also enhances doctors' trust and adoption rate of prediction results through personalized modeling and explainable analysis, providing an efficient and safe solution for precision medicine.

[0116] The above description is only a preferred embodiment of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the scope of the patent application of the present invention.

Claims

1. A method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, characterized in that: The method includes: S1: collecting at least one first multi-source heterogeneous data and performing a first normalization process to obtain at least one target multi-source heterogeneous data, and performing a first clustering process on the at least one target multi-source heterogeneous data to obtain a first patient portrait; S2: performing a first dynamic weight adjustment process according to the first patient portrait and at least one target multi-source heterogeneous data to determine at least one first dynamic weight value; S3: training at least one first federated prediction model based on at least one target multi-source heterogeneous data and at least one first dynamic weight value, and determining and distributing a first target federated prediction model based on the first patient profile; S4: Obtain first prediction information according to the first target federated prediction model, perform first difference analysis processing on the first prediction information to obtain a first dynamic weight adjustment rule, and return to S2.

2. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 1, characterized in that: The first multi-source heterogeneous data includes imaging data, clinical test data and genomic test data.

3. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 2, characterized in that: Said S1 comprises: S11: Collect at least one first multi-source heterogeneous data, and perform a first preprocessing according to the type of the first multi-source heterogeneous data to obtain at least one second multi-source heterogeneous data; S12: performing a first spatiotemporal alignment process on at least one of the second multi-source heterogeneous data to obtain at least one third multi-source heterogeneous data; S13: performing a first semantic consistency process on at least one of the third multi-source heterogeneous data to obtain at least one target multi-source heterogeneous data; S14: Inputting at least one target multi-source heterogeneous data into a patient portrait construction model to output a first patient portrait.

4. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 3, characterized in that: The first spatiotemporal alignment process includes a time axis alignment process and a spatial registration process.

5. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 1, characterized in that: The S2 includes: S21: Constructing a first weight rule base for different patient types; S22: Determine a first weight rule item from the first weight rule library according to the first patient portrait, and perform a first dynamic weight adjustment process on at least one of the target multi-source heterogeneous data based on the first weight rule item to obtain at least one first initial dynamic weight value.

6. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 5, characterized in that: After S22, the method further includes: S23: Perform weight normalization processing on at least one of the first initial dynamic weight values ​​to determine at least one first dynamic weight value.

7. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 1, characterized in that: The S3 includes the following sub-steps: S31: Obtaining at least one first federated prediction model based on XGBoost algorithm training according to at least one target multi-source heterogeneous data and at least one first dynamic weight value; S32: Performing a first security aggregation process on at least one of the first federated prediction models to obtain at least one target federated prediction model; S33: Determine a first target federated prediction model from at least one of the target federated prediction models according to the first patient portrait, and perform a first verification and distribution process on the first target federated prediction model.

8. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 1, characterized in that: The S4 includes the following sub-steps: S41: Obtaining a first confidence level from the first prediction information output by the first target federated prediction model, and determining a first human-machine collaborative decision based on the first confidence level; S42: Performing a first difference analysis process on the first prediction information to obtain at least one first dispute feature and at least one first weight adjustment suggestion; S43: Determine at least one first optimization parameter based on at least one first dispute feature and at least one first weight adjustment suggestion and feed it back to S2.

9. The method for predicting the effect of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to claim 1, characterized in that: The first patient profile is one of metabolism-dominant, immune-sensitive and gene-driven.

10. A system for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer, used to implement the method for predicting the effects of neoadjuvant chemotherapy and immunotherapy for esophageal cancer according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cancer survival prediction method and device, equipment and storage medium

    CN118136191A

  • Medical decision-oriented multi-modal data dynamic fusion and labeling method and system

    CN119377894A

  • Real-time health risk prediction method and system based on dynamic knowledge graph

    CN120280136A

  • Slow obstructive pulmonary disease patient leaving hospital monitoring system

    CN120376014A

  • Automatic prediction system for tumor chemoradiotherapy reaction

    CN120412908A