A multi-modal gastric cancer neoadjuvant therapy efficacy prediction method and device based on image and liquid biopsy genomic data, equipment and storage medium

By employing a multimodal prediction method combining imaging and liquid biopsy genomic data, the problem of non-responsiveness in neoadjuvant therapy for gastric cancer has been addressed. This approach achieves highly accurate and low-risk efficacy prediction, enabling personalized treatment plans.

CN122177497APending Publication Date: 2026-06-09BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL
Filing Date
2026-03-23
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Currently, approximately 30%-50% of gastric cancer patients do not respond or respond poorly to neoadjuvant therapy, delaying the optimal treatment time, increasing surgical difficulty and medical costs, and traditional tissue biopsy carries clinical risks.

Method used

A multimodal prediction method using imaging and liquid biopsy genomic data was employed. Imaging data reflects tumor morphological characteristics, while liquid biopsy genomic data reflects biological characteristics. Combined with clinical baseline data, first and second prediction models were trained to predict the efficacy of neoadjuvant therapy for gastric cancer. A modal modeling and post-fusion strategy was adopted to avoid the risks associated with tissue biopsy.

Benefits of technology

It significantly improves the accuracy of efficacy prediction, reduces clinical risks, and facilitates rapid deployment in existing clinical processes to provide personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177497A_ABST
    Figure CN122177497A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of multi-modal gastric cancer neoadjuvant therapy curative effect prediction method, device and equipment based on image and liquid biopsy genomic data and storage medium, it is related to gastric cancer neoadjuvant therapy technical field, the method comprises: obtaining the image data of tumor focus of gastric cancer patient before receiving gastric cancer neoadjuvant therapy, liquid biopsy genomic data and clinical baseline data;First key feature is extracted from image data and input into first prediction model to obtain first curative effect;Second key feature is extracted from genomic data and input into second prediction model to obtain second curative effect;According to first curative effect and second curative effect, the predicted curative effect of gastric cancer patient after receiving gastric cancer neoadjuvant therapy is determined.The above-mentioned three kinds of information of the gastric cancer patient to be predicted are input into the curative effect prediction model trained based on the information of sample gastric cancer patient by the above scheme, i.e.The curative effect of the gastric cancer patient to be predicted after receiving gastric cancer neoadjuvant therapy can be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neoadjuvant therapy technology for gastric cancer, and in particular to a method, device, equipment, and storage medium for predicting the efficacy of multimodal neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data. Background Technology

[0002] The comprehensive treatment model of neoadjuvant therapy (chemotherapy, targeted therapy, immunotherapy, etc.) combined with radical surgery and adjuvant therapy for gastric cancer is recognized by clinical treatment guidelines such as the NCCN (National Comprehensive Cancer Network) and ESMO (European Society for Medical Oncology) and is a preferred treatment option, representing the current standard model for cancer treatment. Neoadjuvant therapy can reduce tumor size and lower tumor stage, creating conditions for subsequent surgical resection or radical treatment.

[0003] However, in clinical practice, approximately 30%-50% of gastric cancer patients do not respond or respond poorly to neoadjuvant therapy, which not only delays the optimal treatment time but may also increase the difficulty of surgery due to drug side effects, reduce patients' quality of life, and increase medical costs. Therefore, accurately predicting the efficacy of neoadjuvant therapy for gastric cancer patients before surgery, screening potential beneficiaries, and developing individualized treatment plans are key needs for improving the effectiveness of tumor treatment. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, device, and storage medium for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data, so as to predict the efficacy of neoadjuvant therapy for gastric cancer patients. The specific technical solution is as follows:

[0005] In a first aspect, embodiments of this application provide a method for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data, the method comprising:

[0006] Before a patient with gastric cancer to be predicted receives neoadjuvant therapy, imaging data of the tumor lesions, liquid biopsy genomic data, and clinical baseline data of the patient to be predicted are obtained.

[0007] The tumor lesion region is segmented from the image data, and various imaging features of the tumor lesion region are extracted. A first preset type of feature is extracted from each imaging feature as the first key feature.

[0008] Genomic features of various dimensions are extracted from the genomic data, and features of a second preset type are extracted from the genomic features and the clinical baseline data as second key features;

[0009] The first key feature is input into a pre-trained first prediction model to obtain the first therapeutic effect predicted by the first prediction model, wherein the first prediction model is trained based on the imaging features of tumor lesions of the sample gastric cancer patients and the therapeutic effect of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0010] The second key feature is input into a pre-trained second prediction model to obtain the second efficacy predicted by the second prediction model. The second prediction model is trained based on the data features of liquid biopsy genomic data and clinical baseline data of the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0011] Based on the first efficacy and the second efficacy, the predicted efficacy of neoadjuvant therapy for gastric cancer in the patient to be predicted for gastric cancer is determined.

[0012] Optionally, the second prediction model is a random forest model. After inputting the second key feature into the pre-trained second prediction model and obtaining the second therapeutic effect predicted by the second prediction model, the method further includes:

[0013] An interpretability analysis was performed on the random forest model to determine the contribution value of each second key feature to the second therapeutic effect. Features whose contribution values ​​meet preset conditions were identified as key influencing features.

[0014] The treatment plan corresponding to the key influencing features is queried in the pre-set knowledge base of neoadjuvant therapy for gastric cancer, and is used as the recommended neoadjuvant therapy plan for the patient to be predicted to have gastric cancer.

[0015] Optionally, the genomic data is ctDNA sequencing data, and the extraction of genomic features from the genomic data in various dimensions includes:

[0016] The ctDNA sequencing data were subjected to unique molecular identifier deduplication to obtain the first sequence set;

[0017] The first sequence set is compared with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome;

[0018] A repetitive sequence masking tool is used to mark repetitive sequence regions in the human reference genome, and sequences located in the repetitive sequence regions in the first sequence set are excluded based on the location information to obtain a second sequence set;

[0019] Sequences corresponding to the same genomic location in the second sequence set are grouped together, and the mutation frequency at each genomic location is calculated based on the number of sequences carrying mutations in each group and the total number of sequences.

[0020] Genomic locations with mutation frequencies not less than a preset threshold are selected as candidate mutation sites, and genomic features of various dimensions are extracted based on these candidate mutation sites.

[0021] Optionally, the genomic features include at least one of: driver gene mutation status, hematologic malignancy mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism.

[0022] Optionally, the first preset type and the second preset type can be predetermined in the following manner:

[0023] To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy;

[0024] The tumor lesion region is segmented from the sample image data, and various sample imaging features of the tumor lesion region are extracted. Genomic features of various dimensions are extracted from the sample genomic data.

[0025] Based on the treatment efficacy of the gastric cancer patients in the sample, the Lasso algorithm is used to determine the coefficients of the imaging features of each sample. The feature types of the imaging features of the samples with non-zero coefficients are determined as the first preset type. The coefficients are used to characterize whether the features contribute to the prediction of efficacy.

[0026] Based on the treatment efficacy of the gastric cancer patients in the sample, the random forest algorithm is used to calculate the importance score of each feature in the sample's genomic features and clinical baseline data in predicting efficacy. A preset number of features are selected based on the importance scores, and the feature type of the selected features is determined as the second preset type.

[0027] Optionally, the first and second prediction models can be trained in the following manner:

[0028] To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy;

[0029] The sample tumor lesion region is segmented from the sample image data, and each sample imaging feature of the sample tumor lesion region is extracted. Then, a first preset type of feature is extracted from each sample imaging feature as the first sample key feature.

[0030] Extract sample genomic features from various dimensions from the sample genomic data, and extract features of a second preset type from each sample genomic feature and the sample clinical baseline data as second sample key features;

[0031] The key features of the first sample are input into the first model to be trained. The parameters of the first model to be trained are adjusted according to the difference between the predicted efficacy output by the first model to be trained and the treatment efficacy, until the first model to be trained converges, and the first prediction model is obtained.

[0032] The key features of the second sample are input into the second training model. The parameters of the second training model are adjusted according to the difference between the predicted efficacy output by the second training model and the treatment efficacy, until the second training model converges, thus obtaining the second prediction model.

[0033] Secondly, embodiments of this application provide a multimodal neoadjuvant therapy efficacy prediction device for gastric cancer based on imaging and liquid biopsy genomic data, the device comprising:

[0034] The data acquisition module is used to acquire imaging data of tumor lesions, liquid biopsy genomic data and clinical baseline data of the gastric cancer patient to be predicted before the patient receives neoadjuvant therapy for gastric cancer.

[0035] The feature extraction module is used to segment the tumor lesion region from the image data, extract various imaging features of the tumor lesion region, and extract a first preset type of feature from each imaging feature as a first key feature; extract genomic features of various dimensions from the genomic data, and extract a second preset type of feature from each genomic feature and the clinical baseline data as a second key feature;

[0036] The data processing module is used to input the first key feature into a pre-trained first prediction model to obtain the first efficacy predicted by the first prediction model, wherein the first prediction model is trained based on the imaging characteristics of tumor lesions in the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer; and to input the second key feature into a pre-trained second prediction model to obtain the second efficacy predicted by the second prediction model, wherein the second prediction model is trained based on the data characteristics of liquid biopsy genomic data and clinical baseline data of the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0037] The prediction module is used to determine the predicted efficacy of neoadjuvant therapy for gastric cancer in the patient to be predicted, based on the first efficacy and the second efficacy.

[0038] Optionally, the second prediction model is a random forest model, and the apparatus further includes:

[0039] The key influence feature determination module is used to perform interpretability analysis on the random forest model, determine the contribution value of each second key feature to the second therapeutic effect, and determine the features whose contribution values ​​meet the preset conditions as key influence features.

[0040] The treatment plan recommendation module is used to query the treatment plan corresponding to the key influencing features in a preset knowledge base of neoadjuvant therapy plans for gastric cancer, and use it as the recommended neoadjuvant therapy plan for the gastric cancer patient to be predicted.

[0041] Optionally, the genomic data is ctDNA sequencing data, and the feature extraction module is specifically used for:

[0042] The ctDNA sequencing data were subjected to unique molecular identifier deduplication to obtain the first sequence set;

[0043] The first sequence set is compared with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome;

[0044] A repetitive sequence masking tool is used to mark repetitive sequence regions in the human reference genome, and sequences located in the repetitive sequence regions in the first sequence set are excluded based on the location information to obtain a second sequence set;

[0045] Sequences corresponding to the same genomic location in the second sequence set are grouped together, and the mutation frequency at each genomic location is calculated based on the number of sequences carrying mutations in each group and the total number of sequences.

[0046] Genomic locations with mutation frequencies not less than a preset threshold are selected as candidate mutation sites, and genomic features of various dimensions are extracted based on these candidate mutation sites.

[0047] Optionally, the genomic features include at least one of: driver gene mutation status, hematologic malignancy mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism.

[0048] Optionally, the first preset type and the second preset type can be predetermined in the following manner:

[0049] To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy;

[0050] The tumor lesion region is segmented from the sample image data, and various sample imaging features of the tumor lesion region are extracted. Genomic features of various dimensions are extracted from the sample genomic data.

[0051] Based on the treatment efficacy of the gastric cancer patients in the sample, the Lasso algorithm is used to determine the coefficients of the imaging features of each sample. The feature types of the imaging features of the samples with non-zero coefficients are determined as the first preset type. The coefficients are used to characterize whether the features contribute to the prediction of efficacy.

[0052] Based on the treatment efficacy of the gastric cancer patients in the sample, the random forest algorithm is used to calculate the importance score of each feature in the sample's genomic features and clinical baseline data in predicting efficacy. A preset number of features are selected based on the importance scores, and the feature type of the selected features is determined as the second preset type.

[0053] Optionally, the first and second prediction models can be trained in the following manner:

[0054] To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy;

[0055] The sample tumor lesion region is segmented from the sample image data, and each sample imaging feature of the sample tumor lesion region is extracted. Then, a first preset type of feature is extracted from each sample imaging feature as the first sample key feature.

[0056] Extract sample genomic features from various dimensions from the sample genomic data, and extract features of a second preset type from each sample genomic feature and the sample clinical baseline data as second sample key features;

[0057] The key features of the first sample are input into the first model to be trained. The parameters of the first model to be trained are adjusted according to the difference between the predicted efficacy output by the first model to be trained and the treatment efficacy, until the first model to be trained converges, and the first prediction model is obtained.

[0058] The key features of the second sample are input into the second training model. The parameters of the second training model are adjusted according to the difference between the predicted efficacy output by the second training model and the treatment efficacy, until the second training model converges, thus obtaining the second prediction model.

[0059] Thirdly, embodiments of this application provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0060] Memory, used to store computer programs;

[0061] When a processor executes a program stored in memory, it implements any of the methods described in the first aspect above.

[0062] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the methods described in the first aspect above.

[0063] Beneficial effects of the embodiments in this application:

[0064] In the solution provided in this application embodiment, imaging data can reflect the morphological characteristics of the tumor, liquid biopsy genomic data can reflect the biological characteristics of the tumor, and clinical baseline data can provide basic information of gastric cancer patients. By combining the above three types of information of the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer, a first prediction model and a second prediction model are trained. In the actual prediction stage, the three types of information of the gastric cancer patient to be predicted are input into the efficacy prediction model, so as to predict the efficacy of the gastric cancer patient after receiving neoadjuvant therapy for gastric cancer.

[0065] Furthermore, predictions based on multimodal data fusion significantly improve accuracy compared to methods based on single-modal data. This application also utilizes liquid biopsy instead of traditional tissue biopsy, requiring only peripheral blood samples to obtain genomic data, thus avoiding the clinical risks associated with tissue biopsy, such as bleeding and perforation risks associated with invasive procedures like gastroscopy in gastric cancer. Moreover, this application employs a multimodal modeling and post-fusion strategy, allowing the first and second prediction models to be trained and optimized independently. Imaging detection and genome sequencing can be completed separately at different times and on different platforms, without requiring modification or joint debugging of detection equipment, facilitating rapid deployment within existing clinical workflows.

[0066] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0068] Figure 1 A flowchart illustrating a method for predicting the efficacy of multimodal neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data, as provided in this application embodiment;

[0069] Figure 2 A schematic diagram of a specific process for predicting the efficacy of multimodal neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data, as provided in this application embodiment;

[0070] Figure 3 For based on Figure 1 A flowchart illustrating one method for extracting genomic features in the embodiment shown;

[0071] Figure 4 For based on Figure 1 A flowchart illustrating a method for filtering feature types in the embodiment shown;

[0072] Figure 5 For based on Figure 1 A flowchart illustrating a model training method in the embodiment shown;

[0073] Figure 6 This is a schematic diagram of a multimodal neoadjuvant therapy efficacy prediction device for gastric cancer based on image and liquid biopsy genomic data, provided in an embodiment of this application.

[0074] Figure 7This is a schematic diagram of an electronic device provided based on an embodiment of this application. Detailed Implementation

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0076] To predict the efficacy of neoadjuvant therapy for gastric cancer in patients, this application provides a multimodal method, device, electronic device, computer-readable storage medium, and computer program product for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data. The following is a description of the multimodal method for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data provided in this application.

[0077] The multimodal neoadjuvant therapy efficacy prediction method for gastric cancer based on imaging and liquid biopsy genomic data provided in this application can be applied to electronic devices such as computers and servers. This application does not specifically limit this method; for clarity, it will be referred to as an electronic device below.

[0078] A multimodal method for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data includes: acquiring imaging data, liquid biopsy genomic data, and clinical baseline data of gastric cancer patients before they receive neoadjuvant therapy; segmenting the tumor lesion region from the imaging data, extracting various imaging features of the tumor lesion region, and extracting features of a first preset type from each imaging feature as a first key feature; extracting genomic features of various dimensions from the genomic data, and extracting features of a second preset type from each genomic feature and clinical baseline data as a second key feature; and inputting the first key feature into a pre-trained first prediction model. The process involves obtaining the first predictive efficacy predicted by a first predictive model, which is trained based on the imaging characteristics of tumor lesions in sample gastric cancer patients and the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy; inputting a second key feature into a pre-trained second predictive model to obtain the second predictive efficacy predicted by the second predictive model, which is trained based on the data characteristics of liquid biopsy genomic data and clinical baseline data of sample gastric cancer patients and the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy; and using a weighted fusion method based on the first and second efficacy to determine the multimodal predicted efficacy of gastric cancer patients after receiving neoadjuvant therapy.

[0079] In the solution provided in this embodiment, imaging data can reflect the morphological characteristics of the tumor, liquid biopsy genomic data can reflect the biological characteristics of the tumor, and clinical baseline data can provide basic information about gastric cancer patients. By combining the above three types of information from the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy, a first prediction model and a second prediction model are trained. In the actual prediction stage, the three types of information of the gastric cancer patient to be predicted are input into the efficacy prediction model, so that the efficacy of the gastric cancer patient after receiving neoadjuvant therapy can be predicted.

[0080] Furthermore, predictions based on multimodal data fusion significantly improve accuracy compared to methods based on single-modal data. This application also utilizes liquid biopsy instead of traditional tissue biopsy, requiring only peripheral blood samples to obtain genomic data, thus avoiding the clinical risks associated with tissue biopsy, such as bleeding and perforation risks associated with invasive procedures like gastroscopy in gastric cancer. Moreover, this application employs a multimodal modeling and post-fusion strategy, allowing the first and second prediction models to be trained and optimized independently. Imaging detection and genome sequencing can be completed separately at different times and on different platforms, without requiring modification or joint debugging of detection equipment, facilitating rapid deployment within existing clinical workflows.

[0081] For example, the following will combine Figure 1 The specific steps in the method for predicting the efficacy of neoadjuvant therapy for gastric cancer provided in the embodiments of this application will be described in detail, such as... Figure 1 As shown, the method for predicting the efficacy of neoadjuvant therapy for gastric cancer may include steps S101-S106:

[0082] S101: Before a patient with gastric cancer is given neoadjuvant therapy, imaging data of the tumor lesion, liquid biopsy genomic data, and clinical baseline data are obtained.

[0083] When it is necessary to predict the efficacy of neoadjuvant therapy for gastric cancer in patients with unpredictable gastric cancer, imaging data of tumor lesions, liquid biopsy genomic data, and clinical baseline data of the patients before receiving neoadjuvant therapy can be collected, and then the collected data can be entered into an electronic device.

[0084] "Imaging data of tumor lesions" refers to medical images containing the tumor lesion area of ​​gastric cancer patients, such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), PET-CT (Positron Emission Tomography-Computed Tomography), etc.

[0085] "Liquid biopsy genomic data" refers to ctDNA (circulating tumor DNA) sequencing data obtained by collecting peripheral blood samples. Compared with traditional tissue biopsy, it has the advantages of being non-invasive, dynamically monitorable, and able to overcome tumor heterogeneity.

[0086] "Clinical baseline data" includes the patient's age, biopsy pathology type, pathological stage, Lauren classification, etc.

[0087] S102, the tumor lesion area is segmented from the image data, various imaging features of the tumor lesion area are extracted, and features of a first preset type are extracted from the various imaging features as the first key features.

[0088] After acquiring image data, the electronic device can first segment the tumor lesion region from the original image data, which will then be used as the feature extraction region for subsequent feature extraction. The method for segmenting the tumor lesion region can employ existing image segmentation algorithms; this application does not specifically limit this approach. For example, the U-Net++ deep learning model can be used to automatically segment the tumor lesion region, followed by manual verification, to accurately segment the tumor lesion region.

[0089] After segmentation, the electronic device can further extract various imaging features (radiomics features) of the tumor lesion area, which may include first-order statistical features (such as mean and standard deviation), morphological features (such as lesion volume and maximum diameter), texture features (such as gray-level co-occurrence matrix entropy), etc.

[0090] Furthermore, for each extracted imaging feature, the electronic device can extract only those feature types that are pre-determined to contribute to efficacy prediction, i.e., the "first preset type features," as the first key features. The first preset type is pre-determined through analysis of sample data; the specific determination method will be discussed later. Figure 4 The embodiments shown are described in detail, and will not be repeated here.

[0091] Step S103: Extract genomic features from various dimensions of genomic data, and extract features of a second preset type from each genomic feature and clinical baseline data as the second key feature.

[0092] For genomic data, electronic devices can extract genomic features from various dimensions. The specific methods for extracting these genomic features will be discussed later. Figure 3The implementation details are provided in the examples section and will not be repeated here. After extracting the various genomic features, the electronic device can extract only those feature types that are pre-determined to contribute to efficacy prediction, i.e., "features of the second preset type," from the various genomic features and clinical baseline data, based on pre-determined feature types, as the second key features. The second preset type is pre-determined through analysis of sample data; the specific determination method will be discussed later. Figure 4 The embodiments shown are described in detail, and will not be repeated here.

[0093] Step S104: Input the first key feature into the pre-trained first prediction model to obtain the first therapeutic effect predicted by the first prediction model.

[0094] A first predictive model can be pre-trained based on the imaging characteristics of tumor lesions in sample gastric cancer patients and the treatment efficacy of these patients after neoadjuvant therapy. This model is specifically designed to predict treatment efficacy based on imaging characteristics and can then be deployed on electronic devices. The specific training method for the model will be discussed later. Figure 5 The embodiments are described in detail in the section on examples, and will not be repeated here.

[0095] Once the electronic device identifies the first key feature, it can input it into the first prediction model and obtain the output of the first prediction model as the predicted first therapeutic effect. The specific form of the first therapeutic effect can be set according to actual usage needs, and this application embodiment does not impose specific limitations on it. For example, the first therapeutic effect can be a binary classification result of "whether it is effective," or a numerical result of a "score," where a higher score indicates a higher response rate to neoadjuvant therapy for gastric cancer. It can also be a binary classification result of "whether it is effective" plus the confidence level of the binary classification result, or a numerical result of a "score" plus the confidence level of the numerical result.

[0096] The first prediction model is a model such as CNN (Convolutional Neural Network), ResNet3D (3D Residual Network), or DenseNet3D (3D Dense Convolutional Network). This application does not limit the specific type of model.

[0097] Step S105: Input the second key feature into the pre-trained second prediction model to obtain the second therapeutic effect predicted by the second prediction model.

[0098] A second predictive model can be pre-trained based on liquid biopsy genomic data and clinical baseline data from gastric cancer patients, as well as the treatment efficacy of these patients after neoadjuvant therapy. This second model is specifically designed to predict treatment efficacy based on genomic and clinical baseline data and can then be deployed on electronic devices. The specific training method for this model will be discussed later. Figure 5 The embodiments are described in detail in the section on examples, and will not be repeated here.

[0099] Once the electronic device identifies the second key feature, it can input it into the second prediction model and obtain the output of the second prediction model as the predicted second therapeutic effect. The specific form of the second therapeutic effect can be set according to actual usage requirements, and this application embodiment does not impose specific limitations on it. However, in a preferred embodiment, the specific form of the second therapeutic effect needs to be consistent with that of the first therapeutic effect so that the two can be subsequently fused.

[0100] The second prediction model is a random forest, XGBoost (Extreme Gradient Boosting), LightGBM (Light Gradient Boosting Machine), etc. The specific type of model is not limited in the embodiments of this application.

[0101] Step S106: Determine the predicted efficacy of neoadjuvant therapy for gastric cancer in patients with gastric cancer based on the first and second efficacy outcomes.

[0102] After obtaining the first and second therapeutic effects, electronic devices can use fusion strategies such as maximum value fusion, average value fusion, or weighted fusion to fuse the first and second therapeutic effects, and then determine the fusion result as the predicted therapeutic effect of the gastric cancer patient after receiving neoadjuvant therapy.

[0103] • In one possible implementation, based on the method for predicting the efficacy of neoadjuvant therapy for gastric cancer in the above embodiments, the method may further include: performing interpretability analysis on the random forest model to determine the contribution value of each second key feature to the second efficacy, and identifying features whose contribution values ​​meet preset conditions as key influencing features; querying the treatment plan corresponding to the key influencing features in a preset knowledge base of neoadjuvant therapy for gastric cancer as a recommended neoadjuvant therapy plan for the gastric cancer patient to be predicted.

[0104] In the solution provided in this embodiment, after predicting the second therapeutic effect using a second prediction model (random forest), further interpretability analysis is performed to quantify the contribution of each feature to the prediction result, thereby identifying the key influencing features that play a decisive role in predicting the therapeutic effect for current gastric cancer patients. These features have clear biological significance and can help clinicians understand the basis of the prediction results. Based on this, combined with a pre-set treatment plan knowledge base, individualized treatment recommendations can be matched for gastric cancer patients, forming a complete closed loop from "prediction" to "decision-making".

[0105] For example, the following will combine Figure 2 The specific steps in the method for predicting the efficacy of neoadjuvant therapy for gastric cancer provided in the embodiments of this application will be described in detail, such as... Figure 2 As shown, the above-mentioned method for predicting the efficacy of neoadjuvant therapy for gastric cancer may further include steps S107-S108:

[0106] S107. Perform interpretability analysis on the random forest model to determine the contribution value of each second key feature to the second therapeutic effect, and identify the features whose contribution values ​​meet the preset conditions as key influencing features.

[0107] After the random forest model outputs the second therapeutic effect, interpretability analysis tools can be used to perform "interpretability analysis" on the random forest model to determine which features played a key role in the prediction process when the random forest model predicted the "second therapeutic effect". Existing tools can be used for interpretability analysis, and this application does not specifically limit them. For example, the SHAP (SHapley Additive exPlanations) algorithm or LIME (Local Interpretable Model-agnostic Explanations) algorithm can be used.

[0108] Taking the analysis of a random forest model using the SHAP algorithm as an example. For each gastric cancer patient to be predicted, the SHAP algorithm calculates the contribution value of each secondary key feature to the secondary efficacy. The sum of these contribution values ​​plus the baseline value equals the final predicted value. The magnitude of the contribution value indicates the degree to which the feature drives the prediction result; a positive contribution indicates that the prediction result is pushed in the "response" direction, and a negative contribution indicates that it is pushed in the "non-response" direction. Features with the highest contribution values ​​(e.g., absolute values ​​greater than a certain threshold, or ranked in the top N) are identified as the "key influencing features" for that gastric cancer patient. For example, for a gastric cancer patient, the analysis results show that the contribution value of TP53 mutation status is +0.30, the contribution value of bTMB value is +0.20, and the contribution value of ERBB2 amplification status is +0.15, while the contribution values ​​of other features are close to 0. These three are the key influencing features.

[0109] S108: Search the pre-defined knowledge base of neoadjuvant therapy for gastric cancer for treatment plans corresponding to key influencing features, and use them as recommended neoadjuvant therapy plans for patients with gastric cancer to be predicted.

[0110] The treatment knowledge base is pre-built based on authoritative sources such as clinical practice guidelines, drug instructions, and clinical trial evidence. It stores the correspondence between key features and recommended treatment options. For example, "HER2 positive" corresponds to "anti-HER2 targeted therapy," "high TMB" corresponds to "immune checkpoint inhibitors may be effective," and "high PD-L1 expression" corresponds to "immunotherapy." After obtaining the key influencing features of the gastric cancer patient to be predicted, these are used as query conditions to search the knowledge base and match the corresponding treatment options.

[0111] In one implementation, the prediction results that the electronic device ultimately outputs to the user may include: recommended treatment plans, predicted treatment efficacy results, and key influencing features, thereby providing clinicians with intuitive and actionable decision-making support.

[0112] One possible implementation, the extraction of genomic features from genomic data in various dimensions, may include: performing unique molecular identifier deduplication on ctDNA sequencing data to obtain a first sequence set; comparing the first sequence set with a human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome; using a repetitive sequence masking tool to mark repetitive sequence regions in the human reference genome, and excluding sequences located in repetitive sequence regions from the first sequence set according to the position information to obtain a second sequence set; grouping sequences in the second sequence set corresponding to the same genomic position into a group, and calculating the mutation frequency of each genomic position based on the number of sequences carrying mutations in each group and the total number of sequences; selecting genomic positions with mutation frequencies not less than a preset threshold as candidate mutation sites, and extracting genomic features in various dimensions based on the candidate mutation sites.

[0113] In the solution provided in this embodiment, ctDNA has a low content in blood, and the sequencing process suffers from amplification bias, repetitive sequence interference, and background noise. If the raw sequencing data is used directly, it will be difficult to extract the true tumor signal from the massive noise. This technical solution eliminates amplification noise through UMI deduplication, locates the tumor through sequence alignment, eliminates interference through repetitive sequence masking, and captures the true signal through low-frequency mutation screening. Ultimately, it obtains high-quality, high-reliability candidate mutation sites, laying a solid foundation for subsequent genomic feature extraction and efficacy prediction.

[0114] For example, the following will combine Figure 3 The specific implementation method of step S103 above will be described, such as... Figure 3As shown, when the electronic device performs step S103, it can specifically achieve this through steps S301-S305:

[0115] S301, the ctDNA sequencing data is deduplicated by unique molecular identifiers to obtain the first sequence set.

[0116] During ctDNA sequencing library construction, each original DNA molecule is given a unique nucleotide sequence, known as a UMI (Unique Molecular Identifier). In the subsequent PCR (Polymerase Chain Reaction) amplification step, all copies of the same original molecule will carry the same UMI. After sequencing, a large number of sequences are generated, many of which are repetitive copies from the same original molecule. If these repetitive sequences are not processed, the calculated mutation frequency will deviate significantly from the true value. This step utilizes UMI information to merge sequences with the same UMI, retaining only one representative sequence for each UMI, thereby eliminating repetitive sequences generated by PCR amplification and obtaining a first set of sequences representing the original DNA molecule. The deduplication process can be implemented using existing deduplication tools; this application does not specifically limit this step. For example, it can be performed using tools like umi-tools (Unique Molecular Identifier-tools).

[0117] S302, compare the first sequence set with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome.

[0118] The first sequence set obtained after UMI deduplication is a series of DNA fragment sequences, but the specific locations of these fragments on the genome are not yet known. Therefore, electronic devices can compare these sequences with a human reference genome to find the best matching position for each sequence. The alignment results include the genomic coordinates (chromosome number, start position, end position), alignment direction, and other information for each sequence. The alignment process can be implemented using existing alignment tools, and this application does not specifically limit this; for example, the BWA-MEM (Burrows-Wheeler Aligner - Maximum ExactMatch) alignment tool can be used to complete this step.

[0119] S303 uses a repetitive sequence masking tool to mark repetitive sequence regions in the human reference genome, and excludes sequences located in repetitive sequence regions from the first sequence set based on location information to obtain the second sequence set.

[0120] The human genome contains numerous repetitive sequence regions, such as Alu and LINE sequences. These regions are highly similar, making accurate alignment difficult and prone to false positive mutations. Therefore, electronic devices can use repetitive sequence masking tools (such as RepeatMasker) to pre-mark all repetitive sequence regions in the human reference genome. Then, based on the location information obtained in step S302, the electronic device can exclude sequences located within repetitive sequence regions from the first sequence set, obtaining the second sequence set.

[0121] S304. Group the sequences in the second sequence set that correspond to the same genomic location into one group, and calculate the mutation frequency of each genomic location based on the number of sequences carrying mutations in each group and the total number of sequences.

[0122] Sequences aligned to the same genomic location in the second sequence set are grouped together. For each location, the total number of sequences in the group and the number of sequences carrying mutations are counted. The mutation frequency is defined as the percentage of sequences carrying mutations out of the total number of sequences at that location. For example, if a location is covered by 100 sequences, and 3 of them differ from the reference genome at a certain base, then the mutation frequency at that location is 3%. This step transforms the raw sequence information into a mutation frequency value for each location.

[0123] S305 identifies genomic locations with mutation frequencies not less than a preset threshold as candidate mutation sites and extracts genomic features of various dimensions based on these candidate mutation sites.

[0124] The proportion of tumor-derived DNA in ctDNA is extremely low; therefore, real tumor mutations typically exist at a low frequency, generally between 0.1% and 5%. The inherent error rate during sequencing is usually below 0.1%. Therefore, a preset threshold (e.g., 0.1%) is set to distinguish between real signals and background noise: genomic locations with a mutation frequency of at least 0.1% are retained as candidate mutation sites, while those below this threshold are discarded as noise. Finally, based on these reliable candidate mutation sites, genomic features of various dimensions are extracted for subsequent model training. The method for extracting genomic features from candidate mutation sites can be found in existing technologies, and will not be elaborated further in this embodiment.

[0125] Through practical verification, the inventors have demonstrated that the meticulous processing of the five steps in this embodiment can stably detect low-frequency mutation signals from mixed blood samples, increasing the detection rate of low-frequency mutation features by 20%, and providing high-quality input for subsequent efficacy prediction.

[0126] In one possible implementation, the aforementioned genomic features may include at least one of: driver gene mutation status, hematologic malignancy mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism.

[0127] The solution provided in this embodiment comprehensively characterizes the molecular biological features of tumors by extracting multidimensional genomic features from candidate mutation sites. These features cover different levels, from individual genes to the entire genome, signaling pathways, and genetic background, providing rich and comprehensive input information for subsequent efficacy prediction models and helping the models to more accurately capture the correlation between efficacy and molecular features.

[0128] For example, the meaning and extraction method of each type of feature are explained below:

[0129] Driver gene mutation status: refers to the mutation or amplification status of key genes closely related to tumor development and progression. The extraction method is as follows: check whether candidate mutation sites contain mutations of these genes; if so, the mutation status feature of that gene is set to 1; otherwise, it is set to 0. This is a binary feature. For example, when applying the proposed solution to gastric cancer, the focus is on common gastric cancer driver genes such as TP53, PIK3CA, and ERBB2.

[0130] Hematologic tumor mutation burden (bTMB): refers to the number of mutations detected per million bases, reflecting the overall mutation level of the tumor. The extraction method is as follows: count the total number of candidate mutation sites and divide by the total size of the sequencing region (in Mb). For example, in this embodiment, the sequencing panel covers 825 tumor-related genes, with a total region size of approximately 1.5 Mb. If 15 candidate mutation sites are detected, the bTMB value is 10 muts / Mb. This is a continuous value feature.

[0131] Genomic alteration score (bFGA): This refers to the proportion of altered bases out of the total number of bases in the sequenced region, and is also an indicator reflecting the overall degree of variation in the genome. Unlike bTMB, which counts the number of mutations, bFGA considers the range of bases covered by the mutation. The extraction method is as follows: count the number of bases covered by all candidate mutation sites and divide by the total number of bases in the sequenced region.

[0132] Copy number amplification (CNV): refers to an abnormal increase in the copy number of certain genes, such as ERBB2 gene amplification (i.e., HER2 positive). The extraction method involves comparing the sequencing depth of the target gene region with that of the control region to infer the presence of amplification. If amplification is present, the amplification status feature of the gene is set to 1; otherwise, it is set to 0. This is also a type of binary feature.

[0133] KEGG signaling pathways refer to the functional pathways involved in candidate mutation sites. The extraction method involves mapping candidate mutation sites to a KEGG pathway database and identifying which pathways have mutations. For example, TP53 mutations indicate p53 signaling pathway abnormalities, while PIK3CA mutations indicate PI3K-AKT pathway activation. This information can serve as indicators of pathway activity scores or abnormal pathway states.

[0134] Single nucleotide polymorphisms (SNPs) are genetically polymorphic sites present in the genome. Unlike somatic mutations, SNPs are inherited and may affect the activity of drug-metabolizing enzymes, drug target affinity, etc., thereby affecting efficacy. The extraction method involves screening for pre-defined SNP sites among candidate mutation sites and recording their genotypes (e.g., AA, AG, GG).

[0135] One possible implementation involves determining the first and second preset types in the following manner: acquiring sample image data of tumor lesions from gastric cancer patients before receiving neoadjuvant therapy, genomic data of liquid biopsy samples, and sample clinical baseline data, as well as acquiring the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy; segmenting the sample tumor lesion region from the sample image data, extracting each sample imaging feature of the sample tumor lesion region, and extracting sample genomic features of each dimension from the sample genomic data; determining the coefficients of each sample imaging feature using the Lasso algorithm based on the treatment efficacy of the sample gastric cancer patients, and determining the feature type of the sample imaging features with non-zero coefficients as the first preset type, wherein the coefficients are used to characterize whether the feature contributes to predicting the efficacy; calculating the importance score of each feature in the sample genomic features and sample clinical baseline data in predicting efficacy using the random forest algorithm based on the treatment efficacy of the sample gastric cancer patients, selecting a preset number of features based on the importance scores, and determining the feature type of the selected features as the second preset type.

[0136] In this embodiment, the solution employs both the Lasso and Random Forest algorithms for feature selection on sample data labeled with therapeutic efficacy. This allows for an objective and data-driven determination of which feature types are truly valuable for efficacy prediction. The Lasso algorithm uses L1 regularization to compress the coefficients of irrelevant features to zero; features with non-zero coefficients are identified as key feature types contributing to image prediction. The Random Forest algorithm calculates the importance score of each feature during decision tree construction, selecting the most important feature types for genomic and clinical prediction. This feature selection method based on machine learning algorithms avoids the subjectivity and limitations of manual feature selection, ensuring that subsequent prediction models are built on the optimal feature set.

[0137] For example, the following will combine Figure 4The determination method of the first preset type and the second preset type is explained in detail, such as Figure 4 As shown, the first preset type and the second preset type can be determined through steps S401-S404:

[0138] S401, obtain sample imaging data of tumor lesions, genomic data of liquid biopsy samples, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy for gastric cancer, and obtain the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0139] To determine which feature types are truly valuable for predicting treatment efficacy, we can collect imaging data of tumor lesions, genomic data of liquid biopsy samples, and clinical baseline data of gastric cancer patients before they receive neoadjuvant therapy. Simultaneously, we collect efficacy assessment results of these gastric cancer patients after receiving neoadjuvant therapy (i.e., the treatment efficacy of the sample gastric cancer patients after neoadjuvant therapy) as efficacy labels.

[0140] S402, segment the tumor lesion area from the sample image data, extract the various sample imaging features of the tumor lesion area, and extract the sample genomic features of various dimensions from the sample genomic data.

[0141] The methods for extracting imaging features and genomic features have been described in detail in the foregoing embodiments and will not be repeated here.

[0142] S403, based on the treatment efficacy of the sample gastric cancer patients, the Lasso algorithm is used to determine the coefficients of each sample's imaging features, and the feature types of the sample imaging features with non-zero coefficients are determined as the first preset type, wherein the coefficients are used to characterize whether the features contribute to the prediction of efficacy.

[0143] The Lasso (Least Absolute Shrinkage and Selection Operator) algorithm is a linear model with L1 regularization. During training, it automatically shrinks the coefficients of irrelevant or redundant features to zero, while retaining the non-zero coefficients of important features.

[0144] In practice, the sample imaging features are used as input, and the efficacy labels of the sample gastric cancer patients are used as output to train the Lasso model. After training, the coefficients corresponding to each feature are checked. Features with non-zero coefficients indicate that these features contribute to predicting efficacy. The feature types of these features (such as "entropy", "volume", "texture", etc.) are recorded as the "first preset type".

[0145] S404. Based on the treatment efficacy of the gastric cancer patients in the sample, the random forest algorithm is used to calculate the importance score of each feature in the genomic features and clinical baseline data of each sample when predicting efficacy. Based on the importance score, a preset number of features are selected, and the feature type of the selected features is determined as the second preset type.

[0146] The random forest algorithm consists of multiple decision trees, and each tree selects the optimal feature for node splitting during its construction. Based on this mechanism, an "importance score" for each feature in the entire forest can be calculated; a higher score indicates a greater contribution of the feature to the prediction.

[0147] In practice, the genomic and clinical features of the samples are used as input, and the efficacy labels of the gastric cancer patients are used as output to train a random forest model. After training, the importance score of each feature is calculated and sorted from high to low. Based on a preset number, the top-ranked features are selected, and their feature types are recorded as the "second preset type".

[0148] By using the first and second preset types determined in the above steps, features can be extracted directly according to these types when training the prediction model and when making predictions for gastric cancer patients based on the prediction model, without the need for repeated screening processes.

[0149] One possible implementation involves obtaining a first prediction model and a second prediction model in advance as follows: acquiring sample image data of tumor lesions from gastric cancer patients before receiving neoadjuvant therapy, genomic data of liquid biopsy samples, and sample clinical baseline data, as well as acquiring the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy; segmenting the sample tumor lesion region from the sample image data, extracting various sample imaging features from the sample tumor lesion region, and extracting features of a first preset type from each sample imaging feature as first sample key features; extracting sample genomic features of various dimensions from the sample genomic data, and extracting features of a second preset type from each sample genomic feature and sample clinical baseline data as second sample key features; inputting the first sample key features into the first training model, adjusting the parameters of the first training model according to the difference between the predicted efficacy and treatment efficacy output by the first training model until the first training model converges, thus obtaining the first prediction model; inputting the second sample key features into the second training model, adjusting the parameters of the second training model according to the difference between the predicted efficacy and treatment efficacy output by the second training model until the second training model converges, thus obtaining the second prediction model.

[0150] In this embodiment, two prediction models are trained separately on labeled sample data, enabling independent modeling of two different modalities of data. The first prediction model is trained based on image features, capturing the morphological information and spatial heterogeneity of tumors; the second prediction model is trained based on genomic and clinical features, capturing the molecular biological information of tumors and baseline information of gastric cancer patients. The two models are trained and optimized independently, and then integrated in the post-fusion stage. This modal modeling strategy ensures the full extraction of features from each modality of data and facilitates the independent operation of different detection platforms in clinical applications.

[0151] For example, the following will combine Figure 5 The training methods for the first and second prediction models are explained in detail, such as... Figure 5 As shown, the first and second prediction models can be obtained through training steps S501-S505:

[0152] S501, obtain sample imaging data of tumor lesions, genomic data of liquid biopsy samples, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy for gastric cancer, and obtain the treatment efficacy of sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0153] Similar to step S401 above, it will not be repeated here.

[0154] S502, the sample tumor lesion area is segmented from the sample image data, the sample tumor lesion area is extracted, and the first preset type of features is extracted from the sample image features as the first sample key features.

[0155] S503, extract sample genomic features from sample genomic data of various dimensions, and extract features of a second preset type from each sample genomic feature and sample clinical baseline data as the second sample key features.

[0156] Steps S502-S503 are the same as the feature extraction methods in steps S102-S103 above, except that the data targeted by feature extraction is different, which will not be described again here.

[0157] S504, input the key features of the first sample into the first training model, adjust the parameters of the first training model according to the difference between the predicted efficacy and the treatment efficacy output by the first training model, until the first training model converges, and obtain the first prediction model.

[0158] The key features of the first sample are used as input, and the efficacy labels of the gastric cancer patients in the sample are used as supervision signals. These are then input into the first model to be trained. During training, the model outputs the predicted efficacy for each sample, compares it with the true efficacy label, and calculates the loss function (such as cross-entropy loss). Then, the weight parameters of each layer in the network are adjusted through the backpropagation algorithm to minimize the loss function. This process is iterated repeatedly until the model converges (e.g., the loss function no longer decreases, or the model's performance on the validation set reaches a preset standard), resulting in the trained first prediction model.

[0159] S505, input the key features of the second sample into the second training model, adjust the parameters of the second training model according to the difference between the predicted efficacy and the treatment efficacy output by the second training model, until the second training model converges, and obtain the second prediction model.

[0160] The key features of the second sample are used as input, and the efficacy labels of the gastric cancer patients in the sample are used as supervision signals. These are then input into the second training model for training. During training, the model constructs multiple decision trees, each growing based on different subsets of samples and features. The final prediction is obtained through voting or averaging. By adjusting hyperparameters such as the number of trees and the maximum depth, the model performance is optimized, leading to model convergence and ultimately obtaining the trained second prediction model.

[0161] In one implementation, after obtaining two trained models, their outputs can be post-fused on a validation dataset using different weighted fusion weights to determine the accuracy corresponding to each fusion weight, thereby determining the optimal weighted fusion weight between the outputs of the two models.

[0162] It should be noted here that... Figure 4 and Figure 5 The embodiments shown can be executed in combination. Specifically, steps S401-S404 are executed first to obtain relevant data of gastric cancer patients and extract the first key feature of the sample (i.e., the imaging features of the sample with a coefficient that is not zero) and the second key feature of the sample (i.e., a preset number of features selected according to the importance score). Then, on the one hand, the first preset type and the second preset type can be recorded so that the first key feature and the second key feature can be quickly extracted in the subsequent prediction process. On the other hand, based on the first key feature of the sample and the second key feature of the sample obtained in steps S401-S404, the model training process can continue to be executed, that is, steps S504-S505 can continue to be executed.

[0163] To better understand the technical solution provided in this application, a specific application example will be used to provide an overall explanation below.

[0164] For example, in a specific implementation scenario, this application collected 268 gastric cancer patients as a sample dataset for model training and validation. The specific data structure is as follows:

[0165] Imaging data: Enhanced CT images of 268 patients with gastric cancer were collected before treatment.

[0166] Liquid biopsy genomic data: Peripheral blood ctDNA samples were collected from 115 gastric cancer patients before treatment and sequenced using a targeted sequencing panel (covering 825 tumor-related genes). The sequencing depth was ≥30000× and the effective depth was ≥3000×.

[0167] Clinical baseline data: Information such as age, pathological type, pathological stage, and Lauren classification of gastric cancer patients is collected.

[0168] Efficacy label: The efficacy was assessed by imaging examination and pathological biopsy 3 months after treatment, and 141 patients were identified as responders and 127 as nonresponders. Among the gastric cancer patients with liquid biopsy data, there were 47 responders and 68 nonresponders.

[0169] During the model training phase, the following procedure should be followed:

[0170] 1. Data preprocessing and feature extraction.

[0171] Image data: The gastric cancer tumor lesion area was automatically segmented using the U-Net++ model and manually reviewed; about 200 dimensions of imaging features were extracted, including first-order statistical features (mean, standard deviation, etc.), morphological features (volume, maximum diameter, etc.), and texture features (gray-level co-occurrence matrix entropy, etc.).

[0172] ctDNA data: UMI deduplication was performed using umi-tools, BWA-MEM was used to align with the reference genome, RepeatMasker was used to mask repetitive sequence regions, and low-frequency mutation sites with a mutation frequency ≥0.1% were screened; genomic features of multiple dimensions, including driver gene mutation status (TP53, PIK3CA, etc.), bTMB, bFGA, CNV (such as ERBB2 amplification), KEGG signaling pathway, and SNPs, were extracted.

[0173] Clinical data: four key features including age of preservation, biopsy pathological type, pathological stage, and Lauren classification.

[0174] 2. Feature filtering.

[0175] The Lasso algorithm was used to select features with non-zero coefficients from approximately 200 image features, and their feature types were determined as the first preset type.

[0176] The random forest algorithm was used to calculate feature importance scores from genomic and clinical features, and the nine features with the highest importance scores were selected and their feature types were determined as the second preset type.

[0177] 3. Model training.

[0178] The first prediction model: The selected image features were input into the CNN for training. The training set consisted of 230 cases (125 responders and 105 non-responders), and the validation set consisted of 38 cases (16 responders and 22 non-responders). The accuracy of the model on the validation set after training was 72.4%.

[0179] The second prediction model: The selected genomic features and clinical features were input into a random forest for training. The training set consisted of 77 cases (31 responders and 46 non-responders), and the validation set consisted of 38 cases (16 responders and 22 non-responders). The accuracy of the model on the validation set after training was 78.9%.

[0180] 4. Model fusion.

[0181] The outputs of the two models were then fused, and the weighted fusion weights were determined through validation set optimization. The fused model achieved a prediction accuracy of 84.2% on the validation set, which is significantly better than the single model.

[0182] In the predictive application phase, for a new patient with gastric cancer to be predicted, pre-treatment imaging data, ctDNA data, and clinical baseline data are collected following the same procedure. After the same data preprocessing and feature extraction, the pre-determined first and second pre-determined type features are directly extracted and input into the trained first and second prediction models, respectively, to obtain the first and second efficacy outcomes. These are then fused according to the pre-determined weighted fusion weights to obtain the final predicted efficacy outcome for the gastric cancer patient. Simultaneously, SHAP interpretability analysis is performed on the random forest model to identify the key influencing features that contribute most to the second efficacy outcome (such as TP53 mutation, high bTMB, and HER2 positivity). Corresponding treatment recommendations (such as "anti-HER2 combined with chemoimmunotherapy for gastric cancer") are queried from a pre-determined treatment regimen knowledge base. Finally, a clinical report containing efficacy prediction scores, confidence levels, key influencing features, and recommended treatment regimens is generated.

[0183] This embodiment verifies the effectiveness of the technical solution of this application, with a prediction accuracy of 84.2%, and the detection rate of low-frequency mutations is increased by 20% through ctDNA optimization, which fully demonstrates the technical effect and clinical practical value of this application.

[0184] Corresponding to the above-mentioned method for predicting the efficacy of neoadjuvant therapy for gastric cancer based on image and liquid biopsy genomic data, this application also provides a device for predicting the efficacy of neoadjuvant therapy for gastric cancer based on image and liquid biopsy genomic data. The following is a description of the device for predicting the efficacy of neoadjuvant therapy for gastric cancer based on image and liquid biopsy genomic data provided in this application.

[0185] like Figure 6 As shown, a multimodal neoadjuvant therapy efficacy prediction device for gastric cancer based on imaging and liquid biopsy genomic data includes:

[0186] The data acquisition module 601 is used to acquire imaging data of tumor lesions, liquid biopsy genomic data and clinical baseline data of patients with gastric cancer to be predicted before they receive neoadjuvant therapy for gastric cancer.

[0187] The feature extraction module 602 is used to segment the tumor lesion region from the image data, extract various imaging features of the tumor lesion region, and extract a first preset type of feature from each imaging feature as the first key feature; extract genomic features of various dimensions from the genomic data, and extract a second preset type of feature from each genomic feature and clinical baseline data as the second key feature;

[0188] The data processing module 603 is used to input a first key feature into a pre-trained first prediction model to obtain the first therapeutic effect predicted by the first prediction model, wherein the first prediction model is trained based on the imaging characteristics of tumor lesions in the sample gastric cancer patients and the therapeutic effect of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer; and to input a second key feature into a pre-trained second prediction model to obtain the second therapeutic effect predicted by the second prediction model, wherein the second prediction model is trained based on the data characteristics of liquid biopsy genomic data and clinical baseline data of the sample gastric cancer patients and the therapeutic effect of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer.

[0189] The prediction module 604 is used to determine the predicted efficacy of neoadjuvant therapy for gastric cancer in patients to be predicted, based on the first efficacy and the second efficacy.

[0190] As one embodiment of this application, the second prediction model is a random forest model, and the apparatus further includes:

[0191] The key impact feature determination module is used to perform interpretability analysis on the random forest model, determine the contribution value of each second key feature to the second therapeutic effect, and determine the features whose contribution values ​​meet the preset conditions as key impact features.

[0192] The treatment plan recommendation module is used to search for treatment plans corresponding to key influencing features in a pre-set knowledge base of neoadjuvant therapy plans for gastric cancer, and to recommend neoadjuvant therapy plans for patients with gastric cancer to be predicted.

[0193] As one embodiment of this application, the genomic data is ctDNA sequencing data, and the feature extraction module is specifically used for:

[0194] The ctDNA sequencing data were deduplicated by removing unique molecular identifiers to obtain the first sequence set;

[0195] The first sequence set is compared with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome;

[0196] The repetitive sequence masking tool was used to mark the repetitive sequence regions in the human reference genome, and the sequences located in the repetitive sequence regions in the first sequence set were excluded based on the location information to obtain the second sequence set;

[0197] Sequences corresponding to the same genomic location in the second sequence set are grouped together, and the mutation frequency at each genomic location is calculated based on the number of sequences carrying mutations in each group and the total number of sequences.

[0198] Genomic locations with mutation frequencies not less than a preset threshold are selected as candidate mutation sites, and genomic features of various dimensions are extracted based on these candidate mutation sites.

[0199] As one embodiment of this application, the genomic features include at least one of the following: driver gene mutation status, hematologic tumor mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism.

[0200] As one embodiment of this application, the first preset type and the second preset type are predetermined in the following manner:

[0201] To obtain tumor lesion imaging data, liquid biopsy genomic data, and clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of gastric cancer patients after receiving neoadjuvant therapy.

[0202] The tumor lesion region of the sample is segmented from the sample image data, and various sample imaging features of the sample tumor lesion region are extracted. The sample genomic features of various dimensions are extracted from the sample genomic data.

[0203] Based on the treatment efficacy of the sample gastric cancer patients, the Lasso algorithm is used to determine the coefficients of each sample's imaging features. The feature types of the sample imaging features with non-zero coefficients are determined as the first preset type. The coefficients are used to characterize whether the features contribute to the prediction of efficacy.

[0204] Based on the treatment efficacy of the gastric cancer patients in the sample, the random forest algorithm was used to calculate the importance score of each feature in the genomic features and clinical baseline data of each sample in predicting efficacy. Based on the importance score, a preset number of features were selected, and the feature type of the selected features was determined as the second preset type.

[0205] As one embodiment of this application, the first prediction model and the second prediction model are trained in the following manner:

[0206] To obtain tumor lesion imaging data, liquid biopsy genomic data, and clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of gastric cancer patients after receiving neoadjuvant therapy.

[0207] The tumor lesion area of ​​the sample is segmented from the sample image data, and the imaging features of each sample tumor lesion area are extracted. The first preset type of features are extracted from each sample imaging feature as the first sample key features.

[0208] Extract sample genomic features from various dimensions of sample genomic data, and extract second-preset type features from each sample genomic feature and sample clinical baseline data as second sample key features;

[0209] The key features of the first sample are input into the first model to be trained. The parameters of the first model to be trained are adjusted according to the difference between the predicted efficacy and the treatment efficacy output by the first model to be trained until the first model to be trained converges, thus obtaining the first prediction model.

[0210] The key features of the second sample are input into the second training model. The parameters of the second training model are adjusted according to the difference between the predicted efficacy and the treatment efficacy output by the second training model until the second training model converges, thus obtaining the second prediction model.

[0211] This application also provides an electronic device, such as... Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0212] Memory 703 is used to store computer programs;

[0213] When the processor 701 executes the program stored in the memory 703, it implements the multimodal neoadjuvant therapy efficacy prediction method for gastric cancer based on image and liquid biopsy genomic data as described in any of the above embodiments.

[0214] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0215] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0216] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0217] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0218] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of the multimodal neoadjuvant therapy efficacy prediction method for gastric cancer based on image and liquid biopsy genomic data described in any of the above embodiments.

[0219] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the multimodal neoadjuvant therapy efficacy prediction methods for gastric cancer based on image and liquid biopsy genomic data in the above embodiments.

[0220] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0221] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0222] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0223] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A multimodal method for predicting the efficacy of neoadjuvant therapy for gastric cancer based on imaging and liquid biopsy genomic data, characterized in that, The method includes: Before a patient with gastric cancer to be predicted receives neoadjuvant therapy, imaging data of the tumor lesions, liquid biopsy genomic data, and clinical baseline data of the patient to be predicted are obtained. The tumor lesion region is segmented from the image data, and various imaging features of the tumor lesion region are extracted. A first preset type of feature is extracted from each imaging feature as the first key feature. Genomic features of various dimensions are extracted from the genomic data, and features of a second preset type are extracted from the genomic features and the clinical baseline data as second key features; The first key feature is input into a pre-trained first prediction model to obtain the first therapeutic effect predicted by the first prediction model, wherein the first prediction model is trained based on the imaging features of tumor lesions of the sample gastric cancer patients and the therapeutic effect of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer. The second key feature is input into a pre-trained second prediction model to obtain the second efficacy predicted by the second prediction model. The second prediction model is trained based on the data features of liquid biopsy genomic data and clinical baseline data of the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer. Based on the first efficacy and the second efficacy, the predicted efficacy of neoadjuvant therapy for gastric cancer in the patient to be predicted for gastric cancer is determined.

2. The method according to claim 1, characterized in that, The second prediction model is a random forest model. After inputting the second key feature into the pre-trained second prediction model and obtaining the second therapeutic effect predicted by the second prediction model, the method further includes: An interpretability analysis was performed on the random forest model to determine the contribution value of each second key feature to the second therapeutic effect. Features whose contribution values ​​meet preset conditions were identified as key influencing features. The treatment plan corresponding to the key influencing features is queried in the pre-set knowledge base of neoadjuvant therapy for gastric cancer, and is used as the recommended neoadjuvant therapy plan for the patient to be predicted to have gastric cancer.

3. The method according to claim 1, characterized in that, The genomic data is ctDNA sequencing data, and the extraction of genomic features from the genomic data in various dimensions includes: The ctDNA sequencing data were subjected to unique molecular identifier deduplication to obtain the first sequence set; The first sequence set is compared with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome; A repetitive sequence masking tool is used to mark repetitive sequence regions in the human reference genome, and sequences located in the repetitive sequence regions in the first sequence set are excluded based on the location information to obtain a second sequence set; Sequences corresponding to the same genomic location in the second sequence set are grouped together, and the mutation frequency at each genomic location is calculated based on the number of sequences carrying mutations in each group and the total number of sequences. Genomic locations with mutation frequencies not less than a preset threshold are selected as candidate mutation sites, and genomic features of various dimensions are extracted based on these candidate mutation sites.

4. The method according to claim 3, characterized in that, The genomic features include at least one of the following: driver gene mutation status, hematologic tumor mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism.

5. The method according to any one of claims 1-4, characterized in that, The first preset type and the second preset type are determined in advance in the following manner: To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy; The tumor lesion region is segmented from the sample image data, and various sample imaging features of the tumor lesion region are extracted. Genomic features of various dimensions are extracted from the sample genomic data. Based on the treatment efficacy of the gastric cancer patients in the sample, the Lasso algorithm is used to determine the coefficients of the imaging features of each sample. The feature types of the imaging features of the samples with non-zero coefficients are determined as the first preset type. The coefficients are used to characterize whether the features contribute to the prediction of efficacy. Based on the treatment efficacy of the gastric cancer patients in the sample, the random forest algorithm is used to calculate the importance score of each feature in the sample's genomic features and clinical baseline data in predicting efficacy. A preset number of features are selected based on the importance scores, and the feature type of the selected features is determined as the second preset type.

6. The method according to any one of claims 1-4, characterized in that, The first and second prediction models were trained in the following manner: To obtain tumor lesion image data, liquid biopsy sample genomic data, and sample clinical baseline data of gastric cancer patients before receiving neoadjuvant therapy, and to obtain the treatment efficacy of the gastric cancer patients after receiving neoadjuvant therapy; The sample tumor lesion region is segmented from the sample image data, and each sample imaging feature of the sample tumor lesion region is extracted. Then, a first preset type of feature is extracted from each sample imaging feature as the first sample key feature. Extract sample genomic features from various dimensions from the sample genomic data, and extract features of a second preset type from each sample genomic feature and the sample clinical baseline data as second sample key features; The key features of the first sample are input into the first model to be trained. The parameters of the first model to be trained are adjusted according to the difference between the predicted efficacy output by the first model to be trained and the treatment efficacy, until the first model to be trained converges, and the first prediction model is obtained. The key features of the second sample are input into the second training model. The parameters of the second training model are adjusted according to the difference between the predicted efficacy output by the second training model and the treatment efficacy, until the second training model converges, thus obtaining the second prediction model.

7. A multimodal neoadjuvant therapy efficacy prediction device for gastric cancer based on imaging and liquid biopsy genomic data, characterized in that, The device includes: The data acquisition module is used to acquire imaging data of tumor lesions, liquid biopsy genomic data and clinical baseline data of the gastric cancer patient to be predicted before the patient receives neoadjuvant therapy for gastric cancer. The feature extraction module is used to segment the tumor lesion region from the image data, extract various imaging features of the tumor lesion region, and extract a first preset type of feature from each imaging feature as a first key feature; extract genomic features of various dimensions from the genomic data, and extract a second preset type of feature from each genomic feature and the clinical baseline data as a second key feature; The data processing module is used to input the first key feature into a pre-trained first prediction model to obtain the first efficacy predicted by the first prediction model, wherein the first prediction model is trained based on the imaging characteristics of tumor lesions in the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer; and to input the second key feature into a pre-trained second prediction model to obtain the second efficacy predicted by the second prediction model, wherein the second prediction model is trained based on the data characteristics of liquid biopsy genomic data and clinical baseline data of the sample gastric cancer patients and the treatment efficacy of the sample gastric cancer patients after receiving neoadjuvant therapy for gastric cancer. The prediction module is used to determine the predicted efficacy of neoadjuvant therapy for gastric cancer in the patient to be predicted, based on the first efficacy and the second efficacy.

8. The apparatus according to claim 7, characterized in that, The second prediction model is a random forest model. The device further includes: a key influence feature determination module, used to perform interpretability analysis on the random forest model, determine the contribution value of each second key feature to the second efficacy, and determine the features whose contribution values ​​meet preset conditions as key influence features; and a treatment plan recommendation module, used to query the treatment plan corresponding to the key influence feature in a preset knowledge base of neoadjuvant treatment plans for gastric cancer, and use it as the recommended neoadjuvant treatment plan for the gastric cancer patient to be predicted. And / or, The genomic data is ctDNA sequencing data. The feature extraction module is specifically used for: performing unique molecular identifier deduplication on the ctDNA sequencing data to obtain a first sequence set; comparing the first sequence set with the human reference genome to obtain the position information of each sequence in the first sequence set on the human reference genome; using a repetitive sequence masking tool to mark repetitive sequence regions in the human reference genome, and excluding sequences in the first sequence set located in the repetitive sequence regions according to the position information to obtain a second sequence set; grouping sequences in the second sequence set corresponding to the same genomic position into a group, and calculating the mutation frequency of each genomic position based on the number of sequences carrying mutations and the total number of sequences in each group; selecting genomic positions with mutation frequencies not less than a preset threshold as candidate mutation sites, and extracting genomic features of various dimensions based on the candidate mutation sites; And / or, The genomic features include at least one of the following: driver gene mutation status, hematologic malignancy mutation burden, genomic alteration fraction, copy number amplification, KEGG signaling pathway, and single nucleotide polymorphism; And / or, The first and second preset types are determined in advance as follows: Image data of tumor lesions from gastric cancer patients before neoadjuvant therapy, genomic data of liquid biopsy samples, and clinical baseline data of samples are acquired; the treatment efficacy of the gastric cancer patients after neoadjuvant therapy is also acquired. Tumor lesion regions are segmented from the image data, and various imaging features of the tumor lesion regions are extracted. Genomic features of various dimensions are extracted from the genomic data. Based on the treatment efficacy of the gastric cancer patients, the Lasso algorithm is used to determine the coefficients of each imaging feature. The feature types of the imaging features with non-zero coefficients are determined as the first preset type, where the coefficients characterize the contribution of the feature to predicting efficacy. Based on the treatment efficacy of the gastric cancer patients, the random forest algorithm is used to calculate the importance score of each genomic feature and each feature in the clinical baseline data when predicting efficacy. A preset number of features are selected based on the importance scores, and the feature types of the selected features are determined as the second preset type. And / or, The first and second prediction models are trained as follows: Image data of tumor lesions from gastric cancer patients before neoadjuvant therapy, genomic data of liquid biopsy samples, and clinical baseline data of the samples are acquired; the treatment efficacy of the gastric cancer patients after neoadjuvant therapy is also acquired. The tumor lesion region is segmented from the image data, and various imaging features of the tumor lesion region are extracted. A first preset type of feature is extracted from each of the imaging features as the first key feature. Genomic features of various dimensions are extracted from the genomic data, and a second preset type of feature is extracted from each genomic feature and the clinical baseline data as the second key feature. The first key feature is input into the first training model, and the parameters of the first training model are adjusted according to the difference between the predicted efficacy and the treatment efficacy output by the first training model until convergence is achieved, thus obtaining the first prediction model. The second key feature is input into the second training model, and the parameters of the second training model are adjusted according to the difference between the predicted efficacy and the treatment efficacy output by the second training model until convergence is achieved, thus obtaining the second prediction model.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.