Treatment method and device for prognosis prediction and treatment scheme of hypopharyngeal cancer induced chemotherapy and medium
By training preset models combined with CT images and clinical text information, we predict the risk of chemotherapy induced by hypopharyngeal carcinoma and recommend personalized plans, which solves the problem that specific chemotherapy plans cannot be recommended in the prior art, and improves the matching degree of the plans and the consistency of the patient's response.
Patent Information
- Application Number
- CN202510506759.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art cannot recommend specific chemotherapy regimens for hypopharyngeal carcinoma induction based on actual conditions, resulting in patients' condition delays due to tumor progression and large differences in chemotherapy response.
By obtaining multimodal information, training preset models, including CT images and clinical text information, using deep learning networks and multimodal feature fusion, predict chemotherapy risks and recommend target solutions.
It improves the matching degree between chemotherapy plans and actual situations, helps identify patients who respond to chemotherapy, reduces side effects of chemotherapy, and improves quality of life.
Smart Images

Figure CN120565048A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data processing technology, and in particular to a method, device, and medium for predicting the prognosis and treating a hypopharyngeal cancer by induction chemotherapy. Background Art
[0002] Head and neck squamous cell carcinoma (HNSCC) is the sixth most common malignant tumor worldwide. Hypopharyngeal squamous cell carcinoma (HPSCC), while relatively rare, is often diagnosed at an advanced stage due to its hidden location and aggressive nature. The traditional surgical treatment for locally advanced HPSCC is total laryngectomy, but this procedure compromises the patient's voice function and can cause difficult postoperative complications such as pharyngeal fistulas. Induction chemotherapy (IC), a laryngeal preservation strategy, has been increasingly used as the initial treatment for locally advanced HPSCC to improve patients' quality of life. However, responses to induction chemotherapy vary among patients. Patients who do not respond to induction chemotherapy face the financial burden of medication and the potential for chemotherapy toxicity and side effects, and may even face delayed treatment due to tumor progression. Therefore, early identification of patients with HPSCC who are responsive to induction chemotherapy is crucial. Currently, the method of predicting the efficacy of induction chemotherapy for hypopharyngeal cancer by constructing a model based on pre-induction chemotherapy imaging images of hypopharyngeal cancer patients can only assist in determining whether to perform induction chemotherapy, but cannot recommend a specific induction chemotherapy regimen based on actual conditions.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a method, device and medium for predicting the prognosis and treatment plan of induction chemotherapy for hypopharyngeal cancer, which can improve the matching degree between the recommended chemotherapy plan and the actual situation.
[0005] To achieve the above objectives, one aspect of the present invention provides a method for predicting the prognosis of hypopharyngeal cancer after induction chemotherapy and for processing a treatment plan, the method comprising the following steps:
[0006] Obtaining historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, the historical multimodal information comprising a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen;
[0007] Training a preset model using the historical multimodal information, the preset model including a prognosis prediction model for induction chemotherapy of hypopharyngeal cancer and a risk assessment model for treatment plans;
[0008] Determine the corresponding scoring cutoff value after the preset model training is completed;
[0009] Acquiring current multimodal information of a currently untreated hypopharyngeal cancer patient before induction chemotherapy, the current multimodal information comprising a second CT image and second clinical text information;
[0010] Inputting the current multimodal information into the trained preset model to predict the current risk value;
[0011] A target recommendation regimen for induction chemotherapy is determined based on the current risk value and the score cutoff value.
[0012] In some embodiments, the preset model includes a CT image preprocessing module, a CT image feature extraction module, a multimodal feature screening module, and a multimodal feature fusion and prediction module.
[0013] In some embodiments, the training of a preset model using the historical multimodal information includes:
[0014] Preprocessing the first CT image by the CT image preprocessing module to obtain a region of interest mask file corresponding to the first CT image and a third CT image;
[0015] performing a first feature extraction on the region of interest mask file and the third CT image by the CT image feature extraction module to obtain a first CT radiomics feature;
[0016] Performing a second feature extraction on the region of interest mask file and the third CT image using a preset deep learning network model through the CT image feature extraction module to obtain a deep learning feature;
[0017] screening the first CT radiomics feature using the multimodal feature screening module to obtain a second CT radiomics feature, and screening the first clinical text information according to the overall survival period to obtain third clinical text information;
[0018] Inputting the second CT imaging omics feature, the deep learning feature, and the third clinical text information into the multimodal feature fusion and prediction module to perform overall risk prediction of induction chemotherapy prognosis and independent risk prediction of induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen, to obtain an overall risk value of first induction chemotherapy prognosis and an independent risk value of first induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen; and calculating an overall loss function of the multimodal feature fusion and prediction module;
[0019] The model parameters of the multimodal feature fusion and prediction module are adjusted according to the overall loss function.
[0020] In some embodiments, the preset deep learning network model includes an input layer, an initial convolutional layer, a first-stage processing layer, a second-stage processing layer, a third-stage processing layer, a fourth-stage processing layer, a global average pooling layer, a fully connected layer, and an output layer;
[0021] The input layer receives the region of interest mask file and the third CT image;
[0022] The first-stage processing layer includes three residual attention blocks, each of which is obtained by connecting a residual block and a channel attention mechanism module in series;
[0023] The second stage processing layer includes 4 residual attention blocks;
[0024] The third stage processing layer includes 6 residual attention blocks;
[0025] The fourth stage processing layer includes 3 residual attention blocks;
[0026] The input layer, the initial convolution layer, the first stage processing layer, the second stage processing layer, the third stage processing layer, the fourth stage processing layer, the global average pooling layer, the fully connected layer and the output layer are connected in series in sequence;
[0027] The output layer is used to output deep learning features.
[0028] In some embodiments, the multimodal feature fusion and prediction module includes a fully connected layer, a group-aware attention layer, a gated cross-modal fusion layer, an overall prediction head, a first group prediction head, a second group prediction head, a third group prediction head, and a fourth group prediction head;
[0029] The fully connected layer is used to receive the second CT radiomics feature, the deep learning feature and the third clinical text information;
[0030] The fully connected layer, the group-aware attention layer, and the gated cross-modal fusion layer are connected in series in sequence;
[0031] Input ends of the overall prediction head, the first group prediction head, the second group prediction head, the third group prediction head, and the fourth group prediction head are all connected to the output end of the gated cross-modal fusion layer.
[0032] In some embodiments, determining the corresponding score cutoff value after the preset model training is completed includes:
[0033] After determining that the multimodal feature fusion and prediction module in the preset model has completed training, generating a ROC curve according to the overall risk value of the first induction chemotherapy prognosis;
[0034] Calculate the Youden index corresponding to each point in the ROC curve;
[0035] The overall risk value of the first induction chemotherapy prognosis corresponding to the point with the largest Youden index among all the Youden indices is selected as the score cutoff value.
[0036] In some embodiments, determining a target recommended regimen for induction chemotherapy based on the current risk value and the score cutoff value comprises:
[0037] When the overall risk value of the current induction chemotherapy prognosis in the current risk value is less than the score cutoff value, obtaining all second induction chemotherapy prognosis independent risk values in the current risk value, each second induction chemotherapy prognosis independent risk value corresponding to one second induction chemotherapy regimen;
[0038] Rank all the independent risk values of the second induction chemotherapy prognosis;
[0039] A target recommendation regimen for the induction chemotherapy corresponding to the currently untreated hypopharyngeal cancer patient is generated based on the ranking result and the second induction chemotherapy prognostic independent risk value.
[0040] To achieve the above objectives, another aspect of the present invention provides a device for predicting the prognosis of hypopharyngeal cancer induction chemotherapy and processing a treatment plan, the device comprising:
[0041] A first module is configured to obtain historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, the historical multimodal information including a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen;
[0042] A second module is used to train a preset model using the historical multimodal information, wherein the preset model includes a prognosis prediction model for induction chemotherapy of hypopharyngeal cancer and a risk assessment model for treatment plans;
[0043] The third module is used to determine the corresponding score cutoff value after the preset model training is completed;
[0044] A fourth module is configured to obtain current multimodal information of a currently untreated hypopharyngeal cancer patient before induction chemotherapy, the current multimodal information including a second CT image and second clinical text information;
[0045] A fifth module is configured to input the current multimodal information into the trained preset model to predict a current risk value;
[0046] The sixth module is used to determine the target recommendation scheme of induction chemotherapy based on the current risk value and the score cutoff value.
[0047] To achieve the above objectives, another aspect of the present application provides a computer device, including:
[0048] at least one processor;
[0049] at least one memory for storing at least one program;
[0050] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0051] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.
[0052] The embodiments of the present application include at least the following beneficial effects: The present application provides a method, device, and medium for predicting the prognosis and processing treatment plans of induction chemotherapy for hypopharyngeal cancer. The scheme obtains historical multimodal information of multiple historical hypopharyngeal cancer patients before induction chemotherapy, including a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the historically first-treated hypopharyngeal cancer patients, and the overall survival period corresponding to the first induction chemotherapy regimen. Then, a preset model is trained using the historical multimodal information. After determining the corresponding score cutoff value after the preset model training is completed, the current multimodal information of the current first-treated hypopharyngeal cancer patient before induction chemotherapy is input into the trained preset model to predict the current risk value. Then, based on the current risk value and the score cutoff value, a target recommended regimen for induction chemotherapy is determined, so that the determined target recommended regimen is closer to the actual situation, thereby improving the matching degree between the recommended chemotherapy regimen and the actual situation. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a flow chart of a method for predicting the prognosis and treating a hypopharyngeal cancer by induction chemotherapy provided in an embodiment of the present application;
[0054] Figure 2 This is a schematic diagram of a module of a preset model provided in an embodiment of the present application;
[0055] Figure 3 3D SE-Resnet-50 deep learning model module provided in the embodiment of the present application;
[0056] Figure 4 is a module schematic diagram of a multimodal feature fusion and prediction module provided in an embodiment of the present application;
[0057] Figure 5 This is a flowchart of the preset model provided in the embodiment of the present application for recommending solutions;
[0058] Figure 6 Schematic diagram of the structure of a device for predicting the prognosis and treating the hypopharyngeal cancer by induction chemotherapy provided in an embodiment of the present application;
[0059] Figure 7 Schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application.
[0061] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0062] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" in the context of the present invention, and "at least one" or "at least one" includes one, two or more, "plurality" or "any one" includes two or more, "each" or "each one" in the context of the present invention, and "any" or "any one
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0064] Before describing the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:
[0065] Induction chemotherapy (also known as neoadjuvant chemotherapy) is a systemic chemotherapy given before surgery or radiotherapy to reduce tumor size, relieve symptoms, and create more favorable conditions for subsequent local treatment.
[0066] The prognosis of induction chemotherapy refers to a comprehensive assessment of disease progression, treatment effect, and long-term survival of patients after completing induction chemotherapy (neoadjuvant chemotherapy).
[0067] Head and neck squamous cell carcinoma (HNSCC) is the sixth most common malignant tumor worldwide. Hypopharyngeal squamous cell carcinoma (HPSCC), while relatively rare, is often diagnosed at an advanced stage due to its hidden location and aggressive nature. The traditional surgical treatment for locally advanced HPSCC is total laryngectomy, but this procedure compromises the patient's voice function and can cause difficult postoperative complications such as pharyngeal fistulas. Induction chemotherapy (IC), a laryngeal preservation strategy, has been increasingly used as the initial treatment for locally advanced HPSCC to improve patients' quality of life. However, responses to induction chemotherapy vary among patients. Patients who do not respond to induction chemotherapy face the financial burden of medication and the potential for chemotherapy toxicity and side effects, and may even face delayed treatment due to tumor progression. Therefore, early identification of patients with HPSCC who are responsive to induction chemotherapy is crucial. Currently, the main drugs used in induction chemotherapy for hypopharyngeal cancer include traditional chemotherapy drugs (paclitaxels, platinums, 5-FU), immunotherapies (PD-1 inhibitors), and targeted therapies (EGFR inhibitors). Commonly used induction chemotherapy regimens derived from these regimens include the TP regimen (paclitaxel + platinum), the TPF regimen (paclitaxel + platinum + 5-FU), the TPK regimen (paclitaxel + platinum + PD-1 inhibitors), and the TPR regimen (paclitaxel + platinum + EGFR inhibitors). Different regimens have different rates of laryngeal function preservation, pathological response rates, and survival benefits. However, currently, methods that use models constructed based on pre-induction chemotherapy imaging to predict the efficacy of hypopharyngeal cancer induction chemotherapy can only assist in determining whether induction chemotherapy should be performed, but cannot recommend specific induction chemotherapy regimens based on actual conditions.
[0068] In view of this, the embodiments of the present application provide a method, device and medium for predicting the prognosis and treatment plan of induction chemotherapy for hypopharyngeal cancer, which can effectively improve the matching degree between the recommended chemotherapy plan and the actual situation.
[0069] The methods for predicting the prognosis and treating a treatment plan for induction chemotherapy for hypopharyngeal cancer provided in the embodiments of the present application relate to the field of big data processing technology. The methods for predicting the prognosis and treating a treatment plan for hypopharyngeal cancer provided in the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, etc., but are not limited thereto. The server can be configured as an independent physical server, or as a server cluster or distributed system consisting of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application that implements the methods for predicting the prognosis and treating a treatment plan for hypopharyngeal cancer using induction chemotherapy, but is not limited to the above forms.
[0070] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0071] The following is a detailed description of the embodiments of the present application with reference to the accompanying drawings:
[0072] Figure 1 This is an optional flow chart of the method for predicting the prognosis and treating the hypopharyngeal cancer by induction chemotherapy provided in the embodiments of the present application. Figure 1 The method may include but is not limited to steps S110 to S160:
[0073] Step S110: Acquire historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, wherein the historical multimodal information includes a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen;
[0074] Step S120: training a preset model using historical multimodal information, wherein the preset model includes a prognosis prediction model for induction chemotherapy of hypopharyngeal cancer and a risk assessment model for treatment plans;
[0075] Step S130: determining the corresponding score cutoff value after the preset model training is completed;
[0076] Step S140: obtaining current multimodal information of a hypopharyngeal cancer patient undergoing initial treatment before induction chemotherapy, wherein the current multimodal information includes a second CT image and second clinical text information;
[0077] Step S150: input the current multimodal information into the trained preset model to predict the current risk value;
[0078] Step S160: Determine a target recommended regimen for induction chemotherapy based on the current risk value and the score cutoff value.
[0079] It is understood that the first and second CT images involved in this embodiment can both be portal venous phase enhanced 3D CT images of the head and neck, and these CT images can be stored in DICOM format. The first and second clinical text information can include, but are not limited to, age, gender, tumor differentiation grade, T stage, N stage, M stage, clinical stage (according to the 8th edition of the AJCC staging system), white blood cell count, red blood cell count, hemoglobin, platelet count, neutrophil count, lymphocyte count, monocyte count, eosinophil count, basophil count, alanine aminotransferase, aspartate aminotransferase, serum albumin, C-reactive protein, and procalcitonin, and inflammatory indicators thereof can be calculated, including the systemic immune inflammation index (= platelet count × neutrophil count / lymphocyte count), neutrophil / lymphocyte ratio, platelet / lymphocyte ratio, and lymphocyte / monocyte ratio. The first and second induction chemotherapy regimens can be one or more of the TP regimen, TPF regimen, TPK regimen, and TPR regimen. Overall survival (OS) can be recorded as the death event and death date if the patient dies during the follow-up period, or as the censored event and follow-up date if the patient survives.
[0080] After collecting historical multimodal data from several previously untreated hypopharyngeal cancer patients before induction chemotherapy, the data was divided into a training set and a test set in a 7:3 ratio. The training set was used to train the pre-defined models, including the prognosis prediction and treatment plan risk assessment model for hypopharyngeal cancer induction chemotherapy, while the test set was used to test the pre-defined models during the training process.
[0081] It is understandable that if Figure 2 As shown, the preset model of this embodiment may include but is not limited to a CT image preprocessing module, a CT image feature extraction module, a multimodal feature screening module, and a multimodal feature fusion and prediction module. Figure 2 The preset model structure shown in FIG. 1 is used in this embodiment to train the preset model using a training set in historical multimodal information. Specifically, the training process of this embodiment includes but is not limited to the following steps:
[0082] Preprocessing the first CT image by a CT image preprocessing module to obtain a region of interest mask file corresponding to the first CT image and a third CT image;
[0083] Performing a first feature extraction on the region of interest mask file and the third CT image using a CT image feature extraction module to obtain a first CT radiomics feature;
[0084] A CT image feature extraction module is used to extract second features from the region of interest mask file and the third CT image using a preset deep learning network model to obtain deep learning features;
[0085] The first CT radiomics feature is screened by the multimodal feature screening module to obtain a second CT radiomics feature, and the first clinical text information is screened according to the overall survival period to obtain a third clinical text information;
[0086] The second CT imaging omics features, the deep learning features, and the third clinical text information are input into the multimodal feature fusion and prediction module to predict the overall risk of induction chemotherapy prognosis and the independent risk of induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen, and the overall risk value of the first induction chemotherapy prognosis and the independent risk value of the first induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen are obtained; and the overall loss function of the multimodal feature fusion and prediction module is calculated;
[0087] The model parameters of the multimodal feature fusion and prediction modules are adjusted according to the overall loss function.
[0088] It is understood that the CT image preprocessing module preprocesses the first CT image by manually delineating a region of interest (ROI) in the hypopharyngeal tumor region within successive slices of the first CT image using 3D Slicer software to obtain a ROI mask file; and simultaneously performing image resampling, image intensity normalization, and image cropping on the first CT image to obtain a third CT image. Image resampling refers to uniformly resampling the voxel size of the first CT image to 1*1*1nm; image intensity normalization refers to limiting the HU value (Hounsfield unit) range of the first CT image to a certain range using the maximum-minimum normalization method; and image cropping refers to cropping out CT image slices that do not contain the manually delineated ROI.
[0089] It is understandable that the process of the CT image feature extraction module performing the first feature extraction on the region of interest mask file and the third CT image can be to extract low-order feature morphological features, intensity histogram features, texture features from the ROI area in the region of interest mask file through the PyRadiomics Python package, and convert the above low-order feature morphological features into high-order features after wavelet transform and Laplace Gaussian filter (LoG) conversion to form the first CT imaging omics feature. Among them, CT imaging omics features refer to quantifiable structural and functional information extracted from medical images through high-throughput technology, which is used to describe the disease state and assist clinical decision-making. The essence is to convert visual images into a computable feature space, and to explore potential diagnostic or predictive value through statistical analysis and machine learning models. The main feature types of CT imaging omics features include but are not limited to shape features, intensity features, texture features, topological features, etc.
[0090] The CT image feature extraction module uses a preset deep learning network model to extract the second feature from the region of interest mask file and the third CT image. In this embodiment, the preset deep learning network model can be a 3D SE-Resnet-50 deep learning model. Taking the 3D SE-Resnet-50 deep learning model as an example for feature extraction, this embodiment inputs the region of interest mask file and the third CT image into the 3D SE-Resnet-50 deep learning model to extract features from the region of interest mask file and the third CT image to obtain deep learning features.
[0091] Specifically, the 3D SE-Resnet-50 deep learning model can be a model obtained by introducing a 3D convolution operation and a SE (Squeeze-and-Excitation) attention module on the pre-trained Resnet-50 model framework. This model can be applied to process 3D CT image data, capture image feature information, and improve the model's ability to focus on key features. It is understandable that, if Figure 3 As shown, the preset deep learning network model includes an input layer, an initial convolution layer, a first-stage processing layer, a second-stage processing layer, a third-stage processing layer, a fourth-stage processing layer, a global average pooling layer, a fully connected layer and an output layer; the input layer receives the region of interest mask file and the third CT image; the first-stage processing layer includes 3 residual attention blocks, each of which is obtained by connecting a residual block and a channel attention (SE) mechanism module in series; the second-stage processing layer includes 4 residual attention blocks; the third-stage processing layer includes 6 residual attention blocks; the fourth-stage processing layer includes 3 residual attention blocks; the input layer, the initial convolution layer, the first-stage processing layer, the second-stage processing layer, the third-stage processing layer, the fourth-stage processing layer, the global average pooling layer, the fully connected layer and the output layer are connected in series in sequence; the output layer is used to output deep learning features. In this embodiment, the multiple residual attention blocks in each stage processing layer can be connected in series or in parallel, and can be adaptively adjusted according to actual conditions. The residual attention block can be an SE attention module introduced after the residual block.
[0092] It is understandable that the process of the multimodal feature screening module processing the first CT imaging omics feature and the first clinical text can be to perform Z-score normalization, Spearman correlation analysis and LASSO-Cox regression on all the first CT imaging omics features in sequence. Among them, Spearman correlation analysis can be used to exclude the second CT imaging omics features with a correlation coefficient greater than 0.95. LASSO-Cox regression refers to introducing the L1 regularization term in LASSO (Least Absolute Shrinkage and Selection Operator) into the Cox regression model, λ is the regularization parameter, and the partial likelihood deviance (partial likelihood deviance) under different λ is calculated by 10-fold cross validation. The λ value that minimizes the cross-validation partial likelihood deviation is selected as the optimal value, and the imaging omics feature with a non-zero coefficient at this time is screened out as the fourth CT image. At the same time, the first clinical text information is subjected to univariate and multiple Cox regression analysis on overall survival (OS) in sequence, and the text feature information that shows statistically significant prognostic predictive significance (P<0.05) is screened as the third clinical text information.
[0093] In the embodiments of this application, Figure 4 As shown, the multimodal feature fusion and prediction module includes a fully connected layer, a group-aware attention layer, a gated cross-modal fusion layer, an overall prediction head, a first group prediction head, a second group prediction head, a third group prediction head and a fourth group prediction head; the fully connected layer is used to receive the fourth CT image, deep learning features and the third clinical text information; the fully connected layer, the group-aware attention layer and the gated cross-modal fusion layer are connected in series in sequence; the input ends of the overall prediction head, the first group prediction head, the second group prediction head, the third group prediction head and the fourth group prediction head are all connected to the output end of the gated cross-modal fusion layer. It can be understood that the fully connected layer is used to receive the second CT imaging omics features, deep learning features and the third clinical text information, and then fuse the second CT imaging omics features, deep learning features and the third clinical text information. The group-aware attention layer is used to assign an independent attention head subset to each induction chemotherapy group to dynamically calculate the interaction between intra-group and cross-group features. Among them, the group-aware attention layer can be expressed by the following formula:
[0094]
[0095] Among them, Q g is the query matrix of induction chemotherapy group g, K g is the bond matrix of induction chemotherapy group g, K overall is the global bond matrix, V g is the value matrix of induction chemotherapy group g, d k are the dimensions of the query matrix and key matrix, Matrix concatenation operation.
[0096] The gated cross-modal fusion layer is used to dynamically adjust the weight of each modality in each induction chemotherapy group. The adjustment formula is as follows:
[0097]
[0098] Among them, W m is the gating weight matrix of modality m, is the embedding vector of modality m, h group is the embedding vector of the induction chemotherapy group, and σ is the sigmoid activation function.
[0099] It is understandable that this embodiment further includes an output layer consisting of an overall prediction head, a first group prediction head, a second group prediction head, a third group prediction head, and a fourth group prediction head. The overall prediction head is used to output the overall risk value S of induction chemotherapy. overall The first group prediction head is used to output the independent risk value S of the induction chemotherapy prognosis of the induction chemotherapy regimen TP TPThe second group prediction head is used to output the independent risk value S of induction chemotherapy prognosis of induction chemotherapy regimen TPF TPF The third group prediction head is used to output the independent risk value S of the induction chemotherapy prognosis of the induction chemotherapy regimen TPK TPK The fourth group prediction head is used to output the induction chemotherapy prognosis independent risk value S of the induction chemotherapy regimen TPR TPR .
[0100] In this embodiment, when training the multimodal feature fusion and prediction module in the preset model, the Cox partial log-likelihood function can also be constructed as the loss function of each prediction head.
[0101]
[0102] In the formula, δ i is the event indicator (the value is 1 for death event and 0 for censoring event); i is the survival time of the i-th individual; X i is the multimodal vector of the i-th individual; β is the regression coefficient vector.
[0103] Then construct the hierarchical Cox loss function As the loss function of the entire model:
[0104]
[0105] In the formula, is the overall prognostic prediction loss function, is the prognostic prediction loss function for the kth induction chemotherapy regimen.
[0106] During the training of the preset model, this embodiment can also output the overall risk value S of induction chemotherapy according to the training set. overall The receiver operating characteristic (ROC) curve was drawn and the cutoff value S was selected after maximizing Youden's Index. cutoff This is used to divide the induction chemotherapy prognosis into high-risk and low-risk groups. In addition, this embodiment also tests the preset model using a test set, and during the test, draws an ROC curve and calculates the AUC value, draws a calibration curve, and performs decision curve analysis (DCA) and KM (Kaplan-Meier) curve to evaluate the predictive effect of the preset model. Based on the evaluated predictive effect, the model parameters are adjusted to improve the model accuracy of the preset model.
[0107] It is understandable that, in this embodiment, after determining the corresponding score cutoff value S after the preset model training is completed, cutoff When determining that the multimodal feature fusion and prediction module in the preset model have completed training, an ROC curve is generated according to the overall risk value of the prognosis of the first induction chemotherapy; then the Youden index corresponding to each point in the ROC curve is calculated; and then the overall risk value of the prognosis of the first induction chemotherapy corresponding to the point with the largest Youden index among all Youden indices is selected as the score cutoff value.
[0108] In an embodiment of the present application, after completing the training of the preset model and determining the cutoff value, this embodiment determines the target recommendation regimen for induction chemotherapy based on the current risk value and the score cutoff value. Specifically, this embodiment can obtain all the second induction chemotherapy prognosis independent risk values in the current risk value when the overall risk value of the current induction chemotherapy prognosis is less than the score cutoff value, and then sort all the second induction chemotherapy prognosis independent risk values, and then generate the target recommendation regimen for induction chemotherapy corresponding to the currently initially treated hypopharyngeal cancer patient based on the sorting results and the second induction chemotherapy prognosis independent risk values. Wherein, each second induction chemotherapy prognosis independent risk value corresponds to a second induction chemotherapy regimen. For example, Figure 5 As shown, for patients with newly diagnosed hypopharyngeal cancer, after inputting the current multimodal information into the preset model, the preset model outputs the overall risk value S for induction chemotherapy. overall After that, the overall risk value of induction chemotherapy was S overall With the score cutoff value S cutoff Compare to determine whether the current patient belongs to the high-risk group or the low-risk group. overall >S cutoff ) is not recommended for induction chemotherapy. overall <S cutoff ) then induction chemotherapy is recommended, and the patient's S output by the model TP 、S TPF 、S TPK 、S TPR The four treatment risks are ranked from low to high, and the recommended order of treatment options and the target recommendation plan for induction chemotherapy composed of the corresponding risk scores are output to assist clinicians in decision-making.
[0109] From the above content, it can be seen that the method of the embodiment of the present application can promote cross-group and cross-modal interaction by integrating enhanced CT and clinical text multimodal feature information by introducing the attention mechanism, thereby making the recommendation scheme output by the preset model obtained based on multimodal information training closer to the actual situation, and can provide clinical decision support for the precise treatment of hypopharyngeal cancer patients.
[0110] Reference Figure 6The present invention provides a device for predicting the prognosis of hypopharyngeal cancer induced chemotherapy and processing a treatment plan, the device comprising:
[0111] The first module 610 is configured to obtain historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, wherein the historical multimodal information includes a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen;
[0112] The second module 620 is configured to train a preset model using historical multimodal information, wherein the preset model includes a risk assessment model for the prognosis prediction and treatment plan of hypopharyngeal cancer induction chemotherapy;
[0113] The third module 630 is used to determine the corresponding score cutoff value after the preset model training is completed;
[0114] The fourth module 640 is configured to obtain current multimodal information of a hypopharyngeal cancer patient undergoing initial treatment before induction chemotherapy, wherein the current multimodal information includes a second CT image and second clinical text information;
[0115] The fifth module 650 is used to input the current multimodal information into the trained preset model to predict the current risk value;
[0116] The sixth module 660 is used to determine the target recommendation regimen for induction chemotherapy based on the current risk value and the score cutoff value.
[0117] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0118] The present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, or the like.
[0119] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0120] See also Figure 7 , Figure 7 The hardware structure of a computer device according to another embodiment is shown. The computer device includes:
[0121] The processor 710 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0122] The memory 720 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 720 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 720 and is called by the processor 710 to execute the above-mentioned methods of the embodiments of this application.
[0123] Input / output interface 730, used to implement information input and output;
[0124] Communication interface 740, used to implement communication interaction between the computer device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0125] bus 750 , which transmits information between various components of the computer device (e.g., processor 710 , memory 720 , input / output interface 730 , and communication interface 740 );
[0126] The processor 710 , the memory 720 , the input / output interface 730 , and the communication interface 740 are connected to each other within the device via a bus 750 .
[0127] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program implements the above method when executed by a processor.
[0128] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0130] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0132] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0133] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0134] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0136] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0139] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for predicting the prognosis of hypopharyngeal cancer after induction chemotherapy and treating the disease, characterized in that: The method comprises the following steps: Obtaining historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, the historical multimodal information comprising a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen; Training a preset model using the historical multimodal information, the preset model including a prognosis prediction model for induction chemotherapy of hypopharyngeal cancer and a risk assessment model for treatment plans; Determine the corresponding scoring cutoff value after the preset model training is completed; Acquiring current multimodal information of a currently untreated hypopharyngeal cancer patient before induction chemotherapy, the current multimodal information comprising a second CT image and second clinical text information; Inputting the current multimodal information into the trained preset model to predict the current risk value; A target recommendation regimen for induction chemotherapy is determined based on the current risk value and the score cutoff value.
2. The method according to claim 1, characterized in that The preset model includes a CT image preprocessing module, a CT image feature extraction module, a multimodal feature screening module and a multimodal feature fusion and prediction module.
3. The method according to claim 2, characterized in that The training of the preset model using the historical multimodal information includes: Preprocessing the first CT image by the CT image preprocessing module to obtain a region of interest mask file corresponding to the first CT image and a third CT image; performing a first feature extraction on the region of interest mask file and the third CT image by the CT image feature extraction module to obtain a first CT radiomics feature; Performing a second feature extraction on the region of interest mask file and the third CT image using a preset deep learning network model through the CT image feature extraction module to obtain a deep learning feature; screening the first CT radiomics feature using the multimodal feature screening module to obtain a second CT radiomics feature, and screening the first clinical text information according to the overall survival period to obtain third clinical text information; Inputting the second CT imaging omics feature, the deep learning feature, and the third clinical text information into the multimodal feature fusion and prediction module to perform overall risk prediction of induction chemotherapy prognosis and independent risk prediction of induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen, to obtain an overall risk value of first induction chemotherapy prognosis and an independent risk value of first induction chemotherapy prognosis corresponding to each first induction chemotherapy regimen; and calculating an overall loss function of the multimodal feature fusion and prediction module; The model parameters of the multimodal feature fusion and prediction module are adjusted according to the overall loss function.
4. The method according to claim 3, characterized in that The preset deep learning network model includes an input layer, an initial convolutional layer, a first-stage processing layer, a second-stage processing layer, a third-stage processing layer, a fourth-stage processing layer, a global average pooling layer, a fully connected layer, and an output layer; The input layer receives the region of interest mask file and the third CT image; The first-stage processing layer includes three residual attention blocks, each of which is obtained by connecting a residual block and a channel attention mechanism module in series; The second stage processing layer includes 4 residual attention blocks; The third stage processing layer includes 6 residual attention blocks; The fourth stage processing layer includes 3 residual attention blocks; The input layer, the initial convolution layer, the first stage processing layer, the second stage processing layer, the third stage processing layer, the fourth stage processing layer, the global average pooling layer, the fully connected layer and the output layer are connected in series in sequence; The output layer is used to output deep learning features.
5. The method according to claim 3, characterized in that The multimodal feature fusion and prediction module includes a fully connected layer, a group-aware attention layer, a gated cross-modal fusion layer, an overall prediction head, a first group prediction head, a second group prediction head, a third group prediction head, and a fourth group prediction head; The fully connected layer is used to receive the second CT radiomics feature, the deep learning feature and the third clinical text information; The fully connected layer, the group-aware attention layer, and the gated cross-modal fusion layer are connected in series in sequence; Input ends of the overall prediction head, the first group prediction head, the second group prediction head, the third group prediction head, and the fourth group prediction head are all connected to the output end of the gated cross-modal fusion layer.
6. The method according to claim 3, characterized in that Determining the corresponding score cutoff value after the preset model training is completed includes: After determining that the multimodal feature fusion and prediction module in the preset model has completed training, generating a ROC curve according to the overall risk value of the first induction chemotherapy prognosis; Calculate the Youden index corresponding to each point in the ROC curve; The overall risk value of the first induction chemotherapy prognosis corresponding to the point with the largest Youden index among all the Youden indices is selected as the score cutoff value.
7. The method according to claim 1, characterized in that Determining a target recommended regimen for induction chemotherapy according to the current risk value and the score cutoff value includes: When the overall risk value of the current induction chemotherapy prognosis in the current risk value is less than the score cutoff value, obtaining all second induction chemotherapy prognosis independent risk values in the current risk value, each second induction chemotherapy prognosis independent risk value corresponding to one second induction chemotherapy regimen; Rank all the independent risk values of the second induction chemotherapy prognosis; A target recommendation regimen for the induction chemotherapy corresponding to the currently untreated hypopharyngeal cancer patient is generated based on the ranking result and the second induction chemotherapy prognostic independent risk value.
8. A device for predicting the prognosis of hypopharyngeal cancer induction chemotherapy and processing treatment plans, characterized in that: The device comprises: A first module is configured to obtain historical multimodal information of a plurality of previously untreated hypopharyngeal cancer patients before induction chemotherapy, the historical multimodal information including a first CT image, first clinical text information, a first induction chemotherapy regimen selected by the previously untreated hypopharyngeal cancer patients, and an overall survival period corresponding to the first induction chemotherapy regimen; A second module is used to train a preset model using the historical multimodal information, wherein the preset model includes a prognosis prediction model for induction chemotherapy of hypopharyngeal cancer and a risk assessment model for treatment plans; The third module is used to determine the corresponding score cutoff value after the preset model training is completed; A fourth module is configured to obtain current multimodal information of a currently untreated hypopharyngeal cancer patient before induction chemotherapy, the current multimodal information including a second CT image and second clinical text information; A fifth module is configured to input the current multimodal information into the trained preset model to predict a current risk value; The sixth module is used to determine the target recommendation scheme of induction chemotherapy based on the current risk value and the score cutoff value.
9. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Nasopharynx cancer patient survival prediction model, construction method, equipment and storage medium
CN114613502A
Nasopharyngeal cancer prognosis prediction method capable of fusing incomplete multiple modes and crossing medical sites
CN115938532A
Cancer risk prediction method based on deep learning
CN118507048A
Advanced nasopharynx cancer treatment effect prediction system based on deep learning
CN119132582A
Multimodal spatiotemporal deep learning system for prediction of cancer therapy outcomes
WO2024173368A1
Cited By
Auxiliary intelligent chip for nursing pci postoperative patient
CN121528479A