A lung cancer lymph node metastasis multi-modal prediction method and prediction device

CN122615751APending Publication Date: 2026-08-21FIRST PEOPLES HOSPITAL OF KUNMING +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610915950.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种肺癌淋巴结转移多模态预测方法及预测装置,解决现有方法在跨模态语义关联、小病灶识别及临床-影像融合方面的不足,最终实现无创、精准且高效的术前辅助诊断

Benefits of technology

(1)本方案不需要做穿刺活检,只用术前常规的CT、MR影像就能做到精准预测,减少患者的痛苦和穿刺并发症风险,实现了无创性检测,提升了转移预测精度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122615751A_ABST
    Figure CN122615751A_ABST
Patent Text Reader

Abstract

The application discloses a lung cancer lymph node metastasis multi-modal prediction method and a prediction device, relates to the cross technical field of lung cancer metastasis prediction and medical artificial intelligence, and comprises the following steps: acquiring image data, clinical data and gold standard data through a data acquisition module; processing the image data, the clinical data and the gold standard data through a data preprocessing module; performing feature extraction and enhancement processing on the image data through an improved BiModal-Transformer architecture; performing fusion processing and metastasis condition prediction on the image data and the clinical data through a multi-modal feature fusion prediction module; and displaying the prediction result, a prediction probability value and a heat map through a prediction result display module; cross-modal feature deep fusion is realized, lung cancer lymph node metastasis conditions can be accurately identified, the blank of existing technologies in the aspect of multi-modal image collaborative prediction of lymph node metastasis is filled, and the development of precise diagnosis and treatment of lung cancer to cross-modal and explainable directions is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of lung cancer metastasis prediction and medical artificial intelligence, and in particular to a multimodal prediction method and device for lung cancer lymph node metastasis. Background Technology

[0002] The status of lymph node metastasis in lung cancer is the core basis for TNM staging of lung cancer and directly determines the treatment strategy: patients without metastasis only need to undergo lobectomy and lymph node sampling, while patients with metastasis need to undergo extended lymph node dissection or preoperative neoadjuvant therapy. Misdiagnosis can lead to overtreatment or undertreatment.

[0003] In existing technologies, auxiliary diagnostic techniques mainly include single-modality deep learning models, traditional radiomics methods, and clinical indicators combined with single images. First, integrating CT and MR features through feature stitching or result voting fails to establish a semantic association between CT and MR, making it impossible to improve prediction accuracy by leveraging the complementarity of the two modalities. Second, metastatic lymph nodes in lung cancer are often less than 5mm in diameter, and the fixed-window attention mechanism of existing models struggles to focus on these tiny lesions. Furthermore, the different scanning resolutions of CT and MR can lead to feature alignment deviations, further reducing the ability to identify small lesions. In addition, most existing models treat clinical indicators as independent feature inputs, failing to establish an image-pathology association and resulting in low information utilization. Finally, most deep learning models are like "black boxes," unable to clearly explain the contribution of features from different regions of CT and MR to metastasis prediction.

[0004] Therefore, a multimodal prediction method and device for lung cancer lymph node metastasis are provided to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to provide a multimodal prediction method and device for lung cancer lymph node metastasis, which addresses the shortcomings of existing methods in cross-modal semantic association, small lesion identification and clinical-image fusion, and ultimately achieves non-invasive, accurate and efficient preoperative auxiliary diagnosis.

[0006] To achieve the above objectives, the present invention provides a multimodal prediction method for lung cancer lymph node metastasis, comprising the following steps: S1: Acquire imaging data, clinical data, and gold standard data through the data acquisition module and transmit them to the data preprocessing module; S2: The data preprocessing module processes the image data, clinical data, and gold standard data, and transmits the processed data to the dual-modal feature extraction module and the multimodal feature fusion prediction module; S3: The image data is processed by the dual-modal feature extraction module to extract and enhance features, and the processed data is then transmitted to the multimodal feature fusion prediction module. S4: The multimodal feature fusion prediction module performs fusion processing on image data and clinical data and predicts metastasis, and transmits the prediction results to the prediction result display module. S5: The prediction results, prediction probability values, and heatmaps are displayed through the prediction results display module.

[0007] Preferably, step S1 specifically includes the following steps: S11: Acquire the patient's preoperative chest CT and chest MR images through the data acquisition module; S12: Acquire clinical data through the data acquisition module. Clinical data includes patient age, smoking history, tumor pathological subtype, and tumor markers. S13: Obtain gold standard data through the data acquisition module. The gold standard data is set as the lymph node metastasis result confirmed by postoperative pathology.

[0008] Preferably, step S2 specifically includes the following steps: S21: Spatial alignment processing is performed on the image data. Based on the registration algorithm of anatomical tumor markers, the interpolation of the chest MR images is adjusted to the resolution of the chest CT images, which are set to 1×1×1mm. 3 ; S22: Perform lesion localization processing on image data, and automatically segment the lung tumor area and mediastinal lymph node area using the improved U-Net model to generate the region of interest (ROI); S23: Perform structured processing on clinical data, and transform clinical data into feature vectors through the clinical BERT model; S24: Perform Z-score standardization on the gold standard data.

[0009] Preferably, step S3 specifically includes the following steps: S31: Key features of chest CT and chest MR images are extracted using the improved BiModal-Transformer architecture to obtain key feature maps; S32: The key feature map is processed through a cross-modal attention interaction layer to obtain a bimodal semantic association feature map; S33: The region of interest (ROI) is locally magnified by using a small lesion enhancement layer.

[0010] Preferably, step S31 specifically includes the following steps: Step 1: Based on the multi-scale window attention mechanism, key features are extracted from the chest CT image through a window of size 3×3×3 to obtain the key feature map of the chest CT image; Step 2: Introduce edge enhancement units to enhance the boundary between lymph nodes and surrounding tissues using the Sobel operator; Step 3: Based on the multi-scale window attention mechanism, key features are extracted from the chest MR image through a window of size 5×5×5 to obtain the key feature map of the chest MR image; Step 4: Introduce a signal intensity correction unit. First, calculate the mean and standard deviation of the signal intensity of normal lymph nodes in the same batch, and then standardize the chest MR images.

[0011] Preferably, step S32 specifically includes the following steps: Step 1: Calculate the attention weights of chest CT images on chest MR images using a learnable weight matrix; Step 2: Calculate the attention weights of chest MR images on chest CT images using a learnable weight matrix; Step 3: Add the key feature maps of the weighted chest CT image and the weighted chest MR image element-wise to obtain a bimodal semantic association feature map.

[0012] Preferably, step S33 specifically includes the following steps: Step 1: Perform size analysis on the mediastinal lymph nodes and mark small lesions with a diameter of less than 5 mm; Step II: Weight the key features of small lesions, setting the weight to 1.5 times the attention weight; Step III: Introduce a local feature magnification unit to upsample the key feature map of small lesions by a factor of two, and extract the local features of small lesions through a convolution kernel of size 3×3×3.

[0013] Preferably, step S4 specifically includes the following steps: S41: Adapt the bimodal semantic association feature map to the clinical data feature vector; S42: Introduce a feature interaction layer to perform element-wise multiplication and residual connection processing on the adapted clinical features and bimodal features to establish image-clinical association and obtain fused features; S43: Introduce an L1 regularization term into the XGBoost classifier, iteratively train the fused features, set the number of iterations to 100 rounds, and obtain the prediction results and prediction probability values; S44: Introduce the Grad-CAM++ algorithm to generate a bimodal feature attribution map, and use heatmaps to annotate the regions that contribute the most to the prediction results.

[0014] A prediction device for a multimodal prediction method for lymph node metastasis in lung cancer includes a data acquisition module, a data preprocessing module, a dual-modal feature extraction module, a multimodal feature fusion prediction module, and a prediction result display module connected in sequence. The data preprocessing module is connected to the multimodal feature fusion prediction module.

[0015] Therefore, the present invention employs the above-mentioned multimodal prediction method and device for lung cancer lymph node metastasis, which has the following beneficial effects: (1) This approach does not require a puncture biopsy. It can make accurate predictions using only routine preoperative CT and MR images, reducing patient pain and the risk of puncture complications, achieving non-invasive detection and improving the accuracy of metastasis prediction. (2) This scheme uses feature attribution maps and modal contribution quantification as the basis for doctors to clearly judge the model. The clinical acceptance is more than 60% higher than that of the black box model, making it easier to promote in actual diagnosis and treatment and more interpretable. (3) The model of this scheme only takes 0.5 minutes to process the data of a single patient, which is 60 times more efficient than doctors manually reviewing images. It can quickly provide decision support for preoperative discussions, shorten the patient's preoperative waiting time, reduce the hospital's treatment pressure, and optimize treatment efficiency.

[0016] The method of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] Figure 1 This is a flowchart of a multimodal prediction method for lung cancer lymph node metastasis according to the present invention; Figure 2 This is a structural diagram of a multimodal prediction device for lung cancer lymph node metastasis according to the present invention, wherein: 1. Data acquisition module; 2. Data preprocessing module; 3. Dual-modal feature extraction module; 4. Multimodal feature fusion prediction module; 5. Prediction result display module; Figure 3 This is the ROC performance curve of the BiModal-Transformer dual-modal model of this invention; Figure 4 A confusion matrix diagram of the prediction results of the BiModal-Transformer model; Figure 5 Attribution plot for global feature contribution of multimodal fusion model. Detailed Implementation

[0018] The method of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0019] Unless otherwise defined, the methodological or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0020] The terms "comprising" or "including" as used in this invention mean that the element preceding the term encompasses the element listed after the term, and do not exclude the possibility of encompassing other elements. Terms such as "inner," "outer," "upper," and "lower" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In this invention, unless otherwise explicitly specified and limited, the term "attached" and similar terms should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can refer to a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication of two elements or the interaction relationship between two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0021] Example like Figures 1-5 As shown, this invention provides a multimodal prediction method for lymph node metastasis in lung cancer, specifically referring to... Figure 1 This includes the following steps: S1: Acquire imaging data, clinical data, and gold standard data through the data acquisition module and transmit them to the data preprocessing module; Step S1 specifically includes the following steps: S11: Acquire the patient's preoperative chest CT and chest MR images through the data acquisition module; S12: Acquire clinical data through the data acquisition module. Clinical data includes patient age, smoking history, tumor pathological subtype, and tumor markers. S13: Obtain gold standard data through the data acquisition module. The gold standard data is set as the lymph node metastasis result confirmed by postoperative pathology.

[0022] S2: The data preprocessing module processes the image data, clinical data, and gold standard data, and transmits the processed data to the dual-modal feature extraction module and the multimodal feature fusion prediction module; Step S2 specifically includes the following steps: S21: Spatial alignment processing is performed on the image data. Based on the registration algorithm of anatomical tumor markers, such as the registration algorithm of anatomical trachea and aortic arch, the interpolation of the chest MR image is adjusted to the resolution of the chest CT image to eliminate the spatial deviation between the two modalities. The resolution of the chest CT image is set to 1×1×1mm. 3 ; S22: The imaging data is processed to locate lesions. The lung tumor area and mediastinal lymph node area are automatically segmented by the improved U-Net model to generate the region of interest (ROI) and reduce the interference of background noise on subsequent analysis. S23: Perform structured processing on clinical data, and transform clinical data into feature vectors through the clinical BERT model; S24: Perform Z-score standardization on the gold standard data.

[0023] S3: The image data is processed by the dual-modal feature extraction module to extract and enhance features, and the processed data is then transmitted to the multimodal feature fusion prediction module. Step S3 specifically includes the following steps: S31: In order to adapt to the different characteristics of CT and MR modalities, the key features of chest CT images and chest MR images are extracted separately through the improved BiModal-Transformer architecture to obtain key feature maps, avoiding feature loss caused by using a unified branch. Step S31 specifically includes the following steps: Step 1: Considering that the slice thickness of chest CT images is 1mm, the resolution is high, and it can clearly show the lung structure and calcifications, based on the multi-scale window attention mechanism, key features of chest CT images are extracted through a window of size 3×3×3. This size can not only cover the fine structure of chest CT images, but also efficiently capture the density differences of lymph nodes, such as the uneven density features that often appear in metastatic lymph nodes. This branch can extract them more accurately and obtain the key feature map of chest CT images. Step 2: Introduce edge enhancement units to enhance the boundary between lymph nodes and surrounding tissues using the Sobel operator, which helps to better distinguish between normal and abnormal lymph nodes in the future; Step 3: The chest MR image slice thickness is 3mm, which has a slightly lower resolution but a stronger ability to distinguish soft tissues and can show signal changes inside lymph nodes. Based on the multi-scale window attention mechanism, key features are extracted from the chest MR image through a window of size 5×5×5. A larger window can cover more soft tissue information in the chest MR image, and the key feature map of the chest MR image is obtained. Step 4: The chest MR image scan sequences of different patients may differ, which may cause signal fluctuations. Therefore, a signal intensity correction unit is introduced. First, the mean and standard deviation of the signal intensity of normal lymph nodes in the same batch are statistically analyzed. Then, the chest MR images are standardized to eliminate signal interference caused by scan differences and make the extracted features more stable.

[0024] S32: To solve the problem of "simple splicing" in traditional models, and to enable CT and MR features to "communicate" with each other and mine their complementary information, a cross-modal attention interaction layer is used to process key feature maps to obtain bimodal semantic association feature maps. Step S32 specifically includes the following steps: Step 1: Calculate the attention weights of chest CT images on chest MR images using a learnable weight matrix; Step 2: Calculate the attention weights of chest MR images on chest CT images using a learnable weight matrix; For example, if CT features show calcification in a lymph node, it may indicate metastasis and will give higher weight to the signal intensity features of that area in MR features; conversely, if MR features show blurred lymph node boundaries, it may indicate metastasis and will also enhance the density features of that area in CT features.

[0025] Step 3: Add the key feature maps of the weighted chest CT image and the weighted chest MR image element-wise to obtain a bimodal semantic association feature map. Such features contain more comprehensive information than individual CT or MR features.

[0026] The dimensions of the key feature maps of chest CT images, chest MR images, and bimodal semantic association feature maps were all set to 1024.

[0027] S33: The region of interest (ROI) is locally magnified by using a small lesion enhancement layer.

[0028] Step S33 specifically includes the following steps: Step 1: To address the issue that metastatic lymph nodes are often smaller than 5mm and easily overlooked, the size of the mediastinal lymph nodes is statistically analyzed, and small lesions with a diameter of less than 5mm are marked. Step II: Weight the key features of small lesions, setting the weight to 1.5 times the attention weight; Step III: Introduce a local feature magnification unit to upsample the key feature map of small lesions by a factor of two. Extract local features of small lesions, such as signal inhomogeneity within small lesions, through a 3×3×3 convolution kernel to ensure that the features of small metastases are not masked by large lesions or background information.

[0029] S4: The multimodal feature fusion prediction module performs fusion processing on image data and clinical data and predicts metastasis, and transmits the prediction results to the prediction result display module. Step S4 specifically includes the following steps: S41: Adapt the bimodal semantic association feature map with the clinical data feature vector. First, encode the clinical data and concatenate it to obtain a 64-dimensional clinical feature vector. Then, use a single fully connected layer with an input dimension of 64 and an output dimension of 1024 to map it to a 1024-dimensional feature through linear transformation in the nn library. This ensures consistency with the bimodal semantic association feature of the same dimension, and then carry out subsequent feature fusion operations. S42: Introducing a feature interaction layer, the adapted clinical features and bimodal features are processed by element-wise multiplication and residual connection to preserve the original feature information and establish image-clinical association. For example, when the clinical features show that the patient has squamous cell carcinoma, the common imaging features of squamous cell carcinoma metastasis, such as "irregular lymph node margins", in the bimodal features will be strengthened to further improve the targeting of the features and obtain fused features. S43: Introduce an L1 regularization term into the XGBoost classifier to reduce the risk of model overfitting. Iterate the fused features and set the number of iterations to 100 rounds to obtain prediction results and prediction probability values, such as "transfer" with a probability ≥ 0.5 or "no transfer" with a probability < 0.5, for doctors' reference. S44: Introducing the Grad-CAM++ algorithm to generate a dual-modal feature attribution map, using heatmaps to annotate the regions that contribute the most to the prediction results. For example, in transfer prediction, high-signal areas on MR images and density-abnormal areas on CT images will be displayed in red.

[0030] S5: The prediction results, prediction probability values, and heatmaps are displayed through the prediction results display module.

[0031] like Figure 2 As shown, a prediction device for a multimodal prediction method for lymph node metastasis in lung cancer includes a data acquisition module 1, a data preprocessing module 2, a dual-modal feature extraction module 3, a multimodal feature fusion prediction module 4, and a prediction result display module 5 connected in sequence. The data preprocessing module 2 is connected to the multimodal feature fusion prediction module 4.

[0032] Therefore, the present invention adopts the above-mentioned multimodal prediction method and device for lung cancer lymph node metastasis, which integrates medical image processing, deep learning and clinical data. It is a cutting-edge application of artificial intelligence in the field of thoracic tumor staging prediction, and also provides an innovative technical path for the mining of imaging biomarkers for lung cancer lymph node metastasis.

[0033] Finally, it should be noted that the above embodiments are only used to illustrate the method of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the method of the present invention, and these modifications or equivalent substitutions should not cause the modified method to deviate from the spirit and scope of the method of the present invention.

Claims

1. A multimodal prediction method for lymph node metastasis in lung cancer, characterized in that, Includes the following steps: S1: Acquire imaging data, clinical data, and gold standard data through the data acquisition module and transmit them to the data preprocessing module; S2: The data preprocessing module processes the image data, clinical data, and gold standard data, and transmits the processed data to the dual-modal feature extraction module and the multimodal feature fusion prediction module; S3: The image data is processed by the dual-modal feature extraction module to extract and enhance features, and the processed data is then transmitted to the multimodal feature fusion prediction module. S4: The multimodal feature fusion prediction module performs fusion processing on image data and clinical data and predicts metastasis, and transmits the prediction results to the prediction result display module. S5: The prediction results, prediction probability values, and heatmaps are displayed through the prediction results display module.

2. The multimodal prediction method for lung cancer lymph node metastasis according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11: Acquire the patient's preoperative chest CT and chest MR images through the data acquisition module; S12: Acquire clinical data through the data acquisition module. Clinical data includes patient age, smoking history, tumor pathological subtype, and tumor markers. S13: Obtain gold standard data through the data acquisition module. The gold standard data is set as the lymph node metastasis result confirmed by postoperative pathology.

3. The multimodal prediction method for lung cancer lymph node metastasis according to claim 2, characterized in that, Step S2 specifically includes the following steps: S21: Spatial alignment processing is performed on the image data. Based on the registration algorithm of anatomical tumor markers, the interpolation of the chest MR images is adjusted to the resolution of the chest CT images, which are set to 1×1×1mm. 3 ; S22: Perform lesion localization processing on image data, and automatically segment the lung tumor area and mediastinal lymph node area using the improved U-Net model to generate the region of interest (ROI); S23: Perform structured processing on clinical data, and transform clinical data into feature vectors through the clinical BERT model; S24: Perform Z-score standardization on the gold standard data.

4. The multimodal prediction method for lung cancer lymph node metastasis according to claim 3, characterized in that, Step S3 specifically includes the following steps: S31: Key features of chest CT and chest MR images are extracted using the improved BiModal-Transformer architecture to obtain key feature maps; S32: The key feature map is processed through a cross-modal attention interaction layer to obtain a bimodal semantic association feature map; S33: The region of interest (ROI) is locally magnified by using a small lesion enhancement layer.

5. The multimodal prediction method for lung cancer lymph node metastasis according to claim 4, characterized in that, Step S31 specifically includes the following steps: Step 1: Based on the multi-scale window attention mechanism, key features are extracted from the chest CT image through a window of size 3×3×3 to obtain the key feature map of the chest CT image; Step 2: Introduce edge enhancement units to enhance the boundary between lymph nodes and surrounding tissues using the Sobel operator; Step 3: Based on the multi-scale window attention mechanism, key features are extracted from the chest MR image through a window of size 5×5×5 to obtain the key feature map of the chest MR image; Step 4: Introduce a signal intensity correction unit. First, calculate the mean and standard deviation of the signal intensity of normal lymph nodes in the same batch, and then standardize the chest MR images.

6. The multimodal prediction method for lung cancer lymph node metastasis according to claim 4, characterized in that, Step S32 specifically includes the following steps: Step 1: Calculate the attention weights of chest CT images on chest MR images using a learnable weight matrix; Step 2: Calculate the attention weights of chest MR images on chest CT images using a learnable weight matrix; Step 3: Add the key feature maps of the weighted chest CT image and the weighted chest MR image element-wise to obtain a bimodal semantic association feature map.

7. The multimodal prediction method for lung cancer lymph node metastasis according to claim 4, characterized in that, Step S33 specifically includes the following steps: Step 1: Perform size analysis on the mediastinal lymph nodes and mark small lesions with a diameter of less than 5 mm; Step II: Weight the key features of small lesions, setting the weight to 1.5 times the attention weight; Step III: Introduce a local feature magnification unit to upsample the key feature map of small lesions by a factor of two, and extract the local features of small lesions through a convolution kernel of size 3×3×3.

8. The multimodal prediction method for lung cancer lymph node metastasis according to claim 1, characterized in that, Step S4 specifically includes the following steps: S41: Adapt the bimodal semantic association feature map to the clinical data feature vector; S42: Introduce a feature interaction layer to perform element-wise multiplication and residual connection processing on the adapted clinical features and bimodal features to establish image-clinical association and obtain fused features; S43: Introduce an L1 regularization term into the XGBoost classifier, iteratively train the fused features, set the number of iterations to 100 rounds, and obtain the prediction results and prediction probability values; S44: Introduce the Grad-CAM++ algorithm to generate a bimodal feature attribution map, and use heatmaps to annotate the regions that contribute the most to the prediction results.

9. A predictive apparatus for the multimodal prediction method for lung cancer lymph node metastasis as described in any one of claims 1-8, characterized in that, It includes a data acquisition module, a data preprocessing module, a dual-modal feature extraction module, a multimodal feature fusion prediction module, and a prediction result display module, which are connected in sequence. The data preprocessing module is connected to the multimodal feature fusion prediction module.