Classification method of high-risk factors for lung cancer based on multimodal fusion

Through deep learning and multimodal data analysis, combined with fast Fourier transform and multi-head self-attention mechanism, the problem of predicting high-risk factors for early lung cancer was solved, and accurate identification of high-risk factors and personalized treatment decision support were achieved.

CN119206333BActive Publication Date: 2025-10-03BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411266662.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-10-03
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively predict high-risk factors in patients with early-stage lung cancer in clinical practice, which affects the formulation of surgical strategies and treatment plans.

Method used

A method based on deep learning and multimodal data, combined with fast Fourier transform and multi-head self-attention mechanism, is used to perform multi-label classification of chest CT images, medical records and laboratory test data to predict whether patients have high-risk factors such as poorly differentiated tumors, vascular invasion, and visceral pleural invasion.

Benefits of technology

It achieves accurate prediction of high-risk factors for early lung cancer with an AUC of 70%, providing intelligent decision-making support for clinicians and improving the accuracy and personalization of treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206333B_ABST
    Figure CN119206333B_ABST
Patent Text Reader

Abstract

This paper proposes a lung cancer risk factor classification method based on multimodal fusion. This method is a method for predicting high-risk factors for early lung cancer based on deep learning and multimodal data. It belongs to the fields of medical image processing (G06T) and medical diagnosis (A61B). It is particularly related to the use of artificial intelligence technology for early diagnosis of lung cancer. It aims to analyze multimodal medical data using deep learning models to identify and predict high-risk factors for early lung cancer. Experiments have shown that the multimodal method complies with the NCCN guidelines. It can assist in preoperative planning and auxiliary treatment decision-making, ultimately contributing to personalized patient care.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is a method for predicting high-risk factors for early lung cancer based on deep learning and multimodal data, belonging to the fields of medical image processing (G06T) and medical diagnosis (A61B). Background Art

[0002] Lung cancer includes non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC), of which NSCLC accounts for approximately 85% and SCLC accounts for approximately 15%. The prognosis for SCLC patients is extremely poor, and two-thirds of patients already have distant metastatic disease at the time of initial diagnosis. The survival rate of NSCLC patients has improved in recent years, which is related to recent advances in screening, minimally invasive techniques for diagnosis and treatment, radiotherapy (RT) including stereotactic ablative radiotherapy (SABR), and new targeted therapies and immunotherapies. Overall, most patients with advanced (stage IV) lung cancer die within 5 years of diagnosis, while the 5-year survival rate for patients with early (stage IA) lung cancer is over 75%. For patients, early detection, early diagnosis, and early treatment can bring better prognosis to lung cancer patients. Therefore, research on early lung cancer is particularly important.

[0003] According to the NCCN Clinical Practice Guidelines in Oncology, radical surgical resection is the recommended preferred local treatment for patients with early-stage lung cancer. However, some patients still face secondary surgery or postoperative adjuvant therapy due to high-risk factors (including poorly differentiated tumors (excluding well-differentiated neuroendocrine tumors), vascular invasion, visceral pleural invasion, and unclear lymph node status). These high-risk factors are associated with tumor aggressiveness and metastatic propensity. Therefore, using preoperative low-dose computed tomography (LDCT) imaging data to predict a patient's potential high-risk factors is crucial for guiding thoracic surgeons in selecting the extent of lung resection and surgical approach, as well as the need for preoperative (neoadjuvant) chemotherapy or radiotherapy to reduce the tumor.

[0004] In clinical treatment, the discovery of these lung cancer risk factors is typically based on pathological observations and is not reflected in chest CT images or laboratory tests. Therefore, to date, no major breakthroughs have been achieved in diagnosing lung cancer risk factors using clinical data. However, thanks to the rapid development of artificial intelligence technology in recent years, and its excellent performance in microscopic observation, it is now possible to use AI to discover these "invisible" factors. Summary of the Invention

[0005] The purpose of this invention is to predict in advance whether lung cancer patients have high-risk factors based on deep learning and machine learning theories and methods, when this cannot be achieved clinically, so as to improve the patient's prognosis.

[0006] The present invention ultimately incorporated multimodal data (including chest CT, medical records, and laboratory test data) from 1,113 patients collected between 2019 and 2022. A deep learning method based on fast Fourier transform and multi-head self-attention mechanism was proposed to predict high-risk factors for early lung cancer. This method is a multi-label classification method that, through comprehensive analysis of multimodal data, can determine the patient's performance in various high-risk situations (whether there is a poorly differentiated tumor, whether there is vascular invasion, whether there is visceral pleural invasion, and whether there is wedge resection).

[0007] The proposed method demonstrates excellent performance, with an AUC (area under the receiver operating characteristic (ROC) curve) of 70%. We conducted extensive comparative and ablation experiments to demonstrate the effectiveness of our method. Our method provides an intelligent multimodal solution for predicting high-risk factors for early-stage lung cancer, potentially serving as an intelligent decision support system for clinicians, ultimately facilitating the development of rapid and precise treatment strategies for early-stage lung cancer.

[0008] The present invention comprises the following steps,

[0009] Step 1: Build training data:

[0010] 1. Through exclusion screening, data from 1,113 patients were collected, including 393 patients with high-risk lung cancer and 720 patients without high-risk lung cancer. Each patient's input data included 16 chest CT images, a set of medical records, and a set of laboratory test data. Exclusion criteria included: 1) patients with non-stage I or II lung cancer; 2) patients who did not undergo chest LDCT, laboratory tests, or pathology examinations; and 3) patients who lacked a final diagnosis for any of the four high-risk factors.

[0011] 2. In this study, patients with high-risk lung cancer were defined as those with one or more of the following conditions: poorly differentiated tumors, vascular invasion, visceral pleural involvement, or wedge resection. Ultimately, 393 patients with high-risk lung cancer and 720 patients without high-risk lung cancer were enrolled in the study. Of the 393 high-risk patients, 282 underwent wedge resection, 109 with poorly differentiated tumors, 75 with visceral pleural involvement, and 39 with vascular invasion. Demographic characteristics, medical history, and laboratory test results were collected from the patients' medical records.

[0012] 3. Four radiologists reviewed all chest CT images to ensure image quality and clarity and to mark lesion locations. Based on the radiologists' markings, 16 spatially contiguous chest CT images (contiguous 1.25 mm interval CT plain scan images) were selected for each patient. Because this study focused on patients with early-stage lung cancer with relatively small lesions, for patients with fewer than 16 CT images containing the lesion, a total of 16 images centered around the lesion were selected as model input. Medical record data included gender, age, smoking history, alcohol consumption history, history of lung disease, history of malignancy, family history, pulmonary nodule size, and location. Nodule size and location were processed to comprehensively represent lung information. Final medical record data was generated, containing 11 variables. Laboratory test data included six variables: white blood cell (WBC), granulocyte ratio (GR), carcinoembryonic antigen (CEA), cytokeratin 19 fragment (CY), neuron-specific enolase (NSE), and progastrin-releasing peptide (Pro-GRP). These data were included based on previous studies that revealed its association with cancer.

[0013] 4. To ensure balanced distribution across different classes, we randomly split the dataset into training and test sets in an 8:2 ratio, using five-fold cross-validation for both training and testing. The training set included 225 patients who underwent wedge resection, 87 patients with poorly differentiated tumors, 60 patients with pleural invasion, and 31 patients with vascular invasion. The test set included 57 patients who underwent wedge resection, 22 patients with poorly differentiated tumors, 15 patients with pleural invasion, and 8 patients with vascular invasion.

[0014] Step 2: Design and train a multimodal model:

[0015] 1. To integrate diverse medical data, the proposed model consists of three components, enabling it to adapt to and effectively integrate various data modalities: data mapping, feature extraction, and modality fusion. As a universal approach, these three components are plug-and-play modules that can be freely switched to meet different task requirements. Furthermore, to optimize the workflow, the model training process is performed end-to-end, eliminating the need to independently optimize each component. The data is processed and mapped to the input format. Next, the data passes through a feature extraction module for each modality to obtain a high-dimensional semantic feature representation. Next, a feature fusion module performs semantic feature integration analysis.

[0016] 2. Feature Extraction and Visual Embedding. To effectively extract semantic information from CT images, we use the encoder portion of the Swin Transformer V2 model (Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., Wei, F., Guo, B.: Swin transformer v2: Scaling up capacity and resolution. In: 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11999–12009 (2022)) as the backbone for image feature extraction. We collect patients' medical records and laboratory test data and organize them into a structured and standardized data format. Medical records include both discrete values ​​(e.g., gender and smoking history) and continuous values ​​(e.g., lung nodule size). The laboratory test data consists of continuous real numbers. To enhance the model's fit and generalization capabilities and accelerate training convergence, we preprocessed both discrete and continuous data. For discrete values, we used one-hot encoding. For continuous data, we performed min-max normalization on each feature column, normalizing it using the maximum and minimum values ​​of each feature. We then concatenated the medical records and laboratory test data. Using a feature adapter, we aligned the feature records of the medical records and the laboratory test data to have the same length as the image features.

[0017] 3. Missing data is common in medical datasets. In this study, we excluded patients without chest CT images, so missing data only occurred in medical records and laboratory test data. To address this issue, we used polynomial interpolation to handle missing data.

[0018] 4. Multimodal fusion:

[0019] (1. We use an intermediate fusion strategy to fuse the high-dimensional features of each modality in the feature space, and then use a neural network to learn and predict the fused features. We propose a feature fusion method that combines linear attention and fast Fourier transform (FFT). The multimodal features are concatenated and then first input into the linear attention module to discover the dependencies between features. Then, the output of the linear attention module is input into the fast Fourier transform module to transform the spatial domain of the features into the frequency domain, thereby discovering the fused frequency domain information of the features.

[0020] By using the linear attention mechanism (Linear Attention with Double Softmax), the time and space complexity is O(n) compared to the traditional Attention (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, AN, Kaiser, L., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30) 2 ), which uses the joint law of matrix multiplication and matrix transformation to reduce the time and space complexity to O(n). The specific equation can be expressed as:

[0021]

[0022] where Q,K∈R n×d , which is the same as the definition of traditional Attention. Consider Q, K as n×d vectors, and then cluster them into m categories, obtaining a matrix consisting of m cluster centers Represents the transpose of a matrix. Indicates that when the matrix When reversible, the inverse of the matrix, when the matrix When it is not invertible, the pseudo-inverse of the matrix.

[0023] (2. Fast Fourier Transform. Discrete Fourier Transform (DFT) plays a very important role in the field of digital signal processing. The Fast Fourier Transform algorithm uses the symmetry and periodicity of the signal to reduce the computational complexity of DFT from O(n 2 ) is reduced to O(nlogn). The form of the inverse DFT is similar to the DFT and can also be efficiently calculated by the fast inverse Fourier transform (IFFT). Our implementation is similar to the given feature x∈R H×W×D , we first perform a 2D FFT on the spatial dimension to transform it into the frequency domain:

[0024]

[0025] where F[·] is the 2D FFT. Note that X is a complex tensor representing the spectrum of x. We can then filter X with a learnable filter K∈C H×W×D Multiply to adjust the spectrum of x:

[0026]

[0027] where ⊙ is the element-wise product (also called the Hadamard product). The filter K has the same dimensions as X and is therefore called a global filter. X can represent any filter in the frequency domain. Finally, we use IFFT to transform the modulation spectrum X back to the spatial domain and update the features:

[0028]

[0029] (3. Perform deep feature compression and linear regression on the data after FFT operation, and operate the four variables with different multi-layer perceptrons respectively.

[0030] Step 3: Set training parameters

[0031] 1. Use the Adam optimizer, setting the learning rate and maximum number of iterations. The training process includes forward propagation to calculate the loss function and backpropagation to update the neural network weights. Training continues until the preset convergence conditions are met.

[0032] 2. After training is complete, the final trained model parameters are saved for subsequent automatic pathology diagnosis and classification tasks. Using the feature vectors to map labels, you can diagnose whether the patient has wedge resection, poor differentiation, vascular invasion, or pleural invasion, with the labels representing 0 and 1, respectively.

[0033] The present invention seeks to leverage advances in artificial intelligence and deep learning to enhance the detection and treatment of high-risk factors for early lung cancer. Although there are clinical challenges in determining the association between these data and high-risk factors for lung cancer, the present method has achieved promising results with an AUC of 70%.

[0034] Furthermore, the proposed method highlights the potential of combining multiple data modalities to improve diagnostic performance. Compared with single-modality approaches, a multimodal model integrating lung CT images, medical record text, and laboratory test data achieved higher area under the curve (AUC) and accuracy. This suggests that integrating different types of data can lead to more reliable and accurate assessments of lung cancer risk factors.

[0035] This paper addresses the problem of uncovering potential correlations and "hidden" features in multimodal transport by designing different feature extractors to capture the characteristics of each modality separately. These features are then fused using Fast Fourier Transform (FFT) and multi-head self-attention, enabling the discovery of frequency domain features and attentional features associated with high-risk factors in lung cancer.

[0036] The classification method provided by the present invention is entirely implemented through a computer program. The computer program is executed to process data, train the designed multimodal fusion model, and then execute the trained multimodal fusion model on the computer to process lung CT images, medical records, and laboratory test data to obtain classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 The overall process of the multimodal fusion classification method is demonstrated, from data input to design and training model, multimodal fusion, visualization and statistical analysis;

[0038] Figure 2 Shows the structural diagram of the multimodal fusion model;

[0039] Figure 3 Demonstrates the steps for data screening in the early stages of an experiment. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation methods and drawings.

[0041] The specific processing steps are as follows:

[0042] Step 1: Build training data:

[0043] 1. Such as Figure 3 As shown, we initially collected 4351 patients before applying strict exclusion criteria. The exclusion criteria include: 1) patients with non-lung cancer stage I or II; 2) patients who did not undergo chest LDCT examination, laboratory test or pathological examination; 3) patients who lacked a final diagnosis of any of the four high-risk factors. In this study, patients with lung cancer with high-risk factors were defined as patients with one or more of the following conditions: wedge resection, poor differentiation, vascular invasion, pleural invasion, that is, four labels, and each label was divided into two categories. These four labels are also the prediction result categories to be output by the multimodal model of the present invention. Ultimately, 393 patients with lung cancer with high-risk factors and 720 patients with lung cancer without high-risk factors were included in the study.

[0044] 2. Among 393 high-risk patients, 282 underwent wedge resection, 109 had poorly differentiated tumors, 75 had visceral pleural involvement, and 39 had vascular invasion. We collected demographic characteristics, medical history, and laboratory test results from the patients' medical records.

[0045] 3. Four radiologists reviewed all chest CT images to ensure image quality and clarity and to mark lesion locations. Based on the radiologists' markings, 16 spatially contiguous chest CT images were selected for each patient. Because this study focused on patients with early-stage lung cancer with relatively small lesions, for patients with fewer than 16 CT images containing the lesion, we selected a total of 16 images centered around the lesion as model input. Medical record data included gender, age, smoking history, alcohol consumption history, history of lung disease, history of malignancy, family history, pulmonary nodule size, and pulmonary nodule location. We processed pulmonary nodule size and location into four variables to comprehensively represent information about pulmonary nodules, resulting in a final medical record dataset containing 11 variables. Laboratory test data included six variables: white blood cell (WBC), granulocyte ratio (GR), carcinoembryonic antigen (CEA), cytokeratin 19 fragment (CY), neuron-specific enolase (NSE), and progastrin-releasing peptide (Pro-GRP). These data were included based on previous studies demonstrating their association with cancer.

[0046] 4. To ensure balanced distribution across different classes, we randomly split the dataset into training and test sets in an 8:2 ratio, using five-fold cross-validation for both training and testing. The training set included 225 patients who underwent wedge resection, 87 patients with poorly differentiated tumors, 60 patients with pleural invasion, and 31 patients with vascular invasion. The test set included 57 patients who underwent wedge resection, 22 patients with poorly differentiated tumors, 15 patients with pleural invasion, and 8 patients with vascular invasion.

[0047] Step 2: Design and train a multimodal model:

[0048] 1. To integrate diverse medical data, our proposed model consists of three components, enabling it to adapt to and effectively fuse various data modalities: data mapping, feature extraction, and modality fusion. As a general approach, these three components are plug-and-play modules that can be freely switched to meet different task requirements. Furthermore, considering the optimized workflow, the model training process is performed end-to-end, eliminating the need to independently optimize each component. The data is processed and mapped to an input format. The data then passes through a feature extraction module for each modality to obtain a high-dimensional semantic feature representation. Next, a feature fusion module performs semantic feature integration analysis.

[0049] 2. In order to effectively extract semantic information from CT images, we use the encoder part of the Swin Transformer V2 model (Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., Wei, F., Guo, B.: Swin transformer v2: Scaling up capacity and resolution. In: 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11999–12009 (2022)) as the backbone of image feature extraction, that is, Figure 2 The image encoder shown in Figure 1. This model adopts a variant of the Transformer architecture, capturing local and global features of the image through layering and shifting mechanisms. The Swin Transformer V2 model uses a self-attention mechanism to handle long-range dependencies in the image and maintains sensitivity to the order of elements in the image through positional encoding. It gradually extracts and refines features through multiple encoder layers, each of which includes a multi-head self-attention module and a feedforward network to achieve efficient feature learning and representation. In addition, the model is designed with a scalable architecture that allows it to adapt to different tasks and computing resource requirements by increasing the number of layers and heads. Swin Transformer V2 also specifically introduces a shifting window mechanism that periodically shifts the processing window to obtain a more comprehensive image context, thereby enhancing the model's ability to capture information at different locations in the image. This design not only improves the model's sensitivity to multi-scale and multi-angle features, but also enables the model to exhibit higher computational efficiency when processing high-resolution images.

[0050] The image encoder pseudo code is as follows:

[0051]

[0052] The input feature vectors from medical record data and laboratory test data are respectively divided into three different weight vector matrices ω Q 、ω K 、ω V , multiplied together to form Q, K, V, where ω Q 、ω K 、ω V Is a fixed constant. In the text encoder, Q and K are first multiplied, then passed through a linear layer, and V is convolved and combined with the value that passed through the previous linear layer.

[0053] Through the image encoder and text encoder (Encoder), the image and text features are converted into fixed-length encoding vectors, which facilitates subsequent multimodal fusion and various processing.

[0054] 3. We collect patients’ medical records and laboratory test data and organize them into a structured and standardized data format, namely Figure 2 Normalization operation in the dataset. Medical records include discrete values ​​(such as gender and smoking history) and continuous values ​​(such as lung nodule size). Laboratory test data consists of continuous real numbers. In order to enhance the fitting and generalization capabilities of the model and speed up the convergence of the training process, we preprocessed both discrete and continuous data. For discrete values, we used one-hot encoding. For continuous data, we performed min-max normalization on each feature column and standardized it using the maximum and minimum values ​​of each feature. Subsequently, we connected the medical records and laboratory test data. Through the feature adapter, we aligned the feature records of the medical records and the laboratory test data to have the same length as the image features.

[0055] 4. Missing data is common in medical datasets. In this study, we excluded patients without chest CT images, so missing data occurred only in medical records and laboratory test data. To address this issue, we used polynomial interpolation to handle missing data.

[0056] 5. Combine the case data and laboratory test data and then combine them with the image data. The three types of data are combined to form the following Figure 2 The eigenvector shown is used for subsequent LADS operations.

[0057] 6. Multimodal fusion:

[0058] (1. We use an intermediate fusion strategy to fuse the high-dimensional features of each modality in the feature space, and then use a neural network to learn and predict the fused features. Figure 2 As shown in Figure 2, we propose a feature fusion method that combines linear attention and fast Fourier transform (FFT). The multimodal features are concatenated and first input into LADS to discover the dependencies between features. The output of the linear attention module is then input into the fast Fourier transform module to transform the spatial domain of the features into the frequency domain, thereby discovering the fused frequency domain information of the features.

[0059] The time and space complexity of traditional Attention is O(n 2 ), and use the linear attention mechanism (LinearAttention with Double Softmax) to reduce the time and space complexity to O(n) by utilizing the joint law of matrix multiplication and matrix transformation. The specific equation can be expressed as:

[0060]

[0061] where Q,K∈R n×d , which is the same as the definition in traditional Attention, Q (Query) is the query vector, which represents the information you want to find in other elements (i.e., Key), and K (Key) is the key vector, which represents the "identity" or feature of each element in the sequence, used to match the query vector (Query). The operation is the same as before, multiplying the feature vector by the weight vector matrix ω Q 、ω K , get Q, K, treat Q, K as n × d vectors, and then cluster them into m categories to get a matrix consisting of m cluster centers Represents the transpose of a matrix. Indicates that when the matrix When reversible, the inverse of the matrix, when the matrix When it is not invertible, the pseudo-inverse of the matrix.

[0062] (2. Fast Fourier Transform. Discrete Fourier Transform (DFT) plays a very important role in the field of digital signal processing. The Fast Fourier Transform algorithm uses the symmetry and periodicity of the signal to reduce the computational complexity of DFT from O(n 2 ) is reduced to O(nlogn). The inverse DFT transform is similar to the DFT and can also be efficiently calculated by the fast inverse Fourier transform (IFFT). Our implementation sets the LADS-processed feature here to x, where x∈R H×W×D , where H is the height, W is the width, and D is the depth, we first perform a 2D FFT on the spatial dimensions to transform it into the frequency domain:

[0063]

[0064] where F[·] is the 2D FFT. X is a complex tensor representing the spectrum of x, i.e. the processed form of feature x. We can then h×w×D , where h is the height, w is the width, D is the depth, and C is a complex number that is multiplied to adjust the spectrum of x:

[0065]

[0066] Where ⊙ is the element-wise product (also called the Hadamard product). The filter K has the same depth as X (D is the same), so it is called a global filter. Its code is as follows:

[0067] self.complex_weight=nn.Parameter(torch.randn(h,w,dim,2,dtype=torch.float32)*0.02)

[0068] Wherein h=4 and w=3 are fixed values.

[0069] Finally, we use IFFT to transform the modulation spectrum X back to the spatial domain to update the features, that is, update the feature x:

[0070]

[0071] (3. Perform deep feature compression and linear regression on the data after FFT operation. The feature compression pseudo code is as follows:

[0072]

[0073] The input features are mapped to a lower dimensional space through a linear layer. Subsequently, the Tanh activated features are mapped into weight vectors through another linear layer. These weights are normalized by the Softmax function and multiplied with the concatenated features to obtain weighted features.

[0074] (4. The four categories (wedge resection, poor differentiation, vascular invasion, pleural invasion) are identified by different multi-layer perceptrons, i.e. Figure 2 Each variable is classified into 0-1 categories to indicate whether the patient has the disease under the label variable.

[0075] Step 3: Set training parameters

[0076] 1. Use the Adam optimizer, setting the learning rate and maximum number of iterations. The training process includes forward propagation to calculate the loss function and backpropagation to update the neural network weights. Training continues until the preset convergence conditions are met.

[0077] 2. After training is complete, the final trained model parameters are saved for subsequent automatic pathology diagnosis and classification tasks. Using the feature vectors corresponding to labels, it is possible to identify whether the patient belongs to the four categories of wedge resection, poor differentiation, vascular invasion, and pleural invasion, with the labels representing 0 and 1 in the data respectively.

[0078] Compared with other AI-based methods for classifying high-risk factors for lung cancer (such as visceral pleural involvement and unknown lymph node status), the present invention can predict multiple high-risk factors simultaneously. In addition, the research cohort of the present invention is larger and more comprehensive. This is the first multi-label study on high-risk factors for lung cancer, with a cohort of more than one thousand cases. We calculated AUC, sensitivity, specificity, and accuracy. The 95% confidence intervals for sensitivity and specificity were calculated using the "exact" Clopper-Pearson confidence interval. The 95% confidence interval for AUC was obtained using the Delong method (DeLong, ER, DeLong, DM, Clarke-Pearson, DL: Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics 44(3), 837–845). When comparing AUCs, the P value was also calculated using the Delong test. The ROC curve was generated by changing the threshold of the model output prediction. The threshold is used to binarize the real-number output of the model. Different thresholds may result in different binary model predictions for each image, leading to different sensitivity and specificity values ​​on the test dataset for plotting the ROC curve.

[0079] This invention has several limitations. First, the sensitivity of multimodal models is relatively low, resulting in a high false-negative rate. This is particularly problematic in cancer diagnosis, where missing high-risk cases can have serious consequences. Second, deep learning is a data-driven analytical technique that lacks transparent reasoning. To mitigate this issue, the present invention employs visualization techniques and conducts comprehensive ablation experiments. However, these methods focus on interpreting the results rather than analyzing the model's internal reasoning process.

[0080] In summary, this multimodal approach is consistent with NCCN guidelines, which emphasize the complexity of treatment decisions for lung cancer patients and the importance of considering a range of high-risk factors. By providing a comprehensive risk assessment, this invention can assist in preoperative planning and adjuvant treatment decisions, ultimately contributing to personalized patient care. In summary, this invention adds to the growing evidence that deep learning models, particularly those that leverage multimodal data, play an important role in the future of oncology diagnostics.

Claims

1. A lung cancer high-risk factor classification method based on multimodal fusion, characterized by: The following steps are involved: Step 1: Build a training dataset: (1) The collected patient data were divided into lung cancer patients with high-risk factors and lung cancer patients without high-risk factors; lung cancer patients with high-risk factors were defined as having one or more of the following: wedge resection, poor differentiation, vascular invasion, and pleural invasion. These four variables were set as four labels, and these four labels were classified into 0-1 binary categories; (2) Screening patients' chest CT images with clearly marked lesion locations; based on the markings, 16 consecutive chest CT images were selected for each patient, centered on the lesion area; medical records included gender, age, smoking history, drinking history, history of lung disease, history of malignant tumors, family history, size of lung nodules, and location of lung nodules; (3) Laboratory test data included six variables: white blood cell (WBC), granulocyte ratio (GR), carcinoembryonic antigen (CEA), cytokeratin 19 fragment (CY), neuron-specific enolase (NSE), and progastrin-releasing peptide (Pro-GRP); (4) The training dataset was randomly divided into a training set and a test set in a ratio of 8:2, and five-fold cross validation was used for training and testing; Step 2: Design and train a multimodal model: (1) The multimodal model includes a data mapping module, a feature extraction module, and a modality fusion module; (2) The data mapping module is used to normalize the case data and laboratory test data separately, and then combine them as the input of the text encoder, using polynomial interpolation to handle missing data, and using the CT image as the input of the image encoder; (3) Feature extraction module, which is used to convert case data and laboratory test data into feature vectors through the text encoder, and combine them with the feature vectors converted from the CT image data through the image encoder as the input of the LADS module; the encoder part of the Swin Transformer V2 model is used as the backbone network of the image encoder; (4) Modal fusion module, which uses an intermediate fusion strategy to fuse the high-dimensional features of each modality in the feature space, and then uses a neural network to learn and predict the fused features. The features output by the feature extraction module are first input into the LADS module, and then the fast Fourier transform is used. First, 2DFFT is used in the spatial dimension, and then the filter is used to perform Hadamard product, and finally the fast inverse Fourier transform is used to process the features. The entire process from LADS input to fast inverse Fourier transform output is stacked. (5) For the features processed by the modal fusion module, deep feature compression and linear regression are performed. After the output, the four categories to be identified are processed by multi-layer perceptrons respectively, and 0-1 binary classification is performed respectively. The multi-layer perceptrons include MLP1, MLP2, MLP3 and MLP4; (6) Using the training data set to train the multimodal model and save the final trained model parameters, the trained multimodal model is obtained and used to receive actual chest CT images, laboratory test data, and medical record data and output classification results; Among them, LADS uses the joint law of matrix multiplication and matrix transformation to reduce the time and space complexity to , the specific equation can be expressed as: ; in , is the query vector, which represents the information you want to find in other elements. is a key vector representing the "identity" or feature of each element in the sequence, which is used to match the query vector Q. Q and K are generated based on the input of the LADS module. Considered as vector, and then clustered into Class, obtained by The matrix composed of cluster centers , represents the transpose of the matrix, Indicates that when the matrix When reversible, the inverse of the matrix, when the matrix When it is not invertible, the pseudo-inverse of the matrix.

2. The method according to claim 1, wherein Perform a 2D FFT on the spatial dimensions to transform it into the frequency domain: ; in is a 2D FFT, is a complex tensor, representing The spectrum of the LADS-processed feature is set as ,here , where H is the height, W is the width, and D is the depth; Then, by With learnable filters Multiply to adjust The spectrum: ; in is the element-wise product, also known as the Hadamard product, filter and have the same dimension, called global filters; Finally, the modulation spectrum is transformed into Transform back to the spatial domain and update the features : 。 3. An information processing device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to claim 1 or 2 is implemented.

Citation Information

Patent Citations

  • Early-stage lung cancer classification method based on ResNet and Transformer

    CN118072081A

  • Construction method of depression risk prediction model

    CN118430820A