Lung cancer subtype classification method based on visual language concept bottleneck model
Through the bottleneck model based on the concept of visual language, the problem of lack of interpretability of the NSCLC subtype classification model is solved, and more accurate and reliable classification of lung cancer subtypes is achieved, and a transparent decision-making process is provided.
Patent Information
- Application Number
- CN202510098788.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-24
AI Technical Summary
The existing NSCLC subtype classification model lacks interpretability and it is difficult to understand the decision-making process of the model, resulting in unreliable classification results.
A lung cancer subtype classification method based on the concept bottleneck model of visual language is proposed. By obtaining CT images and diagnostic reports, data preprocessing and visual feature extraction, medical concept prediction is used using visual language models, and concept bottleneck model is constructed to achieve interpretable subtype classification.
Enhanced interpretability of the model, provides a transparent decision-making process, and improves the accuracy of medical concept prediction and the reliability of subtype classification.
Smart Images

Figure CN120198907A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to a method for classifying lung cancer subtypes based on a visual language concept bottleneck model. Background Art
[0002] With the development of artificial intelligence technology, using deep learning methods to automatically classify lung cancer pathological subtypes through CT images has become a research hotspot. A series of scientific studies have integrated deep learning algorithms with CT image analysis and explored in the field of NSCLC pathological subtype classification, achieving certain results. In traditional NSCLC subtype classification methods based on CT images, by training a deep learning model on tumor CT images and corresponding pathological subtype labels, it can identify and learn fine features and complex associations that are difficult to capture by human vision, so as to classify NSCLC subtypes. However, most existing subtype classification models adopt complex network structures and train the models in an end-to-end manner, which belong to black-box models. Such models have serious defects: their interpretability is lacking, and it is difficult for researchers to understand how the models work. Even if they can accurately describe their network structures, researchers still cannot explain the decision-making process of the models, resulting in unreliable classification results of the models. Aiming at the lack of interpretability of existing NSCLC subtype classification methods, the present application proposes a method for classifying lung cancer subtypes based on a visual language concept bottleneck model. Summary of the Invention
[0003] Based on this, in view of the lack of interpretability of existing NSCLC subtype classification methods, it is necessary to propose a method for classifying lung cancer subtypes based on a visual language concept bottleneck model.
[0004] The present application relates to a method for classifying lung cancer subtypes based on a visual language concept bottleneck model, including:
[0005] Obtaining non-small cell lung cancer CT images and non-small cell lung cancer diagnosis reports;
[0006] Performing data preprocessing on the collected non-small cell lung cancer CT images;
[0007] Extracting visual features of the preprocessed non-small cell lung cancer CT images based on a three-dimensional visual adapter;
[0008] Performing image concept extraction using the visual features;
[0009] Verifying the extracted image concepts based on a visual language concept bottleneck model;
[0010] Constructing a loss function to form a loss training model;
[0011] Train the model based on the loss, and correct the parameters of the vision-language concept bottleneck model to obtain the pathological subtype classification result.
[0012] This application relates to a method for classifying lung cancer subtypes based on a vision-language concept bottleneck model. A vision-language model suitable for CT images is constructed to make reliable medical concept predictions, and a concept bottleneck model is constructed based on medical concepts, aiming to achieve interpretable NSCLC pathological subtype classification in the absence of dense concept annotations. This application designs an innovative vision-language concept bottleneck model to solve the problem of the lack of dense concept annotations in the NSCLC dataset. First, this application utilizes the cross-modal ability of the vision-language model to provide reliable medical concept predictions for CT images, effectively applying the concept bottleneck model to the NSCLC pathological subtype classification task and enhancing the interpretability of the model. Specifically, this application designs a three-dimensional vision adapter to fully exploit the spatial dependence relationship between slices of three-dimensional CT images, enabling the general vision-language model trained based on two-dimensional natural images to be applicable to three-dimensional CT images. This method can capture more comprehensive visual features and improve the accuracy of medical concept prediction. In addition, to enhance the performance of the general vision-language model in the field of imaging, this application proposes a concept extraction and alignment strategy to fine-tune the learnable adapter. Standard medical concepts are extracted from the original diagnostic reports through a large language model, and the medical concepts are aligned with visual features one by one, thus solving the confusion problem caused by differential expressions in diagnostic reports and establishing an accurate association between visual features and medical concepts. Description of the Drawings
[0013] Figure 1 It is the method flow intention diagram of a method for classifying lung cancer subtypes based on a vision-language concept bottleneck model provided by an embodiment of this application. Detailed Embodiments
[0014] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to clarify this application and are not used to limit this application.
[0015] This application provides a method for classifying lung cancer subtypes based on a vision-language concept bottleneck model.
[0016] As Figure 1 shown, in an embodiment of this application, a method for classifying lung cancer subtypes based on a vision-language concept bottleneck model includes:
[0017] S100, obtain non-small cell lung cancer CT images and non-small cell lung cancer diagnostic reports.
[0018] S200, preprocess the collected CT images.
[0019] S300, extract the visual features of the preprocessed CT images based on a 3D vision adapter.
[0020] S400, perform concept extraction using the visual features.
[0021] S500, confirm the extracted concepts based on a visual language concept bottleneck model.
[0022] S600, construct a loss function to form a loss training model.
[0023] S700, based on the loss training model, correct the parameters of the visual language concept bottleneck model to obtain the pathological subtype classification result.
[0024] Specifically, deep learning models have shown broad application prospects in the field of NSCLC pathological subtype classification. However, most existing subtype classification models adopt end-to-end black-box models, lacking interpretability. As an emerging technical means, the concept bottleneck model has significant advantages in enhancing the interpretability of deep learning. The basic idea of this model is that given an input, a set of intermediate specified concepts are first predicted, and then this set of concepts is directly used to predict the target.
[0025] For traditional concept bottleneck models, it is necessary to confirm the predicted concepts. However, the lung cancer subtype classification method based on the visual language concept bottleneck model used in this embodiment uses a visual language model to establish the semantic association between CT images and diagnostic reports, thereby realizing automatic concept prediction and reducing the high cost of medical concept annotation work.
[0026] The visual language concept bottleneck model can provide a transparent decision-making process based on medical concepts, enabling humans to understand how the model makes decisions, enhancing the interpretability of the model, and realizing transparent and reliable lung cancer subtype classification. The visual language concept bottleneck model has good classification performance, and their decision-making processes are transparent, which can provide decision-making bases that are understandable to doctors and patients.
[0027] This embodiment relates to a method for classifying lung cancer subtypes based on a visual language concept bottleneck model. A visual language model suitable for CT images is constructed to perform reliable medical concept prediction, and a concept bottleneck model is constructed based on medical concepts, aiming to achieve interpretable NSCLC pathological subtype classification in the absence of dense concept annotations. This application designs an innovative visual language concept bottleneck model to solve the problem of the lack of dense concept annotations in the NSCLC dataset. This application first utilizes the cross-modal ability of the visual language model to provide reliable medical concept prediction for CT images, effectively applying the concept bottleneck model to the NSCLC pathological subtype classification task and enhancing the interpretability of the model. Specifically, this application designs a three-dimensional visual adapter to fully exploit the spatial dependence relationship between slices of three-dimensional CT images, enabling the general visual language model trained based on two-dimensional natural images to be applicable to three-dimensional CT images. This method can capture more comprehensive visual features and improve the accuracy of medical concept prediction. In addition, to enhance the performance of the general visual language model in the field of imaging, this application proposes a concept extraction and alignment strategy to fine-tune the learnable adapter. Standard medical concepts are extracted from the original diagnostic reports through a large language model, and the medical concepts are aligned with visual features one by one, thus solving the confusion problem caused by differential expressions in the diagnostic reports and establishing an accurate association between visual features and medical concepts.
[0028] In one embodiment of this application, S100 includes:
[0029] S111, request data from the dataset based on the data request.
[0030] S112, determine whether the data request is a request to obtain NSCLC CT images.
[0031] S113, if the data request is a request to obtain NSCLC CT images, then call the public dataset in the dataset to obtain the NSCLC CT images in the public dataset.
[0032] S114, if the requested dataset is not a request to obtain NSCLC CT images, then determine that the data request is a request to obtain an NSCLC diagnostic report, and obtain the NSCLC diagnostic report in the dataset.
[0033] S115, obtaining the NSCLC diagnostic report in the dataset includes:
[0034] Determine whether the target of the requested data request is a private dataset in the dataset.
[0035] S116, if the target of the data request is a private dataset in the dataset, then call the private dataset in the dataset to obtain the NSCLC diagnostic report in the private dataset.
[0036] S117. If the target of the requested dataset is not the private dataset in the dataset, then call the public dataset in the dataset to obtain the non-small cell lung cancer diagnosis reports in the public dataset.
[0037] In an embodiment of the present application, S100 further includes:
[0038] S121. Create a dataset.
[0039] S122. Receive the inclusion criteria for the dataset for pathological subtype prediction.
[0040] S123. Based on the inclusion criteria, include non-small cell lung cancer CT images and the corresponding non-small cell lung cancer diagnosis reports in the dataset.
[0041] S124. Receive the data exclusion criteria for the dataset for pathological subtype prediction.
[0042] S125. Select a non-small cell lung cancer CT image in the dataset.
[0043] S126. Based on the data exclusion criteria, determine whether the image quality of the selected non-small cell lung cancer CT image in the dataset is less than the preset criteria in the data exclusion criteria.
[0044] S127. If the image quality of the selected non-small cell lung cancer CT image in the dataset is less than the preset criteria in the data exclusion criteria, then delete both the selected non-small cell lung cancer CT image and the corresponding non-small cell lung cancer diagnosis report, and return to select a non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected.
[0045] S128. If the image quality of the selected non-small cell lung cancer CT image in the dataset is greater than or equal to the preset criteria in the data exclusion criteria, then return to select a non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected.
[0046] This embodiment relates to a method for obtaining non-small cell lung cancer CT images and non-small cell lung cancer diagnosis reports.
[0047] Obtain the collected non-small cell lung cancer CT image data and diagnosis reports, and perform relevant annotations.
[0048] The dataset for NSCLC pathological subtype prediction comes from two NSCLC collections, including a private dataset and a public dataset.
[0049] The inclusion criteria include primary NSCLC with pathological diagnosis.
[0050] Exclusion criteria include large or poor-quality artifacts in CT images. Other NSCLC than squamous cell carcinoma and adenocarcinoma.
[0051] Among them, the private dataset was collected from eligible patients diagnosed or treated in a domestic hospital. Finally, 758 cases (422 adenocarcinoma cases and 336 squamous cell carcinoma cases) were collected as the internal dataset, and this dataset has obtained KQ. The CT image slice thickness of the Regional Chest Hospital is 1.0 mm, and the tumor regions were manually marked by two radiologists.
[0052] The public dataset was provided by the Cancer Imaging Archive and included 211 original cases from SA Medical College and AA System. After reviewing the data according to the same inclusion criteria, 140 cases were found with CT image slice thicknesses between 0.625 and 3.0 mm.
[0053] The database includes a public dataset and a private dataset. The public dataset includes multiple groups of public case data, and each group of public case data includes a corresponding non-small cell lung cancer CT image and a non-small cell lung cancer diagnosis report. The private dataset includes multiple groups of private case data, and each group of private case data includes a corresponding non-small cell lung cancer CT image and a non-small cell lung cancer diagnosis report.
[0054] In an embodiment of the present application, S200 includes:
[0055] S211, select a non-small cell lung cancer CT image.
[0056] S212, resample the non-small cell lung cancer CT image using the trilinear interpolation algorithm.
[0057] S213, adjust the pixels of the resampled non-small cell lung cancer CT image.
[0058] S214, screen out the non-small cell lung cancer CT images with pixel intensity values in the range of greater than or equal to -1000 Hounsfield and less than or equal to 400 Hounsfield, and use the non-small cell lung cancer CT images with pixel intensity values in the range of greater than or equal to -1000 Hounsfield and less than or equal to 400 Hounsfield as the first preprocessed non-small cell lung cancer CT images.
[0059] S215, return the selected non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected.
[0060] In an embodiment of the present application, S200 further includes:
[0061] S221, select a first preprocessed non-small cell lung cancer CT image.
[0062] S222, crop the first preprocessed non-small cell lung cancer CT image.
[0063] S223, obtain the cropped first preprocessed non-small cell lung cancer CT image.
[0064] S224, adjust the cropped first preprocessed non-small cell lung cancer CT image to match the pixel standard, and use the obtained CT image as the second preprocessed non-small cell lung cancer CT image.
[0065] S225, return the selected one first preprocessed non-small cell lung cancer CT image until all the first preprocessed non-small cell lung cancer CT images are selected.
[0066] S226, incorporate all the second preprocessed non-small cell lung cancer CT images into the data preprocessing set of CT images.
[0067] Specifically, at least one non-small cell lung cancer CT image corresponds to a non-small cell lung cancer diagnosis report. These non-small cell lung cancer CT images form a three-dimensional image, and this three-dimensional image corresponds to a non-small cell lung cancer diagnosis report.
[0068] After preprocessing all the non-small cell lung cancer CT images, they are incorporated into the data preprocessing set of CT images.
[0069] Perform data preprocessing on the collected CT image data to make it conform to the standard input of the vision-language model. In this embodiment, the CT images of 758 non-small cell lung cancer patients are collected as the sample set, and the trilinear interpolation algorithm is used to resample the lung CT image data.
[0070] Adjust the voxel size to 1mm×1mm×1mm, and clip its pixel intensity to -1000HU to 400HU; subsequently, to exclude the regions irrelevant to non-small cell lung cancer in the CT image, crop the CT image so that it only contains the chest region, and adjust its pixel size to 384×384×288 pixels. 384×384×288 pixels is the pixel standard in S224.
[0071] The preprocessed CT image of the i-th case is denoted as x i , x i The corresponding class label is denoted as y i . And y i ∈{0, 1}. In the study of NSCLC pathological subtype classification by this method, the tumor CT image of adenocarcinoma is positive (y i = 1), while the tumor CT image of squamous cell carcinoma is negative (y i = 0).
[0072] In one embodiment of the present application, S300 includes:
[0073] S311, call a local fully connected neural network and a global encoder.
[0074] S312, generate a three-dimensional visual adapter.
[0075] S313, select a non-small cell lung cancer CT image from the data preprocessing set.
[0076] S314, use the three-dimensional visual adapter to determine the spatial positions between CT slices.
[0077] S315, return a non-small cell lung cancer CT image from the selected data preprocessing set until all non-small cell lung cancer CT images are selected.
[0078] S316, obtain a dataset of the spatial positions between CT slices for the data preprocessing set.
[0079] In one embodiment of the present application, S300 further includes:
[0080] S321, generate a three-dimensional CT image spatial information set network based on the local fully connected neural network of the three-dimensional visual adapter and the global encoder of the three-dimensional visual adapter.
[0081] S322, call the dataset of the spatial positions between CT slices.
[0082] S323, select a spatial position data from the dataset of the spatial positions between CT slices.
[0083] S324, incorporate the spatial position data into the three-dimensional CT image spatial information set network.
[0084] S325, return a spatial position data from the selected dataset of the spatial positions between CT slices until all spatial position data are selected.
[0085] S326, obtain the three-dimensional CT image spatial data for the non-small cell lung cancer diagnosis report.
[0086] In one embodiment of the present application, S300 further includes:
[0087] S331, select a non-small cell lung cancer diagnosis report.
[0088] S332, call the three-dimensional CT image spatial data corresponding to the selected non-small cell lung cancer diagnosis report.
[0089] S333, call the trained linear projection module in the three-dimensional visual adapter.
[0090] S334. Extract the visual features of the spatial data of 3D CT images based on a linear projection module.
[0091] It can be understood that the spatial position data set in S316 is the spatial position data, and the spatial position data is the spatial position data of the three-dimensional virtual spatial positions between CT slices.
[0092] The spatial data of the 3D CT images in S326 is the data obtained after incorporating the spatial position data into the three-dimensional CT image spatial information aggregation network.
[0093] Design a 3D visual adapter to extract visual features, and adopt a local fully connected neural network and a global Transformer encoder to comprehensively capture the spatial dependency relationships between CT slices.
[0094] Vision-language models represented by CLIP (Contrastive Language-Image Pretraining) are only trained on 2D natural images and cannot capture the comprehensive visual features of 3D CT images, which poses a huge challenge to medical concept prediction. For example, identifying "holes" and "pleural retraction" requires the integration and analysis of multiple consecutive slices. Since CLIP cannot learn the context dependency relationships between slices, the concept prediction results are unreliable. Therefore, in order to improve the reliability of medical concept prediction, it is necessary to adapt the visual representation of CLIP to the radiology field. Specifically, concepts within a tumor, such as "necrosis" and "spiculation", are related to short-range dependency relationships in adjacent slices, while concepts between tumors, such as "pleural retraction", involve long-range dependency relationships across the entire CT image. Thus, this embodiment designs a 3D visual adapter to capture short-range and long-range dependency relationships from local and global perspectives to achieve comprehensive visual feature modeling.
[0095] Specifically, for a given 3D CT image of the lungs, this embodiment first divides it into consecutive slices along the axial direction and encodes these slices into slice-level features using a frozen CLIP image encoder. After that, instead of simply concatenating these features, this embodiment introduces a local fully convolutional network on top of these features to capture short-range dependency relationships. It should be noted that this layer is different from a regular fully convolutional network because it restricts the aggregation calculation within a local window rather than the global scope. Specifically, a local window is constructed around each slice feature, and its N nearest neighbor slices are incorporated into it. Within each window, a fully connected layer is added to enhance the representation of the local slice features. This process can be summarized as follows:
[0096]
[0097] where S i 、 represent the original slice features and the enhanced local slice features respectively, where i represents the number of the CT slice, and N i is the N nearest neighbors of the slice. Such a convolutional operation has a local receptive field, thus ensuring effective modeling of short-range dependencies.
[0098] Next, in order to achieve a wider receptive field between different slices, this embodiment further utilizes a global Transformer encoder to explicitly model long-range spatial dependencies. The Transformer encoder adopts a self-attention mechanism, which can focus on the pairwise correlations between each slice in the three-dimensional CT image, and the addition of position information further enhances the utilization of slice spatial information.
[0099] Specifically, in order to model the relationships between all slices, this embodiment uses a trainable linear projection to map the local slice features to a latent dimensional embedding space. Then, in order to encode the slice spatial information, this embodiment learns specific position embeddings and adds them to the slice embeddings to retain the position information, as follows:
[0100]
[0101] where E is the slice embedding projection, E pos represents the position embedding. The Transformer encoder consists of L layers of multi-head self-attention (MSA) and multi-layer perceptron (MLP) modules (formulas (3), (4)). Therefore, the output of each layer can be written in the following form:
[0102] I′ i = MSA(LN(I i-1 )) + I i-1 Formula (3)
[0103] I i = MLP(LN(I′ i )) + I′ i Formula (4)
[0104] where LN(·) represents the layer normalization operation, and I i is the encoded image representation of the i-th layer. By capturing visual features containing local and global spatial dependencies, the three-dimensional visual adapter can more comprehensively understand the complex spatial information of the three-dimensional CT image in medical concept prediction.
[0105] In an embodiment of the present application, S400 includes:
[0106] S410, receiving an instruction table. The instruction table is about the concept extraction task and the expected output structure.
[0107] S420, based on the instruction table, obtain a detailed list of the extracted concepts and at least one concept expression table.
[0108] S430, using radiology report examples, determine clear preset values and consistency preset values for concept labels.
[0109] A concept extraction and alignment strategy is proposed to fine-tune the learnable adapter, solve the confusion problem caused by differential expressions in diagnostic reports, and establish a fine-grained connection between visual features and medical concepts.
[0110] After adapting the visual representation through a 3D visual adapter, the present embodiment now further adjusts the learnable adapter to improve the adaptability of CLIP to the medical field, thereby achieving more accurate medical concept prediction. Previous studies have shown that adapter fine-tuning has many benefits for domain adaptation through the alignment between domain-specific images and text pairs. However, this direct alignment strategy is difficult to effectively align the association between subtle medical concepts and corresponding visual features.
[0111] There are two problems:
[0112] First, some medical concepts have multiple text descriptions in radiology reports, which may cause confusion in the concept alignment process.
[0113] Second, due to the limitations of the vocabulary learned by CLIP, this strategy cannot solve the adaptation problem between the vocabulary of CLIP and medical concepts in the radiology field. Therefore, the present embodiment designs an automatic concept extraction module to solve the problem of expression diversity and applies a concept alignment method to associate medical concepts in the radiology field with visual features.
[0114] In a clinical environment, different radiologists may use different terms for the same medical concept. For example, the concept of "pleural traction" may be described as "traction" or "adhesion" in radiology reports. Therefore, in the joint feature space, aligning various expressions of medical concepts related to the same visual feature may cause confusion between competing concept embeddings, seriously reducing the performance of medical concept prediction. To solve the problem of expression diversity, the present embodiment designs a concept extraction module to collect predefined medical concepts from radiology reports through large language models (LLMs).
[0115] Specifically, the present embodiment designs a prompt template for the large language model, which mainly consists of three parts.
[0116] General instructions on the concept extraction task and the expected output structure.
[0117] A detailed list of concepts to be extracted, as well as various concept expressions.
[0118] An example of a radiology report and its corresponding output to ensure the clarity and consistency of concept tagging. Leveraging the rich knowledge of the large language model and the effective constraints of the prompt template, this embodiment converts the original radiology report containing various concept expressions into a set of unified medical concepts.
[0119] The medical concepts obtained by the large language model are finally represented as a concept vector C = (c1,..., c k ,..., c n ), where c k ∈ {1, 0}, indicating the presence or absence of the k-th medical concept in the radiology report.
[0120] After extracting the medical concepts, this embodiment leverages the association between medical concepts and visual features to extend the vocabulary of CLIP to the radiology domain. Since CLIP is not trained on medical data, it often performs poorly in processing medical image-text pairs due to vocabulary limitations.
[0121] This embodiment aligns the medical concepts in the radiology domain with the corresponding visual features to fine-tune all learnable adapters. Specifically, radiology reports usually mention the presence or absence of specific concepts simultaneously, so the vectorized medical concepts can be converted into a predefined template: "There is / are(no)[CONCEPT]", and text embeddings are obtained through the frozen CLIP text encoder. Then, this embodiment attaches a two-layer multi-layer perceptron (MLP) on top of these concept embeddings as a simpler text adapter. Research has shown that this lightweight adapter has advantages in adapting models similar to CLIP. Given a medical concept embedding encoded by CLIP, this embodiment obtains an adaptive concept embedding through a two-layer MLP:
[0122]
[0123] where W1 and W2 represent the weights of the two layers in the adapter. Finally, this embodiment fine-tunes all learnable adapters through the concept contrast loss. When aligning concept prompts with image features, it is difficult for the model to distinguish the presence or absence of concepts because their expressions may be very similar. Therefore, instead of setting a threshold based on the similarity between concept prompts and images, this embodiment provides the model with both the original prompt and its negative example for each concept and compares their probabilities to align them with visual features. More specifically, for each medical concept, this embodiment compares the adaptive visual features with the features of the original prompt and the negative example and calculates the similarity score. The concept contrast loss is as follows:
[0124]
[0125] where τ1 is a scaling temperature parameter used to balance the learning and generalization capabilities of the model. represents the cosine similarity between the adaptive CT image features and the concept features. The concept contrast loss aims to increase the similarity between the adaptive visual features I a and the original concept features, while reducing the similarity between them and the negative example features, so as to perform more discriminative learning on medical concepts. In summary, through steps three and four, this embodiment achieves successful visual adaptation and effective adapter fine-tuning, and constructs a domain-adaptive vision-language model (DA-VLM). The DA-VLM can effectively adapt CLIP to the radiology field by leveraging the complete spatial context and fine-grained semantic associations, thereby providing reliable medical concept predictions for interpretable non-small cell lung cancer subtype classification.
[0126] In one embodiment of the present application, S500 includes:
[0127] Construct a vision-language concept bottleneck model.
[0128] The method for constructing the vision-language concept bottleneck model includes:
[0129] S510, using fine-grained semantic associations to construct a domain-adaptive vision-language model.
[0130] S520, based on the domain-adaptive vision-language model, train a vision-language concept bottleneck model for non-small cell lung cancer subtype classification.
[0131] Implement interpretable non-small cell lung cancer subtype classification based on the vision-language concept bottleneck model.
[0132] After constructing the DA-VLM, the next step is to train an interpretable concept bottleneck model for non-small cell lung cancer subtype classification. Given an input three-dimensional CT image, this embodiment expects the model to provide a classification result, a concept vector indicating the relevant medical concepts that led to the classification, and the importance of each concept.
[0133] Specifically, this embodiment first uses the DA-VLM to create a bottleneck layer, denoted as which maps the input image i to an n-dimensional vector, where n is the number of medical concepts. In particular, this embodiment calculates the probability of the existence of each predefined medical concept in the three-dimensional CT image. This embodiment defines positive and negative example cues for all medical concepts and calculates the probabilities P pos (c i ), P neg (c i) By applying the softmax function to the probabilities of positive and negative examples, this embodiment calculates the final probability of the existence of the medical concept P(c i ) and constructs a bottleneck layer.
[0134] P(c i ) = max(softmax(P pos (c i ))), softmax(P neg (c i ))) Equation (7)
[0135] After that, in view of the fact that the similarity score between the adaptive visual features and the medical concept embedding can already intuitively reflect the evaluation of each medical concept by the model, this embodiment draws on the idea of human experts' comprehensive multi-dimensional medical concepts for final diagnosis and introduces a linear classification layer:
[0136]
[0137] Among them, concat(,) represents the concatenation operation, and W is the weight in the linear layer, which inherently reflects the importance of the contribution of each medical concept to the overall classification of non-small cell lung cancer subtypes. After the linear classifier is trained, the importance of each concept to the classification target can be understood by analyzing the learned weights. This linear layer accurately predicts the final subtype category of non-small cell lung cancer by aggregating the similarity scores from all relevant concepts, thus simulating the decision-making process of human experts in complex medical diagnosis and improving the scientificity and reliability of the prediction.
[0138] In an embodiment of the present application, S600 includes:
[0139] Construct a loss function and train the model jointly with the concept alignment loss and the subtype classification loss.
[0140] Finally, the entire bottleneck classification model is trained in an end-to-end manner. This embodiment optimizes a joint objective that covers both the cross-entropy loss for final classification and the total concept contrast loss:
[0141]
[0142] During the training process, the cross-entropy loss can accurately quantify the class probabilities output by the model The degree of deviation from the true label y. By minimizing this loss, the model can learn to more accurately distinguish different non-small cell lung cancer subtypes, enabling the prediction results to approach the true situation to the greatest extent and improving the accuracy and reliability of classification. The total concept contrast loss Lcc focuses on enhancing the model's learning of the internal correlation between medical concepts and visual features. For each independent medical concept, this method carefully calculates its corresponding loss value, enabling the model to capture the visual features closely corresponding to different medical concepts with higher precision, thereby improving the model's prediction performance for medical concepts. Finally, linearly combining these two parts of the loss according to the established optimization strategy and multiplying by the appropriate weight coefficient α can effectively avoid a certain loss function dominating the optimization direction during the training process and ensure the collaborative optimization of the two losses during the model iteration process. During the iterative loop process of model training, all learnable parameters are updated based on the gradient-based backpropagation algorithm.
[0143] Analysis of medical concept prediction and pathological subtype classification results.
[0144] To quantitatively evaluate the performance of the proposed subtype classification method, the present invention is compared with other NSCLC pathological subtype classification methods on both public and private datasets.
[0145] In addition, to evaluate the performance of the medical concept prediction of this method, the present invention annotates medical concepts (116 cases) for the public dataset under the guidance of radiologists and predicts 7 common lung cancer-related medical concepts.
[0146] In an embodiment of the present application, S700 includes:
[0147] The parameter correction of the vision-language concept bottleneck model is realized based on the judgment of the degree of conformity between the sample prediction value and the true value.
[0148] Based on whether the sample prediction value and the true value match, 4 results can be obtained:
[0149] TP (True positive), the sample prediction value matches the true value and both are positive, that is, true positive.
[0150] FP (False positive), the sample prediction value is positive while the true value is negative, that is, false positive.
[0151] FN (False negative), the sample prediction value is negative while the true value is positive, that is, false negative.
[0152] TN (True negative), the sample prediction value matches the true value and both are negative, that is, true negative.
[0153] The ROC is the receiver operating characteristic curve, a common method for evaluating the performance of a classifier:
[0154] The abscissa is the False positive rate (FPR).
[0155] FPR = FP / [FP + TN] represents the probability of misclassifying negative samples as positive samples, the false alarm rate.
[0156] The ordinate is the True positive rate (TPR), and TPR = TP / [TP + FN] represents the probability of correctly classifying positive samples, the hit rate.
[0157] In this study, the classification results of non-small cell lung cancer tissue subtypes were evaluated based on the following four metrics:
[0158] Accuracy Acc = [TP + TN] / [TP + TN + FP + FN].
[0159] Sensitivity Sen = TP / [TP + FN].
[0160] Specificity Spe = TN / [TN + FP].
[0161] AUC, the area under the ROC curve, ranges from 0 to 1, and the closer it is to 1, the better the performance of the classifier.
[0162] The technical features of the above-described embodiments can be combined arbitrarily, and there is no limitation on the execution order of the method steps. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should all be considered as within the scope described in this specification.
[0163] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A lung cancer subtype classification method based on a visual language concept bottleneck model, characterized in that: include: Obtain CT images of non-small cell lung cancer and diagnosis reports of non-small cell lung cancer; Preprocessing the collected non-small cell lung cancer CT images; Extract visual features of preprocessed non-small cell lung cancer CT images based on 3D visual adapter; Extract image concepts using visual features; Based on the visual language concept bottleneck model, the extracted image concepts are confirmed; Construct a loss function and form a loss training model; Based on the loss training model, the parameters of the visual language concept bottleneck model are modified to obtain the pathological subtype classification results.
2. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 1 is characterized in that: The obtaining of non-small cell lung cancer CT images and non-small cell lung cancer diagnosis reports includes: requesting data from the dataset based on the data request; Determine whether the data request is a request for obtaining a CT image of non-small cell lung cancer; If the data request is a request for obtaining a CT image of non-small cell lung cancer, then a public data set in the data set is called to obtain the CT image of non-small cell lung cancer in the public data set; If the requested data set is not for obtaining a CT image of non-small cell lung cancer, then the data request is determined to be a request for obtaining a diagnosis report of non-small cell lung cancer, and the diagnosis report of non-small cell lung cancer in the data set is obtained; The non-small cell lung cancer diagnosis reports in the acquisition dataset include: Determine whether the target of the data request is a private data set in the data set; If the target of the data request is a private data set in the data set, the private data set in the data set is called to obtain the non-small cell lung cancer diagnosis report in the private data set; If the target of the requested dataset is not a private dataset in the dataset, the public dataset in the dataset is called to obtain the non-small cell lung cancer diagnosis report in the public dataset.
3. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 2 is characterized in that: Before obtaining the non-small cell lung cancer CT images and non-small cell lung cancer diagnosis report, it also includes: Create a dataset; Inclusion criteria for datasets receiving NSCLC pathological subtype prediction; Based on the inclusion criteria, non-small cell lung cancer CT images and non-small cell lung cancer diagnosis reports corresponding to non-small cell lung cancer CT images were included in the data set; Data exclusion criteria for datasets receiving prediction of pathological subtypes of non-small cell lung cancer; Select a non-small cell lung cancer CT image in the dataset; Based on the data elimination criteria, determining whether the image quality of the selected non-small cell lung cancer CT images in the data set is less than the preset standard in the data elimination criteria; If the image quality of the selected non-small cell lung cancer CT image in the data set is less than the preset standard in the data rejection standard, the selected non-small cell lung cancer CT image and the non-small cell lung cancer diagnosis report corresponding to the selected non-small cell lung cancer CT image are deleted, and the process returns to the step of selecting a non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected; If the image quality of the selected non-small cell lung cancer CT image in the data set is greater than or equal to the preset standard in the data rejection standard, the method returns to selecting a non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected.
4. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 3 is characterized in that: The data preprocessing of the collected CT images includes: Select a non-small cell lung cancer CT image; The trilinear interpolation algorithm was used to resample the CT images of non-small cell lung cancer; Adjust the pixels of the resampled non-small cell lung cancer CT images; Screening out non-small cell lung cancer CT images with pixel intensity values in the range of greater than or equal to -1000 Henry and less than or equal to 400 Henry, and using the non-small cell lung cancer CT images with pixel intensity values in the range of greater than or equal to -1000 Henry and less than or equal to 400 Henry as the first pre-processed non-small cell lung cancer CT images; Return to the step of selecting a non-small cell lung cancer CT image until all non-small cell lung cancer CT images are selected.
5. The pathological subtype classification method based on the visual language concept bottleneck model according to claim 4 is characterized in that: The data preprocessing of the collected CT images also includes: Select a first pre-treatment NSCLC CT image; Cropping the first pre-processed non-small cell lung cancer CT image; Obtaining a cropped first pre-processed non-small cell lung cancer CT image; Adjusting the cropped first pre-processed non-small cell lung cancer CT image so that the first pre-processed non-small cell lung cancer CT image matches the pixel standard, and using the obtained CT image as the second pre-processed non-small cell lung cancer CT image; Returning to the step of selecting a first pre-processed non-small cell lung cancer CT image until all first pre-processed non-small cell lung cancer CT images are selected; All second preprocessed non-small cell lung cancer CT images are included in the data preprocessing set of CT images.
6. The lung cancer subtyping method based on the visual language concept bottleneck model according to claim 5, characterized in that: The method of extracting visual features of the preprocessed CT image based on the three-dimensional visual adapter includes: calling a local fully connected neural network and a global encoder; Generate a 3D vision adapter; Select a non-small cell lung cancer CT image in the data preprocessing set; Using a three-dimensional vision adapter, determine the spatial position between CT slices; Returning a non-small cell lung cancer CT image in the selected data preprocessing set until all non-small cell lung cancer CT images are selected; A spatial position data set between CT slices is obtained with respect to a data preprocessing set.
7. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 6 is characterized in that: The method of extracting visual features of the preprocessed CT image based on the three-dimensional visual adapter further includes: Generate a 3D CT image spatial information aggregation network based on a local fully connected neural network of a 3D vision adapter and a global encoder of a 3D vision adapter; Call the spatial position data set between CT slices; Selecting one of the spatial position data in the spatial position data set between CT slices; Incorporate spatial position data into the spatial information collection network of three-dimensional CT images; Returning one of the spatial position data in the spatial position data set between the selected CT slices until all the spatial position data are selected; Obtain three-dimensional CT image spatial data for non-small cell lung cancer diagnosis report.
8. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 7, characterized in that: The method of extracting visual features of the preprocessed CT image based on the three-dimensional visual adapter further includes: Select a non-small cell lung cancer diagnosis report; Calling the three-dimensional CT image spatial data corresponding to the selected non-small cell lung cancer diagnosis report; Call the trained linear projection module in the 3D vision adapter; The visual features of 3D CT image spatial data are extracted based on the linear projection module.
9. The lung cancer subtype classification method based on the visual language concept bottleneck model according to claim 8, characterized in that: The concept extraction method using visual features includes: receiving a specification table; the specification table regarding a concept extraction task and a desired output structure; Based on the description table, obtaining a detailed list of extracted concepts and at least one concept expression table; Using the radiology report example, clear and consistent defaults for concept labels were determined.
10. The lung cancer subtype classification method based on visual language concept bottleneck model according to claim 9, characterized in that: The loss-based training model modifies the parameters of the visual language concept bottleneck model to obtain the pathological subtype classification results, including: Constructing a visual language concept bottleneck model; The method for constructing a visual language concept bottleneck model comprises: Utilize fine-grained semantic associations to build domain-adaptive visual language models; Training a visual-language concept bottleneck model for non-small cell lung cancer subtype classification based on a domain-adapted visual-language model.