Interpretable multi-view lung nodule image classification methods, systems, and devices

By employing a multi-view image classification method and a deep learning network with a self-attention mechanism, the problems of insufficient feature extraction and uninterpretable models in pulmonary nodule CT images were solved, achieving efficient and accurate classification of benign and malignant pulmonary nodules and interpretable analysis.

CN116563629BActive Publication Date: 2025-11-14WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310524660.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-11-14
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

Existing technologies for CT imaging diagnosis of pulmonary nodules suffer from insufficient feature extraction or excessive noise, leading to decreased classification performance. At the same time, artificial intelligence models lack interpretability and are difficult to provide diagnostic evidence.

Method used

A multi-view image classification method is adopted, which extracts multi-view features from CT images through a deep learning network with self-attention mechanism, and combines the AIM module to analyze nodule attributes and internal structure, providing interpretable benign and malignant classification results.

Benefits of technology

It improves the efficiency and accuracy of lung nodule feature extraction, enhances the credibility of classification results, provides interpretability analysis of nodule attributes and internal structure, and improves the reliability of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563629B_ABST
    Figure CN116563629B_ABST
Patent Text Reader

Abstract

This invention discloses an interpretable multi-view lung nodule image classification method, system, and device. First, a sub-cubic space containing the target nodule is generated based on lung CT images. The obtained three-view cross-sectional information is then processed into blocks and vectorized for computation. These vectorized vectors are then input into a deep learning network based on a self-attention mechanism for feature extraction, yielding visual features representing the nodule. These features are then input into an AIM module, where corresponding prediction heads obtain nodule attribute analysis results and nodule internal structure analysis results. The parameters of two linear mapping layers in the nodule attribute analysis module are fused into the input features and input into a benign / malignant classification module to obtain the classification result. This invention addresses the shortcomings of existing lung nodule classification feature extraction methods and the poor interpretability of model prediction results, overcoming the lack of effective solutions in current methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to medical image processing technology under artificial intelligence, and relates to a multi-view lung nodule image classification method, system, and device, particularly an interpretable multi-view lung nodule image benign and malignant classification method, system, and device. Background Technology

[0002] CT imaging is an effective method for examining lung nodules. However, in clinical practice, the sheer volume of CT images to review demands a high level of clinical experience, leading to subjective interpretations and a tendency for missed or incorrect diagnoses, thus hindering the consistent efficiency and accuracy of diagnosis. In this context, AI-based assisted diagnostic systems play a significant role. These systems automatically extract features from CT images and classify them as benign or malignant, effectively reducing human intervention in the diagnostic process. This decreases the overall workload for clinicians and reduces the likelihood of missed or incorrect diagnoses.

[0003] In current AI-based auxiliary diagnosis of benign and malignant pulmonary nodules, feature extraction is primarily performed on single CT scans (2D) or complete CT scans (3D). The number and representational power of the extracted features directly affect the classification performance of the trained AI model. Although 3D feature extraction yields a larger number of features, it also contains excessive noise and interference, reducing model classification performance. Therefore, using more reasonable feature extraction methods should maximize the number of features extracted while minimizing interference from redundant features, thereby improving the overall classification performance of the model and reducing overfitting caused by noise and interference during training. Furthermore, existing pulmonary nodule benign / malignant diagnostic techniques suffer from low reliability, mainly due to the black-box nature of AI models, which cannot provide interpretable analysis of the results. In clinical applications, disease diagnosis requires specific evidence; a diagnosis alone cannot guarantee complete trust from both physicians and patients. Therefore, auxiliary diagnostic systems for benign / malignant nodules should provide certain criteria to facilitate interpretability and understanding of the results. Overall, in the diagnosis of benign and malignant pulmonary nodules, it is worthwhile to conduct in-depth research on how to extract pulmonary nodule features from CT images more effectively and rationally, and provide interpretable analysis during classification. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method, system, and device for classifying benign and malignant lung nodules from multiple perspectives, offering interpretability.

[0005] The technical solution adopted by the method of the present invention is: a multi-view lung nodule image classification method with interpretability, comprising the following steps:

[0006] Step 1: Obtain the subspace image containing the nodule from the CT image and generate a subcube space containing the nodule;

[0007] Step 2: Extract the cross section, coronal plane and sagittal plane in the sub-cube space to obtain the cross section containing the three perspectives of the nodule;

[0008] Step 3: The obtained three viewpoint section information is processed into blocks, and a vector containing all viewpoint information is obtained through linear mapping;

[0009] Step 4: Input the vector information into the Transformer deep learning network based on the self-attention mechanism to extract multi-view features and obtain the multi-view nodule feature vector representing the vision of the nodule.

[0010] Step 5: Input the obtained multi-view nodule feature vector into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through the two linear mapping layers in the module, respectively;

[0011] Step 6: Concatenate the nodule attribute analysis result vector, the nodule internal structure analysis result vector, and the multi-view feature vector obtained in Step 4 to obtain a vector for benign / malignant classification. Input the vector into the benign / malignant classification module for classification to obtain the final classification result.

[0012] The technical solution adopted by the system of the present invention is: a multi-view lung nodule image classification system with interpretability, comprising the following modules:

[0013] The first module is used to obtain a subspace image containing a nodule from a CT image and generate a subcube space containing the nodule.

[0014] The second module is used to extract the cross section, coronal plane and sagittal plane in the sub-cube space to obtain the cross section containing the three perspectives of the nodule;

[0015] The third module is used to perform block-based processing on the obtained three viewpoint cross-sectional information and obtain a vector containing all viewpoint information through linear mapping.

[0016] The fourth module is used to input vector information into the Transformer deep learning network based on self-attention mechanism to extract multi-view features and obtain multi-view nodule feature vectors representing the vision of the nodule.

[0017] The fifth module is used to input the obtained multi-view nodule feature vectors into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through two linear mapping layers within the module, respectively.

[0018] The sixth module is used to concatenate the nodule attribute analysis result vector, the nodule internal structure analysis result vector, and the multi-view feature vector obtained in the fourth module to obtain a vector for benign and malignant classification. This vector is then input into the benign and malignant classification module for classification to obtain the final classification result.

[0019] The technical solution adopted by the device of the present invention is: a multi-view lung nodule image classification device with interpretability, comprising:

[0020] One or more processors;

[0021] A storage device is provided for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the interpretable multi-view lung nodule image classification method described above. The beneficial effects of this invention are:

[0022] (1) This invention proposes a method to extract lung nodule feature information from three perspectives used by clinicians during diagnosis, in order to solve the problems of insufficient features in two-dimensional extraction and excessive noise in three-dimensional extraction, and to make more reasonable use of CT images to extract features.

[0023] (2) This invention proposes a method for learning the direct association of features from three perspectives based on a self-attention mechanism. By calculating attention values, the spatial association of features from different perspectives can be learned efficiently.

[0024] (3) In view of the black box nature of artificial intelligence models, this invention proposes a module that can analyze and predict nodule attributes and internal structure. By associating this module with the benign and malignant judgment module, it makes causal relationship with classification, achieves interpretability of classification results, solves the black box problem, and makes the diagnosis results more credible. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the method principle framework of an embodiment of the present invention;

[0026] Figure 2 This is a deep learning network structure diagram of the self-attention mechanism in an embodiment of the present invention;

[0027] Figure 3 This is a structural diagram of the AIM module according to an embodiment of the present invention;

[0028] Figure 4This is a structural diagram of the benign / malignant classification module according to an embodiment of the present invention;

[0029] Figure 5 This is an interpretable classification result derived from the specific application of the present invention. Detailed Implementation

[0030] To facilitate understanding and implementation of the technical solutions of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only for illustration and explanation of the present invention and are not intended to limit the present invention.

[0031] This invention performs block processing on the subspace containing the nodule portion of CT images of lung nodules, dividing it into transverse, coronal, and sagittal perspectives. It then utilizes a deep neural network based on a pure self-attention mechanism to extract features and learn associations from the information from different perspectives. After feature extraction, a benign / malignant classification module and an AIM module are used to analyze the category and attribute structure of the nodules. By using a weight-sharing method, a causal relationship is established between the two modules, providing interpretability for the final benign / malignant classification result.

[0032] Please see Figure 1 The present invention provides an interpretable multi-view lung nodule image classification method, comprising the following steps:

[0033] Step 1: Obtain the subspace image containing the nodule from the CT image. Based on the subspace containing the lung nodule in the complete lung CT image, generate a subcube space containing the nodule.

[0034] In this embodiment, a subspace image containing the nodule is obtained from the CT image, and with the nodule location as the center point, it is expanded left and right by N voxel units along the x, y, z axes of the subspace image to obtain the subspace coordinates of the cube containing the nodule [x min ,x max ,y min ,y max ,z min ,z max ]; where N is a preset value, which is 64 in this embodiment, x min ,x max ,y min ,y max ,z min ,z max These represent the minimum and maximum values ​​of the coordinates along the x, y, and z axes of the cube subspace, respectively.

[0035] Step 2: Extract the cross section, coronal plane and sagittal plane in the sub-cube space to obtain the cross section containing the three perspectives of the nodule;

[0036] In this embodiment, the cross-section, coronal plane, and sagittal plane are extracted from the obtained subcubic space to obtain a cross-section containing three views of the nodule: cross-section information Axial_View, coronal plane information Coronal_View, and sagittal plane information Sagittal_View.

[0037] Step 3: The obtained three viewpoint section information is processed into blocks, and a vector containing all viewpoint information is obtained through linear mapping;

[0038] In this embodiment, after step 2, three perspectives of the nodule are obtained. The image containing these three perspectives is denoted as x, then x∈R 224x224×3 Divide the image from each viewpoint into 16×16 blocks p, denoted as p. A fully connected layer performs a linear mapping on the image, mapping each slice p to a 192-dimensional vector E, denoted as . and

[0039] Step 4: Input the vector information into a deep learning network based on a self-attention mechanism to extract multi-view features and obtain a multi-view nodule feature vector representing the vision of the nodule.

[0040] Please see Figure 2 The deep learning network based on the self-attention mechanism in this embodiment consists of 12 identical self-attention modules connected in series.

[0041] The self-attention module in this embodiment is used for feature extraction and association learning of nodule CT images from multiple perspectives. It consists of four layers connected in series. The first layer is a batch normalization layer, which normalizes the input information from the three perspectives. The second layer is a multi-head attention layer, which includes eight self-attention heads that extract features from the input information to obtain feature maps. The third layer is also a batch normalization layer, which performs batch normalization on the feature maps obtained from the multi-head attention layer. The fourth layer is a multilayer perceptron (MLP). In this embodiment, the MLP includes a GELU activation layer and two fully connected layers. The feature information obtained from the previous self-attention module is input into the next self-attention module and undergoes the same processing. After calculation by 12 self-attention modules, the multi-view nodule features are obtained.

[0042] In this embodiment, after step 3, a vector containing nodule information is obtained, and batch normalization is performed on it. Specifically, the overall mean μ and standard deviation σ of the information are 0 and 1, respectively. The formula for batch normalization is as follows:

[0043]

[0044] Where x is the input information, ε is the hyperparameter, E[x] is the variance of the input information, and Var[x] is the expected value of the input information;

[0045] After obtaining the normalized information, self-attention is used to calculate the relevance of the information self-check. The formula for self-attention is as follows:

[0046]

[0047] Where Q is the computational parameter Query in the self-attention mechanism, used to calculate the correlation with other features; K is the computational parameter Key in the self-attention mechanism, used for comparison with other features; V is the computational parameter Value in the self-attention mechanism, used to weight the importance of the relationship; and d... k This represents the total number of dimensions of the input information, used to avoid additional computational overhead caused by an excessive number of dimensions.

[0048] To enhance the overall learning capacity of the model, this embodiment employs a multi-head self-attention mechanism. Each attention head... i Each will learn an independent matrix W i The calculation method for multi-head attention can be expressed by the formula:

[0049] MSA(Q,K,V)=Concat(head1,...,head h W O ;

[0050] head i =Attention(QW i Q ,KW i K VW i V ), W i Q W is an independent matrix that is multiplied by Q. i K W is an independent matrix that is multiplied by K. i V V represents an independent matrix that is multiplied by V; Concat is used to concatenate the results obtained from each self-attention head; W represents... O This is an independent matrix that is compatible with the multi-head attention after being spliced ​​together.

[0051] After passing through a multi-head self-attention layer, the resulting features are fed into another batch normalization (BN) layer, and then into a multilayer perceptron (MLP). The MLP contains a GELU (Gaussian Error Linear Units) activation layer and two fully connected layers. GELU activates the input data by mapping the cumulative Gaussian distribution function Φ(), as detailed in the following formula:

[0052] GELU(x) = x * Φ(x);

[0053] Therefore, the expression formula for MLP is as follows:

[0054] MLP(x) = Φ(xW1+b1)W2+b2;

[0055] Where W1 represents the parameters in fully connected layer 1, b1 represents the bias in fully connected layer 1, W2 represents the parameters in fully connected layer 2, and b2 represents the bias in fully connected layer 2.

[0056] The complete feature extraction process can be expressed by the following formula:

[0057] x′ i =MSA(BN(x))+x;

[0058] x i =MLP(BN(x′) i ))+x′ i ;

[0059] y = BN(x) i );

[0060] Finally, a 192-dimensional feature vector x is obtained. features This feature space contains feature information from three perspectives and their correlations.

[0061] Step 5: Input the obtained multi-view nodule feature vector into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through the two linear mapping layers in the module, respectively;

[0062] Please see Figure 3 The multi-view nodule feature vector is input into the AIM module. Two linear mapping layers are used to predict nodule attributes and internal structure respectively. The features used to predict nodule attributes and nodule structure are obtained through the two linear mapping layers respectively.

[0063] The formula for the linear mapping layer is:

[0064] y = xA T +b;

[0065] Where y is the output of the linear mapping layer, x is the input value, and A T b is the learnable parameter of this layer, and b is the bias parameter of this layer;

[0066] After obtaining the feature x used for attribute prediction attritbutes After x is used to predict the structure structures This will be input into the attribute prediction head and the internal structure prediction head, respectively;

[0067] The formula for the attribute prediction head is as follows:

[0068] y attributes =x attributes A T +b;

[0069] Among them, y attributes For the output of the attribute prediction head, x attributes For input values;

[0070] The formula for the structural prediction head is as follows:

[0071] y structures =x structures A T +b;

[0072] Among them, y structures For the output of the attribute prediction head, x structures For input values;

[0073] Simultaneous feature x attritbutes With x structures It will also be output and used for feature concatenation in the next module.

[0074] In this embodiment, x features It will be input into the AIM module and obtained as a 64-dimensional x through dimensionality reduction. attritbutes And a 34-dimensional x structures These two features will be fed into linear layer classifiers for predicting attributes and structure, respectively, to obtain the prediction results. and

[0075] Step 6: Concatenate the nodule attribute analysis result vector, the nodule internal structure analysis result vector, and the multi-view feature vector obtained in Step 4 to obtain a vector for benign / malignant classification. Input the vector into the benign / malignant classification module for classification to obtain the final classification result.

[0076] Please see Figure 4 The benign / malignant classification module in this embodiment includes a splicing layer and a benign / malignant classification layer;

[0077] The stitching layer in this embodiment is used to combine multi-view nodule features x attritbutes The two features x output by the AIM module attritbutes With x structures Concatenate the data to obtain the features x used for classification. classification ;

[0078] x classification =Concat(x) features +x attributyes +x structures )

[0079] Next feature x classification Input into the benign / malignant classification layer;

[0080] The benign / malignant classification layer in this embodiment includes a linear mapping layer; wherein the formula for the linear mapping layer is:

[0081] y = x classification A T +b;

[0082] Where y represents the result of benign / malignant classification prediction, and x represents the result of malignant classification prediction. classification For input value, A T b is the learnable parameter of this layer, and b is the bias parameter of this layer.

[0083] In this embodiment, a deep learning network based on a self-attention mechanism, an AIM module, and a benign / malignant classification module constitute a multi-view lung nodule image classification network; the multi-view lung nodule image classification network in this embodiment is a pre-trained multi-view lung nodule image classification network.

[0084] The training includes the following sub-steps:

[0085] (1) Acquire several CT images, obtain subspace images containing nodules from the CT images, and generate a subcube space containing the nodules based on the subspace containing the lung nodules in the complete lung CT images.

[0086] (2) In the subcube space, the cross section, coronal plane and sagittal plane are extracted to obtain the cross section containing the three perspectives of the nodule;

[0087] (3) The obtained three viewpoint section information is processed into blocks, and a vector containing all viewpoint information is obtained by linear mapping.

[0088] (4) Input the vector information into a deep learning network based on self-attention mechanism to extract multi-view features and obtain multi-view feature vectors representing the vision of the nodule.

[0089] (5) Input the obtained multi-view feature vector into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through the two linear mapping layers in the module respectively;

[0090] (6) The vector of nodule attribute analysis result, the vector of nodule internal structure analysis result, and the multi-view feature vector obtained in step 4 are concatenated to obtain a vector for benign and malignant classification, and input into the benign and malignant classification module for classification to obtain the final classification result.

[0091] (7) Compare the analysis results obtained by the AIM module with the true values ​​labeled in the dataset, and calculate the loss value by using the attribute analysis loss function and the internal structure loss function;

[0092] The attribute analysis loss function in this embodiment is:

[0093] The internal structure loss function in this embodiment is:

[0094]

[0095] Where n is the number of tag types, y i This is the truth value for this type of label;

[0096] This embodiment uses the prediction results and Square root truth value y structures With y attributes By comparison, the loss value can be obtained. attributes With Loss structures .

[0097] (8) Compare the classification prediction results obtained by the benign and malignant classification module with the benign and malignant labels in the dataset, and obtain the classification loss value through the benign and malignant loss function;

[0098] The benign / malignant loss function in this embodiment is:

[0099] Where y is the truth value of the benign / malignant label, The result of predicting benign or malignant malignancy;

[0100] (9) Weight the three loss values ​​to obtain the total loss value, and then perform backpropagation;

[0101] The total loss value is Loss total =αLoss ce +βLoss mse +λLoss bce α, β, and λ are weights, and in this embodiment, their values ​​are 0.7, 0.2, and 0.1, respectively.

[0102] This embodiment uses the loss value for good and bad performance. classification , and Loss attributes and Loss structures The total loss value is obtained by weighting the values, and the calculation method is as follows:

[0103] Loss total =0.7 Loss classification +0.2 Loss attributes +0.1 Loss structures ;

[0104] (10) Repeat steps (3) to (9) until the multi-view lung nodule image classification network converges.

[0105] This embodiment is implemented using the Python platform based on the PyTorch library. The basic data reading and writing, basic mathematical operations, and optimization solutions are well-known technologies in this field. At the same time, the implementation is carried out on four GTX 2080 GPU graphics cards, and the command and other operations are well-known technologies in Linux operation, which will not be described in detail here.

[0106] The following comparative experiments demonstrate the beneficial effects of the present invention.

[0107] The dataset contains a total of 1301 solid lung nodules, including 612 benign nodules and 689 malignant nodules, with diameters ranging from 3 mm to 54 mm. The data used for training and validation of this invention are shown in Table 1.

[0108] Table 1. Data used for training and validation.

[0109]

[0110]

[0111] This experiment sets the network's objective as extracting features from CT images and determining their benign or malignant nature. The threshold for the determination is set at 0.5. The network is evaluated using metrics commonly used in lung nodule classification, including accuracy, precision, recall, specificity, F1 score, and area under the ROC curve (AUC). Higher values ​​for all evaluation metrics indicate better model performance. The comparison methods primarily include networks used for classification and nodule benign / malignant determination tasks: ResNet50, DenseNet121, Vit, MCCNN, HESAM, CrossViT, Deit, and the proposed method MMViT. The network's test results on the dataset are shown in Table 2.

[0112] Table 2 shows the test results of the network on the dataset.

[0113]

[0114] As shown in Table 2, the present invention consistently outperforms all compared methods in the main metrics of Accuracy, AUC, and Precision. It also demonstrates a significant performance advantage over traditional methods such as ResNet50 and ViT-Tiny. Although the CrossVit and Deit models, which are also based on a self-attention mechanism, achieve similar results in Accuracy and AUC, they lag significantly behind in other metrics. This fully validates the effectiveness of the proposed method in classification.

[0115] Figure 5 This demonstrates the interpretability of the prediction results in this invention. Figure 5 The example shown is a correct judgment of benignity or malignancy. If the judgment only indicates a possibility of benignity or malignancy, it's impossible to know the specific characteristics the network used to reach its conclusion. This can be achieved through observation. Figure 5 (a) The network's predictions of nodule attributes and internal structure show that the network successfully determined that the nodule had no visible lobes, no visible spicules, clearly visible spines, and that the overall nodule was nearly spherical; at the same time, the nodule possessed a ground-glass internal structure, explaining why the network judged it to be malignant. Figure 5 In (b), the network successfully determined that the nodule possessed the characteristics of no visible lobulation, no obvious spicules, no obvious spinous processes, and an oval shape; simultaneously, the nodule exhibited the internal structure of soft tissue, explaining why the network judged it to be malignant. Figure 5In (c), the network successfully determined that the nodule possessed shallow lobulation, subtle spiculations, subtle spinous processes, and an overall nearly elliptical shape; simultaneously, the nodule exhibited the internal structure of soft tissue, explaining why the network classified it as benign. Figure 5 In (d), the network successfully predicted that the nodule had no visible lobes, no obvious spicules, no obvious spinous processes, and was elliptical in shape; at the same time, the nodule had the internal structure of soft tissue, which explains why the network judged it to be malignant. Figure 5 The examples illustrate that the network of this invention can correctly identify nodule attribute features and internal structural features, and use these as one of the bases for reasoning. Therefore, the prediction of nodule attributes and internal structure greatly increases the reliability of the results.

[0116] This invention addresses the shortcomings in feature extraction during the classification and diagnosis of benign and malignant pulmonary nodules. It extracts nodule features from multiple perspectives, simulating the diagnostic methods used by clinicians, thereby improving the efficiency and quality of feature extraction and significantly increasing classification accuracy. Furthermore, this invention solves the black-box problem of artificial intelligence models in diagnosis by providing interpretability for overall reasoning through outputting predictions of nodule attributes and internal structures that are causally related to the benign / malignant classification results, making the reasoning results more credible.

[0117] It should be understood that all parts not described in detail in this specification belong to the prior art. The above description of the preferred embodiments is relatively detailed, but it should not be regarded as a limitation on the scope of protection of this invention. Those skilled in the art can make substitutions or modifications under the guidance of this invention without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A multi-view lung nodule image classification method with interpretability, characterized in that, Includes the following steps: Step 1: Obtain the subspace image containing the nodule from the CT image and generate a subcube space containing the nodule; Step 2: Extract the cross section, coronal plane and sagittal plane in the sub-cube space to obtain the cross section containing the three perspectives of the nodule; Step 3: The obtained three viewpoint section information is processed into blocks, and a vector containing all viewpoint information is obtained through linear mapping; Step 4: Input the vector information into a deep learning network based on a self-attention mechanism to extract multi-view features and obtain a multi-view nodule feature vector representing the vision of the nodule. Step 5: Input the obtained multi-view nodule feature vector into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through the two linear mapping layers in the module, respectively; Step 6: Concatenate the nodule attribute analysis result vector, the nodule internal structure analysis result vector, and the multi-view feature vector obtained in Step 4 to obtain a vector for benign / malignant classification. Input the vector into the benign / malignant classification module for classification to obtain the final classification result.

2. The interpretable multi-view lung nodule image classification method according to claim 1, characterized in that: In step 1, a subspace image containing the nodule is obtained from the CT image. Using the nodule location as the center point, the image is expanded left and right along the x, y, and z axes by N voxels to obtain the subspace coordinates of the cube containing the nodule. Where N is a preset value, x min ,x max ,y min ,y max ,z min ,z max These represent the minimum and maximum values ​​of the coordinates along the x, y, and z axes of the cube subspace, respectively.

3. The interpretable multi-view lung nodule image classification method according to claim 1, characterized in that: In step 3, considering the three perspectives of the nodule obtained after step 2, and denoting the image containing the three perspectives as x, then x∈R 224x224 ×3 Divide the image from each viewpoint into 16×16 blocks p, denoted as p. A fully connected layer performs a linear mapping on the image, mapping each slice p to a 192-dimensional vector E, denoted as . and 4. The interpretable multi-view lung nodule image classification method according to claim 1, characterized in that: In step 4, the deep learning network based on the self-attention mechanism consists of 12 identical self-attention modules connected in series. The self-attention module is used to extract features and learn associations from nodule CT images from multiple perspectives. It consists of four layers connected in series. The first layer is a batch normalization layer, which is used to normalize the input information from the three perspectives. The second layer is a multi-head attention layer, which includes eight self-attention heads that extract features from the input information to obtain feature maps. The third layer is a batch normalization layer, which performs batch normalization on the feature maps obtained by the multi-head attention layer. The fourth layer is a multilayer perceptron (MLP), which includes a GELU activation layer and two fully connected layers. The feature information obtained from the previous self-attention module will be input into the next self-attention module and processed in the same way. After calculation by 12 self-attention modules, multi-view nodule features are obtained.

5. The interpretable multi-view lung nodule image classification method according to claim 1, characterized in that: In step 5, the multi-view nodule feature vector is input into the AIM module. Two linear mapping layers are used to predict the nodule attributes and internal structure, respectively. The features used to predict the nodule attributes and the features used to predict the nodule structure are obtained through the two linear mapping layers. in The formula for the linear mapping layer is: y=xA T +b; Where y is the output of the linear mapping layer, x is the input value, and A T b is the learnable parameter of this layer, and b is the bias parameter of this layer; After obtaining the feature x used for attribute prediction attritbutes After x is used to predict the structure structures This will be input into the attribute prediction head and the internal structure prediction head, respectively; The formula for the attribute prediction head is as follows: y attributes =x attributes A T +b; Among them, y attributes For the output of the attribute prediction head, x attributes Input value; The formula for the structural prediction head is as follows: y structures =x structures A T +b; Among them, y structures For the output of the attribute prediction head, x structures For input values; Simultaneous feature x attritbutes With x structures It will also be output and used for feature concatenation in the next module.

6. The interpretable multi-view lung nodule image classification method according to claim 1, characterized in that: In step 6, the benign / malignant classification module includes a splicing layer and a benign / malignant classification layer; The splicing layer is used to combine multi-view nodule features x attritbutes The two features x output by the AIM module attritbutes With x structures The features x are concatenated to obtain the features used for classification. classification ; x classification =Concat(x features +x attributyes +x structures ) Next feature x classification Input into the benign / malignant classification layer; The benign / malignant classification layer includes a linear mapping layer; wherein the formula for the linear mapping layer is: y=x classification A T +b; Where y represents the result of benign / malignant classification prediction, and x represents the result of malignant classification prediction. classification For input value, A T b is the learnable parameter of this layer, and b is the bias parameter of this layer.

7. The interpretable multi-view lung nodule image classification method according to any one of claims 1-6, characterized in that: The deep learning network based on the self-attention mechanism, the AIM module, and the benign / malignant classification module constitute a multi-view lung nodule image classification network; the multi-view lung nodule image classification network is a pre-trained multi-view lung nodule image classification network. The training includes the following sub-steps: (1) Acquire several CT images, extract subspace images containing nodules from the CT images, and generate a subcube space containing the nodules. (2) In the subcube space, the cross section, coronal plane and sagittal plane are extracted to obtain the cross section containing the three perspectives of the nodule; (3) The obtained three viewpoint section information is processed into blocks, and a vector containing all viewpoint information is obtained by linear mapping. (4) Input the vector information into a deep learning network based on self-attention mechanism to extract multi-view features and obtain multi-view feature vectors representing the vision of the nodule. (5) Input the obtained multi-view feature vector into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through the two linear mapping layers in the module respectively; (6) The vector of nodule attribute analysis result, the vector of nodule internal structure analysis result, and the multi-view feature vector obtained in step 4 are concatenated to obtain a vector for benign and malignant classification, and input into the benign and malignant classification module for classification to obtain the final classification result. (7) Compare the analysis results obtained by the AIM module with the true values ​​labeled in the dataset, and calculate the loss value by using the attribute analysis loss function and the internal structure loss function; The attribute analysis loss function is: The internal structure loss function is: Where n is the number of tag types, y i This is the truth value for this type of label; (8) Compare the classification prediction results obtained by the benign and malignant classification module with the benign and malignant labels in the dataset, and obtain the classification loss value through the benign and malignant loss function; The benign / malignant loss function is: Where y is the truth value of the benign / malignant label, The result of predicting benign or malignant malignancy; (9) Weight the three loss values ​​to obtain the total loss value, and then perform backpropagation; The total loss value is Loss total =αLoss ce +βLoss mse +λLoss bce α, β, and λ are weights; (10) Repeat steps (3) to (9) until the multi-view lung nodule image classification network converges.

8. A multi-view lung nodule image classification system with interpretability, characterized in that, Includes the following modules: The first module is used to obtain a subspace image containing a nodule from a CT image and generate a subcube space containing the nodule. The second module is used to extract the cross section, coronal plane and sagittal plane in the sub-cube space to obtain the cross section containing the three perspectives of the nodule; The third module is used to perform block-based processing on the obtained three viewpoint cross-sectional information and obtain a vector containing all viewpoint information through linear mapping. The fourth module is used to input vector information into the Transformer deep learning network based on self-attention mechanism to extract multi-view features and obtain multi-view nodule feature vectors representing the vision of the nodule. The fifth module is used to input the obtained multi-view nodule feature vectors into the AIM module, and obtain the nodule attribute analysis result vector and the nodule internal structure analysis result vector through two linear mapping layers within the module, respectively. The sixth module is used to concatenate the nodule attribute analysis result vector, the nodule internal structure analysis result vector, and the multi-view feature vector obtained in the fourth module to obtain a vector for benign and malignant classification. This vector is then input into the benign and malignant classification module for classification to obtain the final classification result.

9. A multi-view lung nodule image classification device with interpretability, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement an interpretable multi-view lung nodule image classification method as described in any one of claims 1 to 7.