Pulmonary nodule malignant tumor classification method based on double-graph space-time attention network
Through the method based on the dual-graph spatiotemporal attention network, the multi-scale characteristics of multi-phase 3D lung nodule data are integrated, and the problem of insufficient accuracy in early diagnosis of lung cancer in the prior art is solved, and efficient early diagnosis of lung cancer and multimodal medical analysis are achieved.
Patent Information
- Application Number
- CN202510510235.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to effectively integrate multi-scale features of multi-phase 3D lung nodule data, and the relationship between objects is complex and interdependent, resulting in insufficient accuracy in early diagnosis of lung cancer.
Using a method based on a dual-graph spatiotemporal attention network, a multi-scale feature extraction encoder, a dual-modal graph construction module, a graph feature fusion module and a full connection layer are used to integrate the complex relationships within and between modes to build a network model for the classification of pulmonary nodules malignant tumors.
It significantly improved the accuracy of early diagnosis of lung cancer, increased the classification accuracy by at least 3.57%, and reduced the number of model parameters by more than 28%, reducing the radiation risk of multiple CT scans, and providing new ideas for cross-modal modeling.
Smart Images

Figure CN120448894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal graph fusion and spatiotemporal feature technology, and specifically to a lung nodule malignant tumor classification method based on a dual-graph spatiotemporal attention network. Background Art
[0002] Lung cancer is the leading cause of cancer-related deaths worldwide, with approximately 2.5 million new cases each year, accounting for 12.4% of all newly diagnosed cancer cases, and approximately 1.8 million deaths, representing 18.7% of all cancer deaths. Unfortunately, lung cancer has a poor prognosis, with a low five-year survival rate. Survival rates for patients diagnosed in the early stages are approximately 50% to 60%, while those diagnosed in the later stages have a survival rate of less than 10%. Although usually benign, lung nodules have the potential to develop into lung cancer. If not monitored promptly, these nodules can potentially develop into cancer. Therefore, early detection and diagnosis of these nodules is crucial to improving survival rates for lung cancer patients.
[0003] Current research, such as CSF-Net and SAGVI, has attempted to use attention fusion mechanisms or graph construction to study inter-modal feature fusion. However, these methods fail to capture the dynamic changes of lesions. With increasing interest in lesion dynamics, research has shifted to longitudinal multi-time point analysis. Effectively integrating multimodal and multi-temporal data and deeply exploring the complex relationships between features within and between modalities has become a major challenge in predicting malignant lung nodules.
[0004] In general, the current technical problems are as follows:
[0005] 1) It is difficult to extract multi-scale features from multi-temporal 3D lung nodule data.
[0006] 2) The relationships between objects are complex and highly interdependent. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a method for classifying lung nodules into malignant tumors based on a dual-image spatiotemporal attention network. This method uses the input lung nodule ROI image as the data source and integrates the complex relationships within and between modalities through advanced algorithms to automatically predict the probability of developing into a malignant tumor, thereby significantly improving the diagnostic accuracy.
[0008] In order to achieve the above object, the technical solutions specifically adopted by the present invention are as follows:
[0009] A method for classifying malignant lung nodules based on a dual-image spatiotemporal attention network includes the following steps:
[0010] Step 1: Collect and preprocess image data and corresponding text data. The image data is a patient's CT scan sequence image, and the text data includes lung nodule annotation data and clinical data. After preprocessing, the image data is obtained to obtain a dual-time series lung nodule ROI image.
[0011] Step 2: Construct a dual-image spatiotemporal attention network for classifying the malignancy of pulmonary nodules, wherein the dual-image spatiotemporal attention network includes a multi-scale feature extraction encoder, a bimodal graph construction module, a graph feature fusion module, and a fully connected layer;
[0012] Step 3: Use the preprocessed dual-time series lung nodule ROI images and the text data corresponding to each time series as input to repeatedly train the dual-image spatiotemporal attention network, optimize the network parameters, and continuously perform iterative optimization;
[0013] Step 4: The pre-processed lung nodule ROI image is input into the trained dual-image spatiotemporal attention network to output accurate malignant tumor classification results.
[0014] Preferably, the preprocessing method is: extracting a 3D region of interest with a size of 16x 64x64 according to the nodule position of the image data, and dividing it into malignant and benign subtypes, and saving the corresponding nodule diameter as a text feature.
[0015] Preferably, the multi-scale feature extraction encoder includes a local feature extraction branch, a global feature extraction branch and a feature fusion branch.
[0016] Preferably, in the local feature extraction branch, the dual-time series lung nodule ROI image obtained by preprocessing is input, and first a 3D convolution operation is performed through Conv3D, and then it passes through 4 layers of local feature extraction networks in sequence. The output of the previous layer of local feature extraction network is used as the input of the next layer of local feature extraction network, wherein the first layer of local feature extraction network is composed of a regularization layer and a local feature extraction module, and the other three layers of local feature extraction networks are composed of a regularization layer, Conv3D and a local feature extraction module. The last layer of local feature extraction network is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to finally obtain the final local feature.
[0017] Preferably, in the global feature extraction branch, first, the dual-time series pulmonary nodule ROI image obtained by input preprocessing is divided into blocks of 4×4×2, and then passes through 4 layers of global feature extraction networks in sequence. The output of the previous global feature extraction network is used as the input of the next global feature extraction network. The first global feature extraction network is transformed into a Token sequence suitable for Transformer processing through a linear encoder, and then the global feature is extracted through the global feature extraction module. The other three global feature extraction networks are merged through Patches and then the global feature is extracted through the global feature extraction module. The last global feature extraction network is connected to the adaptive graph channel attention module, and through the adaptive graph channel attention module, adaptive average pooling, flattening, and linear layer operations are performed in sequence, and finally the final global feature is obtained.
[0018] Preferably, the feature fusion branch includes 4 adaptively connected feature fusion blocks in sequence. The 4 adaptively connected feature fusion blocks respectively correspond to the 4 layers of networks in the global feature extraction branch and the local feature extraction branch. The adaptively connected feature fusion block fuses the local feature output by the corresponding local feature extraction network, the global feature output by the global feature extraction network, and the output of the previous adaptively connected feature fusion block, and serves as the input of the next adaptively connected feature fusion block; the last adaptively connected feature fusion block is connected to the adaptive graph channel attention module, and through the adaptive graph channel attention module, adaptive average pooling, flattening, and linear layer operations are performed in sequence, and finally the final fusion feature is obtained.
[0019] Preferably, in the dual-modal graph construction module, the local and global features of the dual-time series t0 and t1 from the multi-scale feature extraction encoder are obtained in the form of an intra-modal graph through a graph construction method, and the text features from the fully connected layer and the respective fusion features of t0 and t1 are obtained in the form of an inter-modal graph through a graph construction method; the two graphs will be input into the subsequent graph feature fusion module.
[0020] Preferably, the graph construction method is as follows: Let the feature representation of each node be v0, v1,..., v n-1 , where the dimension of each feature vector is d, and the graph G is represented as G=(V, E, X), where v is a set of n nodes, E is the set of edges, defined as E={(i,j)|0≤i,j<n,i≠j}; the node feature matrix X is given by X=[v0 v1... v n-1 , where v n-1 is the feature vector of the i-th node; finally, GAT is used to capture the interactions within the two modal graphs to obtain the respective fusion features between and within the modalities.
[0021] Preferably, the graph feature fusion module includes two self-attention blocks and one cross-attention block.
[0022] Preferably, in the graph feature fusion module, the initial layer adopts a dual-path parallel self-attention block to independently process the inter-modal and intra-modal bimodal input features obtained through graph construction, and the initial layer adopts an adaptive weight distribution mechanism; the middle layer adopts a cross-attention block, and uses the key-value query mechanism of the cross-attention to create a modal association matrix for the intra-modal and inter-modal features, respectively; the last layer inputs the inter-modal and intra-modal fusion features, and uses the self-attention block to perform global recalibration.
[0023] Preferably, the output of the graph feature fusion module is output as a final classification result through a fully connected layer.
[0024] The present invention has the following characteristics and beneficial effects:
[0025] The beneficial effects of this method include: 1) Experiments on the NLST-cmst and CLST datasets show that its classification accuracy is at least 3.57% higher than the existing optimal method, and the number of model parameters is reduced by more than 28%, making it both efficient and lightweight; 2) Multimodal fusion reduces dependence on single-modality data and reduces the patient radiation risk caused by multiple CT scans; 3) The constructed NLST-cmst multimodal dataset and dual-image fusion framework provide new ideas for cross-modal modeling in medical image processing and have the potential for broad promotion; 4) Breakthroughs have been made in the application of graph convolutional networks and spatiotemporal attention mechanisms, laying a technical foundation for early diagnosis of lung cancer and multimodal medical analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the process of the present invention;
[0027] Figure 2 The overall architecture of the dual-graph spatiotemporal attention network model for pulmonary nodule malignancy prediction in this invention DGSAN;
[0028] Figure 3 Schematic diagram of correlation in lung nodule diagnosis, (a) inter-modality correlation, (b) intra-modality correlation, (c) relationship diagram containing inter-modality and intra-modality information.
[0029] Figure 4 Five schemes for constructing bimodal graphs.
[0030] Figure 5 ROC curves of DGSAN and other methods on the NIST-cmst dataset. DETAILED DESCRIPTION
[0031] The present invention is described in detail below in conjunction with specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0032] A method for classifying malignant pulmonary nodules based on dual-graph spatiotemporal attention network, such as Figure 1 As shown, the following steps are included:
[0033] Step 1: Collect and preprocess image data and corresponding text data. The image data is a CT serial scan image of the patient, and the text data includes lung nodule annotation data and clinical data. After preprocessing, the image data obtains a dual-time series lung nodule ROI image.
[0034] In this embodiment, the collected image datasets include a public dataset and a private dataset. The private dataset, the NLST-smst dataset, originates from the National Lung Screening Trial (NLST) initiated by the National Cancer Institute of the United States and contains physician-annotated region-of-interest (ROI) annotations of lung nodules from 433 subjects. These annotations were performed by physicians. This data was rigorously screened, with each participant having at least two longitudinal CT scans and clinical data (age, gender, smoking status, and screening results). The public dataset, CLST, includes 317 CT sequences and 2,295 annotated nodules from 109 patients. Nodules are categorized as invasive adenocarcinoma, minimally invasive adenocarcinoma, adenocarcinoma in situ, other malignant subtypes, inflammatory, and other benign subtypes. For the CLST, we categorize the subtypes into malignant (invasive adenocarcinoma, minimally invasive adenocarcinoma, adenocarcinoma in situ, and other malignant subtypes) and benign (inflammatory and other benign subtypes), and store the corresponding nodule diameter as a clinical feature. Regarding data partitioning, in this embodiment, the NLST-cmst dataset is divided into 347 training datasets and 86 test datasets, and all 36 cases of the screened CLST dataset are used for testing.
[0035] Furthermore, data preprocessing was performed to ensure data consistency. A 3D region of interest (ROI) of 16 x 64 x 64 was extracted based on the nodule locations in the two datasets. The nodules were then classified into malignant and benign subtypes, and the corresponding nodule diameters were saved as text features.
[0036] Step 2: Construct a dual-image spatiotemporal attention network for classifying the malignancy of pulmonary nodules, wherein the dual-image spatiotemporal attention network includes a multi-scale feature extraction encoder, a bimodal graph construction module, a graph feature fusion module, and a fully connected layer;
[0037] Specifically, such as Figure 2As shown, the multi-scale feature extraction encoder includes a local feature extraction branch, a global feature extraction branch and a feature fusion branch.
[0038] Furthermore, in the local feature extraction branch, the dual-time series lung nodule ROI image obtained by preprocessing is input, and first a 3D convolution operation is performed through Conv3D, and then it passes through 4 layers of local feature extraction networks in sequence. The output of the previous layer of local feature extraction network is used as the input of the next layer of local feature extraction network, wherein the first layer of local feature extraction network consists of a regularization layer and a local feature extraction module, and the other three layers of local feature extraction networks consist of a regularization layer, Conv3D and a local feature extraction module. The last layer of local feature extraction network is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to finally obtain the final local feature.
[0039] In the global feature extraction branch, first, the dual-time series lung nodule ROI image obtained by input preprocessing is divided into 4×4×2 blocks, and then passes through 4 layers of global feature extraction networks in sequence. The output of the previous layer of global feature extraction network is used as the input of the next layer of global feature extraction network. The first layer of global feature extraction network is converted into a Token sequence suitable for Transformer processing through a linear encoder, and then the global feature is extracted through the global feature extraction module. The other three layers of global feature extraction networks are merged through Patch and then the global features are extracted through the global feature extraction module. The last layer of global feature extraction network is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to obtain the final global feature.
[0040] The feature fusion branch includes four layers of adaptive feature fusion blocks connected in sequence. The four adaptive feature fusion blocks correspond to the four-layer networks of the global feature extraction branch and the local feature extraction branch respectively. The adaptive feature fusion block fuses the local features output by the corresponding local feature extraction network, the global features output by the global feature extraction network, and the output of the adaptive feature fusion block of the previous layer, and serves as the input of the adaptive feature fusion block of the next layer; the last layer of adaptive feature fusion blocks is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to finally obtain the final fusion feature.
[0041] It can be understood that the characterization of lung nodules at different scales by the multi-scale feature extraction encoder is crucial for accurately predicting their malignancy. In order to achieve effective multi-scale feature extraction, MSFEE adopts a three-branch structure. Each branch is used to capture local features, global features, and fusion features, and both contain four stages. In terms of local feature extraction, the original image is directly input into the local feature map to mix and reorganize the channels. This process captures the correlation between different channels, thereby enhancing the feature representation. In the global feature extraction process, the original image is divided into 4×4×2 blocks, which are fed into the global feature extraction branch. The global feature block adopts the
[0042] AGCA(m)=m·sigmoid(F′ r (ReLU(AGCM(F r (GAP(m)),A)))),
[0043] AGCM(I,A)=W·I·(A0×A1+A2),
[0044] Windowed Multi-Head Self-Attention (W-MSA) and Shifted Windowed Multi-Head Self-Attention (SW-MSA) modules. These modules enhance local information interaction while maintaining computational efficiency. To obtain fully fused features, the Adaptive Feature Fusion (AFF) block merges local features with global features at each stage. After the second stage, the fused features of the previous stage are connected to the AFF block of the current stage. The Adaptive Graph Channel Attention (AGCA) module is placed after both the local and global feature extraction stages. It reduces computational redundancy through adaptive graph convolution and enhances feature representation at a lower computational cost. The calculation formula of AGCA is defined as follows:
[0045] Where m represents the feature map and GAP represents the global average pooling operation. Function F r and F r ′ is a linear embedding function, usually implemented through 1×1 convolution. The adaptive graph convolution module (AGCM) calculates the weights between feature vertices. Where I is the input of the AGCM, and the adaptive adjacency matrix A consists of three parts: A0 (identity matrix), A1 (diagonal matrix), and A2 (learnable adjacency matrix). The weight matrix W helps learn the relationship between feature vertices. The calculation formula for fusion features is:
[0046] Among them L i、 G i and F i They represent the local, global and fusion features of the i-th stage (1≤i≤4), respectively.
[0047] F′i-1 = AvgPool3D(Conv3D(F i-1 )),
[0048]
[0049] Furthermore, as Figure 3 shown, the graph construction module can closely relate to predicting the malignancy of pulmonary nodules and the complex relationships within and between multimodal features. To comprehensively explore these relationships. This includes treating the features extracted from each modality as different nodes and constructing intra-modal and inter-modal graphs. The inter-modal graph contains text features and fused features from two time points t0 and t1. In contrast, the intra-modal graph contains local and global features from t0 and t1. This design of the bimodal graph can effectively capture the rich and complex information contained in multimodal data, and this method can minimize the risk of information loss. To cover all possible feature relationships within the same modal graph, all nodes are connected by fully connected edges. This practice enhances the representation of the interaction between features. Let the feature representation of each node be v0, v1,..., v n-1 , where the dimension of each feature vector is d. The graph G is represented as G = (V, E, X), where v is a set of n nodes, E is the set of edges, defined as E = {(i,j)|0 ≤ i,j < n, i ≠ j}. The node feature matrix X is given by X = [v0 v1... v n-1 , where v n-1 is the feature vector of the i-th node. Finally, GAT is used to capture the interactions within the modal graph.
[0050] Finally, as Figure 4 shown, the feature fusion module HCMGFM consists of two self-attention blocks (SAB) and one cross-attention block (CAB). The initial layer adopts a dual-path parallel SAB to independently process the bimodal input features. This layer uses an adaptive weight allocation mechanism to focus on the high-order semantic relationships within each modality, thereby enhancing and refining the specific features of each modality. The intermediate CAB establishes a bidirectional information channel between modalities. It uses the key-value query mechanism of multi-head cross-attention to create a modal association matrix and promote the interaction between different modalities. In the last layer, another SAB is used to globally recalibrate the fused joint features. This recalibration optimizes the attention weights in the fully connected feature space, reduces the potential alignment noise between modalities, and enhances the semantic consistency and spatial coherence of the fused features. This module follows a structured approach of "intra-modal optimization → cross-modal interaction → global optimization" to gradually achieve the deep fusion of bimodal semantics while maintaining the independence of single-modal features. Subsequent ablation experiments show that this design significantly improves the performance of multimodal tasks.
[0051] Step 3: Use the preprocessed dual time series lung nodule ROI images as input to repeatedly train the dual image spatiotemporal attention network, optimize the network parameters, and continuously perform iterative optimization;
[0052] Specifically, in this embodiment, our method is implemented on the GeForce RTX3090Ti GPU using the torch-2.1.0-cu12.1-cudnn8.9 framework. First, we pre-train MSFEE using the cross entropy loss function, Adam optimizer, a learning rate of 0.0001, and 200 cycles. Then, the entire DGSAN model is trained under the same settings. The momentum parameters are set to β1=0.5 and β2=0.999, and the parameters are updated every 20 cycles. Given the limitations of the dataset, we use 5-fold cross-validation to evaluate the performance of the model to ensure the reliability and generalization ability of the results, where the ROC curves of DGSAN and other methods on the NLST-cmst dataset are shown as follows. Figure 5 shown.
[0053] Step 4: The pre-processed lung nodule ROI image is input into the trained dual-image spatiotemporal attention network to output accurate malignant tumor classification results.
[0054] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network, characterized in that: The steps include: Step 1: Collect and preprocess image data and corresponding text data. The image data is a patient's CT scan sequence image, and the text data includes lung nodule annotation data and clinical data. After preprocessing, the image data is obtained to obtain a dual-time series lung nodule ROI image. Step 2: Construct a dual-image spatiotemporal attention network for classifying the malignancy of pulmonary nodules, wherein the dual-image spatiotemporal attention network includes a multi-scale feature extraction encoder, a bimodal graph construction module, a graph feature fusion module, and a fully connected layer; Step 3: Use the preprocessed dual-time series lung nodule ROI images and the text data corresponding to each time series as input to repeatedly train the dual-image spatiotemporal attention network, optimize the network parameters, and continuously perform iterative optimization; Step 4: The pre-processed lung nodule ROI image is input into the trained dual-image spatiotemporal attention network to output accurate malignant tumor classification results.
2. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 1, characterized in that: The preprocessing method is as follows: extracting a 3D region of interest with a size of 16x 64x 64 according to the nodule location of the image data, classifying it into malignant and benign subtypes, and saving the corresponding nodule diameter as a text feature.
3. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 1, characterized in that: The multi-scale feature extraction encoder includes a local feature extraction branch, a global feature extraction branch and a feature fusion branch.
4. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 3, characterized in that: In the local feature extraction branch, the dual-time series lung nodule ROI image obtained by preprocessing is input, and first a 3D convolution operation is performed through Conv3D, and then it passes through 4 layers of local feature extraction networks in sequence. The output of the previous layer of local feature extraction network is used as the input of the next layer of local feature extraction network, where the first layer of local feature extraction network consists of a regularization layer and a local feature extraction module, and the other three layers of local feature extraction networks consist of a regularization layer, Conv3D and a local feature extraction module. The last layer of local feature extraction network is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to finally obtain the final local feature.
5. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 4, characterized in that: In the global feature extraction branch, first, the dual-time series lung nodule ROI image obtained by input preprocessing is divided into 4×4×2 blocks, and then passes through 4 layers of global feature extraction networks in sequence. The output of the previous layer of global feature extraction network is used as the input of the next layer of global feature extraction network. The first layer of global feature extraction network is converted into a Token sequence suitable for Transformer processing through a linear encoder, and then the global feature is extracted through the global feature extraction module. The other three layers of global feature extraction networks are merged through Patch and then the global features are extracted through the global feature extraction module. The last layer of global feature extraction network is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to obtain the final global feature.
6. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 5, characterized in that: The feature fusion branch includes four layers of adaptive feature fusion blocks connected in sequence. The four adaptive feature fusion blocks correspond to the four-layer networks of the global feature extraction branch and the local feature extraction branch respectively. The adaptive feature fusion block fuses the local features output by the corresponding local feature extraction network, the global features output by the global feature extraction network, and the output of the adaptive feature fusion block of the previous layer, and serves as the input of the adaptive feature fusion block of the next layer; the last layer of adaptive feature fusion blocks is connected to the adaptive graph channel attention module, and adaptive average pooling, flattening and linear layer operations are performed in sequence through the adaptive graph channel attention module to finally obtain the final fusion feature.
7. The method for classifying malignant pulmonary nodules based on a dual-image spatiotemporal attention network according to claim 6, characterized in that: In the dual-modal graph construction module, the local features and global features of the dual time series t0 and t1 from the multi-scale feature extraction encoder are converted into the form of an intra-modal graph through a graph construction method, and the text features from the fully connected layer and the fusion features of t0 and t1 are converted into the form of an inter-modal graph through a graph construction method; the two graphs will be input into the subsequent graph feature fusion module.
8. The method for classifying malignant pulmonary nodules based on a dual-image spatiotemporal attention network according to claim 7, characterized in that: The graph construction method is as follows: Let the feature representation of each node be v0, v1,..., v n-1 , where the dimension of each feature vector is d. The graph G is represented as G = (V, E, X), where V is the set of n nodes, E is the set of edges, defined as E = {(i, j)|0 ≤ i, j < n, i ≠ j}; the node feature matrix X is given by X = [v0 v1... v n-1 , where v n-1 is the feature vector of the i-th node; finally, GAT is used to capture the interactions within the two modal graphs to obtain the respective fusion features between and within the modalities.
9. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 8, characterized in that: The graph feature fusion module includes two self-attention blocks and one cross-attention block.
10. The method for classifying malignant pulmonary nodules based on a dual-graph spatiotemporal attention network according to claim 9, characterized in that: In the graph feature fusion module, the initial layer adopts a dual-path parallel self-attention block to independently process the inter-modal and intra-modal bimodal input features obtained through graph construction, and the initial layer adopts an adaptive weight distribution mechanism; the middle layer adopts a cross-attention block, and uses the key-value query mechanism of cross-attention to create modal association matrices for intra-modal and inter-modal features, respectively; the last layer inputs the inter-modal and intra-modal fusion features, and uses the self-attention block to perform global recalibration.
11. The method for classifying malignant pulmonary nodules based on a dual-image spatiotemporal attention network according to claim 10, characterized in that: The output of the graph feature fusion module is output through a fully connected layer to output the final classification result.