Pulmonary nodule malignant risk dynamic prediction method and system based on space-time attention and multi-modal data guide fusion
Through the method of guiding fusion based on space-time attention and multimodal data, the problem of insufficient information fusion in CT imaging diagnosis of pulmonary nodules is solved, and efficient dynamic prediction of malignant risks of pulmonary nodules is achieved, which improves diagnostic accuracy and efficiency.
Patent Information
- Application Number
- CN202510597935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-26
AI Technical Summary
The prior art has problems such as insufficient single-modal information fusion in the diagnosis of pulmonary nodules, difficulty in capturing the spatiotemporal heterogeneity of malignant nodules, low medical diagnosis efficiency and poor consistency in the diagnosis of pulmonary nodules, difficult to distinguish benign and malignant nodules.
Using a method based on spatiotemporal attention and multimodal data guidance fusion, we use an open-source medical big model to extract image features, combine clinical data and CT images to perform feature weighted fusion and dynamic prediction, including discretization processing, mask marking, slice processing, spatiotemporal coding and gating mechanisms, to achieve intelligent interaction and adaptive integration of multimodal information.
It improves the prediction performance of malignant risk of lung nodules, overcomes the diagnostic bottlenecks of traditional methods, improves the prediction accuracy and efficiency of the model, alleviates the problems of equipment differences and inconsistencies in scanning protocols, and realizes dynamic spatiotemporal and spatial characteristics capture of lung nodules.
Smart Images

Figure CN120544869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and in particular to a method and system for dynamically predicting the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data-guided fusion. Background Art
[0002] Lung cancer is the leading cause of cancer-related deaths worldwide, and its early screening and accurate diagnosis face significant clinical needs. Differentiation between benign and malignant lung nodules is a core challenge in the early diagnosis of lung cancer, and its accurate identification is directly related to patients' treatment decisions and prognosis. Lung nodule puncture biopsy is the gold standard, but as an invasive procedure, it may cause complications such as pneumothorax, and it is difficult to sample sub-centimeter nodules. CT, with its high spatial resolution and rapid imaging characteristics, can clearly display the morphological characteristics of nodules. By deconstructing the morphological changes of nodules in time and space, the malignancy of lung lesions can be objectively reflected, which helps doctors understand the severity of the disease and rationally choose the timing of surgical intervention or targeted drug treatment.
[0003] Although the prevalence of CT technology has significantly improved the detection rate of lung nodules, traditional diagnostic methods rely on the visual evaluation of CT images by radiologists. Faced with benign and malignant nodules with highly similar morphological features, they have inherent defects such as strong subjectivity and high misdiagnosis rate. With the popularization of CT screening, the number of lung nodules detected has increased exponentially, and the workload of physicians and the risk of missed diagnosis have risen simultaneously. The annual number of lung nodules detected is large, but it takes a long time for professional physicians to evaluate a single case, and due to differences in physician experience, there are problems with clinical diagnostic efficiency and consistency; benign and malignant nodules show high similarity in key morphological features, and traditional measurement methods have difficulty capturing the spatiotemporal heterogeneity unique to malignant nodules; existing open source large models are insufficient in the extraction of image-specific features; single-time point analysis cannot capture the spatiotemporal evolution characteristics of nodules; single-modality diagnostic systems lack a multimodal information fusion mechanism, resulting in limited predictive performance. Therefore, how to efficiently and effectively distinguish between benign and malignant nodules remains a difficult problem that needs to be solved. Summary of the Invention
[0004] To this end, the present invention provides a dynamic prediction method and system for the malignant risk of lung nodules based on spatiotemporal attention and multimodal data-guided fusion, which solves the problems of insufficient feature extraction capabilities of existing lung nodule CT images and limited prediction performance due to lack of multimodal information.
[0005] According to the design scheme provided by the present invention, on the one hand, a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion is provided, comprising:
[0006] Acquiring multi-source pulmonary nodule data of the patient, wherein the multi-source pulmonary nodule data includes pulmonary CT images and clinical data collected during multiple follow-up visits;
[0007] Clinical data is discretized and standardized to obtain clinical features. Lung CT images are masked and sliced to obtain lung CT slice images collected at multiple follow-up visits. General medical visual features describing image text features in lung CT slice images are extracted using an open-source medical large model. Image features in lung CT slice images are extracted using a pre-trained visual Transformer model. Feature-guided optimization of image features is performed based on clinical and general medical visual features to obtain patient-specific image features. These features are then combined with the time intervals of multiple follow-up visits and the spatiotemporal attention mechanism to jointly map the patient-specific features to obtain spatiotemporal enhanced features of the patient images.
[0008] The gating mechanism is used to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features. The patient's weighted fusion features are input into a pre-trained classifier, and the classifier is used to obtain the patient's malignant pulmonary nodule risk prediction probability.
[0009] As a dynamic prediction method for malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion of the present invention, the clinical data is further discretized and standardized to obtain clinical features, including:
[0010] The clinical data collected at each follow-up visit of the patients were divided into discrete variables and continuous variables;
[0011] Discrete variables were coded and continuous variables were Z-score standardized to obtain the clinical characteristics of each follow-up clinical data of the patients.
[0012] As a dynamic prediction method for malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion of the present invention, further, mask marking and slicing of lung CT images are performed, including:
[0013] The lung nodules in the lung CT images collected at each follow-up visit of the patient were masked and image slices of a specified size were cut with the coordinates of the center point of the lung nodule as the center, and the slice images of the area of interest were selected;
[0014] The slice images are stitched together to obtain a plurality of continuous stitched slice images.
[0015] As a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, the present invention further uses an open source medical large model to extract general medical visual features of image text feature descriptions in lung CT slice images, including:
[0016] Set a prompt template, use the prompt template and the spliced slice image as input to the open source medical big model, and use the open source medical big model to extract common medical features related to patients and lung nodules. The prompt template is used to guide the model to obtain feature output related to lung nodules in the image. The open source medical big model uses Radiology-GPT;
[0017] The language model is used to embed general medical features to obtain general medical visual features.
[0018] As a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data-guided fusion, the present invention further extracts image features from lung CT slice images through a pre-trained visual Transformer model, and performs feature-guided optimization of the image features based on clinical features and general medical visual features, including:
[0019] A pre-trained ViT model with a dual-level spatiotemporal coding mechanism of temporal position coding and standard position coding is used as a visual Transformer model to extract image feature embedding of lung CT slice images in multiple follow-up acquisitions using the visual Transformer model. The inputs of the ViT model with a dual-level spatiotemporal coding mechanism are multiple lung nodule image slices with standard position coding and first-scan lung CT slice images, temporal position coding, and lung CT slice images in multiple follow-up acquisitions, and the position coding process of temporal position coding is expressed as: P st =P s +ΔT·e -λΔT , P s is the standard position encoding of the ViT model, ΔT is the follow-up interval, the decay coefficient λ is the model learning parameter, and λ=sigmoid(WλΔT), W λ is the time decay weight matrix;
[0020] The image feature embedding is mapped to the Value vector, and the clinical features and general medical visual features are mapped to the Query vector and Key vector respectively. The Value vector, Query vector and Key vector are used to interact to generate attention weights, and the image feature embedding is weighted by the attention weights. The weighted image feature embedding is residually connected and layer-normalized with the original image feature embedding to obtain the final patient image-specific features.
[0021] As a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion of the present invention, the time intervals of multiple follow-up visits are further combined and the spatiotemporal attention mechanism is used to jointly map the patient image-specific features to obtain the spatiotemporal enhancement features of the patient images, including:
[0022] Obtain follow-up time attenuation information based on follow-up time interval and attenuation coefficient;
[0023] Using the follow-up time decay information, the patient image-specific features corresponding to one of the follow-ups are time-decayed to obtain a first intermediate image feature, and the first intermediate image feature is combined with the patient image-specific features corresponding to another follow-up and subjected to feature projection mapping to obtain a spatiotemporal representation Query vector;
[0024] Using the follow-up time decay information, the patient image-specific features corresponding to another follow-up are temporally decayed to obtain a second intermediate image feature, and the second intermediate image feature is combined with the patient image-specific features corresponding to one of the follow-ups and subjected to feature projection mapping to obtain a spatiotemporal representation Key vector;
[0025] The patient image-specific features corresponding to the two follow-up visits are combined and feature-projected to obtain the spatiotemporal representation Value vector;
[0026] The spatiotemporal representation Query vector, spatiotemporal representation Key vector and spatiotemporal representation Value vector are used to calculate the spatiotemporal attention weighted features that capture the dynamic correlation between patient image features at different time points and different spatial locations. The spatiotemporal attention weighted features are combined with the patient image-specific features corresponding to the two follow-ups through residual connection and layer normalization to obtain the spatiotemporal enhanced features of the patient images.
[0027] As a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, the gating mechanism is further used to weight the fusion process of the patient's clinical features, general medical visual features and spatiotemporal enhancement features as follows: fused =α⊙F gen +β⊙F enh +γ⊙F clin ,in, ⊙ is the Hadamard product, σ is the Softmax function, For the splicing operation, W g is the gating weight matrix, F fused is the weighted fusion feature, F gen is a general medical visual feature, F enh is the spatiotemporal enhancement feature, F clin Clinical characteristics of patients.
[0028] On the other hand, the present invention also provides a dynamic prediction system for malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, comprising: a data collection module, a feature extraction module and a classification prediction module, wherein:
[0029] A data collection module is used to obtain multi-source pulmonary nodule data of patients, wherein the multi-source pulmonary nodule data includes pulmonary CT images and clinical data collected during multiple follow-up visits;
[0030] The feature extraction module is used to discretize and standardize clinical data to obtain clinical features; mask and slice lung CT images to obtain lung CT slice images collected at multiple follow-up visits; use an open source medical large model to extract general medical visual features that describe image text features in lung CT slice images; extract image features from lung CT slice images using a pre-trained visual Transformer model; and perform feature-guided optimization of image features based on clinical features and general medical visual features to obtain patient image-specific features. The module then combines the time intervals of multiple follow-up visits and uses a spatiotemporal attention mechanism to jointly map the patient image-specific features to obtain spatiotemporal enhanced features of the patient images;
[0031] The classification prediction module is used to use a gating mechanism to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features, input the patient's weighted fusion features into a pre-trained classifier, and use the classifier to obtain the patient's malignant pulmonary nodule risk prediction probability.
[0032] Beneficial effects of the present invention:
[0033] 1. The present invention constructs a multimodal guided fusion architecture and a two-level spatiotemporal coding attention, which is suitable for predicting the malignant risk of lung nodules in multiple follow-up CT images. The lung nodule CT images and follow-up clinical parameters are preprocessed to obtain 64×64×64 lung nodule sequence images and clinical parameter features. A two-level spatiotemporal coding mechanism of standard position coding and time position coding with time attenuation characteristics is used to guide the visual Transformer to extract specific features of the follow-up images based on general visual medical features and clinical parameter features. At the same time, the extracted follow-up image specific features are calculated using spatiotemporal attention to obtain spatiotemporal enhancement features of the follow-up images that fuse time and space characteristics. A dynamic gating mechanism is used to realize the multimodal fusion of general image features, spatiotemporal enhancement features and clinical parameter features. The fused features are passed through a multi-layer perceptron to output the malignant probability prediction of the lung nodules. This overcomes the contradiction between the surge in annual lung nodule examinations and the bottleneck of physician diagnostic efficiency, the difficulty in quantifying the spatiotemporal heterogeneity of benign and malignant nodules, the insufficient cross-modal feature fusion capability of traditional models, and the lack of time series modeling in single-time point diagnosis systems, thereby improving the model prediction performance.
[0034] 2. The present invention uses an open source large model to extract universal image features of CT images. Its core advantage lies in its ability to learn generalized representation capabilities across data sets through large-scale pre-training. This ability stems from the fact that the large model establishes a deep feature association system when pre-training on diverse medical imaging data. By integrating millions of multi-center, multi-device CT scan data, the large model constructs a universal understanding framework for anatomical structure, tissue density and morphological features in medical images. This pre-training mechanism enables the model to capture key visual patterns such as contrast differences between lung nodules and surrounding lung parenchyma, edge texture features, and internal structural heterogeneity. At the same time, it effectively alleviates the feature drift problem caused by differences in equipment parameters and inconsistent scanning protocols in traditional methods.
[0035] 3. The use of the spatiotemporal attention mechanism in the present invention can break through the static processing mode of temporal information of traditional methods. Although traditional methods such as three-dimensional convolutional neural networks can extract spatial features, they cannot effectively model the relationship between images at different follow-up time points. The spatiotemporal attention mechanism models images with different follow-up intervals through a learnable time decay factor, and associates the follow-up interval with the image feature representation, so that the model can adaptively adjust the attention intensity of image features with different time spans and capture the temporal change pattern of image features.
[0036] 4. The present invention utilizes a gated fusion mechanism to break through the static limitations of traditional feature fusion methods, and realizes intelligent interaction and adaptive integration of multimodal features through a dynamic weight allocation mechanism. Traditional methods such as feature splicing or weighted averaging often ignore the nonlinear correlation between different data modalities. In particular, when there are complex synergistic effects between image features and clinical parameters, static fusion can easily lead to information redundancy or suppression of key features. The gated fusion mechanism realizes adaptive dynamic weighting of general image features, specific image features and clinical parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Schematic diagram of the dynamic prediction process of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion in the embodiment;
[0038] Figure 2 This is a schematic diagram of the principle framework of the dynamic prediction algorithm for malignant risk of pulmonary nodules in the embodiment;
[0039] Figure 3 This is a schematic diagram of the open source large model feature extraction structure in the embodiment;
[0040] Figure 4 This is a schematic diagram of the visual Transformer feature extraction structure in the embodiment;
[0041] Figure 5 This is a schematic diagram of the workflow of the spatiotemporal attention mechanism in the embodiment;
[0042] Figure 6 Schematic diagram of the gated fusion mechanism workflow in the embodiment. DETAILED DESCRIPTION
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below with reference to the accompanying drawings and technical solutions.
[0044] In order to address the problems encountered in the diagnosis of benign and malignant pulmonary nodules, such as the time-consuming evaluation of single cases by professional physicians, the problems of clinical diagnostic efficiency and consistency due to differences in physician experience, the difficulty of traditional measurement methods in capturing the spatiotemporal heterogeneity unique to malignant nodules due to the high similarity in key morphological features between benign and malignant nodules, the insufficient ability of existing open source large models to extract image-specific features, the inability of single-time point analysis to capture the spatiotemporal evolution characteristics of nodules, and the limited prediction performance due to the lack of multimodal information fusion mechanism in single-modal diagnostic systems, an embodiment of the present invention provides a dynamic prediction method for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data-guided fusion, which specifically includes the following contents:
[0045] S101. Acquire multi-source lung nodule data of a patient, where the multi-source lung nodule data includes lung CT images and clinical data collected during multiple follow-up visits.
[0046] In actual clinical scenarios, conventionally acquired chest CT images, in addition to containing lung parenchyma and lesion areas, often contain anatomical components irrelevant to diagnosis, such as chest wall soft tissue, bone structure, and mediastinum. This redundant information can introduce artifact interference, significantly affecting the accuracy of extracting morphological features of lung nodules and distinguishing microinvasive growth. Specifically, a radiologist can first perform ROI annotation on continuous CT slices under lung window conditions, and then extract the region of interest based on the doctor's annotations, cropping the image of the region of interest to reduce interference from irrelevant background factors.
[0047] S102. Discretize and standardize the clinical data to obtain clinical features; mask and slice the lung CT images to obtain lung CT slice images collected at multiple follow-up visits; use the open source medical large model to extract general medical visual features of the image text feature description in the lung CT slice images; extract the image features in the lung CT slice images through the pre-trained visual Transformer model; and perform feature-guided optimization on the image features based on the clinical features and general medical visual features to obtain patient image-specific features; combine the time intervals of multiple follow-up visits and use the spatiotemporal attention mechanism to jointly map the patient image-specific features to obtain patient image spatiotemporal enhanced features.
[0048] like Figure 1 and 2As shown in the figure, by integrating multimodal information such as common features extracted from open source large models, multi-period CT image features and clinical parameters, a spatiotemporal evolution model is constructed, which can intelligently analyze the malignancy risk of lung nodules, provide decision support for clinicians' diagnosis, and help achieve the transformation from empirical judgment to data-driven clinical decision-making.
[0049] Among them, the clinical data is discretized and standardized to obtain clinical characteristics, which can be designed to include:
[0050] The clinical data collected at each follow-up visit of the patients were divided into discrete variables and continuous variables;
[0051] Discrete variables were coded and continuous variables were Z-score standardized to obtain the clinical characteristics of each follow-up clinical data of the patients.
[0052] Clinical parameters were discretized, such as EGFR mutation, gender, height, weight, age, nodule nature, nodule location, spiculation characteristics, vacuolar characteristics, lobulation characteristics, bronchial inflation characteristics, calcification, pleural traction or convexity, vascular typing, pathology, smoking status, personal cancer history, history of underlying diseases, family cancer history, seven lung cancer antibiotics, and dust exposure history. Discrete variables were converted to decimal codes, and continuous variables were standardized using the Z-score.
[0053] Mask labeling and slice processing of lung CT images can be designed to include:
[0054] The lung nodules in the lung CT images collected at each follow-up visit of the patient were masked and image slices of a specified size were cut with the coordinates of the center point of the lung nodule as the center, and the slice images of the area of interest were selected;
[0055] The slice images are stitched together to obtain a plurality of continuous stitched slice images.
[0056] Specifically, the open source medical big model is used to extract general medical visual features for describing image text features in lung CT slice images, which can be designed to include:
[0057] Set a prompt template, use the prompt template and the spliced slice image as input to the open source medical big model, and use the open source medical big model to extract common medical features related to patients and lung nodules. The prompt template is used to guide the model to obtain feature output related to lung nodules in the image. The open source medical big model uses Radiology-GPT;
[0058] The language model is used to embed general medical features to obtain general medical visual features.
[0059] Use open source medical models such as Radiology-GPT to extract general medical features, such as Figure 3 As shown in the figure, we can first select multiple slice images with ROI areas from two 64×64×64 CT image slice sequences, and then stitch these multiple images into a 64×64×N image in order of the number of layers, where n is the number of layers of the lung nodule image; then input this stitched image into the open source large model Radiology-GPT model, and set the prompt words as: "The following figures are multiple consecutive slices of lung nodule CT images. Please focus on analyzing the nodule properties, burr features, lobulation features, pleural traction features, whether the edge is smooth, solid components, vascular ingrowth, cavitation features, bronchial inflation features, and whether there is calcification." Through deep semantic understanding, we extract general medical features related to the lesion. These features are further embedded and processed by language models such as BERT to form general medical visual features F. gen1 and F gen2 .
[0060] Specifically, the image features in lung CT slice images are extracted through a pre-trained visual Transformer model, and the image features are feature-guided and optimized based on clinical features and general medical visual features. The pre-trained ViT model with a dual-level spatiotemporal coding mechanism of temporal position coding and standard position coding is used as the visual Transformer model. The visual Transformer model corresponds to different inputs, and its model weights are shared to use the visual Transformer model to extract image feature embedding of lung CT slice images in multiple follow-up acquisitions. Among them, the inputs of the ViT model with a dual-level spatiotemporal coding mechanism are standard position coding and multiple lung nodule image slices of the first scanned lung CT slice image, temporal position coding and lung CT slice images in multiple follow-up acquisitions, and the position coding process of the temporal position coding is expressed as: P st =P s +ΔT·e -λΔT , P s is the standard position encoding of the ViT model, ΔT is the follow-up interval, the decay coefficient λ is the model learning parameter, and λ=sigmoid(W λ ΔT), W λ is the time-attenuated weight matrix; the image feature embedding is mapped to the Value vector, the clinical features and general medical visual features are mapped to the Query vector and the Key vector respectively, and the Value vector, Query vector and Key vector are used to interact to generate the attention weight, so as to use the attention weight to weight the image feature embedding, and the weighted image feature embedding is residually connected and layer-normalized with the original image feature embedding to obtain the final patient image-specific features.
[0061] The image features extracted by ViT are embedded and feature mapped to obtain the Value vector. The general visual features and clinical parameter features are feature mapped and projected to obtain the Query vector and Key vector. The Key, Query and Value are interactively generated to generate attention weights for weighted image features. At the same time, the weighted image features are residually connected and layer-normalized with the original image features to obtain the final image-specific features F. spec The calculation process is as follows: V = F ViT W v , Q=F gen W q , K=F clin W k , where W k 、W q 、W v is the projection matrix; the attention calculation is Finally, we get image-specific features
[0062] F spec =LayerNorm(Attention+F ViT ).
[0063] In the extraction of general medical features, the image sequence is first arranged in spatial sequence intervals and spliced into a unified image. The images collected from multiple follow-up visits are input into the pre-trained open source medical model to generate a general medical description, which is further embedded through BERT to form a high-level semantic representation.
[0064] In the analysis of pulmonary nodule CT images, in order to efficiently extract common medical features and enhance the model's ability to understand complex lesions, such as Figure 3 As shown in the figure, first, the lung CT image slice sequences collected at multiple follow-up visits are arranged at spatial intervals, and the slices of the region of interest contained in each slice are extracted. Subsequently, these slices are spliced into a continuous image according to the spatial hierarchical order, and the semantic features of lesion changes at different time points are extracted.
[0065] To further enrich the expressive power of image features, the system feeds the stitched images into an open-source medical model for processing. This model, combined with a pre-trained medical knowledge base, extracts common medical features associated with lesions, such as nodule spiculation, lobulation, solid components, and the presence of important signs such as pleural traction, vascular ingrowth, or air bronchus. These features are further embedded in the BERT language model to form a unified feature representation for subsequent feature-guided fusion and malignancy probability prediction.
[0066] In the extraction of specific medical features, we focus on the specific features and spatial relationship modeling of images. The image slices collected from multiple follow-up visits are converted into Token sequences. The two-level spatiotemporal coding visual Transformer model with position coding and time-position coding with time attenuation characteristics is input. Based on the general visual medical features and clinical parameter features, the visual Transformer is guided to extract the specific features of the follow-up images. Subsequently, the interaction relationship between image tokens at different times is calculated through the self-attention mechanism to extract specific medical features, so that the model can better capture the complex structure and dynamic changes of the lesion area.
[0067] In order to effectively extract specific medical features with diagnostic value in CT image analysis of lung nodules, Figure 4 In the process, first, the lung CT image sequences collected during multiple follow-up visits are sorted according to time intervals to form multi-input slice image data. Subsequently, these slices are serialized into tokens, input into the weight-sharing visual Transformer model, and these tokens are serialized and encoded. The visual Transformer model has a two-level spatiotemporal encoding mechanism of time position encoding and position encoding. It takes multiple follow-up lung nodule CT image slice sequence data as input and models the image information at different time intervals at the same time. The Transformer can use the self-attention mechanism to capture the global correlation between slices, and at the same time combine multi-layer feature mapping to construct cross-space contextual information interaction, thereby generating image features with specific diagnostic significance.
[0068] By integrating the general semantic features and clinical parameter features of open-source medical models, these two features are used to guide the effective extraction of image features. Image features are embedded and mapped to obtain a value vector. General visual features and clinical parameter features are projected through feature mapping to obtain a query vector and a key vector. Attention weights are generated through the interaction of key, query, and value, and used to weight image features. The weighted image features are then residually connected with the original image features and layer-normalized to obtain the final image-specific features.
[0069] To effectively utilize general visual medical features and clinical parameter features to guide efficient image feature extraction, the image features extracted by the visual Transformer model are embedded and feature mapped to obtain a Value vector. The general visual features and clinical parameter features are feature mapped and projected to obtain a Query vector and a Key vector. The Key, Query, and Value are interactively generated to generate attention weights for weighting image features. The weighted image features are then residually connected and layer-normalized with the original image features to obtain the final image-specific features. By integrating the feature-guided representation module into each feature level of the visual Transformer model, the image feature representation is optimized to obtain multiple follow-up image-specific features.
[0070] Among them, the time intervals of multiple follow-up visits are combined and the spatiotemporal attention mechanism is used to jointly map the patient image-specific features to obtain the spatiotemporal enhancement features of the patient images, which can be designed to include:
[0071] Obtain follow-up time attenuation information based on follow-up time interval and attenuation coefficient;
[0072] Using the follow-up time decay information, the patient image-specific features corresponding to one of the follow-ups are time-decayed to obtain a first intermediate image feature, and the first intermediate image feature is combined with the patient image-specific features corresponding to another follow-up and subjected to feature projection mapping to obtain a spatiotemporal representation Query vector;
[0073] Using the follow-up time decay information, the patient image-specific features corresponding to another follow-up are temporally decayed to obtain a second intermediate image feature, and the second intermediate image feature is combined with the patient image-specific features corresponding to one of the follow-ups and subjected to feature projection mapping to obtain a spatiotemporal representation Key vector;
[0074] The patient image-specific features corresponding to the two follow-up visits are combined and feature-projected to obtain the spatiotemporal representation Value vector;
[0075] The spatiotemporal representation Query vector, spatiotemporal representation Key vector and spatiotemporal representation Value vector are used to calculate the spatiotemporal attention weighted features that capture the dynamic correlation between patient image features at different time points and different spatial locations. The spatiotemporal attention weighted features are combined with the patient image-specific features corresponding to the two follow-ups through residual connection and layer normalization to obtain the spatiotemporal enhanced features of the patient images.
[0076] The two time decay calculations of the spatiotemporal attention mechanism are and Where ΔT is the follow-up interval, and the decay coefficient λ 12 and λ 21is a learnable parameter, λ=sigmoid(W λ ΔT), W λ is the time decay weight matrix. Feature F spe1 With F spe2 Combine the time decay information separately, perform joint and feature mapping, and generate the spatiotemporal representation Query and Key, that is, The two image-specific features F spe1 and F spe2 The image feature F obtained by combining spe12 The feature projection mapping is used to represent the value in space and time, which is calculated as follows: The spatiotemporal attention is calculated by Query, Key and Value. The spatiotemporal attention is calculated as Attention results and feature F spe12 Through residual connection and layer normalization processing, the spatiotemporal enhanced feature F is finally obtained enh =LN(Attention+F spe12 ), ⊙ is the Hadamard product, For the splicing operation, W v ′、W v ′ and W v ′ is the projection matrix and LN is the layer normalization operation.
[0077] Combining convolutional coding with the self-attention mechanism enables efficient modeling of spatiotemporal features. By combining multiple follow-up specific medical images with different time attenuation information, preliminary spatiotemporal feature representations (Query and Key) are generated. The spatiotemporal joint image features are obtained by combining the specific features of multiple follow-up images and mapped to the spatiotemporal representation (Value) through feature projection. Subsequently, the correlation between features is normalized using the Softmax function. The attention results and the specific spatiotemporal joint image features are processed through residual connections and layer normalization to generate spatiotemporal enhanced features. This module effectively leverages the spatiotemporal modeling advantages of the attention mechanism, providing more comprehensive feature support for clinical diagnosis.
[0078] like Figure 5As shown in the figure, in order to fully explore the spatiotemporal characteristics of the lung nodule lesion area, the image features obtained by combining the time-attenuated features of one image with the features of another follow-up image are mapped to the spatiotemporal representation Query through feature projection. Similarly, the image features obtained by combining the time-attenuated features of one image with the features of another image are mapped to the spatiotemporal representation Key through feature projection. Through the effective association of image features and time information, the model can better model the temporal variation of follow-up image features. At the same time, the joint image features obtained by combining the specific features of the two images are mapped to the spatiotemporal representation Value through feature projection. The spatiotemporal attention is calculated from the Query, Key, and Value. Through the self-attention mechanism, the module can calculate the interaction between features, thereby capturing the dynamic association between different time points and different spatial locations. Finally, the attention results and the joint image features are connected through residual connection and normalization to generate spatiotemporal enhanced features.
[0079] S103. Use a gating mechanism to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features, input the patient's weighted fusion features into a pre-trained classifier, and use the classifier to obtain the patient's malignant lung nodule risk prediction probability.
[0080] For the spatiotemporal enhancement feature F enh , general image feature F gen and clinical featuresF clin , feature fusion is achieved through adaptive weight allocation, and its fusion formula is F fused =α⊙F gen +β⊙F enh +γ⊙F clin ,in, ⊙ is the Hadamard product, σ is the Softmax function, For the splicing operation, W g is the gating weight matrix.
[0081] In dynamic gated fusion, an adaptive gating mechanism is introduced to achieve efficient fusion of multimodal information. By uniformly encoding spatiotemporal features, general medical features, and clinical parameter features, the gating weight matrix is used to generate fusion weights, and the Softmax function is used to generate adaptive gating coefficients. The gating mechanism is used to effectively adjust the information weight distribution between modalities, so that key features can be given priority and the ability to collaboratively understand data from different modalities is improved. Figure 6As shown in the figure, through adaptive weight allocation, a unified feature representation is generated by combining spatiotemporal features, universal semantic features, and clinical parameters, thereby enhancing the collaborative understanding of multimodal data. The module receives spatiotemporal enhancement features, universal medical features generated by open source medical large models, and clinical auxiliary features. It uses a gating weight matrix to fuse these three features and generates an adaptive gating coefficient through the Softmax activation function. This coefficient dynamically adjusts the weight distribution between universal imaging features and spatiotemporal features to generate a multimodal fusion feature representation.
[0082] Furthermore, based on the above method, an embodiment of the present invention also provides a dynamic prediction system for the malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, comprising: a data collection module, a feature extraction module and a classification prediction module, wherein:
[0083] A data collection module is used to obtain multi-source pulmonary nodule data of patients, wherein the multi-source pulmonary nodule data includes pulmonary CT images and clinical data collected during multiple follow-up visits;
[0084] The feature extraction module is used to discretize and standardize clinical data to obtain clinical features; mask and slice lung CT images to obtain lung CT slice images collected at multiple follow-up visits; use an open source medical large model to extract general medical visual features that describe image text features in lung CT slice images; extract image features from lung CT slice images using a pre-trained visual Transformer model; and perform feature-guided optimization of image features based on clinical features and general medical visual features to obtain patient image-specific features. The module then combines the time intervals of multiple follow-up visits and uses a spatiotemporal attention mechanism to jointly map the patient image-specific features to obtain spatiotemporal enhanced features of the patient images;
[0085] The classification prediction module is used to use a gating mechanism to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features, input the patient's weighted fusion features into a pre-trained classifier, and use the classifier to obtain the patient's malignant pulmonary nodule risk prediction probability.
[0086] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0087] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0088] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0089] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.
[0090] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A dynamic prediction method for malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, characterized by: Include: Acquiring multi-source pulmonary nodule data of the patient, wherein the multi-source pulmonary nodule data includes pulmonary CT images and clinical data collected during multiple follow-up visits; Discretize and standardize clinical data to obtain clinical characteristics; Lung CT images are masked and sliced to obtain lung CT slice images collected over multiple follow-up visits. Open-source medical models are used to extract general medical visual features that describe image text features in lung CT slice images. Image features in lung CT slice images are extracted using a pre-trained visual Transformer model. Feature-guided optimization of these features is performed based on clinical and general medical visual features to obtain patient-specific image features. These features are then combined with the time intervals of multiple follow-up visits and the spatiotemporal attention mechanism to jointly map the patient-specific features, resulting in spatiotemporal enhancement of patient image features. The gating mechanism is used to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features. The patient's weighted fusion features are input into a pre-trained classifier, and the classifier is used to obtain the patient's malignant pulmonary nodule risk prediction probability.
2. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1, characterized in that: Discretize and standardize clinical data to obtain clinical characteristics, including: The clinical data collected at each follow-up visit of the patients were divided into discrete variables and continuous variables; Discrete variables were coded and continuous variables were Z-score standardized to obtain the clinical characteristics of each follow-up clinical data of the patients.
3. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1, characterized in that: Perform mask marking and slice processing on lung CT images, including: The lung nodules in the lung CT images collected at each follow-up visit of the patient were masked and image slices of a specified size were cut with the coordinates of the center point of the lung nodule as the center, and the slice images of the area of interest were selected; The slice images are stitched together to obtain a plurality of continuous stitched slice images.
4. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1 or 3, characterized in that: Use open source medical big models to extract general medical visual features for image text feature descriptions in lung CT slice images, including: Set a prompt template, use the prompt template and the spliced slice image as input to the open source medical big model, and use the open source medical big model to extract common medical features related to patients and lung nodules. The prompt template is used to guide the model to obtain feature output related to lung nodules in the image. The open source medical big model uses Radiology-GPT; The language model is used to embed general medical features to obtain general medical visual features.
5. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1 or 3, characterized in that: The pre-trained visual Transformer model is used to extract image features from lung CT slice images. Feature-guided optimization of these image features is performed based on clinical features and general medical visual features, including: A pre-trained ViT model with a dual-level spatiotemporal coding mechanism of temporal position coding and standard position coding is used as a visual Transformer model to extract image feature embedding of lung CT slice images in multiple follow-up acquisitions using the visual Transformer model. The inputs of the ViT model with a dual-level spatiotemporal coding mechanism are multiple lung nodule image slices with standard position coding and first-scan lung CT slice images, temporal position coding, and lung CT slice images in multiple follow-up acquisitions, and the position coding process of temporal position coding is expressed as: P st =P s +ΔT·e -λΔT , P s is the standard position encoding of the ViT model, ΔT is the follow-up interval, the decay coefficient λ is the model learning parameter, and λ=sigmoid(W λ ΔT), W λ is the time decay weight matrix; The image feature embedding is mapped to the Value vector, and the clinical features and general medical visual features are mapped to the Query vector and Key vector respectively. The Value vector, Query vector and Key vector are used to interact to generate attention weights, and the image feature embedding is weighted by the attention weights. The weighted image feature embedding is residually connected and layer-normalized with the original image feature embedding to obtain the final patient image-specific features.
6. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1, characterized in that: By combining the time intervals of multiple follow-up visits and using the spatiotemporal attention mechanism to jointly map the patient image-specific features, we can obtain the spatiotemporal enhancement features of the patient images, including: Obtain follow-up time attenuation information based on follow-up time interval and attenuation coefficient; Using the follow-up time decay information, the patient image-specific features corresponding to one of the follow-ups are time-decayed to obtain a first intermediate image feature, and the first intermediate image feature is combined with the patient image-specific features corresponding to another follow-up and subjected to feature projection mapping to obtain a spatiotemporal representation Query vector; Using the follow-up time decay information, the patient image-specific features corresponding to another follow-up are temporally decayed to obtain a second intermediate image feature, and the second intermediate image feature is combined with the patient image-specific features corresponding to one of the follow-ups and subjected to feature projection mapping to obtain a spatiotemporal representation Key vector; The patient image-specific features corresponding to the two follow-up visits are combined and feature-projected to obtain the spatiotemporal representation Value vector; The spatiotemporal representation Query vector, spatiotemporal representation Key vector and spatiotemporal representation Value vector are used to calculate the spatiotemporal attention weighted features that capture the dynamic correlation between patient image features at different time points and different spatial locations. The spatiotemporal attention weighted features are combined with the patient image-specific features corresponding to the two follow-ups through residual connection and layer normalization to obtain the spatiotemporal enhanced features of the patient images.
7. The method for dynamic prediction of malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion according to claim 1, characterized in that: The process of weighted fusion of patient clinical features, general medical visual features and spatiotemporal enhancement features using the gating mechanism is expressed as: fused =α☉F gen +β☉F enh +γ☉F clin in, ⊙ is the Hadamard product, σ is the Softmax function, For the splicing operation, W g is the gating weight matrix, F fused is the weighted fusion feature, F gen is a general medical visual feature, F enh is the spatiotemporal enhancement feature, F clin Clinical characteristics of patients.
8. A dynamic prediction system for malignant risk of pulmonary nodules based on spatiotemporal attention and multimodal data guided fusion, characterized by: Contains: data collection module, feature extraction module and classification prediction module, among which, A data collection module is used to obtain multi-source pulmonary nodule data of patients, wherein the multi-source pulmonary nodule data includes pulmonary CT images and clinical data collected during multiple follow-up visits; The feature extraction module is used to discretize and standardize clinical data to obtain clinical features; mask and slice lung CT images to obtain lung CT slice images collected at multiple follow-up visits; use an open source medical large model to extract general medical visual features that describe image text features in lung CT slice images; extract image features from lung CT slice images using a pre-trained visual Transformer model; and perform feature-guided optimization of image features based on clinical features and general medical visual features to obtain patient image-specific features. The module then combines the time intervals of multiple follow-up visits and uses a spatiotemporal attention mechanism to jointly map the patient image-specific features to obtain spatiotemporal enhanced features of the patient images; The classification prediction module is used to use a gating mechanism to perform weighted fusion of the patient's clinical features, general medical visual features, and spatiotemporal enhancement features, input the patient's weighted fusion features into a pre-trained classifier, and use the classifier to obtain the patient's malignant pulmonary nodule risk prediction probability.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.
Citation Information
Cited By
Pulmonary arterial hypertension detection method fusing image features and tricuspid regurgitation velocity
CN121304657A
Lung interstitial image analysis method and system based on clinical prior guidance feature fusion
CN121482029A
A Method and System for Interstitial Lung Imaging Analysis Based on Clinical Prior Guidance Feature Fusion
CN121482029B
Pulmonary nodule benign and malignant identification and prediction system based on multi-modal feature fusion
CN121617603A
Multi-modal fusion processing method based on multi-source data, and case report generation method and system
CN121839160A