Multi-task lung cancer brain metastasis lifetime prediction method based on multi-modal data fusion

Through multimodal data fusion and cross-modal attention fusion methods, the problem of insufficient knowledge sharing between non-small cell lung cancer foci segmentation and brain metastasis prediction tasks is solved, and the accuracy and stability of survival prediction are improved.

CN120147291APending Publication Date: 2025-06-13TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510307409.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The non-small cell lung cancer lesion segmentation task and brain metastasis classification judgment and survival prediction are usually regarded as independent tasks, and the commonality of their joint focus on tumor lesion areas is ignored, resulting in the inability to effectively share knowledge among tasks, affecting the model prediction performance.

Method used

A multi-task lung cancer brain metastasis survival prediction method is proposed based on multi-modal data fusion. By obtaining CT image data, gene data and clinical data of non-small cell lung cancer, lesion segmentation, feature extraction and cross-modal attention fusion are performed, and survival prediction is finally carried out through conditional guidance diffusion.

Benefits of technology

By integrating multimodal data and sharing knowledge between tasks, the accuracy and stability of survival prediction are improved, the generalization ability of the model is enhanced, and the problem of insufficient knowledge sharing among multiple tasks is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147291A_ABST
    Figure CN120147291A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task lung cancer brain metastasis lifetime prediction method based on multi-modal data fusion, and relates to the field of medical image processing, and the method comprises the steps: obtaining non-small cell lung cancer CT image data, gene data and clinical data; performing focus segmentation on the non-small cell lung cancer CT image data; feature extraction is carried out based on the focus segmentation result, and radiomics features and depth image features are obtained; performing feature extraction on the gene data to obtain gene features; missing value deletion and numeralization preprocessing are carried out on the clinical data to obtain clinical features; local cross-modal attention fusion is carried out on the radiomics features and the depth image features; performing global cross-modal attention fusion on the preliminary mixed image features, the gene features and the clinical features to obtain original mixed features; and performing conditional guidance diffusion based on the original mixed features to obtain a lung cancer brain metastasis classification judgment result and a lifetime prediction risk. According to the invention, the lifetime prediction accuracy and stability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical image processing, and particularly to a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion. Background Art

[0002] Non-small cell lung cancer (NSCLC) is the most common malignant tumor and a common cause of death related to diseases globally. Its high recurrence and metastasis rates result in a relatively short overall survival period. Therefore, accurately predicting the survival period classification results of lung cancer brain metastasis based on multimodal data is crucial. However, the performance of survival period classification is affected by multiple factors such as gene mutations, lesion size, and clinical data (such as KPS score). For doctors, the amount of cancer-related data is huge, making it difficult to comprehensively master and fully utilize. Computer-aided diagnosis technology can simulate the doctor's diagnosis process, reduce the doctor's workload, and improve the accuracy of prognosis prediction. Foreign research started earlier. Fujita et al. analyzed the impact of EGFR mutation status on the survival period of non-small cell lung cancer brain metastasis through the Kaplan-Meier method and the Cox proportional hazards regression model, thus achieving precise stratified management. Zhou et al. developed a deep learning model based on pathological sections to predict non-small cell lung cancer brain metastasis. This model uses ResNet-18 combined with pre-trained weights to extract H&E stained pathological section features and evaluates the progression risk of brain metastasis through median pooling. This method effectively differentiates lung cancer brain metastasis from disease-free progression, thus providing a new auxiliary decision-making tool for precision medicine. Haim et al. developed an image-based deep learning model to predict the mutation status of the driver gene EGFR in non-small cell lung cancer brain metastasis. This model uses ResNet-50 combined with pre-trained weights to extract MRI image lesion features and cascades MLP (multi-layer perceptron) for EGFR mutation judgment, with high accuracy, achieving non-invasive molecular typing and providing a new technical means for precision medicine. Yu et al. constructed a multi-task learning model based on SegFormer. The tumor image data is input into the VGG network to extract key features, and then the key features are input into the multi-layer perceptron MLP and the retrieval module composed of a hash coding layer and a binary coding layer for classification. At the same time, cross-layer fusion of feature maps and lesion segmentation are performed. Compared with single-task prognosis prediction, the system can diagnose cancer more accurately. Chu et al. constructed a multi-task learning model based on Transformer. First, four-modal MRI images are input into ResNet to extract key features, and then the features under different modalities are input into Transformer for feature fusion. Finally, the fused features are input into the multi-layer perceptron MLP for microvascular invasion MVI classification and recurrence-free survival prediction, thus achieving personalized management. Gu et al. constructed a multi-task survival model DeepMTS based on SegNet. First, the patient's PET / CT image is input into the cascaded network of SegNet and DenseNet to obtain the lesion segmentation result and the deep radiomics score. Then, the patient's PET / CT image is input into Lasso-Cox for regression analysis to obtain the traditional radiomics score. Finally, the traditional radiomics score, the deep radiomics score, and clinical data are sent to the multi-layer perceptron MLP, and the survival period prediction result is obtained in combination with the nomogram.Both image lesion segmentation and survival prediction show high accuracy, significantly improving the diagnostic efficiency. Tan et al. constructed DPTBNet, which uses SwinUnetR as a shared backbone network to extract common features for lung cancer classification and segmentation tasks. In the encoder part, SwinTE is used to extract high-level features while preserving local and global spatial relationships. In the decoder part, Unetr layers are used to upsample these features to the original resolution and combine them with the corresponding features in the encoder through skip connections. The classification task can provide additional constraints and guidance to help the model better understand and learn features related to tuberculosis lesions, thus improving the performance of the semantic segmentation task. Fu et al. constructed a multi-task survival model SG-Fusion, which jointly uses WSI pathological images, clinical data, and gene data for glioma staging and survival prediction. First, SwinTE and GCN are used to extract pathological image features and gene features respectively, and then cross-attention is used to enhance information interaction, and contrastive learning is used to enhance the model's ability to identify common and different features. Cancer staging and survival prediction show high accuracy, significantly improving the diagnostic efficiency compared to single-task single-modal prognostic models. Currently, relevant domestic research is also gradually emerging. Xiao Xiaojiao et al. constructed the Mt-C-Mmf multi-task model, which uses Unet as the segmentation branch and FastRcnn as the detection branch, and combines multi-modal MRI images to achieve accurate tumor segmentation and classification, thus providing a safe, time-saving, and accurate auxiliary diagnostic tool for clinicians. Wen Han et al. constructed an MTL model to simultaneously perform tumor classification and segmentation on multi-modal CE-MRI images. The segmentation subnet uses Unet combined with boundary-aware attention to solve the problem of tumor over-segmentation, and the classification subnet uses FCNN for classification prediction, which can provide reference for the clinical diagnosis and treatment of cancer patients. Zhu Zhengqun et al. extracted radiomics features and deep learning features from CT images. First, the least absolute shrinkage and selection operator method is used to reduce the dimensions of radiomics features and deep learning features respectively and calculate the radiomics score (Radscore) and deep learning score (Deepscore). Then, a multi-factor logistic regression analysis is used to establish a prediction model to achieve the prediction of tumor radiotherapy and chemotherapy efficacy. Compared with single-task single-modal prediction models, it provides a more effective, fast, and non-invasive prediction method for clinical practice. Generally speaking, with the continuous development of deep learning technology, especially in the field of multi-modal multi-task learning, the research space and application field for predicting the survival period of non-small cell lung cancer brain metastases are very broad.

[0003] However, the task of non-small cell lung cancer lesion segmentation and brain metastasis classification and survival prediction are often regarded as independent tasks, ignoring their commonality of focusing on the tumor lesion area, resulting in the ineffective sharing of knowledge between tasks, which in turn affects the prediction performance of the model. The irregular morphology and fuzzy boundaries of non-small cell lung cancer lesions make segmentation difficult, affecting the accurate extraction of image features. The heterogeneity of genetic, clinical and imaging data makes it difficult to fuse multimodal data features, affecting the effective integration of key information by the model, and making feature mining lack of comprehensiveness and rationality. Different modal data (such as imaging, genomic, clinical data, etc.) may have a high degree of correlation or repetitiveness in the feature fusion process, resulting in the introduction of data redundancy and noise, which is not conducive to the extraction of key features by the model and reduces the accuracy and stability of survival prediction.

[0004] In summary, the multi-task and multi-modal survival prediction model faces challenges such as insufficient knowledge sharing between different tasks, difficult lesion segmentation and feature extraction, and difficulty in fusing multi-modal heterogeneous data features. Therefore, it is necessary to further explore effective learning strategies and data fusion methods to improve the accuracy of the model's survival prediction. Summary of the invention

[0005] The purpose of this application is to provide a multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion, which can improve the accuracy and stability of survival prediction.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] The present application provides a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion, and the multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion includes:

[0008] Obtain CT imaging data, genetic data and clinical data of non-small cell lung cancer.

[0009] Perform lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result.

[0010] Based on the lesion segmentation result, image data features are extracted to obtain radiomics features and deep image features; the radiomics features are features extracted based on the pyradiomics package; the deep image features are features extracted based on the combination of the SwinTF encoder and the DenseNet encoder.

[0011] Feature extraction is performed on the gene data to obtain gene features; the gene features are features extracted based on the SGCN encoder.

[0012] The clinical data is first preprocessed by deleting missing values and numericalizing, and then feature extraction is performed to obtain clinical features; the clinical features are features extracted based on an MLP encoder.

[0013] The radiomics features and the depth image features are subjected to local cross-modal attention fusion to obtain preliminary mixed features.

[0014] The preliminary mixed image features, the gene features, and the clinical features are subjected to global cross-modal attention fusion to obtain original mixed features.

[0015] Based on the original mixed features, conditional guided diffusion is performed to obtain the classification judgment result of lung cancer brain metastasis and the survival period prediction risk.

[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application:

[0017] The present application provides a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion. The method includes: obtaining non-small cell lung cancer CT image data, gene data, and clinical data; performing lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result; extracting image data features based on the lesion segmentation result to obtain radiomics features and depth image features; the radiomics features are features extracted based on the pyradiomics package; the depth image features are features extracted by combining a SwinT F encoder and a DenseNet encoder; extracting features from the gene data to obtain gene features; the gene features are features extracted based on an SGCN encoder; performing missing value deletion and numerical preprocessing on the clinical data first, and then extracting features to obtain clinical features; the clinical features are features extracted based on an MLP encoder; performing local cross-modal attention fusion on the radiomics features and the depth image features to obtain a preliminary mixed feature; performing global cross-modal attention fusion on the preliminary mixed image feature, the gene feature, and the clinical feature to obtain an original mixed feature; performing conditional guided diffusion based on the original mixed feature to obtain a lung cancer brain metastasis classification judgment result and a survival prediction risk. By integrating non-small cell lung cancer CT image data, gene data, and clinical data, more comprehensive and effective tumor information can be obtained. Using the lesion segmentation result as the input for brain metastasis classification and survival prediction for knowledge sharing can enhance task relevance, thereby jointly optimizing each task and enhancing model performance. By performing lesion segmentation on non-small cell lung cancer CT image data, the non-small cell lung cancer lesion area can be effectively focused on, and more comprehensive image features can be obtained. Furthermore, through local cross-modal attention fusion and global cross-modal attention fusion, the features of different modalities of non-small cell lung cancer can be better fused, thereby improving the accuracy of survival prediction. By breaking the correlation of multi-modal data through conditional guided diffusion, optimizing the feature distribution, reducing information redundancy, and enhancing the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is an application environment diagram of a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion in an embodiment of the present application.

[0020] Figure 2Schematic flowchart of a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion provided by an embodiment of the present application.

[0021] Figure 3 Schematic diagram of the structure of a multi-task survival prediction model for non-small cell lung cancer brain metastasis based on multi-modal data provided by an embodiment of the present application.

[0022] Figure 4 Schematic diagram of feature extraction by fusing frequency domain and statistical moments under the Swin Transformer attention mechanism provided by an embodiment of the present application.

[0023] Figure 5 Schematic diagram of a local attention feature fusion method provided by an embodiment of the present application.

[0024] Figure 6 Schematic diagram of a global attention feature fusion method provided by an embodiment of the present application.

[0025] Figure 7 Schematic diagram of a conditional-guided diffusion method for eliminating data redundancy provided by an embodiment of the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0027] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0028] In non-small cell lung cancer, the lesion area in CT image data usually appears as nodules or masses in the lung parenchyma, reflecting the morphological characteristics of the tumor. Its gene data contains driver gene mutation information such as EGFR, ALK, ROS1, etc., reflecting the molecular biological characteristics of the tumor. Clinical data includes KPS score, age, gender, etc., which can reflect the patient's physical condition and disease progression. Using traditional radiomics or genomics alone to mine the data characteristics of non-small cell lung cancer lacks comprehensiveness and rationality. The effect of feature extraction is not satisfactory, and lesion segmentation is independent of brain metastasis classification and survival prediction tasks. The common feature of simultaneously focusing on the tumor region characteristics of the three tasks is not considered, resulting in low data utilization efficiency, difficult lesion segmentation, and low accuracy of brain metastasis classification and survival prediction. This application proposes a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion, enhancing task relevance through multi-task learning, and fusing multi-modal features through local and global attention mechanisms to improve the segmentation fineness of the lesion area in non-small cell lung cancer image data, the accuracy of brain metastasis classification judgment and survival prediction.

[0029] Currently, tumor prediction and prognosis technologies represented by deep learning have been widely studied. Commonly used multi-task multi-modal prognosis prediction models include DeepMTS, DPTBNet, etc. In addition, the research on attention mechanisms has further enhanced the network prediction performance, and the research on diffusion mechanisms is beneficial to enhancing the generalization ability of the model. Therefore, there are also different model solutions for the non-small cell lung cancer survival prediction problem. However, due to differences in multi-task model learning strategies, multi-modal feature fusion, data redundancy removal, etc., the performance of the models also varies to a certain extent. In this application, CT image features, gene features, and clinical features are characterized by SwinTF encoder, SGCN encoder, and MLP encoder. The DenseUnetR decoder automatically segments the lesion area. The extracted features are fused across modalities using GAT to obtain mixed features. The multi-modal data redundancy is eliminated through conditional guided diffusion CFD. Finally, the FC layer outputs the non-small cell carcinoma brain metastasis classification judgment and death risk, and the survival period is predicted in combination with Nomograph.

[0030] The multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send non-small cell lung cancer CT image data, gene data, and clinical data to the server 104. After receiving the non-small cell lung cancer CT image data, gene data, and clinical data, the server 104 performs lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result; based on the lesion segmentation result, image data feature extraction is performed to obtain radiomics features and depth image features; the radiomics features are features extracted based on the pyradiomics package; the depth image features are features extracted based on the combination of the SwinT F encoder and the DenseNet encoder; feature extraction is performed on the gene data to obtain gene features; the gene features are features extracted based on the SGCN encoder; the clinical data is first preprocessed by deleting missing values and numericalizing, and then feature extraction is performed to obtain clinical features; the clinical features are features extracted based on the MLP encoder; the radiomics features and the depth image features are fused by local cross-modal attention to obtain preliminary mixed features; the preliminary mixed image features, the gene features, and the clinical features are fused by global cross-modal attention to obtain original mixed features; based on the original mixed features, conditional guided diffusion is performed to obtain a lung cancer brain metastasis classification judgment result and a survival period prediction risk. The server 104 can feedback the obtained lung cancer brain metastasis classification judgment result and survival period prediction risk to the terminal 102. In addition, in some embodiments, the multi-task lung cancer brain metastasis survival period prediction method based on multi-modal data fusion can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly perform survival period prediction on non-small cell lung cancer CT image data, gene data, and clinical data, or the server 104 can obtain non-small cell lung cancer CT image data, gene data, and clinical data from the data storage system and perform survival period prediction on the non-small cell lung cancer CT image data, gene data, and clinical data.

[0031] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, and tablet computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0032] In an exemplary embodiment, such as Figure 2As shown, a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion is provided. This method is executed by a computer device, which can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiment of the present application, taking the application of this method to Figure 1 server 104 in

[0033] S1: Obtain non-small cell lung cancer CT image data, gene data, and clinical data.

[0034] S2: Perform lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result.

[0035] S3: Extract image data features based on the lesion segmentation result to obtain radiomics features and depth image features; the radiomics features are features extracted based on the pyradiomics package; the depth image features are features extracted by combining the SwinT F encoder and the DenseNet encoder.

[0036] S4: Extract features from the gene data to obtain gene features; the gene features are features extracted based on the SGCN encoder.

[0037] S5: First, perform missing value deletion and numerical preprocessing on the clinical data, and then extract features to obtain clinical features; the clinical features are features extracted based on the MLP encoder.

[0038] S6: Perform local cross-modal attention fusion on the radiomics features and the depth image features to obtain a preliminary mixed feature.

[0039] S7: Perform global cross-modal attention fusion on the preliminary mixed image feature, the gene feature, and the clinical feature to obtain an original mixed feature.

[0040] S8: Perform conditional guided diffusion based on the original mixed feature to obtain a lung cancer brain metastasis classification judgment result and a survival prediction risk.

[0041] By implementing the above steps S1 to S8, the present application improves the prediction accuracy and generalization ability of the model by integrating multi-modal information and sharing knowledge between tasks. Specifically, by integrating non-small cell lung cancer CT image data, gene data, and clinical data, more comprehensive and effective tumor information can be obtained. The lesion segmentation results are used as the input for brain metastasis classification and survival prediction for knowledge sharing, enhancing task relevance, thereby jointly optimizing each task and enhancing the model performance. By performing lesion segmentation on non-small cell lung cancer CT image data, the lesion area of non-small cell lung cancer can be effectively focused on, and more comprehensive image features can be obtained. Furthermore, through local cross-modal attention fusion and global cross-modal attention fusion, the feature of different modalities of non-small cell lung cancer data can be better integrated, thus improving the accuracy of survival prediction. By conditional-guided diffusion, the correlation of multi-modal data is broken, the feature distribution is optimized, the information redundancy is reduced, and the generalization ability of the model is improved.

[0042] As Figure 3 shown, a multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion includes the following steps:

[0043] (1) Extract radiomics features through the Pyradiomics package, extract deep image features by combining the SwinT F encoder and the DenseNet encoder, extract gene features through the SGCN encoder, and extract clinical features through the MLP encoder. Among them, the SGCN encoder extracts features from gene data through two graph convolutional layers, gradually transforming the feature dimension from (636, 16) to (636, 64) and (636, 32). Finally, the features of all nodes are flattened and mapped through the FC layer to obtain gene features with a dimension of (1, 32). The MLP encoder extracts clinical data features through three fully connected layers, gradually transforming the feature dimension from (1, 6) to (1, 64) and (1, 32). Finally, the output layer features are used as clinical features.

[0044] (2) Perform cross-layer fusion through the DenseUnetR decoder to automatically segment the lesion.

[0045] (3) Perform local cross-modal attention fusion on the extracted traditional radiomics features and deep image features. Then, the obtained mixed image features, gene features, and clinical features are sent into the heterogeneous graph and global cross-modal attention fusion is performed using the multi-head multi-order GAT to obtain mixed features.

[0046] (4) Then, eliminate data redundancy through tumor TNM stage-guided diffusion CFD to avoid overfitting during the training process.

[0047] (5) Finally, the FC layer outputs the classification judgment result of lung cancer brain metastasis and the survival prediction risk.

[0048] Based on the above steps, the objectives of this application are as follows:

[0049] 1. In view of the problem that existing methods often regard lesion segmentation, brain metastasis classification judgment, and survival period prediction as independent tasks without considering the correlation between tasks, this application proposes a multi-task learning method (SGAFN) based on the commonality of simultaneously focusing on tumor region features in the three tasks of non-small cell lung cancer imaging data lesion segmentation, brain metastasis classification, and survival period prediction. The lesion segmentation results of the imaging data are used as the input for classification judgment and survival period prediction for knowledge sharing, enhancing task correlation, and thus jointly optimizing each task.

[0050] 2. In view of the problem that the lesion morphology of non-small cell lung cancer imaging data is irregular and the boundary is blurred, resulting in greater segmentation difficulty and affecting the accurate extraction of image features, this application proposes a feature extraction method FMCA based on 3D Swin-DenseUnetR for lesion segmentation and frequency domain-statistical moment fusion. Local and global information is extracted through the attention sliding window mechanism, and multi-scale features are densely connected and cross-layer fused to effectively focus on the lesion area of non-small cell lung cancer imaging data and improve the segmentation fineness. Statistical moments and frequency domain information are extracted through FMCA to obtain more comprehensive image features.

[0051] 3. In view of the problem that the mining of non-small cell lung cancer imaging genomics data features lacks comprehensiveness and rationality, resulting in unsatisfactory feature extraction effects, this application proposes a non-small cell lung cancer survival period prediction method (MHMOGAT) based on a local-global attention mechanism fusion of multi-modal features for heterogeneous graphs. Feature fusion through the attention mechanism can select effective and complementary multi-modal features from multi-modal lung cancer data, which is beneficial to improving the accuracy of lung cancer survival period prediction.

[0052] 4. In view of the problem that in the analysis of non-small cell lung cancer multi-modal data, different modalities of data (such as imaging, genomics, clinical data, etc.) may have high correlations or repetitions in certain features, resulting in overfitting, low computational efficiency, and may affect the generalization ability of the model, this application proposes to eliminate redundant information in multi-modal data through a conditional guided diffusion mechanism, avoid overfitting during the training process, and enhance the generalization ability of the model.

[0053] As an alternative implementation, in step S2, it specifically includes:

[0054] S21: Use the SwinTF encoder to extract image features at different positions of the non-small cell lung cancer CT imaging data through the sliding window SW-MSA to obtain a number of depth feature maps.

[0055] S22: Convert a number of the depth feature maps to obtain a number of channel feature maps.

[0056] S23: Extract the in-channel statistical moment information of several of the channel feature maps using EMA.

[0057] S24: Extract the frequency domain information of several of the channel feature maps using FFT.

[0058] S25: After splicing and dimensionality reduction of the in-channel statistical moment information and the frequency domain information, obtain the lesion segmentation result.

[0059] Specifically, the feature extraction method FMCA for lesion segmentation and frequency domain-statistical moment fusion based on 3D Swin-DenseUnetR is as follows:

[0060] The SwinTF encoder extracts image features at different positions through the sliding window SW-MSA, and uses 4 blocks for downsampling to capture global multi-scale information, where H, W, and D represent the height, width, and depth of the CT image data respectively.

[0061] The DenseUnetR decoder introduces dense connectivity on the basis of the classical Unet, and performs cross-layer feature fusion during upsampling to restore the scale of the lesion image.

[0062] Introduction to FMCA: Generate feature maps with n channels from the depth feature maps obtained by downsampling through a 1*1*1 convolutional kernel, capture the global spatial features of the channel feature maps using EMA, calculate the mean, variance, and skewness of each channel feature map, splice different-order statistical moments and combine with ReLU activation, perform Fourier transform on the channel feature maps using FFT, extract the frequency domain combined features of each channel feature map, splice the in-channel statistical moment information and the frequency domain information first and then reduce the dimension through a fully connected layer to obtain the depth image features with a dimension of (1, n / 2).

[0063] Description of the composition of the lesion segmentation model architecture:

[0064] The SwinTF encoder is based on the Swin Transformer architecture and uses the Shifted Window Multi-Head Self-Attention (SW-MSA) mechanism for feature extraction. It consists of four parts: Patch Partition, Linear Embedding, Merging, and Block, which achieve multi-scale feature extraction with progressive downsampling. The CT image data (H×W×D×2) is first divided into patches of a fixed size of 2×2×2 by Patch Partition to form a token sequence, and then linearly projected into a high-dimensional feature space through Linear Embedding. Subsequently, in Stage1-4, progressive downsampling is performed through layer-by-layer Merging, expanding from the dimension of H / 2×W / 2×D / 2×8 to H / 32×W / 32×D / 32×128, enhancing the global information perception ability. Inside the Swin Transformer Block, the W-MSA (window attention) calculates the attention within a local window, while the SW-MSA (shifted window attention) enhances global information extraction through cross-window interaction. Finally, a multi-scale deep feature map is output and passed to DenseUnetR for lesion segmentation. The DenseUnetR decoder is improved based on U-Net, introducing the Dense Connectivity mechanism, enabling each layer of features to be not only passed to the next layer but also cross-layer connected to higher layers to achieve feature reuse. At the same time, the Res-Block is used to retain key information to prevent gradient disappearance, and combined with upsampling, the resolution is gradually restored. The image is restored from H / 32×W / 32×D / 32×128 to H×W×D×2, where 2 represents the background and lesion channel regions, thus achieving accurate lesion segmentation.

[0065] The feature extraction and the obtained features are described in detail as follows:

[0066] As Figure 4 shown, the feature extraction process is as follows: First, the input CT image data is patch-segmented to obtain the feature map X 0 , and through the FMCA (Statistical Moment-Frequency Domain Fusion Feature Extraction) module, the statistical moment-frequency domain fusion features y 0 (mean, variance, skewness, amplitude, phase) of the feature map with a channel depth of 8 and a dimension of (1, 4) are obtained. Subsequently, feature extraction is performed through 4 Swin Transformer Blocks. In each stage, downsampling is performed through Merging, gradually reducing the feature maps X 1 , X 2 , X 3 , X 4The size is increased and the number of channels is increased. The cascaded FMCA separately obtains the statistical moments of the feature maps with channel depths of 16, 32, 64, and 128 - the frequency - domain fusion feature y 1 , y 2 , y 3 , y 4 , whose dimensions are (1, 8), (1, 16), (1, 32), and (1, 64) respectively. The five statistical moment - frequency - domain fusion features are concatenated to obtain an image feature with a dimension of (1, 120). The CT image data is multiplied point - by - point with the lesion segmented by 3D Swin - DenseUnetR to obtain the lesion area. The DenseNet network composed of three cascaded Denseblocks with 64, 128, and 256 output channels is used to extract features from it, and then the feature maps X 5 , X 6 , X 7 output by the three blocks are compressed in channels. After batch normalization (BN), ReLU activation, and FMCA operations, image features y 5 , y 6 , y 7 with dimensions of (1, 16), (1, 32), and (1, 64) respectively are obtained. Then they are concatenated to obtain an image feature with a dimension of (1, 112). Finally, the image feature extracted by the SwinTF encoder is concatenated with the image feature extracted by DenseNet to obtain a depth image feature X g .

[0067] The expression of the depth image feature is as follows:

[0068]

[0069] y l = concat(EMA(X l ), FFT(X l ))), l = 0, 1…, 7 (3);

[0070] X g = concat(y 0 , y 1 , y 2 , y 3 , y 4 , y 5 , y 6 , y 7 ) (4);

[0071] Among them, EMA(X l ) is the statistical feature composed of the mean, variance, and skewness of the feature map; X lFeature maps output at different channel depths; α k is the k-th weight parameter, used to assign different weights to different moment terms in the extended moment aggregation formula; E(X l ) is the mean of the channel feature map; M k (X l ) is the variance and skewness; FFT(X l ) is the frequency-domain feature composed of different frequency components of the feature map; is the weight parameter in the frequency-domain aggregation formula; N is the number of pixels; X l (k) is the k-th frequency component of the channel feature map; is the rotation factor, representing the periodic oscillation of the complex exponential; X l (n) is the pixel value of the feature map; y l is the image feature formed by splicing the statistical features and frequency-domain features of the channel feature maps with different depths; X g is the depth image feature.

[0072] As an optional implementation manner, in step S6, it specifically includes:

[0073] S61: Perform a linear transformation on the radiomics features to generate a first query vector, a first key vector, and a first value vector.

[0074] S62: Perform a linear transformation on the depth image features to generate a second query vector, a second key vector, and a second value vector.

[0075] S63: Use the Softmax function to calculate the matrix product of the first query vector and the first key vector to obtain a first attention score.

[0076] S64: Use the Softmax function to calculate the matrix product of the second query vector and the second key vector to obtain a second attention score.

[0077] S65: Weightedly aggregate the first attention score and the first value vector, and perform an element-wise sum with the radiomics features to obtain a first enhanced radiomics feature.

[0078] S66: Weightedly aggregate the second attention score and the second value vector, and perform an element-wise sum with the depth image features to obtain a first enhanced depth image feature.

[0079] S67: Perform a linear transformation on the first enhanced radiomics feature to generate a new first query vector, a new first key vector, and a new first value vector.

[0080] S68: Perform a linear transformation on the first enhanced depth image features to generate new second query vectors, new second key vectors, and new second value vectors.

[0081] S69: Use the Softmax function to calculate the matrix product of the new first query vector and the new first key vector to obtain a new first attention score.

[0082] S610: Use the Softmax function to calculate the matrix product of the new second query vector and the new second key vector to obtain a new second attention score.

[0083] S611: Weightedly aggregate the new first attention score with the new first value vector and perform an element-wise summation with the first enhanced radiomics features to obtain second enhanced radiomics features.

[0084] S612: Weightedly aggregate the new second attention score with the new second value vector and perform an element-wise summation with the first enhanced depth image features to obtain second enhanced depth image features.

[0085] S613: Concatenate the second enhanced radiomics features and the second enhanced depth image features to obtain preliminary hybrid features.

[0086] As an optional implementation, in step S7, it specifically includes:

[0087] S71: Input the preliminary hybrid image features, the gene features, and the clinical features into the heterogeneous graph and combine them with the graph neural network.

[0088] S72: In the heterogeneous graph, use multi-head multi-order graph attention fusion to fuse the preliminary hybrid image features, the gene features, and the clinical features to obtain fused features.

[0089] S73: React the fused features back to the preliminary hybrid image features, the gene features, and the clinical features to obtain original hybrid features.

[0090] Specifically, the method for local and global attention feature fusion based on the heterogeneous graph MHMOGAT is as follows:

[0091] As Figure 5 shown, the local feature fusion process is based on a cross-modal attention cascaded self-attention mechanism. First, perform a linear transformation on the radiomics features X P with a dimension of (1, 236) to generate a first query vector Q p with a dimension of (1, 32), a first key vector K p , and a first value vector V p, for the depth image feature X with dimension (1, 236) g perform a linear transformation to generate a second query vector Q with dimension (1, 32) g , a second key vector K g , and a second value vector V g , calculate the matrix product of the queries and keys of the two modalities and scale according to the dimension of the keys, and use the Softmax function to calculate the first attention score A p and the second attention score A g , these attention scores are weighted and aggregated with the first value vector V p , the second value vector V g , and after being linearly transformed to restore to the original feature dimension (1, 236), they are element-wise summed with the original values (radiomics feature X P and depth image feature X g ) respectively, to obtain the first enhanced radiomics feature X p ′ and the first enhanced depth image feature X g ′, realizing cross-modal feature selection. Then, perform a linear transformation on the features X p ′ and X g ′, calculate the matrix product of the queries and keys of the same modality and scale according to the dimension of the keys, and use the Softmax function to calculate the new first attention score A′ p and the new second attention score A′ g , these attention scores are weighted and aggregated with the new first value vector V′ p , the new second value vector V′ g , and are element-wise summed with the original values (the first enhanced radiomics feature X p ′ and the first enhanced depth image feature X g ′) respectively to obtain the second enhanced radiomics feature h p and the second enhanced depth image feature h g , and finally perform feature concatenation to obtain a CT image feature with dimension (1, 472), and through linear transformation to obtain a feature h with dimension (1, 32) 2 .

[0092] As Figure 6 shown, the global feature fusion process is based on a heterogeneous graph, taking genes, clinical, and CT features with dimension (1, 32) and the fusion features obtained by direct concatenation and dimensionality reduction as nodes h 1 , h 2 , h 3 , h 4 , the edges adopt a fully connected manner, and combine multi-head multi-order GAT to obtain the features of multi-hop neighbor nodes through a multi-order adjacency matrix, allowing each attention head to automatically calculate the neighbor node h using a trainable weight vector a and a trainable weight matrix Wj with the central node h i the importance score Enhance the expressive power of the attention calculation through the non-linear activation function LeakyReLU, and use Softmax normalization to obtain the n-th order attention weight of the i-th node in the k-th attention head to the neighbor j Subsequently, the features of each neighbor node are weighted and summed according to the attention weights, and after multi-head concatenation, the fusion result of the central node at a specific layer is obtained Then generate the final fused feature h of the central node through an average (Avg) operation i Finally, output the adaptively weighted aggregated node h 4 feature y 0 Compared with traditional methods such as feature concatenation and weighted fusion, this method can better capture the complex dependencies between different modalities, learn global feature representations, and improve the performance of the model in multi-modal learning tasks

[0093] Among them, the expression of the preliminary mixed feature is:

[0094]

[0095] Among them, X′ g is the first enhanced depth image feature; Q p is the first query vector; K g is the second key vector; V g is the second value vector; X g is the depth image feature; d k is the dimension size of the key vector, used to scale the attention; X p ′ is the first enhanced radiomics feature; Q g is the second query vector; K p is the first key vector; V p is the first value vector; X P is the radiomics feature; h p is the second enhanced radiomics feature; Q′ p is the new first query vector; K′ p is the new first key vector; V′ p is the new first value vector; h g is the second enhanced depth image feature; Q′ p is the new second query vector; K′ g is the new second key vector; V′ g is the new second value vector; h 2 is the preliminary mixed feature

[0096] The expression of the original mixed feature is:

[0097]

[0098] Among them, is the n-th order attention weight of the i-th node in the k-th attention head to neighbor j; a is the attention weight vector; W k is a trainable linear transformation matrix; h i is the feature vector of node i; h j is the feature vector of node j; h m is the feature vector of node m; k is the number of heads of multi-head attention; n is the order; is the n-th order neighborhood of the i-th node; LeakyReLU is an activation function used to introduce non-linearity to enhance the expressive power of attention calculation; σ is the ReLU activation function; ∥ is the feature concatenation of the central node and neighbor nodes; is the new node representation formed by the cross-order aggregation of k attention heads in the n-th channel; y 0 is the original mixed feature.

[0099] As an optional implementation manner, in step S8, it specifically includes:

[0100] S81: Add normally distributed noise that changes with time steps to the original mixed feature to obtain Gaussian distributed noise that approaches isotropy.

[0101] S82: Restore the Gaussian distributed noise to the original mixed feature data distribution.

[0102] S83: Gradually remove the noise from the original mixed feature data distribution to obtain the feature after removing redundancy.

[0103] S84: Based on the feature after removing redundancy, obtain the brain metastasis classification judgment result and survival period prediction risk.

[0104] Specifically, the conditional guided diffusion to eliminate data redundancy method CFD is as follows:

[0105] As Figure 7 shown, conditional guided diffusion (CFD) is based on Markov diffusion. In this application, according to the non-small cell lung cancer tumor stage, first add normally distributed noise ε that changes with time step t to the original mixed feature y 0 to make it approach an isotropic Gaussian distribution y t (y t is the feature obtained after forward diffusion). Then restore from the Gaussian noise y t to the original feature data distribution ( is the feature after removing redundancy), and gradually remove the noise ε θ, redundancy removal is achieved. Finally, the classification task branch passes through an FC linear layer, and the feature dimension is reduced from (1, 32) to y with a dimension of (1, 2). The dimension is reduced to y with a dimension of (1, 2). P , and after passing through softmax, the brain metastasis classification result Y is output. The prediction task branch passes through an FC linear layer, and the feature dimension is reduced from (1, 32) to y with a dimension of (1, 8). The dimension is reduced to y with a dimension of (1, 8). C , and after passing through Cox regression, the death risk h(t|y C ) is output, and combined with the nomograph, the overall survival OS of non-small cell lung cancer brain metastasis is predicted.

[0106] Among them, the expression of the feature after removing redundancy is:

[0107]

[0108] Among them, q(y t |y 0 , f φ (x)) is the forward diffusion probability distribution function; is a Gaussian distribution; y t is the feature obtained after forward diffusion; is the mean of the distribution; is the covariance matrix; I is the identity matrix; α t is the retention ratio of the original data at each time step; β t is the fixed noise rate at time step t, {β t} t=1...T ∈(0, 1); is the cumulative noise factor; α s is the s-th noise factor; p θ (y 0:T-1 |y t , f φ (x)) is the reverse diffusion probability distribution function; y t-1 is the feature obtained after forward diffusion from time step t to t - 1; f φ (x) is the conditional guiding feature; is the feature obtained after reverse diffusion from time step t to t - 1; ∈ θ are the parameters of the denoising model UNet. The denoising model UNet is used to learn the noise distribution, and through denoising guided by the tumor TNM staging features, the final prediction is generated. t is the time step; is the feature after removing redundancy.

[0109] The expression of the lung cancer brain metastasis classification judgment result and the survival period prediction risk is:

[0110] h(t|yC ) = h 0 (t) exp(β T y C ) (14);

[0111]

[0112] where h(t|y C ) is the instantaneous hazard rate of the death event occurring at time t under the prediction condition y C , that is, the survival prediction risk; h 0 (t) is the baseline hazard, which is used to provide a survival risk reference independent of the feature variables; β is the regression coefficient; y C is the prediction condition; P(y = k|y P ) is the probability that the sample belongs to category k under the classification condition y P ; is the k-th classification condition; Y is the classification judgment result of lung cancer brain metastasis.

[0113] In summary, this application has the following advantages:

[0114] (1) The multi-modal multi-task learning model SGAFN is proposed. By integrating non-small cell lung cancer CT image data, gene data, and clinical data, more comprehensive and effective tumor information is obtained. The lesion segmentation result of the image data is used as the input for brain metastasis classification and survival prediction for knowledge sharing, enhancing the multi-task correlation, thereby jointly optimizing each task and enhancing the model performance.

[0115] Compared with the method that regards image data lesion segmentation, prognostic stratification, and survival prediction as independent tasks and uses single-modal data for survival prediction, the multi-modal multi-task learning model SGAFN extracts the data features of non-small cell lung cancer in different modalities, fully captures the multi-dimensional information of cancer, and conducts knowledge sharing among the image data lesion segmentation, prognostic stratification, and survival prediction tasks, enhancing the overall prognostic prediction effect.

[0116] (2) A feature extraction method FMCA based on 3D Swin-DenseUnetR for image data lesion segmentation and frequency domain-statistical moment fusion is proposed, effectively focusing on the lesion area of non-small cell lung cancer image data, improving the segmentation fineness, and being able to obtain more comprehensive image features.

[0117] Compared with the method of using 3D Segnet lesion segmentation and FC layer for feature extraction, the FMCA (Feature extraction method of lesion segmentation and frequency domain-statistical moment fusion) of 3DSwin-DenseUnetR can enhance the perception and segmentation ability of non-small cell lung cancer image data lesions through multi-scale feature fusion and cross-window information interaction, and can provide a more comprehensive expression of lesion features in image data by fusing statistical moment features and frequency domain features.

[0118] (3) Proposed a local and global attention feature fusion method MHMOGAT based on heterogeneous graph, which can better fuse the features of different modalities of non-small cell lung cancer, thus improving the accuracy of survival prediction.

[0119] Compared with the methods of directly concatenating and weighting features, the local and global attention feature fusion method MHMOGAT makes full use of the correlation between different modalities, can capture the relationship between features within each modality, and can identify which features are more important for the current task, so as to emphasize or suppress them.

[0120] (4) Proposed a method CFD (Condition-guided diffusion to eliminate data redundancy) based on tumor TNM staging conditions to eliminate multi-modal data redundancy, avoid overfitting during training, and enhance the generalization ability of the model.

[0121] Compared with directly predicting survival by combining mixed features through MLP, condition-guided diffusion adds noise to multi-modal data features with the tumor staging of non-small cell lung cancer as a condition in the forward process, and denoises under the guidance of tumor staging conditions in the reverse process, which can automatically identify and eliminate redundant information and reduce overfitting.

[0122] The present application also provides an application scenario, which applies the above multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion. Specifically: The multi-task lung cancer brain metastasis survival prediction method based on multi-modal data fusion provided in this embodiment can be applied in a medical image processing scenario. The medical image processing scenario includes: a data acquisition link, a segmentation link, a feature extraction link, a local cross-modal attention fusion link, a global cross-modal attention fusion link, and a prediction link; First, obtain the non-small cell lung cancer CT image data, gene data, and clinical data of the patient; Secondly, perform lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result; Thirdly, perform image data feature extraction based on the lesion segmentation result to obtain radiomics features and depth image features; The radiomics features are features extracted based on the pyradiomics package; The depth image features are features extracted by combining the SwinTF encoder and the DenseNet encoder; Perform feature extraction on the gene data to obtain gene features; The gene features are features extracted based on the SGCN encoder; Perform missing value deletion and numerical preprocessing on the clinical data first, and then perform feature extraction to obtain clinical features; The clinical features are features extracted based on the MLP encoder; Furthermore, perform local cross-modal attention fusion on the radiomics features and the depth image features to obtain a preliminary mixed feature; Then, perform global cross-modal attention fusion on the preliminary mixed image feature, the gene feature, and the clinical feature to obtain an original mixed feature; Finally, perform conditional guided diffusion based on the original mixed feature to obtain the lung cancer brain metastasis classification judgment result and the survival prediction risk.

[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0124] Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion, characterized in that: The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion includes: Obtain CT imaging data, genetic data, and clinical data of non-small cell lung cancer; Performing lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result; Based on the lesion segmentation result, image data features are extracted to obtain radiomics features and deep image features; the radiomics features are features extracted based on the pyradiomics package; the deep image features are features extracted based on the combination of the SwinTF encoder and the DenseNet encoder; Performing feature extraction on the gene data to obtain gene features; the gene features are features extracted based on the SGCN encoder; The clinical data is first subjected to missing value deletion and numerical preprocessing, and then feature extraction is performed to obtain clinical features; the clinical features are features extracted based on the MLP encoder; Performing local cross-modal attention fusion on the radiomics feature and the deep image feature to obtain a preliminary mixed feature; Performing global cross-modal attention fusion on the preliminary mixed image features, the gene features and the clinical features to obtain original mixed features; Conditional guided diffusion is performed based on the original mixed features to obtain the classification judgment results of lung cancer brain metastasis and the predicted risk of survival period.

2. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: Performing lesion segmentation on the non-small cell lung cancer CT image data to obtain a lesion segmentation result specifically includes: The SwinTF encoder is used to extract image features at different positions of the non-small cell lung cancer CT image data through sliding window SW-MSA to obtain several depth feature maps; Converting a plurality of the depth feature maps to obtain a plurality of channel feature maps; EMA is used to extract the intra-channel statistical moment information of a plurality of the channel feature maps; Using FFT to extract frequency domain information of a plurality of channel feature graphs; After the intra-channel statistical moment information and the frequency domain information are concatenated and dimensionally reduced, a lesion segmentation result is obtained.

3. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: The radiomics features are locally fused with the deep image features through cross-modal attention to obtain preliminary mixed features, including: performing a linear transformation on the radiomics feature to generate a first query vector, a first key vector, and a first value vector; Performing a linear transformation on the depth image features to generate a second query vector, a second key vector, and a second value vector; Calculate the matrix product of the first query vector and the first key vector using a Softmax function to obtain a first attention score; Calculate the matrix product of the second query vector and the second key vector using the Softmax function to obtain a second attention score; Performing weighted aggregation on the first attention score and the first value vector, and performing element-wise summation on the first attention score and the first value vector with the radiomics feature to obtain a first enhanced radiomics feature; Performing weighted aggregation on the second attention score and the second value vector, and performing element-wise summation on the second attention score and the depth image feature to obtain a first enhanced depth image feature; performing a linear transformation on the first enhanced radiomics feature to generate a new first query vector, a new first key vector, and a new first value vector; Performing a linear transformation on the first enhanced depth image feature to generate a new second query vector, a new second key vector, and a new second value vector; Calculate the matrix product of the new first query vector and the new first key vector using the Softmax function to obtain a new first attention score; Calculate the matrix product of the new second query vector and the new second key vector using the Softmax function to obtain a new second attention score; Performing weighted aggregation on the new first attention score and the new first value vector, and performing element-wise summation on the new first attention score and the first enhanced radiomics feature to obtain a second enhanced radiomics feature; Performing weighted aggregation on the new second attention score and the new second value vector, and performing element-wise summation on the new second attention score and the first enhanced depth image feature to obtain a second enhanced depth image feature; The second enhanced radiomics feature is concatenated with the second enhanced depth image feature to obtain a preliminary mixed feature.

4. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: The preliminary mixed image features, the gene features and the clinical features are fused with global cross-modal attention to obtain original mixed features, which specifically include: Inputting the preliminary mixed image features, the gene features and the clinical features into a heterogeneous graph and combining them with a graph neural network; In the heterogeneous graph, multi-head multi-order graph attention fusion is used to fuse the preliminary mixed image features, the gene features and the clinical features to obtain fused features; The fusion feature is reacted to the preliminary mixed image feature, the gene feature and the clinical feature to obtain the original mixed feature.

5. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: Conditional guided diffusion is performed based on the original mixed features to obtain the classification judgment results of lung cancer brain metastasis and the predicted survival risk, specifically including: Adding a normally distributed noise that varies with time steps to the original mixed feature to obtain a Gaussian distributed noise that approaches isotropic; Restoring the Gaussian distribution noise to the original mixed feature data distribution; gradually removing noise from the original mixed feature data distribution to obtain features after redundancy is removed; Based on the redundant features, the classification judgment results of lung cancer brain metastasis and the predicted survival risk are obtained.

6. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: The expression of the depth image feature is: y l =concat(EMA(X l ),FFT(X l )),l=0,1…,7; X g =concat(and 0 ,and 1 ,and 2 ,and 3 ,and 4 ,and 5 ,and 6 ,and 7 ); Among them, EMA(X l ) is the statistical feature composed of the mean, variance and skewness of the characteristic graph; X l is the feature map output at different channel depths; α k is the kth weight parameter, which is used to assign different weights to different moment terms in the extended moment aggregation formula; E(X l ) is the mean of the channel feature map; M k (X l ) is the variance and skewness; FFT(X l ) is the frequency domain feature composed of different frequency components of the feature map; is the weight parameter in the frequency domain aggregation formula; N is the number of pixels; X l (k) is the kth frequency component of the channel feature map; is the rotation factor, representing the periodic oscillation of the complex exponential; X l (n) is the pixel value of the feature map; y l X is the image feature formed by combining the statistical features of the feature maps of different depth channels with the frequency domain features; g is the deep image feature.

7. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: The expression of the preliminary mixed feature is: Among them, X′ g is the first enhanced depth image feature; Q p is the first query vector; K g is the second bond vector; V g is the second value vector; X g is the deep image feature; d k is the dimension size of the key vector, used to scale the attention; X p ' is the first enhanced radiomics feature; Q g is the second query vector; K p is the first bond vector; V p is the first value vector; X P is the radiomics feature; h p is the second enhanced radiomics feature; Q′ p is the new first query vector; K′ p is the new first bond vector; V p ′ is the new first value vector; h g is the second enhanced depth image feature; Q′ p is the new second query vector; K′ g is the new second bond vector; V g ′ is the new second value vector; h2 is the preliminary mixed feature.

8. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1 is characterized in that: The expression of the original mixed feature is: in, is the nth-order attention weight of the ith node in the kth attention head to its neighbor j; a is the attention weight vector; W k is a trainable linear transformation matrix; h i is the feature vector of node i; h j is the feature vector of node j; h m is the feature vector of node m; k is the number of heads of multi-head attention; n is the order; is the nth-order neighborhood of the ith node; LeakyReLU is the activation function, which is used to introduce nonlinearity to enhance the expression ability of attention calculation; σ is the ReLU activation function; ∥ is the feature concatenation of the central node and the neighboring nodes; is a new node representation formed by the cross-sequential aggregation of k attention heads in the nth channel; y0 is the original mixed feature.

9. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 5, characterized in that: The expression of the feature after removing redundancy is: Among them, q(y t |y0,f φ (x)) is the forward diffusion probability distribution function; is Gaussian distribution; t is the feature obtained after forward diffusion; is the mean of the distribution; is the covariance matrix; I is the identity matrix; α t is the retention ratio of the original data for each time step; β t is the fixed noise rate at time step t, {β t } t=1...T ∈(0,1); is the cumulative noise factor; α s is the sth noise factor; p θ (y 0:T-1 |y t ,f φ (x)) is the reverse diffusion probability distribution function; y t-1 is the feature obtained by forward diffusion at time step t-1; f φ (x) is the conditional guidance feature; is the feature obtained after back diffusion at time step t-1; ∈ θ is the parameter of the denoising model UNet; t is the time step; is the feature after removing redundant features.

10. The multi-task lung cancer brain metastasis survival prediction method based on multimodal data fusion according to claim 1, characterized in that: The expression of the classification judgment result of lung cancer brain metastasis and the predicted risk of survival period is: h(t|y C )=h0(t)exp(β T y C ); Among them, h(t|y C ) is the prediction condition y C In the case of , the instantaneous risk rate of death at time t is the predicted survival risk; h0(t) is the baseline risk, which is used to provide a reference for survival risk that is independent of characteristic variables; β is the regression coefficient; y C is the prediction condition; P(y=k|y P ) is the sample under classification condition y P The probability of belonging to category k in the case of is the kth classification condition; Y is the classification judgment result of lung cancer brain metastasis.

Citation Information

Cited By

  • Lightweight multi-feature fusion thyroid nodule classification method

    CN120318613A

  • A lightweight multi-feature fusion method for thyroid nodule classification

    CN120318613B

  • Cornea image death time prediction method based on multi-modal Transform-diffusion model

    CN120823623A

  • Gastric cancer immunotherapy curative effect prediction method based on multi-mode fusion

    CN121237451A

  • Multi-modal medical image segmentation method based on frequency domain perception fusion

    CN121685970A