A prognostic prediction method based on co-attention modality fusion
Through the prognosis prediction method of co-attention modal pre-fusion, combined with graph neural network and co-attention mechanism, the interaction characteristics of images and pathological images are extracted, which solves the subjectivity and inconsistency problems of tumor prognosis assessment in existing technologies and achieves more accurate tumor prognosis prediction.
Patent Information
- Application Number
- CN202411734974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing tumor prognosis assessment methods rely on manual analysis, which is subjective and inconsistent, and fail to effectively capture the complex predictive information between imaging and pathological images. Existing multimodal models fail to deeply explore the interaction between macroscopic imaging and microscopic pathology.
A prognosis prediction method based on co-attention modality pre-fusion is adopted. Graph neural networks and pre-trained convolutional neural networks are used to extract cell interaction features and morphological features in pathological images. The co-attention mechanism is combined to capture the interaction between imaging omics and pathology omics, and a multi-instance learning module is used for prognosis prediction.
It improves the accuracy and consistency of tumor prognosis prediction, is applicable to different types of cancer, and significantly improves the prognosis prediction accuracy of traditional grading and staging indicators, single-modality models, and multimodality post-fusion methods.
Smart Images

Figure CN119904409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor image analysis, and in particular to a prognosis prediction method based on co-attention modality front fusion. Background Art
[0002] Prognosis plays an important role in cancer research. Currently, the prognostic assessment methods widely used in clinical practice mainly rely on tumor staging systems, which assess the severity of tumors by quantifying the size of the tumor, the degree of invasion of adjacent tissues, lymph node metastasis, and distant metastasis. However, this assessment method relies on manual analysis of radiological and pathological images. This process may not only introduce subjectivity, but also has significant inconsistencies in judgment between different assessors. In addition, traditional manually defined features are often relatively simple and cannot effectively capture the complex predictive information embedded in the image. Therefore, there is an urgent need to develop more objective and sophisticated assessment methods to improve the accuracy and consistency of tumor prognostic assessment.
[0003] Deep learning models can automatically extract comprehensive prognostic information from medical images, significantly enhancing predictive capabilities. In recent years, a number of studies based on pathological sections and radiological scans have applied deep learning to tumor prognosis prediction, demonstrating the potential of deep learning in the field of survival prediction. On this basis, researchers have gradually recognized the advantages of integrating multiple data modes and actively developed multimodal prognostic models for different types of tumors. Currently, most existing multimodal models adopt a post-fusion strategy, that is, each modality is trained independently, and finally the decision results of each modality are integrated to generate the final prediction. This decision-level fusion strategy shows high flexibility in dealing with scenarios with incomplete modalities, but ignores the possible synergy between different modalities.
[0004] To more effectively capture the interactions between modalities, a growing number of studies are leaning towards feature-level fusion strategies. In this pre-feature fusion approach, heterogeneous multimodal features can be integrated at an early stage to form a more informative cross-modal representation. However, early fusion places high demands on data integrity, meaning that each sample must contain all complete modalities. Some studies, based on the Cancer Genome Atlas (TCGA) project, have adopted early fusion strategies to integrate genomic, transcriptomic, proteomic, and pathological data to predict patient survival. However, research on the integration of other medical data modalities is still insufficient and urgently needs further exploration.
[0005] Histopathology and radiology images play a crucial role in clinical decision-making regarding cancer treatment. As the most commonly used imaging techniques in clinical practice, these two imaging modalities provide complementary information. Histopathology characterizes the cellular and tissue morphology within a tumor at a microscopic level, ex vivo; while radiology provides macroscopic insights into tumor size and texture through non-invasive in vivo imaging.
[0006] Currently, several studies have shown that there is a significant correlation between radiographic patterns and specific pathological features, which are in turn closely related to tumor prognosis. For example, in hepatocellular carcinoma, the irregular edge enhancement pattern observed on CT scans is directly associated with the histopathological features that promote its aggressive tumor behavior. In breast cancer studies, dense mammographic images have been shown to be associated with higher histological grades. However, existing technologies have failed to deeply explore the interaction between macroscopic imaging and microscopic pathology, nor have they conducted tumor prognosis analysis, and the prediction accuracy needs to be improved. Summary of the Invention
[0007] The purpose of the present invention is to provide a prognosis prediction method based on co-attention modal pre-fusion. This method extracts cell interaction features and morphological features in pathological images through graph neural networks and pre-trained convolutional neural networks to construct pathological genomics representation, and uses artificial predefined feature groups and pre-trained radiomic models to extract shape texture and morphological features in image images to form radiomics representation. The co-attention mechanism is used to capture the interaction between pathological genomics and radiomics, and a cross-modal pathological genomics representation based on image co-attention is obtained. Finally, two multi-instance learning modules are input to obtain prognosis prediction results, thereby improving the accuracy of prognosis prediction in different scenarios.
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] A prognosis prediction method based on co-attention modality pre-fusion, comprising the following steps:
[0010] Step S1, image acquisition and preprocessing: acquiring postoperative pathological images and three-dimensional preoperative image sequences and preprocessing them separately, wherein the postoperative pathological images are 40x magnified hematoxylin-eosin (HE) stained postoperative pathological images, performing segmentation, filtering, and staining normalization preprocessing on the postoperative pathological images, and performing tumor region delineation preprocessing on the three-dimensional preoperative image sequences;
[0011] Step S2: Extracting pathomic representations: Applying graph neural networks and pre-trained convolutional neural networks to each pathological patch in the postoperative pathological image, the cell interaction features and morphological features of all pathological image patches are extracted to form a pathomic representation. The pathological features of a single pathological patch are composed of the splicing of cell interaction features and morphological features. The pathomic representation of a single patient is represented as a bag of all pathological patch instances in the pathological tissue.
[0012] Step S3, extracting radiomics representation: using the preoperative image sequence and the delineated tumor area, shape, texture, and morphological features are extracted from the image sequence based on a manually predefined feature set and a pre-trained radiomics model to form a radiomics representation;
[0013] Step S4: Capturing the interaction between pathology and radiomics: Based on the extracted pathology and radiomics representations, a co-attention mechanism module is used to capture the association and interaction between the two, thereby obtaining a cross-modal pathology representation based on image co-attention.
[0014] Step S5, prognosis prediction under multi-instance learning: the imaging omics representation and the cross-modal pathology omics representation based on image co-attention are input into two multi-instance learning modules respectively, the omics information is aggregated, and the prognosis prediction result is obtained.
[0015] The preprocessing of the postoperative pathological images specifically includes:
[0016] Obtaining pathological blocks: cutting the postoperative pathological images into pathological blocks of preset size;
[0017] Image patch filtering: Use the OTSU method to extract the foreground tissue area in the pathological patch, calculate the ratio of the foreground tissue area to the corresponding image patch, and filter the image patches whose ratio is lower than a preset value;
[0018] Brightness and color normalization: For each pathological block obtained after filtering, the brightness of the scanning field was normalized using the LuminosityStandardizer function in the staintools package, and the color was normalized using the macenko method.
[0019] The preoperative imaging sequences were obtained via computed tomography (CT) or magnetic resonance imaging (MRI). The imaging sequence type was selected based on clinical experience to suit different types of cancer. Brain tumor imaging used gadolinium-enhanced T1-weighted (T1-Gd) sequences and T2 fluid-attenuated inversion-recovery (T2-FLAIR) sequences. Kidney tumors used enhanced arterial phase CT abdominal and pelvic scan sequences. Lung tumors used low-dose chest CT scans.
[0020] The preprocessing of the preoperative image sequence specifically includes:
[0021] Format conversion: convert DICOM format image sequences into NII format;
[0022] Resampling: resampling the converted image sequence to a set resolution, which is determined by the image sequence type;
[0023] Image registration: Advanced Normalization Tools (ANTs) were used to register T2 fluid-attenuated inversion recovery sequences to gadolinium-enhanced T1-weighted sequences; CT images did not require registration.
[0024] Tumor area delineation: Outline the tumor contour for each 2D scan section in the 3D sequence;
[0025] Region of interest extraction: Based on the outlined tumor area, a three-dimensional region of interest containing the tumor is cropped from the full-resolution image.
[0026] Specifically, the tumor area delineation is as follows: for kidney and lung tumors, the deep segmentation learning model nnU-Net pre-trained on a public dataset is used for automatic contour delineation, and manual image reading is performed to correct the tumor boundaries segmented by the model; for brain tumors, manual delineation is performed directly due to the lack of corresponding pre-trained models.
[0027] The extraction of cell interaction features in step S2 relies on graph neural networks, and the extraction process specifically includes:
[0028] Cell nucleus classification and segmentation: The Hover-UNet model pre-trained on the PanNuke dataset is used to segment and classify cell nuclei. The Hover-UNet model obtains a mask for each individual cell nucleus through cell nucleus segmentation and classifies the cell nuclei into tumor cells, inflammatory cells, junctional cells, dead cells, and normal epithelial cells.
[0029] Use HistomicsTK in Python to extract pathological features of single cell nuclei;
[0030] Nuclear graph network construction: Pathological patches were modeled using a nuclear graph network. The nodes of the nuclear graph network represent cell nuclei. Each node features the pathological omics features of a single cell nucleus extracted using HistomicsTK and the probability of the cell nucleus belonging to a certain category. The connections between nodes in the graph network use the eight-nearest neighbor method, and the attributes of the edges include the connected cell type, the distance between cells, and the degree of parallelism between cells.
[0031] The constructed cell nucleus graph network is trained to extract intermediate feature representations as cell interaction features. The graph network includes four edge-conditioned convolution layers under cell spatial interactions, followed by a subgroup average pooling layer to aggregate information from all nodes to obtain the required intermediate feature representations.
[0032] The morphological features of the pathological image in step S2 are the output results generated by the last layer of the ResNet34 model, and the ResNet34 model is pre-trained on ImageNet.
[0033] In step S3, feature extraction is performed using Pyradiomics based on a manually predefined shape and texture feature group. The shape and texture feature group includes five dimensions: shape, first order, texture, filtering, and wavelet. The features of each dimension are mapped to the same dimension through a linear layer. Morphological features are extracted using a visual encoder in the image-based model RadFM pre-trained on a large-scale imaging dataset. The radiomics of a single patient is represented as an aggregated package of the shape and texture feature group and the morphological features.
[0034] The co-attention mechanism module in step S4 receives the pathogenomics representation and the radiomics representation extracted in steps S2 and S3, and based on the multi-head attention mechanism, uses the radiomics representation as a query and the pathogenomics representation as a key-value pair to obtain a cross-modal pathogenomics representation based on image co-attention.
[0035] The multi-instance learning module in step S5 consists of a set-based Transformer, a global attention pooling layer, and a linear layer. The multi-instance learning module is used to aggregate omics features for the imaging omics representation and the cross-modal pathology omics representation based on image co-attention to obtain two feature vectors, and the feature vectors are concatenated and connected to the linear layer to output the prognosis prediction result.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) Compared with existing pathology and imaging image representation methods that are mostly limited to morphological features, the present invention considers both cell interactions and morphological features when constructing pathology omics representation, and considers both shape texture and morphological features when constructing imaging omics representation. The features considered are more comprehensive and can improve the accuracy of the final prediction results.
[0038] (2) Compared with existing prognostic models that are mostly designed for specific cancer types, the present invention uses a co-attention mechanism multi-instance learning framework to achieve pan-cancer prognosis prediction.
[0039] (3) Compared with the existing single-modality or multi-modality post-fusion methods, the present invention adopts a pre-fusion method and uses a co-attention mechanism to capture the interaction between pathological genomics and imaging genomics, which significantly improves the accuracy of traditional grading and staging indicators, single-modality models and multi-modality post-fusion methods in prognosis prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flow chart of the method of the present invention;
[0041] Figure 2 The construction and model structure of the cell nucleus map network in the embodiment. Part A shows the construction process of the cell nucleus map network in the pathological patch. This process can be subdivided into three steps: pathological patch cutting, cell nucleus segmentation and map construction; Part B is the graph convolution model structure under cell space interaction;
[0042] Figure 3 The detailed construction process of the pathology omics package in the embodiment;
[0043] Figure 4 Schematic diagram of the structure of the comparison model in the embodiment, Part A is the unimodal comparison model in the embodiment that only uses the image as input, Part B is the unimodal comparison model in the embodiment that only uses the pathological image as input, and the model structure in Part C is the same as the multi-instance learning framework of the co-attention mechanism described in the embodiment except that the cell interaction characteristics are not considered. DETAILED DESCRIPTION
[0044] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0045] This embodiment provides a prognosis prediction method based on co-attention modality pre-fusion, which aims to provide a co-attention mechanism multi-instance learning framework to capture the interaction between pathology and imaging omics and build a pan-cancer prognosis prediction model. Figure 1 As shown, the following steps are included:
[0046] Step S1, image acquisition and preprocessing: Obtain postoperative pathological images and three-dimensional preoperative image sequences and preprocess them separately, wherein the postoperative pathological images are 40-fold magnified hematoxylin-eosin (HE) stained postoperative pathological images. The postoperative pathological images are preprocessed by slicing, filtering, and staining normalization, and the three-dimensional preoperative image sequences are preprocessed by tumor area delineation.
[0047] This embodiment uses three cancer multimodal datasets (including kidney tumors, lung tumors, and brain tumors) as examples for detailed description.
[0048] S11. Collect postoperative pathological images and perform preprocessing, wherein the preprocessing specifically includes:
[0049] Obtaining pathological blocks: Cut the postoperative pathological images into 768×768 pathological blocks;
[0050] Image patch filtering: Use the OTSU method to extract the foreground tissue area in the pathology patch, calculate the ratio of the foreground tissue area to the corresponding image patch, filter out image patches with a ratio below 0.7, and only retain image patches with a foreground tissue area ratio greater than or equal to 0.7;
[0051] Brightness and color normalization: For each pathological block obtained after filtering, the brightness of the scanning field was normalized using the LuminosityStandardizer function in the staintools2.1.2 package in sPython, and the color was normalized using the macenko method.
[0052] S12. Preoperative Image Sequence Acquisition and Preprocessing
[0053] Preoperative imaging sequences were obtained using computed tomography (CT) or magnetic resonance imaging (MRI). The type of imaging sequence was selected based on clinical experience to suit different types of cancer. For brain tumors, gadolinium-enhanced T1-weighted (T1-Gd) sequences and T2 fluid-attenuated inversion-recovery (T2-FLAIR) sequences were used. For renal tumors, enhanced arterial phase CT abdominal and pelvic scan sequences were used. For lung tumors, low-dose chest CT was used.
[0054] Pre-operative image sequence preprocessing specifically includes:
[0055] Format conversion: convert DICOM format image sequences into NII format;
[0056] Resampling: resample the converted image sequence to a set resolution. For example, the CT image is resampled to a resolution of 0.8 mm × 0.8 mm × 5 mm, and the MRI image is resampled to a resolution of 1 mm × 1 mm × 1 mm.
[0057] Image registration: Advanced Normalization Tools (ANTs) were used to register T2 fluid-attenuated inversion recovery sequences to gadolinium-enhanced T1-weighted sequences; CT images did not require registration.
[0058] Tumor region delineation: Tumor contours are outlined for each 2D scan section in the 3D sequence. For kidney and lung tumors, nnU-Net, a deep segmentation learning model pre-trained on public datasets, is used for automatic delineation, and manual image reading is performed to correct the tumor boundaries segmented by the model. For brain tumors, manual delineation is performed directly due to the lack of corresponding pre-trained models.
[0059] Region of interest extraction: Based on the outlined tumor area, a three-dimensional region of interest containing the tumor is cropped from the full-resolution image. The size of the region is padded with zeros to an integer multiple of 32×32×12 to meet the input requirements of the RadFM model.
[0060] Step S2, extracting the pathomic representation: Applying a graph neural network and a pre-trained convolutional neural network to each pathological patch in the postoperative pathological image, respectively extracting the cell interaction features and morphological features in all pathological image patches to form a pathomic representation; the pathological features of a single pathological patch are composed of the splicing of cell interaction features and morphological features. The pathomic representation of a single patient is represented as a bag of all pathological patch instances in the pathological tissue.
[0061] The extraction of cell interaction features relies on graph neural networks. The extraction process specifically includes:
[0062] Cell nucleus classification and segmentation: The Hover-UNet model pre-trained on the PanNuke dataset is used to segment and classify cell nuclei. The Hover-UNet model obtains a mask for each individual cell nucleus through cell nucleus segmentation and classifies the cell nuclei into tumor cells, inflammatory cells, junctional cells, dead cells, and normal epithelial cells.
[0063] Use HistomicsTK in Python to extract pathological features of single cell nuclei;
[0064] Nuclear graph network construction: The pathological patches were modeled using a nuclear graph network. The nodes of the nuclear graph network represent the cell nuclei. The characteristics of each node include the pathological omics features of a single cell nucleus extracted by HistomicsTK and the probability of the category to which the cell nucleus belongs. The connections between the nodes in the graph network use the eight-nearest neighbor method. The attributes of the edges include the connected cell type, the distance between cells, and the degree of parallelism between cells, such as Figure 2 As shown in part A.
[0065] The constructed cell nucleus graph network is trained to extract intermediate feature representations as cell interaction features. The graph network consists of four cell spatial interaction-conditioned graph convolutional (CSIGC) layers, followed by a subgroup average pooling layer to aggregate information from all nodes, thereby obtaining the required 64-dimensional intermediate feature representation. The specific form of the cell spatial interaction-conditioned graph convolutional layer is the edge-conditioned convolution layer.
[0066] During the pre-training stage of the cell nucleus map network, a linear layer is added to the intermediate features to predict the discrete survival risk of each pathological patch. For the prediction results of all pathological patches of a single patient, in order to ensure that the intermediate feature representation extracted by the CSIGC layer can effectively reflect the prognostic information, the present invention uses the maximum pooling method to aggregate the risk prediction results of the pathological patches, thereby obtaining patient-level survival prediction results. The model is implemented using PyTorch1.12.1 (cu113) and PyG 2.3.1 and trained on 4×NVIDIA A100 (40G) GPUs. The model optimization loss function is obtained by calculating the negative logarithmic loss of the discretized survival label of each patient. During the training process, the batch size is 1, the learning rate is 5e-4, the weight decay is 1e-4, and the optimizer uses Adam.
[0067] The loss calculation method for discretization of survival labels is as follows: 0 ,x 1 ,…,x n ) The observed survival time is discretized into the interval [a0,a1), [a1,a2), [a2,a3) [a3,a4] according to the 25%, 50% and 75% quantiles, where a0 <a1<a2<a3<a4∈R + Patient x i The survival time falls on [a n-1 ,a n ), its discrete time t i =n. Thus, the hazard function can be defined as:
[0068] h(t|X)=P(T=t|t≥t,X) (1)
[0069] The survival function is:
[0070]
[0071] The negative log loss of discretized survival labels can be expressed as:
[0072]
[0073] where t i ∈{1,2,3,4}.
[0074] The optimal nuclear map network model is selected as the weight with the highest Harrell's consistency index (c-index) on the validation set. Using the selected pre-trained nuclear map network weights, the intermediate feature representation is derived to obtain the cell interaction features, such as Figure 2 As shown in part B in , the feature is 64-dimensional.
[0075] The morphological features of the pathological images are the output results generated by the last layer of the ResNet34 model, which is pre-trained on ImageNet.
[0076] Step S3: Extract radiomics representation: Using the preoperative image sequence and the delineated tumor area, shape, texture, and morphological features are extracted from the image sequence based on the manually predefined feature set and the pre-trained radiomics model to form a radiomics representation.
[0077] The imaging omics of a single patient is represented as an aggregate of shape and texture feature groups and morphological features. The shape and texture feature groups are extracted by Pyradiomics. The required input is the image data and its corresponding tumor area, and the output is five-dimensional imaging omics features, including shape features (14 dimensions), first-order statistical features (18 dimensions), texture features (75 dimensions), filter features (279 dimensions), and wavelet features (744 dimensions). The features of each dimension are mapped to the same dimension through a linear layer. The morphological features are extracted by the visual encoder in the image-based model RadFM pre-trained on a large-scale imaging dataset. The input is a three-dimensional region of interest containing the tumor, and the output is a 768-dimensional feature vector. The imaging omics of a single patient is represented as an aggregate package of shape and texture feature groups and morphological features.
[0078] Step S4: Capturing the interaction between pathology and radiomics: Based on the extracted pathology and radiomics representations, a co-attention mechanism module is used to capture the association and interaction between the two, and a cross-modal pathology representation based on radiomics co-attention is obtained.
[0079] The co-attention mechanism module receives the pathomic representation and radiomics representation extracted in steps S2 and S3, and based on the multi-head attention mechanism, takes the radiomics representation as the query and the pathomic representation as the key-value pair to obtain a cross-modal pathomic representation based on radiomics co-attention.
[0080] First, the pathology and imaging features are reconstructed to form the representation of pathology packages and imaging packages.
[0081] For each pathological image patch, two learnable linear layers are used to map the 512-dimensional pathological image morphological features to 192 dimensions, which are then concatenated with the 64-dimensional cell interaction features to form a 256-dimensional pathological feature. A patient's pathological package consists of the features of all pathological image patches and can be represented as Where n is the number of small patches of the patient's pathological image. Figure 3 As shown in the figure, in addition to the morphological features extracted from ResNet-34, the features of the pathological patches also include intermediate features from the pre-trained nuclear map network. By applying a linear layer, the 512-dimensional ResNet-34 features are mapped to 192 dimensions and concatenated with the 64-dimensional cell-cell interaction features in the pathological patches to obtain 256-dimensional pathological patch features.
[0082] The patient's imaging omics is a set of shape texture features and morphological features. The features of each dimension are mapped to 256 dimensions through two learnable linear layers, and the final imaging package can be represented as
[0083] Secondly, a Transformer-based co-attention mechanism module is used to model the interaction between pathology packages and imaging packages.
[0084] In the Transformer-based co-attention module, the imaging omics package is used as the query and the pathology omics package as the key-value to obtain a cross-modal pathology representation based on imaging co-attention.
[0085] Specifically, it can be written as
[0086]
[0087] in, is the learnable neural network weight. The cross-modal pathomic representation based on image co-attention can be written as
[0088] Step S5, prognosis prediction under multi-instance learning: the imaging omics representation and the cross-modal pathology omics representation based on image co-attention are input into two multi-instance learning modules respectively, the omics information is aggregated, and the prognosis prediction result is obtained.
[0089] The multi-instance learning module consists of a set-based Transformer, a global attention pooling layer, and a linear layer. The multi-instance learning module is used to aggregate omics features for radiomics representation and cross-modal pathology representation based on image co-attention, respectively, to obtain two 256-dimensional feature vectors. The 256-dimensional feature vectors are then concatenated and connected to a linear layer to output the prognosis prediction results.
[0090] In this example, the model was implemented using PyTorch 1.12.1 (cu113) and trained on a 1× NVIDIA A100 (40G) GPU. For model optimization, the discretized negative logarithmic loss (Formula (3)) of each patient's survival label was calculated, with a training batch size of 1, a learning rate of 1e-4, a weight decay of 2e-4, and the Adam optimizer.
[0091] The model was evaluated using the c-index and the time-dependent area under the receiver operating characteristic curve (Td-AUROC). In view of the differences in follow-up time for different cancer types, this embodiment selected 2 years as the milestone time, which is closest to the median survival time of all patients. For the five-fold cross-training validation, the cross-validated c-index and Td-AUROC and the corresponding standard deviation were estimated by the weighted average of the predictions of each validation set. The evaluation indicators were implemented using survcomp 1.40.0, timeROC 0.4, and survival 3.1.12 in R 4.0.2.
[0092] To verify the accuracy of the present invention, the example verification content of this embodiment is as follows:
[0093] 1) Dataset
[0094] In this embodiment, the dataset used is derived from internal and public resources, covering three types of cancer. The data collection for each type can be divided into two parts: the first part is used to develop a cell nuclear map network model to extract cell interaction features in pathological genomics; the second part is used to develop a multi-instance learning framework of a co-attention mechanism to predict survival. The dataset used for the development of the cell nuclear map network model includes H&E-stained pathological diagnostic sections and their corresponding survival information. The dataset used for the co-attention mechanism multi-instance learning framework includes preoperative radiological imaging, postoperative H&E-stained pathological sections and related survival information. There is no sample overlap between these two parts of data. Specifically as follows:
[0095] For renal tumors, the nuclear graph network model was developed using data from 503 patients in the TCGA-KIRC project. To develop a multi-instance learning framework with a shared attention mechanism, a dataset was used containing information on 777 patients with clear cell renal cell carcinoma who underwent curative resection at a single hospital between 2014 and 2020. The survival outcome in this dataset was recurrence after curative surgery, and the imaging data used contrast-enhanced arterial CT scans of the abdomen and pelvis.
[0096] For lung tumors, the nuclear graph network model was developed using 772 patients from TCGA-LUSC and TCGA-LUAD, excluding any patients with available radiographic imaging. The dataset used to develop the co-attention multi-instance learning framework included 278 patients: 17 from TCGA-LUAD, 31 from TCGA-LUSC, and 230 from the National Lung Screening Trial. The survival outcome for this dataset was overall survival. Imaging data used low-dose chest CT scans.
[0097] For brain tumors, a nuclear graph network model was developed based on data from 646 patients without radiographic imaging from the TCGA-GBM and TCGA-LGG projects. The dataset used to develop the shared-attention multi-instance learning framework included 165 patients with radiographic imaging from these two projects. The survival outcome for this dataset was overall survival. Radiographic imaging data used gadolinium-enhanced T1-weighted (T1-Gd) and T2 fluid-attenuated inversion-recovery (T2-FLAIR) sequences.
[0098] For all public datasets, histological images and their corresponding survival information are available from the Genomic Data Commons data portal, and radiomic imaging data are available from the Cancer Imaging Archive.
[0099] Tables 1 and 2 below show detailed information about the datasets used to develop the cell nucleus graph network model and the co-attention mechanism multi-instance learning framework, including the number of patients for each cancer type, the censoring rate, the median survival time, the number of pathology images and pathology patches used, and the total number of imaging images.
[0100] Table 1 Dataset information used for developing the cell nucleus map network model
[0101] Tumor type Number of patients Censorship rate Median survival (years) Number of WSIs Number of small pieces Average number of small blocks Kidney tumor 503 0.664 2.863 508 2,611,366 4,983 Lung tumors 772 0.687 1.459 860 4,290,825 4,477 brain tumors 646 0.529 1.381 1,213 4,385,005 4,575
[0102] Table 2 Dataset information for the development of the co-attention mechanism multi-instance learning framework
[0103]
[0104]
[0105] 2) Cell interaction feature extraction based on cell nucleus graph network
[0106] The nuclear map network was validated using five-fold cross-training, training on renal, lung, and brain tumors. To assess the effectiveness of the nuclear map network, the prognostic indicators c-index and Td-AUROC were used for validation. The results showed that the nuclear map network model demonstrated significant prognostic predictive ability in all three cancer types. Detailed results for the c-index and Td-AUROC indicators are shown in Table 3.
[0107] Table 3 Five-fold cross validation performance of the cell nucleus map network model
[0108]
[0109] 3) Verification of the superiority of the multi-instance learning framework based on the co-attention mechanism
[0110] To validate the advantages of the co-attention mechanism multi-instance learning framework (hereafter referred to as "PaRa-MIL") in terms of multimodal input, multi-level features, and attention-based pre-fusion modal interactions, this study used the following control models to compare the prognostic prediction effects: a hierarchical staging information model, a unimodal model, a model that did not include pathological cell interaction features, and a post-fusion model.
[0111] Tumor staging is the gold standard for clinically assessing cancer severity. For kidney and brain tumors, staging information is extracted. Because staging information is lacking for lung tumors, grade is used for assessment. Tumor grade reflects the aggressiveness of tumor cells, while staging is assessed according to the American Joint Committee on Cancer criteria, and grade relies on histopathological assessment. The distribution of staging or grade information for different cancer types is shown in Table 4.
[0112] Table 4 Stage / grade information of different cancer types
[0113]
[0114]
[0115] Unimodal models include models that use only images as input (Rad-MIL, Figure 4 Part A) and the model that only uses pathological images as input (Path-MIL, Figure 4 Part B). The structure of Rad-MIL is the same as that of the radiological imaging branch of PaRa-MIL, and the structure of Path-MIL is the same as that of the pathological branch of PaRa-MIL. To verify the validity of the pathological cell interaction characteristics, PaRa-MIL-wo (see Figure 4 The experiment was conducted in part C). The overall structure of PaRa-MIL-wo is the same as that of PaRa-MIL, except for the pathology part. Among them, PaRa-MIL-wo does not consider the cell interaction characteristics in pathology, while PaRa-MIL takes this part of the characteristics into account. In addition, in order to verify the superiority of the front fusion based on the co-attention mechanism, the back fusion-based PaRa-Cox was also compared. PaRa-Cox uses the proportional hazard model Cox to integrate the predicted risks of Rad-MIL and Path-MIL to obtain the final multimodal survival prediction results. All hyperparameters and training strategies required for deep learning training are consistent with PaRa-MIL to ensure the rationality of the experiment.
[0116] Table 5 shows the c-index and Td-AUROC results of all compared models. Specifically, the c-index of PaRa-MIL for kidney tumors, lung tumors, and brain tumors was 0.793±0.021, 0.682±0.028, and 0.725±0.022, respectively, and the Td-AUROC results were 0.818±0.029, 0.819±0.045, and 0.795±0.044. Compared with the single-modal models Rad-MIL and Path-MIL, PaRa-MIL showed superior results in overall performance indicators, further validating the advantages of multimodal input. In addition, compared with the PaRa-Cox post-fusion strategy, the PaRa-MIL model based on the co-attention mechanism significantly improved the prognostic indicators of all three cancers, indicating the potential advantages of the pre-fusion strategy. Compared to PaRa-MIL-wo, which does not consider pathological cell-cell interactions, PaRa-MIL demonstrated superior performance, highlighting the importance of multi-angle feature input. Furthermore, using clinical grade and staging information for prognosis prediction showed that the c-index for renal cancer staging was 0.662±0.025; the c-index for lung cancer grade prognosis prediction was 0.687±0.033; and the c-index for brain tumor staging was 0.694±0.026. Overall, PaRa-MIL, based on multimodal deep learning, achieved superior prognostic prediction performance compared to using grade and staging information.
[0117] Table 5 Prognostic prediction ability of different models
[0118]
[0119]
[0120] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A prognosis prediction method based on co-attention modality fusion, characterized in that: The following steps are involved: Step S1, image acquisition and preprocessing: acquiring a postoperative pathological image and a three-dimensional preoperative image sequence and preprocessing them separately, wherein the postoperative pathological image is an amplified HE-stained postoperative pathological image, performing segmentation, filtering, and staining normalization preprocessing on the postoperative pathological image, and performing tumor region delineation preprocessing on the three-dimensional preoperative image sequence; Step S2: Extracting pathomic representations: Applying graph neural networks and pre-trained convolutional neural networks to each pathological patch in the postoperative pathological image, the cell interaction features and morphological features of all pathological image patches are extracted to form a pathomic representation. The pathological features of a single pathological patch are composed of the splicing of cell interaction features and morphological features. The pathomic representation of a single patient is represented as a package of all pathological patch instances in the pathological tissue. Step S3, extracting radiomics representation: using the preoperative image sequence and the delineated tumor area, shape, texture, and morphological features are extracted from the image sequence based on a manually predefined feature set and a pre-trained radiomics model to form a radiomics representation; Step S4: Capturing the interaction between pathology and radiomics: Based on the extracted pathology and radiomics representations, a co-attention mechanism module is used to capture the association and interaction between the two, thereby obtaining a cross-modal pathology representation based on image co-attention. Step S5, prognosis prediction under multi-instance learning: the imaging omics representation and the cross-modal pathology omics representation based on image co-attention are input into two multi-instance learning modules respectively, the omics information is aggregated, and the prognosis prediction result is obtained.
2. A prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The preprocessing of the postoperative pathological image specifically includes: Obtaining pathological blocks: cutting the postoperative pathological images into pathological blocks of preset size; Image patch filtering: Use the OTSU method to extract the foreground tissue area in the pathological patch, calculate the ratio of the foreground tissue area to the corresponding image patch, and filter the image patches whose ratio is lower than a preset value; Brightness and color normalization: For each pathological patch obtained after filtering, brightness normalization and color normalization of the scan field are performed.
3. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The preoperative imaging sequence is obtained by CT scanning or magnetic resonance imaging. The imaging sequence type is selected based on clinical experience. Brain tumor imaging uses gadolinium-enhanced T1-weighted image sequence and T2 fluid-attenuated inversion recovery sequence, kidney tumors use enhanced arterial phase CT abdominal and pelvic scanning sequence, and lung tumors use low-dose chest CT scan.
4. A prognosis prediction method based on co-attention modality pre-fusion according to claim 3, characterized in that: The preprocessing of the preoperative image sequence specifically includes: Format conversion: convert DICOM format image sequences into NII format; Resampling: resampling the converted image sequence to a set resolution, which is determined by the image sequence type; Image registration: Advanced normalization tools were used to register T2 fluid-attenuated inversion recovery sequences to gadolinium-enhanced T1-weighted sequences. Tumor area delineation: Outline the tumor contour for each 2D scan section in the 3D sequence; Region of interest extraction: Based on the outlined tumor area, a three-dimensional region of interest containing the tumor is cropped from the full-resolution image.
5. The prognosis prediction method based on co-attention modality pre-fusion according to claim 4 is characterized in that: The tumor area delineation is specifically as follows: for kidney tumors and lung tumors, the deep segmentation learning model nnU-Net pre-trained on a public dataset is used for automatic contour delineation, and manual image reading is performed to correct the tumor boundary segmented by the model; for brain tumors, manual delineation is performed directly.
6. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The extraction of cell interaction features in step S2 relies on graph neural networks, and the extraction process specifically includes: Cell nucleus classification and segmentation: The Hover-UNet model pre-trained on the PanNuke dataset is used to segment and classify cell nuclei. The Hover-UNet model obtains a mask for each individual cell nucleus through cell nucleus segmentation and classifies the cell nuclei into tumor cells, inflammatory cells, junctional cells, dead cells, and normal epithelial cells. HistomicsTK was used to extract the pathological features of individual cell nuclei; Nuclear graph network construction: Pathological patches were modeled using a nuclear graph network. The nodes of the nuclear graph network represent cell nuclei. Each node features the pathological omics features of a single cell nucleus extracted using HistomicsTK and the probability of the cell nucleus belonging to a certain category. The connections between nodes in the graph network use the eight-nearest neighbor method, and the attributes of the edges include the connected cell type, the distance between cells, and the degree of parallelism between cells. The constructed cell nucleus graph network is trained to extract intermediate feature representations as cell interaction features. The graph network includes four edge-conditioned convolutional layers under cell spatial interactions, followed by a subgroup average pooling layer to aggregate information from all nodes to obtain the required intermediate feature representations.
7. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The morphological features of the pathological image in step S2 are the output results generated by the last layer of the ResNet34 model, and the ResNet34 model is pre-trained on ImageNet.
8. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: In step S3, feature extraction is performed using Pyradiomics based on a manually predefined shape and texture feature group. The shape and texture feature group includes five dimensions: shape, first order, texture, filtering, and wavelet. The features of each dimension are mapped to the same dimension through a linear layer. Morphological features are extracted using a visual encoder in the image-based model RadFM pre-trained on a large-scale imaging dataset. The radiomics of a single patient is represented as an aggregated package of the shape and texture feature group and the morphological features.
9. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The co-attention mechanism module in step S4 receives the pathology omics representation and the imaging omics representation extracted in steps S2 and S3, and based on the multi-head attention mechanism, takes the imaging omics representation as a query and the pathology omics representation as a key-value pair to obtain a cross-modal pathology omics representation based on imaging co-attention.
10. The prognosis prediction method based on co-attention modality pre-fusion according to claim 1, characterized in that: The multi-instance learning module in step S5 consists of a set transformer, a global attention pooling layer, and a linear layer; the multi-instance learning module is used to aggregate omics features for the imaging omics representation and the cross-modal pathology omics representation based on image co-attention, respectively, to obtain two feature vectors, and the feature vectors are spliced and connected to the linear layer to output the prognosis prediction result.
Citation Information
Patent Citations
Cancer pathological image survival prognosis model construction method based on deep learning
CN113947607A
Weakly supervised gastric cancer tissue pathology image classification method and system based on multi-mode comparative learning
CN116704245A