Method for predicting gene mutation information from lung cancer histopathology image
By constructing the NAVF-Bio model, multi-scale feature fusion and adaptive cross-view knowledge complementary strategies were used to extract multi-view information from histopathological images, solving the limitations of existing models in predicting lung cancer gene mutations, mutation subtypes and mutated exons, and achieving high-accurate gene mutation prediction.
Patent Information
- Application Number
- CN202510224833.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing deep learning models have limitations in predicting lung cancer gene mutations, mutated subtypes and mutated exons, and cannot meet the needs of accurate predictions.
A deep learning model of adaptive multi-view feature fusion, NAVF-Bio model, was constructed. Through multi-scale feature fusion and adaptive cross-view knowledge complementary strategies, multi-view information is extracted from histopathological images and predicted gene mutation information.
The NAVF-Bio model showed high accuracy in predicting key gene mutations and tumor mutation burden of lung cancer, with AUC values reaching 0.9162-0.9293, which can predict the subtype of TP53 mutation and the exons of gene mutation for the first time, achieving clinical-level performance.
Smart Images

Figure CN120148029A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and biomedical engineering, and particularly relates to a method for predicting gene mutations, mutant subtypes, and mutant exons from lung cancer histopathological images. Background Art
[0002] Lung cancer is one of the deadliest malignant tumors globally, with both the incidence and mortality rates ranking first. Accurately predicting gene mutations in lung cancer, including mutant subtypes and their exon sites, is crucial for personalized precision targeted therapy and prognosis of lung cancer patients. Gene mutations such as TP53, EGFR, KRAS, and ALK, and tumor mutational burden (TMB) are closely related to tumorigenesis, treatment response, and prognosis. Traditional gene mutation identification mainly relies on polymerase chain reaction (PCR), Sanger sequencing, fluorescence in situ hybridization (Fish), and next-generation sequencing (NGS) for determination and evaluation. However, there are two drawbacks to current gene detection for lung cancer: First, since the most commonly used gene detection method NGS today involves complex steps, requires expensive equipment and reagents, and usually takes several days to weeks to complete the complex gene detection process and result analysis, NGS is both expensive and time-consuming. Second, primary hospitals or regions with limited medical resources cannot meet the needs of patients for personalized targeted therapy. Whole-slide imaging (WSI) has revolutionized traditional pathology, providing comprehensive visual data at the cellular and tissue levels. The rich information provided in WSI, including multi-scale features from micro to macro, offers a unique opportunity to explore complex relationships within tumors.
[0003] Currently, some deep learning-based models have been applied to tumor biomarker prediction research, including ABMIL, CLAM, DSMIL, TransMIL, DTFT-MIL, IBMIL, HIPT, MHIM, R2T_MIL, and Patch-GCN. However, these models are not specifically designed for lung cancer gene prediction, so they can only roughly predict whether there are gene mutations in lung cancer and cannot study the detailed information of gene variations (such as mutant subtypes, mutant exon sites, etc.) at a deeper level. The AUC values of these models for predicting the presence or absence of mutations in lung cancer genes (TP53, EGFR, KRAS) in WSI are approximately 0.733 - 0.856, which cannot meet the requirements of precise prediction. Moreover, traditional methods often rely on manually crafted features or limited deep learning architectures, which cannot fully utilize the inherent complex spatial relationships in WSI. Finally, the interpretability of the models has only been visually analyzed at the patch level.
[0004] Therefore, the limitations of existing models result in an unsatisfactory lung cancer gene prediction function, and there is a need to design a more advanced, comprehensive, and targeted method for precisely predicting lung cancer gene mutation information. Summary of the Invention
[0005] The object of the present invention is to provide a method for predicting gene mutation information from lung cancer histopathological images, so as to solve the problem that the current deep learning models still have certain limitations in predicting gene mutations, mutation subtypes, and mutated exons as mentioned in the background art.
[0006] To achieve the above object, the present invention provides a method for predicting gene mutation information from lung cancer histopathological images, including the following steps:
[0007] S1. Collect the whole slide images of histopathology of lung cancer patients and the corresponding gene detection reports;
[0008] S2. Perform image preprocessing on the whole slide images of histopathology to construct a pathomics dataset;
[0009] S3. Construct a NAVF-Bio model for predicting gene mutation information from the whole slide images of lung cancer histopathology; the NAVF-Bio model includes a pre-training module, a pathological space topological map representation learning module, a multi-scale feature fusion module, and an adaptive cross-view knowledge supplementation module;
[0010] The pre-training module is used for feature extraction of the preprocessed images; the pathological space topological map representation learning module is used for further processing the extracted features to form one-dimensional vectors; the multi-scale feature fusion module is used for further extraction and fusion of features at different scales; the adaptive cross-view knowledge supplementation module is used for incorporating multi-view features to improve the prediction performance of the model;
[0011] S4. Use the pathomics dataset to train the constructed NAVF-Bio model, and constrain the NAVF-Bio model based on multi-scale loss and weighted fusion loss to obtain the final NAVF-Bio model;
[0012] S5. Use the final NAVF-Bio model to predict gene mutation information from the whole slide images of lung cancer histopathology, and the gene mutation information includes gene mutations, tumor mutation burden, gene mutation subtypes, and protein functional domains.
[0013] In a specific embodiment, in the step S2, the preprocessing of the whole slide images of histopathology includes digitizing the whole slide images of histopathology, segmenting tissue regions, detecting background and blurred regions.
[0014] In a specific embodiment, in step S3, the pre-training module uses the Otsu algorithm to distinguish the background area of the whole-slide image, and uses the strategy of sliding window to cut the whole-slide image into patches of large, medium and small scales, and uses the pre-trained model of CTransPath to extract feature representations with a dimension of 768 for each patch.
[0015] In a specific embodiment, in the pathological spatial topology graph representation learning module, the KNN algorithm is used to construct the points and edges of the spatial topology graph of patches at different scales; the HD-Yolo algorithm is used to segment and classify cells in the whole-slide image, and quantify the tumor microenvironment indicators. The KNN algorithm is also used to construct the cell spatial topology graph of the tumor microenvironment. The nodes of the cell spatial topology graph include the labels of cells, the probability of classification, and the area of the cell nucleus. The relationship between the same type of cells is used as the edge.
[0016] In a specific embodiment, the pathological spatial topology graph representation learning module includes an SGAEConv module and a multi-layer perceptron module;
[0017] The SGAEConv module uses a linear transformation to combine the features of the node itself and the neighbor features, which is expressed as:
[0018]
[0019] where W m , W n are both learnable weight matrices, is the embedding of node v at layer k, is the aggregated feature of neighbor nodes, and σ is the RuLU activation function;
[0020] The multi-layer perceptron module includes a ReLU activation function, layer normalization, and dropout regularization; the ReLU activation function is used to perform a non-linear transformation on the input features; based on the function of layer normalization to stabilize the training of the model and accelerate the convergence of the model; the multi-layer perceptron module uses dropout regularization to discard a part of neurons with a certain probability in each training, so as to help the model train better and prevent overfitting.
[0021] In a specific embodiment, the multi-scale feature fusion module includes an adaptive pooling layer, a splicing layer, and a self-attention feature importance extractor; the adaptive pooling layer is used to dynamically adjust the length of the input sequence to match the length of the given target size; the splicing layer is used to perform splicing in a unified feature space; the self-attention feature importance extractor uses the attention mechanism to obtain the features that are relatively important to the model according to the feature importance.
[0022] In a specific embodiment, the adaptive cross-view knowledge supplementation module includes an adaptive feature fusion module, an adaptive weight module, and a classifier; the adaptive feature fusion module uses an attention mechanism to process the input data and generate a global representation of the input features; the adaptive weight module includes two fully connected layers, and the adaptive weight module is used for feature scaling to form adaptive weights; the classifier is used to obtain the representation of each label.
[0023] In a specific embodiment, the multi-scale loss is optimized separately at different scales; the weighted fusion loss is used to integrate the losses from different scales and reflects the importance of the losses at different scales to the overall task through a weighted manner.
[0024] In a specific embodiment, predicting gene mutation information includes:
[0025] Predicting whether there is a TP53 gene mutation, whether there is an EGFR mutation, whether there is a KRAS mutation, and whether there is an ALK mutation;
[0026] Predicting the high or low tumor mutation burden;
[0027] Predicting the TP53 gene mutation subtype;
[0028] Predicting the exons of TP53, EGFR, KRAS, and ALK mutations.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The present invention aims to predict key gene mutations in lung cancer, including gene mutation subtypes and exons, by simulating the film reading steps of pathologists to observe the multi-scale features and TME features of WSI. For the first time, the present invention constructs a deep learning model for adaptive multi-view feature fusion, which extracts multi-view information such as TME from WSI by adopting multi-scale feature fusion and adaptive cross-view knowledge complementation strategies to predict key gene mutations and tumor mutation burden (TMB) status. This model is named the NAVF-Bio model. The process of multi-scale feature fusion fully simulates the mode of pathologists' film reading for feature extraction and fusion, that is, obtaining information such as tumor distribution and boundary under low magnification, obtaining the tumor tissue structure framework and cell morphological characteristics under medium magnification, and obtaining the changes in tumor cell karyotype under high magnification; the adaptive cross-view knowledge complementation (ACKC) strategy flexibly incorporates multi-view features with the help of the attention mechanism and adaptive weight module to improve the prediction performance of the model. NAVF-Bio solves the following challenges: First, in predicting the presence or absence of gene mutations and TMB status, the AUC of the NAVF-Bio model reaches 0.9162 - 0.9293 in the LCSXH-CSU and The Cancer Genome Atlas-Lung Adenocarcinoma (TCGA-LUAD) datasets, with a very high accuracy. Second, the NAVF-Bio model predicts the subtypes of TP53 mutations and the exons of gene (TP53, EGFR, KRAS, ALK) mutations for the first time, achieving clinical-grade performance, providing a reference for precision targeted drug use for lung cancer patients who fail to detect gene mutations for various reasons.
[0031] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below. Brief Description of the Drawings
[0032] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0033] Figure 1 is the training flow chart of the NAVF-Bio model according to an embodiment of the present invention;
[0034] Figure 2 is a schematic diagram of an attention mechanism according to an embodiment of the present invention, where the Kronecker product represents the operation between two arbitrary-sized matrices;
[0035] Figure 3 is the architecture diagram of the NAVF-Bio model according to an embodiment of the present invention;
[0036] Figure 4 It is the architecture diagram of the ACKC module of the NAVF-Bio model in an embodiment of the present invention;
[0037] Figure 5 It is the ROC curve diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation subtypes based on the LCSXH-CSU dataset;
[0038] Figure 6 It is the confusion matrix diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation subtypes based on the LCSXH-CSU dataset;
[0039] Figure 7 It is the ROC curve diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation subtypes based on the TCGA-LUAD dataset;
[0040] Figure 8 It is the confusion matrix diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation subtypes based on the TCGA-LUAD dataset;
[0041] Figure 9 It is the AUCROC curve diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation exons based on the LCSXH-CSU dataset;
[0042] Figure 10 It is the confusion matrix diagram of the NAVF-Bio model in an embodiment of the present invention for predicting TP53 mutation exons based on the LCSXH-CSU dataset;
[0043] Figure 11 It is the AUCROC curve diagram of the NAVF-Bio model in an embodiment of the present invention for predicting EGFR mutation exons based on the LCSXH-CSU dataset;
[0044] Figure 12 It is the confusion matrix diagram of the NAVF-Bio model in an embodiment of the present invention for predicting EGFR mutation exons based on the LCSXH-CSU dataset;
[0045] Figure 13 It is the AUCROC curve diagram of the NAVF-Bio model in an embodiment of the present invention for predicting KRAS mutation exons based on the LCSXH-CSU dataset;
[0046] Figure 14 It is the confusion matrix diagram of the NAVF-Bio model in an embodiment of the present invention for predicting KRAS mutation exons based on the LCSXH-CSU dataset;
[0047] Figure 15The AUC-ROC curve of the NAVF-Bio model of an embodiment of the present invention for predicting ALK mutant exons based on the LCSXH-CSU dataset;
[0048] Figure 16 It is the confusion matrix diagram of the NAVF-Bio model of an embodiment of the present invention for predicting ALK mutant exons based on the LCSXH-CSU dataset. Detailed implementation manners
[0049] The following is a detailed description of the embodiments of the present invention. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] The present invention provides a method for predicting gene mutation information from lung cancer histopathological images, including the following steps:
[0051] S1. Collect the whole slide images (WSI) of lung cancer patients' histopathology and the corresponding gene detection reports;
[0052] S2. Perform image preprocessing on the whole slide images of histopathology to construct a pathomics dataset;
[0053] S3. Construct a NAVF-Bio model for predicting gene mutation information from lung cancer histopathological whole slide images; the NAVF-Bio model includes a pre-training module, a pathological space topology graph representation learning module, a multi-scale feature fusion module, and an adaptive cross-view knowledge supplementation module;
[0054] The pre-training module is used to extract features from the preprocessed images; the pathological space topology graph representation learning module is used to further process the extracted features to form one-dimensional vectors; the multi-scale feature fusion module is used for further extraction and fusion of features at different scales; the adaptive cross-view knowledge supplementation module is used to incorporate multi-view features to improve the prediction performance of the model;
[0055] S4. Use the pathomics dataset to train the constructed NAVF-Bio model, and constrain the NAVF-Bio model based on multi-scale loss and weighted fusion loss to obtain the final NAVF-Bio model;
[0056] S5. Use the final NAVF-Bio model to predict gene mutation information from lung cancer histopathological whole slide images, and the gene mutation information includes gene mutation, tumor mutation burden, gene mutation subtype, and protein functional domain.
[0057] Embodiment 1
[0058] A method for predicting biomarkers from lung cancer histopathological images, including the following steps:
[0059] S1. Collect the histopathological whole-slide images of lung cancer patients and the corresponding gene detection reports.
[0060] S2. Perform image preprocessing on the histopathological whole-slide images to construct a pathomics dataset.
[0061] The present invention collected the histopathological images and high-throughput sequencing gene detection reports of 1576 lung cancer patients and constructed a large histopathological image dataset.
[0062] The step S2 specifically includes: digitalizing the histopathological whole-slide images, segmenting tissue regions, detecting background and blurred regions, and constructing a large pathomics dataset.
[0063] S3. Construct a NAVF-Bio model for predicting gene mutation information from histopathological whole-slide images of lung cancer; the NAVF-Bio model includes a pre-training module, a pathological space topology map representation learning module, a multi-scale feature fusion module, and an adaptive cross-view knowledge supplementation module;
[0064] The pre-training module uses the Otsu algorithm to distinguish the background region of the WSI and uses the sliding window strategy to cut the WSI into many patches of large, medium, and small scales. Preferably, the pixel size of the large-scale patch is 1024×1024, the pixel size of the medium-scale patch is 512×512, and the pixel size of the small-scale patch is 256×256, and the spatial position of the patch is marked. Use the pre-trained model of CTransPath to extract one-dimensional features for each patch at multiple scales to enrich the diversity feature representation of the image. CTransPath is composed of three convolutional layers and four Swin Transformer modules based on the Swin Transformer architecture, which can extract local features in the image data, capture global features and long-range dependencies. Based on CTransPath, the embedded features of the patches are extracted, and each patch embedding is unified into a 768-dimensional vector representation. Therefore, all patches at each scale in the WSI are represented in the form of feature vectors.
[0065] However, the patches in the WSI have spatial position relationships. The present invention uses the KNN algorithm to construct the points and edges of the spatial topology graph (G l , G m , G s ). Secondly, in order to construct the tumor microenvironment (TME) of the WSI, the present invention uses the HD-Yolo algorithm to segment and classify the cells in the WSI and quantify the TME indicators. The present invention also uses the KNN algorithm to construct the cell spatial topology graph of the TME (GT ), where the nodes of the graph contain the labels of cells, the probabilities of classification, and the areas of cell nuclei, and the relationships between cells of the same type are used as edges. Therefore, the multi-scale features and TME at different field-of-view scales are respectively represented as G l =(V l , E l ), G m =(V m , E m ), G s =(V s , E s ), G T =(V T , E T ).
[0066] Spatial topological graph representation learning. The spatial topological graphs at different field-of-view scales and the TME spatial topological graph in the present invention are respectively represented as G l =(V l , E l ), G m =(V m , E m ), G s =(V s , E s ), G T =(V T , E T ). The present invention proposes a pathological spatial topological graph representation learning module for learning, including the SGAEConv module and the MultilayerReception module (multilayer perception module). Each node v ∈ V of the multi-scale G has a feature vector representing the node embedding at the k-th layer. The purpose of SGAEConv is to generate the node embedding at the k + 1-th layer. First, the SGAEConv module samples a subset N S (v) from its set of neighbor nodes N(v), and then averages and aggregates the features of these neighbor nodes. The process is as follows:
[0067]
[0068] where N S (v) is the neighbor aggregation of node v, and the features of the sampled neighbor nodes are averaged and aggregated. Second, the feature of node v itself is combined with the aggregated neighbor features to obtain a new embedding representation of node v. The SAGEConv module uses a linear transformation to combine the features of the node itself and the neighbor features, which is expressed as:
[0069]
[0070] where Wm ,W n are all learnable weight matrices, is the embedding of node v at layer k, is the aggregated feature of neighbor nodes, and σ is the RuLU activation function. To stabilize the training process, it is preferable to normalize the aggregated features so that the embeddings of each node have a stable numerical range. Here, to prevent gradient explosion or vanishing, the model performs L2 normalization on the output of the graph neural network:
[0071]
[0072] where ||.|| represents the norm of the input feature vector.
[0073] The features processed by the SGAEConv module need to be further processed by a multi-layer perceptron module (MPB). The multi-layer perceptron module includes a ReLU activation function, layer normalization, and dropout regularization. To enable the model to learn more complex feature representations, the present invention uses the ReLU activation function to perform a non-linear transformation on the input features. Secondly, based on the function of layer normalization to stabilize the training of the model and accelerate the convergence of the model, the model makes the input features more robust. Finally, the multi-layer perceptron module uses dropout regularization to discard a part of neurons with a probability of 20% in each training, thereby helping the model to train better and preventing overfitting. Especially when dealing with ultra-large graphs, dropout regularization forces the model to learn more robustly and improves the generalization ability.
[0074] The features extracted by the pathological space topology graph representation learning module are extremely long one-dimensional vectors The process of multi-scale feature fusion fully simulates the process of a pathologist's film reading experience for feature extraction and fusion. Feature fusion is to fuse the TME features with the small field of view features and the medium field of view features respectively. Therefore, the present invention proposes a feature fusion module to guide pairwise fusion of different features. This module consists of an adaptive pooling layer, a concatenation layer, and a self-attentive feature importance extractor.
[0075] Different from ordinary pooling layers, Adaptive Average Pooling can automatically adjust the size and stride of the pooling window according to the given target output size, so that the size of the output feature map is consistent with the expectation. The purpose of this layer is to dynamically adjust the length of the input sequence so that the output length matches the given target size. Its core idea is to divide the input sequence into several intervals and then apply the max pooling operation to each interval. Secondly, the present invention uses a concatenation layer to concatenate in a unified feature space and and
[0076] Given the huge number of patches and cells in WSI, the fused features need to further extract key information. Therefore, a key step in feature fusion is to use an attention mechanism to obtain features that are relatively important to the model according to feature importance. Let and and The merged features are respectively and To save computing resources, the present invention proposes an attention mechanism with an algorithmic complexity of o(N) to learn and The present invention first maps and to query, key, and value vectors respectively, and the process is as follows:
[0077]
[0078] where W Q ∈R d×d , W K ∈R d×d , W V ∈R d×d are trainable parameters respectively. The present invention uses L2 normalization to process the key and query vectors. The present invention reduces the complexity of attention by swapping the order of matrix multiplications. With the Taylor expansion of the exponential function, it can be used to motivate a new attention function:
[0079]
[0080] where Ι N×1 is an all-one vector. In the above calculations, the present invention realizes the aggregation of full attention, avoids the N×N attention matrix, and only requires a linear complexity of o(Nd 2 ). For large vectors and N is usually several orders of magnitude larger than d, so the computational efficiency can be significantly improved in practice.
[0081] The adaptive cross-view knowledge supplementation module can flexibly incorporate features from multiple views to enhance the prediction performance of the model.
[0082] The adaptive cross-view knowledge supplementation (ACKC) module consists of an adaptive feature fusion (Adaptive Feature Fusion) module, an adaptive weight module, and a classifier. The role of the adaptive feature fusion module is to adaptively fuse features from different sources or different levels, thereby generating a more effective feature representation. The goal of this module is to dynamically select and combine features in order to obtain better performance in downstream tasks, including classification and regression. For the large, medium, and small field-of-view features and TME feature representations input to the ACKC module and The multilayer perceptron (MLP) first maps the features of different branches into the same dimensional space, and then uses a concatenation layer to concatenate and combine and
[0083]
[0084] Then, the adaptive feature fusion module uses the attention mechanism to process the input data and generate a global representation of the input features. Given that the features of different branches are in the same dimensional space, the attention mechanism can capture the global features of all features and the context dependencies, and further extract important features. The idea is to perform linear mapping and positional encoding of features with one-dimensional features. The attention mechanism can capture the dependencies between each feature and the global in long sequence data. The stacked multi-layer attention mechanism layers can extract multi-level inputs of features, from low-level features to high-level global semantic features. Secondly, the attention mechanism can dynamically assign weights to the input features during feature extraction, allowing the model to focus on important features, thereby effectively extracting key information. Specifically, the input feature X first undergoes linear projection, and the process is as follows:
[0085] X p = W emb ·X + b emb
[0086] where is the weight matrix, b emb is the bias, and X p is the result of the linear mapping of feature X. To adapt to the calculation of the attention mechanism module, the present invention uses one-dimensional feature X pThe length S is divided into multiple blocks, each with a length of C. Assume that X p is divisible, then the number of blocks N = S / C. Secondly, positional encoding is added to each input sequence block so that the model can understand the order information of each position in the sequence. Next, X is processed through the Transformer encoder p . Each layer of the transformer includes an attention mechanism and a feed-forward neural network. The process is as follows:
[0087]
[0088] FFN(X) = W 2 · ReLU(W 1 · X + b 1 ) + b 2
[0089] To enhance the expressive ability of key features and balance the contributions of different features, the features after feature importance screening need to be further processed by an adaptive weight module. This step is also to suppress irrelevant and redundant features, supplement the deficiencies of the attention mechanism, and improve the generalization performance of the model. Specifically, the adaptive weight module consists of two fully connected layers to perform feature scaling to form adaptive weights. Assume the weight matrices of the fully connected layers are respectively where r is the scaling factor. Then the feature after full connection is:
[0090] s = W m · ReLU(W n · x)
[0091] where ReLU represents the ReLU activation function. Then, its feature needs to pass through the Sigmoid activation function to generate weights and limit the output to [0, 1]. Finally, the generated weights are used to perform element-wise multiplication on the input features. Therefore, the working process of the adaptive weight module can be summarized as:
[0092] y = x ⊙ σ(W m · ReLU(W n · x))
[0093] where σ is the Sigmoid activation function and ⊙ is element-wise multiplication. To obtain the final output result, the present invention uses a Linear layer as the backbone of the feature classifier to obtain the representation of each label.
[0094] Based on the multi-view loss and weighted fusion loss, the model is constrained and multiple biomarkers are predicted, including gene mutations, tumor mutation burden, gene mutation subtypes, and protein functional domains.
[0095] Since the model obtains features from multi-scale images of WSI, the present invention proposes a weighted multi-scale and fused Loss to optimize the performance of the model. The multi-scale loss is optimized separately at different scales to fully exploit the diverse information in the data. The weighted fused Loss is used to integrate these Losses from different scales and reflect their importance to the overall task through weighting. The combination of the two helps the model understand the data from multiple scales during the learning process, improve the comprehensiveness of feature learning, and ensure that the contributions of each scale are reasonably treated, thereby improving the overall performance, robustness, and generalization of the model. For the features at each scale, the present invention defines the corresponding loss function Loss ms . In the present invention, the biomarker prediction task can be reduced to a classification task, and the cross-entropy loss function can be used to optimize the performance of the model. Secondly, the present invention assigns different weights to the loss functions according to the contributions of each scale. Assuming there are N ∈ [1,4], the process can be expressed as:
[0096]
[0097] where λ is the weight coefficient. x and y represent the input data and labels at different scales respectively. Each Loss adp is the loss after multi-scale fusion. Through these weights, the influence of the loss of each scale on the overall optimization process can be dynamically adjusted.
[0098] All models are trained on a Linux host, based on the pytorch2.1 platform and on 4 V100s and 1 H800. -5 . The learning rate is set to 1*10 -4 . The training step size is 100. The present invention uses a five-fold cross-validation strategy to train the model, and when the model achieves good prediction results in each fold, the best-performing model in that fold is saved.
[0099] The present invention uses the Area Under the Receiver Operating Characteristic Curve (AUCROC) curve, Precision-Recall Curve (PRC) curve, Accuracy, recall, precision, F1-score, AUCROC as evaluation metrics for the model. The larger the area under the AUCROC curve and the PRC curve, the larger the evaluation metric value, and the better the model performance.
[0100] Experimental results:
[0101] Currently, there are 11 AI models that can perform crude cancer gene mutation prediction. These models are: ABMIL, CLAM-SB, CLAM-MB, DSMIL, TransMIL, DTFT-MIL, IBMIL, HIPT, MHIM, RRT_MIL, and Patch-GCN. The NAVF-Bio model of the present invention has opened up a new track in the field of lung cancer gene prediction, successfully improving the prediction accuracy of key lung cancer gene mutation subtypes and mutant exons to the clinical level. Specifically, in the two datasets of LCSXH-CSU and TCGA-LUAD, compared with other models, the NAVF-Bio model of the present invention shows extremely high accuracy in the tasks of predicting the presence or absence of TP53 gene mutation (Tables 1 and 6), the presence or absence of EGFR mutation (Tables 2 and 7), the presence or absence of KRAS mutation (Tables 3 and 8), the presence or absence of ALK mutation (Tables 4 and 9), and distinguishing high and low TMB (Tables 5 and 10), reaching the clinical application level and having strong generalization performance, which can be widely applied in different datasets.
[0102] In addition, the NAVF-Bio model of the present invention can achieve good results in the classification of TP53 gene mutation subtypes. The results of ROC curve analysis show that the NAVF-Bio model has good prediction performance for TP53 gene mutation subtypes in both the LCSXH-CSU dataset and the TCGA-LUAD dataset. The confusion matrix shows that NAVF-Bio obtains good accuracy for each type of TP53 mutation subtype. Especially in the prediction of wild-type expression, the accuracy rates reach 0.8728 and 0.9312 respectively. The prediction results of the NAVF-Bio model for the exons of TP53, EGFR, KRAS, and ALK mutations show that the ROC curves of each key gene mutation exon exhibit high prediction performance, and their AUCs reach 0.9313, 0.8772, 0.9476, and 0.9467 respectively. The confusion matrix shows that NAVF-Bio obtains good accuracy in the localization of each gene mutation exon.
[0103] Table 1. Performance comparison of the present invention and 11 models based on the LCSXH-CSU dataset in the TP53 binary classification task.
[0104] AUCROC ACC Precision Recall F1-Score ABMIL 80.9±7.28 71.16±9.23 78.38±5.87 79.05±3.39 78.06±4.59 CLAM-SB 88.53±8.82 85.61±8.89 86.43±8.43 86.5±8.31 85.47±9.2 CLAM-MB 88.47±8.47 85.74±8.5 86.37±8.33 86.86±8.52 85.65±9.07 DSMIL 83.96±8.13 78.02±6.07 81.3±6 81.26±5.41 79.03±7.44 TransMIL 92.15±7.02 91.61±7.43 91.95±7.13 91.61±7.43 91.73±7.32 DTFD-MIL 66.96±4.35 65.15±2.44 65.18±2.43 65.22±2.42 61.61±4.71 IBMIL 64.72±3.88 62.25±1.46 62.46±1.45 62.69±1.43 59±4.32 HIPT 79.35±6.46 72.23±4.61 73.08±4.73 74.82±4.99 73.66±4.75 MHIM 51.4±1.29 45.09±2.19 58.49±0.57 52.39±2.29 46.43±3.53 RRT_MIL 52.69±1.85 52.23±1.53 58.05±1.12 62.28±2.75 55.72±2.4 Patch-GCN 87.98±6.78 83.15±5.67 83.3±5.69 83.59±5.65 83.08±5.77 The present invention 92.93±6.27 92.3±6.24 92.3±6.23 92.38±6.18 92.3±6.24
[0105] Table 2. Performance comparison of the present invention and 11 models based on the LCSXH-CSU dataset in the EGFR binary classification task.
[0106] AUCROC ACC Precision Recall F1-Score ABMIL 83.71±5.1 79.04±4.74 79.7±4.49 81.07±4.15 79.72±4.35 CLAM-SB 82.35±5.54 73.09±4.53 74.42±4.46 76.64±4.74 75.87±4.64 CLAM-MB 69.22±13.64 76±4.85 77.12±4.33 79.33±5.15 76.82±4.6 DSMIL 85.84±6.37 78.6±5.55 81.51±5.08 82.27±5.21 81.18±5.34 TransMIL 90.44±7.15 88.16±6.19 88.2±6.15 88.16±6.19 87.42±6.66 DTFD-MIL 66.76±2.88 64.54±1.99 64.6±1.97 64.6±1.97 63.38±2.02 IBMIL 66.1±2.79 63.97±1.14 64.07±1.13 64.07±1.13 62.47±1.61 HIPT 82.02±5.75 74.36±4.59 75.58±4.49 74.55±4.46 74.74±4.41 MHIM 51.98±0.76 47.06±2.57 60.51±0.8 54.52±1.88 46.82±1.6 RRT_MIL 53.69±1.85 52.25±0.98 57.79±0.87 58.28±1.49 54.38±1.3 Patch-GCN 75.43±5.55 71.45±4.01 72.12±3.99 72.12±3.99 70.18±4.45 The present invention 92.32±6.69 91.15±6.71 91.46±6.43 91.32±6.55 91.61±6.3
[0107] Table 3. Performance comparison of the present invention and 11 models based on the LCSXH-CSU dataset in the KRAS binary classification task.
[0108] AUCROC ACC Precision Recall F1-Score ABMIL 77.96±8.69 59.8±9.43 62.93±7.75 60.66±9.46 52.09±6.6 CLAM-SB 87.7±9.61 94.63±2.06 94.99±1.8 95.01±1.98 84.23±7.58 CLAM-MB 87.41±9.69 90.72±4.56 91.44±4.38 91.73±4 82.83±7.68 DSMIL 80.8±8.91 77.53±3.34 84.16±2.78 83.71±3.93 64.6±3.4 TransMIL 88.43±10.35 97.98±1.6 97.98±1.6 98.11±1.62 89.75±8.71 DTFD-MIL 69.03±5.93 69.44±3.41 70.79±3.14 72.65±2.82 55.79±2.37 IBMIL 66.14±7.12 76.82±2.66 77.1±2.61 77.11±2.67 58.8±3 HIPT 74.43±7.48 58.22±10.21 58.24±10.19 58.26±10.21 48.64±7.8 MHIM 53.07±2.86 65..06±2.68 78.69±2.82 79.84±3.89 54.31±0.87 RRT_MIL 53.7±1.83 65.83±5.76 77.98±4.07 71.07±6.06 51.65±1.82 Patch-GCN 72.57±6.73 78.79±4.05 79.49±4.09 79.85±4.4 60.13±3.15 The present invention 90.6±7.74 96.7±2.32 96.64±2.41 96.54±2.48 96.77±2.34
[0109] Table 4. Performance comparison between the present invention and 11 models on the ALK binary classification task based on the LCSXH-CSU dataset.
[0110] AUCROC ACC Precision Recall F1-Score ABMIL 79±5.5 71.51±1.42 73.76±1.17 72.2±1.52 54.99±2.01 CLAM-SB 90.53±8.08 94.5±2.92 94.77±2.87 95.07±2.54 86.18±7.54 CLAM-MB 90.58±8.01 94.5±2.19 94.53±2.18 94.75±2.09 83.08±6.78 DSMIL 84.74±5.98 77.51±3.5 84.81±2.14 84.01±4.35 64.75±3.8 TransMIL 91.21±7.22 96.02±1.97 96.08±1.97 96.02±1.97 87.08±7.04 DTFD-MIL 71.58±3.19 75.49±1.72 75.9±1.83 75.55±1.77 55.8±1.95 IBMIL 69.51±3.11 76.56±3.4 77.26±3.56 76.69±3.4 56.55±3.2 HIPT 78±3.65 78.9±2.35 78.95±2.34 78.9±2.35 58.77±2.21 MHIM 48.88±1.86 81.83±4.38 88.06±3.95 87.46±1.12 55.53±1.06 RRT_MIL 49.89±3.29 84.71±4.48 87.06±3.72 86.16±4.64 52.69±1.13 Patch-GCN 72.78±4.36 78.6±4.98 79.67±4.57 78.66±4.94 58.12±3.26 The present invention 92.58±7.32 98.83±0.97 98.41±1.35 98.4±1.36 98.83±0.97
[0111] Table 5. Performance comparison between the present invention and 11 models on the TMB binary classification task based on the LCSXH-CSU dataset.
[0112] AUCROC ACC Precision Recall F1-Score ABMIL 81.31±6.11 74.84±4.75 76.61±4.45 80.43±5.08 66.07±4.78 CLAM-SB 90.43±8.48 95.99±2.87 95.99±2.87 95.99±2.87 89.78±7.35 CLAM-MB 90.49±8.47 95.87±3.02 96.12±3.05 95.87±3.02 90.31±7.6 DSMIL 86.44±7.66 80.71±6.04 83.33±4.92 90.18±6.35 74.96±5.41 TransMIL 91.1±7.96 95.74±3.81 95.93±3.64 96.65±3 92.08±7.08 DTFD-MIL 77.28±5.54 77.7±2.84 77.76±2.82 77.95±2.78 64.57±2.52 IBMIL 73.94±4.95 72.5±5.12 72.85±4.97 72.65±5.18 59.98±2.97 HIPT 79.41±6.1 78.87±3.94 79.05±3.83 78.87±3.94 66.04±3.36 MHIM 50.98±2.61 69.8±5.6 77.38±4.41 83.01±0.85 57.1±0.92 RRT_MIL 49.8±1.92 73.63±2.25 80.13±2.1 78.7±3.33 54.53±1.86 Patch-GCN 80.42±7.67 86.91±2.58 86.91±2.58 86.91±2.58 71.58±6.02 The present invention 91.62±7.75 96.6±2.5 96.67±2.42 96.59±2.5 96.77±2.33
[0113] Table 6. Performance comparison between the present invention and 11 models on the TP53 binary classification task based on the TCGA-LUAD dataset.
[0114] AUCROC ACC Precision Recall F1-Score ABMIL 80.9±16.3 71.16±20.65 65.86±2.76 66.93±4 78.06±10.27 CLAM-SB 88.53±19.72 85.61±19.88 82.92±7.37 83.57±6.87 85.47±20.57 CLAM-MB 88.47±18.95 85.74±19 80.46±7 80.99±6.86 85.65±20.3 DSMIL 71.45±5.99 63.37±3.33 68.76±3.88 69.32±4.46 68.29±3.84 TransMIL 91.19±7.39 89.51±8.07 90.64±7.06 92.28±5.6 91.34±6.43 DTFD-MIL 58.88±2.12 62.78±0.35 63.97±0.55 63.57±0.178 60.72±1.85 IBMIL 59.1±2.29 61.79±0.53 63.22±0.78 63.57±0.52 59.38±1.96 HIPT 69.09±2.68 65.35±2.08 67.26±2.33 65.35±2.08 65.99±2.23 MHIM 55.82±0.55 38.28±2.49 60.2±1.21 60.74±4.99 59.68±1.99 RRT_MIL 56.59±0.85 51.6±1.6 62.06±1.12 59.13±3.18 57.19±2.47 Patch-GCN 64.28±2.93 63.96±2.29 64.91±1.71 65.94±0.99 63.09±2.54 The present invention 90.18±8.78 92.2±6.98 92.34±6.85 92.28±6.91 92.4±6.8
[0115] Table 7. Performance comparison between the present invention and 11 models on the EGFR binary classification task based on the TCGA-LUAD dataset.
[0116] AUCROC ACC Precision Recall F1-Score ABMIL 72.21±1.35 71.09±2.85 79.11±2.44 74.06±3.58 66.26±1.02 CLAM-SB 68.64±2.11 80.2±1.96 82.23±1.45 81.59±1.76 66.25±1.48 CLAM-MB 72.21±1.35 71.09±2.85 79.11±2.44 74.05±3.57 66.26±1.02 DSMIL 73.39±5.32 78.02±2.25 82.8±2.06 85.35±1.57 69.44±2.93 TransMIL 94.84±4.55 92.87±5.51 93.13±5.28 93.27±5.16 91.1±6.44 DTFD-MIL 58.89±1.99 57.23±5.3 58.25±5.21 59.61±5.21 51.62±3.53 IBMIL 60.16±3.33 51.73±7.35 52.15±7.37 51.98±7.55 46.79±4.61 HIPT 63.71±4.51 73.47±5 75.32±4.33 75.84±3.47 62.48±3.12 MHIM 56.15±1.53 60.51±5.99 76.74±3.85 71.23±5.28 56.51±2.31 RRT_MIL 61.71±2.68 71.86±4.39 78.17±2.22 85.28±4.39 66.56±1.24 Patch-GCN 60.76±3.09 77.63±3.81 77.77±3.78 77.63±3.81 61.81±1.86 The present invention 87.72±7.49 77.44±11.02 86.69±8.46 86.18±8.23 88.43±8.2
[0117] Table 8. Performance comparison between the present invention and 11 models on the KRAS binary classification task based on the TCGA-LUAD dataset.
[0118] AUCROC ACC Precision Recall F1-Score ABMIL 90.33±2.83 85.35±3.16 86.23±3.3 85.35±3.16 65.2±4.85 CLAM-SB 96.29±0.53 96.61±0.87 96.61±0.87 96.61±0.87 92.08±5.07 CLAM-MB 96.54±0.21 96.4±0.22 96.41±0.22 96.41±0.22 94.74±2.05 DSMIL 90.15±2.62 91.19±4.73 84.94±4.17 83.76±5.03 64.32±5.49 TransMIL 96.3±1.32 96.23±2.05 96.23±2.05 96.23±2.05 89.97±6.94 DTFD-MIL 83.42±5.02 78.02±4.89 78.19±4.95 78.02±4.89 58.28±5.27 IBMIL 83.87±5 77.43±5.69 77.6±5.69 77.63±5.78 58.59±5.65 HIPT 87.93±1.86 81.59±2.97 81.68±2.98 82.18±3.1 60.08±3.04 MHIM 88.77±2.08 83.74±0.94 85.6±1.07 87.72±2.4 62.64±0.9 RRT_MIL 69.97±3.94 69.48±4.46 75.86±4.4 77.26±3.92 57.15±2.86 Patch-GCN 65.25±5.65 69.9±15.64 71.69±15.59 71.88±15.8 51.83±11.06 The present invention 96.67±2.85 97.09±2.39 97.03±2.48 96.92±2.55 97.15±2.41
[0119] Table 9. Performance comparison between the present invention and 11 models on the ALK binary classification task based on the TCGA-LUAD dataset.
[0120] AUCROC ACC Precision Recall F1-Score ABMIL 74.94±3.98 89.11±2.79 89.54±2.78 89.51±2.64 71.54±3.92 CLAM-SB 75.88±4.7 79.21±4.41 83.28±4.57 85.35±4.54 62.74±4.29 CLAM-MB 83.9±5.06 81.98±4.65 82.04±3.96 80±4.47 70.03±5.5 DSMIL 80.65±6.73 84.36±3.58 88.01±3.23 90.3±3.4 71.96±4.88 TransMIL 89.68±7.41 97.43±1.67 97.44±1.67 98.02±1.77 91.07±5.72 DTFD-MIL 64.32±3.8 80.4±2.99 83.13±3.42 80.4±3 57.03±1.47 IBMIL 65.74±3.45 81.19±3.35 81.77±3.51 81.59±3.78 59.97±3.52 HIPT 62.15±3.95 80.6±4.5 80.98±4.25 81.78±3.77 59.4±4.2 MHIM 66.43±4.36 73.79±5.05 83.96±3.42 78.95±4.87 59.07±1.56 RRT_MIL 69.97±3.94 69.49±4.46 75.86±4.4 77.26±3.92 57.15±2.86 Patch-GCN 64.23±4.65 78.22±6.37 78.58±6.27 78.42±6.26 60.34±6.28 The present invention 89.86±8.64 98.83±0.97 98.41±1.35 98.83±0.97 98.4±1.36
[0121] Table 10. Performance comparison between the present invention and 11 models on the TMB binary classification task based on the TCGA-LUAD dataset.
[0122] AUCROC ACC Precision Recall F1-Score ABMIL 65.1±3.87 67.54±3.5 69±3.13 69.14±2.68 66.67±2.57 CLAM-SB 86.1±8.04 85.47±6.06 85.47±6.06 85.47±6.06 85.16±6.14 CLAM-MB 86.98±7.29 84.64±5.46 85.24±5.01 85.85±4.82 85.25±4.9 DSMIL 80.41±6.95 72.82±4.1 81.91±5.09 76.89±4.39 78.09±4.19 TransMIL 90.76±8.19 91.58±6.63 91.73±6.49 91.98±6.27 91.83±6.39 DTFD-MIL 63.29±2.53 67.09±1.39 67.89±1.01 67.09±1.39 65.23±1.28 IBMIL 63.05±2.76 65.07±1.72 66.79±1.19 65.48±1.77 65.17±1.29 HIPT 71.44±5.02 70.78±2.73 71.02±2.72 71.59±2.75 70.7±2.61 MHIM 53.97±1.71 45.33±1.97 57.7±1.59 59.59±3.22 53.38±3.45 RRT_MIL 55.6±2.3 51.87±1.69 61.25±1.07 66.47±4.29 61.22±3.27 Patch-GCN 80.42±7.67 86.91±2.58 86.91±2.58 86.91±2.58 71.58±6.02 The present invention 92.52±6.69 94.61±4.82 95.02±4.46 95.61±3.93 94.61±4.82
[0123] Table 11. Performance comparison between the present invention and 11 models on the TP53 multi-classification task based on the LCSXH-CSU dataset.
[0124] AUCROC ACC Precision Recall F1-Score ABMIL 59.97±2.59 25±2.12 57.81±0.65 46.15±2.47 45.24±3.82 CLAM-SB 86.35±7.95 64.34±7.27 74.96±6.44 84.54±4.18 75.3±7.48 CLAM-MB 86.37±7.95 64.21±7.23 74.51±6.25 85.04±4.24 75.23±7.46 DSMIL 59.58±2.3 25.95±1.75 50.21±3.72 51.65±4.99 43.08±3.85 TransMIL 86.84±7.18 77.72±10.22 82.0±8.32 86.21±5.76 83.62±8.27 DTFD-MIL 61.76±5.1 36.2±1.13 47.71±2.61 62.07±5.25 47.03±4.88 IBMIL 63.43±4.03 37±3.98 49.02±1.65 64.2±2.2 48.83±2.91 HIPT 81.13±6.76 59.08±5.69 62.99±5.4 72.86±6.49 64.04±6.2 MHIM 49.76±1.27 33.58±3.18 48.94±2.04 56.54±3.53 36.55±2.68 RRT_MIL 50.95±1.86 39.15±4.39 51.45±1.68 55.44±3.6 37.11±2.55 Patch-GCN 76.59±5.22 49.81±3.35 57.93±2.93 71.91±6.78 59.36±5.1 The present invention 87.33±6.45 79.28±6.95 82.41±6.06 86.95±5.98 85.06±5.76
[0125] Table 12. Performance comparison between the present invention and 11 models on the TP53 multi-classification task based on the TCGA-LUAD dataset.
[0126] AUCROC ACC Precision Recall F1-Score ABMIL 51.39±1.27 13.12±4.13 26.24±2.42 63.87±8.29 30.44±1.96 CLAM-SB 49.22±0.81 8.51±3.69 28.42±8.97 27.33±11.84 18.27±5.92 CLAM-MB 48.31±1.71 1.78±0.59 35.25±15.22 33.86±12.23 20.41±7.46 DSMIL 47.80±1.53 9.31±5.74 24.8±2.58 71.09±6.01 32.25±1.68 TransMIL 90.96±8.05 80.99±15.48 84.08±13.02 91.49±7.4 85.13±12.16 DTFD-MIL 61.99±3.58 12.12±2.57 30.09±2.8 61.52±5.07 36.02±1.9 IBMIL 62.78±3.8 8.33±2.31 31.31±2.75 66.09±5.61 36.59±2.33 HIPT 53.95±3.09 26.18±1.81 35.27±1.22 55.34±4.3 32.47±1.46 MHIM 56.81±1.07 11.73±4.48 32.75±3.23 62.24±9.04 36.04±1.63 RRT_MIL 90.95±8.05 80.2±15.29 83.49±12.88 91.29±7.35 84.84±12.09 Patch-GCN 65.29±5.08 23.03±3.8 27.31±2.64 55.78±2.89 34.37±1.86 The present invention 91.21±8.74 89.6±9.08 91.08±7.84 89.55±9.07 93.07±6.2
[0127] Table 13. Performance comparison between the present invention and 11 models based on the LCSXH-CSU dataset for the TP53 functional domain prediction task.
[0128] AUCROC ACC Precision Recall F1-Score ABMIL 84.79±7.99 38.02±8.37 52.25±7.03 83.26±8.07 63.46±7.81 CLAM-SB 89.94±8.78 78.93±16.34 82.81±12.99 91.43±7.33 84.9±12.17 CLAM-MB 89.91±8.89 78.78±16.67 82.68±13.07 91.43±7.17 84.75±12.19 DSMIL 89.02±8.1 67.11±12.95 74.49±10.9 87.03±6.07 78.57±9.99 TransMIL 90.04±8.91 81.8±16.27 84.91±13.5 88.97±9.86 86.27±12.28 DTFD-MIL 80.66±6.39 27.26±5.93 44.73±4.46 77.06±8.28 55.96±5.48 IBMIL 75.38±5.7 16.95±3.49 38.22±3.16 77.95±6.17 49.23±5.17 HIPT 88.04±8.46 61.06±13.51 66.6±10.34 88.68±4.26 74.18±9.45 MHIM 51.51±1.29 4.09±1.47 24.28±0.58 65.84±2.71 33.45±1.01 RRT_MIL 53.11±0.9 13.18±1.31 27.89±0.68 41.36±2.72 31.59±1.52 Patch-GCN 86.2±7.4 47.87±11.31 56.88±8.27 85.96±4.51 67.1±7.73 The present invention 93.13±6.8 27.99±8.01 68.05±7.79 59.16±6.92 87.72±7.67
[0129] Table 14. Performance comparison between the present invention and 11 models based on the LCSXH-CSU dataset for the EGFR functional domain prediction task.
[0130] AUCROC ACC Precision Recall F1-Score ABMIL 67.99±3.55 41.89±4.83 71.33±8.29 56.73±3.57 57.88±4.96 CLAM-SB 56.31±2.78 27.38±1.98 56.45±2.91 49.02±1.02 41.45±1.03 CLAM-MB 59.72±2.49 23.7±4.44 53.71±3.98 50.96±5.03 41.99±2.51 DSMIL 65.88±3.41 35.44±2.52 66.1±5.83 59.15±1.59 54.77±4.05 TransMIL 89.94±8.66 82.18±11.71 85.96±9.41 88.15±7.4 85.37±10.11 DTFD-MIL 74.2±4.21 53.28±0.81 61.75±1.28 69.52±4.66 56.48±2.5 IBMIL 72.43±3.68 47.65±3.85 54.77±2.22 67.78±2.43 52.98±1.92 HIPT 70.54±3.19 33.49±2.38 53.83±3.07 81.11±6.05 61.42±4.49 MHIM 50.08±0.98 23.98±4.33 39.74±2.99 54.02±1.8 38.37±0.7 RRT_MIL 53.3±1.26 34.91±5.09 49.46±2.31 59.8±3.88 42.6±2.25 Patch-GCN 79.99±5.58 56.29±5.45 60.99±4.62 71.96±4.32 58.51±4.49 The present invention 90.72±7.49 77.44±11.02 86.69±8.46 86.18±8.43 88.43±8.2
[0131] Table 15. Performance comparison between the present invention and 11 models based on the LCSXH-CSU dataset for the KRAS functional domain prediction task.
[0132] AUCROC ACC Precision Recall F1-Score ABMIL 94.11±4.48 86.41±10.69 96.99±1.42 98.67±0.73 97.5±1 CLAM-SB 93.62±5.71 86.45±12.12 94.94±4.53 91.92±7.22 94.94±4.53 CLAM-MB 82.59±15.58 81.94±16.16 83.51±14.75 98.71±1.15 81.94±16.16 DSMIL 89.14±9.72 87.74±10.96 92.97±6.29 91.16±7.9 95.48±4.04 TransMIL 85.69±12..8 81.29±16.73 80.27±17.64 80.15±17.75 81.29±16.73 DTFD-MIL 89.31±9.56 87.74±10.96 89.82±9.11 98.83±1.04 87.74±10.96 IBMIL 89.64±8.87 87.72±10.25 89.33±8.81 91.71±6.69 87.72±10.25 HIPT 93.64±0.32 93.33±0.6 93.39±0.54 93.56±0.4 93.33±0.6 MHIM 48.46±4.81 43.69±9.49 61.07±11.07 91.77±0.88 55.07±9.75 RRT_MIL 60.51±7.7 64.3±7.52 93.97±1.27 77.13±5.09 69.01±6.65 Patch-GCN 80.61±8.6 90.71±2.57 92.27±2.03 95.21±1.34 90.71±2.57 The present invention 94.69±4.53 97.86±1.92 98.05±1.74 98.3±1.52 97.86±1.92
[0133] Table 16. Performance comparison between the present invention and 11 models based on the LCSXH-CSU dataset for the ALK functional domain prediction task.
[0134] AUCROC ACC Precision Recall F1-Score ABMIL 91.39±7.1 95±3.46 95.67±2.7 97.21±1.84 95±3.46 CLAM-SB 94.67±2.98 95±4.47 96.59±3.05 96.19±3.41 97±2.68 CLAM-MB 92.5±6.71 89±9.84 90±8.94 91.25±7.83 89±9.84 DSMIL 94.17±5.22 90±8.94 95±4.47 92.86±6.39 98±1.79 TransMIL 93.19±6.09 91±8.05 91.71±7.42 91.43±7.67 92±7.16 DTFD-MIL 93.61±5.71 98±1.79 98.51±1.33 98.05±1.74 99±0.89 IBMIL 94.44±4.97 92±7.16 92±7.16 92±7.16 92±7.16 HIPT 93.89±5.47 90±8.94 92.9±6.35 98.18±1.63 90±8.94 MHIM 85.6±4.54 87.07±5.55 92.16±2.74 96.68±1.4 90.4±3.54 RRT_MIL 87.14±4.56 86.12±4.69 90.59±3.61 95.41±1.07 89.23±5.04 Patch-GCN 94.67±2.98 94±3.58 94±3.58 94±3.58 94±3.58 The present invention 94.72±4.72 99±0.89 99±0.89 99±0.89 99±0.89
[0135] The NAVF-Bio model of the present invention extracts multi-view information from whole slide images (WSIs) by adopting multi-scale feature fusion and an adaptive cross-view knowledge supplementation module to predict various gene mutation information, including gene mutation subtype and gene mutation exon prediction. The multi-scale feature fusion simulates the steps of a pathologist's reading of slides for feature extraction and fusion, and the adaptive cross-view knowledge supplementation module can flexibly incorporate multi-view features to improve the prediction performance of the model. In the task of predicting the mutations of key genes TP53, EGFR, KRAS, and ALK in lung cancer, NAVF-Bio achieved an AUCROC value of over 90 in the five-fold cross-validation of the dataset of the Second Xiangya Hospital of Central South University (the AUCROC values were 92.93±6.27, 92.32±6.69, 90.6±7.74, and 92.58±7.32 respectively), which is better than the currently available state-of-the-art methods. Moreover, the NAVF-Bio prediction for this task can also be generalized to the TCGA_LUAD dataset. More importantly, in predicting the subtypes of TP53 mutations, NAVF-Bio obtained AUCROC values of 90.72±7.49 and 87.72±7.49 on the datasets of the Second Xiangya Hospital of Central South University and the TCGA_LUAD dataset respectively, and for the first time predicted the exons of TP53, EGFR, KRAS, and ALK mutations and the tumor mutation burden (TMB), reaching the performance level of clinical grade. The NAVF-Bio model solves the technical problem that the current AI models cannot effectively predict gene mutation subtypes and mutation exon localization.
[0136] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for predicting gene mutation information from lung cancer tissue pathology images, characterized in that: The following steps are involved: S1. Collect histopathological whole slide images and corresponding gene testing reports of lung cancer patients; S2, preprocess the whole slide images of histopathology and construct the pathology dataset; S3. Construct a NAVF-Bio model for predicting gene mutation information from lung cancer tissue pathology whole slide images; the NAVF-Bio model includes a pre-training module, a pathology space topology representation learning module, a multi-scale feature fusion module, and an adaptive cross-view knowledge supplementation module; The pre-training module is used to extract features from the pre-processed images; the pathological space topology representation learning module is used to further process the extracted features to form a one-dimensional vector; The multi-scale feature fusion module is used to further extract and fuse features of different scales; the adaptive cross-view knowledge supplement module is used to incorporate features from multiple views to improve the prediction performance of the model; S4. Using the pathological omics dataset to train the constructed NAVF-Bio model, constraining the NAVF-Bio model based on multi-scale loss and weighted fusion loss, to obtain the final NAVF-Bio model; S5. The final NAVF-Bio model is used to predict gene mutation information from whole slide images of lung cancer tissue pathology, wherein the gene mutation information includes gene mutation, tumor mutation load, gene mutation subtype and protein functional domain.
2. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: In the step S2, preprocessing the histopathology whole slide image includes digitizing the histopathology whole slide image, segmenting tissue regions, and detecting background and blurred regions.
3. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: In step S3, the pre-training module uses the Otsu algorithm to distinguish the background area of the whole slide image, and uses the sliding window strategy to divide the whole slide image into large, medium and small scale patches, and uses the CTransPath pre-training model to extract a feature representation with a dimension of 768 for each patch.
4. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 3, characterized in that: The pathological spatial topology map represents that in the learning module, the KNN algorithm is used to construct the points and edges of the spatial topology map of patches of different scales; the HD-Yolo algorithm is used to segment and classify the cells in the whole slide image, and the tumor microenvironment indicators are quantified. The KNN algorithm is also used to construct the cell spatial topology map of the tumor microenvironment. The nodes of the cell spatial topology map include the cell label, the probability of classification, the area of the cell nucleus, and the relationship between the same cells is used as the edge.
5. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: The pathological space topology representation learning module includes an SGAEConv module and a multi-layer perception module; The SGAEConv module uses a linear transformation to combine the node's own features and neighbor features, expressed as: Among them, W m ,W n are all learnable weight matrices, Is a node v In the k-th layer of embedding, is the aggregated feature of neighbor nodes, σ is the RuLU activation function; The multi-layer perception module includes ReLU activation function, layer normalization and information discarding regularization; ReLU activation function is used to perform nonlinear transformation on input features; layer normalization is used to stabilize model training and accelerate model convergence; the multi-layer perception module uses information discarding regularization to discard a part of neurons with a certain probability in each training, thereby helping the model to be better trained and preventing overfitting.
6. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: The multi-scale feature fusion module includes an adaptive pooling layer, a splicing layer and a self-attention feature importance extractor; The adaptive pooling layer is used to dynamically adjust the length of the input sequence so that the output length matches the given target size; the concatenation layer is used to concatenate in a unified feature space; the self-attention feature importance extractor uses the attention mechanism to obtain features that are relatively important to the model based on feature importance.
7. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: The adaptive cross-view knowledge supplementation module includes an adaptive feature fusion module, an adaptive weight module and a classifier; The adaptive feature fusion module uses the attention mechanism to process the input data and generate a global representation of the input features; The adaptive weight module includes two fully connected layers. The adaptive weight module is used to perform feature scaling to form adaptive weights; the classifier is used to obtain the representation of each label.
8. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: The multi-scale loss is optimized at different scales respectively; the weighted fusion loss is used to integrate losses from different scales, and reflects the importance of losses of different scales to the overall task in a weighted manner.
9. The method for predicting gene mutation information from lung cancer tissue pathology images according to claim 1, characterized in that: Predicted gene mutation information includes: Predict whether there is TP53 gene mutation, EGFR mutation, KRAS mutation, or ALK mutation; Predicting the level of tumor mutation burden; Predict TP53 gene mutation subtypes; Predict TP53, EGFR, KRAS, and ALK mutant exons.
Citation Information
Patent Citations
Method for predicting intratracheal dissemination from lung cancer histopathologic image
CN118507062A
PET-CT lung cancer image segmentation method and system based on deep learning
CN119380027A
Generalizable and Interpretable Deep Learning Framework for Predicting MSI from Histopathology Slide Images
US20190347557A1
Pulmonary nodule automatic detection method, apparatus and computer system
WO2022063199A1