Model and method for predicting lung adenocarcinoma gene mutation based on HE pathological image

By constructing a deep neural network model based on the VisionTransformer architecture, combining self-supervised learning and multi-instance learning, the problem of extracting mutation information of lung adenocarcinoma genes from HE pathological images is solved, and accurate prediction of mutations in key genes such as EGFR, KRAS, and TP53 is achieved, providing an efficient alternative.

CN119942212APending Publication Date: 2025-05-06THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510080083.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-19
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract information related to lung adenocarcinoma gene mutations from HE pathological images, and faces problems such as color differences and background noise interference, resulting in insufficient accuracy and stability in predicting mutations of key genes such as EGFR, KRAS, and TP53.

Method used

A deep neural network model was constructed using DINO self-supervised learning algorithm based on VisionTransformer architecture. Through pre-training and fine-tuning, common features were learned from unlabeled WSIs, and fine-tuned on the annotated data to capture features related to lung oncogene mutations. Combining multi-instance learning, instance-level clustering and pseudo-label supervised learning, the feature space is optimized to improve prediction accuracy.

Benefits of technology

Accurate prediction of key gene mutations such as EGFR, KRAS, TP53 in HE pathological images is achieved, providing a fast, economical and effective alternative, overcoming the time-consuming and cost-effective limitations of traditional gene sequencing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942212A_ABST
    Figure CN119942212A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of medical auxiliary diagnosis, and particularly relates to a model and method for predicting lung adenocarcinoma gene mutation based on an HE pathological image, and the diagnosis method comprises the following steps: data collection: collecting Hamp from TCGA, CPTAC and GTEx public databases; e, dyeing the digital pathological full-section image, constructing a pathological section pre-training data set, and collecting Hamp from a plurality of hospital centers; the method comprises the following steps: establishing a lung cancer-gene mutation data set, preprocessing, establishing and training an instance feature extraction model, establishing and predicting an instance feature aggregation model, and verifying and evaluating the model according to the mutation condition of hematoxylin-eosin (Hematoxylin-Eosin, Hamp; the method comprises the following steps: analyzing lung cancer pathological full-slice images (WSIs, E) dyed by using a sample to effectively predict mutation of key genes such as EGFR (epidermal growth factor receptor), KRAS, TP53 and the like, so as to effectively predict mutation of the key genes such as EGFR, KRAS, TP53 and the like. The invention provides a quicker, more economical and more effective alternative scheme, and can effectively overcome the limitations of long time consumption and high cost of the existing gene sequencing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical auxiliary diagnosis, and specifically relates to a model and method for predicting gene mutation of lung adenocarcinoma based on HE pathological images. Background Art

[0002] In the current field of lung adenocarcinoma gene mutation prediction, traditional methods face many challenges. The diagnosis and treatment of lung adenocarcinoma are highly dependent on the accurate understanding of gene mutations, and accurate acquisition of gene mutation information usually requires complex and expensive genetic testing technology.

[0003] Pathological images are an important basis for reflecting the disease status, among which HE pathological images are widely used. However, it has always been a difficult problem to extract effective information related to gene mutations in lung adenocarcinoma from HE pathological images. HE pathological images from different sources have problems such as color differences and background noise interference, making it difficult to analyze gene mutations directly from images.

[0004] Previous models had difficulty in extracting features comprehensively and accurately when processing such images. For example, traditional neural network models lacked generalization capabilities when faced with complex pathological image features, and could not effectively cope with diverse image data collected by different hospitals and different equipment. At the same time, a large amount of unlabeled pathological image data was not fully utilized, and the value of the data was not fully mined, which limited the performance and accuracy of the model. In addition, when predicting mutations in key genes such as EGFR, KRAS, and TP53, the accuracy and stability of existing models could not meet clinical needs, and could not provide strong support for the accurate diagnosis and personalized treatment of lung adenocarcinoma.

[0005] To this end, the present invention provides a model and method for predicting lung adenocarcinoma gene mutations based on HE pathological images. Summary of the invention

[0006] In order to make up for the deficiencies of the prior art, at least one technical problem raised in the background technology is solved.

[0007] The technical solution adopted by the present invention to solve the technical problem is: the model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images described in the present invention, the diagnostic method comprises the following steps:

[0008] S1. Data collection: collect H&E stained digital pathology whole slide images from TCGA, CPTAC, and GTEx public databases to build a pathology slide pre-training dataset; collect H&E stained lung cancer WSIs and their key gene mutations from multiple hospital centers to build a lung cancer-gene mutation dataset;

[0009] S2, preprocessing, standardize all WSIs, normalize colors, remove background information with saturation <15, and cut them into patches of uniform size of 256×256 pixels;

[0010] S3. Instance feature extraction model construction and training: A deep neural network model for instance feature extraction is constructed. The model uses the DINO self-supervised learning algorithm based on the VisionTransformer architecture and is pre-trained on a pathology slice pre-training dataset, so that the model can learn general feature representations of pathology images from unlabeled WSIs.

[0011] Afterwards, the model is fine-tuned on the lung cancer gene mutation dataset so that the model can further learn feature representations that are more relevant to the task of predicting lung cancer gene mutations from each patch;

[0012] S4, instance feature aggregation model construction and prediction, after obtaining the feature representation of all patches, a deep neural network model for instance feature aggregation is constructed for WSI category prediction;

[0013] First, we introduce the Multiple Instance Learning (MIL) algorithm based on the attention mechanism, calculate the attention weight of each patch and multiply it with the patch to learn the contribution of each patch to the entire WSI category;

[0014] Next, instance-level clustering is used to constrain and optimize the feature space, while pseudo labels are generated for each patch for supervised learning and optimized using a smoothed SVM loss function;

[0015] Finally, the classification of patches is aggregated using traditional aggregation functions to obtain the classification of the entire pathological image. Based on the final binary classification probability prediction results, it is determined whether mutations in specific genes such as EGFR, KRAS, and TP53 occur.

[0016] S5. Model validation and evaluation. The lung cancer-gene mutation dataset was randomly divided into a training set, a parallel set, and a test set in a ratio of 8:1:1, and validated by a 10-fold Monte Carlo cross-validation method. The main evaluation indicators included accuracy, sensitivity, specificity, and AUC (area under the curve).

[0017] Preferably, in the S1 data collection step, for the H&E-stained lung cancer WSIs and their key gene mutations collected from multiple hospital centers, the data is further cleaned to remove data samples containing blurred images, missing key information, or incorrect gene mutation annotations, so as to ensure the quality of the constructed lung cancer-gene mutation dataset.

[0018] Preferably, in the S2 preprocessing step, the color normalization process adopts a color standardization technology based on the Macenko method, which can more accurately unify the color differences between different images by analyzing and adjusting the color components of the image.

[0019] Preferably, in the S2 preprocessing step, after removing the background information with a saturation <15, a morphological opening operation is performed on the remaining image area to further eliminate possible tiny noise interference and make the pathological tissue area clearer.

[0020] Preferably, in the S3 instance feature extraction model construction and training steps, the deep neural network model based on the VisionTransformer architecture includes multiple multi-head self-attention modules, and the number of heads of each module is set to 8 to enhance the model's ability to extract different feature dimensions of the image.

[0021] Preferably, during the pre-training process of the S3 instance feature extraction model, a contrastive learning strategy is adopted, and the same image after different transformations is used as a positive sample pair, and different images are used as a negative sample pair to enhance the model's ability to distinguish pathological image features.

[0022] Preferably, during the fine-tuning stage of the S3 instance feature extraction model, a learning rate decay strategy is adopted, with the initial learning rate set to 0.001, and the learning rate decays to the original 0.9 after every 5 training rounds to ensure that the model can converge to the optimal solution more stably during the fine-tuning process.

[0023] Preferably, in the S4 instance feature aggregation model construction and prediction step, the multi-instance learning (MIL) algorithm based on the attention mechanism adopts a multi-head attention mechanism with 4 heads, and calculates the attention weight of each patch from different angles, further improving the evaluation accuracy of the contribution degree of each patch.

[0024] Preferably, in the S4 instance feature aggregation model construction and prediction step, the instance-level clustering adopts the DBSCAN algorithm, and the feature space is clustered by reasonably setting the neighborhood radius and the minimum number of samples to explore the close connection between the features.

[0025] Preferably, in the S4 instance feature aggregation model construction and prediction step, the process of generating pseudo labels is based on the prediction confidence of the model in the current training stage. For prediction results with confidence higher than a set threshold (such as 0.8), they are used as pseudo labels for supervised learning to increase effective supervision information.

[0026] Preferably, in the S5 model validation and evaluation step, in addition to using the 10-fold Monte Carlo cross-validation method, a leave-one-out cross-validation method is also introduced for comparative validation to more comprehensively evaluate the performance stability of the model.

[0027] The beneficial effects of the present invention are as follows:

[0028] The model and method for predicting gene mutations in lung adenocarcinoma based on HE pathological images described in the present invention effectively predicts key gene mutations such as EGFR, KRAS, and TP53 by analyzing whole slide images (WSIs) of lung cancer pathology stained with hematoxylin-eosin (H&E). It provides a faster, more economical and more effective alternative that can effectively overcome the time-consuming and high-cost limitations of existing gene sequencing technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention will be further described below in conjunction with the accompanying drawings.

[0030] Figure 1 It is a flow chart of the method in the present invention. DETAILED DESCRIPTION

[0031] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0032] like Figure 1 As shown, the model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images in an embodiment of the present invention, the diagnostic method comprises the following steps:

[0033] S1. Data collection: collect H&E stained digital pathology whole slide images from TCGA, CPTAC, and GTEx public databases to build a pathology slide pre-training dataset; collect H&E stained lung cancer WSIs and their key gene mutations from multiple hospital centers to build a lung cancer-gene mutation dataset;

[0034] S2, preprocessing, standardize all WSIs, normalize colors, remove background information with saturation <15, and cut them into patches of uniform size of 256×256 pixels;

[0035] S3. Instance feature extraction model construction and training: A deep neural network model for instance feature extraction is constructed. The model uses the DINO self-supervised learning algorithm based on the VisionTransformer architecture and is pre-trained on a pathology slice pre-training dataset, so that the model can learn general feature representations of pathology images from unlabeled WSIs.

[0036] Afterwards, the model is fine-tuned on the lung cancer gene mutation dataset so that the model can further learn feature representations that are more relevant to the task of predicting lung cancer gene mutations from each patch;

[0037] S4, instance feature aggregation model construction and prediction, after obtaining the feature representation of all patches, a deep neural network model for instance feature aggregation is constructed for WSI category prediction;

[0038] First, we introduce the Multiple Instance Learning (MIL) algorithm based on the attention mechanism, calculate the attention weight of each patch and multiply it with the patch to learn the contribution of each patch to the entire WSI category;

[0039] Next, instance-level clustering is used to constrain and optimize the feature space, while pseudo labels are generated for each patch for supervised learning and optimized using a smoothed SVM loss function;

[0040] Finally, the classification of patches is aggregated using traditional aggregation functions to obtain the classification of the entire pathological image. Based on the final binary classification probability prediction results, it is determined whether mutations in specific genes such as EGFR, KRAS, and TP53 occur.

[0041] S5. Model validation and evaluation. The lung cancer-gene mutation dataset was randomly divided into a training set, a parallel set, and a test set in a ratio of 8:1:1, and validated by a 10-fold Monte Carlo cross-validation method. The main evaluation indicators included accuracy, sensitivity, specificity, and AUC (area under the curve).

[0042] In the S1 data collection step, the H&E-stained lung cancer WSIs and their key gene mutations collected from multiple hospital centers were further cleaned to remove data samples with blurred images, missing key information, or incorrect gene mutation annotations to ensure the quality of the constructed lung cancer-gene mutation dataset.

[0043] In the S2 preprocessing step, the color normalization process adopts the color standardization technology based on the Macenko method. This technology can more accurately unify the color differences between different images by analyzing and adjusting the color components of the image.

[0044] In the S2 preprocessing step, after removing the background information with saturation <15, the remaining image area is subjected to morphological opening operation to further eliminate possible minor noise interference and make the pathological tissue area clearer.

[0045] In the S3 instance feature extraction model construction and training steps, the deep neural network model based on the VisionTransformer architecture contains multiple multi-head self-attention modules, and the number of heads in each module is set to 8 to enhance the model's ability to extract different feature dimensions of the image.

[0046] During the pre-training process of the S3 instance feature extraction model, a contrastive learning strategy was adopted, and the same image after different transformations was used as a positive sample pair, and different images were used as negative sample pairs to enhance the model's ability to distinguish pathological image features.

[0047] During the fine-tuning stage of the S3 instance feature extraction model, a learning rate decay strategy was adopted. The initial learning rate was set to 0.001, and the learning rate decayed to the original 0.9 after every 5 training rounds to ensure that the model could converge to the optimal solution more stably during the fine-tuning process.

[0048] In the S4 instance feature aggregation model construction and prediction step, the multi-instance learning (MIL) algorithm based on the attention mechanism adopts a multi-head attention mechanism with 4 heads to calculate the attention weight of each patch from different angles, further improving the evaluation accuracy of the contribution degree of each patch.

[0049] In the S4 instance feature aggregation model construction and prediction step, the instance-level clustering uses the DBSCAN algorithm to cluster the feature space and explore the close connections between features by reasonably setting the neighborhood radius and the minimum number of samples.

[0050] In the S4 instance feature aggregation model construction and prediction step, the process of generating pseudo labels is based on the prediction confidence of the model in the current training stage. For prediction results with confidence higher than the set threshold (such as 0.8), they are used as pseudo labels for supervised learning to increase effective supervision information.

[0051] In the S5 model validation and evaluation step, in addition to the 10-fold Monte Carlo cross-validation method, the leave-one-out cross-validation method was also introduced for comparative validation to more comprehensively evaluate the performance stability of the model.

[0052] Working principle:

[0053] This method is mainly based on a large amount of H&E-stained digital pathology whole slide images (WSIs) data. First, relevant images and corresponding gene mutations are collected from public databases and multiple hospital centers to construct a corresponding data set.

[0054] In the preprocessing stage, WSIs are color normalized to eliminate image color differences caused by different sources, ensuring unified data standards; background information with a saturation less than 15 is removed to reduce irrelevant "noise" and highlight key areas of pathological tissue; and the image is cut into patches of fixed size to facilitate subsequent efficient and accurate feature extraction and analysis.

[0055] In terms of feature extraction, a deep neural network model based on the VisionTransformer architecture and the DINO self-supervised learning algorithm is used to first self-supervise the learning of common features of pathological images from a large number of unlabeled WSIs, and then fine-tune the model based on data annotated with gene mutations, focusing on capturing subtle features related to the prediction of lung cancer gene mutations, thereby achieving effective learning from general to specific task features.

[0056] When aggregating features, the multi-instance learning (MIL) algorithm based on the attention mechanism is introduced to calculate the attention weight for each patch, highlight the important local area and measure its contribution to the overall image category. Then, instance-level clustering, pseudo-label generation and smoothed SVM loss function are combined to optimize the feature space and reduce redundancy and noise. Finally, the classification of patches is integrated through the aggregation function to form a comprehensive judgment of the entire pathological image, so as to determine whether a specific gene is mutated.

[0057] In the model verification and evaluation phase, the collected data set is divided into training set, validation set and test set in proportion. The 10-fold Monte Carlo cross-validation method is used to measure the model performance in many aspects with the help of indicators such as accuracy, sensitivity, specificity, AUC, etc., to ensure that the model can be reliably and effectively used for intelligent diagnosis of digital pathology images. The whole process simulates the thinking of pathologists analyzing slices, and the purpose of intelligent diagnosis is achieved through the coordinated cooperation of various links.

[0058] First, it demonstrates high efficiency in data utilization. On the one hand, it widely collects H&E-stained digital pathology full-slice images from public databases such as TCGA, CPTAC, GTEx, and multiple hospital centers, integrates multi-source data, and greatly enriches the diversity of pathology image samples. In this way, the model is able to learn the pathological characteristics in different scenarios, and then has excellent generalization capabilities, and can accurately diagnose images regardless of which hospital or device is used to collect the images. On the other hand, with the help of weakly supervised learning models, partial labeled data is cleverly used in collaborative training with a large amount of unlabeled data. It not only reduces the reliance on massive, expensive and time-consuming labeled data, but also fully taps the potential of the data and reduces the cost of data preparation.

[0059] Second, the preprocessing stage has achieved remarkable results. Color normalization is like a "calibration ruler", which unifies the image colors that vary due to different staining conditions and scanning equipment, so that the model can completely get rid of the interference caused by color differences and focus on the essential characteristics of pathological tissues. Removing background information with a saturation of less than 15 is like clearing the battlefield and removing irrelevant "debris", allowing the model to have a keen eye and directly hit the key pathological areas, greatly improving the efficiency and accuracy of feature extraction. In addition, cutting the full slice image into blocks of uniform size is like breaking the whole into parts, which reduces the computing cost while facilitating the accurate capture of local lesion characteristics. The subsequent aggregation operation can perfectly realize the pathological judgment from local to overall.

[0060] Third, feature extraction and learning are accurate and correct. The deep neural network model built based on the VisionTransformer architecture and the DINO self-supervised learning algorithm plays an important role. VisionTransformer performs well in the field of images, and the DINO self-supervised learning helps the model first gain insights into the common features of pathological images from massive unlabeled data, just like building a solid foundation; then fine-tuning on the lung cancer-gene mutation dataset, accurately locking in subtle features closely related to lung cancer gene mutations, achieving a gorgeous transformation from general to special, and deeply exploring the intrinsic relationship between cell morphology, tissue structure changes and gene mutations, laying a solid foundation for diagnosis.

[0061] Fourth, feature aggregation is reasonable and orderly. A multi-instance learning algorithm based on the attention mechanism is introduced to accurately "weight" each tile. The model is like an experienced pathologist, which can keenly focus on the local area with the most diagnostic value and scientifically measure the contribution of the tile to the overall image category. At the same time, optimization strategies such as instance-level clustering, pseudo-label generation, and the use of smoothed SVM loss function are used in a multi-pronged manner. Clustering mines feature connections, pseudo-labels provide additional supervision, and smoothed SVM optimizes classification boundaries, which jointly eliminate feature space noise and redundancy to ensure robust and reliable diagnosis.

[0062] Fifth, the model validation and evaluation are scientific and rigorous. The data set is finely divided into training set, validation set and test set in a ratio of 8:1:1, and the 10-fold Monte Carlo cross-validation method is used to comprehensively test the model performance. It not only avoids overfitting misjudgment, but also reduces the accidental impact of data division with the help of multiple cross-validations. In addition, a comprehensive evaluation of multiple indicators such as accuracy, sensitivity, specificity, AUC, etc. is used to consider the diagnostic effect of the model from different dimensions, accurately diagnose the advantages and disadvantages of the model, and point out the direction for subsequent optimization.

[0063] It should be noted that in the prediction task of mutant genes related to lung adenocarcinoma, the TCGA-LUAD dataset was used to perform model training and cross-validation operations. In terms of specific gene selection, the three most relevant genes, EGFR, KRAS, and TP53, were selected for mutation prediction tasks. For these genes, the mutations in the patient population are as follows: in the EGFR gene, the ratio of the number of patients with mutations to the number of patients without mutations is 70:452; for the KRAS gene, the corresponding ratio is 139:383; and for the TP53 gene, the ratio is 245:277. The constructed model was evaluated using the ten-fold cross-validation method. In the prediction tasks of the above three genes, the AUC (area under the curve) scores were 0.688, 0.568, and 0.706, respectively.

[0064] In order to further verify the effectiveness of the method, the EGFR and KRAS genes that performed relatively well in the previous ten-fold cross-validation were selected to carry out zero-sample validation on the clinical data collected by the hospital. These clinical data contain information on 65 patients, totaling 122 H&E-stained digital pathology whole-slice images (WSI). Finally, in the zero-sample validation process, the AUC score obtained for the prediction results of the EGFR gene was 0.558, while the corresponding score for the KRAS gene was 0.593. The above results show that the digital pathology image intelligent diagnosis method we proposed has certain feasibility and application potential in the task of predicting gene mutations related to lung adenocarcinoma, and provides a valuable reference for further research and clinical application. Detailed statistics were conducted on the mutations of these three genes in the patient population, and the specific data are shown in the following table:

[0065]

[0066]

[0067] Through the display of these data, we can not only clearly see the prediction effects of different genes at different verification stages, but also verify the effectiveness of the method from multiple dimensions. At the same time, when dealing with the prediction of gene mutations related to lung adenocarcinoma, the prediction performance of different genes has certain differences, which also provides a direction for improvement and optimization for subsequent research, prompting us to explore more deeply how to further improve the performance of the model in different gene prediction tasks, so that it can play a more stable and reliable role in clinical applications.

[0068] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A model and method for predicting gene mutations in lung adenocarcinoma based on HE pathological images, characterized by: The diagnostic method comprises the following steps: S1. Data collection: collect H&E stained digital pathology whole slide images from TCGA, CPTAC, and GTEx public databases to build a pathology slide pre-training dataset; collect H&E stained lung cancer WSIs and their key gene mutations from multiple hospital centers to build a lung cancer-gene mutation dataset; S2, preprocessing, standardize all WSIs, normalize colors, remove background information with saturation <15, and cut them into patches of uniform size of 256×256 pixels; S3. Instance feature extraction model construction and training: A deep neural network model for instance feature extraction is constructed. The model uses the DINO self-supervised learning algorithm based on the VisionTransformer architecture and is pre-trained on a pathology slice pre-training dataset, so that the model can learn general feature representations of pathology images from unlabeled WSIs. Afterwards, the model is fine-tuned on the lung cancer gene mutation dataset so that the model can further learn feature representations that are more relevant to the task of predicting lung cancer gene mutations from each patch; S4, instance feature aggregation model construction and prediction, after obtaining the feature representation of all patches, a deep neural network model for instance feature aggregation is constructed for WSI category prediction; First, we introduce the Multiple Instance Learning (MIL) algorithm based on the attention mechanism, calculate the attention weight of each patch and multiply it with the patch to learn the contribution of each patch to the entire WSI category; Next, instance-level clustering is used to constrain and optimize the feature space, while pseudo labels are generated for each patch for supervised learning and optimized using a smoothed SVM loss function; Finally, the classification of patches is aggregated using traditional aggregation functions to obtain the classification of the entire pathological image. Based on the final binary classification probability prediction results, it is determined whether mutations in specific genes such as EGFR, KRAS, and TP53 occur. S5. Model validation and evaluation. The lung cancer-gene mutation dataset was randomly divided into a training set, a parallel set, and a test set in a ratio of 8:1:1, and validated by a 10-fold Monte Carlo cross-validation method. The main evaluation indicators included accuracy, sensitivity, specificity, and AUC (area under the curve).

2. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: In the S1 data collection step, the H&E-stained lung cancer WSIs and their key gene mutations collected from multiple hospital centers were further cleaned to remove data samples with blurred images, missing key information, or incorrect gene mutation annotations to ensure the quality of the constructed lung cancer-gene mutation dataset.

3. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: In the S2 preprocessing step, the color normalization process adopts the color standardization technology based on the Macenko method. This technology can more accurately unify the color differences between different images by analyzing and adjusting the color components of the image.

4. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: In the S2 preprocessing step, after removing the background information with saturation <15, the remaining image area is subjected to morphological opening operation to further eliminate possible minor noise interference and make the pathological tissue area clearer.

5. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: In the S3 instance feature extraction model construction and training steps, the deep neural network model based on the VisionTransformer architecture contains multiple multi-head self-attention modules, and the number of heads in each module is set to 8 to enhance the model's ability to extract different feature dimensions of the image.

6. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: During the pre-training process of the S3 instance feature extraction model, a contrastive learning strategy was adopted, and the same image after different transformations was used as a positive sample pair, and different images were used as negative sample pairs to enhance the model's ability to distinguish pathological image features.

7. The model and method for predicting gene mutations in lung adenocarcinoma based on HE pathological images according to claim 1, characterized in that: During the fine-tuning stage of the S3 instance feature extraction model, a learning rate decay strategy was adopted. The initial learning rate was set to 0.001, and the learning rate decayed to the original 0.9 after every 5 training rounds to ensure that the model could converge to the optimal solution more stably during the fine-tuning process.

8. The model and method for predicting lung adenocarcinoma gene mutation based on HE pathological images according to claim 1, characterized in that: In the S4 instance feature aggregation model construction and prediction step, the multi-instance learning (MIL) algorithm based on the attention mechanism adopts a multi-head attention mechanism with 4 heads to calculate the attention weight of each patch from different angles, further improving the evaluation accuracy of the contribution degree of each patch.

9. The model and method for predicting gene mutations in lung adenocarcinoma based on HE pathological images according to claim 1, characterized in that: In the S4 instance feature aggregation model construction and prediction step, the instance-level clustering uses the DBSCAN algorithm to cluster the feature space and explore the close connections between features by reasonably setting the neighborhood radius and the minimum number of samples.

10. The model and method for predicting gene mutations in lung adenocarcinoma based on HE pathological images according to claim 1, characterized in that: In the S4 instance feature aggregation model construction and prediction step, the process of generating pseudo labels is based on the prediction confidence of the model in the current training stage. For prediction results with confidence higher than the set threshold (such as 0.8), they are used as pseudo labels for supervised learning to increase effective supervision information.