Lung adenocarcinoma gene mutation prediction method, device, equipment, medium and product

By identifying tumor areas and extracting features from lung slice images, and using the DINO model and a two-stage multi-instance classification model, the time-consuming and costly problem of traditional lung adenocarcinoma gene testing was solved, and fast and accurate gene mutation prediction was achieved.

CN120673843APending Publication Date: 2025-09-19CHINA JAPAN FRIENDSHIP HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510756700.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional gene detection methods for lung adenocarcinoma are time-consuming and expensive, resulting in wasted treatment time and economic burden, making them difficult to be widely used in clinical practice.

Method used

A method based on the DINO model and a two-stage multi-instance classification model was used to identify, segment, and extract features of tumor regions in lung slice images, and machine learning technology was used to quickly predict the mutation status of genes such as EGFR, KRAS, and ALK.

Benefits of technology

It improves the accuracy and efficiency of gene mutation prediction in lung adenocarcinoma, reduces detection time and cost, and is suitable for rapid gene mutation identification in clinical settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673843A_ABST
    Figure CN120673843A_ABST
Patent Text Reader

Abstract

The invention discloses a lung adenocarcinoma gene mutation prediction method and device, equipment, a medium and a product, and relates to the field of image recognition and machine learning, and the method comprises the steps: obtaining a lung full-slice image of a user; performing tumor region identification on the lung slice image to obtain a region-of-interest image; segmenting the region-of-interest image to obtain a plurality of image blocks, and randomly combining the image blocks to obtain a plurality of slice packets; inputting the slice packets into the trained DINO model, and extracting packet features corresponding to the slice packets; inputting the packet features corresponding to all the slice packets into a two-stage multi-instance classification model to obtain a lung adenocarcinoma gene mutation prediction result; the two-stage multi-instance classification model comprises a plurality of single-stage multi-instance models and a two-stage multi-instance model. The accuracy and efficiency of lung adenocarcinoma gene mutation prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of image recognition and machine learning, and in particular to a method, apparatus, device, medium and product for predicting gene mutations in lung adenocarcinoma. Background Art

[0002] Lung cancer is the leading cause of malignancy-related mortality. Lung adenocarcinoma (LUAD) is the most common subtype of non-small cell lung cancer (NSCLC). The development of LUAD is influenced by both genetic factors (internal influences) and environmental factors (external influences). Therefore, considering genetic information, particularly tumor gene mutations, is crucial in the study and management of LUAD. Targeted therapies, targeting specific key molecular or cellular mechanisms, have significantly improved patient survival. Identifying a patient's genetic mutations is crucial for guiding targeted therapy. Several key targets have been identified, such as EGFR, HER-2, KRAS, and ALK. Most of these genes can activate related downstream pathways, promoting tumor cell proliferation, survival, and other functions. For example, EGFR mutations and overexpression can activate multiple downstream signaling pathways, such as the MAPK, PI3K / AKT, PKC, and STAT pathways. KRAS mutations can trigger aberrant activation of downstream signaling pathways, such as RAF / MEK / ERK and PI3K / AKT / mTOR. PIK3CA mutations upregulate the PI3K / AKT signaling pathway. These mutations are highly prevalent in lung adenocarcinoma, with most patients exhibiting abnormalities in one or more genes. Targeted therapies targeting these genes have the potential to significantly improve patient survival. Therefore, mutations in these genes play a key role in determining treatment options and patient prognosis. Related targeted drugs are either approved for clinical use or are under development.

[0003] Traditional genetic testing methods, such as immunohistochemistry (IHC) and next-generation sequencing, are time-consuming and expensive. Specifically, patients often have to wait a week or more to receive genetic test results, resulting in precious treatment time being wasted. In addition, the high cost imposes a considerable financial burden on patients. These factors together make these methods inefficient in routine clinical practice. Therefore, their widespread application in clinical practice is hindered. Therefore, there is an urgent need for a simple method that can quickly and effectively assess the status of target genes. Summary of the Invention

[0004] The purpose of this application is to provide a method, device, equipment, medium and product for predicting lung adenocarcinoma gene mutations, which can improve the accuracy and efficiency of lung adenocarcinoma gene mutation prediction.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for predicting gene mutations in lung adenocarcinoma, comprising:

[0007] Obtain a full-slice image of the user's lungs;

[0008] Performing tumor region recognition on the lung slice image to obtain a region of interest image;

[0009] Segmenting the image of the region of interest to obtain a number of image blocks, and randomly combining the image blocks to obtain a number of slice packages;

[0010] For each of the slice packages, the slice package is input into the trained DINO model to extract the package features corresponding to the slice package;

[0011] The package features corresponding to all the slice packages are input into the two-stage multi-instance classification model to obtain the lung adenocarcinoma gene mutation prediction results; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

[0012] In a second aspect, the present application provides a device for predicting gene mutations in lung adenocarcinoma, comprising:

[0013] A lung full-slice image acquisition module is used to acquire a user's lung full-slice image;

[0014] a tumor region recognition module, configured to perform tumor region recognition on the lung slice image to obtain a region of interest image;

[0015] An image segmentation and combination module is used to segment the image of the region of interest to obtain a number of image blocks, and randomly combine the image blocks to obtain a number of slice packages;

[0016] A packet feature extraction module is used to input each slice packet into the trained DINO model to extract the packet features corresponding to the slice packet;

[0017] A lung adenocarcinoma gene mutation prediction module is used to input the package features corresponding to all the slice packages into a two-stage multi-instance classification model to obtain lung adenocarcinoma gene mutation prediction results; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract features from the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

[0018] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for predicting gene mutations in lung adenocarcinoma.

[0019] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for predicting gene mutations in lung adenocarcinoma.

[0020] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the above-mentioned method for predicting lung adenocarcinoma gene mutations when executed by a processor.

[0021] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0022] The present application provides a method, apparatus, device, medium and product for predicting gene mutations in lung adenocarcinoma. The method comprises the following steps: identifying a tumor region on a lung slice image to obtain a region of interest image; segmenting the region of interest image to obtain a plurality of image blocks; and randomly combining the image blocks to obtain a plurality of slice packages. Feature extraction is performed based on a trained DINO model, and a two-stage multi-instance classification model is used to predict gene mutations in lung adenocarcinoma using the package features corresponding to the extracted slice packages. A more comprehensive set of feature information can be extracted, and lung adenocarcinoma gene mutation prediction results are obtained by predicting gene mutations in lung adenocarcinoma using all single-stage features. The correlation between instances in the package is effectively utilized, thereby improving the accuracy of lung adenocarcinoma gene mutation prediction. A machine learning model is used for tumor region identification, feature extraction and lung adenocarcinoma gene mutation prediction, thereby improving the efficiency of lung adenocarcinoma gene mutation prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 This is a diagram of the application environment of a method for predicting gene mutations in lung adenocarcinoma in one embodiment of the present application.

[0025] Figure 2 A flowchart of a method for predicting gene mutations in lung adenocarcinoma provided in one embodiment of the present application is provided.

[0026] Figure 3 This is a schematic diagram of the overall algorithm flow of a method for predicting gene mutations in lung adenocarcinoma provided in one embodiment of the present application, wherein: Figure 3 a in the figure is a schematic diagram of the ROI recognition process in the self-supervised model. Figure 3 b in the figure is the framework for predicting gene mutations. Figure 3 The c in the table is the performance comparison between different models and the external test results. Figure 3 The d in the figure is the comparison result between the weakly supervised model for predicting gene mutations and pathologists.

[0027] Figure 4 A schematic diagram of the network structure of the DINO model and the two-stage multiple instance classification (GAMIL) model provided in one embodiment of the present application, wherein: Figure 4 a in is the network structure of the DINO model. Figure 4 The b in the figure is the network structure of the GAMIL model.

[0028] Figure 5 The ROC curves of EGFR, KRAS, ALK, HER2 and some rare genes in lung adenocarcinoma in the training set and test set provided in one embodiment of the present application are shown in FIG. Figure 5 ab in the figure is the performance of the GAMIL model at the slide level. Figure 5 The cd in is the performance of the patient-level GAMIL model, Figure 5 The e in is the performance of the GAMIL model under triple cross validation.

[0029] Figure 6 This is a performance comparison result diagram of the GAMIL model provided in an embodiment of the present application and other models. Figure 6Where ac is the classification performance of the GAMIL model, Clustering-constrained Attention Multiple Instance Learning (CLAM) model, and Inceptionv3 model on EGFR gene status using features extracted from the UNI base model. Figure 6 de in the figure represent the external test performance of the GAMIL model using the Y hospital dataset and The cancer genome atlas (TCGA) dataset, respectively.

[0030] Figure 7 This is an example image of a high-weight region of gene mutation predicted by the GAMIL model provided in one embodiment of the present application. Figure 7 a in the figure is a high-weight region associated with EGFR gene mutation. Figure 7 b in the figure is a high-weight region associated with KRAS gene mutation. Figure 7 The c in the figure indicates a high-weight region associated with ALK gene mutation.

[0031] Figure 8 This is a schematic diagram of the results of pathologists and the GAMIL model predicting EGFR gene mutations according to an embodiment of the present application, wherein: Figure 8 a in the figure is the score comparison evaluation result between pathologists with different proficiency levels and the GAMIL model. Figure 8 (b) shows the performance comparison between the GAMIL model and pathologists on 100 slides through ROC curve analysis.

[0032] Figure 9 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] Lung cancer is widely considered a common malignant tumor. Traditional genetic testing methods have limitations such as high cost and lengthy procedures. Predicting clinically relevant gene mutations through histopathological images can facilitate the rapid identification of gene mutations in clinical settings.

[0035] Obtaining pathology hematoxylin and eosin (H&E) sections is relatively straightforward. Studies have shown that genetic alterations can induce changes in histomorphology; for example, most studies highlight a significant association between EGFR mutations and the squamous subtype. Correlations with other subtypes, such as micropapillary and acinar, remain a topic of ongoing debate, and the association between histological morphology and KRAS mutations primarily corresponds to solid and invasive mucinous adenocarcinoma subtypes. Furthermore, ALK mutations are known to be associated with a solid histological pattern and may contribute to the presence of mucinous cells. However, current studies have shown only limited correlations between morphology and genetic mutations. The accuracy of identifying genetic alterations based on current pathological morphology remains inferior to that achieved by molecular pathology. This effort is important because it has the potential to influence subsequent patient care and improve prognosis.

[0036] Artificial intelligence (AI) is a revolutionary technology that enables computers to process vast amounts of data with human-like cognitive abilities. Image processing and classification are well within the capabilities of AI, and this technology is increasingly being used in medicine, particularly in the classification of histopathology images. Nicholas Coudray et al. used the deep convolutional neural network Inception-v3 to predict the mutation status of 10 common genes, achieving AUC values ​​ranging from 0.733 to 0.856. Similarly, Chen Mayer et al. achieved remarkable success in predicting the mutation status of the ALK and ROS1 genes in lung adenocarcinoma by applying AI techniques.

[0037] This finding suggests that patterns reflecting gene mutations in tumor tissue morphological changes can be summarized. Although these patterns have not yet been fully elucidated by humans and cannot be recognized by the naked eye, they can be detected by computer vision-based algorithms. Using artificial intelligence methods to predict gene mutations in pathological images is of great significance. However, limitations in sample collection may lead to limited data sets, thereby reducing model accuracy. In addition, some models are trained and tested through databases such as TCGA, which may introduce differences from real-world scenarios, thereby limiting their clinical application. Therefore, further research using larger clinical cohorts is needed to better clarify the association between gene mutations and lung adenocarcinoma, thereby improving the accuracy of prediction models.

[0038] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0039] The lung adenocarcinoma gene mutation prediction method provided in the present application example can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the user's lung full slice image to the server 104. After the server 104 receives the user's lung full slice image, for the user's lung full slice image, the server 104 performs tumor area recognition on the lung slice image to obtain a region of interest image, divides the region of interest image to obtain several image blocks, and randomly combines the image blocks to obtain several slice packages. The slice package is input into the trained DINO model to extract the package features corresponding to the slice package, and the package features corresponding to all slice packages are input into the two-stage multi-instance classification model to obtain the lung adenocarcinoma gene mutation prediction result. The server 104 can feedback the obtained lung adenocarcinoma gene mutation prediction result for the user's lung full slice image to the terminal 102. In addition, in some embodiments, the lung adenocarcinoma gene mutation prediction method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform lung adenocarcinoma gene mutation prediction on the user's lung whole slice image, or the server 104 can obtain the user's lung whole slice image from the data storage system and perform lung adenocarcinoma gene mutation prediction on the user's lung whole slice image.

[0040] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0041] In an exemplary embodiment, Figure 2 As shown, a method for predicting gene mutations in lung adenocarcinoma is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, including the following steps 201 to 205.

[0042] Step 201: Obtain a full lung slice image of the user. The full lung slice image is an H&E slice.

[0043] Step 202: performing tumor region recognition on the lung slice image to obtain a region of interest image.

[0044] Step 203 : segment the image of the region of interest to obtain a number of image blocks, and randomly combine the image blocks to obtain a number of slice packages.

[0045] Step 204: For each of the slice packets, the slice packet is input into the trained DINO model to extract the packet features corresponding to the slice packet.

[0046] Step 205: input the package features corresponding to all the slice packages into a two-stage multi-instance classification model to obtain a lung adenocarcinoma gene mutation prediction result; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract features from the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

[0047] The prediction results of lung adenocarcinoma gene mutation are no gene mutation and gene mutation status; the gene mutation status includes EGFR, KRAS, ALK, HER2, ROS1, RET, BRAF, PIK3CA and NRAS gene mutations.

[0048] By implementing the above-mentioned steps 201 to 205, a region of interest image is obtained by performing tumor region identification on the lung slice image, and then the region of interest image is segmented to obtain a number of image blocks, and the image blocks are randomly combined to obtain a number of slice packages. Feature extraction is performed based on the trained DINO model, and a two-stage multi-instance classification model is used to predict lung adenocarcinoma gene mutations using the package features corresponding to the extracted slice packages. A more comprehensive set of feature information can be extracted, and lung adenocarcinoma gene mutation prediction results are obtained by performing lung adenocarcinoma gene mutation prediction on all single-stage features. The correlation between instances in the package is effectively utilized, and the accuracy of lung adenocarcinoma gene mutation prediction is improved. A machine learning model is used for tumor region identification, feature extraction, and lung adenocarcinoma gene mutation prediction, thereby improving the efficiency of lung adenocarcinoma gene mutation prediction.

[0049] This example collects 2221 slices from 1999 patients diagnosed with lung adenocarcinoma. The data includes full-slice image data and gene mutation information and related clinical data of EGFR, KRAS, ALK, HER2 and other rare genes (ROS1, RET, BRAF, PIK3CA, NRAS). The self-supervised models DINO model and GAMIL model are used to accurately identify the mutation status of 9 genes related to tumor occurrence and cancer progression. The comparison of model performance involves the use of various base models (UNI), classification models (CLAM and Inception v3), external data sets (TCGA and other medical institutions), and comparative analysis with human pathologists.

[0050] The gene mutation prediction framework includes the tumor region recognition model, the DINO model and the GAMIL model.

[0051] This example initially collected 2221 whole-slice images (WSIs) from 1999 patients and the corresponding clinical information of Hospital X, trained a ResNet model for tumor localization, extracted features through the DINO model, and trained a two-stage multi-instance classification model to predict gene mutations. At the same time, this example performed a comparative effect analysis with the basic model (UNI) and classification models (CLAM and inception v3). The remaining model capability evaluations included external testing (using 256 slices from Hospital Y and 541 slices from the TCGA database). The remaining model capability evaluations also included morphological analysis of high-weight areas annotated by the model, and comparisons between pathologists and the model in identifying gene mutations on WSIs. The overall process is as follows: Figure 3 shown.

[0052] First, sample collection is carried out to determine the data set.

[0053] Model training and testing: This example collected data from patients diagnosed with lung adenocarcinoma, whose visits ranged from September 2015 to April 2023. Each patient provided 1 to 3 representative tumor slice images. After excluding slices that did not meet the eligibility criteria (e.g., slices with too small tumor areas), a total of 2221 lung adenocarcinoma slices from 1999 patients were included in the study. This included 1756 surgical samples and 465 biopsy samples. All 2221 slices were used for training and testing of the tumor region model. However, one patient (2 slices) had multifocal tumors, and the second tumor lesion was not genetically tested. In order to maintain rigor, the slices corresponding to these tumors were excluded from the establishment and testing of the gene mutation prediction model. Therefore, the total number of slices used to establish and test the gene mutation prediction model was 2219 (1998 patients) slices. There was no overlap between the training set and the test set. All slices will be converted to svs format for model training and testing.

[0054] Sample Information: All slides were diagnosed as invasive adenocarcinoma by an experienced pathologist. Specimens included partial tumor tissue obtained by computed tomography-guided needle biopsy / bronchoscopy, as well as surgically resected tumors. The status of nine genes (EGFR, KRAS, ALK, HER2, ROS1, RET, BRAF, PIK3CA, and NRAS) was recorded. The number of slides per patient with each gene mutation is shown in Table 2. Some patients harbored two different types of EGFR mutations. Pathological H&E slides were digitized using a slide scanning system (SQS-600P or KF-PRO-400-HI). These WSI cases were scanned at 20x magnification (resolutions of 0.206 / 0.208 / 0.203 / 0.246 / 0.4 μm per pixel). Resolution was normalized before segmentation into blocks. Sample blocks were free of overlap.

[0055] Model generalization ability test: This example uses lung adenocarcinoma WSIs from Y Hospital and TCGA. The patients' visits were between 2023 and 2024. The inclusion and exclusion criteria were consistent with those mentioned above. A total of 256 slices from 256 patients were collected. Pathological H&E slices were digitized by a slice scanning system (model: KF-PRO-040). All slices contain EGFR and ALK gene status information. TCGA dataset: This dataset contains 54 WSIs, providing information on 9 different gene statuses. Among them, there are 540 cases of primary lung adenocarcinoma and 1 case of metastatic lung adenocarcinoma. Related resources can be downloaded from the following website: https: / / portal.gdc.cancer.gov / analysis_page?app=Downloads.

[0056] WSS S4LUAD Dataset: This dataset contains 87 holographic tomographic images, 67 of which are from Z Hospital and 20 from The Cancer Genome Atlas (TCGA). Annotation information includes tumor region data, including 10,091 plaque-level annotations (training set) and over 130 million labeled pixels (validation and test sets). Related resources can be downloaded from the following website: https: / / wsss4luad.grand-challenge.org / WSSS4LUAD / .

[0057] WSI Annotation: A total of 297 slides were annotated, and all annotated regions were carefully reviewed by expert pathologists (e.g. Figure 3 (As shown). The annotation software used was Aperio Scope. A total of 203 lung adenocarcinoma sections were included, covering 5 different lung adenocarcinoma subtypes: 49 cases were mainly acinar, 81 cases were lobular, 40 cases were papillary, 17 cases were micropapillary, and 16 cases were solid. The annotation process included marking non-tumor areas and tumor areas. In addition, this example selected 94 representative sections in which tumor areas were difficult to identify due to the presence of abundant bronchi, blood vessels, and bronchial cartilage. These sections were also annotated according to the same standards.

[0058] The following describes the process of specimen pretreatment and acquisition of genetic status.

[0059] Specimen preparation: The samples undergo a series of pretreatment steps to ensure optimal analysis. These steps include fixation with 10% neutral formalin, dehydration, paraffin embedding and other routine procedures commonly used in pathology.

[0060] Genetic Testing: To detect EGFR, KRAS, HER-2, ROS1, RET, BRAF, PIK3CA, and NRAS genes, nine mutation detection kits were used. These kits can simultaneously identify mutations in EGFR, ALK, KRAS, HER-2, ROS1, RET, BRAF, PIK3CA, and NRAS genes. Genetic Testing: To detect EGFR, this example used a human EGFR gene mutation detection kit.

[0061] However, according to a meta-analysis, the area under the receiver operating characteristic (ROC) curve for ALK status detected by IHC was 0.996 compared with ALK gene rearrangement, indicating high sensitivity and specificity, so the immunohistochemistry results of ALK can be considered equivalent to the results of ALK gene rearrangement. In addition, previous reports have shown that changes in ALK status detected by IHC are more closely correlated with clinical drug sensitivity. IHC can predict tumor response and survival in patients with advanced NSCLC treated with crizotinib. Given that ALK protein expression detected by IHC is considered equivalent to ALK rearrangement status, the ALK status of the sections was assessed by IHC method, and the ALK gene test results were not included in this study.

[0062] This example uses an ALK antibody (clone D5F3, Ventana anti-ALK rabbit monoclonal primary antibody, Roche) and an automated IHC machine (Roche, model: Roche Ultra) to ensure standardization and reduce human error. Each patient sample is accompanied by a positive control, in which strong nuclear IHC staining positive results were observed in the nerve cells of the appendicular muscle layer, and a negative control is also included. According to the manufacturer's scoring algorithm, a binary scoring system (ALK status is positive and negative) is used to evaluate the staining results: the determination of ALK positivity relies on the observation of cell nuclei with significant positive staining intensity, which must be comparable to the positive control to be considered positive.

[0063] The fully automated IHC VentanaBenchmark XT stainer and rabbit monoclonal primary antibody ALK (clone D5F3, Ventana) were used. Immunohistochemical interpretation criteria were the same as those at Hospital X.

[0064] The comprehensive algorithm consists of three main components: (1) ResNet: a tumor region recognition model used to identify regions of interest for subsequent prediction of mutant genes; (2) DINO model: a self-supervised model that generates features used by the mutant gene classification network; (3) GAMIL model: a two-stage multi-instance classification model used to predict the presence of gene mutations in WSI. The workflow is as follows Figure 3 Hardware: CPU: Intel(R) Xeon(R) CPU; GPU: NVIDIA GeForce RTX 3090; CUDA 11.0; cuDNN 8.0.4.30-1. Model training time: A single training run (1 gene) using 147 million image patches took 7 days on four NVIDIA GeForce RTX 3090s, with 100 epochs per run.

[0065] In step 202, tumor region recognition is performed on the lung slice image to obtain a region of interest image, which specifically includes: using a tumor region recognition model, taking the lung slice image as input, to recognize and obtain the region of interest image. The tumor region recognition model is a ResNet model.

[0066] ResNet: Tumor region recognition: The ResNet architecture (RRID: SCR_002121) was selected as the tumor region recognition network. The network's initialization parameters were derived from the pre-trained model on the public ImageNet dataset, and the corresponding weight parameters can be obtained from the following URL: https: / / download.pytorch.org / models / resnet18-f37072fd.pth. Each image block does not overlap with each other. The dataset contains a total of 297 slices, including 297 sample full-slice images, which are divided into training sets and test sets in a ratio of approximately 8:2, with no patient overlap between the two groups. The sample full-slice images are divided into image blocks (size 256×256 pixels), representing non-tumor areas and tumor areas, and the tumor area is recorded as the region of interest image.

[0067] Before inputting each slice package into a trained DINO model and extracting package features corresponding to the slice package, the lung adenocarcinoma gene mutation prediction method further includes: a DINO model training process, specifically as follows:

[0068] Acquire a data set; the data set includes a plurality of sample tumor image blocks;

[0069] Performing different processing on the sample tumor image block to obtain a first sample image and a second sample image;

[0070] The first sample image is input into the student model to obtain the student prediction output; the student prediction output is transformed by softmax to obtain the first eigenvector; the DINO model includes a teacher model and a student model;

[0071] The second sample image is input into the teacher model to obtain the teacher prediction output; the teacher prediction output is transformed by softmax to obtain the second eigenvector;

[0072] Calculating a loss function value based on the first eigenvector and the second eigenvector, and performing backpropagation and parameter update on the student model based on the loss function value to obtain updated student model parameters;

[0073] The model parameters of the teacher model are updated using the exponential moving average, and the stop gradient operator of the teacher model propagation gradient prevents the model parameters of the teacher model from being updated until the training stop condition is reached to obtain the trained student model and the trained teacher model; the trained student model and the trained teacher model constitute the trained DINO model.

[0074] DINO model: A feature extraction network. The self-supervised model used in this example is DINO (RRID: SCR_013497). To address the limited amount of available data for certain mutant genes, this example uses a self-supervised network model to extract features and reduce interference caused by class imbalance and lack of supervisory information. Figure 4 Figure a shows the architecture of the DINO (RRID: SCR_013497) network. DINO consists of two identical components: a teacher model and a student model. The student model continuously learns to mimic the teacher model's output, thereby enhancing feature representation and improving model performance. The training set and test set ratio is approximately 8:2, with no patient overlap between the two groups. Both the teacher and student models can be Transformer models.

[0075] The output distribution of the teacher model tends to be biased towards a certain mean, resulting in a decrease in diversity, so the centering operation is introduced. The formula for centering is z t\ ←z t -mean(z t ), where z t\ is the value after the centering operation, z t is the value before the centering operation, mean(z t ) is the batch average of the teacher model output. This operation can force the output to be distributed around zero and prevent network degradation.

[0076] Single-stage multi-instance model (referring to Figure 4 GAMIL-1 in : used for ROI recognition in pathological image self-supervision model. Two-stage multi-instance model (referring to Figure 4 GAMIL-2) in

[15] : A weakly supervised model for predicting gene mutations. The package features extracted from each WSI by the single-stage multi-instance model GAMIL-1 are merged to generate a final instance package. This is then input into a two-stage multi-instance model for lung adenocarcinoma gene mutation prediction, resulting in lung adenocarcinoma gene mutation prediction results.

[0077] The single-stage multi-instance model includes a first data combination module, a first gated attention mechanism layer, and a first fully connected layer. The first data combination module is used to concatenate the package features corresponding to the slice package to obtain a first overall feature representation F1, which includes features of several image blocks. The first overall feature representation F1 is input into the first gated attention mechanism layer, which assigns different weight values ​​to different instance features to obtain the weight value corresponding to each instance feature. The first overall feature representation F1 is multiplied by the weight value corresponding to the instance feature to obtain the first feature representation N2. The first fully connected layer is used to make predictions based on the first feature representation N2 to obtain the category information of the slice package. The category information of the slice package is a single-stage feature. Different instance features refer to features corresponding to different image blocks.

[0078] The two-stage multi-instance model includes a second data combination module, a second gated attention mechanism layer, and a second fully connected layer. The second data combination module is used to concatenate all single-stage features to obtain a second overall feature representation F2. The second overall feature representation F2 is input into the second gated attention mechanism layer, which assigns different weights to different instance features to obtain the weight value corresponding to each instance feature. The second overall feature representation F2 is multiplied by the weight value corresponding to the instance feature to obtain the second feature representation N2. The second fully connected layer uses the second feature representation N2 to predict lung adenocarcinoma gene mutations, obtaining the lung adenocarcinoma gene mutation prediction results.

[0079] GAMIL model: Two-stage multi-instance classification model: This embodiment adopts a multi-instance learning method to predict mutant genes and integrates a gated attention mechanism. Therefore, this embodiment is named GAMIL. In order to predict mutant genes, it is crucial to combine a more comprehensive set of feature information extracted from WSI. As a specific weakly supervised learning method, the multi-instance learning method operates by learning from package-level features rather than explicit instance labels. This method effectively utilizes the correlation between instances within a package. Compared with the single-stage multi-instance model, the two-stage multi-instance model not only improves the ability to capture nonlinear changes by increasing the number of model layers, but also improves the model accuracy by selecting a larger number of patches for each WSI slice. See the schematic diagram of the relevant process for details. Figure 4 The b in.

[0080] The following compares the GAMIL model provided in this application with the UNI basic model, CLAM model, and Inception v3 model.

[0081] UNI: The UNI model, based on the ViT architecture, is pre-trained using the DINOV2 self-supervised model and trained on more than 1 billion image patches from 20 different tissue types, totaling approximately 77TB of data, achieving excellent performance in various pathology tasks.

[0082] CLAM: The Cluster Constrained Attention Multi-Instance Learning (CLAM) model was proposed in 2021. As a weakly supervised model, it only requires sliding-level labels required for binary or multi-classification on WSI. The model is able to achieve high-precision classification with limited training data and performs well in pathological image analysis. In this example, the gene prediction effect is compared with the model of this example. The CLAM hyperparameters are set as follows: the dropout value is 0.25, the learning rate is 0.0001, the maximum number of training steps is 200, and early stopping is enabled to prevent overfitting. The selected CLAM model is CLAM_SB.

[0083] Inception v3: A deep learning model developed by Google for image recognition and classification. This model contains 48 convolutional layers, capable of extracting richer, higher-level image features. Through the Inception module, it utilizes multi-scale convolutional kernels to capture image features at different levels. Fusion of features at multiple scales improves image understanding capabilities. Inception v3 is a supervised network model. First, the Inception v3 model obtains the category of an image patch. Then, it calculates the average probability of positive samples within a sliding window and uses this as the prediction result at the sliding window level.

[0084] In addition to using different metrics to evaluate model performance, this example also evaluates its performance by comparing it with the performance of pathologists. Six pathologists participated in this study, all of whom had formal pathology training and different levels of certification (senior, intermediate, and junior; two at each level).

[0085] For this study, this example selected 150 annotated lung adenocarcinoma WSI samples, with a particular focus on the presence of EGFR mutations. Fifty of these samples came from the EGFR training set of the GAMIL model, and the remaining 100 came from the EGFR test set. There was no overlap between the training and test sets. Due to the limited number of samples in the EGFR test set used to train GAMIL, and the scarcity of samples of certain subtypes (papillary, micropapillary, and solid), this example could only collect all samples of a few subtypes, and strived to balance the number of different subtypes (lamellar, acinar, papillary, micropapillary, and solid). In addition, the cases were evenly divided into two groups: EGFR mutations and wild-type, each accounting for 50% (50 vs. 50).

[0086] In the initial stage of the study, a set of 50 annotated training sets WSI was used to assist the pathologists in the learning and training process. The main goal was to enable pathologists to independently analyze and infer the morphological patterns present in the slices. The GAMIL model has been established for the "Prediction of Common Genetic Mutations" section. In the subsequent stage of the experiment, a set of 100 lung adenocarcinoma pathology slides (randomly selected from the test set) in which the annotations of EGFR gene mutations were hidden were provided for pathologists and GAMIL models to use. The task was to identify and predict the presence of EGFR gene mutations. Pathologists relied on their existing expertise and learned patterns to interpret them. Each correctly diagnosed slice received 1 point, while there was no deduction for incorrect diagnosis.

[0087] The experiment used sensitivity, specificity, accuracy, AUC value, and ROC curve as indicators to evaluate model performance. Sensitivity measures the accuracy of identifying positive samples, while specificity measures the accuracy of identifying negative samples. Accuracy represents the overall recognition accuracy of all samples. The AUC value comprehensively considers the model's ability to classify both positive and negative samples. The ROC curve is a graphical tool used to demonstrate the performance of a model at different classification thresholds. In addition, this example used the K-fold cross-validation method to evaluate the model, where K = 3. There was no overlap between the patients in the training and test sets.

[0088] Gene mutations are mainly expressed in tumor regions. This example uses the ResNet model to build a model for identifying tumor regions. Subsequently, this example conducts comparative analysis on lung adenocarcinoma datasets, which include the WSSS4LUAD competition dataset and the dataset annotated by Hospital X. The model performs significantly on the test set, achieving an impressive AUC value of 0.995 (e.g. Figure 4 The model trained on the WSSS4LUAD dataset showed higher specificity, while the model trained on the doctor-annotated dataset performed well in both sensitivity and specificity. In addition, this example used the model to predict the tumor probability of each slice and generated a heat map by color coding to evaluate the unlabeled slices. Figure 4 The b in.

[0089] There is significant variation in mutation frequencies between genes. Some genes exhibit unusually low mutation rates, and the number of samples available for modeling is insufficient. Therefore, genes with significantly lower mutation rates, such as ROS1, RET, BRAF, PIK3CA, and NRAS, are classified as "rare genes" for model training.

[0090] This example uses the GAMIL model, a two-stage multi-instance learning method, to predict the status of EGFR, KRAS, ALK, HER2, and other rare genes. The WSIs dataset was divided into a training set and a test set with a ratio of approximately 8:2 to build the model. There was no overlap between patients in the training set and the test set. After algorithm optimization, this example evaluated the model performance at the slice level. The AUC values ​​of all genes exceeded 0.825, showing good sensitivity and specificity, as shown in Table 1 and Figure 5 As shown in a. Of particular note is the significant potential shown by ALK, with an impressive AUC value of 0.987. These findings highlight the excellent predictive power of the GAMIL model.

[0091] Considering that evaluating the model performance at the patient level may better reflect its practicality in actual clinical practice, the subsequent evaluation of this example focuses on the model performance at this level. For patients with multifocal tumors, this example retains the slices with a higher positive probability to represent the predicted probability of the patient. As shown in Table 1 and Figure 5 As shown in the cd, the sensitivity, specificity, and other indicators of the five gene sets at the patient level were consistent with the results at the slice level, with AUC values ​​ranging from 0.823 to 0.987. These findings indicate that the model performance in the patient-level evaluation is consistent with that observed at the slice level.

[0092] Table 1. Prediction results of EGFR, KRAS, ALK, HER2 and other rare genes using the GAMIL model

[0093] Sensitivity Specificity ACC Accounting Standards AUC EGFR-Slides 0.842 0.749 0.797 0.825 KRAS-Slideshow 0.786 0.972 0.942 0.911 ALK-Slides 0.972 0.982 0.978 0.987 HER2-Slideshow 0.882 0.989 0.981 0.934 Rare Genes - Slideshow 0.810 0.945 0.959 0.900 EGFR-patients 0.852 0.739 0.797 0.823 KRAS patients 0.769 0.974 0.942 0.905 ALK patients 0.972 0.982 0.978 0.987 HER2 patients 0.882 0.988 0.980 0.934 Rare Gene-Patients 0.850 0.948 0.965 0.907

[0094] In addition, this example performed a three-fold cross-validation on the GAMIL model. Again, there was no overlap between the test and training sets. As shown in Table 2, the average AUC values ​​for the GAMIL model ranged from 0.748 to 0.965. The variance ranged from 0.01 to 0.07, indicating minimal data fluctuations and demonstrating the stability of the GAMIL model.

[0095] Table 2 3-fold cross validation results of GAMIL model

[0096]

[0097] This example establishes various comparison methods, including the basic model (UNI), classification models (CLAM, Inceptionv3), and external datasets (Y hospital dataset, TCGA database) for comparative analysis.

[0098] Initially, this example replaced the DINO model with the base model UNI for feature extraction, and then used the GAMIL model for gene status classification. Since the EGFR gene had the lowest AUC value in the GAMIL model and had a high positive rate in the slices, EGFR was selected as the training object for the UNI+GAMIL model, using the same training set as the UNI+GAMIL model. As shown in Table 3 and Figure 6 As shown in Figure 3, the AUC value of the GAMIL model using the UNI base model can reach 0.799.

[0099] Subsequently, this example compared and analyzed the CLAM, Inceptionv3, and GAMIL models. The AUC value of CLAM (EGFR gene) was only 0.679, while the AUC value of the Inception v3 model was 0.665 (both at the slice level). This indicates that the specificity is low but the sensitivity index is acceptable, as shown in Table 3 and Figure 6 In comparison, the model of this embodiment obviously has higher comprehensive advantages.

[0100] Table 3 Comparison of gene status prediction performance of the GAMIL model using other models and external datasets

[0101]

[0102]

[0103] Finally, to evaluate the generalization ability of the GAMIL model, this example used 256 WSIs from Y Hospital and 541 slices from the TCGA database as test sets. The GAMIL model of this example achieved satisfactory results on the Y Hospital dataset, with AUC values ​​of 0.843 (ALK gene) and 0.800 (EGFR gene), respectively. However, the AUC value on the TCGA dataset only reached 0.508-0.716, as shown in Table 3 and Figure 6 As shown, Figure 6 The performance of the GAMIL model was compared with other models and validated on external datasets. This may be attributed to the generally low degree of differentiation of the sections in the TCGA dataset, which is mainly concentrated in surgical specimens of advanced lung adenocarcinoma. In addition, there are differences in the color intensity standards of HE staining, resulting in poor quality of some sections.

[0104] This example uses the high-weight regions identified by the model on the slices to study and establish the correlation between mutated genes and pathological tissue morphological characteristics, which helps pathologists gain a deeper understanding of the tissue morphology of lung adenocarcinoma with gene mutations.

[0105] Analysis of histomorphological characteristics of patients with gene mutation lung adenocarcinoma: This example also uses the model to generate an attention map to show the distribution of WSI prediction results. This example selects representative areas for display (such as Figure 7 The model of this embodiment effectively locates the relevant areas and highlights them with different colors in the heat map. For example, this embodiment detects the manifestation of squamous cell carcinoma, which is closely related to EGFR mutation (such as Figure 7 In addition, this example identifies regions corresponding to invasive mucinous adenocarcinoma and solid subtypes caused by KRAS mutations (e.g., Figure 7 In addition, regions characterized by mucous cells and solid subtypes were observed in this example, and these features were attributed to ALK mutations (e.g. Figure 7 Furthermore, in terms of noise and systematicity, the model of this embodiment performs well, with fewer noise points and fewer scattered false-positive regions in the prediction results. This demonstrates that the model of this embodiment can effectively filter out irrelevant information and more accurately identify the range of gene mutations. However, the prediction performance in certain areas is slightly insufficient, such as identifying certain lymphocytes as high-weight regions. These cases have been documented in this embodiment. In subsequent research, this embodiment will focus on optimizing the model's recognition and prediction capabilities in these areas.

[0106] In this example, 150 WSI samples were selected and 6 pathologists were recruited to evaluate the presence of EGFR mutations. For detailed information on the allocation of training and test sets and the scoring criteria, please refer to the Materials and Methods section. This example aims to compare the effectiveness of human pathologists and artificial intelligence models in identifying EGFR mutations in lung adenocarcinoma by histological morphology. A detailed scoring analysis was performed on the pathologists and the GAMIL model, as shown in Figure 2. Figure 8 As shown in a in Figure 3. The analysis results clearly showed that the GAMIL model scored significantly higher than the pathologists.

[0107] Despite the pathologists’ extensive expertise and advanced knowledge in morphology-based diagnosis of lung disease, as well as their training in identifying EGFR mutation-positive and -negative sections, they demonstrated limited sensitivity and specificity in their ability to accurately distinguish EGFR mutation-positive from wild-type sections. This discrepancy highlights the challenges pathologists face in accurately localizing such samples. Compared to junior pathologists, mid-level pathologists demonstrated higher sensitivity and specificity, while senior pathologists performed slightly better. However, all metrics remained significantly insufficient. In contrast, the GAMIL model achieved higher sensitivity, specificity, and AUC values ​​than human pathologists, highlighting its superior performance (e.g., Figure 8 (as shown in b in the figure).

[0108] To assess the agreement between pathologists at different levels, a concordance test was performed, with the kappa coefficient between groups set empirically. The results clearly showed a lack of consensus among pathologists at different levels, suggesting that each pathologist used different criteria when assessing the presence of EGFR gene mutations.

[0109] This example uses the GAMIL model to predict several common gene mutations, demonstrating excellent accuracy, with an AUC value of 0.987 for the ALK gene. The advantage of this example's approach lies in the large dataset of tissue sections, including 2,219 sections from 1,998 patients, which was used to train and test the GAMIL model to predict gene mutations, thereby improving its performance. The high-weight regions identified by the model provide insights into the morphological characteristics of gene mutations, offering pathologists a new perspective that links macroscopic morphological changes with microscopic molecular levels, helping to more accurately understand this association. In addition, this example also conducted comparative analyses with other models, pathologists, and this example's own model.

[0110] This example further uses data from Y Hospital and TCGA to verify the generalization ability and robustness of the model. The AUC value reached 0.508 to 0.716. The prediction error of the model may be related to the quality of the tissue slices, such as unclear local scanning of the slices and tissue blockage caused by residual sealing glue on the coverslip. In this example, every effort was made to clean the slices during the data collection process to reduce the impact of slice quality. However, the quality of external datasets is more difficult to control. In addition, there are fewer cases of poorly differentiated lung adenocarcinoma in the training set, resulting in a lack of characteristics for some poorly differentiated tumors, which may lead to recognition difficulties and reduce the AUC value of the model.

[0111] In interviews with pathologists, they expressed the challenge of linking morphological changes to genetic mutations. One mid-level pathologist stated, "When initially examining several WSIs, I thought I had identified a specific pattern based on the cell morphology and the presence of inflammatory cells in the background. However, upon further examination, I observed similar morphological changes in WSIs with wild-type EGFR." Two senior physicians reported no clear pattern among the WSIs. The remaining physicians experienced varying difficulties in identifying the pattern, with some even questioning its existence. Nevertheless, the GAMIL model's AUC value of 0.810 does indicate the presence of a detectable pattern.

[0112] Currently, the main treatments for lung adenocarcinoma include surgical intervention and chemotherapy. Although significant progress has been made in the diagnosis and treatment of lung cancer in recent years, several challenges still exist. Traditional treatments are no longer sufficient to meet individualized medical needs. Research on lung adenocarcinoma has shifted from the macroscopic level to the microscopic level, prompting people to explore the molecular mechanisms in depth. Given the importance and clinical significance of specific gene mutations in LUAD, the study in this example focuses on nine common genes. The development of personalized targeted treatment regimens for different gene mutations is a current trend, some of which already have mature targeted treatment options, such as the use of EGFR inhibitors to treat patients with EGFR mutations. Multi-generation ALK inhibitors targeting ALK rearrangements are crucial. Despite the continuous progress in targeted treatment of lung cancer, further research is still needed to develop specific therapeutic drugs for certain mutations, such as NRAS.

[0113] Compared with more expensive methods, the use of whole genome sequencing of patients to evaluate biomarkers is relatively easier to implement in future clinical practice. The artificial intelligence model (predictive gene mutation framework) trained in this embodiment can quickly identify the gene status immediately after the H&E stained sections are prepared from the patient's pathological tissue. This method significantly reduces the time required for the patient and saves costs. Especially for patients who undergo puncture or interventional surgical biopsy and are not suitable for surgical resection of the tumor, the amount of tissue obtained is usually limited. Even if necessary, the IHC analysis required for diagnosis will not be performed, and the remaining tissue may not be sufficient for molecular testing. Using this model to predict gene status does not require additional tissue, which can minimize human intervention in the process of molecular pathology gene testing. This embodiment believes that in the future, artificial intelligence will be used in clinical practice and become a tool for auxiliary diagnosis and treatment.

[0114] After the above verification, the method of this embodiment is superior to the CLAM model and the initial v3 model (i.e., the Inception v3 model), with AUC values ​​ranging from 0.825 to 0.987 in predicting gene mutations. The AUC values ​​on the external test data set were 0.516 to 0.843. In addition, when comparing the prediction of EGFR gene mutations between pathologists and the GAMIL model, the GAMIL model showed a significantly higher AUC value of 0.810, exceeding the pathologist average AUC value of 0.508. The GAMIL model performed well in the delineation of lung adenocarcinoma tumor areas and the prediction of gene mutations. The use of these models has great potential in significantly improving the efficiency of molecular detection and opening up new avenues for personalized treatment.

[0115] The present application also provides an application scenario, which applies the above-mentioned lung adenocarcinoma gene mutation prediction method. Specifically: the lung adenocarcinoma gene mutation prediction method provided in this embodiment can be applied in the lung adenocarcinoma gene mutation prediction scenario. The lung adenocarcinoma gene mutation prediction scenario includes an image generation link and a lung adenocarcinoma gene mutation prediction link; the user's lung full slice image enters the lung adenocarcinoma gene mutation prediction link from the image generation link, and the corresponding lung adenocarcinoma gene mutation prediction result is obtained through human-computer collaboration. The lung adenocarcinoma gene mutation prediction method provided in this embodiment belongs to the lung adenocarcinoma gene mutation prediction link. Specifically, in the lung adenocarcinoma gene mutation prediction link process for the lung full slice image, the lung slice image can be subjected to tumor area identification to obtain an image of the region of interest, the image of the region of interest can be segmented to obtain several image blocks, and the image blocks can be randomly combined to obtain several slice packages, the slice package is input into the trained DINO model, the package features corresponding to the slice package are extracted, and the package features corresponding to all slice packages are input into the two-stage multi-instance classification model to obtain the lung adenocarcinoma gene mutation prediction result.

[0116] Based on the same inventive concept, the present application also provides a lung adenocarcinoma gene mutation prediction device for implementing the aforementioned lung adenocarcinoma gene mutation prediction method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more lung adenocarcinoma gene mutation prediction device embodiments provided below can be found in the above-mentioned limitations of the lung adenocarcinoma gene mutation prediction method and will not be repeated here.

[0117] In an exemplary embodiment, a device for predicting gene mutations in lung adenocarcinoma is provided, comprising the following modules:

[0118] A lung full-slice image acquisition module is used to acquire a user's lung full-slice image;

[0119] a tumor region recognition module, configured to perform tumor region recognition on the lung slice image to obtain a region of interest image;

[0120] An image segmentation and combination module is used to segment the image of the region of interest to obtain a number of image blocks, and randomly combine the image blocks to obtain a number of slice packages;

[0121] A packet feature extraction module is used to input each slice packet into the trained DINO model to extract the packet features corresponding to the slice packet;

[0122] A lung adenocarcinoma gene mutation prediction module is used to input the package features corresponding to all the slice packages into a two-stage multi-instance classification model to obtain lung adenocarcinoma gene mutation prediction results; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract features from the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

[0123] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store lung adenocarcinoma gene mutation prediction data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for predicting lung adenocarcinoma gene mutation is implemented.

[0124] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0125] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0126] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0127] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0128] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0129] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for predicting gene mutations in lung adenocarcinoma, characterized in that: The lung adenocarcinoma gene mutation prediction method comprises: Obtain a full-slice image of the user's lungs; Performing tumor region recognition on the lung slice image to obtain a region of interest image; Segmenting the image of the region of interest to obtain a number of image blocks, and randomly combining the image blocks to obtain a number of slice packages; For each of the slice packages, the slice package is input into the trained DINO model to extract the package features corresponding to the slice package; The package features corresponding to all the slice packages are input into the two-stage multi-instance classification model to obtain the lung adenocarcinoma gene mutation prediction results; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

2. The method for predicting lung adenocarcinoma gene mutation according to claim 1, characterized in that: Performing tumor region identification on the lung slice image to obtain a region of interest image specifically includes: The tumor region recognition model is used to identify and obtain a region of interest image by taking the lung slice image as input.

3. The method for predicting lung adenocarcinoma gene mutation according to claim 2, characterized in that: The tumor region recognition model is a ResNet model.

4. The method for predicting lung adenocarcinoma gene mutation according to claim 1, wherein: Before inputting each slice package into a trained DINO model to extract package features corresponding to the slice package, the lung adenocarcinoma gene mutation prediction method further includes: the DINO model includes a teacher model and a student model; Acquire a data set; the data set includes a plurality of sample tumor image blocks; Performing different processing on the sample tumor image block to obtain a first sample image and a second sample image; Input the first sample image into the student model to obtain the student prediction output; the student prediction output is transformed by softmax to obtain the first eigenvector; The second sample image is input into the teacher model to obtain the teacher prediction output; the teacher prediction output is transformed by softmax to obtain the second eigenvector; Calculating a loss function value based on the first eigenvector and the second eigenvector, and performing backpropagation and parameter update on the student model based on the loss function value to obtain updated student model parameters; The model parameters of the teacher model are updated using the exponential moving average, and the stop gradient operator of the teacher model propagation gradient prevents the model parameters of the teacher model from being updated until the training stop condition is reached to obtain the trained student model and the trained teacher model; the trained student model and the trained teacher model constitute the trained DINO model.

5. The method for predicting lung adenocarcinoma gene mutation according to claim 1, wherein: The whole-slice images of the lungs are H&E slices.

6. The method for predicting gene mutations in lung adenocarcinoma according to claim 1, wherein: The lung adenocarcinoma gene mutation prediction results are no gene mutation and gene mutation status; the gene mutation status includes EGFR, KRAS, ALK, HER2, ROS1, RET, BRAF, PIK3CA and NRAS gene mutations.

7. A device for predicting gene mutations in lung adenocarcinoma, characterized in that: The lung adenocarcinoma gene mutation prediction device comprises: A lung full-slice image acquisition module is used to acquire a user's lung full-slice image; a tumor region recognition module, configured to perform tumor region recognition on the lung slice image to obtain a region of interest image; An image segmentation and combination module is used to segment the image of the region of interest to obtain a number of image blocks, and randomly combine the image blocks to obtain a number of slice packages; A packet feature extraction module is used to input each slice packet into the trained DINO model to extract the packet features corresponding to the slice packet; A lung adenocarcinoma gene mutation prediction module is used to input the package features corresponding to all the slice packages into a two-stage multi-instance classification model to obtain lung adenocarcinoma gene mutation prediction results; the two-stage multi-instance classification model includes several single-stage multi-instance models and one two-stage multi-instance model; the number of single-stage multi-instance models is the number of slice packages, and the single-stage multi-instance model is used to extract features from the package features corresponding to the slice packages to obtain single-stage features; the two-stage multi-instance model is used to predict lung adenocarcinoma gene mutations based on all the single-stage features to obtain lung adenocarcinoma gene mutation prediction results.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the lung adenocarcinoma gene mutation prediction method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting gene mutations in lung adenocarcinoma according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for predicting gene mutations in lung adenocarcinoma according to any one of claims 1 to 6 is implemented.