Lung cancer EGFR (Epidermal Growth Factor Receptor) gene mutation prediction method based on multi-instance learning and Transform technology

By segmenting lung CT images into image blocks and combining multi-example learning and Transformer technology, the problem that traditional methods fail to fully consider the complexity of internal tumors is solved, achieving higher accuracy of EGFR mutation prediction and model generalization.

CN119993265AActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411818764.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Traditional CT-based EGFR mutation prediction methods fail to fully consider the complexity and diversity of the internal structure of the tumor, resulting in low prediction accuracy.

Method used

Using a method based on multi-example learning and Transformer technology, by segmenting lung CT images into image blocks for instance-level feature extraction, soft pseudo-marking technology provides additional supervision signals for instance-level features, and integrates unique relative spatial position information into the self-attention mechanism to mine dependencies between instances.

Benefits of technology

This improves the representation ability of heterogeneous tumors and the prediction accuracy of EGFR mutation status, and enhances the generalization of the model on different data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993265A_ABST
    Figure CN119993265A_ABST
Patent Text Reader

Abstract

The invention relates to a lung cancer EGFR (Epidermal Growth Factor Receptor) gene mutation prediction method based on multi-instance learning and Transform technology, which comprises the following steps of: segmenting a lung CT (Computed Tomography) image into image blocks for instance-level feature extraction, providing an extra supervision signal for instance-level features by utilizing a soft pseudo-marking technology in a self-generating manner, guiding a multi-instance learning model to more effectively distinguish instances, and predicting the mutation of the lung cancer EGFR gene mutation. The identification capability of instance-level features is enhanced; unique relative spatial position information is integrated into a self-attention mechanism, so that the dependency relationship between instances can be flexibly and clearly mined, and the expression capability of heterogeneity tumors can be greatly improved; the importance scores of different instance features are automatically learned through the self-adaptive gating instance aggregation module, dynamic weighting of feature aggregation instances is facilitated, and the generalization of the model on different data sets is improved through the method. According to the prediction method provided by the invention, heterogeneity characteristics in the tumor can be caught and analyzed finely, and it is ensured that prediction of the EGFR mutation state is more accurate and high in pertinence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for predicting EGFR gene mutations in lung cancer based on multi-instance learning and Transformer technology. Background Art

[0002] Tumors play a key role in cancer-related deaths, and non-small cell lung cancer (NSCLC), as its main pathological classification, occupies a prominent position. In recent years, the leap in genomics and precision medicine has promoted the innovation of the NSCLC treatment paradigm. In particular, tyrosine kinase inhibitors (TKIs) therapy targeting epidermal growth factor receptor (EGFR) has become one of the first choice for the treatment of NSCLC patients. Clinical evidence clearly shows that compared with EGFR wild-type cases, mutant cases carrying EGFR mutations show a higher treatment response rate to TKIs, which greatly improves the prognostic indicators and disease-free survival of these NSCLC patients, thereby emphasizing the clinical significance of predicting EGFR mutation status before treatment, and laying the foundation for the precise selection of EGFR-TKI treatment and personalized medical pathways.

[0003] Traditionally, histological samples obtained through tissue biopsy or tumor surgery are regarded as the gold standard for EGFR mutation prediction. However, this method has several limitations, including sampling bias, DNA degradation, high cost and long processing cycle, which together restrict its widespread application in the field of EGFR mutation screening. In addition, the invasiveness of the biopsy operation itself and the possible complications, such as pneumothorax and bleeding, are particularly uncomfortable for patients in advanced disease states. In view of this, it is imperative to develop efficient and non-invasive EGFR mutation prediction technology to overcome the shortcomings of existing methods and broaden the feasibility boundaries of pre-treatment molecular typing.

[0004] Recently, computed tomography (CT) has demonstrated its extraordinary value as a noninvasive imaging tool in the diagnosis and treatment of non-small cell lung cancer (NSCLC), especially in the noninvasive prediction of EGFR mutations.

[0005] However, the prediction models used in traditional CT-based EGFR mutation prediction methods are mostly built on a fully supervised learning framework. These models generally assume that the EGFR mutation status inside the tumor is evenly distributed and directly extract features from the overall CT image to predict the mutation status. This traditional method has a fatal flaw: it fails to fully consider the complexity and diversity of the internal structure of the tumor, namely tumor heterogeneity.

[0006] Different subregions of the tumor may show completely different mutation characteristics. Traditional CT-based EGFR mutation prediction methods fail to consider and utilize this spatial variation information, resulting in low prediction accuracy. Summary of the invention

[0007] Based on this, it is necessary to provide a lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology to address the problem that the traditional lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology fails to fully consider the complexity and diversity of the internal structure of the tumor.

[0008] The present application provides a method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology, the method comprising:

[0009] Obtain multiple lung CT images;

[0010] Each lung CT image is preprocessed to obtain an image of a region of interest in each lung CT image;

[0011] Divide each region of interest image into multiple image blocks of equal size;

[0012] Create an EGFR mutation prediction model;

[0013] The EGFR mutation prediction model is trained k times to obtain a trained EGFR mutation prediction model; k is a positive integer and k is greater than 1;

[0014] Acquiring a lung CT image to be tested, and inputting the lung CT image to be tested into the trained EGFR mutation prediction model;

[0015] Starting the trained EGFR mutation prediction model to obtain a prediction result output by the trained EGFR mutation prediction model;

[0016] The training of the EGFR mutation prediction model for k times to obtain the trained EGFR mutation prediction model comprises:

[0017] Using the feature encoder to encode all image blocks one by one, and obtaining a preliminary prediction result of each image block based on the encoding result of each image block;

[0018] The importance score of each image block is embedded in the preliminary prediction result of each image block to generate a soft pseudo label for each image block;

[0019] Create a spatial perception Transformer module with a multi-layer encoder structure. Based on the spatial perception Transformer module with a multi-layer encoder structure, the spatial relative position relationship between different image blocks is mined to obtain the comprehensive features of each image block.

[0020] Based on the adaptive gated instance aggregation module, the comprehensive features of each image block and the importance score of each image block are aggregated using a dynamic weighting method to obtain the packet-level CT image features of each image block;

[0021] Construct the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function respectively;

[0022] Construct an overall loss function using the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function;

[0023] In the next training, the value of the overall loss function is minimized as the training goal.

[0024] The present application relates to a lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology. By segmenting lung CT images into image blocks for instance-level feature extraction, the soft pseudo-labeling technology is used in a self-generated manner to provide additional supervision signals for instance-level features, guide the multi-instance learning model to distinguish instances more effectively, and enhance the identification ability of instance-level features; the unique relative spatial position information is integrated into the self-attention mechanism, which can flexibly and clearly mine the dependencies between instances, greatly improve the representation ability of heterogeneous tumors, and take into account the importance of sub-regional interactions in the prediction process; the importance scores of different instance features are automatically learned through the adaptive gated instance aggregation module, which facilitates the dynamic weighting of feature aggregation instances. This method ensures that more information-rich features are emphasized in heterogeneous tumors, thereby improving the generalization of the model on different data sets. The prediction method provided in this application can delicately capture and analyze the heterogeneous characteristics inside the tumor, ensuring that the prediction of the EGFR mutation status is more accurate and targeted. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flowchart of a method for predicting EGFR gene mutations in lung cancer based on multi-instance learning and Transformer technology is provided in one embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of this application more clear, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.

[0027] The present application provides a method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology. It should be noted that the method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology provided by the present application is applied to the prediction of gene mutations in EGFR cases of non-small cell lung cancer (NSCLC).

[0028] In addition, the lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology provided in the present application does not limit its execution subject. Optionally, the execution subject of the lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology provided in the present application can be a lung cancer EGFR gene mutation prediction system based on multi-instance learning and Transformer technology. Specifically, the execution subject of the lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology provided in the present application can be one or more processing terminals in the lung cancer EGFR gene mutation prediction based on multi-instance learning and Transformer technology.

[0029] It should be noted that all the steps and processes described in this application are not intended to treat or diagnose diseases. The lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology provided in this application only predicts lung cancer EGFR gene mutation by establishing a prediction model, which belongs to the image processing and computer data analysis of lung CT images, and has no direct effect on any qualitative diagnosis, but only plays an auxiliary and reference role.

[0030] That is, the lung cancer EGFR gene mutation prediction result obtained in the present application is only an intermediate information for any subsequent application, or an intermediate data processing result for subsequent analysis.

[0031] like Figure 1 As shown, in one embodiment of the present application, the lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology includes:

[0032] S100, acquiring multiple lung CT images.

[0033] S200, preprocessing is performed on each lung CT image to obtain an image of a region of interest in each lung CT image.

[0034] S300, dividing each region of interest image into a plurality of image blocks of equal size.

[0035] S400, create an EGFR mutation prediction model.

[0036] S500, training the EGFR mutation prediction model k times to obtain a trained EGFR mutation prediction model, where k is a positive integer and k is greater than 1.

[0037] S600, obtaining a lung CT image to be tested, and inputting the lung CT image to be tested into the trained EGFR mutation prediction model.

[0038] S700, starting the trained EGFR mutation prediction model, and obtaining the prediction result output by the trained EGFR mutation prediction model.

[0039] The S500 includes:

[0040] S510, using a feature encoder to encode all image blocks one by one, and obtaining a preliminary prediction result of each image block according to the encoding result of each image block.

[0041] S520, embedding the importance score of each image block in the preliminary prediction result of each image block, and generating a soft pseudo label for each image block.

[0042] S530, creating a spatial perception Transformer module with a multi-layer encoder structure, based on the spatial perception Transformer module with a multi-layer encoder structure, mining the spatial relative position relationship between different image blocks, and obtaining the comprehensive features of each image block.

[0043] S540, based on the adaptive gated instance aggregation module, the comprehensive features of each image block and the importance score of each image block are aggregated by a dynamic weighting method to obtain the package-level CT image features of each region of interest image.

[0044] S550, constructing a prediction loss function of the bag feature, a prediction loss function of the class label feature, and a supervision loss function respectively.

[0045] S560, constructing an overall loss function using the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function.

[0046] S570, in the next training, minimizing the value of the overall loss function is used as a training objective.

[0047] Specifically, S100 acquires a plurality of lung CT images, each of which is a sample. The lung CT images constitute a data set for predicting the EGFR mutation status.

[0048] The datasets used to predict EGFR mutation status come from two NSCLC collections, including a private dataset and a public dataset. Inclusion criteria can include: (1) pathologically diagnosed primary NSCLC. (2) Detection of EGFR mutation. Exclusion criteria include: (1) large artifacts or poor quality in CT images. (2) Lack of genetic testing or poor tissue quality, resulting in detection failure.

[0049] Among them, the private dataset was collected from cases of eligible patients diagnosed or treated in a tertiary hospital in a certain province in China. Finally, 420 cases (228 EGFR mutants and 192 EGFR wild-type cases) were collected as an internal dataset, which has been approved by the Institutional Review Board of the domestic tertiary hospital (No.k2021-014). The collected CT image slices were 1.0 mm thick, and the tumor areas were manually annotated by two experienced radiologists. The public dataset is provided by the Cancer Imaging Archive and includes 211 original cases from the Stanford University School of Medicine and the Palo Alto Veterans Health Care System. After reviewing the data according to the same inclusion criteria, 140 cases were found with CT image slice thickness ranging from 0.625 to 3.0 mm.

[0050] S200, preprocessing is performed on each lung CT image to obtain an image of a region of interest in each lung CT image. The preprocessing process includes:

[0051] S210, collected CT images of 560 patients with non-small cell lung cancer EGFR mutation as a sample set, used the trilinear interpolation algorithm to resample the lung CT image data, and adjusted the voxel size of the lung CT image to 1mm×1mm×1mm. The purpose of this is to unify the size.

[0052] S220, pixel intensity of each lung CT image of 1mm×1mm×1mm size is clipped to -1000~400HU.

[0053] S230, according to the tumor center point marked by the doctor, the region of interest of the tumor is cropped and its pixel size is adjusted to an image of 224×224 pixels. The tumor center point is the equivalent physical center of the tumor area, because the tumor area is an irregular shape. The region of interest is the tumor area.

[0054] Next, S300 is executed.

[0055] S300 uses a sliding window of 16×16 pixels to segment the tumor region of interest into multiple sub-regions, and each sub-region is used as an example of the model input in this method.

[0056] Therefore, each 224×224 pixel image can be cut into 196 16×16 pixel blocks. Each image block can be denoted by x i,j , then each region of interest image has an image patch dataset X i ,,X i ={x i,1 , x i,2 ,...,x i,j}. i is the serial number of the region of interest image, and j is the serial number of the image block. i = 1, 2, 3, ..., N. j = 1, 2, 3, ..., M. N and M are both positive integers. N is the number of region of interest images, and M is the number of image blocks contained in a single region of interest image.

[0057] Assign a true label Y to each image of size 224×224 pixels i , giving each 16×16 pixel image block a true label y i,j .

[0058] The relationship between the true label of the region of interest image and the true label of the region of interest image of the image block can be expressed as follows based on the constraints of multi-instance learning:

[0059]

[0060] Among them, Y i is the true label of the image of the region of interest, y i,j is the true label of the image block, M is the number of image blocks contained in a single region of interest image, i is the serial number of the region of interest image, and j is the serial number of the image block.

[0061] The meaning of this constraint expression is that the mutation result of the region of interest image with a size of 224×224 pixels is 1, so if there is at least one small image block in the region of interest image with a size of 224×224 pixels, the mutation result is 1. 1 means there is a mutation, and 0 means there is no mutation. This is a binary expression method.

[0062] Because these lung CT images used for training are mature cases, they already have mutation results, so the labels generated here are called true labels.

[0063] It should be noted that in the subsequent description of this application, the region of interest image = package, and the image block = instance. That is, the concept of package refers to the category of region of interest image, and the concept of instance refers to the category of image block. No additional explanation will be given later.

[0064] In this embodiment, the lung CT image is divided into image blocks for instance-level feature extraction, and the soft pseudo-labeling technology is used in a self-generated manner to provide additional supervision signals for instance-level features, guide the multi-instance learning model to distinguish instances more effectively, and enhance the identification ability of instance-level features. Integrating the unique relative spatial position information into the self-attention mechanism can flexibly and clearly mine the dependencies between instances, which can greatly improve the representation ability of heterogeneous tumors and take into account the importance of sub-regional interactions in the prediction process. The importance scores of different instance features are automatically learned through the adaptive gated instance aggregation module, which facilitates the dynamic weighting of feature aggregation instances. This method ensures that more information-rich features are emphasized in heterogeneous tumors, thereby improving the generalization of the model on different data sets. The prediction method provided in this application can delicately capture and analyze the heterogeneous characteristics inside the tumor, ensuring that the prediction of EGFR mutation status is more accurate and targeted.

[0065] In one embodiment of the present application, S510 includes, that is, using the feature encoder to encode all image blocks one by one, and obtaining a preliminary prediction result of each image block according to the encoding result of each image block, including:

[0066] S511, encoding all image blocks one by one using a feature encoder based on Formula 1 to generate feature encoding information for each image block.

[0067] x′ i,j =F r (x i,j :θ r ) Formula 1.

[0068] Among them, x′ i,j For image block x i,j The feature encoding information, x i,j is the jth image block in the i-th region of interest image, i is the serial number of the region of interest image, j is the serial number of the image block, F r (·:θ r ) is a symbol for encoding behavior.

[0069] S521, based on Formula 2, the feature encoding information of each image block is input into the instance prediction head for processing, and a preliminary prediction result of each image block output by the instance prediction head is obtained.

[0070]

[0071] in, For image block x i,j The preliminary prediction results, x′ i,j For image block x i,j The feature encoding information, ψ p(·:θ p ) is a symbol for expressing the processing behavior performed by the instance prediction head.

[0072] Specifically, the feature encoder used is a feature encoder based on the ResNet architecture. The prediction head is the last layer of the feature encoder, which is responsible for converting the previously extracted features into the final prediction results.

[0073] In this embodiment, CT images of small cell lung cancer (NSCLC) often have low contrast and high variability, which may lead to insufficient discrimination ability of the extracted instance features for non-invasive EGFR mutation prediction. Especially in heterogeneous tumor regions, the intensity difference between EGFR mutant and wild-type tumor tissues may be extremely small, which significantly increases the difficulty of learning identification features from complex CT images. This step can generate features by encoding image blocks, and further processing by the prediction head can complete feature refinement and classification.

[0074] In one embodiment of the present application, S520 includes, that is, embedding the importance score of each image block in the preliminary prediction result of each image block to generate a soft pseudo label for each image block, including:

[0075] S521, generating an importance score for each image block.

[0076] S522, creating an image block importance score set, and including all image block importance scores into the image block importance score set one by one.

[0077] S523, based on the image block importance score set, generate a soft pseudo label package for each image block using Formula 3.

[0078]

[0079] in, For image block x i,j Soft fake labels, For image block x i,j The importance score of is the set of image block importance scores of the i-th region of interest image, is the indicator function, and k is the symbol representing the current number of training times. i is the true label of the image of the region of interest.

[0080] Specifically, when Y i=1 hour, The value of is 1. R is the range of the given set. A higher importance score indicates that the instance is embedded with richer mutation information, which reflects the contribution of the instance to EGFR prediction to a certain extent. For CT images of wild-type EGFR, we use their mutation status to initialize the pseudo labels of each instance, which conforms to the constraints in multi-instance learning. According to formula 0, it can be understood that the true labels of all image blocks in the lung CT images without EGFR mutations are non-mutated image blocks. For CT images with EGFR mutations, due to the heterogeneity within the tumor, all image blocks in the lung CT images may contain both mutated image blocks and non-mutated image blocks. Because we first calculate the importance score of each image block instance to indicate the contribution of each image block instance to the entire region of interest image, and then scale the image blocks in the region of interest image by formula 4, and obtain soft pseudo labels from the importance scores.

[0081] In this embodiment, soft pseudo labels are generated based on the importance scores of instance features, which can further enhance the learning ability of instance features.

[0082] In one embodiment of the present application, S530 includes, that is, creating a spatial perception Transformer module with a multi-layer encoder structure, based on the spatial perception Transformer module with a multi-layer encoder structure, mining the spatial relative position relationship between different image blocks, and obtaining the comprehensive features of each image block, including:

[0083] S531, create τ transformer encoder layers. τ is a positive integer and τ is greater than 2.

[0084] S532, selecting a region of interest image.

[0085] S533, inputting feature coding information of all image blocks of an image of a region of interest into a first transformer encoder layer, and calculating an output result of the first transformer encoder layer according to Formula 4.

[0086]

[0087] in, is the output of the first transformer encoder layer, x′ i,j For image block x i,j The feature encoding information of is, i is the serial number of the image of the region of interest, j is the serial number of the image block, is the class label feature of the i-th region of interest image, E c is the linear projection matrix, P i ab Encodes the absolute position of the i-th region of interest image.

[0088] S534, input the output result of the previous transformer encoder layer to the next transformer encoder layer, and calculate the encoding intermediate amount of the next transformer encoder layer according to Formula 5.

[0089]

[0090] in, is the encoded intermediate quantity for the next transformer encoder layer, is the output of the previous transformer encoder layer, SA_SRPE[·] is the self-attention expression including spatial relative position encoding, for The output after the first normalization layer in the next Transformer Encoder layer.

[0091] S535, calculate the output result of the next transformer encoder layer according to Formula 6.

[0092]

[0093] in, is the output of the next transformer encoder layer, is the encoded intermediate quantity for the next transformer encoder layer, for The output after the second normalization layer in the next transformer encoder layer, for The output of the feed-forward network layer after the next transformer encoder layer.

[0094] S536, return to S534, that is, return to the step of inputting the output result of the previous transformer encoder layer to the next transformer encoder layer, calculating the encoding intermediate amount of the next transformer encoder layer according to Formula 5, until the output result of the last transformer encoder is obtained, and the output result of the last transformer encoder is used as the image block comprehensive feature set of the region of interest image i. Each element in the image block comprehensive feature set is a comprehensive feature of each image block of the region of interest image i.

[0095] S537, returning to S532, that is, returning to the step of selecting an image of a region of interest, until all images of the region of interest are selected.

[0096] Specifically, the comprehensive feature of each image block of the region of interest image i finally obtained in S536 is h i,j A comprehensive feature set H of the region of interest image i can be generated i .H i ={h i,1 ,h i,2 ,...,...h i,j},H i ∈R M×dx .H i ∈A matrix with M rows and dx columns. R M×dx It means a matrix with M rows and dx columns. If there is no special explanation, the symbol R at the end means the matrix.

[0097] Each transformer encoder layer of the encoder includes a feedforward network (FFN), two layers of normalization (LN), and a self-attention layer with spatial relative position encoding. The specific connection relationship is: first normalization layer-SRP layer-second normalization layer-feedforward network layer.

[0098] Formula 5 is performed through the first normalization layer and the SRP layer, and Formula 6 is performed through the second normalization layer and the feedforward network layer. The SRP layer is also called the SA_SRPE layer.

[0099] To obtain the overall characteristics of heterogeneous tumors, radiologists often consider the relationship between different pathological subregions of CT images. For example, the spatial relationship between the tumor and the surrounding areas, such as peritumoral pleural traction or vascular convergence, is crucial for inferring the elusive EGFR mutation status in complex CT images. Therefore, failure to capture the relationship between instances may limit the overall representation of the tumor. In this embodiment, the representation ability of heterogeneous tumors can be improved by designing a spatially aware Transformer to actively explore the dependencies between instances.

[0100] In one embodiment of the present application, S530 also includes, that is, creating a spatial perception Transformer module with a multi-layer encoder structure, mining the spatial relative position relationship between different image blocks based on the spatial perception Transformer module with a multi-layer encoder structure, and obtaining the comprehensive features of each image block, and further includes:

[0101] S538, construct a self-attention expression including spatial relative position encoding to couple the spatial relative position encoding information and self-attention. The self-attention expression including spatial relative position encoding is shown in Formula 7.

[0102]

[0103] where SA_SRPE[·] is the self-attention expression including the spatial relative position encoding, X′ i is a set of feature encoding information of each image block in the i-th region of interest image, x′ i,j For image block x i,j The feature encoding information of , S(·) is the expression of the softmax function, W Q is the query parameter projection matrix, WK is the parameter projection matrix of the key, W V is the parameter projection matrix of the value, M is the number of image blocks contained in a single region of interest image, dx is the first preset matrix attribute parameter, H is the height dimension symbol, W is the width dimension symbol, x′ i,A and x′ i,B is the feature encoding information of the randomly selected image block A and image block B in the same region of interest image i, T is the transposed matrix symbol, is the height parameter matrix, is the width parameter matrix, is a learnable encoding of the relative height of image patch A and image patch B, is a learnable encoding of the relative widths of image patch A and image patch B.

[0104] Specifically, the extracted patches,image blocks are first arranged according to their spatial positions, and their relative heights and widths are encoded according to the spatial information of the tumor CT image. Next, the relative position encoding information between two random instances is used to embed the attention features.

[0105] S538 can be executed at the beginning of the entire S500. S538 can also be executed at the end of the entire S500 by substituting Formula 7 into Formula 5. The purpose of Formula 7 is to define SA_SRPE[·] in Formula 5.

[0106] One limitation of the ordinary Transformer adopted by existing methods is that it only adopts absolute position encoding, which may not fully capture subtle position dependency information, which is indispensable in exploring the spatial relationship between highly complex tumors and their surrounding areas. To overcome this challenge, we integrate spatial relative position encoding into the self-attention mechanism to promote the exploration of relationships between instances. We first arrange the extracted patches and image blocks according to their spatial positions and encode their relative height and width according to the spatial information of the tumor CT image. Then, by randomly selecting two image blocks: image block A and image block B, their direct relative forged encoding information is embedded into the attention feature. W Q is the query parameter projection matrix, W K is the parameter projection matrix of the key, W V is the parameter projection matrix of the value, corresponding to the query, key and value in the attention mechanism respectively.

[0107] In this embodiment, the spatial relative position encoding is integrated into the self-attention mechanism, which can effectively promote the exploration of the relationship between instances. Different from the absolute position encoding that only captures static position information and implicitly implies spatial relationships, the spatial relative position encoding we proposed is good at flexibly and explicitly mining the dependencies between instances, ensuring that key spatial relationships are not ignored, so that the model can skillfully identify potential interactions between tumors and surrounding peritumoral areas, thereby improving the characterization ability of heterogeneous tumors and improving the prediction accuracy of EGFR mutations.

[0108] In one embodiment of the present application, S521 includes, that is, generating the importance score of each image block includes:

[0109] S521a, generating a weight coefficient for each image block according to Formula 8.

[0110]

[0111] Among them, a i,j For image block x i,j , W is the first learnable parameter matrix, U is the second learnable parameter matrix, μ is the third learnable parameter matrix, dx is the first preset matrix attribute parameter, dz is the second preset matrix attribute parameter, ⊙ is the element multiplication symbol, σ is the expression symbol of the S-type activation function, For image block x i,j The comprehensive characteristics of h i,j The transpose of .

[0112] S521b, generating an importance score for each image block according to Formula 8.

[0113]

[0114] in, For image block x i,j The importance score, a i,j For image block x i,j The weight coefficient is , j is the sequence number of the image block, i is the sequence number of the region of interest image, and M is the number of image blocks contained in a single region of interest image.

[0115] Specifically, due to the intra-tumor heterogeneity of NSCLC, the mutation status of EGFR varies significantly between different tumor subregions, resulting in uneven contributions to the overall EGFR mutation status of the patient. Therefore, a key aspect when generating bag-level features is to prioritize informative instances to ensure that the aggregated CT image features produce a high signal-to-noise ratio, thereby maintaining robustness to unseen tumors. Otherwise, the prediction model may rely on uninformative noisy instances, resulting in degraded generalization performance on unseen samples. In this method, this problem is addressed by an adaptive gating procedure that identifies valuable instances from different tumor subregions and highlights their contribution to feature aggregation.

[0116] In this embodiment, the comprehensive feature set H of the region of interest image i obtained based on the above content i By introducing the adaptive gated instance aggregation module from H i By learning the feature importance scores of specific image blocks, the generalization ability of the model can be improved.

[0117] In one embodiment of the present application, S521 further includes, that is, generating the importance score of each image block further includes:

[0118] S521c, generating a weight coefficient vector set a based on the weight coefficients of each image block in each region of interest image i , a i ={a i,1 ,a i,2 ,a i,3 ,...,...a i,j}.

[0119] S521d, mapping the weight coefficient vector set into a unit vector space with L2 normalization to generate a standardized scalar set v i , v i The expression of is shown in formula 10.

[0120]

[0121] Among them, v i is a standardized scalar set, a i is a weight coefficient vector set, j is the serial number of the image block, i is the serial number of the region of interest image, and M is the number of image blocks contained in a single region of interest image.

[0122] S521e, converting the standardized scalar set into a modified importance score set according to formula 11.

[0123]

[0124] in, is the image block importance score set of the i-th region of interest image, v i is a standardized scalar set and ⊙ is the element-wise multiplication symbol.

[0125] Specifically, one limitation of the adaptive gating module is that the softmax function may not sufficiently sharpen the importance scores. Due to the high heterogeneity of tumors, this may hinder the signal-to-noise ratio of the bag feature representation aggregated in the feature space, because the mutation status informed by the positive instances (positive instances: refers to image patches with mutations) and their relationships may be overwhelmed by the useless information in the large number of negative instances (negative instances: refers to image patches without mutations) from the background or tumor subregions with insufficient information (such as stromal tissue). Therefore, in order to promote the learning of instance scores by the adaptive gating module in this embodiment, we first map the set of weight coefficient vectors into a unit vector space with L2 normalization, v i Each element of represents a normalized scalar indicating the relative contribution of the corresponding instance.

[0126] Next, we convert these elements into importance scores in the range of [0,1] using Formula 11, calculate the modified importance score vector by applying the Hadamard product, and use the modified importance score vector as the image block importance score set.

[0127] In this embodiment, compared with a single softmax function, the processing method of this embodiment can better amplify the variance of the importance score, thereby encouraging the gating mechanism to focus more on information instances that are valuable for feature aggregation.

[0128] In one embodiment of the present application, S540 includes, that is, the adaptive gated instance aggregation module aggregates the comprehensive features of each image block and the importance score of each image block in a dynamic weighted manner to obtain the package-level CT image features of each region of interest image, including:

[0129] S541, obtaining the package-level CT image features of each region of interest image according to formula 12.

[0130]

[0131] in, is the packet-level CT image feature of the image block of the i-th region of interest image, For x i,j The importance score of the image block, h i,j For image block x i,j comprehensive characteristics.

[0132] Specifically, in this way, the packet-level CT image features of each region of interest image may include both the comprehensive features of the image block and the importance score of the image block.

[0133] The adaptive gated instance aggregation module proposed in this embodiment can effectively emphasize the contribution of valuable instance features and reduce the adverse effects of noise instances in heterogeneous tumors.

[0134] In one embodiment of the present application, S550 includes, that is, the prediction loss function of the package feature, the prediction loss function of the class label feature, and the supervision loss function are constructed separately, including:

[0135] S551, obtaining the true label of each region of interest image. The true label includes the actual gene mutation prediction result.

[0136] S552, define an expression of a loss function, the expression of the loss function is shown in Formula 13.

[0137]

[0138] Where N is the number of lung CT images, Y i is the true label of the image in the region of interest, λ is the weighting factor to balance the positive and negative samples, is a factor that balances the weight of difficult samples, ξ is a smoothing parameter, is the true label after smoothing, P i is the prediction result of the image of the region of interest.

[0139] S553, construct a prediction loss function of the class label feature, and the expression of the prediction loss function of the class label feature is shown in Formula 14.

[0140]

[0141] in, is the prediction loss function of the class label feature, is the true label after smoothing, P i cla is the prediction result of the class label feature.

[0142] S554, construct a prediction loss function of the packet feature. The expression of the prediction loss function of the packet feature is shown in Formula 15.

[0143]

[0144] in, is the prediction loss function of the package feature, is the true label after smoothing, is the prediction result of the package feature.

[0145] S555, construct a supervised loss function, the expression of which is shown in Formula 16.

[0146]

[0147] in, is the supervised loss function, N is the number of region of interest images, M is the number of image patches contained in a single region of interest image, For image block x i,j Soft fake labels, For image block x i,j The importance score of .

[0148] Specifically, the loss function introduced in this embodiment is a focal loss function, which can eliminate the class imbalance problem that may exist in EGFR prediction, and adjust the focus on samples of different categories through weighting factors, thereby improving the robustness of the model.

[0149] Formula 14 and Formula 15 use the true label. The lung CT images used for training are already mature cases, so they already have mutation results and therefore have true labels. i cla is the prediction result of the class label feature. The prediction result of the class label feature is the model based on the class label feature. The prediction result obtained in formula 15 The prediction result of the package feature is the prediction result of the model based on the package feature. The prediction results are obtained. Package features are package-level CT image features

[0150] The prediction results of the class marker features are obtained through the MLP (MLP, Multilayer Perceptron) module in the model, and the prediction results of the package features are obtained through the package classifier head in the model. How the prediction results of the class marker features and the prediction results of the package features are specifically obtained is not the focus of protection of this application. The principle also adopts the more existing and common means in artificial intelligence machine learning, which will not be described in detail here.

[0151] In addition, label smoothing technology is used to further improve the generalization ability of the model. to achieve.

[0152] The supervised loss function can enhance the model's ability to capture subtle features, explicitly supervise and enhance instance feature learning of tumor CT images, and cooperate with pseudo labels to provide additional supervisory signals, effectively improving the overall performance of EGFR mutation prediction.

[0153] In one embodiment of the present application, S560 includes, that is, constructing an overall loss function using the prediction loss function of the package feature, the prediction loss function of the class label feature, and the supervision loss function, including:

[0154] S561, construct an overall loss function according to Formula 17.

[0155]

[0156] in, is the overall loss function, is the prediction loss function of the package feature, is the prediction loss function of the class label feature, is the supervised loss function, α is the first trade-off factor, and β is the second trade-off factor.

[0157] Specifically, α is the first trade-off factor, and β is the second trade-off factor, both of which can be preset constants. The purpose of the k-times training of the model is to minimize the value of the overall loss function after each training, and finally minimize the value of the overall loss function after the last training.

[0158] In S570, in the next training, minimizing the value of the overall loss function is used as the training goal, that is, the training goal is: in the process of k training times, making the value of the overall loss function smaller and smaller.

[0159] In this embodiment, we use class marker features and package features, as well as supervisory factors, to comprehensively predict the EGFR mutation status, and the model has high prediction accuracy.

[0160] The technical features of the above-described embodiments may be arbitrarily combined, and the execution order of the method steps is not limited. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0161] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology, characterized in that: The method comprises: Obtain multiple lung CT images; Each lung CT image is preprocessed to obtain an image of a region of interest in each lung CT image; Divide each region of interest image into multiple image blocks of equal size; Create an EGFR mutation prediction model; The EGFR mutation prediction model is trained k times to obtain a trained EGFR mutation prediction model; k is a positive integer and k is greater than 1; Acquiring a lung CT image to be tested, and inputting the lung CT image to be tested into the trained EGFR mutation prediction model; Starting the trained EGFR mutation prediction model to obtain a prediction result output by the trained EGFR mutation prediction model; The training of the EGFR mutation prediction model for k times to obtain the trained EGFR mutation prediction model comprises: Using the feature encoder to encode all image blocks one by one, and obtaining a preliminary prediction result of each image block based on the encoding result of each image block; The importance score of each image block is embedded in the preliminary prediction result of each image block to generate a soft pseudo label for each image block; Create a spatial perception Transformer module with a multi-layer encoder structure. Based on the spatial perception Transformer module with a multi-layer encoder structure, the spatial relative position relationship between different image blocks is mined to obtain the comprehensive features of each image block. Based on the adaptive gated instance aggregation module, the comprehensive features of each image block and the importance score of each image block are aggregated using a dynamic weighting method to obtain the packet-level CT image features of each region of interest image; Construct the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function respectively; Construct an overall loss function using the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function; In the next training, the value of the overall loss function is minimized as the training goal.

2. The lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology according to claim 1, characterized in that: The method of encoding all image blocks one by one using a feature encoder and obtaining a preliminary prediction result of each image block according to the encoding result of each image block includes: Based on formula 1, all image blocks are encoded one by one using a feature encoder to generate feature encoding information for each image block; x' i,j =F r (x i,j :θ r ) Formula 1; Among them, x' i,j For image block x i,j The feature encoding information, x i,j is the jth image block in the i-th region of interest image, i is the serial number of the region of interest image, j is the serial number of the image block, F r (·:θ r ) is the expression symbol of the coding behavior; Based on formula 2, the feature encoding information of each image block is input into the instance prediction head for processing, and the preliminary prediction result of each image block output by the instance prediction head is obtained; in, For image block x i,j The preliminary prediction results, x' i,j For image block x i,j The feature encoding information, ψ p (·:θ p ) is a symbol for expressing the processing behavior performed by the instance prediction head.

3. The lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology according to claim 2 is characterized in that: The step of embedding the importance score of each image block in the preliminary prediction result of each image block and generating a soft pseudo label of each image block comprises: Generate an importance score for each image patch; Creating an image block importance score set, and including all image block importance scores into the image block importance score set one by one; Based on the image block importance score set, use Formula 3 to generate a soft pseudo label package for each image block; in, is the soft pseudo label of image block xi,j, is the importance score of image block xi,j, is the image block importance score set of the i-th region of interest image, is the indicator function, k is the symbol expressing the current number of training times, Y i is the true label of the image of the region of interest.

4. The lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology according to claim 3 is characterized in that: The spatial perception Transformer module with a multi-layer encoder structure is created. Based on the spatial perception Transformer module with a multi-layer encoder structure, the spatial relative position relationship between different image blocks is mined to obtain the comprehensive features of each image block, including: Create τ transformer encoder layers; τ is a positive integer and τ is greater than 2; Select an area of ​​interest image; Input the feature encoding information of all image blocks of an area of ​​interest image into the first transformer encoder layer, and calculate the output result of the first transformer encoder layer according to Formula 4; in, is the output of the first transformer encoder layer, x' i,j For image block x i,j The feature encoding information of is, i is the serial number of the image of the region of interest, j is the serial number of the image block, is the class label feature of the i-th region of interest image, E c is the linear projection matrix, P i ab Encode the absolute position of the i-th region of interest image; The output result of the previous transformer encoder layer is input to the next transformer encoder layer, and the encoding intermediate quantity of the next transformer encoder layer is calculated according to Formula 5; in, is the encoded intermediate quantity for the next transformer encoder layer, is the output of the previous transformer encoder layer, SA_SRPE[·] is the self-attention expression including spatial relative position encoding, for The output after the first normalization layer in the next transformer encoder layer; Calculate the output of the next transformer encoder layer according to Formula 6; in, is the output result of the next transformer encoder layer, is the encoded intermediate quantity for the next transformer encoder layer, for The output after the second normalization layer in the next transformer encoder layer, for The output after the feed-forward network layer in the next transformer encoder layer; Return the output result of the previous transformer encoder layer to the next transformer encoder layer, calculate the encoding intermediate amount of the next transformer encoder layer according to Formula 5, until the output result of the last transformer encoder is obtained, and use the output result of the last transformer encoder as the image block comprehensive feature set of the region of interest image i; each element in the image block comprehensive feature set is a comprehensive feature of each image block of the region of interest image i; Return to the step of selecting a region of interest image until all region of interest images have been selected.

5. The method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology according to claim 4, characterized in that: The creating of a spatial perception Transformer module with a multi-layer encoder structure, mining the spatial relative position relationship between different image blocks based on the spatial perception Transformer module with a multi-layer encoder structure, and obtaining the comprehensive features of each image block, further includes: A self-attention expression including spatial relative position encoding is constructed to couple the spatial relative position encoding information and self-attention. The self-attention expression including spatial relative position encoding is shown in Formula 7. Where SA_SRPE[·] is the self-attention expression including the spatial relative position encoding, X' i is a set of feature encoding information of each image block in the i-th region of interest image, x' i,j For image block x i,j The feature encoding information of , S(·) is the expression of the softmax function, W Q is the query parameter projection matrix, W K is the parameter projection matrix of the key, W V is the parameter projection matrix of the value, M is the number of image blocks contained in a single region of interest image, dx is the first preset matrix attribute parameter, H is the height dimension symbol, W is the width dimension symbol, x' i,A and x' i,B is the feature encoding information of the randomly selected image block A and image block B in the same region of interest image i, T is the transposed matrix symbol, is the height parameter matrix, is the width parameter matrix, is a learnable encoding of the relative height of image patch A and image patch B, is a learnable encoding of the relative widths of image patch A and image patch B.

6. The lung cancer EGFR gene mutation prediction method based on multi-instance learning and Transformer technology according to claim 5, characterized in that: Generating the importance score of each image block includes: Generate the weight coefficient of each image block according to formula 8; Among them, a i,j For image block x i,j , W is the first learnable parameter matrix, U is the second learnable parameter matrix, μ is the third learnable parameter matrix, dx is the first preset matrix attribute parameter, dz is the second preset matrix attribute parameter, ⊙ is the element multiplication symbol, σ is the expression symbol of the S-type activation function, For image block x i,j The comprehensive characteristics of h i,j The transpose of Generate the importance score of each image block according to formula 8; in, For image block x i,j The importance score, a i,j For image block x i,j The weight coefficient is , j is the sequence number of the image block, i is the sequence number of the region of interest image, and M is the number of image blocks contained in a single region of interest image.

7. The method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology according to claim 6, characterized in that: Generating the importance score of each image block also includes: A weight coefficient vector set a is generated based on the weight coefficients of each image block in each region of interest image i , a i ={a i,1 ,a i,2 ,a i,3 ,...,...a i,j }; Map the weight coefficient vector set into a unit vector space with L2 normalization to generate a standardized scalar set v i , v i The expression of is shown in formula 10; Among them, v i is a standardized scalar set, a i is a weight coefficient vector set, j is the serial number of the image block, i is the serial number of the region of interest image, and M is the number of image blocks contained in a single region of interest image; According to formula 11, the standardized scalar set is converted into a modified importance score set; in, is the image block importance score set of the i-th region of interest image, v i is a standardized scalar set and ⊙ is the element-wise multiplication symbol.

8. The method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology according to claim 7, characterized in that: The adaptive gated instance aggregation module aggregates the comprehensive features of each image block and the importance score of each image block in a dynamic weighted manner to obtain the package-level CT image features of each region of interest image, including: The packet-level CT image features of each region of interest image are obtained according to formula 12; in, is the packet-level CT image feature of the i-th region of interest image, For image block x i,j The importance score, h i,j For image block x i,j comprehensive characteristics.

9. The method for predicting lung cancer EGFR gene mutation based on multi-instance learning and Transformer technology according to claim 8, characterized in that: The respectively constructing the prediction loss function of the package feature, the prediction loss function of the class label feature, and the supervision loss function include: Get the true label of each region of interest image; the true label contains the actual gene mutation prediction result; Define the expression of the loss function. The expression of the loss function is shown in Formula 13; Where N is the number of lung CT images, Y i is the true label of the image in the region of interest, λ is the weighting factor for balancing positive and negative samples, is a factor that balances the weight of difficult samples, ξ is a smoothing parameter, is the true label after smoothing, P i is the prediction result of the image of the region of interest; Construct a prediction loss function of the class label feature. The expression of the prediction loss function of the class label feature is shown in Formula 14; in, is the prediction loss function of the class label feature, is the true label after smoothing, P i cla is the prediction result of the class label feature; Construct a prediction loss function of the packet feature. The expression of the prediction loss function of the packet feature is shown in Formula 15. in, is the prediction loss function of the package feature, is the true label after smoothing, is the prediction result of the package feature; Construct a supervised loss function. The expression of the supervised loss function is shown in Formula 16. in, is the supervised loss function, N is the number of region of interest images, M is the number of image patches contained in a single region of interest image, For image block x i,j Soft fake labels, For image block x i,j The importance score of .

10. The lung cancer EGFR gene mutation prediction method based on multiple example learning and Transformer technology according to claim 9, characterized in that: The method of constructing an overall loss function using the prediction loss function of the bag feature, the prediction loss function of the class label feature, and the supervision loss function includes: Construct the overall loss function according to formula 17; in, is the overall loss function, is the prediction loss function of the package feature, is the prediction loss function of the class label feature, is the supervised loss function, α is the first trade-off factor and β is the second trade-off factor.

Citation Information

Patent Citations

  • EGFR gene mutation detection method and system based on lung CT image

    CN115861303A

  • Multi-modal image-based glioma patient prognosis lifetime prediction method and system

    CN116530965A

  • Breast cancer full-slice image classification method combining self-supervision and weak supervision learning

    CN117237733A

  • Auxiliary detection system for non-small cell lung cancer histological image EGFR gene mutation

    CN117408997A

  • EGFR (epidermal growth factor receptor) gene mutation state prediction method, system, equipment and medium

    CN118098360A