AI pulmonary epithelial cell region identification method based on MAE training ViT and SAM network

Through the combination of MAE pre-trained ViT model and SAM network, the problem of insufficient labeling data in the recognition of lung epithelial cell regions is solved, and efficient and accurate lesion recognition is achieved, reducing the amount of calculation and improving the recognition effect of the model.

CN120471826APending Publication Date: 2025-08-12WUXI APPTEC (SHANGHAI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510380347.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art cannot effectively predict designated lesion targets on H&E stained sections, especially the recognition effect of lung epithelial cell regions is limited, and the trainingable annotation data is very limited, resulting in pathological diagnostic accuracy and inefficiency.

Method used

The MAE pre-trained ViT model was used to pre-train the unlabeled lung H&E stained sections, embedded the SAM model, and fine-tuned it with a small amount of annotated data, and feature extraction and recognition were performed using the encoder-decoder structure and hierarchical learning rate.

Benefits of technology

Efficient identification of lung epithelial cell region at low labeling costs is achieved, reducing calculation amount and improving identification accuracy and accuracy, reducing labor costs and improving the practicality and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471826A_ABST
    Figure CN120471826A_ABST
Patent Text Reader

Abstract

The invention discloses an AI lung epithelial cell region identification method based on an MAE pre-training ViT and SAM network, and the method is characterized in that the method comprises the following steps: S1, employing an unlabeled lung Hamp; e, pre-training a ViT model by adopting MAE for the dyed section; s2, a backbone network in the pre-trained ViT model is embedded into an SAM model; s3, the marked lung Hamp is utilized; e, performing fine adjustment on the SAM model by the dyeing section; and S4, realizing pulmonary epithelial cell region identification through the fine-tuned SAM model. The method is applied to pulmonary epithelial cell region identification for the first time, efficient focus identification is achieved under limited annotation data, and the identification prediction effect of the model is ensured while the calculated amount is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of biology and artificial intelligence, and in particular to a method for applying artificial intelligence to identify lung epithelial cell regions. Background Art

[0002] The identification of lung epithelial cells is of great significance in pathology, with applications in drug development and toxicity assessment, tissue repair and regeneration, and the study of disease and physiological mechanisms. Furthermore, for pathologists, dynamic lesion identification within a designated area is more convenient in practical applications. By adjusting the target area, doctors can precisely identify lesions in a specific region, which plays a crucial role in improving diagnostic accuracy and efficiency.

[0003] However, existing technologies face two major difficulties. First, current interactive segmentation networks (such as SAM and MedSAM) cannot effectively predict designated lesion targets on H&E stained sections, and the prediction effect on epithelial cell areas is limited. Second, in actual work, the available training annotated data is very limited. Therefore, how to achieve efficient lesion recognition with limited annotated data becomes a key issue. These challenges urgently need new methods and technologies to address in order to improve the accuracy and efficiency of pathological diagnosis. Summary of the Invention

[0004] To solve at least one of the above technical problems, the present invention provides an AI lung epithelial cell region recognition method based on MAE pre-trained ViT and SAM network, comprising the following steps:

[0005] S1: Using unlabeled lung H&E stained sections to pre-train the ViT model based on MAE;

[0006] S2: Embed the backbone network in the pre-trained ViT model into the SAM model;

[0007] S3: Fine-tuning the SAM model using annotated lung H&E-stained sections;

[0008] S4: Lung epithelial cell region recognition via a fine-tuned SAM model.

[0009] In some embodiments, step S1 uses unlabeled lung H&E stained sections to pre-train the ViT model using MAE, and the specific steps are:

[0010] S11: Segment the unlabeled H&E-stained lung sections into fixed-size image blocks, randomly select a portion of the image blocks for masking, divide the image blocks into masked image blocks and visible image blocks, and convert each image block into an embedding vector;

[0011] S12: Input the visible image block into the encoder, the encoder processes the embedding vector of the visible image block and performs absolute position encoding, and outputs the potential representation;

[0012] S13: Input the potential representation and the mask mark into the decoder, and the decoder reconstructs the image based on the potential representation and the mask mark;

[0013] S14: Calculate the loss between the reconstructed image and the original image;

[0014] S15: Update the model’s parameters through backpropagation to minimize the loss.

[0015] In some embodiments, the backbone network in the pre-trained ViT model of step S2 is embedded in the SAM model, specifically in the following steps:

[0016] S21: Load the ViT model pre-trained by MAE, extract its Attention layer and FFN layer and embed them into SAM as the core module for feature extraction;

[0017] S22: The position encoding adopts the relative position encoding of SAM, and the FFN layer adopts the structure of the ViT model pre-trained by MAE.

[0018] In some embodiments, the fine-tuning in step S3 refers to fine-tuning using labeled lung H&E stained sections by setting a learning rate and in an interactive manner, and the range of the learning rate is 1e-6.

[0019] In some embodiments, the loss of step S14 is calculated by a loss function, which is

[0020] Among them, y j represents the true value of the jth sample, It represents the predicted value of the model for the jth sample, and M represents the total number of samples.

[0021] In some embodiments, pre-training is updated using a layer-wise learning rate, where the layer-wise learning rate is 1e-6 for the encoder part and 1e-4 for other parts.

[0022] In some embodiments, the MedSAM model is used instead of the SAM model.

[0023] In some embodiments, S11 further includes the following steps:

[0024] S111: Convert the H&E-stained lung sections from MPP=0.136 to MPP=0.5.

[0025] S112: The H&E stained lung sections were cut from the upper left corner to a size of (1024, 1024), with the batch size set to 4 and the weight decay set to 0.01.

[0026] In some embodiments, the effect identified in S4 is evaluated by calculating the Dice coefficient, which is

[0027] Among them, |A∩B| represents the number of pixels where the prediction result completely overlaps with the true annotation, |A| represents the total number of pixels in the true annotation area, and |B| represents the total number of pixels in the model prediction area.

[0028] In some embodiments, the MAE includes an asymmetric encoder-decoder architecture, where the decoder is more lightweight than the encoder.

[0029] In some embodiments, the random mask ratio of S11 is 75%.

[0030] The MAE in this invention is called Masked Autoencoder in English. It is a deep learning method based on the autoencoder architecture for self-supervised learning. The core idea is to force the model to learn efficient representation of the image by masking image blocks and reconstructing the original pixels, thereby completing pre-training without manually annotating data.

[0031] ViT in this invention stands for Vision Transformer, which is a deep learning model architecture for computer vision tasks. It introduces the Transformer architecture widely used in the field of natural language processing into image processing to extract image features and complete various visual tasks such as image classification, object detection and segmentation.

[0032] The Transformer architecture in this invention is a deep learning architecture based on the self-attention mechanism.

[0033] The SAM in this invention is called SegmentAnything Model in English. It is an advanced computer vision model that aims to achieve the segmentation of any object in images and videos. It consists of an image encoder, a hint encoder, and a mask decoder.

[0034] The H&E stained sections in the present invention refer to tissue sections treated with hematoxylin and eosin (H&E) staining technology, which are used to observe and analyze the structure and morphology of tissues or cells under a microscope.

[0035] The Attention layer in this paper refers to the attention layer, which is a key component in deep learning models. It is used to dynamically focus on the most important parts of the input data. It is the core part of the Transformer encoder. It enhances the expressiveness of the model through the multi-head attention mechanism, enabling the model to capture long-distance dependencies between different areas in the image.

[0036] The FFN layer in the present invention stands for Feed-Forward Network layer in English. It is a feed-forward neural network layer, which is a key component in the Transformer architecture and is used to perform nonlinear transformation on the feature vector at each position.

[0037] The full name of MPP in the present invention is Microns Per Pixel, which means microns per pixel, indicating the actual physical size corresponding to each pixel in the image, with the unit being microns, μm.

[0038] The batch size in this invention refers to the number of samples that the model processes simultaneously during each iterative training.

[0039] Weight Decay in this invention is weight attenuation, which refers to a regularization technique that penalizes the absolute value of model parameters by adding an L2 regularization term to the loss function to prevent overfitting.

[0040] MedSAM in this invention, whose full name is Medical Segmentation Assistant Model, is a deep learning model designed specifically for medical image segmentation. It aims to process anatomical structures and lesions in various medical imaging modes through universal segmentation capabilities.

[0041] Epoch in the present invention is a training round, that is, a complete forward and backward propagation process on the entire data set.

[0042] Compared with the prior art, the beneficial effects of the present invention are embodied in:

[0043] This experiment is based on a large number of unlabeled lung H&E-stained sections and a small number of lung H&E-stained sections with epithelial cell areas labeled. It achieves the goal of training an accurate and interactive epithelial cell recognition model and tool at the lowest possible labeling cost. Reducing the amount of labeling can significantly reduce labor costs and improve efficiency. In addition, this design of training the model on a large number of unlabeled images allows the model to learn a large amount of potential information, thereby greatly reducing the amount of computation while ensuring the model's recognition and prediction effect.

[0044] Furthermore, the present invention masks part of the input image and then uses an encoder-decoder structure to reconstruct these masked parts. The encoder only processes the unmasked parts, while the decoder is responsible for reconstructing the complete image. Through this method, a backbone network with image coding capabilities can be trained without labeling.

[0045] Furthermore, the MAE architecture consists of an asymmetric encoder-decoder, where the encoder only processes visible image blocks, while the lightweight decoder reconstructs the image based on the latent representation and mask labels. This design greatly reduces the amount of computation, thereby ensuring the model effect, allowing it to be trained on massive unlabeled images, thereby enabling the model to learn a large amount of potential information.

[0046] Furthermore, the present invention embeds the above-mentioned backbone network into the SAM model and fine-tunes the SAM model using a small amount of labeled data, thereby achieving interactive recognition of epithelial cell areas in H&E-stained sections. This method not only reduces the labeling cost but also improves the recognition accuracy of the model, making it more practical in actual applications.

[0047] Furthermore, by using a high ratio (75%-80%) of random masks, MAE effectively reduces the information redundancy in the image, creating a more challenging self-supervised task and prompting the model to learn more useful and generalized features, thereby improving prediction accuracy and precision.

[0048] Furthermore, the position encoding of the SAM model follows the relative position encoding of SAM, thereby weakening the absolute correlation of positions, improving generalization, and thus improving the prediction accuracy of the model.

[0049] Furthermore, during pre-training, a hierarchical learning rate is adopted, with the encoder part adopting a learning rate of 1e-6 and the other parts being updated at a learning rate of 1e-4. This not only retains some feature recognition capabilities learned through pre-training, but also enhances the accuracy of the model through manual labeling.

[0050] In another aspect, the present invention discloses a computing system, wherein the computing system is configured to execute the above method.

[0051] The present invention also discloses a computer-readable storage medium, which stores instructions that can be executed by one or more processors of a computing system to implement the above method.

[0052] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is the curve of MAE training loss changing with the number of iterations in Example 1.

[0054] Figure 2 This is a diagram showing the prediction effect of the MAE pre-trained model on image A in Example 2.

[0055] Figure 3 This is a diagram showing the prediction effect of the MAE pre-trained model on image B in Example 2.

[0056] Figure 4 This is the prediction result of the control group A for image C in Example 3.

[0057] Figure 5 This is the prediction result of the experimental group for image C in Example 3.

[0058] Figure 6 This is the prediction result of the control group A for image D in Example 3.

[0059] Figure 7 It is the prediction result of the experimental group for image D in Example 3. DETAILED DESCRIPTION

[0060] In order to make the technical means, creative features, objectives and effects of the invention easier to understand, the present invention is further explained with reference to specific illustrations, but the present invention is not limited to the following implementation cases.

[0061] It should be noted that the structures, proportions, sizes, etc. illustrated in the drawings in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them. They are not used to limit the conditions under which the present invention can be implemented. Therefore, they have no substantive technical significance. Any modification of the structure, change in the proportion relationship or adjustment of the size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.

[0062] Example 1

[0063] This experiment was carried out through the following steps:

[0064] The ViT model was pre-trained using the MAE method using 100 unlabeled lung H&E stained sections. Specifically, the unlabeled lung H&E stained sections were divided into fixed-size image blocks, converted from MPP=0.136 to MPP=0.5, cut from the upper left corner to a size of (1024,1024), and 75% of the image blocks were randomly selected for masking. The image blocks were divided into 75% masked image blocks and 25% visible image blocks. Each image block was converted into an embedding vector, and the 25% visible image blocks were input into the encoder. The learning rate of the encoder was set to 1e-6, which only converted the embedding vector of the visible image block and performed absolute position encoding. The latent representation was output, and the latent representation and mask label were input into the decoder. The learning rate of the decoder was set to 1e-4, and the decoder reconstructed the image based on the latent representation and mask label. The loss between the reconstructed image and the original image was calculated by first normalizing the masked pixels and then calculating the L2 distance, as shown in the following example: Figure 1 As shown in the figure, the loss quickly converges to 0.2 during the pre-training process.

[0065] The backbone network in the pre-trained ViT model is embedded in the SAM model. Specifically, the ViT model pre-trained by MAE is loaded, and its Attention layer and FFN layer are extracted and embedded into the SAM model as the core module of feature extraction. The position encoding adopts the relative position encoding of SAM, and the FFN layer adopts the structure of the VIT model pre-trained by MAE.

[0066] The SAM model was fine-tuned using labeled lung H&E-stained sections. Specifically, the SAM model was fine-tuned using 10 labeled lung H&E-stained sections, and the learning rate was set to 1e-6.

[0067] Finally, the fine-tuned SAM model was applied to lung epithelial cell region recognition.

[0068] Example 2

[0069] In order to explore the effect of pre-training, the MAE pre-trained model is used to predict the effect of the masked image.

[0070] The prediction effect of the MAE pre-trained model on image A and image B is as follows Figure 2 、 Figure 3 As shown, Figure 2 、 Figure 3 The upper left of the image is the original image, the upper right is the masked image, the lower left is the image generated using the masked image, and the lower right is the joint display of the generated image and the prompt image. The closer the upper left and lower right are, the stronger the model capability is. Figure 2 、 Figure 3 It can be seen that the MAE pre-trained model has strong understanding and prediction capabilities for images.

[0071] Example 3

[0072] In order to verify the technical effect of the technical solution of the present invention, that is, the SAM model after MAE pre-training and fine-tuning, a control group was set up for comparative experiments.

[0073] Experimental group: SAM model pre-trained with MAE and fine-tuned using 10 labeled lung H&E-stained sections with a learning rate set to 1e-6.

[0074] Control group A: without any pre-training, the SAM model was fine-tuned using 10 labeled lung H&E stained sections with a learning rate set to 1e-4.

[0075] The control group B was set as: SAM model without pre-training and fine-tuning.

[0076] Experimental method: The experimental group, control group A, and control group B were tested for lung epithelial cell region recognition on images C and D, and the recognition effect was calculated using the Dice value.

[0077] The experimental results are as follows Figure 4-Figure 7 As shown. Among them, Figure 4 The prediction results of the control group A for image C (Dice = 82%), Figure 5 is the prediction result of the experimental group for image C (Dice=90%), Figure 6 is the prediction result of control group A for image D (Dice=82%), Figure 7 This is the prediction result of the experimental group for image D (Dice = 90%). The prediction result of the control group B is very poor and does not serve as a comparison, so there is no need to use the prediction result for comparison. As can be seen, the SAM model constructed by the present invention (experimental group) is significantly better than the untrained, fine-tuned SAM model (control group A) in recognizing tissue edges and some epithelial cells.

[0078] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention can be made by one of ordinary skill in the art without inventive effort. Therefore, any technical solution that can be derived by one of ordinary skill in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. An AI lung epithelial cell region recognition method based on MAE pre-trained ViT and SAM network, characterized by: The following steps are involved: S1: MAE pre-trained ViT model using unlabeled H&E-stained lung sections; S2: embedding the backbone network in the pre-trained ViT model into the SAM model; S3: Fine-tuning the SAM model using annotated lung H&E-stained sections; S4: Lung epithelial cell region recognition is achieved through the fine-tuned SAM model.

2. The method according to claim 1, characterized in that In step S1, the ViT model is pre-trained using MAE using unlabeled lung H&E stained sections. The specific steps are as follows: S11: Segment the unlabeled H&E-stained lung sections into fixed-size image blocks, randomly select a portion of the image blocks for masking, divide the image blocks into masked image blocks and visible image blocks, and convert each image block into an embedding vector; S12: Input the visible image block into the encoder, the encoder processes the embedding vector of the visible image block and performs absolute position encoding, and outputs a potential representation; S13: Inputting the potential representation and the mask mark output by the encoder into the decoder, and the decoder reconstructs the image based on the potential representation and the mask mark; S14: Calculate the loss between the reconstructed image and the original image; S15: Update the model’s parameters through backpropagation to minimize the loss.

3. The method according to claim 1, characterized in that The backbone network in the pre-trained ViT model in step S2 is embedded in the SAM model, and the specific steps are: S21: Load the ViT model pre-trained by MAE, extract its Attention layer and FFN layer and embed them into the SAM model as the core module for feature extraction; S22: The position encoding adopts the relative position encoding of the SAM model, and the FFN layer adopts the structure of the VIT model pre-trained by MAE.

4. The method according to claim 1, wherein The fine-tuning in step S3 refers to fine-tuning using the labeled lung H&E stained sections by setting a learning rate and in an interactive manner, wherein the range of the learning rate is 1e-6.

5. The method according to claim 2, characterized in that The loss in step S14 is calculated using a loss function, which is: in, y j represents the true value of the jth sample, It represents the predicted value of the model for the jth sample, and M represents the total number of samples.

6. The method according to claim 1, characterized in that The pre-training is updated using a layered learning rate, where the encoder part uses a learning rate of 1e-6 and the other parts use a learning rate of 1e-4.

7. The method according to claim 1, characterized in that The effect identified in S4 was evaluated by calculating the Dice coefficient, which is Among them, |A∩B| represents the number of pixels where the prediction result completely overlaps with the true annotation, |A| represents the total number of pixels in the true annotation area, and |B| represents the total number of pixels in the model prediction area.

8. The method according to claim 2, characterized in that The MAE includes an asymmetric encoder-decoder architecture, and the decoder is lighter than the encoder; the random mask ratio of the S11 is 75%.

9. A computing system, characterized in that The computing system is configured to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions executable by one or more processors of a computing system to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • SAM large model-based goiter ultrasonic image identification method and system

    CN118230077A

  • Pulmonary nodule classification and segmentation method based on Gaussian mixture model

    CN118587490A

  • Mask recovery enhancement-based low-resolution weak and small target detection method and system

    CN118864826A

  • Terahertz image breast tumor detection method based on SAM and curriculum learning

    CN118967559A