Gynecological tumor pathology generation type pre-training Transform large model and application method
Through generative pre-training of Transformer large model, image features are extracted using mask mask module and double-arm architecture, the existing gynecological tumor pathological section model has been solved, and high accuracy and stable gynecological tumor pathological prediction is achieved.
Patent Information
- Application Number
- CN202510102706.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing gynecological tumor pathological slice model has poor performance, lacks generalization, and pixel-based image processing technology consumes a lot of computing power, which is not conducive to the deployment of small workstations.
The generative pre-trained Transformer model is adopted to extract image features through the mask mask module and the double-arm architecture, error calculation and weight update are performed, the weight of the target arm is frozen, and the exponential moving average and the weight of the content arm are updated.
A high-accurate gynecological oncology pathology prediction was achieved, with the accuracy of classification tasks reaching 99.1%, the accuracy of gene prediction tasks reaching 80.5%, and the accuracy of survival prediction tasks reaching 79.9%, and it performed stably in external data sets.
Smart Images

Figure CN120014292A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning image processing, and more specifically to a large-scale generative pre-trained Transformer model for the pathology of gynecological tumors and an application method thereof. Background Art
[0002] Artificial intelligence technology is used in the diagnosis and treatment guidance of various types of cancer. Existing models can classify pathological sections, automatically outline and label tissues or cells, predict genes, and perform prognosis classification tasks through the analysis of digital pathological sections.
[0003] Model performance in the fields of chemotherapy efficacy prediction, gene prediction, targeted drugs, and immunotherapy efficacy prediction for gynecological tumors is poor, or there is a lack of corresponding clinical data and artificial intelligence models. Most of the existing models for gynecological malignancies are task-specific and can only solve a specific task, lacking generalization. Supervised models often have problems such as difficulty in collecting training data, difficulty in using unlabeled data, and difficulty in passing external verification.
[0004] Currently, many pathological slice models for gynecological tumors often divide WSIs into small images for model construction instead of using the complete image. This only contains the local characteristics of the tumor tissue, not the global characteristics. HIPT provides an idea, which is trained on small images of different sizes and aggregates the features of these levels. This process suggests that training at the WSI level is effective for improving model performance, and integrating the features of multiple layers of images can improve the performance of the model.
[0005] In addition, the pixel information of a WSI often requires a huge digital matrix, which becomes an important obstacle to model building. We found that pixel-based image processing technology consumes more computing power than feature-based technology, which is not conducive to deployment on small workstations. Self-supervised pre-training has become an important method as a method that does not require a large amount of manual annotation. Compared with supervised training, self-supervised training can ignore the heterogeneity of data sets and can be extended to multiple tasks. In the field of computer vision, the two most classic technologies are invariance-based methods such as BYOL and DINO, and generative technologies such as masked autoencoder and I-JEPA (Image Joint Embedding Predictive Architecture). Invariance-based technologies often require artificial data enhancement such as deformation and cropping of images, which can easily cause bias in data sets. Generative training methods do not require a lot of hyperparameters, can maintain the stability of complex models, and simplify training steps. Summary of the invention
[0006] In view of this, the present invention provides a large generative pre-trained Transformer model for the pathology of gynecological tumors and an application method to solve the problems existing in the background technology.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A large Transformer model for gynecological tumor pathology, including a first-layer framework and a second-layer framework;
[0009] Both the first-layer framework and the second-layer framework include: a mask module for extracting an image mask of an input image, including a target image mask and a content image mask;
[0010] The content arm is used to extract a first image feature from the input image and to clip the content image feature according to the content image mask; and is also used to output a first target image feature according to the content image feature and the target image mask;
[0011] a target arm, for extracting a second image feature from the input image, and cropping the second target image feature according to the target image mask;
[0012] Calculate the error between the first target image feature and the second target image feature, and perform gradient backpropagation and weight update on the content arm; freeze the weight of the target arm, update the target arm using the exponential moving average and the weight of the content arm, and select the optimal weight based on the minimum error value of each round.
[0013] Optionally, after the first layer is completed, the weights of the first layer and the Vision Transformer architecture are used to extract and concatenate all patches of the slice into a slice feature map at the image level, and the slice feature map is used as the input value of the second layer.
[0014] Optionally, the content arm includes a first feature extractor and a feature predictor, and the target arm includes a second feature extractor.
[0015] Optionally, both the first feature extractor and the second feature extractor are based on a Vision Transformer architecture and extract the input image according to pixel values of the input image.
[0016] Optionally, the input patches and pseudo labels of the first layer framework, the first and second feature extractors adopt the VisionTransformer architecture.
[0017] Optionally, the second layer framework inputs slice feature maps and pseudo labels, and the first and second feature extractors adopt the Vision Transformer architecture.
[0018] A method for applying a large generative pre-trained Transformer model for gynecological tumor pathology, comprising the following steps:
[0019] Input the image to be identified and / or the true label into the large Transformer model of gynecological tumor pathology generative pre-training;
[0020] The input of the first layer framework is a patch, which uses the Vision Transformer architecture, which is the same as the Vision Transformer architecture in the first layer target feature extractor. The weights of the target feature extractor in the pre-trained first layer best weights are loaded, and the image features are output;
[0021] The input of the second-layer framework is the slice feature map. The Vision Transformer architecture is the same as the Vision Transformer architecture in the second-layer target feature extractor. A multi-layer classification perceptron is used to output the predicted value. The prediction performance is evaluated by comparing the predicted value with the real-world label.
[0022] It can be seen from the above technical solution that, compared with the prior art, the present invention provides a large-scale generative pre-trained Transformer model and application method for the pathology of gynecological tumors, which has the following beneficial effects:
[0023] 1. Model prediction is performed using HE-stained tissue pathology sections, and samples are easy to obtain without the need for resampling;
[0024] 2. The prediction accuracy is high, reaching 99.1% for classification tasks, 80.5% for gene prediction tasks, and 79.9% for survival prediction tasks; and the performance is stable in external data sets.
[0025] 3. As a large model for gynecological cancer treatment, a large amount of unlabeled data can be fully utilized through pre-training.
[0026] 4. The model is generalized, the performance of large sample tasks is stabilized, and the performance of small sample tasks is improved. For example, in the task of predicting the efficacy of targeted drugs, the internal verification accuracy can reach 80% using only data from 100 patients.
[0027] 5. The model architecture and training process are simple. Model migration can be easily achieved using model weights, and the model can be updated using new data and weights. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0029] Figure 1 It is a schematic diagram of the pre-training and application architecture of the present invention;
[0030] Figure 2 It is a schematic diagram of the double-arm structure of the present invention;
[0031] Figure 3 It is a schematic diagram of the pre-training structure of the present invention;
[0032] Figure 4 Schematic diagram of the application architecture of the present invention. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] Example 1
[0035] The embodiment of the present invention discloses a large Transformer model for pathological generative pre-training of gynecological tumors, such as Figure 1-Figure 4 As shown, the model establishes a model pre-training framework based on the dual-arm non-equal-height model of Vision Transformer and mask mask technology. The pre-training is performed on the patch and slice feature layer planes, that is, it includes the first-layer framework and the second-layer framework;
[0036] Both the first-layer framework and the second-layer framework include: a mask module for extracting an image mask of an input image, including a target image mask and a content image mask;
[0037] The content arm is used to extract a first image feature from the input image and to clip the content image feature according to the content image mask; and is also used to output a first target image feature according to the content image feature and the target image mask;
[0038] a target arm, for extracting a second image feature from the input image, and cropping the second target image feature according to the target image mask;
[0039] In this embodiment, the mask masking technology randomly extracts an image mask from an image in a certain size range and a certain aspect ratio, that is, the position of the image relative to the original image. In this model, two interrelated image masks are generated: the target image mask and the content image mask. The target image mask is four masks generated according to the size of 0.15-0.2 and the aspect ratio of 0.75-1.5, and the content image mask is one mask generated according to the size of 0.85-1.0 and the original aspect ratio. The production of the two masks is independent of each other, so there may be some overlap between the two.
[0040] Input the pre-trained patch image or image feature, enter the content arm and the target arm, output the first target image feature and the second target image feature and calculate the mean absolute error between the two, and backpropagate the content arm in the direction of error reduction to achieve weight update. The optimizer is AdamW, the learning rate is 1.5e-4, and the weight decay is 0.2; at the same time, the weight of the target arm is not updated according to the error, but the "Exponential Moving Average (EMA)" technology is used to integrate the updated content arm weight and target arm weight to achieve the update of the target arm weight, and the weighted weight value is 0.9.
[0041] The model connects the first-layer framework with the second-layer framework through the Vision Transformer architecture. After the first layer is completed, the weight of the first-layer target arm and the Vision Transformer architecture are used to input all 384*384 pixel patches of a slice, and the relative position of each patch on the slice is recorded. Each patch is extracted into a 1*768 dimension token, and a two-dimensional slice feature map is output according to the relative position of the patch.
[0042] The input of the first layer framework is a 384*384 pixel patch and a pseudo label, and the image enters the content arm and the target arm at the same time. The content arm includes the first feature extractor and the feature predictor, and the target arm includes the second feature extractor, both based on the transformer architecture. The patch size of the first feature extractor is 16*16, the depth is 12, and the output dimension is 768. It first makes an image patch and a position embedding according to the patch size, inputs the transformer architecture, and outputs 576 tokens. Each token has a dimension of 768, that is, the image feature of the original patch. The content image feature is cropped according to the above content image mask. The feature predictor does not need to make an image patch. The content image feature is padded with random features to 576 tokens, which are input into the transformer architecture together with the position embedding. The depth is 12, and the output dimension is 768, which is composed of 576 tokens of predicted image features. According to the target image mask, the first target image feature is output. The architecture of the second feature extractor is the same as the first feature extractor. After the image enters the target arm feature extractor, the image features of the original patch are output, and the second target image features are cropped according to the above target image mask. The best model weights are selected based on the minimum mean absolute error of the model. The input of the second layer framework is a two-dimensional slice feature map. The patch size of the transformer architecture of the first and second feature extractors is 2*2, the depth is 35, and the output dimension is 768. The rest of the framework is the same as the first layer.
[0043] This example uses a two-layer Vision Transformer and pre-trained model weights to perform supervised training related to outcomes:
[0044] Supervised training uses the classic two-layer Vision Transformer architecture and the optimal weights of the pre-trained target arm feature extractor. The input is a specific task dataset. The dataset is first randomly divided into a training set and an internal validation set with a ratio of 8:2.
[0045] The input of the first layer framework is a 384*384 pixel patch and a real-world label. The VisionTransformer architecture is used, with a patch size of 16*16 and a depth of 12. The weights of the target feature extractor in the pre-trained first layer best weights are loaded to output image features. The image features are classified using a multilayer perceptron (MLP) and a sigmoid function and the predicted labels are output. The predicted values of the training set and the true values of the training set are compared using the cross-entropy loss value (CrossEntropy Loss), and the architecture is gradient-backed and the weights are updated. The penalty weight is determined by the proportion of the true label. The optimizer is AdamW, the learning rate is 3e-4, and the weight decay is 0.05. The output is used to calculate the accuracy and precision in the internal validation set to evaluate its prediction performance and select the best weight. Using the best weights and the same VisionTransformer architecture, all patches of each slice are extracted into tokens of 1*768 dimensions, and the slice feature map is output according to the relative position of the patches. The input of the second layer framework is the slice feature map. The patch size of the Vision Transformer architecture is 2*2 and the depth is 35. The rest of the framework is the same as the first layer. The prediction performance is evaluated and the optimal weight is selected based on the decrease in the cross entropy loss value, the accuracy and precision calculated in the internal validation set, etc.
[0046] This embodiment also discloses a method for applying a large Transformer model for pathology of gynecological tumors, including the following steps:
[0047] Input the images to be recognized and / or real-world labels into a large, synthetic pre-trained Transformer model for gynecological tumor pathology;
[0048] The input of the first layer framework is a 384*384 pixel patch. It uses the Vision Transformer architecture with a patch size of 16*16 and a depth of 12. It loads the weights of the target feature extractor in the pre-trained first layer best weights and outputs a slice feature map.
[0049] The input of the second layer framework is the slice feature map. The patch size of the Vision Transformer architecture is 2*2 and the depth is 35. The output value passes through the MLP layer to obtain the predicted value. The prediction performance is evaluated by comparing the predicted value with the real-world label.
[0050] Example 2
[0051] In this embodiment, a total of 2743 ovarian cancer patients and 1335 cervical cancer patients were retrospectively included, with a total of 12,384 WSI images of the patients included, generating 17,136,892 patches of 384*384 pixels. The basic characteristics of the patients in the training set and the validation set are balanced. The loss value of the first layer of pre-training tends to be stable after about 50 data iterations, and the loss value of the second layer tends to be stable after about 120 data iterations. The loss value and prediction efficiency of the first layer of supervised training tend to be stable after about 30 data iterations, and the loss value and prediction efficiency of the second layer tend to be stable after about 50 data iterations. After 50 iterations, its predictive efficiency is evaluated in the internal validation set.
[0052] The accuracy of the internal validation set for the task of classifying cervical cancer pathological sections (squamous cell carcinoma or adenocarcinoma) was 99.1%. The accuracy of the external validation set retrospectively conducted at Shandong Cancer Hospital was 98.4%. The accuracy of the internal validation set for the task of predicting cervical cancer lymph node metastasis (whether there is lymph node metastasis) was 92.3%, with an OR value of 11.3. The accuracy of the external validation set prospectively conducted at Qilu Hospital of Shandong University was 91.9%.
[0053] The internal validation set accuracy of the three-category survival prediction of ovarian cancer (0-3 years death, 3-5 years death, >5 years survival) was 79.9%, and the balanced accuracy was 62.7%. The KM survival curve also suggests that the survival population can be effectively screened according to the model grouping. The accuracy in the TCGA external validation set data set reached 67%, and the balanced accuracy was 57.3%. The internal validation set accuracy of the ovarian cancer BRCA / HRD prediction task (whether BRCA is mutated, whether it is HRD) was 80.5%, and the OR value was 8.4. The accuracy in the external validation set of TCGA reached 77.9%. In addition, the TCGA data set was used for the targeted drug efficacy guidance task (whether genes such as CD274 and ERBB are mutated), and the accuracy of the internal validation set predicting ERBB was about 80.2%.
[0054] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0055] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large Transformer model for pathology of gynecological tumors, characterized by: It includes a first layer frame and a second layer frame; Both the first-layer framework and the second-layer framework include: a mask module for extracting an image mask of an input image, including a target image mask and a content image mask; The content arm is used to extract a first image feature from the input image and to clip the content image feature according to the content image mask; and is also used to output a first target image feature according to the content image feature and the target image mask; a target arm, for extracting a second image feature from the input image, and cropping the second target image feature according to the target image mask; The loss value is calculated for the first target image feature and the second target image feature, and the gradient is back-propagated and the weight is updated for the content arm; the weight of the target arm is frozen, and the target arm is updated using the exponential moving average and the weight of the content arm.
2. According to claim 1, a gynecological tumor pathology generative pre-trained Transformer large model is characterized in that: After the first layer is completed, the weights of the first layer and the Vision Transformer architecture are used to input the patch, and the patch is extracted into a slice feature map at the image level, and the slice feature map is used as the input value of the second layer.
3. The pathological generative pre-trained Transformer large model of gynecological tumors according to claim 1 is characterized in that: The content arm includes a first feature extractor and a feature predictor, and the target arm includes a second feature extractor.
4. The pathological generative pre-trained Transformer large model of gynecological tumors according to claim 1, characterized in that: The first feature extractor and the second feature extractor are both based on the Vision Transformer architecture and extract the input image according to the pixel values of the input image.
5. The pathological generative pre-trained Transformer large model of gynecological tumors according to claim 1, characterized in that: The input patch of the first layer framework adopts the Vision Transformer architecture, loads the weights of the target feature extractor in the pre-trained first layer best weights, and outputs the slice feature map.
6. The pathological generative pre-trained Transformer large model of gynecological tumors according to claim 1, characterized in that: The input slice feature map of the second layer framework uses the Vision Transformer architecture.
7. A method for applying a large Transformer model for pathology of gynecological tumors, characterized in that: The following steps are involved: Input the image to be identified and / or the true label into the large Transformer model of gynecological tumor pathology generative pre-training; The input of the first layer framework is a patch, which uses the Vision Transformer architecture, which is the same as the Vision Transformer architecture in the first layer target feature extractor. The weights of the target feature extractor in the pre-trained first layer best weights are loaded, and the image features are output; The input of the second-layer framework is the slice feature map. The Vision Transformer architecture is the same as the Vision Transformer architecture in the second-layer target feature extractor. A multi-layer classification perceptron is used to output the predicted value. The prediction performance is evaluated by comparing the predicted value with the real-world label.