A method, apparatus, computer device and medium for interpreting a cytopathological slide

Through self-supervised learning model and Gaussian filtering technology, image enhancement and feature fusion of cytopathic glass slides, combined with lightweight convolutional neural network for object detection, the problem of low degree of automation of cytopathic interpretation and insufficient detection accuracy is solved, and efficient and accurate automated interpretation is achieved.

CN119942541BActive Publication Date: 2025-07-11FIRST PEOPLES HOSPITAL OF YUNNAN PROVINCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510325657.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-11
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In the prior art, the scanning area positioning and interpretation of cell pathological slides mainly relies on manual operations, with low efficiency and insufficient detection accuracy. Deep learning has problems that changes in staining style affect detection performance in cross-domain applications.

Method used

A self-supervised learning model including the first visual Transformer encoder, the second visual Transformer encoder and feature fusion is constructed, and image enhancement and feature fusion is performed on the slide image, combining Gaussian filtering and lightweight convolutional neural network for uniform light processing and object detection, and using the RepVit module for multi-scale feature fusion and target classification.

Benefits of technology

The degree of automation of cell pathological interpretation has been significantly improved, the detection accuracy has been greatly improved, the resource utilization efficiency has been optimized, manual intervention and calculation consumption has been reduced, and the accuracy and efficiency of interpretation has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942541B_ABST
    Figure CN119942541B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, computer equipment and medium for interpreting cytopathological slides, belonging to the technical field of graphic image processing. The method includes the following steps: constructing a self-supervised learning model according to the image features of historical slides, performing image enhancement on the original slides through the self-supervised learning model to generate enhanced slide images; performing uniform illumination processing on the enhanced slide images by using Gaussian filtering to generate corrected images, constructing and training a lightweight convolutional neural network based on the RepVit module, and using the lightweight convolutional neural network to perform target classification and target detection on the corrected images to generate label classification and positioning scan regions corresponding to the original slides; if the label classification is qualified, using a microscope to interpret the positioning scan region of the original slide to obtain the interpretation result of cytopathology. The present invention interprets cytopathological slides through a deep learning model, improving the interpretation efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of graphic image processing, and particularly relates to a method, device, computer device and medium for interpreting cytopathological slides. Background Art

[0002] ROSE, namely Rapid on site evaluation, in various diagnostic interventional operations (such as bronchoscopy, thoracoscopy, gastroscopy, colonoscopy, and various surgical endoscopes), after the small specimens taken are specially made into slides on site, a rapid staining technique is used, and through microscopic observation, the nature of the lesion is evaluated.

[0003] Currently, the scanning area positioning of ROSE slides and the judgment of whether the slides are qualified mainly rely on doctors' manual film reading, with low work efficiency. If the doctor's manual positioning of the slide scanning area and quality inspection are omitted, it will greatly increase disk occupancy and slide reading time. In addition, deep learning has important applications in digital pathology, but its lack of cross-domain generalization (such as staining differences) hinders its actual clinical application. Traditional data augmentation methods (such as random rotation, cropping, etc.) are difficult to effectively handle domain shift problems. For example, in medical image analysis, changes in staining styles will significantly affect detection performance. Summary of the Invention

[0004] In view of this, the present invention provides a method for interpreting cytopathological slides to solve the technical problems of low automation degree and low detection accuracy in existing cytopathological interpretation. The method includes:

[0005] Construct and train a self-supervised learning model including a first Vision Transformer encoder , a second Vision Transformer encoder and a feature fuser according to the image features of the slide images of historical slides, and perform image enhancement on the slide images of the original slides through the self-supervised learning model to generate enhanced slide images, wherein the image features include cell morphological features and staining style features, and the first Vision Transformer encoder is used to encode the cell morphological features, the second Vision Transformer encoder is used to encode the staining style features, and the feature fuser is used to perform feature fusion on the encoding of the cell morphological features and the encoding of the staining style features to generate enhanced slide images;

[0006] Perform equalization processing on the enhanced glass slide image using Gaussian filtering to generate a corrected image, construct and train a lightweight convolutional neural network based on the RepVit module, and use the lightweight convolutional neural network to perform target classification and target detection on the corrected image to generate the label classification and the positioning scan area corresponding to the original glass slide, where the label classification includes qualified and unqualified;

[0007] When the label classification is qualified, use a microscope to interpret the positioning scan area of the original glass slide to obtain the interpretation result of cell pathology.

[0008] Further, constructing and training a self-supervised learning model based on the image features of the glass slide images of historical glass slides includes:

[0009] Take the mean square error of the cell morphological features as the staining style consistency loss and take the mean square error of the staining style features as the cell morphology loss ;

[0010] Set the staining style consistency loss weight and the cell morphology loss weight and the self-reconstruction loss weight ;

[0011] According to the staining style consistency loss and the staining style consistency loss weight and the cell morphology loss and the cell morphology loss weight and the self-reconstruction loss and the self-reconstruction loss weight Construct the loss function of the supervised learning model;

[0012] Adjust the parameters of the self-supervised learning model and optimize the self-supervised learning model until the loss function is minimized and then end the training of the self-supervised learning model.

[0013] Even further, adjusting the parameters of the self-supervised learning model and optimizing the self-supervised learning model includes:

[0014] Train the second Vision Transformer encoder through a self-supervised method based on contrastive learning ;

[0015] After the training of the second Vision Transformer encoder ends, keep the parameters of the second Vision Transformer encoder fixed, and for the first Vision Transformer encoder and the feature fuser are trained.

[0016] Further, the method for performing image enhancement on the slide image of the original slide through the self-supervised learning model to generate an enhanced slide image includes:

[0017] Dividing the slide image of the original slide into a plurality of image patches ;

[0018] Encoding each image patch through the first Vision Transformer encoder to generate cell morphology feature embeddings , and encoding each image patch through the second Vision Transformer encoder to generate staining style feature embeddings ;

[0019] After splicing the cell morphology feature embeddings and the staining style feature embeddings , inputting them into the feature fuser;

[0020] Performing feature fusion on the cell morphology feature embeddings and the staining style feature embeddings through the feature fuser to generate a feature synthesis image, and using the feature synthesis image as the enhanced slide image.

[0021] Further, the method for performing equalization processing on the enhanced slide image through Gaussian filtering to generate a corrected image includes:

[0022] Converting the enhanced slide image to the Lab color space to process the luminance channel to generate a luminance channel image;

[0023] Determining the parameters of the Gaussian kernel of the Gaussian filtering according to the size of the enhanced slide image, where the parameters include the kernel size and the standard deviation;

[0024] Performing a convolution operation on the luminance channel image using the Gaussian filtering, extracting the background estimation range of the luminance channel image, and smoothing the boundary of the background estimation range to generate an estimated range of the illumination background;

[0025] Subtracting the estimated range of the illumination background from the enhanced slide image to generate an equalized image, so that the luminance of the equalized image is in the same domain;

[0026] Performing contrast enhancement processing on the equalized image to generate a corrected image.

[0027] Further, constructing and training a lightweight convolutional neural network based on the RepVit module, and using the lightweight convolutional neural network to perform object classification and object detection on the corrected image, generating label classification and positioning scan regions corresponding to the original glass slide, including:

[0028] Training and optimizing the RepVit module through prior boxes;

[0029] Extracting features of the corrected image through multiple layers of the RepVit module to generate multi-scale feature maps;

[0030] Inputting the multi-scale feature maps into an attention mechanism to enhance the features and generate enhanced feature maps;

[0031] Performing multiple two-way feature fusions on the enhanced feature maps through a series of weighted bidirectional pyramid network modules to generate fused multi-scale feature maps;

[0032] Inputting the fused multi-scale feature maps into a regression and classification network to obtain the label classification and positioning scan regions, where the unqualified label classifications include bleeding, abnormal staining, blurring, and non-human tissue impurities.

[0033] Furthermore, the training and optimization of the RepVit module through prior boxes includes:

[0034] Generating a set of prior boxes according to the annotated images in the training set and the size and aspect ratio of the multi-scale feature maps;

[0035] For each annotation box in the annotated image, traversing all prior boxes to obtain a set of prior boxes that completely enclose the annotation box;

[0036] Taking the prior box with the smallest area in the set of prior boxes as a positive sample, taking the prior boxes with areas greater than a set area threshold in the set of prior boxes as negative samples, and taking the prior boxes with areas that are not the smallest and less than or equal to the set area threshold in the set of prior boxes as irrelevant samples;

[0037] Taking the label classification of the annotation box as the label classification of the positive sample, marking the label classification of the negative sample as the background classification, deleting the irrelevant samples from the currently used training set, and training and optimizing the RepVit module using the dataset composed of the positive samples and label classifications and the negative samples and background classifications.

[0038] The present invention also provides an apparatus for interpreting cell pathological glass slides to solve the technical problems of low automation degree and low detection accuracy in cell pathological interpretation in the prior art. The apparatus includes:

[0039] A slide image enhancement module for constructing and training a self-supervised learning model including a first Vision Transformer encoder according to the image features of the slide images of historical slides , a second Vision Transformer encoder and a feature fusion unit, and performing image enhancement on the slide images of the original slides through the self-supervised learning model to generate enhanced slide images. Wherein, the image features include cell morphological features and staining style features, and the first Vision Transformer encoder is used to encode the cell morphological features, the second Vision Transformer encoder is used to encode the staining style features, and the feature fusion unit is used to perform feature fusion on the encoding of the cell morphological features and the encoding of the staining style features to generate enhanced slide images;

[0040] A classification and object detection module for performing uniform illumination processing on the enhanced slide images by using Gaussian filtering to generate corrected images, constructing and training a lightweight convolutional neural network based on the RepVit module, and performing object classification and object detection on the corrected images by using the lightweight convolutional neural network to generate the label classification and the localization scanning area corresponding to the original slide, wherein the label classification includes qualified and unqualified;

[0041] A cytopathological interpretation module for, when the label classification is qualified, using a microscope to interpret the localization scanning area of the original slide to obtain the cytopathological interpretation result.

[0042] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method for interpreting cytopathological slides is implemented to solve the technical problems of low automation degree and low detection accuracy in the prior art for cytopathological interpretation.

[0043] The present invention further provides a computer-readable storage medium storing a computer program for executing the above-mentioned method for interpreting cytopathological slides to solve the technical problems of low automation degree and low detection accuracy in the prior art for cytopathological interpretation.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] 1. Significantly improved automation: By constructing a self-supervised learning model that includes a first Vision Transformer encoder, a second Vision Transformer encoder, and a feature fusion module, the method of the present invention can automatically enhance and process the original glass slide images without excessive manual intervention. Subsequently, a lightweight convolutional neural network is used for target classification and detection, further reducing the workload of manual interpretation, thereby significantly improving the automation level of cytopathological interpretation.

[0046] 2. Greatly improved detection accuracy: The present invention makes full use of cell morphological features and staining style features. The first Vision Transformer encoder and the second Vision Transformer encoder are used to encode the two types of features respectively, and a feature fusion module is used for feature fusion to generate an enhanced glass slide image with stronger expressive ability. This feature enhancement design significantly improves the accuracy of subsequent target classification and detection, providing a more reliable basis for clinical diagnosis.

[0047] 3. Uniform light processing optimizes image quality: The present invention uses Gaussian filtering to perform uniform light processing on the enhanced glass slide image, which can effectively correct the image quality problems caused by uneven illumination. By processing the brightness channel in the Lab color space and subtracting the estimated range of the illumination background, the generated corrected image has uniform brightness, laying a good foundation for subsequent feature extraction and analysis, thereby improving the overall interpretation effect.

[0048] 4. Lightweight network balances performance and efficiency: The lightweight convolutional neural network designed based on the RepVit module in the present invention not only ensures high performance in target classification and detection, but also significantly reduces the computational complexity and resource consumption of the model. This design makes it easy to deploy to actual application scenarios, meeting both the accuracy requirements and improving the running efficiency.

[0049] 5. Multi-scale feature fusion enhances detection ability: The present invention extracts multi-scale feature maps through multiple RepVit modules, and combines the attention mechanism and the weighted bidirectional pyramid network module for feature enhancement and fusion. This method can effectively capture the features of targets at different scales. This multi-scale feature processing method enhances the model's ability to identify complex pathological features and improves the robustness of detection.

[0050] 6. Optimization of prior boxes improves prediction accuracy: The present invention generates and optimizes prior boxes according to the annotated images in the training set, combined with the division strategy of positive samples, negative samples, and irrelevant samples. This method improves the prediction accuracy of the model for the target position and category. The optimized training of prior boxes enhances the generalization ability of the model, enabling it to adapt to diverse glass slide images.

[0051] 7. Refinement of Unqualified Label Classification: The present invention subdivides unqualified labels into specific categories such as bleeding, abnormal staining, blurring, and non-human tissue impurities, which helps to more accurately identify and process unqualified samples. This targeted classification design improves the interpretability of the interpretation results and provides clear guidance for subsequent manual review or sample processing.

[0052] 8. Optimization of Resource Utilization Efficiency: When the label is classified as qualified in the present invention, only the positioning scan area is microscopically interpreted, avoiding a full scan of the entire glass slide. This strategy effectively reduces unnecessary computational and microscope resource consumption, improves work efficiency, and at the same time ensures the pertinence and accuracy of the interpretation.

[0053] In summary, through innovative designs such as image enhancement by a self-supervised learning model, uniform illumination processing by Gaussian filtering, object detection by a lightweight convolutional neural network, and multi-scale feature fusion, the present invention significantly improves the automation degree and detection accuracy of cell pathological glass slide interpretation, while optimizing the resource utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 is a flowchart of a method for interpreting a cell pathological glass slide of the present invention;

[0056] Figure 2 is a structural block diagram of a computer device of the present invention;

[0057] Figure 3 is a structural block diagram of an apparatus for interpreting a cell pathological glass slide of the present invention.

[0058] In the figure: 201 - Memory, 202 - Processor, 301 - Glass Slide Image Enhancement Module, 302 - Classification and Object Detection Module, 303 - Cell Pathology Interpretation Module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The following will describe the embodiments of the present invention in detail with reference to the drawings.

[0060] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0061] In an embodiment of the present invention, a method for interpreting a cytopathological glass slide is provided. As Figure 1 shown, the method includes:

[0062] Step S101: Construct and train a self-supervised learning model including a first Vision Transformer encoder , a second Vision Transformer encoder and a feature fusion unit according to the image features of historical glass slides. Use the self-supervised learning model to perform image enhancement on the original glass slide (i.e., the glass slide to be interpreted) to generate an enhanced glass slide image. Among them, the image features include cell morphological features and staining style features. The first Vision Transformer encoder is used to encode the cell morphological features, the second Vision Transformer encoder is used to encode the staining style features, and the feature fusion unit is used to perform feature fusion on the encoding of the cell morphological features and the encoding of the staining style features to generate an enhanced glass slide image;

[0063] Step S102: Use Gaussian filtering to perform uniform illumination processing on the enhanced glass slide image to generate a corrected image. Construct and train a lightweight convolutional neural network based on the RepVit module, and use the lightweight convolutional neural network to perform object classification and object detection on the corrected image to generate the label classification and the positioning scan area corresponding to the original glass slide. Among them, the label classification includes qualified and unqualified;

[0064] Step S103: If the label classification is qualified, use a microscope to interpret the positioning scan area of the original glass slide to obtain the interpretation result of cytopathology.

[0065] Specifically, the overall technical solution of the embodiment of the present invention is:

[0066] 1. An image enhancement method based on self-supervised generation:

[0067] Use ViT (Vision Transformer) to separate the cytomorphological features and staining style characteristics of the input image, generate diverse synthetic images, and increase the diversity of the training dataset.

[0068] 2. Efficient Object Detection Method Based on RepViT:

[0069] Employ RepViT as the backbone network of the detection model, perform object detection on the enhanced dataset, locate the cell aggregation area as the area to be scanned, and at the same time classify the glass slides into unqualified (such as bleeding, abnormal staining, blurring, non-human tissue impurities, etc.) and qualified.

[0070] 3. If the glass slide label classification is recognized as qualified, scan the scanned area output by the model on the glass slide through a microscope.

[0071] In specific implementation, in order to enhance the original glass slide image according to the features of the ROSE glass slide, the following steps are used to construct and train a self-supervised learning model based on the image features of the glass slide images of historical glass slides:

[0072] Take the mean square error of the cytomorphological features as the staining style consistency loss , and take the mean square error of the staining style features as the cytomorphological loss ; set the staining style consistency loss weight , the cytomorphological loss weight and the self-reconstruction loss weight ; according to the staining style consistency loss , the staining style consistency loss weight , the cytomorphological loss , the cytomorphological loss weight , the self-reconstruction loss and the self-reconstruction loss weight construct the loss function of the self-supervised learning model; adjust the parameters of the self-supervised learning model, and the parameters of the self-supervised learning model include: learning rate, batch size, optimizer, number of training epochs, etc. Among them, the initial values of the staining style consistency loss weight , the cytomorphological loss weight , the self-reconstruction loss weight are respectively set to 1.0, 0.5, 0.5, the optimizer is configured as: use the Adam optimizer, the learning rate is 0.0001, the number of training epochs is 200, and the batch size is 32; and optimize the self-supervised learning model until the loss function is minimized and then end the training of the self-supervised learning model.

[0073] In specific implementation, the parameters of the self-supervised learning model are adjusted and the self-supervised learning model is optimized through the following steps:

[0074] Train the second Vision Transformer encoder through a self-supervised method based on contrastive learning ; After the training of the second Vision Transformer encoder ends, keep the parameters of the second Vision Transformer encoder fixed and train the first Vision Transformer encoder and the feature fusion unit.

[0075] In specific implementation, the following steps are used to perform image enhancement on the slide image of the original slide through the self-supervised learning model to generate an enhanced slide image:

[0076] Divide the slide image of the original slide into multiple image patches ; Encode each image patch through the first Vision Transformer encoder to generate cell morphology feature embeddings , and encode each image patch through the second Vision Transformer encoder to generate staining style feature embeddings ; After splicing the cell morphology feature embeddings and the staining style feature embeddings (that is, splicing the cell morphology feature embeddings and the staining style feature embeddings of any two image patches), input them into the feature fusion unit; After feature fusion of the cell morphology feature embeddings and the staining style feature embeddings by the feature fusion unit, generate a feature synthesis image, and use the feature synthesis image as the enhanced slide image.

[0077] Specifically, first, preprocess the input image (that is, the slide image of the original slide), and divide the input image into non-overlapping patches of size P×P (that is, image patches ). Each patch is embedded into a high-dimensional space so that , where PS is the patch size, C is the number of channels, is the height of the input image, and W is the width of the input image.

[0078] Second, according to the special features (cell morphology features and staining style features) of the input image (that is, the slide image of the original slide), through two Vision Transformer encoders and (First Vision Transformer Encoder , Second Vision Transformer Encoder ) for each image block Encode and get each image block Cell morphology features embedded in and coloring style feature embedding ,in, and The encoders are and potential dimension.

[0079] In one embodiment, the second visual Transformer encoder Learning coloring styles using the Comparable Learning MoCo model, After the training, in the follow-up During training is frozen. Transformer encoder The vit framework based on self-supervised training can be used. The method of self-supervised training can be MoCo-based contrastive learning. After the training is completed, the second visual Transformer encoder The model parameters are no longer changed, and only the first visual Transformer encoder is trained later. And the feature fuser, through the above training strategy, ensures that only the dyeing features are captured.

[0080] Specifically, pathological slide images collected by the same equipment in the same hospital are used as positive samples, and those with different staining styles are considered as negative samples. The loss of the MoCo comparison model is as follows:

[0081] ,

[0082] Among them, the sign function is based on and Whether they are from the same slice, output 0 and 1, K is the number of negative samples, sim is the cosine similarity calculation, for The positive sample of for Negative samples.

[0083] Again, the loss function of the self-supervised learning model uses three different mean square error (MSE) loss terms, namely, coloring style consistency loss , cell morphology loss and self-reconstruction loss . Loss of consistent cell morphology Used to ensure that the generated image retains the original cell morphological features. The loss of consistency in staining style features is used to ensure the consistency of the staining style characteristics of the generated image.

[0084] Total loss function:

[0085] ,

[0086] Among them, 、 and are the weights of the staining style consistency loss, the cell morphology loss, and the self-reconstruction loss respectively, which are used to adjust the influence of each loss term during training. B is the batch, and N is the number of training samples.

[0087] Specifically, a dataset generator can also be used to generate images with different staining styles but consistent structures to expand the diversity of the original dataset.

[0088] Finally, randomly select the cell morphological features of one image and the staining style characteristics of another image , and convert them into image matrices and .

[0089] In one embodiment, the feature fuser includes a channel mixer Convolutional GLU and an image synthesizer IS. The concatenated and features are fed into the channel mixer Convolutional GLU to calculate the fused features, and then the image synthesizer IS synthesizes the fused features into an image. The image synthesizer IS consists of multiple convolutional modules and is used to gradually restore the image resolution.

[0090] Specifically, when implementing, the following steps are used to perform uniform illumination processing on the enhanced slide image through Gaussian filtering to generate a corrected image:

[0091] Convert the enhanced slide image to the Lab space to process the luminance channel to generate a luminance channel image; determine the parameters of the Gaussian kernel of the Gaussian filter according to the size of the enhanced slide image, where the parameters include the kernel size and the standard deviation; use the Gaussian filter to perform a convolution operation on the luminance channel image, extract the background estimation range of the luminance channel image, and smooth the boundary of the background estimation range to generate the estimated range of the illumination background; subtract the estimated range of the illumination background from the enhanced slide image to generate a uniformly illuminated image so that the luminance of the uniformly illuminated image is in the same domain; perform contrast enhancement processing on the uniformly illuminated image to generate a corrected image.

[0092] Specifically, when the microscope parameters are not adjustable, in order to reduce the problem of uneven illumination in the pathological slide imaging caused by the microscope, the collected pathological images are processed by Gaussian filtering for uniform illumination to improve the image quality and provide a better image basis for subsequent pathological analysis and diagnosis. For the pathological slide images with uneven illumination, a Gaussian kernel is used to perform a convolution operation on the image to achieve a smoothing effect, and then the smoothed image is used as the illumination background estimate, and this illumination background estimate is subtracted from the original image to obtain the corrected image. Among them, the background estimate estimates a value representing the background brightness by analyzing the global or local brightness distribution of the image.

[0093] Specifically, the image is converted to the Lab color space to process the brightness channel for better separation of brightness and color information, and then the background smoothing effect is observed through experiments; in one embodiment, the kernel size of the Gaussian kernel can be selected from 1 / 8 to 1 / 4 of the smaller value of the length and width of the image resolution.

[0094] The expression of the standard deviation of the Gaussian kernel is:

[0095] ,

[0096] where ksize is the size of the Gaussian kernel.

[0097] Specifically, in implementation, the construction and training of a lightweight convolutional neural network based on the RepVit module are achieved through the following steps. The lightweight convolutional neural network is used to perform object classification and object detection on the corrected image to generate the label classification and the localization scanning area corresponding to the original slide:

[0098] The RepVit module is trained and optimized through prior boxes; the configuration of the RepVit module includes 4 stages, and the number of channels is [64, 128, 256, 512] respectively. 3 RepVit blocks are stacked in each stage. RepVit adopts a reparameterized block, and each RepVit block is composed of a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution in series and embedded; during training: a multi-branch structure is adopted, Branch 1: direct connection (Identity Mapping); Branch 2: 1×1 convolution (the number of channels is expanded to 256); Branch 3: 3×3 dilated convolution (dilation = 2, expanding the receptive field); during inference: the convolutions of multiple paths are combined into a single 3×3 convolution kernel to improve the inference efficiency and reduce the computational amount;

[0099] Extract the features of the corrected image through multiple RepVit modules to generate multi-scale feature maps; input the multi-scale feature maps into the attention mechanism (perform global pooling on the input feature maps along the horizontal and vertical axes respectively, and encode each channel; first concatenate the encodings in the two directions, and then send them to a shared 1×1 convolution transformation function to generate an intermediate feature map; split the intermediate feature map back into two independent branches, corresponding to the horizontal and vertical directions respectively; through two independent 1×1 convolutions, generate the attention weights in the horizontal and vertical directions, and weighted adjustment of the original features) to enhance the features and generate enhanced feature maps; through a series of multiple weighted bidirectional pyramid network modules, perform multiple bidirectional feature fusions on the enhanced feature maps to generate the fused multi-scale feature maps; specifically, the feature maps of different scales obtained after passing through multiple RepVit modules are P3, P4, P5 respectively. Perform two downsamplings on P5 to obtain P6, P7, where P3 is a high-resolution feature (small targets are clearer), and P7 is a low-resolution feature (more global information); in the top-down propagation process of the upsampling path, gradually fuse the high-level features into the low-level features to make the low-level features have more semantic information; the upsampling path includes the following steps:

[0100] Step 1. Reduce the dimension through 1×1 convolution so that features of different scales have the same number of channels:

[0101] ,

[0102] where, i is the input feature map, is the feature map after 1×1 convolution transformation;

[0103] Step 2. Use a learnable weighting mechanism to perform weighted summation on features of different scales:

[0104] ,

[0105] where, is the feature map generated by the top-down calculation path, is the feature map after 1×1 convolution transformation, is the deeper feature map after 1×1 convolution, that is, a feature map with richer semantic information but lower spatial resolution than . w1 and w2 are both learnable parameters used to dynamically adjust the influence of different feature layers, and Upsample uses bilinear interpolation;

[0106] During the bottom-up propagation process, the downsampling path transfers low-level features to high-level ones to supplement resolution information. The downsampling path is as follows: dimensionality reduction is performed through 1×1 convolution, and a weighting mechanism is applied:

[0107] ,

[0108] where is the feature map generated by the bottom-up calculation path, is the feature map generated by the top-down calculation path, is the feature map at a lower scale (higher resolution) in the top-down path. Both w3 and w4 are learnable parameters used to adjust the weights of features at different scales in the fusion. Downsample is implemented using a 3×3 convolution with a stride of 2 for downsampling;

[0109] The expression for the feature fusion is:

[0110] ,

[0111] where is the final feature map calculated by the weighted bidirectional pyramid network, is the feature map generated by the top-down calculation path, is the feature map generated by the bottom-up calculation path. Both w5 and w6 are learnable weighting parameters used to weight and fuse the top-down and bottom-up features;

[0112] The fused multi-scale feature map is input into the regression and classification networks to obtain label classification and localization of the scanning region. Among them, the unqualified label classifications include bleeding, abnormal staining, blurring, and non-human tissue impurities. The structure of the classification network is as follows: 4 layers of 3×3 convolution are used for feature extraction, and the same weights are shared for all scales: 3×3 convolution → GroupNorm group normalization → BN. The final output channel number is the number of anchor boxes × the number of classes, and then Sigmoid activation is used to output the classification confidence of each anchor point. The structure of the regression network is as follows: 4 layers of 3×3 convolution identical to the classification network are used for feature extraction, and the final output channel number is the number of anchor boxes × 4, corresponding to the coordinate offsets of the bounding boxes.

[0113] Specifically, during implementation, to provide an accurate object detection network for microscope localization of the scanning region, ensure that the sample is completely covered during the scanning process, innovatively improve the object matching strategy in the traditional prior box-based object detection network, avoid missing sample information during the scanning process, improve the scanning efficiency at the same time, reduce the scanning of irrelevant regions, and enhance the accuracy and resource utilization efficiency of microscope scanning. It is achieved through the following steps to train and optimize the RepVit module through prior boxes:

[0114] Generate a set of prior boxes according to the labeled images in the training set and the sizes and aspect ratios of the multi-scale feature maps; for each labeled box in the labeled image, traverse all the prior boxes to obtain the set of prior boxes that completely enclose the labeled box; use the prior box with the smallest area in the set of prior boxes as the positive sample, use the prior boxes with areas greater than the set area threshold in the set of prior boxes as the negative samples, and use the prior boxes with areas that are not the smallest and areas less than or equal to the set area threshold in the set of prior boxes as irrelevant samples; classify the label of the labeled box as the label classification of the positive sample, label the label classification of the negative sample as the background classification, delete the irrelevant samples from the currently used training set, and use the data set composed of the positive samples and their label classifications and the negative samples and the background classification to train and optimize the RepVit module.

[0115] Specifically, according to the size of the input image and the size of the feature map, generate a series of prior boxes according to different scales and aspect ratios. These prior boxes will be used as candidates for potential scanning regions. For each labeled box, traverse all the prior boxes and filter out the set of prior boxes that completely contain the labeled box. Select the prior box with the smallest area from this set as the positive sample to ensure that only the region that just covers the sample is scanned, avoiding scanning too many irrelevant regions. For the prior boxes with areas exceeding the set threshold, mark them as negative samples to exclude obviously too large irrelevant regions. The remaining prior boxes that do not meet the conditions for positive samples nor negative samples are regarded as irrelevant samples and do not participate in the calculation during training to avoid interfering with the model training. Label the positive samples with the category labels corresponding to the labeled boxes, and label the negative samples with the background classification. Use the classification loss and regression loss of the object detection network to train the positive and negative samples, and adjust the positions and sizes of the prior boxes to make them enclose the target more accurately. In the prediction stage, perform forward propagation on the input image according to the trained model, and use post-processing methods such as non-maximum suppression to obtain the final detection results to achieve accurate scanning region localization.

[0116] Specifically, if the slide label is qualified, use the cell region output by the microscope scanning model; otherwise, the slide is judged as a waste slide and the microscope does not need to scan it. In an embodiment of the present invention, when using the microscope to interpret the positioning and scanning region of the original slide, the general view of the original slide needs to use the same resampling, filling, and normalization operations to ensure that the size of the image input to the network is 1024*1024. Input the processed picture into the network, and the output result passes through the high-probability screening and NMS filtering in the post-processing process. The algorithm finally outputs the position and category of the detection region with the highest probability, and maps the region position back to the general view space of the original slide.

[0117] In this embodiment, a computer device is provided, such as Figure 2As shown in the figure, it includes a memory 201, a processor 202, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for interpreting cytopathological slides.

[0118] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.

[0119] In this embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program for executing any of the above-mentioned methods for interpreting cytopathological slides.

[0120] Specifically, the computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0121] Based on the same inventive concept, an apparatus for interpreting cytopathological slides is also provided in an embodiment of the present invention, as described in the following embodiments. Since the principle of the apparatus for interpreting cytopathological slides to solve problems is similar to that of the method for interpreting cytopathological slides, the implementation of the apparatus for interpreting cytopathological slides can refer to the implementation of the method for interpreting cytopathological slides, and the repeated parts will not be described again. As used hereinafter, the term "unit" or "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0122] Figure 3 is a structural block diagram of the apparatus for interpreting cytopathological slides according to an embodiment of the present invention, as Figure 3 shown, including: a slide image enhancement module 301, a classification and target detection module 302, and a cytopathological interpretation module 303. The structure will be described below.

[0123] The slide image enhancement module 301 is used to construct and train a self-supervised learning model including a first Vision Transformer encoder , a second Vision Transformer encoder and a feature fusion unit based on the image features of the slide images of historical slides. The self-supervised learning model is used to perform image enhancement on the slide images of the original slides to generate enhanced slide images. Among them, the image features include cell morphological features and staining style features. The first Vision Transformer encoder is used to encode the cell morphological features, and the second Vision Transformer encoder is used to encode the staining style features. The feature fusion unit is used to fuse the encoded cell morphological features and the encoded staining style features to generate enhanced slide images;

[0124] The classification and object detection module 302 is used to perform equalization processing on the enhanced slide images by using Gaussian filtering to generate corrected images, construct and train a lightweight convolutional neural network based on the RepVit module, and use the lightweight convolutional neural network to perform object classification and object detection on the corrected images to generate label classification and positioning scan regions corresponding to the original slides, where the label classification includes qualified and unqualified;

[0125] The cytopathological interpretation module 303 is used to, when the label classification is qualified, use a microscope to interpret the positioned scan region of the original slide to obtain the interpretation result of the cytopathology.

[0126] In this embodiment, the slide image enhancement module includes:

[0127] The loss definition unit is used to use the mean square error of the cell morphological features as the staining style consistency loss , and use the mean square error of the staining style features as the cell morphology loss ;

[0128] The weight setting unit is used to set the staining style consistency loss weight , the cell morphology loss weight and the self-reconstruction loss weight ;

[0129] The loss function construction unit is used to construct the loss function of the self-supervised learning model according to the staining style consistency loss , the staining style consistency loss weight , the cell morphology loss , the cell morphology loss weight , the self-reconstruction loss and the self-reconstruction loss weight ;

[0130] A self-supervised model training unit for adjusting the parameters of a self-supervised learning model and optimizing the self-supervised learning model until the training of the self-supervised learning model ends after the loss function is minimized.

[0131] In this embodiment, the self-supervised model training unit is used to train the second Vision Transformer encoder through a self-supervised method based on contrastive learning ; after the training of the second Vision Transformer encoder ends, keep the parameters of the second Vision Transformer encoder fixed, and train the first Vision Transformer encoder and the feature fuser.

[0132] In this embodiment, the glass slide image enhancement module further includes:

[0133] An image division unit for dividing the glass slide image of the original glass slide into multiple image patches ;

[0134] A feature encoding unit for encoding each image patch through the first Vision Transformer encoder to generate cell morphology feature embeddings , and encoding each image patch through the second Vision Transformer encoder to generate staining style feature embeddings ;

[0135] A feature splicing unit for splicing the cell morphology feature embeddings and the staining style feature embeddings and inputting the spliced result into the feature fuser;

[0136] A feature fusion unit for fusing the cell morphology feature embeddings and the staining style feature embeddings through the feature fuser to generate a feature synthesis image, and using the feature synthesis image as the enhanced glass slide image.

[0137] In this embodiment, the classification and object detection module includes:

[0138] A channel conversion unit for converting the enhanced glass slide image to the Lab color space to process the luminance channel and generate a luminance channel image;

[0139] A parameter determination unit for determining the parameters of the Gaussian kernel of Gaussian filtering according to the size of the enhanced glass slide image, where the parameters include the kernel size and the standard deviation;

[0140] An estimation range acquisition unit for performing a convolution operation on the luminance channel image using Gaussian filtering, extracting the background estimation range of the luminance channel image, and smoothing the boundaries of the background estimation range to generate an estimation range of the illumination background;

[0141] An equalization processing unit for subtracting the estimation range of the illumination background from the enhanced glass slide image to generate an equalized image, so that the luminance of the equalized image is in the same domain;

[0142] A contrast enhancement unit for performing contrast enhancement processing on the equalized image to generate a corrected image.

[0143] In this embodiment, the classification and object detection module further includes:

[0144] A prior box optimization unit for training and optimizing the RepVit module through prior boxes;

[0145] A feature extraction unit for extracting the features of the corrected image through a multi-layer RepVit module to generate a multi-scale feature map;

[0146] A feature enhancement unit for inputting the multi-scale feature map into an attention mechanism to enhance the features and generate an enhanced feature map;

[0147] A feature fusion unit for performing multiple bidirectional feature fusions on the enhanced feature map through a series of weighted bidirectional pyramid network modules to generate a fused multi-scale feature map;

[0148] A positioning scan area and classification unit for inputting the fused multi-scale feature map into a regression and classification network to obtain label classification and positioning of the scan area, where the unqualified label classifications include bleeding, staining abnormalities, blurring, and non-human tissue impurities.

[0149] In this embodiment, the prior box optimization unit is configured to generate a set of prior boxes according to the labeled images in the training set and the sizes and aspect ratios of the multi-scale feature maps; for each labeled box in the labeled images, traverse all the prior boxes to obtain a set of prior boxes that completely enclose the labeled box; use the prior box with the smallest area in the set of prior boxes as the positive sample, use the prior boxes with areas greater than a set area threshold (the set area threshold is 1.2 times the area of the labeled box, the formula is: Threshold = 1.2 × AreaAnnotation, dynamic adjustment strategy: if the aspect ratio of the labeled box > 2:1, the threshold is expanded to 1.5 times to adapt to narrow and long targets) in the set of prior boxes as the negative samples, and use the prior boxes with areas that are not the smallest and areas less than or equal to the set area threshold in the set of prior boxes as irrelevant samples; use the label classification of the labeled box as the label classification of the positive sample, mark the label classification of the negative sample as the background classification, delete the irrelevant samples from the currently used training set, and use the data set composed of the positive samples and label classifications and the negative samples and background classifications to train and optimize the RepVit module.

[0150] The embodiments of the present invention achieve the following technical effects:

[0151] By separating the cell morphology and staining style characteristics, the staining style of the image is changed while maintaining the consistency of the foreground content, thereby enhancing the integrity of the data set and improving the generalization ability of the deep learning model for unseen fields; using the deep learning method of object detection to quickly locate the effective microscope scanning area, and the model backbone uses the RepVit module, which has both the local perception ability of CNN and the global abstraction ability of ViT; using Gaussian filtering technology to perform uniform illumination processing on the collected pathological images, reducing the uneven illumination phenomenon of the collected pathological images caused by the parameters of the microscope (such as light source intensity distribution, optical element performance, etc.), which seriously affects the subsequent pathological image analysis and diagnosis work; adding the Coordinate Attention attention mechanism to the multi-scale feature maps in the object detection network, which can aggregate features along two spatial directions respectively, helps to capture long-range dependencies in one direction, and at the same time retains accurate position information in the other direction. This attention mechanism is designed simply and efficiently, with almost no increase in computational overhead; through the innovative improvement of the object matching and training process of the prior box-based object detection network, a more high-quality, efficient, and accurate solution is provided for the scanning area positioning of the microscope, providing a reliable guarantee for sample scanning and subsequent analysis and diagnosis; the embodiments of the present invention provide a method for interpreting cell pathological slides, which can reduce human subjectivity and labor costs, quickly determine the adoption quality and scanning area of the slides, and greatly reduce the use of disk space.

[0152] Obviously, those skilled in the art should understand that the various modules or steps of the above-described embodiments of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

[0153] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for interpreting a cytopathological slide, characterized in that, Including: Construct and train a self-supervised learning model including a first Vision Transformer encoder based on the image features of the slide images of historical slides. , a second Vision Transformer encoder and a feature fusion unit, and enhance the slide images of the original slides through the self-supervised learning model to generate enhanced slide images. Among them, the image features include cell morphological features and staining style features. The first Vision Transformer encoder is used to encode the cell morphological features, and the second Vision Transformer encoder is used to encode the staining style features. The feature fusion unit is used to perform feature fusion on the encoded cell morphological features and the encoded staining style features to generate enhanced slide images; Performing equalization processing on the enhanced slide image using Gaussian filtering to generate a corrected image, constructing and training a lightweight convolutional neural network based on the RepVit module, and using the lightweight convolutional neural network to perform object classification and object detection on the corrected image to generate the label classification and the positioning scan area corresponding to the original slide, where the label classification includes qualified and unqualified; When the label classification is qualified, using a microscope to interpret the positioning scan area of the original slide to obtain the interpretation result of cell pathology; The constructing and training of the lightweight convolutional neural network based on the RepVit module, and using the lightweight convolutional neural network to perform object classification and object detection on the corrected image to generate the label classification and the positioning scan area corresponding to the original slide, includes: Training and optimizing the RepVit module through prior boxes; Extracting the features of the corrected image through multiple layers of the RepVit module to generate multi-scale feature maps; Inputting the multi-scale feature maps into an attention mechanism to perform feature enhancement processing to generate enhanced feature maps; Through a series of multiple weighted bidirectional pyramid network modules, performing multiple bidirectional feature fusions on the enhanced feature maps to generate fused multi-scale feature maps; Inputting the fused multi-scale feature maps into a regression and classification network to obtain the label classification and the positioning scan area, where the unqualified label classification includes bleeding, abnormal staining, blurring, and non-human tissue impurities.

2. The method for interpreting a cytopathological slide according to claim 1, wherein, The constructing and training of the self-supervised learning model according to the image features of the slide images of historical slides, includes: Use the mean square error of the cell morphological features as the staining style consistency loss , and use the mean square error of the staining style features as the cell morphology loss ; Set the weight of the consistency loss of the staining style , the weight of the cell morphology loss and the weight of the self-reconstruction loss ; According to the above-mentioned dyeing style consistency loss 、the weight of the dyeing style consistency loss 、the cell morphology loss 、the weight of the cell morphology loss 、the self-reconstruction loss and the weight of the self-reconstruction loss Construct the loss function of the supervised learning model; Adjusting the parameters of the self-supervised learning model and optimizing the self-supervised learning model until the loss function is minimized and then ending the training of the self-supervised learning model.

3. The method for interpreting a cytopathological glass slide according to claim 2, wherein, The adjusting the parameters of the self-supervised learning model and optimizing the self-supervised learning model, includes: Train the second vision Transformer encoder through a self-supervised method based on contrastive learning ; After the training of the second vision Transformer encoder ends, keep the parameters of the second vision Transformer encoder fixed, and train the first vision Transformer encoder and the feature fuser.

4. The method for interpreting a cytopathological glass slide according to claim 1, wherein The performing image enhancement on the slide image of the original slide through the self-supervised learning model to generate an enhanced slide image, includes: Divide the slide image of the original slide into multiple image blocks ; Through the first vision Transformer encoder Encode each image patch to generate cell morphology feature embeddings ; and through the second vision Transformer encoder Encode each image patch to generate staining style feature embeddings ; Embed the cell morphological features and the staining style features After splicing, input them into the feature fusion device; Embedding the cell morphological features through the feature fuser and the staining style features embedding After feature fusion, a feature synthesis image is generated, and the feature synthesis image is used as the enhanced slide image.

5. The method for interpreting a cytopathological glass slide according to claim 1, wherein The performing equalization processing on the enhanced slide image through Gaussian filtering to generate a corrected image, includes: Converting the enhanced slide image to the Lab space to process the luminance channel to generate a luminance channel image; Determining the parameters of the Gaussian kernel of the Gaussian filtering according to the size of the enhanced slide image, where the parameters include the kernel size and the standard deviation; Using the Gaussian filtering to perform a convolution operation on the luminance channel image, extracting the background estimation range of the luminance channel image, and smoothing the boundary of the background estimation range to generate an estimation range of the illumination background; Subtracting the estimation range of the illumination background from the enhanced slide image to generate an equalized image, so that the luminance of the equalized image is in the same domain; Performing contrast enhancement processing on the equalized image to generate a corrected image.

6. The method for interpreting a cytopathological slide according to claim 1, wherein, The training and optimizing the RepVit module through prior boxes, includes: Generating a set of prior boxes according to the annotated images of the training set and the size and aspect ratio of the multi-scale feature maps; For each annotation box of the annotated image, traverse all prior boxes and obtain the set of prior boxes that completely enclose the annotation box; Take the prior box with the smallest area in the set of prior boxes as the positive sample, take the prior boxes with areas greater than the set area threshold in the set of prior boxes as the negative samples, and take the prior boxes with areas that are not the smallest and are less than or equal to the set area threshold in the set of prior boxes as irrelevant samples; Take the label classification of the annotation box as the label classification of the positive sample, mark the label classification of the negative sample as the background classification, delete the irrelevant samples from the currently used training set, and use the dataset composed of the positive samples and label classifications and the negative samples and background classifications to train and optimize the RepVit module.

7. An interpretation device for the interpretation method of a cytopathological glass slide according to claim 1, characterized in that, Comprising: A glass slide image enhancement module for constructing and training a self-supervised learning model including a first Vision Transformer encoder according to the image features of the glass slide images of historical glass slides , a second Vision Transformer encoder and a feature fuser, and enhancing the glass slide image of the original glass slide through the self-supervised learning model to generate an enhanced glass slide image. Among them, the image features include cell morphological features and staining style features. The first Vision Transformer encoder is used to encode the cell morphological features, the second Vision Transformer encoder is used to encode the staining style features, and the feature fuser is used to perform feature fusion on the encoding of the cell morphological features and the encoding of the staining style features to generate an enhanced glass slide image; A classification and object detection module, configured to perform equalization processing on the enhanced slide image by using Gaussian filtering to generate a corrected image, construct and train a lightweight convolutional neural network based on the RepVit module, and use the lightweight convolutional neural network to perform object classification and object detection on the corrected image to generate the label classification and the positioning scan area corresponding to the original slide, wherein the label classification includes qualified and unqualified; A cytopathological interpretation module, configured to, when the label classification is qualified, use a microscope to interpret the positioning scan area of the original slide to obtain the cytopathological interpretation result.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the computer program, it implements the method for interpreting a cytopathological slide according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for executing the method for interpreting a cytopathological slide according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Coronary artery stenosis recognition method based on depth auto-encoder composition

    CN115984555A

  • Scanning method, device, medium and equipment for rapid cell pathology interpretation

    CN116128856A