Method and device for interpreting cell pathology slide, computer equipment and medium
By constructing a self-supervised learning model and lightweight convolutional neural network, image enhancement and object detection of cytopathic slides are solved, and the problems of low automation and low detection accuracy in the existing technology are achieved, and efficient cytopathic interpretation is achieved.
Patent Information
- Application Number
- CN202510325657.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the prior art, the degree of automation of cell pathology interpretation and low detection accuracy are difficult to effectively deal with the problem of field shift caused by changes in staining style.
通过构建包含第一视觉Transformer编码器、第二视觉Transformer编码器和特征融合器的自监督学习模型,对原始玻片图像进行图像增强和处理,生成增强后图像。 Then, Gaussian filtering is used to uniform light processing, a lightweight convolutional neural network based on the RepVit module is constructed, and the corrected images are classified and targeted to locate the scanning area.
It significantly improves the degree of automation and detection accuracy of cell pathological slide interpretation, optimizes image quality and resource utilization efficiency, and enhances the model's ability to identify complex pathological features.
Smart Images

Figure CN119942541A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graphic image processing, and in particular relates to a method, device, computer equipment and medium for reading cytopathological slides. Background Art
[0002] ROSE, or Rapid on site evaluation, is a technique used in various diagnostic interventional procedures (such as bronchoscopy, thoracoscopy, gastroscopy, colonoscopy, and various surgical laparoscopy) to evaluate the nature of lesions by making slides of small specimens collected on site and then observing them under a microscope using rapid staining technology.
[0003] Currently, the scanning area positioning of ROSE slides and the judgment of whether the slides are qualified mainly rely on manual reading by doctors, which has low work efficiency. If the doctor's manual positioning of the slide scanning area and quality inspection are omitted, the disk usage and slide reading time will increase significantly. In addition, deep learning has important applications in digital pathology, but its lack of cross-domain generalization (such as staining differences) hinders its actual clinical application. Traditional data enhancement methods (such as random rotation, cropping, etc.) are difficult to effectively deal with domain shift problems. For example, in medical image analysis, changes in staining style will significantly affect detection performance. Summary of the invention
[0004] In view of this, the present invention provides a method for interpreting cytopathology slides to solve the technical problems of low automation and low detection accuracy in cytopathology interpretation in the prior art. The method comprises: Construct and train the first visual Transformer encoder based on the image features of the slide images of the historical slides , Second Vision Transformer Encoder and a self-supervised learning model of a feature fusion device, wherein the slide image of the original slide is enhanced by the self-supervised learning model to generate an enhanced slide image, wherein the image features include cell morphology features and staining style features, and the first visual Transformer encoder For encoding cell morphological features, the second visual Transformer encoder Used to encode the staining style features, and the feature fusion device is used to perform feature fusion on the encoding of the cell morphology features and the encoding of the staining style features to generate an enhanced slide image; The enhanced glass slide image is homogenized by using Gaussian filtering to generate a corrected image, a lightweight convolutional neural network based on the RepVit module is constructed and trained, and the lightweight convolutional neural network is used to perform target classification and target detection on the corrected image to generate a label classification and a positioning scanning area corresponding to the original glass slide, wherein the label classification includes qualified and unqualified; When the label classification is qualified, the positioning scanning area of the original slide is interpreted using a microscope to obtain the interpretation result of cytopathology.
[0005] Furthermore, the self-supervised learning model is constructed and trained based on the image features of the slide images of the historical slides, including: The mean square error of the cell morphology features is used as the staining style consistency loss , the mean square error of the staining style feature is used as the cell morphology loss ; Set the coloring style consistency loss weight , cell morphology loss weight and self-reconstruction loss weight ; According to the dyeing style consistency loss , coloring style consistency loss weight , cell morphology loss , cell morphology loss weight , self-reconstruction loss and self-reconstruction loss weight Construct loss functions for supervised learning models; The parameters of the self-supervised learning model are adjusted, and the self-supervised learning model is optimized until the loss function is minimized, and then the training of the self-supervised learning model is terminated.
[0006] Furthermore, the adjusting the parameters of the self-supervised learning model and optimizing the self-supervised learning model include: A self-supervised approach based on contrastive learning for the second visual Transformer encoder Conduct training; For the second visual Transformer encoder After the training is completed, the second visual Transformer encoder The parameters of the first visual Transformer encoder remain fixed. and feature fuser for training.
[0007] Furthermore, the method of performing image enhancement on the slide image of the original slide by the self-supervised learning model to generate an enhanced slide image includes: Divide the slide image of the original slide into a plurality of image blocks ; Through the first visual Transformer encoder For each image block Encode and generate cell morphology feature embedding , through the second visual Transformer encoder For each image block Encode and generate coloring style feature embedding ; Embedding the cell morphology features and coloring style feature embedding After splicing, input to the feature fusion device; The cell morphology features are embedded into the And the dyeing style feature embedding After feature fusion, a feature composite image is generated, and the feature composite image is used as the enhanced slide image.
[0008] Furthermore, the step of performing light homogenization processing on the enhanced slide image by Gaussian filtering to generate a corrected image includes: Converting the enhanced slide image into a Lab space to process a brightness channel to generate a brightness channel image; Determine the parameters of the Gaussian kernel of the Gaussian filter according to the size of the enhanced slide image, wherein the parameters include the kernel size and the standard deviation; Using the Gaussian filter to perform a convolution operation on the brightness channel image, extracting a background estimation range of the brightness channel image, and smoothing the boundary of the background estimation range to generate an estimation range of the illumination background; Deducting the estimated range of the illumination background from the enhanced slide image to generate a homogenized image, so that the brightness of the homogenized image is in the same domain; The uniformly lighted image is subjected to contrast enhancement processing to generate a corrected image.
[0009] Furthermore, the construction and training of a lightweight convolutional neural network based on the RepVit module, using the lightweight convolutional neural network to perform target classification and target detection on the corrected image, and generating label classification and positioning scanning areas corresponding to the original slide, include: Train and optimize the RepVit module through the prior frame; Extracting features of the corrected image through multiple layers of the RepVit module to generate a multi-scale feature map; Inputting the multi-scale feature map into the attention mechanism to enhance the features and generate an enhanced feature map; By connecting a plurality of weighted bidirectional pyramid network modules in series, the enhanced feature map is subjected to multiple bidirectional feature fusions to generate a fused multi-scale feature map; The fused multi-scale feature map is input into the regression and classification network to obtain the label classification and locate the scanning area, wherein unqualified label classifications include bleeding, abnormal staining, blur and non-human tissue impurities.
[0010] Furthermore, the training and optimization of the RepVit module through the prior frame includes: Generate a set of prior frames according to the size and aspect ratio of the labeled images of the training set and the multi-scale feature map; For each annotated frame of the annotated image, traverse all prior frames to obtain a set of prior frames that completely include the annotated frame; The a priori frame with the smallest area in the a priori frame set is taken as a positive sample, the a priori frame with an area greater than a set area threshold in the a priori frame set is taken as a negative sample, and the a priori frame with a non-minimum area and an area less than or equal to the set area threshold in the a priori frame set is taken as an irrelevant sample; The label classification of the marked box is used as the label classification of the positive sample, the label classification of the negative sample is marked as the background classification, the irrelevant sample is deleted from the currently used training set, and the RepVit module is trained and optimized using the data set consisting of the positive sample and label classification and the negative sample and background classification.
[0011] The present invention also provides a device for reading cytopathology slides to solve the technical problems of low automation and low detection accuracy in cytopathology reading in the prior art. The device comprises: The slide image enhancement module is used to construct and train the first visual Transformer encoder based on the image features of the slide image of the historical slide , Second Vision Transformer Encoder and a self-supervised learning model of a feature fusion device, wherein the slide image of the original slide is enhanced by the self-supervised learning model to generate an enhanced slide image, wherein the image features include cell morphology features and staining style features, and the first visual Transformer encoder For encoding cell morphological features, the second visual Transformer encoder Used to encode the staining style features, and the feature fusion device is used to perform feature fusion on the encoding of the cell morphology features and the encoding of the staining style features to generate an enhanced slide image; A classification and target detection module, used to perform light homogenization processing on the enhanced glass slide image using Gaussian filtering to generate a corrected image, construct and train a lightweight convolutional neural network based on the RepVit module, perform target classification and target detection on the corrected image using the lightweight convolutional neural network, and generate a label classification and positioning scanning area corresponding to the original glass slide, wherein the label classification includes qualified and unqualified; The cytopathology interpretation module is used to interpret the positioning scanning area of the original slide using a microscope to obtain the cytopathology interpretation result when the label classification is qualified.
[0012] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned methods for interpreting cytopathology slides when executing the computer program, so as to solve the technical problems of low automation and low detection accuracy in cytopathology interpretation in the prior art.
[0013] The present invention further provides a computer-readable storage medium storing a computer program for executing any of the above-mentioned methods for interpreting cytopathology slides, so as to solve the technical problems of low automation and low detection accuracy in the prior art of cytopathology interpretation.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly improved automation: The present invention constructs a self-supervised learning model including a first visual Transformer encoder, a second visual Transformer encoder and a feature fuser. The method can automatically enhance and process the original slide image without much manual intervention. Subsequently, a lightweight convolutional neural network is used to classify and detect targets, further reducing the workload of manual interpretation, thereby significantly improving the automation level of cytopathology interpretation. 2. Significantly improved detection accuracy: The present invention makes full use of cell morphology features and staining style features, encodes the two features respectively through the first visual Transformer encoder and the second visual Transformer encoder, and fuses the features with the help of the feature fuser to generate enhanced slide images with stronger expression capabilities. This feature enhancement design significantly improves the accuracy of subsequent target classification and detection, providing a more reliable basis for clinical diagnosis; 3. Even light processing to optimize image quality: The present invention uses Gaussian filtering to even light the enhanced slide image, which can effectively correct the image quality problem caused by uneven illumination. By processing the brightness channel in the Lab space and deducting the illumination background estimation range, the generated corrected image has uniform brightness, which lays a good foundation for subsequent feature extraction and analysis, thereby improving the overall interpretation effect; 4. Lightweight network takes into account both performance and efficiency: The lightweight convolutional neural network designed based on the RepVit module of the present invention not only ensures high performance of target classification and detection, but also significantly reduces the computational complexity and resource consumption of the model. This design makes it easy to deploy in actual application scenarios, which not only meets the accuracy requirements but also improves the operating efficiency; 5. Multi-scale feature fusion enhances detection capability: The present invention extracts multi-scale feature maps through a multi-layer RepVit module, and combines the attention mechanism and the weighted bidirectional pyramid network module for feature enhancement and fusion. This method can effectively capture the features of targets of different scales. This multi-scale feature processing method enhances the model's ability to recognize complex pathological features and improves the robustness of detection. 6. Prior frame optimization improves prediction accuracy: The present invention generates and optimizes the prior frame based on the labeled images of the training set, and combines the division strategy of positive samples, negative samples and irrelevant samples. This method improves the prediction accuracy of the model for the target position and category. The optimization training of the prior frame enhances the generalization ability of the model, enabling it to adapt to a variety of slide images. 7. Refine the classification of unqualified labels: The present invention subdivides unqualified labels into specific categories such as bleeding, abnormal staining, blur and non-human tissue impurities, which helps to more accurately identify and process unqualified samples. This targeted classification design improves the interpretability of the interpretation results and provides clear guidance for subsequent manual review or sample processing; 8. Optimization of resource utilization efficiency: When the label classification is qualified, the present invention only performs microscope interpretation on the positioning scanning area, avoiding comprehensive scanning of the entire slide. This strategy effectively reduces unnecessary calculations and microscope resource consumption, improves work efficiency, and ensures the pertinence and accuracy of the interpretation; In summary, the present invention significantly improves the automation level and detection accuracy of cytopathology slide interpretation and optimizes resource utilization efficiency through innovative designs such as image enhancement of self-supervised learning models, uniform light processing of Gaussian filtering, target detection of lightweight convolutional neural networks, and multi-scale feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 It is a flow chart of a method for interpreting a cytopathology slide of the present invention; Figure 2is a structural block diagram of a computer device of the present invention; Figure 3 It is a structural block diagram of a cytopathology slide interpretation device of the present invention.
[0017] In the figure: 201 - memory, 202 - processor, 301 - slide image enhancement module, 302 - classification and target detection module, 303 - cell pathology interpretation module. DETAILED DESCRIPTION
[0018] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0019] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0020] In an embodiment of the present invention, a method for interpreting cytopathological slides is provided, such as Figure 1 As shown, the method includes: Step S101: Construct and train a first visual Transformer encoder based on the image features of the historical slide , Second Vision Transformer Encoder The self-supervised learning model of the feature fusion is used to enhance the original slide (i.e., the slide to be interpreted) and generate an enhanced slide image. The image features include cell morphology features and staining style features. The first visual Transformer encoder Used to encode cell morphological features, the second visual Transformer encoder It is used to encode the staining style features, and the feature fusion device is used to fuse the encoding of cell morphology features and the encoding of staining style features to generate an enhanced slide image; Step S102: using Gaussian filtering to perform light homogenization on the enhanced slide image to generate a corrected image, constructing and training a lightweight convolutional neural network based on the RepVit module, using the lightweight convolutional neural network to perform target classification and target detection on the corrected image, and generating label classification and positioning scanning areas corresponding to the original slide, wherein the label classification includes qualified and unqualified; Step S103: If the label classification is qualified, the positioning scanning area of the original slide is interpreted using a microscope to obtain the interpretation result of cytopathology.
[0021] Specifically, the overall technical solution of the embodiment of the present invention is: 1. Image enhancement method based on self-supervised generation: ViT (Vision Transformer) is used to separate the cell morphology features and staining style characteristics of the input image, generate diverse synthetic images, and increase the diversity of the training data set.
[0022] 2. Efficient target detection method based on RepViT: RepViT is used as the backbone network of the detection model to perform target detection on the enhanced data set, locate the cell aggregation area as the area to be scanned, and classify the slides into unqualified (such as bleeding, abnormal staining, blur, non-human tissue impurities, etc.) and qualified.
[0023] 3. If the slide label classification is identified as qualified, scan the scan area output by the model on the slide through a microscope.
[0024] In the specific implementation, in order to enhance the original slide image through the features of the ROSE slide, the following steps are performed to build and train a self-supervised learning model based on the image features of the slide image of the historical slide: The mean square error of cell morphology features is used as the staining style consistency loss , the mean square error of the staining style feature is used as the cell morphology loss ; Set the coloring style consistency loss weight , cell morphology loss weight and self-reconstruction loss weight ; Consistency loss according to dyeing style , coloring style consistency loss weight , cell morphology loss , cell morphology loss weight , self-reconstruction loss and self-reconstruction loss weight Construct a loss function of the self-supervised learning model; adjust the parameters of the self-supervised learning model, wherein the parameters of the self-supervised learning model include: learning rate, batch size, optimizer, number of training rounds, etc., wherein the coloring style consistency loss weight , cell morphology loss weight , self-reconstruction loss weight The initial values of are set to 1.0, 0.5, and 0.5 respectively, and the optimizer is configured as follows: use the Adam optimizer, the learning rate is 0.0001, the number of training rounds is 200, and the batch size is 32; and optimize the self-supervised learning model until the loss function is minimized and the training of the self-supervised learning model is terminated.
[0025] In specific implementation, the following steps are performed to adjust the parameters of the self-supervised learning model and optimize the self-supervised learning model: A self-supervised approach based on contrastive learning for the second visual Transformer encoder Training; for the second visual Transformer encoder After the training is completed, the second visual Transformer encoder The parameters of the first visual Transformer encoder remain fixed. and feature fuser for training.
[0026] In a specific implementation, the following steps are performed to enhance the slide image of the original slide through the self-supervised learning model to generate an enhanced slide image: Divide the slide image of the original slide into multiple image blocks ; Through the first visual Transformer encoder For each image block Encode and generate cell morphology feature embedding , through the second visual Transformer encoder For each image block Encode and generate coloring style feature embedding ; Embed cell morphology features and coloring style feature embedding After splicing (i.e., embedding the cell morphology features and the dyeing style features of any two image blocks into splicing), the image blocks are input into the feature fusion device; the cell morphology features are embedded into the feature fusion device. and coloring style feature embedding After feature fusion, a feature composite image is generated, and the feature composite image is used as the enhanced slide image.
[0027] Specifically, first, the input image (i.e., the slide image of the original slide) is preprocessed. Divide into non-overlapping patches (i.e., image blocks) of size P×P ). Each patch is embedded into a high-dimensional space so that , where PS is the patch size, C is the number of channels, is the height of the input image, and W is the width of the input image.
[0028] Secondly, according to the special features (cell morphology and staining style features) of the input image (i.e., the slide image of the original slide), two visual Transformer encoders and (First Vision Transformer Encoder , Second Vision Transformer Encoder ) for each image block Encode and get each image block Cell morphology features embedded in and coloring style feature embedding ,in, and The encoders are and potential dimension.
[0029] In one embodiment, the second visual Transformer encoder Learning coloring styles using the Comparable Learning MoCo model, After the training, in the follow-up During training is frozen. Transformer encoder The vit framework based on self-supervised training can be used. The method of self-supervised training can be MoCo-based contrastive learning. After the training is completed, the second visual Transformer encoder The model parameters are no longer changed, and only the first visual Transformer encoder is trained later. And the feature fuser, through the above training strategy, ensures that only the dyeing features are captured.
[0030] Specifically, pathological slide images collected by the same equipment in the same hospital are used as positive samples, and those with different staining styles are considered as negative samples. The loss of the MoCo comparison model is as follows: , Among them, the sign function is based on and Whether they are from the same slice, output 0 and 1, K is the number of negative samples, sim is the cosine similarity calculation, for The positive sample of for Negative samples.
[0031] Again, the loss function of the self-supervised learning model uses three different mean square error (MSE) loss terms, namely, coloring style consistency loss , cell morphology loss and self-reconstruction loss . Loss of consistent cell morphology It is used to ensure that the generated images retain the original cell morphology. The staining style feature consistency loss is used to ensure that the staining style characteristics of the generated images are consistent.
[0032] Total loss function: , in, , and They are the staining style consistency loss weight, cell morphology loss weight and self-reconstruction loss weight, which are used to adjust the influence of each loss item in the training process. B is the batch and N is the number of training samples.
[0033] Specifically, a dataset generator can be used to generate images with different staining styles but consistent structures, thereby expanding the diversity of the original dataset.
[0034] Finally, randomly select the cell morphological features of an image and the coloring style characteristics of another image , converted to an image matrix and .
[0035] In one embodiment, the feature fusion device includes a channel mixer Convolutional GLU and an image synthesizer IS. and The features are sent to the channel mixer Convolutional GLU to calculate the fused features, and then the image synthesizer IS is used to synthesize the fused features into an image. The image synthesizer IS consists of multiple layers of convolution modules, which are used to gradually restore the image resolution.
[0036] In specific implementation, the following steps are performed to perform light homogenization processing on the enhanced slide image through Gaussian filtering to generate a corrected image: The enhanced slide image is converted into a Lab space to process the brightness channel and generate a brightness channel image; the parameters of the Gaussian kernel of the Gaussian filter are determined according to the size of the enhanced slide image, wherein the parameters include the kernel size and the standard deviation; the brightness channel image is convolved with the Gaussian filter to extract the background estimation range of the brightness channel image, and the boundary of the background estimation range is smoothed to generate the estimation range of the illumination background; the estimation range of the illumination background is deducted from the enhanced slide image to generate a homogenized image, so that the brightness of the homogenized image is in the same domain; the homogenized image is contrast enhanced to generate a corrected image.
[0037] Specifically, when the microscope parameters are not adjustable, in order to reduce the problem of uneven illumination of pathological scan images caused by the microscope, the collected pathological images are homogenized using Gaussian filtering to improve image quality and provide a better image basis for subsequent pathological analysis and diagnosis. For pathological slide images with uneven illumination, a Gaussian kernel is used to perform a convolution operation on the image to achieve a smoothing effect, and then the smoothed image is used as the illumination background estimate, which is subtracted from the original image to obtain a corrected image. Among them, the background estimate estimates a value representing the background brightness by analyzing the global or local brightness distribution of the image.
[0038] Specifically, the image is converted into Lab space to process the brightness channel for better separation of brightness and color information, and then the background smoothing effect is observed through experiments; in one embodiment, the kernel size of the Gaussian kernel can be selected as 1 / 8 to 1 / 4 of the smaller value of the length and width of the image resolution.
[0039] The expression for the standard deviation of the Gaussian kernel is: , Among them, ksize is the size of the Gaussian kernel.
[0040] In the specific implementation, the following steps are used to build and train a lightweight convolutional neural network based on the RepVit module, and use the lightweight convolutional neural network to perform target classification and target detection on the rectified image to generate label classification and positioning scanning areas corresponding to the original slide: The RepVit module is trained and optimized through the prior frame; the configuration of the RepVit module includes 4 stages, the number of channels is [64, 128, 256, 512], 3 RepVit blocks are stacked in each stage, RepVit adopts a re-parameterized block, and each RepVit block is composed of 1×1 convolution, 3×3 convolution and 1×1 convolution in series and embedded; during training: a multi-branch structure is adopted, branch 1: direct connection (Identity Mapping); branch 2: 1×1 convolution (channel expanded to 256); branch 3: 3×3 empty convolution (dilation=2, expanding the receptive field); during reasoning: the convolutions of multiple paths are merged into a single 3×3 convolution kernel to improve reasoning efficiency and reduce the amount of calculation; The features of the corrected image are extracted through multi-layer RepVit modules to generate a multi-scale feature map; the multi-scale feature map is input into the attention mechanism (the input feature map is globally pooled along the horizontal and vertical axes respectively, and each channel is encoded; the codes in the two directions are first concatenated, and then sent to the shared 1×1 convolution transformation function to generate an intermediate feature map; the intermediate feature map is split back into two independent branches, corresponding to the horizontal and vertical directions respectively; through two independent 1×1 convolutions, the attention weights in the horizontal and vertical directions are generated, and the original features are weighted and adjusted) to enhance the features and generate an enhanced feature map; through multiple weighted bidirectional pyramid network modules connected in series, the enhanced feature map is subjected to multiple bidirectional feature fusions to generate a fused multi-scale feature map; specifically, the feature maps of different scales obtained after the multi-layer RepVit modules are P3, P4, and P5, and P5 is downsampled twice to obtain P6 and P7, where P3 is a high-resolution feature (small targets are clearer) and P7 is a low-resolution feature (more global information); the upsampling path is top-down (Top-down) During the propagation process, high-level features are gradually integrated into low-level features so that low-level features have more semantic information. The upsampling path includes the following steps: Step 1: Reduce the dimension through 1×1 convolution so that features of different scales have the same number of channels: , Among them, P i is the input feature map, It is the feature map after 1×1 convolution transformation; Step 2: Use a learnable weighting mechanism to perform weighted summation of features of different scales: , in, The feature map generated for the top-down computation path, is the feature map after 1×1 convolution transformation, is a deeper feature map after 1×1 convolution, that is, Feature maps with richer semantic information but lower spatial resolution. Both w1 and w2 are learnable parameters used to dynamically adjust the influence of different feature layers. Upsample uses bilinear interpolation. The downsampling path passes low-level features to high-level layers in the bottom-up propagation process to supplement the resolution information. The downsampling path is: dimensionality reduction through 1×1 convolution and the application of a weighted mechanism: , in, The feature map generated for the bottom-up computation path, The feature map generated for the top-down computation path, It is a feature map of a lower scale (higher resolution) in the top-down path. w3 and w4 are both learnable parameters used to adjust the weights of features of different scales in fusion. Downsample uses a 3×3 convolution with a step size of 2 to achieve downsampling. The expression of the feature fusion is: , in, is the final feature map calculated by the weighted bidirectional pyramid network. The feature map generated for the top-down computation path, The feature map generated by the bottom-up calculation path, w5 and w6 are both learnable weighting parameters used to weightedly fuse the top-down and bottom-up features; The fused multi-scale feature map is input into the regression and classification network to obtain the label classification and locate the scanning area, where unqualified label classifications include bleeding, abnormal staining, blur and non-human tissue impurities; the structure of the classification network is: 4 layers of 3×3 convolution are used for feature extraction, and all scales share the same weights: 3×3 convolution → GroupNorm group normalization → BN, and the final output channel number is the number of anchor boxes × the number of categories, and then Sigmoid activation is used to output the classification confidence of each anchor point; the structure of the regression network is: the same 4 layers of 3×3 convolution as the classification network are used for feature extraction, and the final output channel number is the number of anchor boxes × 4, corresponding to the bounding box coordinate offset.
[0041] In the specific implementation, in order to provide an accurate target detection network for the microscope positioning scanning area and ensure that the sample is completely covered during the scanning process, the target matching strategy in the traditional target detection network based on the prior frame is innovatively improved to avoid missing sample information during the scanning process. At the same time, the scanning efficiency is improved, the scanning of irrelevant areas is reduced, and the accuracy and resource utilization efficiency of the microscope scanning are improved. The following steps are used to train and optimize the RepVit module through the prior frame: A set of prior frames are generated according to the size and aspect ratio of the annotated images and multi-scale feature maps of the training set; for each annotated frame of the annotated image, all prior frames are traversed to obtain a set of prior frames that completely include the annotated frame; the prior frame with the smallest area in the prior frame set is taken as a positive sample, the prior frame with an area greater than the set area threshold in the prior frame set is taken as a negative sample, and the prior frame with a non-minimum area and an area less than or equal to the set area threshold in the prior frame set is taken as an irrelevant sample; the label classification of the annotated frame is taken as the label classification of the positive sample, the label classification of the negative sample is marked as the background classification, and the irrelevant sample is deleted from the currently used training set. The RepVit module is trained and optimized using the dataset consisting of positive samples and label classification and negative samples and background classification.
[0042] Specifically, according to the size of the input image and the size of the feature map, a series of prior frames are generated according to different scales and aspect ratios. These prior frames will be used as potential candidates for the scanning area. For each annotated frame, all prior frames are traversed to filter out the set of prior frames that completely contain the annotated frame. The prior frame with the smallest area is selected from the set as the positive sample to ensure that only the area that just covers the sample is scanned to avoid scanning too many irrelevant areas. For the prior frames whose area exceeds the set threshold, they are marked as negative samples to exclude irrelevant areas that are obviously too large. The remaining prior frames that neither meet the positive sample conditions nor the negative sample conditions are regarded as irrelevant samples and do not participate in the calculation during training to avoid interfering with model training. The positive samples are marked as the category labels of the corresponding annotated frames, and the negative samples are marked as background classification. The classification loss and regression loss of the target detection network are used to train the positive and negative samples, and the position and size of the prior frame are adjusted to make it surround the target more accurately. In the prediction stage, the input image is forward propagated according to the trained model, and the final detection result is obtained using post-processing methods such as non-maximum suppression to achieve accurate scanning area positioning.
[0043] Specifically, if the slide label is qualified, the cell area output by the model is scanned using a microscope; otherwise, the slide is judged as a waste slide and the microscope does not need to scan it. In one embodiment of the present invention, when a microscope is used to interpret the positioning scanning area of the original slide, the general image of the original slide needs to use the same resampling, padding and standardization operations to ensure that the image size of the input network is 1024*1024. The processed image is input into the network, and the output result is subjected to high probability screening and NMS filtering in the post-processing process. The algorithm finally outputs the location and category of the detection area with the highest probability, and maps the area position back to the general image space of the original slide.
[0044] In this embodiment, a computer device is provided, such as Figure 2 As shown, it includes a memory 201, a processor 202 and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, any of the above-mentioned methods for interpreting cytopathological slides is implemented.
[0045] Specifically, the computer device may be a computer terminal, a server or a similar computing device.
[0046] In this embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program for executing any of the above-mentioned methods for interpreting cytopathological slides.
[0047] Specifically, computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, modules of programs or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable storage media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0048] Based on the same inventive concept, an apparatus for reading a cytopathology slide is also provided in an embodiment of the present invention, as described in the following embodiments. Since the principle of solving the problem by the apparatus for reading a cytopathology slide is similar to that of the method for reading a cytopathology slide, the implementation of the apparatus for reading a cytopathology slide can refer to the implementation of the method for reading a cytopathology slide, and the repeated parts will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0049] Figure 3 is a structural block diagram of a cytopathology slide interpretation device according to an embodiment of the present invention, such as Figure 3 As shown, it includes: a slide image enhancement module 301, a classification and target detection module 302 and a cell pathology interpretation module 303. The structure is described below.
[0050] The slide image enhancement module 301 is used to construct and train a first visual Transformer encoder based on the image features of the slide image of the historical slide , Second Vision Transformer Encoder The self-supervised learning model of the feature fusion is used to enhance the slide image of the original slide to generate an enhanced slide image, where the image features include cell morphology features and staining style features. The first visual Transformer encoder Used to encode cell morphological features, the second visual Transformer encoder It is used to encode the staining style features, and the feature fusion device is used to fuse the encoding of cell morphology features and the encoding of staining style features to generate an enhanced slide image; The classification and target detection module 302 is used to perform light homogenization processing on the enhanced slide image by using Gaussian filtering to generate a corrected image, construct and train a lightweight convolutional neural network based on the RepVit module, perform target classification and target detection on the corrected image by using the lightweight convolutional neural network, and generate a label classification and positioning scanning area corresponding to the original slide, wherein the label classification includes qualified and unqualified; The cytopathology interpretation module 303 is used to interpret the positioning scanning area of the original slide using a microscope to obtain the cytopathology interpretation result when the label classification is qualified.
[0051] In this embodiment, the slide image enhancement module includes: Loss definition unit, used to use the mean square error of cell morphology features as the staining style consistency loss , the mean square error of the staining style feature is used as the cell morphology loss ; Weight setting unit, used to set the coloring style consistency loss weight , cell morphology loss weight and self-reconstruction loss weight ; Loss function building block for coloring style consistency loss , coloring style consistency loss weight , cell morphology loss , cell morphology loss weight , self-reconstruction loss and self-reconstruction loss weight Constructing loss functions for self-supervised learning models; The self-supervised model training unit is used to adjust the parameters of the self-supervised learning model and optimize the self-supervised learning model until the loss function is minimized and the training of the self-supervised learning model is terminated.
[0052] In this embodiment, the self-supervised model training unit is used to train the second visual Transformer encoder by a self-supervised method based on contrastive learning. Training; for the second visual Transformer encoder After the training is completed, the second visual Transformer encoder The parameters of the first visual Transformer encoder remain fixed. and feature fuser for training.
[0053] In this embodiment, the slide image enhancement module further includes: An image division unit, used to divide the slide image of the original slide into a plurality of image blocks ; Feature encoding unit, used to pass the first visual Transformer encoder For each image block Encode and generate cell morphology feature embedding , through the second visual Transformer encoder For each image block Encode and generate coloring style feature embedding ; Feature splicing unit, used to embed cell morphological features into and coloring style feature embedding After splicing, input to the feature fusion device; Feature fusion unit, used to embed cell morphological features through feature fusion and coloring style feature embedding After feature fusion, a feature composite image is generated, and the feature composite image is used as the enhanced slide image.
[0054] In this embodiment, the classification and target detection module includes: A channel conversion unit, used for converting the enhanced slide image into a Lab space processing brightness channel to generate a brightness channel image; A parameter determination unit, used to determine the parameters of the Gaussian kernel of the Gaussian filter according to the size of the enhanced slide image, wherein the parameters include the kernel size and the standard deviation; An estimation range acquisition unit is used to perform a convolution operation on the brightness channel image using a Gaussian filter, extract a background estimation range of the brightness channel image, and smooth a boundary of the background estimation range to generate an estimation range of the illumination background; A light homogenization processing unit is used to deduct the estimated range of the illumination background from the enhanced slide image to generate a light homogenized image so that the brightness of the light homogenized image is in the same domain; The contrast enhancement unit is used to perform contrast enhancement processing on the homogenized image to generate a corrected image.
[0055] In this embodiment, the classification and target detection module also includes: A priori box optimization unit, used to train and optimize the RepVit module through the priori box; A feature extraction unit, used to extract features of the corrected image through a multi-layer RepVit module to generate a multi-scale feature map; The feature enhancement unit is used to input the multi-scale feature map into the attention mechanism to enhance the features and generate an enhanced feature map; A feature fusion unit is used to perform multiple bidirectional feature fusions on the enhanced feature map through multiple weighted bidirectional pyramid network modules connected in series to generate a fused multi-scale feature map; The positioning scanning area and classification unit is used to input the fused multi-scale feature map into the regression and classification network to obtain the label classification and locate the scanning area, among which the unqualified label classification includes bleeding, abnormal staining, blur and non-human tissue impurities.
[0056] In this embodiment, the prior frame optimization unit is used to generate a set of prior frames according to the size and aspect ratio of the annotated image and the multi-scale feature map of the training set; for each annotated frame of the annotated image, traverse all the prior frames to obtain a set of prior frames that completely include the annotated frame; take the prior frame with the smallest area in the set of prior frames as a positive sample, take the prior frame with an area greater than a set area threshold (the set area threshold is 1.2 times the area of the annotated frame, the formula is: Threshold=1.2×AreaAnnotation, dynamic adjustment strategy: if the aspect ratio of the annotated frame is greater than 2:1, the threshold is expanded to 1.5 times to accommodate narrow and long targets) as a negative sample, and take the prior frame with a non-minimum area and an area less than or equal to the set area threshold in the set of prior frames as an irrelevant sample; take the label classification of the annotated frame as the label classification of the positive sample, mark the label classification of the negative sample as the background classification, delete the irrelevant sample from the currently used training set, and use the data set consisting of the positive sample and the label classification and the negative sample and the background classification to train and optimize the RepVit module.
[0057] The embodiments of the present invention achieve the following technical effects: By separating cell morphology and staining style characteristics, the staining style of the image is changed while maintaining the consistency of the foreground content, thereby enhancing the integrity of the data set and improving the generalization ability of the deep learning model to unseen fields; the deep learning method of target detection is used to quickly locate the effective microscope scanning area, and the model backbone uses the RepVit module, which has both CNN local perception ability and ViT global abstraction ability; Gaussian filtering technology is used to uniformly process the collected pathological images to reduce the uneven illumination of the collected pathological images caused by the parameters of the microscope (such as light source intensity distribution, optical component performance, etc.), which seriously affects the subsequent pathological image analysis and diagnosis; Coordinate is added to the multi-scale feature map in the target detection network The Attention mechanism can aggregate features along two spatial directions respectively, which helps to capture long-range dependencies in one direction while retaining precise position information in the other direction. The design of the attention mechanism is concise and efficient, and almost does not increase the computational overhead. Through innovative improvements to the target matching and training process of the target detection network based on the prior frame, a better, more efficient and more accurate solution is provided for the scanning area positioning of the microscope, which provides a reliable guarantee for sample scanning and subsequent analysis and diagnosis. The embodiment of the present invention provides a method for interpreting cytopathology slides, which reduces human subjectivity and labor costs, can quickly determine the quality of the slides and the scanning area, and greatly reduces the use of disk space.
[0058] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0059] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for interpreting cytopathological slides, characterized in that: include: Construct and train the first visual Transformer encoder based on the image features of the slide images of the historical slides , Second Vision Transformer Encoder and a self-supervised learning model of a feature fusion device, wherein the slide image of the original slide is enhanced by the self-supervised learning model to generate an enhanced slide image, wherein the image features include cell morphology features and staining style features, and the first visual Transformer encoder For encoding cell morphological features, the second visual Transformer encoder Used to encode the staining style features, and the feature fusion device is used to perform feature fusion on the encoding of the cell morphology features and the encoding of the staining style features to generate an enhanced slide image; The enhanced glass slide image is homogenized by using Gaussian filtering to generate a corrected image, a lightweight convolutional neural network based on the RepVit module is constructed and trained, and the lightweight convolutional neural network is used to perform target classification and target detection on the corrected image to generate a label classification and a positioning scanning area corresponding to the original glass slide, wherein the label classification includes qualified and unqualified; When the label classification is qualified, the positioning scanning area of the original slide is interpreted using a microscope to obtain the interpretation result of cytopathology.
2. The method for interpreting cytopathological slides according to claim 1, characterized in that: The self-supervised learning model is constructed and trained based on the image features of the slide images of the historical slides, including: The mean square error of the cell morphology features is used as the staining style consistency loss , the mean square error of the staining style feature is used as the cell morphology loss ; Set the coloring style consistency loss weight , cell morphology loss weight and self-reconstruction loss weight ; According to the dyeing style consistency loss , coloring style consistency loss weight , cell morphology loss , cell morphology loss weight , self-reconstruction loss and self-reconstruction loss weight Construct loss functions for supervised learning models; The parameters of the self-supervised learning model are adjusted, and the self-supervised learning model is optimized until the loss function is minimized, and then the training of the self-supervised learning model is terminated.
3. The method for interpreting cytopathological slides according to claim 2, characterized in that: The adjusting the parameters of the self-supervised learning model and optimizing the self-supervised learning model include: A self-supervised approach based on contrastive learning for the second visual Transformer encoder Conduct training; For the second visual Transformer encoder After the training is completed, the second visual Transformer encoder The parameters of the first visual Transformer encoder remain fixed. and feature fuser for training.
4. The method for interpreting cytopathological slides according to claim 1, characterized in that: The method of performing image enhancement on the slide image of the original slide by using the self-supervised learning model to generate an enhanced slide image includes: Divide the slide image of the original slide into a plurality of image blocks ; Through the first visual Transformer encoder For each image block Encode and generate cell morphology feature embedding , through the second visual Transformer encoder For each image block Encode and generate coloring style feature embedding ; Embedding the cell morphology features and coloring style feature embedding After splicing, input to the feature fusion device; The cell morphology features are embedded into the And the dyeing style feature embedding After feature fusion, a feature composite image is generated, and the feature composite image is used as the enhanced slide image.
5. The method for interpreting cytopathological slides according to claim 1, characterized in that: The step of performing light homogenization processing on the enhanced slide image by Gaussian filtering to generate a corrected image includes: Converting the enhanced slide image into a Lab space to process a brightness channel to generate a brightness channel image; Determine the parameters of the Gaussian kernel of the Gaussian filter according to the size of the enhanced slide image, wherein the parameters include the kernel size and the standard deviation; Using the Gaussian filter to perform a convolution operation on the brightness channel image, extracting a background estimation range of the brightness channel image, and smoothing the boundary of the background estimation range to generate an estimation range of the illumination background; Deducting the estimated range of the illumination background from the enhanced slide image to generate a homogenized image, so that the brightness of the homogenized image is in the same domain; The uniformly lighted image is subjected to contrast enhancement processing to generate a corrected image.
6. The method for interpreting a cytopathological slide according to any one of claims 1, characterized in that: The method of constructing and training a lightweight convolutional neural network based on the RepVit module, using the lightweight convolutional neural network to perform target classification and target detection on the corrected image, and generating label classification and positioning scanning areas corresponding to the original slide includes: Train and optimize the RepVit module through the prior frame; Extracting features of the corrected image through multiple layers of the RepVit module to generate a multi-scale feature map; Inputting the multi-scale feature map into the attention mechanism to enhance the features and generate an enhanced feature map; By connecting a plurality of weighted bidirectional pyramid network modules in series, the enhanced feature map is subjected to multiple bidirectional feature fusions to generate a fused multi-scale feature map; The fused multi-scale feature map is input into the regression and classification network to obtain the label classification and locate the scanning area, wherein unqualified label classifications include bleeding, abnormal staining, blur and non-human tissue impurities.
7. The method for interpreting cytopathological slides according to claim 6, characterized in that: The RepVit module is trained and optimized through the prior frame, including: Generate a set of prior frames according to the size and aspect ratio of the labeled images of the training set and the multi-scale feature map; For each annotated frame of the annotated image, traverse all prior frames to obtain a set of prior frames that completely include the annotated frame; The a priori frame with the smallest area in the a priori frame set is taken as a positive sample, the a priori frame with an area greater than a set area threshold in the a priori frame set is taken as a negative sample, and the a priori frame with a non-minimum area and an area less than or equal to the set area threshold in the a priori frame set is taken as an irrelevant sample; The label classification of the marked box is used as the label classification of the positive sample, the label classification of the negative sample is marked as the background classification, the irrelevant sample is deleted from the currently used training set, and the RepVit module is trained and optimized using the data set consisting of the positive sample and label classification and the negative sample and background classification.
8. A reading device for the method for reading cytopathological slides according to claim 1, characterized in that: include: The slide image enhancement module is used to construct and train the first visual Transformer encoder based on the image features of the slide image of the historical slide , Second Vision Transformer Encoder and a self-supervised learning model of a feature fusion device, wherein the slide image of the original slide is enhanced by the self-supervised learning model to generate an enhanced slide image, wherein the image features include cell morphology features and staining style features, and the first visual Transformer encoder For encoding cell morphological features, the second visual Transformer encoder Used to encode the staining style features, and the feature fusion device is used to perform feature fusion on the encoding of the cell morphology features and the encoding of the staining style features to generate an enhanced slide image; A classification and target detection module, used to perform light homogenization processing on the enhanced glass slide image using Gaussian filtering to generate a corrected image, construct and train a lightweight convolutional neural network based on the RepVit module, perform target classification and target detection on the corrected image using the lightweight convolutional neural network, and generate a label classification and positioning scanning area corresponding to the original glass slide, wherein the label classification includes qualified and unqualified; The cytopathology interpretation module is used to interpret the positioning scanning area of the original slide using a microscope to obtain the cytopathology interpretation result when the label classification is qualified.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for interpreting the cytopathology slide according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the method for interpreting a cytopathological slide according to any one of claims 1 to 7.
Citation Information
Patent Citations
Deep learning-based lymphoma pathological image intelligent identification method
CN111798464A
Coronary artery stenosis recognition method based on depth auto-encoder composition
CN115984555A
Scanning method, device, medium and equipment for rapid cell pathology interpretation
CN116128856A
Lung cancer cell morphological analysis and identification method and computer readable storage medium
CN117496276A
Method and system for predicting three-level lymphatic structure based on HE staining pathological image
CN118941916A
Cited By
Image processing method and device for fluorescent slide, and storage medium
CN122391230A