Algorithm based on cervical liquid-based cell squamous epithelium lesion diagnosis

The diagnosis of cervical fluid-based cytopathy through deep learning algorithms solved the problems of difficulty in data labeling and low segmentation accuracy, achieved efficient and accurate cell classification and grading, and improved screening efficiency.

CN120340024APending Publication Date: 2025-07-18HANGZHOU YIPAI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342168.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems such as difficulty in data labeling, low cell segmentation accuracy, insufficient classification and grading accuracy, and low negative discharge rate in the diagnosis of cervical fluid-based cells, which cannot meet the needs of large-scale screening.

Method used

The deep learning-based algorithm is used for prospect extraction, training data preprocessing, cell segmentation and classification, and feature extraction and data enhancement are used to use convolutional neural networks. Combined with morphological features assisted classification, the model is optimized through the cross entropy loss function and the optimizer.

Benefits of technology

It improves the accuracy and efficiency of cervical fluid-based cell lesions diagnosis, reduces the probability of low-level cells being misdetected into high-level cells, and improves the negative exclusion rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340024A_ABST
    Figure CN120340024A_ABST
Patent Text Reader

Abstract

The invention particularly relates to an algorithm based on cervical liquid-based cell squamous epithelium lesion diagnosis. The algorithm comprises the following steps: foreground extraction; preprocessing the training data; performing cell segmentation; and classifying cells. According to the method, the algorithm is designed and optimized in a data stream mode, the type of data input into the algorithm for training is effectively controlled, the labeling difficulty is reduced, and the efficiency of negative screening is improved, and the algorithm firstly extracts a foreground region of the cervical liquid-based cell pathology picture and then extracts the foreground region of the cervical liquid-based cell pathology picture. Then, in a training stage, cutting the picture by utilizing a region segmentation mode, strictly controlling data input by utilizing a small picture, and carrying out refined segmentation on positive cells; finally, the cells obtained through segmentation are classified, in the classification process, the accuracy of squamous epithelial lesions is improved, and the situation that low-level cells are mistakenly detected as high-level cells is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical cytopathological analysis, and particularly to an algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells. Background Art

[0002] Cervical cancer is one of the most common malignant tumors in women. Early screening and diagnosis are crucial for reducing the incidence and mortality. Cervical liquid-based cytology examination is a commonly used clinical method for pre-screening cervical cancer. Usually, the diagnosis of cell lesions is carried out under a microscope according to cell morphology. However, the traditional manual film reading method has problems such as strong subjectivity, slow film reading speed, and high misdiagnosis rate, and cannot meet the needs of large-scale screening.

[0003] In recent years, with the rapid development of deep learning, the use of image processing methods and deep learning technologies for automated diagnosis of cervical liquid-based cells has become a research hotspot. Existing studies have tried to use cell segmentation, classification, and object detection algorithms, but still face the following challenges: Difficult data annotation: The background in cervical liquid-based cell images is complex and the picture resolution is high, resulting in difficult annotation.

[0004] Low cell segmentation accuracy: The boundaries between cells are blurred and the morphologies are diverse, resulting in low accuracy of the segmentation algorithm.

[0005] Insufficient classification and grading accuracy: There are subtle differences in the morphologies of different types of diseased cells, and the accuracy of classification and grading needs to be improved.

[0006] Low negative exclusion rate: The low accuracy of cell diagnosis leads to a low negative exclusion rate, and the accuracy of sample negative exclusion becomes the key to improving the efficiency of TCT diagnosis.

[0007] In view of the above problems, the present invention proposes an algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells. Summary of the Invention

[0008] The purpose of the present invention is to propose an algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells in order to solve the above problems.

[0009] In order to achieve the above purpose, the present invention adopts the following technical solutions: An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells, including the following parts: Foreground extraction: Extract the foreground part from the whole picture to remove background noise and irrelevant content; Preprocessing of training data: Crop the picture and filter the negative regions; Cell segmentation: Use a convolutional neural network for feature extraction, and identify and segment the cell regions through a trained model; Cell classification: First, determine whether the lesion is low - grade or high - grade, and then further subdivide the specific lesion. The low - grade lesions are divided into , , and the high - grade lesions are divided into , ; Use model training, perform data augmentation such as image rotation and flipping, use the cross - entropy loss function and optimizer to learn the cell morphological features and achieve classification.

[0010] Preferably, the foreground extraction specifically includes the following parts: Use deep - learning technology to perform preliminary extraction of the foreground area, and a semantic segmentation network such as network can be adopted; After being processed by the network, output the foreground segmentation mask Perform contour detection on the foreground segmentation mask, use the function in to obtain a series of contours, and use a function to perform ellipse fitting on each contour. This function returns the center coordinates, semi - major axis, semi - minor axis, and the rotation angle of the major axis of the ellipse; Take the center of the ellipse as the origin, convert the ellipse to polar coordinates, thereby obtaining a series of polar - coordinate points, and convert the polar - coordinate points back to Cartesian coordinates; and approximate the contour composed of the obtained Cartesian - coordinate points to generate an approximately circular contour area.

[0011] Preferably, the training data pre - processing specifically includes the following parts: Cut the picture into blocks along the circular area and perform cropping using a preset window size; Analyze and judge the negative area by calculating the average gray - value feature of each cut picture, and then locate the area containing diseased cells.

[0012] Preferably, the cell segmentation specifically includes the following parts: Cut the pathological picture into pictures of a preset size and record them as small - block pictures, and input the learning rate, optimizer, and iteration parameters into the deep - learning segmentation model; Assume that the set of input small - block pictures is , and the corresponding set of cell - segmentation labels is ; Use the trained model to predict the input small - block pictures, obtain the mask picture of all cell contours in the picture, and then find the positions of non - zero elements according to the cell - contour mask; Use the contour detection algorithm for post - processing to finally obtain the cell contours.

[0013] Preferably, first determine whether the lesion is low - grade or high - grade, and then further subdivide the specific lesion. The low - grade lesions are divided into , , and the high - grade lesions are divided into , , specifically including the following parts: Judge the severity according to the overall characteristics of the lesion cells, which are the key features for distinguishing low - grade and high - grade lesions; For the cells determined to be low - grade lesions, further subdivide them into and ; where indicates that the cell morphology has slight abnormalities, but the characteristics are not typical, making it difficult to accurately judge the nature of the lesion; indicates that there is already a definite low - grade squamous intraepithelial lesion, and the cell morphology change is more obvious; For the cells determined to be high - grade lesions, continue to subdivide them into and , refers to atypical squamous cells, represents high - grade squamous intraepithelial lesion.

[0014] Preferably, in the model training, data augmentation such as image rotation and flipping is performed, and the cross - entropy loss function and optimizer are used to learn the cell morphology features for classification. Specifically, it includes the following parts: Use the model for training, and perform data augmentation on the input cell images, specifically including: randomly rotating the cell images within a preset angle range; Then horizontally and vertically flip the cell images; And randomly scale the cell images within a preset ratio range; By adjusting color attributes such as brightness, contrast, and saturation of the image, enable the model to learn the influence of color changes on cell features, enhance the classification ability of the model under different color conditions, and train through the cross - entropy loss function and optimizer. Monitor the accuracy and recall metrics of the model during training to make the model reach the preset classification accuracy.

[0015] Preferably, in the cell segmentation step, perform morphological feature analysis on the obtained cell contours; the morphological features include the area, perimeter, circularity, and aspect ratio of the cells; use these morphological features as additional features and input them into the subsequent cell classification step to assist in improving the accuracy of cell classification.

[0016] Preferably, in the cell classification step, the confidence of the classification result is evaluated; the specific method is: after the model outputs the probability of each cell belonging to each category, the maximum probability value is extracted, and the probability value is used as the confidence of the classification result; A confidence threshold is preset, and the obtained confidence is compared with the confidence threshold. If the confidence is lower than the confidence threshold, the cell is marked as an uncertain sample and reclassified by manual review or by using a model trained with multiple different parameters.

[0017] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. The present invention designs and optimizes the algorithm through data flow, effectively controls the type of data input into the algorithm for training, reduces the difficulty of labeling, and improves the efficiency of negative screening. The algorithm first extracts the foreground area of the cervical liquid-based cytopathology image, and then cuts the image by regional segmentation during the training stage, strictly controls data input using small images, and performs fine segmentation on positive cells; finally, the segmented cells are classified. In the classification process, the accuracy of squamous epithelial lesions is improved, and the misdetection of low-level cells as high-level cells is effectively reduced.

[0018] 2. The present invention uses targeted technologies in foreground extraction, cell segmentation and classification through multi-step fine processing. Foreground extraction effectively removes background interference, cell segmentation uses convolutional neural networks to accurately identify and segment cell areas, cell classification combines the characteristics of diseased cells to subdivide types, and introduces morphological features to assist, and also conducts confidence assessment on classification results to reduce misjudgment, comprehensively improving the accuracy of diagnosis of cervical liquid-based cell squamous epithelial lesions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0020] Several embodiments of the present application will be described in more detail below with reference to the accompanying drawings so that those skilled in the art can implement the present application. The present application can be embodied in many different forms and purposes and should not be limited to the embodiments described herein. These embodiments are provided to make the present application comprehensive and complete, and to fully convey the scope of the present application to those skilled in the art. The embodiments do not limit the present application.

[0021] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the relevant art and / or the context of this specification, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0022] Please refer to Figure 1 as shown, the present invention provides a technical solution: An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells, comprising the following parts: Foreground extraction: Extract the foreground part from the whole image, removing background noise and irrelevant content; The foreground extraction specifically includes the following parts: Utilize deep learning technology to perform preliminary extraction of the foreground region, and a semantic segmentation network such as network can be adopted; The network is a symmetric convolutional neural network, consisting of an encoder and a decoder; the encoder part performs downsampling on the input image through multiple convolutional and pooling operations to extract high-level semantic features of the image; the decoder part performs upsampling on the feature map through deconvolution and skip connection operations to restore the spatial resolution of the image, and finally outputs the segmentation mask of the foreground region; Suppose the input image is , where respectively represent the height, width, and number of channels of the image; After being processed by the network, the foreground segmentation mask is output, where each pixel value is 0 or 1, 1 represents a foreground pixel, and 0 represents a background pixel; Perform contour detection on the foreground segmentation mask, using the function in to obtain a series of contours , and perform ellipse fitting on each contour using the function , and this function returns the center coordinates , semi-major axis , semi-minor axis of the ellipse, as well as the rotation angle of the major axis; Taking the center of the ellipse as the origin, convert the ellipse to polar coordinates, thereby obtaining a series of polar coordinate points, and convert the polar coordinate points back to Cartesian coordinates; and approximate the contour composed of the obtained Cartesian coordinate points to generate an approximately circular contour region; Specifically, it includes the following parts: In polar coordinates, the equation of an ellipse is , where is the polar radius, is the polar angle; samples are taken within a preset range of polar angles to obtain a series of polar coordinate points , ; Convert the polar coordinate points into Cartesian coordinates , where , ; Preprocessing of training data: Crop the image and filter the negative regions; The preprocessing of training data specifically includes the following parts: Cut the image into blocks along the circular region and crop using a preset window size; Assume the original image is , and the foreground region obtained after foreground extraction, with as the reference for the circular contour in , sliding and cropping on using the preset window size to obtain a series of small image pieces ; Analyze and judge the negative regions by calculating the average gray value features of each cut image, and then locate the regions containing diseased cells; Specifically, it includes the following parts: For an image with a size of , its gray value matrix is , , and the calculation formula for its average gray value is , traverse the gray values of each pixel point in the image and accumulate them, then divide by the total number of pixels to obtain the average gray value; Preset a gray threshold, compare the average gray value with the gray threshold, if the average gray value is less than the gray threshold, then determine the image corresponding to the average gray value as a negative region; Verify the images determined as negative regions by the average gray value, and extract the image texture features through the gray-level co-occurrence matrix, specifically including the following parts: The gray-level co-occurrence matrix is obtained by statistically counting the frequencies of gray value pairs in the image at a certain distance and direction; for a gray image, calculate its gray-level co-occurrence matrix at different directions (such as 0°, 45°, 90°, 135°) and different distances (such as 1 pixel, 2 pixels, etc.); Multiple texture features can be extracted from the gray-level co-occurrence matrix, such as contrast, correlation, energy, and entropy, etc. Contrast reflects the unevenness of the gray-level distribution in the image, correlation measures the linear correlation of the gray levels in the image, energy represents the uniformity of the gray-level distribution of the image, and entropy reflects the amount of information in the image. For the negative regions, their texture features usually show relatively low contrast, relatively high correlation, relatively high energy, and relatively low entropy. Through a large number of sample data, a threshold range of texture features can be determined. When the texture feature value of the calculated small piece of image exceeds this threshold range, then this small piece of image is more likely to be the region containing diseased cells. Combining the average gray-level value and texture features to judge the negative regions can effectively avoid misjudgments caused by single-feature judgment. Cell segmentation: Use a convolutional neural network for feature extraction, and identify and segment cell regions by training the model. Cell segmentation specifically includes the following parts: Cut the pathological image into images of a preset size and record them as small pieces of images, and input the learning rate, optimizer, and iteration parameter settings into the deep learning segmentation model. Assume that the input set of small piece of images is , and the corresponding set of cell segmentation labels is ; Use the trained model to predict the input small piece of image, obtain the mask image of all cell contours in this image, and then find the positions of non-zero elements according to the cell contour mask. Use the contour detection algorithm for post-processing to finally obtain the contours of the cells. The loss function of the network adopts the binary cross-entropy loss function: , where is the cell segmentation probability map predicted by the network; Use the optimizer to update the network parameters, and the update formula of the optimizer is: Calculate the gradient: ; Calculate the first moment estimate: ; Calculate the second moment estimate: ; Correct the first moment estimate: ; Correct the second moment estimate: ; Update the parameters: ; where is the network parameter of the th iteration, is the learning rate, is a preset constant, and are hyperparameters of the optimizer, usually ; In the cell segmentation step, morphological feature analysis is performed on the obtained cell contours; the morphological features include the area, perimeter, circularity, and aspect ratio of the cells; these morphological features are used as additional features and input into the subsequent cell classification step to assist in improving the accuracy of cell classification; Among them, the cell area is obtained by calculating the number of pixels contained in the contour; the perimeter is calculated using the function in ; the circularity calculation formula is ; the aspect ratio is the ratio of the length to the width of the circumscribed rectangle of the cell; Cell classification: First, determine whether the lesion is low-grade or high-grade, and then further subdivide the specific lesion. The low-grade is divided into , , and the high-grade is divided into , ; Model training is used to perform data augmentation such as image rotation and flipping, and the cross-entropy loss function and optimizer are used to learn the cell morphological features to achieve classification; Specifically, it includes: Judging the severity based on the overall feature differences of the diseased cells, which are the key features for distinguishing low-grade and high-grade lesions; Like the nuclear morphology, in low-grade lesions, although the nucleus is abnormal, it is relatively regular, and the size and shape change limitedly; in high-grade lesions, the nucleus is significantly enlarged and has a strange shape. In terms of chromatin distribution, the chromatin abnormality in low-grade lesions is not very significant, while in high-grade lesions, the chromatin is highly condensed and distributed disorderly. The model judges whether the cell lesion belongs to low-grade or high-grade based on these features; For cells determined to be low-grade lesions, they are further subdivided into and ; among them, means that the cell morphology has slight abnormalities, but the features are not typical and it is difficult to accurately judge the nature of the lesion; means that there is already a clear low-grade squamous intraepithelial lesion, and the cell morphology changes more significantly; such as nuclear enlargement, mild increase in the nuclear-cytoplasmic ratio, etc.; For cells determined to be high-grade lesions, they are further subdivided into and , Atypical squamous cells, with a high suspicion of high-grade squamous intraepithelial lesion, but the evidence is not yet sufficient. The degree of abnormality of these cells is between that of low-grade and typical high-grade lesion cells; Represents high-grade squamous intraepithelial lesion. The cell morphology shows significant abnormalities, with large and deeply stained nuclei, a significantly increased nuclear-cytoplasmic ratio, and disrupted cell polarity. It is a relatively severe lesion state that requires timely treatment intervention; Adopt The model is trained to perform data augmentation on the input cell images, specifically including: randomly rotating the cell images within a preset angle range to simulate different orientations of cells in actual samples, enabling the model to learn the characteristics of cells at various angles and improving the recognition ability of cells in different directions; Then, the cell images are horizontally flipped and vertically flipped; this can increase the diversity of the data, allowing the model to learn the symmetric characteristics of cells and further improving the generalization ability; And randomly scale the cell images within a preset ratio range to simulate the observation effects of cells at different magnifications, enabling the model to adapt to cells of different sizes and improving the robustness to cell size changes; By adjusting color attributes such as the brightness, contrast, and saturation of the images, the model learns the impact of color changes on cell characteristics, enhances the classification ability of the model under different color conditions, and is trained through the cross-entropy loss function and The optimizer. During training, monitor the accuracy and recall metrics of the model to make the model reach the preset classification accuracy.

[0023] Conduct a confidence evaluation on the classification results; The specific method is: after the model outputs the probabilities of each cell belonging to each category, extract the maximum probability value and use this probability value as the confidence of the classification result; Preset a confidence threshold and compare the obtained confidence with the confidence threshold. If the confidence is lower than the confidence threshold, mark the cell as an uncertain sample and use manual review or reclassify it using multiple models trained with different parameters again.

[0024] The above formulas are all obtained through software simulation by collecting a large amount of data and selecting a formula close to the true value. The influence weight factors and specific coefficient values in the formulas are set by those skilled in the art according to the actual situation and can be adjusted and modified later.

[0025] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells, characterized in that, It includes the following parts: Foreground extraction: Extract the foreground part from the entire image, removing background noise and irrelevant content; Training data preprocessing: Crop the pictures and filter the negative regions; Cell segmentation: Use a convolutional neural network for feature extraction and identify and segment the cell regions by training the model; Cell classification: First, determine whether the lesion is low-grade or high-grade, and then further classify the specific lesion. The low-grade lesions are classified into , , and the high-grade lesions are classified into , ; Use for model training, perform data augmentation such as image rotation and flipping, use the cross-entropy loss function and optimizer to learn the morphological features of cells, and achieve classification.

2. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 1, characterized in that, The foreground extraction specifically includes the following parts: For the initial extraction of the foreground region using deep learning techniques, a semantic segmentation network can be adopted, such as network; After network processing, a foreground segmentation mask is output; Perform contour detection on the foreground segmentation mask, using the function in to obtain a series of contours, and perform ellipse fitting on each contour using a function that returns the center coordinates, semi-major axis, semi-minor axis, and major axis rotation angle of the ellipse; Taking the center of the ellipse as the origin, convert the ellipse to polar coordinates to obtain a series of polar coordinate points, and then convert the polar coordinate points back to Cartesian coordinates; and approximate the contour formed by the obtained Cartesian coordinate points to generate an approximately circular contour region.

3. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 1, characterized in that, The training data preprocessing specifically includes the following parts: Cut the picture into blocks along the circular region and crop it using a preset window size; Judge the negative regions by analyzing the average gray value features of each cut picture, and then locate the regions containing diseased cells.

4. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 1, wherein The cell segmentation specifically includes the following parts: Cut the pathological image into images of a preset size and record them as small block images, and input the learning rate, optimizer, and iteration parameter settings into the deep learning segmentation model; assume that the input set of small block images is , and the corresponding cell segmentation label set is ; Use the trained model to predict the input small piece of image, obtain the mask image of all cell contours in the image, and then find the positions of non-zero elements according to the cell contour mask; Utilize The contour detection algorithm is used for post-processing, and finally the contours of the cells are obtained.

5. An algorithm for diagnosing cervical liquid-based cytology squamous epithelial lesions according to claim 1, characterized in that, First, determine whether the lesion is low-grade or high-grade, and then further classify the specific lesion. The low-grade lesions are divided into , , and the high-grade lesions are divided into , . Specifically, it includes the following parts: Judge the severity according to the overall feature differences of diseased cells, which are the key features for distinguishing low-grade and high-grade lesions; For cells determined to be of low-grade lesions, they are further subdivided into and ; among which indicates that the cell morphology has slight abnormalities, but the characteristics are not typical, making it difficult to accurately determine the nature of the lesion; indicates that there is already a clear low-grade squamous intraepithelial lesion, and the cell morphology changes are more obvious; For cells determined to be high-grade lesions, they are further subdivided into and , which refers to atypical squamous cells, representing high-grade squamous intraepithelial lesions.

6. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 5, characterized in that, During the model training, perform data augmentation such as image rotation and flipping, and use the cross-entropy loss function and optimizer to learn the cell morphological features for classification. Specifically, it includes the following parts: Adopt The model is used for training, and data augmentation processing is performed on the input cell images, specifically including: randomly rotating the cell images within a preset angle range; Then horizontally and vertically flip the cell images; And randomly scale the cell images within a preset ratio range; By adjusting color attributes such as the brightness, contrast, and saturation of the image, the model learns the impact of color changes on cell features, enhances the model's classification ability under different color conditions, and is trained using the cross-entropy loss function and the optimizer. During training, monitor the accuracy and recall metrics of the model to enable the model to achieve the preset classification accuracy.

7. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 4, characterized in that, In the cell segmentation step, perform morphological feature analysis on the obtained cell contours; the morphological features include the area, perimeter, circularity, and aspect ratio of the cells; use these morphological features as additional features and input them into the subsequent cell classification step to assist in improving the accuracy of cell classification.

8. An algorithm for diagnosing squamous epithelial lesions based on cervical liquid-based cells according to claim 1, characterized in that, In the cell classification step, evaluate the confidence of the classification results; the specific method is: after the model outputs the probability of each cell belonging to each category, extract the maximum probability value and use this probability value as the confidence of the classification result; Preset a confidence threshold, compare the obtained confidence with the confidence threshold. If the confidence is lower than the confidence threshold, mark the cell as an uncertain sample, and use manual review or reclassify it using models trained with multiple different parameters again.