A Zero-Shot Cervical Cell Detection Method Based on Semantic Alignment
Through a two-stage object detection method based on semantic alignment and sparse annotation, high-precision cervical cell detection without labeling data is achieved, solving the problems of high data labeling cost and misdiagnosis in traditional methods, and improving the accuracy and efficiency of cervical cell detection.
Patent Information
- Application Number
- CN202310375319.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-10
AI Technical Summary
In the prior art, the cervical cell detection model relies on large-scale high-quality annotation data sets, resulting in high data annotation cost, limiting its application in cervical cytology, and traditional methods have problems of misdiagnosis and misdiagnosis.
Using a zero-sample cervical cell detection method based on semantic alignment, the knowledge of the text image alignment model is extracted into the object detection model by collecting open source object detection data sets and using knowledge distillation technology, without any cervical cell annotation, and a sparse annotation object detection model is constructed, and the model is trained using collaborative information mining method to reduce missed recognition.
It realizes high-precision cervical cell detection, reduces data collection costs, and improves the accuracy and reliability of the detection, solving the problem that the zero-sample object detection model is difficult to accurately detect cervical cells.
Smart Images

Figure CN116309525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical image processing, and particularly to the target detection task of cervical cells. Background Art
[0002] Cervical cancer is one of the common malignant tumors among women globally. Early diagnosis and treatment play a crucial role in the treatment effect and survival rate of patients. Cervical cytology is an important means for cervical cancer screening. By observing the morphology of patients' cervical cells, the cervical lesions of patients can be detected. However, traditional cervical cytology examinations require a large amount of manual operations, which are not only time-consuming and laborious, but also may have problems such as missed diagnosis and misdiagnosis, affecting the diagnostic effect and treatment effect. Computer-aided diagnosis technology is an effective measure to solve this problem, and cervical cell target detection is the basis of computer-aided diagnosis. It can locate and identify cervical cells, providing a cell basis for subsequent analysis and processing.
[0003] Currently, high-performance cervical cell target detection models rely on large-scale high-quality labeled datasets. However, cervical cells are dense and extremely numerous in quantity. The annotation of data requires a large amount of time and effort from pathologists. Therefore, the cost of constructing a large-scale dataset is extremely high, which limits the application of target detection methods in cervical cytology. To solve this problem, the present invention proposes a zero-shot cervical cell detection method based on semantic alignment. This method is trained by collecting publicly available target detection datasets, and uses knowledge distillation technology to extract the knowledge of the text-image alignment model into the target detection model, achieving zero-shot cervical cell detection without any cervical cell annotation. In addition, we also propose a sparse annotation target detection model, which is trained using a collaborative information mining method, so as to effectively mine unlabeled target instances and solve the problem of missed recognition in zero-shot cervical cell detection. The present invention combines these two models to construct a two-stage target detection method for high-precision cervical cell detection. This method can achieve high-precision cervical cell detection without any cervical cell target detection annotation, providing a basis for intelligent diagnosis of cytopathology. Summary of the Invention
[0004] The present invention aims at the problem of cervical cell target detection under the lack of labeled data, and proposes a zero-shot cervical cell detection method based on semantic alignment.
[0005] The above invention objective is mainly achieved through the following technical solutions:
[0006] S1. Collect open-source target detection datasets for model training. The specific steps are as follows:
[0007] Collect open-source object detection datasets, including general object detection datasets and open-source datasets specifically for medical images; in addition, perform data augmentation on the collected datasets using operations such as rotation, translation, scaling, cropping, and blurring.
[0008] S2. Collect cervical cell images for model training and validation, and the specific steps are as follows:
[0009] Collect panoramic cervical cell images through an automatic scanner, and crop the panoramic images into small images with both width and height of M; then, roughly segment cervical cells using image processing techniques and discard blank images without cells; finally, obtain an image dataset of unlabeled cervical cells.
[0010] S3. Build a zero-shot object detection model based on semantic alignment for detecting cervical cells, and the specific steps are as follows:
[0011] The model distills the knowledge of the text-image alignment model into the object detection model to achieve zero-shot object detection; first, use a pre-trained text-image alignment model to encode the class text description and the object detection object; then, use the object detection dataset collected in S1 to train a zero-shot object detection model; during the training process, align the detection box region feature encoding of the object detection model with the corresponding text-image alignment model encoding; after training, the detection box region feature encoding of the object detection model has achieved semantic alignment with the class text description; thus, by comparing the distance between the detection box region feature encoding and the text feature of the cervical cell class description, zero-shot object detection of cervical cells is achieved.
[0012] Use the CLIP and ALIGN text-image alignment pre-trained large models as the teacher models, and extract the knowledge in the teacher models into the object detection model through knowledge distillation; the object detection model here selects an object detection model with a Region Proposal Network (RPN); during the training process, input an image, and the image has multiple objects r∈P, where P is the set of objects and r is one of the objects. First, use the teacher model to extract the encoded images F v (r) of multiple objects labeled in the image, and extract the encoded text descriptions t i of C classes in the dataset, and then obtain the region of interest feature encoding F p (r) in the object detection model; finally, use knowledge distillation to achieve knowledge transfer from the teacher model to the object detection model, and the specific loss function L is as follows:
[0013] L = L text +L image (1)
[0014]
[0015] Among them, F bg is a learnable parameter used to represent the background category, t i represents the text description encoding of each category in the dataset, C represents the total number of target categories in the dataset, and y r represents the category label of region r, M represents the number of targets in the input image, and L CE represents the cross-entropy loss, and τ represents the temperature hyperparameter of the softmax function.
[0016] S4. Construct a sparse annotation object detection model to reduce the situation of missed recognition of cervical cells. The specific steps are as follows:
[0017] Adopt a collaborative information mining method to train the sparse annotation object detection model; this model has two object detection branches with shared parameters. Among them, the first branch takes the original image as input, and the second branch takes the corresponding enhanced image as input; then, combine the prediction results of the two branches to mine unlabeled object instances; finally, use the mined unlabeled object instances and the original labels to update the model parameters.
[0018] The two branches of the sparse annotation object detection model are respectively composed of the same object detection model, and the parameters of these two branches are shared; during the training process of the model, input a sparsely labeled image x and the corresponding annotation y, and perform rotation, translation, scaling, cropping, and blurring operations on the image to obtain the enhanced image x a , and input x and x a into the two branches of the model respectively to obtain the prediction results p and p a ; and perform further post-processing on p and p a . First, delete the prediction boxes with confidence lower than T; then, delete the redundant object boxes through non-maximum suppression. At the same time, it is also necessary to delete the object boxes overlapping with the annotation to obtain the processed prediction results p′ and p a ′; in order to train the model, it is also necessary to merge p′ and p′ a with the annotated y to obtain y′ = p′ a ∪y and y′ a = p′ ∪ y two annotated data.
[0019] S5. Propose a high-precision two-stage object detection method for cervical cell detection. The specific steps are as follows:
[0020] S5.1. Use the open-source object detection dataset collected in S1 to train a zero-shot object detection model based on semantic alignment so that the model can detect cervical cells;
[0021] S5.2. Use the zero-shot object detection model based on semantic alignment to predict the cervical cell images collected in S2, and construct an incompletely annotated cervical cell object detection dataset;
[0022] S5.3. Use the incompletely annotated cervical cell object detection dataset to train a sparsely annotated object detection model;
[0023] S5.4. Finally, use the sparsely annotated object detection model to achieve high-precision object detection of cervical cells.
[0024] Advantages of the Invention
[0025] The present invention proposes a zero-shot cervical cell detection method based on semantic alignment, which realizes high-precision zero-shot object detection of cervical cells through semantic alignment between text and images. Traditional object detection methods require a large amount of labeled data to train the model, while the present invention only relies on open-source datasets and realizes high-precision cervical cell detection without annotating cervical cell images, greatly reducing the cost of data collection. In addition, the present invention proposes a sparsely annotated object detection method, which fully utilizes the inaccurate and undetected results predicted by the zero-shot object detection model to train the model, and solves the problem that it is difficult for the zero-shot object detection model to accurately detect all cervical cells because cervical cells are densely distributed and have a small difference from the background. The two-stage object detection method proposed by the present invention combines the zero-shot object detection model based on semantic alignment and the sparsely annotated object detection model, which can detect cervical cells more accurately and improve the accuracy of object detection.
[0026] In summary, the present invention provides a novel and efficient cervical cell detection method, which can achieve high-precision cervical cell object detection without any cervical cell annotation, providing important technical support for intelligent cervical cytology diagnosis and medical research. Brief Description of the Drawings
[0027] Figure 1 It is the overall design diagram of the algorithm;
[0028] Figure 2 It is the structural schematic diagram of the zero-shot object detection model based on semantic alignment;
[0029] Figure 3 It is the structural schematic diagram of the sparsely annotated object detection model.
[0030] Specific Implementation Method Specific Embodiments
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0032] As Figure 1 shown, the zero-shot cervical cell detection method based on semantic alignment proposed in this paper mainly includes the following steps:
[0033] S1. Collect open-source object detection datasets for model training;
[0034] S2. Collect cervical cell images for model training and validation;
[0035] S3. Build a zero-shot object detection model based on semantic alignment;
[0036] S4. Build a sparse annotation object detection model;
[0037] S5. Propose a high-precision two-stage object detection method for cervical cell detection.
[0038] In the embodiments of the present invention, first, it is necessary to collect open-source object detection datasets for training the zero-shot object detection model based on semantic alignment; then use the zero-shot object detection model to predict the collected cervical cell images to obtain an incompletely and inaccurately annotated cervical cell annotation dataset for training the sparse annotation object detection model; finally, use the trained sparse annotation object detection model to achieve high-precision object detection of cervical cells.
[0039] The following will explain the embodiments of the present invention in detail:
[0040] As Figure 1 shown, the implementation of the algorithm includes the steps:
[0041] S1. Collect open-source object detection datasets for model training:
[0042] Collect general object detection datasets such as COCO, Open Images, Objects365, PASCAL VOC, LVIS, and KITTI. Also collect medical image object detection datasets such as Hyper-Kvasir, CBC, FOD-A, ISBI, and CRIC to build a multi-class large-scale object detection dataset; in addition, perform data augmentation on the collected datasets using operations such as rotation, translation, scaling, cropping, and blurring.
[0043] S2. Collect cervical cell images for model training and validation:
[0044] Collect panoramic cervical cell images through an automatic scanner, and crop the panoramic images into small images with a width and height of 512 pixels. Then, use adaptive threshold segmentation to segment the small images, and then remove the noise points with an area less than 25×25 pixels to obtain a rough cell segmentation result. When the number of cells is less than 2, mark the small image as a blank image and discard it. Finally, obtain an image dataset of unannotated cervical cells.
[0045] S3. Build a zero-shot object detection model based on semantic alignment:
[0046] The model structure is as Figure 2 shown. The model extracts the knowledge of the text-image alignment model CLIP into the object detection model Faster R-CNN through knowledge distillation, so that Faster R-CNN can achieve zero-shot object detection. First, use the pre-trained CLIP model to encode the class text description and the object detection object. Then, use the object detection dataset collected in S1 to train the zero-shot object detection model. During the training process, align the region of interest feature encoding in the RPN network of Faster R-CNN with the corresponding CLIP image and text encoding. After training, the region of interest feature encoding of Faster R-CNN has been aligned with the text semantic encoding. Therefore, by comparing the distance between the region of interest feature and the cervical cell description text feature, zero-shot object detection of cervical cells is achieved.
[0047] Use CLIP as the teacher model, and extract the knowledge in the teacher model into Faster R-CNN through knowledge distillation. During the training process, input an image, and the image has multiple targets r∈P, where P is the set of targets and r is one of the targets. First, use the teacher model to extract the encoded images of multiple targets F v (r) marked in the image, and extract the encoded text descriptions t i of C categories in the dataset. Then, obtain the region of interest feature encoding F p (r) in the RPN network of Faster R-CNN. Finally, use knowledge distillation to achieve knowledge transfer from the teacher model to the object detection model. The specific loss function L is as follows:
[0048] L = L text + L image (5)
[0049]
[0050] z(r) = [sim(F p (r), Fbg ), sim(F p (r), t1), …, sim(F p (r), t C )] (7)
[0051]
[0052] Among them, F bg is a learnable parameter used to represent the background category, t i represents the text description encoding of each category in the dataset, C represents the total number of categories of the targets in the dataset, and y r represents the category label of region r, M represents the number of targets in the input image. Here, the upper limit of the number of targets is set to 200, and the excess ones are randomly discarded. L CE represents the cross-entropy loss, and τ represents the temperature hyperparameter of the softmax function, which is set to 0.01 here.
[0053] S4. Construct a sparse annotation object detection model to reduce the situation of missed recognition of cervical cells. The specific steps are as follows:
[0054] The structure of the model is as Figure 3 shown. The model has two object detection branches with shared parameters. Among them, the first branch takes the original image as the input, and the second branch takes the corresponding enhanced image as the input; combines the prediction results of the two branches to mine unlabeled object instances; finally, uses the mined unlabeled and original labels to update the model parameters.
[0055] The two branches of the sparse annotation object detection model are respectively composed of yolov5, and the parameters of these two branches are shared; during the training process of the model, input a sparsely labeled image x and the corresponding annotation y, and perform rotation, translation, scaling, cropping, and blurring operations on the image to obtain the enhanced image x a , and input x and x a into the two branches of the model respectively to obtain the prediction results p and p a ; and perform further post-processing on p and p a . First, delete the prediction boxes with a confidence level lower than 0.65; then delete the redundant object boxes through non-maximum suppression, and at the same time, also need to delete the object boxes overlapping with the annotation to obtain the processed prediction results p′ and p′ a ; in order to train the model, it is also necessary to merge p′ and p′ a with the annotated y to obtain y′ = p′ a ∪y and y′ a = p′ ∪ y for the two annotation data.
[0056] S5. Propose a high-precision two-stage object detection method for cervical cell detection, and the specific steps are as follows:
[0057] S5.1. Use the open-source object detection dataset collected in S1 to train a zero-shot object detection model based on semantic alignment, so that the model can detect cervical cells;
[0058] S5.2. Use the zero-shot object detection model based on semantic alignment to predict the cervical cell images collected in S2, and construct an incompletely annotated cervical cell object detection dataset;
[0059] S5.3. Use the incompletely annotated cervical cell object detection dataset to train a sparse annotation object detection model;
[0060] S5.4. Finally, use the sparse annotation object detection model to achieve high-precision object detection of cervical cells.
Claims
1. A zero-shot cervical cell detection method based on semantic alignment, characterized in that, It includes the following steps: S1. Collect an open-source object detection dataset for model training; S2. Collect cervical cell images for model training and validation; S3. Build a zero-shot object detection model based on semantic alignment, which distills the knowledge of the text-image alignment model into the object detection model to achieve zero-shot object detection; First, use a pre-trained text-image alignment model to encode the class text description and the object detection object; Then, use the object detection dataset collected in S1 to train a zero-shot object detection model; During the training process, align the detection box region feature encoding of the object detection model with the corresponding text-image alignment model encoding; After training, the detection box region feature encoding of the object detection model has achieved semantic alignment with the class text description; Therefore, by comparing the distance between the detection box region feature encoding and the text feature of the cervical cell class description, zero-shot object detection of cervical cells is achieved; S4. Build a sparse annotation object detection model and use a collaborative information mining method to train the sparse annotation object detection model; This model has two object detection branches with shared parameters. Among them, the first branch takes the original image as input, and the second branch takes the corresponding enhanced image as input; Then, combine the prediction results of the two branches to mine unlabeled object instances; Finally, use the mined unlabeled object instances and the original labels to update the model parameters; S5. Combine the models built in S3 and S4 to propose a high-precision two-stage object detection method for cervical cell detection.
2. The zero-shot cervical cell detection method based on semantic alignment according to claim 1, wherein In step S1, an open-source object detection dataset is collected for model training, specifically: Collect an open-source object detection dataset, including general object detection datasets and open-source datasets specifically for medical images; In addition, use operations such as rotation, translation, scaling, cropping, and blurring to perform data augmentation on the collected dataset.
3. The zero-shot cervical cell detection method based on semantic alignment according to claim 1, characterized in that, In step S2, cervical cell images are collected for model training and validation, specifically: Collect panoramic cervical cell images through an automatic scanner, crop the panoramic images into small images with a width and height of M; Then, use image processing techniques to roughly segment cervical cells and discard blank images without cells; Finally, obtain an image dataset of unannotated cervical cells.
4. The zero-sample cervical cell detection method based on semantic alignment according to claim 1, characterized in that, In step S3, build a zero-shot object detection model based on semantic alignment, specifically: Using the CLIP and ALIGN text-image alignment pre-trained large models as teacher models, extract the knowledge in the teacher models into the object detection model through knowledge distillation; the object detection model here selects the object detection model with a Region Proposal Network (RPN); during training, input an image with multiple objects r ∈ P, where P is the set of objects and r is one of the objects. First, use the teacher model to extract the encoded images of the multiple objects marked in the image F v (r), and extract the encoded text descriptions t of C categories in the dataset i , and then obtain the feature encoding F of the region of interest in the object detection model p (r); finally, use knowledge distillation to achieve knowledge transfer from the teacher model to the object detection model. The specific loss function L is as follows: L = L text + L image (1) z(r) = [sim(F p (r), F bg ), sim(F p (r), t1), …, sim(F p (r), t C )] (3) Among them, F bg is a learnable parameter used to represent the background category, t i represents the text description encoding of each category in the dataset, C represents the total number of target categories in the dataset, y r represents the category label of region r, M represents the number of targets in the input image, L CE represents the cross-entropy loss, and τ represents the temperature hyperparameter of the softmax function.
5. The zero-shot cervical cell detection method based on semantic alignment according to claim 1, characterized in that, In step S4, build a sparse annotation object detection model, specifically: The two branches of the sparse annotation object detection model are respectively composed of the same object detection model, and the parameters of these two branches are shared; during the training process of the model, a sparsely labeled image x and the corresponding annotation y are input, and the image is enhanced by rotation, translation, scaling, cropping, and blurring operations to obtain the enhanced image x a , x and x a are respectively input into the two branches of the model to obtain the prediction results p and p a ; and further post-processing is performed on p and p a . First, the prediction boxes with confidence lower than T are deleted; then, the redundant object boxes are deleted through non-maximum suppression, and at the same time, the object boxes overlapping with the annotation also need to be deleted to obtain the processed prediction results p′ and p′ a ; in order to train the model, p′ and p′ a also need to be merged with the annotated y to obtain y′ = p′ a ∪y and y′ a = p′ ∪ y, two annotated data sets.
6. The zero-shot cervical cell detection method based on semantic alignment according to claim 1, characterized in that, In step S5, combine the models built in S3 and S4 to propose a high-precision two-stage object detection method for cervical cell detection, specifically: S5.1 Use the open-source object detection dataset collected in S1 to train a zero-shot object detection model based on semantic alignment so that the model can detect cervical cells; S5.2 Use the zero-shot object detection model based on semantic alignment to predict the cervical cell images collected in S2 to build an incompletely annotated cervical cell object detection dataset; S5.3 Use the incompletely annotated cervical cell object detection dataset to train the sparse annotation object detection model; S5.4 Finally, use the sparse annotation object detection model to achieve high-precision object detection of cervical cells.
Citation Information
Patent Citations
A Mask-RCNN-based cervical cell smear image segmentation method and system
CN109886179A
Semantic enhanced hash method for zero-sample image retrieval
CN111274424A