CNN-based cervical squamous epithelial cell number calculation method
Through CNN-based methods, we learn to identify cervical squamous epithelial cells and generate density maps, which solves the problem of large errors and long time-consuming cervical squamous epithelial cells counting in the prior art, and achieves rapid and accurate cell counting to meet the needs of cervical cancer screening.
Patent Information
- Application Number
- CN202311613073.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately and quickly calculate the number of cervical squamous epithelial cells, resulting in large errors, long time and poor repetition, which cannot meet the needs of cervical cancer screening.
Using a CNN-based method, cervical squamous epithelial cells are identified and counted by learning the nonlinear mapping of input images and output density maps, cervical squamous epithelial cells are excluded, inflammatory cells, cervical canal cells and metaplastic cells are eliminated, and rapid and accurate cell counting is achieved.
It has achieved rapid and accurate counting of the number of cervical squamous epithelial cells, with small errors and strong repeatability, and can calculate the total number of cells on the entire slide to meet the needs of cervical cancer screening.
Smart Images

Figure CN120070298A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision in cytopathology, and particularly relates to a method for calculating the number of cervical squamous epithelial cells based on CNN. Background Art
[0002] Cervical cancer is the most common gynecological malignant tumor. According to WHO statistics, as of 2018, the annual number of new cases of cervical cancer globally was 569,847, of which China accounted for 18.6%, and the mortality rate accounted for 15%. China has approximately 106,430 new cases each year, ranking third among women aged 15 - 44, and approximately 47,739 deaths each year, ranking second among women aged 15 - 44. Approximately 131 Chinese women die of cervical cancer every day. From WHO data, it is known that the lifetime risk of Chinese women suffering from cervical cancer is approximately 0.9%. At present, the incidence and mortality rates of cervical cancer in China are showing an increasing trend year by year.
[0003] The development of cervical cancer is a long - term process. If screening can be carried out in the early stage of the lesion, the occurrence and development of canceration can be prevented. Therefore, cervical cancer screening is of great significance. If it can be detected, diagnosed, and treated early, the incidence and fatality rates of cervical cancer can be reduced, and the impact on the physical health and life of patients can be minimized.
[0004] In the 1940s, the famous Greek scientist Georgios Papanikolaou invented the cervical cancer detection technology called "Pap Smear". Its birth made the early screening of cervical cancer possible. By the end of the last century, in European and American countries where Pap Smear was widely used, the mortality rate dropped by half. Therefore, Pap Smear is called "one of the most important medical discoveries of the 20th century". To this day, Pap Smear remains the gold standard for detecting cervical cancer from cervical exfoliated cells.
[0005] In the mid-late 1990s, the "Second Generation Cytology" technology revolution emerged, namely: liquid-based thin-layer cytology test, abbreviated as TCT. The technology and the dedicated device ThinPrep2000 liquid-based cytometer invented by the American company Cytyc were used for cervical cytology testing methods, which obtained the US FDA certification for clinical use in 1996. It is a technology that preserves cervical exfoliated cells in a liquid and makes a smear by evenly dispersing and attaching the cells to a glass slide through a special device. It uses mechanical, pneumatic and hydrodynamic principles to evenly coat the cells into a thin-layer specimen with a fully automatic device. And it can greatly reduce the number of blood, mucus and inflammatory cells that block the view. The pneumatic and hydrodynamic system controlled by a microcomputer is used to monitor cell collection, making the cell transfer process successfully completed, thus making a thin-layer sample with dispersed cells. The TCT technology is the only one that has obtained the FDA certification and combines the use of the TBS classification standard, which can improve the detection of precancerous lesions.
[0006] The TBS (The Bethesda System) cervical cytology reporting system was proposed at a meeting in Bethesda, USA by the American Cancer Society and the Cervical Cancer Society in 1988, and the TBS classification standard was officially published in 1994. The second edition of the new TBS classification method was proposed in 2001. So far, the application of TBS has gone through nearly 30 years, and its occurrence and development have had a profound impact on cervical cytology and even gynecological clinical practice. With the accumulation and changes of practical experience, the third edition (2004) of TBS supplemented the specimen quality interpretation standard and non-neoplastic cell types, and added descriptions and morphological differentiations of difficult types in epithelial cell abnormalities.
[0007] Specimen quality assessment: TBS-2014 basically maintains the specimen quality assessment standard of TBS-2001, but points out that: A, for women with a cervix, the satisfactory specimen should have no less than 5,000 well-preserved squamous epithelial cells in liquid-based preparations and no less than 8,000 in traditional smears; for women after radiotherapy or chemotherapy, after hysterectomy, or with postmenopausal atrophy, the cell count in the smear does not have to strictly follow this standard. For example, in the liquid-based preparations of such women, if there are no less than 2,000 well-preserved squamous epithelial cells, it can be judged as a satisfactory specimen. Because even if these women have a cervix, their anatomy has changed, and in fact, the specimens taken are from the vagina; B, a satisfactory specimen needs to describe the presence or absence of endocervical cells and metaplastic cells. For atrophic smears, it is really difficult to distinguish metaplastic cells from outer basal cells, and it can be noted in the annotation.
[0008] There are generally two methods for evaluating the sample quality standards of cervical specimens. The first method is that the experimenter observes the slices under a microscope, randomly selects 10 fields of view, and counts them with a handheld press counter. This method has two disadvantages. The first is that the number of cells on the slices varies and is uneven, so there is inaccuracy in such random 10 fields of view. The second is that this method is not repeatable because the slide position is manually moved for observation. The second method is to use a scanner to scan the slides for digital images and save them on a computer, and then use traditional image algorithms to process the images. The steps are as follows: grayscale -> select a threshold for binarization -> connected domain analysis -> obtain the position and segmentation mask of each cell, but these steps cannot handle clustered cell clusters, and cannot distinguish impurities and non-specifically stained cells, so they can only be segmented out. The error rate is still quite high, because it only uses simple grayscale threshold binarization and can only deal with simple cell scenes.
[0009] The TBS system requires that a satisfactory sample must have more than 5,000 squamous epithelial cells. Currently, there is no application on the market that can accurately calculate the squamous epithelial cells in a sample. All counting is done manually under a microscope, which results in large errors, long time consumption, poor repeatability, and the inability to estimate the total number of cells on the entire slide.
[0010] Therefore, there is an urgent need for a method for counting cervical squamous epithelial cells that has small errors, short time consumption, is repeatable, and can estimate the total number. Summary of the invention
[0011] The purpose of the present invention is to provide a method for calculating the number of cervical squamous epithelial cells based on CNN. Based on the nonlinear mapping of CNN learning input images and output density maps, the characteristics of complete cervical squamous epithelial cells can be learned. At the same time, not only the most similar inflammatory cells can be excluded, but also endocervical cells, metaplastic cells, tissue cells, etc. can be excluded. The method has fast speed, strong repeatability, small error, and can also calculate the total number of cells on the entire slide.
[0012] In order to solve the above technical problems, the embodiment of the present invention discloses a CNN method for calculating the number of cervical squamous epithelial cells, comprising the following steps:
[0013] Obtain a pathological image of cervical squamous epithelial cells to be trained, where the height of the pathological image is H and the width is W;
[0014] marking the cervical squamous epithelial cells in the pathological image, and recording the coordinates of the marked cervical squamous epithelial cells on the pathological image;
[0015] According to the pathological image, a single-channel matrix of size H×W is generated, in which the elements at the positions where the coordinates of the marked cervical squamous epithelial cells are located are 1, and the elements at other positions are 0;
[0016] Performing Gaussian filtering on each element 1 in the single-channel matrix, thereby obtaining label data corresponding to the pathological image;
[0017] Training a convolutional neural network model based on the pathological image and label data corresponding to the pathological image to obtain a single-channel density map generation model;
[0018] Inputting the pathological image of the cervical squamous epithelial cells to be detected into the single-channel density map generation model to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cells to be detected;
[0019] The values in the single-channel density map matrix are summed to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cells to be detected.
[0020] In another preferred embodiment, before the step of acquiring the pathological image of the cervical squamous epithelial cells to be trained, the following steps are also included:
[0021] Performing immunocytochemical staining on the cervical squamous epithelial cell sample using an antibody reagent targeting a protein marker;
[0022] The cell sections treated with immunocytochemical staining are scanned by a digital pathology scanner to obtain the pathological digital microscopic image of the cervical squamous epithelial cells.
[0023] In another preferred embodiment, in the step of marking the nuclei of the cervical squamous epithelial cells in the pathological image, only the cervical squamous epithelial cells including the nuclei and the complete cytoplasm are marked.
[0024] In another preferred embodiment, the step of marking the cervical squamous epithelial cells in the pathological image and recording the coordinates of the marked cervical squamous epithelial cells on the pathological image includes the following sub-steps:
[0025] The cell nuclei of the cervical squamous epithelial cells in the pathological image are marked, and the coordinates of the cell nuclei of the marked cervical squamous epithelial cells on the pathological image are recorded.
[0026] In another preferred embodiment, the nucleus of the cervical squamous epithelial cells after immunocytochemical staining is located in the center of the cell and is dark blue, and the cytoplasm is light blue;
[0027] In the step of annotating the nuclei of cervical squamous epithelial cells in the pathological image, the following sub-steps are included:
[0028] Mark a point at the center position of the nucleus of the cervical squamous epithelial cell.
[0029] In another preferred example, the height of the pathological image is 2048 and the width is 2448.
[0030] In another preferred example, in the step of training a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the label data is pre-multiplied by the value 100.
[0031] In another preferred example, in the step of training a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the mean squared error loss function is selected as the loss function of the convolutional neural network model.
[0032] In another preferred example, in the step of training a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the mean absolute error function is selected as the evaluation function of the convolutional neural network model.
[0033] An embodiment of the present invention also discloses a device for calculating the number of cervical squamous epithelial cells based on CNN, including:
[0034] An image acquisition module, configured to acquire a pathological image of cervical squamous epithelial cells to be trained, where the height of the pathological image is H and the width is W;
[0035] An annotation module, configured to annotate the cervical squamous epithelial cells in the pathological image and record the coordinates of the annotated cervical squamous epithelial cells on the pathological image;
[0036] A single-channel matrix generation module, configured to generate a single-channel matrix with a size of H×W according to the pathological image, where the element at the position where the coordinates of the annotated cervical squamous epithelial cells are located in the single-channel matrix is 1, and the elements at the remaining positions are all 0;
[0037] A Gaussian filtering processing module, configured to perform Gaussian filtering processing on each element 1 in the single-channel matrix to obtain label data corresponding to the pathological image;
[0038] A training module, configured to train a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image to obtain a single-channel density map generation model;
[0039] A detection module, configured to input a pathological image of cervical squamous epithelial cells to be detected into the single-channel density map generation model, so as to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cells to be detected;
[0040] A calculation module, configured to sum the values in the single-channel density map matrix, so as to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cells to be detected.
[0041] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here.
[0042] Compared with the prior art, the main differences and effects of the embodiments of the present invention are as follows:
[0043] The method for calculating the number of cervical squamous epithelial cells based on CNN learns the non-linear mapping between the input image and the output density map based on CNN, can learn the complete features of squamous epithelial cells, and can not only exclude the most similar inflammatory cells, but also exclude endocervical cells, metaplastic cells, histiocytes, etc. It has the advantages of fast speed, strong repeatability, small error, and can also calculate the total number of cells on the whole glass slide.
[0044] Learning the mapping from the input image to the density map through CNN is different from the regression from the input image to a single value, different from image segmentation + post-processing, different from object detection, and even more different from traditional image processing without learning ability.
[0045] The network model required for the method for calculating the number of cervical squamous epithelial cells based on CNN has the advantages of small calculation amount, low complexity, and short operation time.
[0046] During training, the label data is multiplied by the value of 100, which makes the network training easier to converge, the oscillation more stable, and the evaluation function value (MAE) of the trained model on the validation set lower, that is, in actual application, the error between the model prediction value and the actual value is smaller. Description of the Drawings
[0047] Figure 1 is a flowchart of a method for calculating the number of cervical squamous epithelial cells based on CNN in the first embodiment of the present invention;
[0048] Figure 2 is an example diagram of cervical squamous epithelial cells after immunocytochemical staining;
[0049] Figure 3 is an example diagram of inflammatory cells after immunocytochemical staining;
[0050] Figure 4 is a flowchart of a preferred embodiment in the first embodiment of the present invention;
[0051] Figure 5 is an example diagram of a selected picture to be labeled;
[0052] Figure 6 is a schematic diagram of a labelme software interface;
[0053] Figure 7 is a schematic diagram of a labeling step;
[0054] Figure 8 is Figure 7 a schematic diagram after labeling;
[0055] Figure 9 is a screenshot of a record file generated by labeling;
[0056] Figure 10 is a visualization comparison diagram of the original input diagram and the density diagram;
[0057] Figure 11 is a schematic diagram of dilated convolution;
[0058] Figure 12 and Figure 13 is a comparison schematic diagram of the original input diagram, the true density diagram, and the model prediction diagram;
[0059] Figure 14 is a schematic diagram of a glass slide snapshot;
[0060] Figure 15 is a schematic diagram of the positions of five selected middle pictures;
[0061] Figure 16 is a schematic diagram of the coordinates selected in the actual code;
[0062] Figure 17 The application effect diagram of the method for calculating the number of cervical squamous epithelial cells based on CNN in software;
[0063] Figure 18 is a comparison diagram of the average of three experiments of 10 visual fields of 10 samples for each researcher;
[0064] Figure 19 is a comparison diagram of the coefficient of variation of different samples for each researcher;
[0065] Figure 20 is a comparison diagram of the average time of three experiments of each researcher and the three experiments of CellQC;
[0066] Figure 21It is a comparison chart of the coefficient of variation between 8 researchers and CellQC on different samples;
[0067] Figure 22 It is a schematic structural diagram of a device for calculating the number of cervical squamous epithelial cells based on CNN in the second embodiment of the present invention. Specific Embodiments
[0068] In the following description, many technical details are presented for the convenience of the reader's understanding of the present application. However, those of ordinary skill in the art can understand that the technical solutions claimed in the various claims of the present application can be implemented even without these technical details and various changes and modifications based on the following embodiments.
[0069] Before describing the present invention, it should be understood that the present invention is not limited to the specific methods and experimental conditions described, as such methods and conditions can vary. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting, and the scope of the present invention will be limited only by the appended claims.
[0070] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0071] The embodiments of the present invention relate to a method for calculating the number of cervical squamous epithelial cells based on CNN. Figure 1 It is a flowchart of the method for calculating the number of cervical squamous epithelial cells based on CNN.
[0072] Specifically, as Figure 1 shown, the method for calculating the number of cervical squamous epithelial cells based on CNN includes the following steps:
[0073] In step 101, a pathological image of cervical squamous epithelial cells to be trained is obtained, and the height of the pathological image is H and the width is W.
[0074] In this embodiment, preferably, before step 101, the following steps may further be included:
[0075] The cervical squamous epithelial cell sample is subjected to immunocytochemical staining treatment using an antibody reagent for a protein marker target.
[0076] Specifically, for the cervical squamous epithelial cell sample, it is stained using an antibody reagent for a protein marker target, that is: immunocytochemistry (ICC) methodology staining.
[0077] The cell section treated by immunocytochemical staining is scanned by a digital pathology scanner, so as to obtain the pathological digital microscopic image of the cervical squamous epithelial cells.
[0078] It should be noted that the heights and widths of the pathological images scanned by different scanners are different.
[0079] In this embodiment, preferably, the height of the pathological image is 2048 and the width is 2448.
[0080] Of course, this is only a preferred embodiment. In some other embodiments of the present invention, other heights and widths can also be adopted, and it is not limited thereto.
[0081] Thereafter, step 102 is entered, the cervical squamous epithelial cells in the pathological image are labeled, and the coordinates of the labeled cervical squamous epithelial cells on the pathological image are recorded.
[0082] In this embodiment, preferably, in step 102, only the cervical squamous epithelial cells containing nuclei and complete cytoplasm are labeled.
[0083] That is to say, when labeling the cervical squamous epithelial cells, it is strictly in accordance with the rule that only the cells with complete cytoplasm are considered as cervical squamous epithelial cells, so that inflammatory cells will not be misjudged as cervical squamous epithelial cells. This is because cervical squamous epithelial cells contain nuclei and complete cytoplasm, while inflammatory cells only have one nucleus, and the nuclei of inflammatory cells are the same as those of cervical squamous epithelial cells.
[0084] Furthermore, preferably, step 102 may further include the following sub-steps:
[0085] The nuclei of the cervical squamous epithelial cells in the pathological image are labeled, and the coordinates of the labeled nuclei of the cervical squamous epithelial cells on the pathological image are recorded.
[0086] The nucleus of the cervical squamous epithelial cell after immunocytochemical staining is located at the center of the cell, showing dark blue, and the cytoplasm shows light blue.
[0087] Figure 2 is an example diagram of a cervical squamous epithelial cell after immunocytochemical staining. The color of the nucleus is much darker than that of the cytoplasm. The nucleus is at the center of the cervical squamous epithelial cell, showing dark blue, and the cytoplasm shows light blue.
[0088] Figure 3 is an example diagram of an inflammatory cell after immunocytochemical staining. Figure 3There are two normal squamous epithelial cells shown, surrounded by inflammatory cells (cells that cause inflammation). The cytoplasm of the inflammatory cells has disappeared, leaving only the cell nuclei. In this embodiment, it is necessary to be able to distinguish Figure 3 the inflammatory cells shown in
[0089] from the cervical squamous epithelial cells and not count the inflammatory cells as cervical squamous epithelial cells.
[0090] Furthermore, preferably, in the step of annotating the cell nuclei of the cervical squamous epithelial cells in the pathological image, the following sub-steps may further be included:
[0091] Mark a point at the center position of the cell nucleus of the cervical squamous epithelial cells.
[0092] In this embodiment, preferably, a dark red point or a point of other colors may be marked at the center position of the cell nucleus of the cervical squamous epithelial cells.
[0093] The color of the marked point is to assist in observing whether it is marked within the cell nucleus area. The most important function of the marking is to record the coordinate position of the center of a cervical squamous epithelial cell in the pathological image for subsequent generation of a Gaussian density map (i.e., label data).
[0093] Thereafter, step 103 is entered. According to the pathological image, a single-channel matrix of size H×W is generated. In this single-channel matrix, the element at the position where the coordinates of the marked cervical squamous epithelial cells are located is 1, and the elements at the remaining positions are all 0.
[0094] In step 103, a single-channel matrix of the same size as the pathological image and with elements only 0 and 1 is generated.
[0095] Thereafter, step 104 is entered. Gaussian filtering is performed on each element 1 in the single-channel matrix to obtain the label data corresponding to the pathological image.
[0096] Thereafter, step 105 is entered. Based on the pathological image and the label data corresponding to the pathological image, a convolutional neural network (CNN) model is trained to obtain a single-channel density map generation model.
[0097] In this embodiment, preferably, in the step of training the convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the label data is pre-multiplied by the value 100.
[0098] During training, the label data is multiplied by the value 100, which makes the network training easier to converge, the oscillation more stable, and the evaluation function value (MAE) of the trained model on the validation set lower. That is, in actual use, the error between the model prediction value and the actual value is smaller.
[0099] Moreover, in the above step of training the convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the mean squared error (MSE) loss function can be selected as the loss function of the convolutional neural network model, and the mean absolute error (MAE) function can be selected as the evaluation function of the convolutional neural network model.
[0100] Thereafter, step 106 is entered, and the pathological image of the cervical squamous epithelial cells to be detected is input into the single-channel density map generation model to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cells to be detected.
[0101] In step 106, using the trained model, the pathological image to be detected is input to obtain a predicted single-channel density map.
[0102] Thereafter, step 107 is entered, and the values in the single-channel density map matrix are summed to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cells to be detected.
[0103] Thereafter, this process ends.
[0104] In order to better understand the technical solution of this specification, a preferred embodiment is described below. The details listed in this preferred embodiment are mainly for easy understanding and do not limit the protection scope of this application.
[0105] Figure 4 is the flowchart of this preferred embodiment.
[0106] Specifically, as Figure 4 shown, the technical solution of this preferred embodiment includes the following steps:
[0107] 1. Dataset preparation
[0108] 1.1 Select the pictures to be labeled, covering the application scenarios as much as possible. Select the pictures that actually appear, with more than 10 pictures of each type, and the more the better. As Figure 5 shown, there are 6 (2×3) pictures to be labeled, including three scenarios: scattered distribution (labeled with numbers 1, 5, 6 in the figure), dense (labeled with numbers 2, 3, 4 in the figure), and excessive inflammatory cells (labeled with number 3 in the figure).
[0109] 1.2 Use the open-source tool labelme. Here, the open-source tool is not necessarily specified, as long as it is an open-source tool that can record the coordinates of the annotation points on the picture. The labelme software interface is as Figure 6In the middle is the image to be annotated. When annotating, we should pay attention to the fact that a cervical squamous epithelial cell contains a nucleus and complete cytoplasm. The nucleus is located in the approximate center of the cell and is dark blue, while the cytoplasm is light blue. Figure 7 As shown in 1, 2, and 3 in the figure, the points in step 2 can be marked on the cell nucleus, which is dark red in the figure. Figure 8 As shown, the saved json file is as follows Figure 12 As shown, the two numbers in brackets in the points in each shape are the x and y coordinates of the point in the figure.
[0110] 1.3 Label preprocessing. This embodiment uses CNN-based density map estimation, that is, the input image is mapped to the density map. CNN needs to learn to generate the density map of cervical squamous epithelial cells from the input image, which can be regarded as a regression algorithm. However, this is not a regression of a specific value, but a fitting estimate. The density map is regressed and then the final total number of cells is obtained by integration and summation. The process of generating a density map from an annotation file is as follows: first, a matrix of the same size as the original image is constructed, and all its elements are set to 0. Then, the element at the position corresponding to each annotated squamous cell nucleus is set to 1. In this way, a matrix with only 0 and 1 elements is obtained. Finally, a continuous density map is obtained by convolution through a Gaussian kernel function. Formula 1 where x i For the location of each cervical squamous epithelial cell nucleus, the label of N cervical squamous epithelial cell nuclei can be expressed as H(x), where x represents a two-dimensional coordinate, that is, H(x) is the matrix with only 0 and 1 mentioned above. The Gaussian kernel function G(x) convolves the H(x) function to obtain the density function F(x), F(x) = H(x)*G σ (x), 2D Gaussian kernel function This formula takes two-dimensional coordinate points (x, y) as input, where the parameter is σ, which controls the radial range of the function. The reason for using the Gaussian kernel function is that it is a normal distribution bell curve. The closer the coordinates are to the center point, the larger the value; conversely, the smaller the value. That is, the closer to the center point, the greater the weight, and the farther from the center point, the smaller the weight. In the actual scenario, each cervical squamous epithelial cell nucleus has a certain size, corresponding to a small area in the picture. It is obviously unreasonable to set the value of a pixel point to "1" in the annotation file to represent this cell nucleus. Therefore, we use the Gaussian kernel function to replace the pixel value of the center point with the weighted average of the pixel values of its surrounding points, and the weights of the surrounding pixel points add up to 1. This not only does not affect the total number of cell nuclei in the generated density map but also can more realistically reflect the position characteristics of each cervical squamous epithelial cell nucleus in space. According to the radius size of the cell nuclei in the image scanned by the scanner, the σ parameter of this preferred embodiment is set to 5. Since the cell nucleus sizes in the pictures scanned by different scanners or with different parameters are inconsistent, the σ parameter needs to be adjusted for adaptation. In actual operation, the label preprocessing can be performed by writing a Python script. The matrix after being processed by the Gaussian filter, also known as the single-channel density map, is saved in the.npy format, which is convenient for directly reading using the Python package of numpy during training to accelerate the training. As Figure 10 is the visualization of the original input image and the corresponding density map. In the density map, blue represents 0, and each light green circle is the result of processing by a Gaussian kernel function. The middle is red, and the innermost is black. The light green -> red -> black represents an increase from 0.0 -> 1.0.
[0111] 1.4 Splitting the dataset. The training set actually used in this solution is 230 pictures, and the validation set is 30 pictures. Each picture has a width and height of 2448 * 2048.
[0112] 2 Model configuration
[0113] 2.1 Network structure construction: The Pytorch framework is selected. Convolutional layers with the simplest 1×1 and 3×3 convolutional kernels are used to extract features. The max pooling layer is used for downsampling to reduce the size of the model, improve the calculation speed, and at the same time enhance the robustness of the extracted features. Pooling generally follows convolution. The convolutional layer is used to extract features. Generally, the positions with corresponding features are relatively large numbers. Max pooling can extract the information of this part with corresponding features. The role of the upsampling layer is to expand the smaller feature maps and output pixel-level results with the same size as the original image. The reason is that when extracting high-level semantic information of the input image, the size of the network's feature maps generally shrinks first and then aggregates. The shrinkage is generally divided into three levels, and the number of channels in each convolutional layer also deepens. This process is also called the encoding process. The decoding process is to gradually expand the size, also in three levels. During this process, different-level feature information will be aggregated, then fused through the convolutional layer, and finally bilinearly upsampled. Dilated convolution can be used here. The characteristics of dilated convolution enable the network to expand the receptive field without increasing the number of parameters or the amount of calculation (for example, adding more convolutional layers can make the receptive field larger, but more operations will be introduced). In this way, larger-scale context information can be fused. As Figure 11 shown, on the far left, the dilation rate is 1, which is the conventional 3×3 convolution; the second from the left, the dilation rate is 2, and its receptive field is actually 5×5; the third from the left, the dilation rate is 3, and its receptive field is 7×7. Through experimental comparison, generally, the effect is better when the dilation rate is 2.
[0114] 2.2 Data reading: Still using the Pytorch framework, data augmentation techniques such as random cropping and random vertical and horizontal mirroring are adopted. The random cropping size is 1024×1024, plus random mirroring. Each time the model reads a different picture during training, and the model can learn more data, improving the robustness of the model.
[0115] 2.3 The loss function also selects the MSEloss of the Pytorch framework, that is, the L2 loss function. That is, the Euclidean distance is used to measure the difference between the real density map and the density map predicted by the model. The mean of the squared differences between the predicted value and the real value is calculated to measure the accuracy of the model, and the parameters of the model are adjusted through backpropagation. The loss function is where N is the size of the training batch (that is, how many pictures are learned in one batch), Z(X i ; θ) is the output generated by the model, θ is the trainable parameter, X i represents the input image, represents the real density map corresponding to the input image. Generally speaking, MAE (Mean Absolute Error) can represent the accuracy of the predicted density map, while MSE (Mean Square Error) can represent the robustness of the model.
[0116] 2.4 Optimizer, Adam optimizer converges quickly.
[0117] 2.5 Training parameters include 1e-4 learning rate, a lower learning rate in the early stage of training, and a stable convergence of the loss function. The training graphics card used in this preferred embodiment is RTX3090, 24g video memory, and the maximum trainable batch size is 8. Each round is 230 / 8=35 batches. Each batch will back propagate and update parameters according to the loss function. A total of 90 rounds are iterated. After each round, the average evaluation function value MAE is tested on the validation set. After training, the round with the lowest MAE is selected as the model weight for actual deployment.
[0118] 2.6 Evaluation Function Like MSE, this metric measures the size of the error without considering the direction. However, unlike MSE, MAE requires more complex tools like linear programming to calculate the gradient. In addition, MAE is more robust to outliers because it does not use squares. Figure 12 and 13 As shown, the original input map, the true density map, and the model prediction map are compared. The original input map is on the far right, the true density map is in the middle, the target_num above is the total number of cells, the model prediction map is on the far left, and predict is the sum of the density maps. From the numerical value and the cell distribution map, the model effect is good, the error is small, and the cell position is correct.
[0119] 3 Actual Deployment
[0120] 3.1 During deployment, the available computing resources are generally CPU and GPU. The CPU uses the OpenVINO framework, and the GPU uses the TensorRT framework. C++ programs are written because the training framework generally uses Python to write programs. After obtaining the best model weights on the validation set, you can leave the training framework and use the C++ language with better running efficiency. At this time, the image size of the inference is the same as that of the validation set test, which is the original image size with a resolution of 2448×2048.
[0121] 3.2 Calculation of cell number in the whole image First, we need to understand the scanning method of the digital pathology scanner, such as Figure 14 It is a snapshot of a sample slide taken by a scanner. The scanner will scan the square area in the picture row by row. Each row is divided into many columns. Finally, it will be a matrix. Each point on the matrix is a middle picture. The size of each middle picture is the size of our inferred picture, 2448×2048 in width and height. The whole square is called the large picture. For example, Figure 14A total of 14 (rows) × 12 (columns) middle-sized images were scanned. Also, since there are no cell areas in the four corners of the square, they need to be avoided. We fixedly selected five middle-sized images within the circular cell area for calculating the average number of cells per unit pixel. The approximate positions of the five middle-sized images are as Figure 15 shown. In the code, fixed values are used to obtain the (x, y) coordinates of the five middle-sized images because the name of each middle-sized image saved by the scanner contains row and column values. For example, for the topmost middle-sized image, x = 12 × 0.5 = 6.0, rounded down = 6, y = 14 × 0.3 = 4.2, rounded down = 4, and the final coordinates are (6, 4), with the upper left vertex of the square being (0, 0).
[0122] Overall image calculation scheme:
[0123] (1) Use the AI model to calculate the number of cells in each of the five middle-sized images at five points. Four points are spread out from the center point with a spacing of 0.2, ensuring that the points are definitely inside the circle;
[0124] (2) Take the average of the five tiles to get cell_mean;
[0125] (3) The number of cells per unit resolution area = cell_mean / (2048 × 2448);
[0126] (4) The resolution corresponding to the circular preparation on the scanned large image = 177mm 2 / (PixelSizeμm**2)*(10**6);
[0127] (5) Finally, multiply cell_mean by the resolution of the circular cell area to obtain the total number of cells on the entire glass slide.
[0128] Among them, 177mm 2 is the actual area size of the glass slide. PixelSize is the actual physical size of each pixel point during scanner scanning, with the unit of μm, provided by the scanner. In this preferred embodiment, it is 0.526μm, and the area is the square of this value. **2 represents square, and **6 represents the sixth power.
[0129] 3.3 Integrate it into the software application APP, called "Cell-QC". At the same time, when the value obtained from the calculation scheme is less than 5000, this value in the software interface will be marked in red as a special reminder that the preparation quality of this sample is unqualified. Such as Figure 17As shown, in actual use, for the cell section shown in the 3rd row, the result of algorithm counting is 2,975 cells, which is lower than the requirement of the TBS standard that each section should contain more than 5,000 cells. Therefore, red is displayed on the display screen. Using the Cell-QC algorithm for analysis, the calculation time for the number of cells in each sample is 4 seconds. The test machine configuration is a Lenovo brand machine Geek Pro, with a CPU of i5-12400F, a GPU of gtx1650s, and 16g of memory.
[0130] It should be noted that in this preferred embodiment:
[0131] 1. The training framework is not necessarily PyTorch. Commonly used frameworks in the market such as paddlepaddle, TensorFlow, Keras, etc. can all be used;
[0132] 2. The training parameters are not necessarily as in the text. It is just a reference value and can be actually adjusted as long as a model weight that performs well on the validation set can be trained;
[0133] 3. The annotation software is not necessarily labelme. It can be any annotation software or can be developed by oneself;
[0134] 4. The structure of the convolutional neural network is diverse and is built from basic modules. Just note that the receptive field of the model, the ability to extract spatial features and semantic information can meet the effect of the model on the validation set;
[0135] 5. The deployment framework is not necessarily TensorRT. Frameworks such as openvino, onnx-runtime, etc. can all be used;
[0136] 6. For pictures from scanners of other manufacturers, if the differences are not significant, the trained model can still be applicable. Otherwise, pictures need to be collected again for model training.
[0137] In summary, in this preferred embodiment, based on the CNN learning the non-linear mapping between the input picture and the output density map, it can learn the characteristics of complete cervical squamous epithelial cells, and at the same time can not only exclude the most similar inflammatory cells, but also exclude endocervical cells, metaplastic cells, histiocytes, etc.; when the error between the model prediction and the true value meets the actual use, the network model to be built by this method has the advantages of small computational amount, low complexity, and short operation time. This enables the technical problems and manual bottlenecks commonly encountered in the mainstream detection method (cytological detection) in cervical cancer screening to be solved by scientific, fast and stable methods, and enables the key quality control standards of the industry standard (TBS system) in this field to be implemented and implemented.
[0138] In addition, this embodiment also shows the technical advancement in the following aspects:
[0139] 1. When we label cervical squamous epithelial cells, it is strictly required that only when they have complete cytoplasm, and at the same time, CNN is used to learn such cell patterns, so CNN will not misjudge inflammatory cells as squamous epithelial cells; while traditional image processing methods cannot recognize inflammatory cells (cells that cause inflammation). They filter out the cell nuclei through gray-scale thresholds, and here inflammatory cells will be included. Post-processing cannot change this either, because inflammatory cells only have one cell nucleus, which is the same as that of ordinary cervical squamous epithelial cells. After passing through the gray-scale threshold, they all become 255 pixel values.
[0140] 2. Compared with the image segmentation and object detection schemes based on CNN, it is time-saving and labor-saving to label when making the dataset in this embodiment; the constructed network structure is simple. A simple model network structure can lead to a small amount of calculation during model inference and a fast operation result; the post-processing is very simple, and only the density map matrix output by the model needs to be directly summed.
[0141] 3. Compared with manually counting cells under a microscope, the computer vision algorithm has the advantages of fast speed, strong repeatability, small error, and can also calculate the total number of cells on the entire glass slide.
[0142] Next, the manual counting experiment will be compared with the CNN-based cervical squamous epithelial cell counting experiment (CellQC cell counting experiment) of this embodiment
[0143] Requirements for the manual counting experiment:
[0144] 1. Select 10 slides, including samples of high, medium, and low types of manually estimated quantities.
[0145] 2. Eight researchers observe these 10 slides in three times. For each slide, randomly select 10 fields of view under a 20-fold microscope for observation, counting, and recording respectively. Each experiment must be carried out under normal conditions, and the three times cannot be consecutive to avoid experimental errors caused by fatigue. A pressing counter can be used to assist in counting, and the total time consumed for each observation also needs to be recorded.
[0146] Requirements for the cell counting experiment of this embodiment of the application:
[0147] Take 10 slides and scan them in three times under a scanner. Each time, calculate the total number of cells by the technical solution of this embodiment of the application, and record the time at the same time.
[0148] The experimental results are as follows:
[0149] Table 1 The average value of 10 fields of view of a single sample is taken by 8 researchers each time, and the average value is taken again three times; and the average value, standard deviation, and coefficient of variation of each researcher in a single sample
[0150]
[0151] The standard deviation (SD) is the most commonly used quantitative form to reflect the degree of dispersion of a set of data and is an important indicator of precision.
[0152] Coefficient of Variation: When it is necessary to compare the degree of dispersion of two sets of data, if the measurement scales of the two sets of data differ greatly, or the data dimensions are different, it is not appropriate to directly use the standard deviation for comparison. At this time, the influence of the measurement scale and dimension should be eliminated, and the coefficient of variation can achieve this. It is the ratio of the standard deviation of the original data to the average of the original data. CV has no dimension, so objective comparison can be made.
[0153] Figure 18 It is a comparison graph of the average of three experiments in 10 visual fields of 10 samples for each researcher. Among them, the X-axis represents the slice numbers 1 - 10; the Y-axis represents the number of cells in 10 visual fields under the microscope, which is the average of the three cell counting experiments of each researcher.
[0154] As Figure 18 shown, the differences among the 8 researchers on the 8th sample are the largest, the differences on the 5th, 6th, 7th, 9th, and 10th samples are smaller, and the differences on the 1st, 2nd, and 3rd samples are moderate.
[0155] Table 2 Coefficient of variation of each researcher on the average number of cells per sample each time
[0156]
[0157] Figure 19 It is a comparison graph of the coefficient of variation of each researcher for different samples.
[0158] From Figure 19 it can be seen that the coefficients of variation of each researcher for each sample vary widely, some are very high, some are very low. The value of researcher 6 is the highest, full of instability, and the repeatability is very poor.
[0159] Table 3 Average time of three experiments of each researcher and three experiments of CellQC
[0160]
[0161] Figure 20 It is a comparison graph of the average time of three experiments of each researcher and three experiments of CellQC. Among them, the Y-axis represents time, unit: (seconds).
[0162] From Figure 20It can be seen that the time consumption of CellQC is extremely low, with an average time of only 40 seconds, and the efficiency is extremely high, more than 1000 times higher than that of manual counting.
[0163] Table 4 Records of the total number of cells in three experiments on 10 samples by CellQC, as well as the average, standard deviation, and coefficient of variation
[0164] Sample 1 Sample 2 Sample 3 Sample 4 Sample 5 Sample 6 Sample 7 Sample 8 Sample 9 Sample 10 First time 21080 42050 28196 36859 28182 55863 16785 32325 12059 12776 Second time 21081 42852 28530 37148 28312 56157 16791 32724 12161 12858 Third time 20857 43446 28279 37012 28354 55898 16505 32523 12201 12796 SD 129.04 700.58 173.90 144.58 89.67 160.59 163.42 199.50 73.22 42.76 Average 21006.00 42782.67 28335.00 37006.33 28282.67 55972.67 16693.67 32524.00 12140.33 12810.00 Coefficient of variation 0.61% 1.64% 0.61% 0.39% 0.32% 0.29% 0.98% 0.61% 0.60% 0.33%
[0165] Combined with Table 4 and Table 2, draw Figure 21 . Figure 21 is a comparison chart of the coefficient of variation of 8 researchers and CellQC on different samples. Among them, the Y-axis represents time, unit: (seconds).
[0166] Through Figure 21 , it can be seen that the stability of the black column-shaped CellQC is very high and the repeatability is strong. In Table 4, the glass slides are scanned three times, and the results will vary due to factors of the scanner itself (such as vibration, etc.), but the overall coefficient of variation is about 1%. It can be seen that the technical solution of this embodiment combined with the scanner is more accurate than manual counting. If the scanned sample is run through CellQC three times, then the result will be the same each time, and the coefficient of variation is close to 0, which can be said to be very accurate.
[0167] The second embodiment of the present invention relates to a device for calculating the number of cervical squamous epithelial cells based on CNN. Figure 22 is a schematic structural diagram of the device for calculating the number of cervical squamous epithelial cells based on CNN.
[0168] Specifically, as Figure 22 shown, the device for calculating the number of cervical squamous epithelial cells based on CNN includes:
[0169] An image acquisition module, configured to acquire a pathological image of cervical squamous epithelial cells to be trained, where the height of the pathological image is H and the width is W;
[0170] A labeling module, configured to label the cervical squamous epithelial cells in the pathological image and record the coordinates of the labeled cervical squamous epithelial cells on the pathological image;
[0171] A single-channel matrix generation module, configured to generate a single-channel matrix with a size of H×W according to the pathological image, where the element at the position of the coordinates of the labeled cervical squamous epithelial cells in the single-channel matrix is 1, and the elements at the remaining positions are all 0;
[0172] A Gaussian filtering processing module, configured to perform Gaussian filtering processing on each element 1 in the single-channel matrix, so as to obtain label data corresponding to the pathological image;
[0173] A training module, configured to train a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, so as to obtain a single-channel density map generation model;
[0174] A detection module, configured to input a pathological image of a cervical squamous epithelial cell to be detected into the single-channel density map generation model, so as to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cell to be detected;
[0175] A calculation module, configured to sum the values in the single-channel density map matrix, so as to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cell to be detected.
[0176] This embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in cooperation with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the first embodiment.
[0177] It should be noted that each module mentioned in the device embodiments of the present invention is a logical module. Physically, a logical module can be a physical module, a part of a physical module, or a combination of multiple physical modules. The physical implementation manner of these logical modules themselves is not the most important. The combination of the functions implemented by these logical modules is the key to solving the technical problems proposed by the present invention. In addition, to highlight the innovative part of the present invention, the above device embodiments of the present invention do not introduce modules that are not closely related to solving the technical problems proposed by the present invention, which does not mean that there are no other modules in the above device embodiments.
[0178] It should be noted that those skilled in the art should understand that the implementation functions of the modules shown in the above device embodiments can be understood with reference to the relevant descriptions of the corresponding methods mentioned above. The functions of the modules shown in the above device embodiments can be implemented by a program (executable instruction) running on a processor or by specific logic circuits. When the above device embodiments of the present specification are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present specification, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the methods described in the embodiments of the present specification. In this way, the embodiments of the present specification are not limited to any specific combination of hardware and software.
[0179] It should be noted that in the claims and the specification of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0180] Although the present invention has been illustrated and described with reference to certain preferred embodiments thereof, those of ordinary skill in the art should understand that various changes may be made thereto in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A method for calculating the number of cervical squamous epithelial cells based on CNN, characterized in that, it includes the following steps: Obtain the pathological image of cervical squamous epithelial cells to be trained, the height of the pathological image is H, and the width is W; Label the cervical squamous epithelial cells in the pathological image, and record the coordinates of the labeled cervical squamous epithelial cells on the pathological image; According to the pathological image, generate a single-channel matrix with a size of H×W. In this single-channel matrix, the elements at the positions where the coordinates of the labeled cervical squamous epithelial cells are located are 1, and the elements at the remaining positions are all 0; Perform Gaussian filtering on each element 1 in the single-channel matrix to obtain the label data corresponding to the pathological image; Train a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image to obtain a single-channel density map generation model; Input the pathological image of the cervical squamous epithelial cells to be detected into the single-channel density map generation model to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cells to be detected; Sum the values in the single-channel density map matrix to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cells to be detected.
2. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, Before the step of obtaining the pathological image of the cervical squamous epithelial cells to be trained, the following steps are further included: Perform immunocytochemical staining on the cervical squamous epithelial cell sample using an antibody reagent for a protein marker target; Scan the cell section after immunocytochemical staining through a digital pathology scanner to obtain the pathological digital microscopic image of the cervical squamous epithelial cells.
3. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, In the step of labeling the nuclei of the cervical squamous epithelial cells in the pathological image, only the cervical squamous epithelial cells containing nuclei and complete cytoplasm are labeled.
4. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 3, characterized in that, In the step of labeling the cervical squamous epithelial cells in the pathological image and recording the coordinates of the labeled cervical squamous epithelial cells on the pathological image, the following sub-steps are included: Label the nuclei of the cervical squamous epithelial cells in the pathological image, and record the coordinates of the nuclei of the labeled cervical squamous epithelial cells on the pathological image.
5. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 4, characterized in that, The nuclei of the cervical squamous epithelial cells after immunocytochemical staining are located at the center of the cells, showing dark blue, and the cytoplasm shows light blue; In the step of labeling the nuclei of the cervical squamous epithelial cells in the pathological image, the following sub-steps are included: Mark a point at the center position of the nucleus of the cervical squamous epithelial cell.
6. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, the height of the pathological image is 2048 and the width is 2448.
7. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, in the step of training the convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the label data is pre-multiplied by the value 100.
8. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, in the step of training the convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the mean squared error loss function is selected as the loss function of the convolutional neural network model.
9. The method for calculating the number of cervical squamous epithelial cells based on CNN according to claim 1, characterized in that, in the step of training the convolutional neural network model based on the pathological image and the label data corresponding to the pathological image, the mean absolute error function is selected as the evaluation function of the convolutional neural network model.
10. A device for calculating the number of cervical squamous epithelial cells based on CNN, characterized in that, comprising: an image acquisition module for acquiring a pathological image of cervical squamous epithelial cells to be trained, the height of the pathological image being H and the width being W; a labeling module for labeling the cervical squamous epithelial cells in the pathological image and recording the coordinates of the labeled cervical squamous epithelial cells on the pathological image; a single-channel matrix generation module for generating a single-channel matrix of size H×W according to the pathological image, where the elements at the positions of the coordinates of the labeled cervical squamous epithelial cells in the single-channel matrix are 1 and the elements at the remaining positions are 0; a Gaussian filtering processing module for performing Gaussian filtering processing on each element 1 in the single-channel matrix to obtain label data corresponding to the pathological image; a training module for training a convolutional neural network model based on the pathological image and the label data corresponding to the pathological image to obtain a single-channel density map generation model; a detection module for inputting a pathological image of cervical squamous epithelial cells to be detected into the single-channel density map generation model to obtain a single-channel density map corresponding to the pathological image of the cervical squamous epithelial cells to be detected; a calculation module for summing the values in the single-channel density map matrix to obtain the number of cervical squamous epithelial cells in the pathological image of the cervical squamous epithelial cells to be detected.