A cell segmentation method based on the Oriented Cascade Mask RCNN network

Through the Oriented Cascade Mask RCNN network, combined with Oriented and Cascade strategies, the Mask RCNN network is optimized, which solves the problems of low accuracy and missed detection in cell segmentation, and achieves efficient and accurate cell segmentation effect, which is suitable for medical image analysis.

CN116563534BActive Publication Date: 2025-07-08ROBOTICS RESEARCH CENTER OF YUYAO CITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310380175.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-07-08
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

The existing Mask RCNN network has problems of low accuracy and missed detection in cell segmentation, especially in densely distributed cell areas, which are difficult to effectively classify and annotate.

Method used

The Oriented Cascade Mask RCNN network is adopted, combined with Oriented and Cascade strategies, and the model is optimized to improve segmentation accuracy and detection rate through steps such as feature extraction, ROIAlign, cascade classification, regression and segmentation branching, and non-maximum suppression algorithm.

Benefits of technology

It realizes efficient and accurate cell segmentation, can automatically complete the analysis and detection of medical images, improves segmentation accuracy and detection rate, reduces noise interference, and is suitable for a variety of terminal systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563534B_ABST
    Figure CN116563534B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image segmentation, and discloses a cell segmentation method based on the Oriented Cascade Mask RCNN network. Based on the Mask RCNN framework, combining Cascade RCNN and Oriented RCNN, the Oriented Cascade Mask RCNN is proposed to achieve precise segmentation and classification of cells in medical images. Among them, Cascade RCNN can achieve a more accurate segmentation effect through cascaded classification and regression branches, while Oriented RCNN introduces Oriented Anchors into the RCNN model, reducing the invalid regions in the bounding boxes, thereby reducing the impact of useless features on the segmentation results. Combining these two parts achieves a more accurate detection effect with a higher detection rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a cell segmentation method based on an Oriented Cascade Mask RCNN network. Background Art

[0002] Classifying and labeling cells with relatively dense distribution in a medical image is a very cumbersome and repetitive task, and it is very time-consuming and laborious to complete the relevant tasks manually. With the progress of computer technology, it has become increasingly common to rely on algorithms and machine vision to complete this task. In the early stage, this technology mainly relied on some image processing algorithms, such as watershed, threshold segmentation algorithms, etc. Although these algorithms are relatively fast in speed, their accuracy is poor, and there are disadvantages of over-segmentation or under-segmentation. With the development of deep learning technology and the progress of computer computing power, cell segmentation based on machine learning models has achieved remarkable results. It can not only achieve better segmentation effects, have higher detection rates and segmentation accuracies, but also have higher efficiency.

[0003] In terms of cell segmentation, the existing relatively good effect is the Mask RCNN network. The Mask RCNN network is a two-stage segmentation network, which can simultaneously predict the category to which the cell belongs, but only using Mask RCNN will have problems of low accuracy and missed detection. Summary of the Invention

[0004] The purpose of the present invention is to provide a cell segmentation method based on an Oriented Cascade Mask RCNN network to solve the above technical problems.

[0005] To solve the above technical problems, the specific technical solution of a cell segmentation method based on an Oriented Cascade Mask RCNN network of the present invention is as follows:

[0006] A cell segmentation method based on an Oriented Cascade Mask RCNN network includes the following steps:

[0007] Step 1: Data processing: including image normalization, image denoising, image segmentation, and image enhancement;

[0008] Step 2: Feature extraction: Use the Resnet-50 network to extract features from the segmented image, and use the ImageNet dataset to pre-train this part;

[0009] Step 3: RPN Region of Interest Proposal: Candidate boxes are output according to the output features of the Resnet-50 network, and then they are labeled as positive or negative samples based on the IOU between these candidate boxes and the ground truth anchor boxes. The candidate boxes are then passed to the detection head to complete the optimization;

[0010] Step 4: ROIAlign Finds the Feature Vectors Corresponding to Each ROI Region from the Original Feature Map: After the RPN outputs candidate boxes, the ROIAlign module is used to find the corresponding feature vectors from the feature map;

[0011] Step 5: The Cascade-Type Classification, Regression, and Segmentation Branches Predict Results Respectively: The cascade-type classification regression and segmentation branches process the feature vectors corresponding to each region of interest output by the ROIAlign module. Using a cascaded design concept, a higher IOU matching threshold is used in each stage, and three embedded convolutional structures are used to model the classification, regression, and segmentation branches respectively;

[0012] Step 6: Use NMS to Filter the Predicted Inclined Boxes: After using the three branches to output the classification and regression results, the NMS algorithm is used to filter the results. After the NMS algorithm, the predicted bounding box with the highest confidence for each target is obtained, so that there is exactly one bounding box to identify and locate each cell.

[0013] Further, in Step 2, an inclined bounding box is used to locate the region where the cell is located. The inclined bounding box is the smallest external rectangle corresponding to the cell, and the upper left and lower right coordinates of the rectangle are obtained using the cv2.minAreaRect function.

[0014] Further, in Step 3, the cascaded detection heads provided by the Cascade strategy are used to screen by gradually increasing the difficulty of the task, so as to gradually reduce the impact of noise on network training. The Cascade strategy divides and filters positive and negative samples according to the level of IOU. Let A, B, and C represent the three detection heads respectively, and the IOU values used are IOU A , IOU B , IOU C respectively. Then there is a relationship of IOU A < IOU B < IOU C . They are set to 0.5, 0.6, and 0.7 respectively, and the samples in each stage are divided into positive and negative samples according to this IOU value.

[0015] Furthermore, the method uses an overlapping sliding window method to segment the image during both training and prediction, and uniformly uses a resolution of 640*640 for prediction. During training, the image is randomly scaled, with a scale between 0.8 and 1.2 and a probability of 0.2.

[0016] Furthermore, step 5 includes a two-stage instance segmentation model. In the first stage, the RPN module is used to generate possible ROI regions, and these ROI regions need to be labeled as positive or negative examples according to the IOU with the ground truth mask. Then, the ROIAlign module is used to transform these ROI regions onto the feature map, so as to extract the corresponding regional features from the feature map. In the classification branch, Cross Entropy is used as the loss function, and its calculation formula is

[0017]

[0018] where, when the predicted category is the same as the ground truth category, y i is 1, otherwise, it is 0.

[0019] In the regression branch, SmoothL1 is used as the loss function, and its calculation formula is:

[0020]

[0021] In the segmentation branch, the Cross Entropy loss function is also used to calculate the difference between the predicted category of each pixel and the predicted category of the ground truth pixel. Here, the category refers to whether each pixel belongs to a cell, that is, the segmentation task is equivalent to a binary classification task. Based on the above classification and regression, the bounding box of the region of interest where cells may exist and the category of cells that may exist in it are obtained. Then, the feature vector in this region is taken out, and through the segmentation branch, the specific region where the cells are located is segmented from this region.

[0022] Furthermore, in the second stage of step 5, the regression branch is used to obtain the final predicted bounding boxes, and these candidate boxes need to be filtered, that is, processed by the non-maximum suppression algorithm to filter out some candidate boxes with a large intersection over union (IOU) between each other. The specific process is as follows:

[0023] S1: Sort the set B of candidate bounding boxes in descending order according to the confidence level;

[0024] S2: Select the first candidate box from the set B, put it into the final set D of bounding boxes and delete it from the set B;

[0025] S3: Traverse each candidate box in set B, calculate the IOU value between them and this candidate box in set D. If the IOU value is greater than the threshold N, then delete it from set B;

[0026] S4: Repeat steps S2 - 3 until set B is empty.

[0027] Furthermore, in the training stage, the labelme software is used to label the cells in the image, and finally the segmentation mask data of each cell is obtained. Each cell in each image is uniquely identified by an integer. Then, opencv is used to obtain the bounding box corresponding to each cell. After these data are filtered and split, each small image is cleaned, and finally a cell segmentation dataset is obtained.

[0028] Furthermore, the method uses a three - fold cross - validation method. The dataset is divided into three parts. Each time, two of them are taken as the training set, and the other one is taken as the validation set. They are trained and validated in turn, and the average value of the test metrics on the validation set is taken as the final test result.

[0029] A cell segmentation method based on the Oriented Cascade Mask RCNN network of the present invention has the following advantages:

[0030] (1) The present invention has high timeliness. The cell segmentation method based on the convolutional neural network has high efficiency.

[0031] (2) The present invention has strong anti - interference performance. In the data processing stage, the data is filtered and pre - processed. In the training process, many data augmentation methods are also used to reduce the risk of model overfitting. When the model is used to predict the test image, it can also ensure high transferability.

[0032] (3) The present invention has relatively high performance indicators. Since the Oriented and Cascade strategies are used to optimize the model, and Mask RCNN belongs to a two - stage instance segmentation network with relatively high accuracy itself, this enables the model to perform better segmentation prediction on the image.

[0033] (4) The present invention has a relatively high detection rate. For the same reason above, the Oriented Cascade Mask RCNN network can analyze medical images well and can detect and identify cells with dense distribution to the greatest extent, which greatly enhances the processing effect of the model.

[0034] (4) The present invention can automatically complete the analysis and detection of medical images, with high convenience and easy to use. Relying on the portability of python, users can install and run the entire network on terminal systems such as laptops, desktops, and servers. Brief Description of the Drawings

[0035] Figure 1 It is a flowchart of model training and detection for the present invention;

[0036] Figure 2 It is a schematic diagram of the Cascade principle;

[0037] Figure 3 It is a schematic diagram of the Long Edge Definition (90) encoding method used by Oriented;

[0038] Figure 4 It is a structural diagram of the Oriented Cascade Mask RCNN model;

[0039] Figure 5 It is an example diagram of the immunohistochemical segmentation dataset used in the experiment;

[0040] Figure 6 It is a Qupath detection result diagram;

[0041] Figure 7 It is a detection result diagram of Oriented Cascade Mask RCNN;

[0042] Figure 8 It is a schematic diagram of how to segment large-resolution images and merge detection results;

[0043] Figure 9 It is a detection sample and label diagram;

[0044] Figure 10 It is a detection result diagram of Mask RCNN;

[0045] Figure 11 It is a detection result diagram of Cascade Mask RCNN;

[0046] Figure 12 It is a detection result diagram of Oriented Mask RCNN

[0047] Figure 13 It is a detection result diagram of Oriented Cascade Mask RCNN. Detailed Description of the Preferred Embodiment

[0048] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a method for cell segmentation based on the Oriented Cascade Mask RCNN network of the present invention with reference to the accompanying drawings.

[0049] Based on the Mask RCNN network, the present invention combines the Oriented RCNN and Cascade RCNN networks. As Figure 4 shown, we first used the Cascade strategy and found that the accuracy of the model did increase. However, in some areas with dense distributions, many cells were missed. Through analysis, we found that this was because the IOU between horizontal bounding boxes would be relatively high. When it was higher than the IOU threshold of the non-maximum suppression algorithm, one of the two relatively close cells would be filtered out. Therefore, based on the Cascade strategy, we added the Oriented strategy. This can effectively avoid the above problems. Moreover, it is worth noting that since the Oriented strategy can convert the bounding boxes into tilted boxes, the method for calculating the IOU also needs to be converted to that for tilted boxes, which also provides a better calculation method for the Cascade strategy, producing an effect where one plus one is greater than two.

[0050] As Figure 1 shown, a high-resolution medical image cell segmentation method based on the Oriented Cascade Mask RCNN network of the present invention includes the following steps:

[0051] Step 1: Data processing;

[0052] Data processing mainly includes image normalization, image denoising, image segmentation, and image enhancement, etc. Since the cell segmentation dataset is small, in order to enhance the training effect of the subsequent model, these processes are very important. Moreover, generally, the resolution of medical images is high. Imagine passing a medical image with a resolution of 1024*1024 to the subsequent network for processing. Due to the large scale of the model, high number of parameters, and high computational complexity, there will be a problem of out-of-memory. Even if multi-cards are used for training and prediction, the time consumed in this process will be very long.

[0053] Step 2: Feature extraction;

[0054] Feature extraction mainly uses the Resnet-50 network to extract features from the segmented images. We first used the ImageNet dataset to pre-train this part.

[0055] Step 3: RPN region of interest proposal;

[0056] The RPN region of interest proposal mainly outputs candidate bounding boxes according to the output features of the Resnet-50 network. Then we can label them as positive or negative samples according to the IOU between these candidate bounding boxes and the true anchor boxes. Since these candidate bounding boxes are still relatively rough and cannot complete the segmentation operation, we need to pass them to the detection head for optimization.

[0057] Step 4: ROIAlign finds the feature vectors corresponding to each ROI region from the original feature map;

[0058] After the RPN outputs the candidate bounding boxes, we need to use the ROIAlign module to find the corresponding feature vectors from the feature map. Different from the ROIPooling module, the ROIAlign module can avoid the accuracy loss problem caused by the errors in the quantization process during the calculation.

[0059] Step 5: The Cascade classification, regression, and segmentation branches predict the results respectively;

[0060] The Cascade classification, regression, and segmentation branches can process the feature vectors corresponding to each region of interest output by the ROIAlign module. To improve the accuracy of the model, we use a cascaded design idea, using higher and higher IOU matching thresholds at each stage, which can sequentially improve the quality of the training samples at each stage, and the output quality of the final model will be higher. We use three embedded convolutional structures to model the three branches of classification, regression, and segmentation respectively.

[0061] Step 6: Use NMS to filter the predicted oriented bounding boxes;

[0062] After using the three branches to output the classification and regression results, we need to use the NMS algorithm to filter the results. After the NMS algorithm, we can obtain the predicted bounding box with the highest confidence for each target, and there is as much as possible one and only one bounding box to identify and locate each cell.

[0063] After the above steps, we can obtain the segmentation mask of cells in a medical image.

[0064] Different from the previous cell segmentation methods, the advantage of Oriented Cascade Mask RCNN is that it can simultaneously ensure a high detection rate and accurate segmentation.

[0065] (1) Precision of segmentation: Different from object detection and segmentation in natural scenes, one of the difficulties in cell segmentation is that cells are densely distributed and their distribution has a certain directionality. It is noted that cells are generally oval-shaped. If we use a horizontal bounding box to enclose an inclined oval, the proportion of the background area in this bounding box will be relatively large. Then, after ROIAlign, the features extracted from each region will also contain a large number of irrelevant features, which will cause certain interference to cell segmentation. Correspondingly, if we use an inclined bounding box to locate the region where the cell is located, the proportion of the cell region will become much larger, and there will not be so much noise in the extracted features. This inclined bounding box is actually the minimum external rectangle corresponding to the cell, and the top-left and bottom-right coordinates of this rectangle can be obtained using the cv2.minAreaRect function. On this basis, the Cascade strategy provides us with cascaded detection heads, as shown in Figure 2 , which can obtain better results by gradually optimizing the detection results. This can screen by gradually increasing the difficulty of the task to gradually reduce the impact of noise on network training. Since the Cascade strategy divides and filters positive and negative samples according to the level of IOU, assuming the three heads are represented by A, B, and C respectively, and the IOUs used are IOU A , IOU B , IOU C respectively, then there is a relationship of IOU A < IOU B < IOU C . We set them to 0.5, 0.6, and 0.7 respectively. According to this IOU value, the positive and negative samples of each stage can be divided. It should be noted that since we have used an inclined bounding box instead of a horizontal bounding box, the IOU calculation method of the inclined bounding box is also used here, which can further improve the detection accuracy of the model. Through the above analysis, we can expect that Oriented Cascade Mask RCNN can achieve better results than Oriented Mask RCNN and Cascade Mask RCNN.

[0066] (2) High detection rate: As mentioned in the precision of segmentation, the Oriented strategy can reduce the area of the irrelevant part in the region of interest, as shown in Figure 3The encoding method used by Oriented, Long Edge Definition (90), is shown. In the cell segmentation task, since cells are densely distributed, when using ordinary horizontal bounding boxes, the Intersection over Union (IOU) between adjacent bounding boxes will be relatively large. If it is greater than the set IOU threshold, one of them will be wrongly filtered out. That is to say, through the non-maximum suppression process, the bounding box originally representing a certain cell is wrongly suppressed. By using the strategy of tilted bounding boxes, the IOU between the bounding boxes corresponding to two adjacent objects will become smaller, so the probability of one of them being wrongly eliminated becomes lower, which also improves the detection rate of the model. At the same time, as mentioned in (1), the detection accuracy of the model becomes higher, so the confidence of the results predicted by the model will also become higher. Then, when finally filtering the results according to the confidence (usually taking the confidence higher than 0.5), the predicted bounding boxes wrongly eliminated will also be fewer, which also improves the detection rate of the model.

[0067] The data processing part includes sample denoising, positive and negative sample matching, etc. Since the medical image cell segmentation dataset is generally small, in order to improve the performance of the subsequent model, we need to screen and augment the data; the backbone part of the model uses the Resnet-50 module, which is a deep neural network. The core idea is to solve problems such as gradient vanishing and gradient explosion through residual blocks, so that a deeper convolutional neural network can be effectively trained, and it has good effects on various tasks; the RPN module in the model is used as the detector in the first stage. First, we will use methods such as clustering to give some possible Regions of Interest (ROIs), and label them as positive or negative examples according to the IOU between them and the true anchor boxes. These samples are used to train the RPN network; the last part of the model uses three branches of classification, regression, and segmentation to complete the corresponding tasks respectively. This multi-task training mode can fuse information from multiple tasks, playing a role in supervising each sub-task and reducing overfitting.

[0068] The Oriented Cascade Mask RCNN model can achieve end-to-end training, can detect cells in medical images of any size, and has high accuracy and detection rate. Since medical images generally have a high resolution, such as 1024*1024 or even larger, for general devices, it is impossible to directly perform cell detection and segmentation on such large images. Therefore, we use an overlapping sliding window method to split the images during both training and prediction, and uniformly use a resolution of 640*640 for prediction. During training, to enhance the training effect, we randomly scale the images, with a ratio between 0.8 and 1.2 and a probability of 0.2. This can not only achieve data augmentation and reduce overfitting, but also perform predictions with a unified size, reducing the complexity of subsequent processing.

[0069] As a two-stage instance segmentation model, in the first stage, we use the RPN module to generate possible ROI regions, and these ROI regions need to be labeled as positive or negative examples according to the IOU with the ground truth mask. Then, we use the ROIAlign module to transform these ROI regions onto the feature map, so as to extract the corresponding regional features from the feature map. Compared with ROIPooling, it can well solve the problem of regional mismatch caused by two quantizations. Because the role of ROIPooling is to pool the corresponding regions in the feature map into a feature map of a fixed size according to the position coordinates of the candidate boxes for subsequent classification, regression, and segmentation. Since these operations are generally floating-point numbers and need to be converted into size values, some quantization operations are required, which will cause a certain deviation between the finally obtained candidate boxes and the positions regressed at the beginning. This deviation will affect the accuracy of detection or segmentation. Regarding this problem, ROIAlign gives a good solution idea. It cancels the quantization operation and uses bilinear interpolation to obtain the image tree value at the pixel points with floating-point coordinates, thus converting the entire feature aggregation process into a continuous operation. In the cell segmentation task, since cells are generally small in volume and occupy fewer pixel points, the error caused by this regional mismatch has a greater impact on detection, and the proposed ROIAlign can better solve this problem.

[0070] In the classification branch, we use Cross Entropy as the loss function, and its calculation formula is

[0071]

[0072] Among them, when the predicted category is the same as the ground truth category, y iIf it is 1, otherwise, it is 0. The cross-entropy loss function can accelerate the convergence of the model during training, accelerate the update of parameters, and avoid the problem of the learning rate decline of the mean squared error loss function. Because if the sigmoid loss function is adopted, there is a problem of gradient disappearance.

[0073] In the regression branch, we use SmoothL1 as the loss function, and its calculation formula is

[0074]

[0075] Compared with L1Loss and L2Loss, SmoothL1Loss is more robust to outliers. Especially in the cell segmentation task, since we need to crop large images, there will be more cell fragments at the edge part. Using SmoothL1Loss can make the loss smoother and contribute to the convergence of the model.

[0076] The calculation formula of L1Loss is

[0077]

[0078] The calculation formula of L2Loss is

[0079]

[0080] In the segmentation branch, we also use the Cross Entropy loss function to calculate the difference between the predicted category of each pixel point and the true pixel point. It should be noted that the category here refers to whether each pixel point belongs to a cell, that is, the segmentation task is equivalent to a binary classification task. Based on the above classification and regression, we obtain the bounding box of the region of interest where cells may exist and the categories of the cells that may exist in it. Then, we extract the feature vectors in this region. After passing through the segmentation branch, the specific region where the cells are located is segmented from this region. The segmentation branch finally provides us with the specific mask of the cells, with a more accurate segmentation result.

[0081] To improve the detection effect of the model, we use a multi-stage cascaded classification branch and regression branch. Mask RCNN is a two-stage segmentation method. In the first stage, we obtain some candidate bounding boxes and use the ROIAlign module to extract the feature vectors in the corresponding regions from the feature map. These feature vectors then go through the classification and regression branches for secondary optimization. Therefore, the effect of Mask RCNN is better than that of one-stage segmentation methods such as the YOLO model. In the second stage, we use the regression branch to obtain the final predicted bounding boxes. These candidate bounding boxes need to be filtered, that is, processed by the non-maximum suppression algorithm to filter out some candidate bounding boxes with a large intersection over union (IOU) between each other. The specific process is as follows:

[0082] (1) Sort the set B of candidate bounding boxes (each candidate bounding box has a confidence level, representing the possibility that there are cells in the candidate bounding box. The higher the confidence level, the more reliable the candidate bounding box) in descending order according to the confidence level

[0083] (2) Select the first candidate bounding box (with the highest confidence level) from the set B, put it into the final set D of bounding boxes (initially an empty set) and remove it from the set B

[0084] (3) Traverse each candidate bounding box in the set B and calculate their IOU values with this candidate bounding box in the set D. If the IOU value is greater than the threshold N, then remove it from the set B

[0085] (4) Repeat steps 2 to 3 until the set B is empty.

[0086] Mask RCNN belongs to a two-stage segmentation algorithm. In the first stage, we will label the candidate bounding boxes as positive or negative examples according to the IOU between the candidate bounding boxes and the true anchor boxes. This requires specifying a threshold in advance, and the determination of this threshold has a greater impact on the effect of the model. If the threshold is large, the number of positive samples will be small and the model is prone to overfitting; if the threshold is small, the accuracy of the positive samples will decrease, and at the same time the training cost of the model will also increase. To obtain better training results, Cascade RCNN provides an idea: adopt a cascaded method to gradually increase the IOU threshold and use training sets with better quality to train the cascaded branches so that the final output of the model can achieve a satisfactory effect.

[0087] To obtain the dataset for training the model, we need to use the labelme software to annotate the cells in the images. Eventually, we can get the segmentation mask data for each cell, and each cell in each image is uniquely identified by an integer. Then we can use opencv to obtain the bounding box corresponding to each cell. After these data are filtered (such as removing cells with small areas) and split, we then clean each small image (such as removing cells that only have a very small part due to splitting at the image edge). Eventually, we obtain the cell segmentation dataset. To accurately evaluate the performance of the model, we also need to use the three - fold cross - validation method. We divide the dataset into three parts, take two of them as the training set each time, and the other as the validation set, train and validate in turn, and take the average of the test metrics on the validation set as the final test result.

[0088] The experimental dataset used below is a private immunohistochemical segmentation dataset. Figure 5 The example figure is shown as follows.

[0089] Example 1:

[0090] Take the example of segmenting cells from a picture with a resolution of 1024 * 1024. As can be seen from the following figure, this picture contains a large number of cells. To mark a bounding box for each cell as uniquely as possible, it will take a lot of time and effort to complete this work manually. We also used an open - source software Qupath to analyze this picture, and the result is as Figure 6 shown. As can be seen from the result, the detection result of Qupath is not satisfactory. While the detection result using Oriented Cascade Mask RCNN is as Figure 7 shown. As can be seen from the following figure, our method can achieve better results.

[0091] To implement this work using Oriented Cascade Mask RCNN generally requires the following steps:

[0092] (1) Split the image. We use the overlapping sliding window method to complete this work. After splitting, a 1024 * 1024 image can be split into 64 images of 256 * 256. After these images are filtered and pre - processed, they can be processed using the Oriented Cascade Mask RCNN model.

[0093] (2) For each 256 * 256 small image after splitting, the model will perform the following processing in turn.

[0094] (3) Resnet-50 extracts features from each image; the RPN module gives possible regions of interest based on these features, that is, some candidate boxes and their categories; the ROIAlign module aligns each candidate box to the feature map extracted by Resnet-50, so that we obtain the feature vectors of all regions of interest in each image; these feature vectors are passed to the three branches of classification, regression and segmentation. Due to the design of Cascade, the classification and regression branches at each stage can obtain better and better prediction results in turn, and the IOU between these results and the real anchor boxes will become larger and larger; after obtaining all possible predicted bounding boxes and corresponding masks for each cell, we also need to use NMS to filter these results, so as to ensure that each cell can be detected as much as possible and can be uniquely identified with only one bounding box, which means reducing the risk of missed detection and over-detection.

[0095] (4) After obtaining the detection results of each 256*256 image, we need to merge these results. We use the following algorithm to complete this task.

[0096] (5) First, the possibility of duplicate detection results in the merging process exists at the intersection of several small images, such as Figure 8 Small Figure 1 We only need to consider areas A and B. We take out the small Figure 1 All predicted anchor boxes and small Figure 2 All the predicted anchor boxes in region A are collected, and their coordinates are calibrated to the original 1024*1024 large image, and then these sets are NMS filtered. For region B, we also need to do the above operation. We traverse each small image in order from top to bottom and from left to right, and perform NMS filtering on the predicted boxes in the intersection area of ​​each image and its upper right, lower left and lower right (if any) small images. After the above operations, we can get the segmentation mask and its bounding box of almost all cells in the 1024*1024 image.

[0097] Embodiment 2:

[0098] We have conducted relevant ablation experiments for the Oriented strategy and the Cascade strategy. The dataset we used is a self-labeled immunohistochemistry image dataset. It contains 60 images with a resolution of 1536*1536, with a total of 41,003 positive cells and 95,422 negative cells. Each image corresponds to a txt file, and the format of each line is the category and the coordinates of the upper left and lower right corners of its bounding box. It is worth noting that when segmenting, we need to recalibrate these coordinates to each small image.

[0099] Figure 9 For the detection samples and label maps, through experiments, we can obtain the comparison results shown in the following table. It can be seen that compared with Mask RCNN, Oriented Mask RCNN, and Cascade Mask RCNN, Oriented CascadeMask RCNN can achieve better results. Figures 10 - 13 These are the detection result maps of the four models respectively.

[0100]

[0101] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A cell segmentation method based on the Oriented Cascade Mask RCNN network, characterized in that, It includes the following steps: Step 1: Data processing: including image normalization, image denoising, image segmentation, and image enhancement; Step 2: Feature extraction: Use the Resnet-50 network to extract features from the segmented images, and use the ImageNet dataset to pre-train this part; Step 3: RPN region of interest proposal: Output candidate boxes according to the output features of the Resnet-50 network, and then label them as positive or negative samples according to the IOU between these candidate boxes and the real anchor boxes, and pass the candidate boxes to the detection head to complete the optimization; Step 4: ROIAlign finds the corresponding feature vectors for each ROI region from the original feature map: After the RPN outputs candidate boxes, use the ROIAlign module to find the corresponding feature vectors from the feature map; Step 5: The Cascade type classification, regression, and segmentation branches respectively predict the results: The Cascade type classification regression and segmentation branches process the feature vectors corresponding to each region of interest output by the ROIAlign module. Using the cascaded design idea, each stage uses an increasingly high IOU matching threshold, and three embedded convolutional structures are used to model the classification, regression, and segmentation branches respectively; Step 6: Use NMS to filter the predicted inclined boxes: After using the three branches to output the classification and regression results, use the NMS algorithm to filter the results. After the NMS algorithm, the predicted bounding box with the highest confidence for each target is obtained, so that there is exactly one bounding box to identify and locate each cell.

2. The cell segmentation method according to claim 1, wherein In step 2, a tilted bounding box is used to locate the area where the cell is located. The tilted bounding box is the smallest external rectangle corresponding to the cell, and the upper left corner coordinates and the lower right corner coordinates of the rectangle are obtained using the cv2.minAreaRect function.

3. The cell segmentation method according to claim 1, wherein The cascaded detection heads provided by the Cascade strategy are used in step 3 to gradually increase the difficulty of the task for screening and gradually reduce the impact of noise on network training. The Cascade strategy divides and filters positive and negative samples according to the level of IOU. Let A, B, and C represent the three detection heads respectively, and the IOUs used are respectively. Then there is relationship. They are set to 0.5, 0.6, and 0.7 respectively, and the samples in each stage are divided into positive and negative samples according to the IOU value.

4. The cell segmentation method according to claim 1, wherein The method uses an overlapping sliding window method to segment the image during both training and prediction, and uniformly uses a resolution of 640*640 for prediction. During training, the image is randomly scaled, with a ratio between 0.8 and 1.2 and a probability of 0.

2.

5. The cell segmentation method according to claim 1, wherein Step 5 includes a two-stage instance segmentation model. In the first stage, use the RPN module to generate possible ROI regions, and these ROI regions need to be labeled as positive or negative examples according to the IOU with the real mask; then, use the ROIAlign module to convert these ROI regions to the feature map, so as to extract the corresponding regional features from the feature map. In the classification branch, use Cross Entropy as the loss function, and its calculation formula is , Among them, when the predicted category is the same as the true category, it is 1, otherwise, it is 0. In the regression branch, use SmoothL1 as the loss function, and its calculation formula is: , In the segmentation branch, the Cross Entropy loss function is also used to calculate the difference between the predicted class of each pixel and the true class of the pixel. Here, the class refers to whether each pixel belongs to a cell, that is, the segmentation task is equivalent to a binary classification task. Based on the above classification and regression, the bounding box of the region of interest where cells may exist and the classes of the cells that may exist therein are obtained. Then, the feature vectors in this region are taken out. After passing through the segmentation branch, the specific region where the cells are located is segmented from this region.

6. The cell segmentation method according to claim 1, wherein In step 5 of the second stage, the regression branch is used to obtain the final predicted bounding boxes. These candidate boxes need to be filtered, that is, processed by the non-maximum suppression algorithm, and some candidate boxes with a large intersection over union (IOU) between each other are filtered out. The specific process is as follows: S1: Sort the set B of candidate bounding boxes in descending order according to the confidence level; S2: Select the first candidate box from the set B, put it into the final set D of bounding boxes and delete it from the set B; S3: Traverse each candidate box in the set B, calculate the IOU value between it and this candidate box in the set D. If the IOU value is greater than the threshold N, then delete it from the set B; S4: Repeat steps S2 to S3 until the set B is empty.

7. The cell segmentation method according to claim 1, characterized in that, In the training stage of the method, the labelme software is used to annotate the cells in the image, and finally the segmentation mask data of each cell is obtained. Each cell in each image is uniquely identified by an integer. Then, the opencv is used to obtain the bounding box corresponding to each cell. After these data are filtered and split, each small image is cleaned, and finally a cell segmentation dataset is obtained.

8. The cell segmentation method according to claim 1, characterized in that, The method uses the three-fold cross-validation method, divides the dataset into three parts, takes two of them as the training set each time, and the other as the validation set, trains and validates in turn, and takes the average value of the test metrics on the validation set as the final test result.