A method for detecting urine sediment cells based on convolutional neural network
By using a convolutional neural network in urine sediment cell detection, combining CSPDarknet-53 and NPANet networks for multi-scale feature fusion, and using AIoULoss loss function, the problem of low detection accuracy of small and medium-sized cells in urine sediment is solved, achieving efficient and low-cost cell detection effect.
Patent Information
- Application Number
- CN202211138511.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The prior art has low detection accuracy in urine sediment, which is prone to missed detection. The strategies or means introduced increase the detection cost and the accuracy is not satisfactory.
The urine sediment cell detection method based on convolutional neural network is adopted, and multi-scale feature fusion is performed through the CSPDarknet-53 backbone network and the NPANet feature fusion network, and the AIoULoss border regression loss function is used to optimize the design of the detection head to improve the detection accuracy.
It effectively improves the accuracy of cell detection in urine sediment, reduces the missed detection rate, and avoids excessive detection costs, achieving more efficient detection results.
Smart Images

Figure CN115578580B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning medical image analysis and processing, and specifically relates to a urine sediment cell detection method based on convolutional neural network. Background Art
[0002] In computer vision, object detection is a basic task. With the continuous development of deep learning technology, many high-performance detectors have been born. These detectors are widely used in face recognition, traffic flow detection, autonomous driving and medical image analysis. The current detectors can achieve good detection results for conventional objects, but the detection accuracy for small objects is relatively low. Especially in the application of medical urine sediment images, since urine sediment cells are generally small, it is easy to cause missed detection.
[0003] At present, the strategies for detecting small targets include: copy-paste data enhancement methods, or generating high-resolution images through GAN, or using a better multi-scale fusion method. Usually, additional means are needed to improve the detection accuracy of small objects, such as using anchor-free methods to avoid the imbalance of positive and negative samples, or using context extraction information to process the correlation between the target and surrounding information, or introducing attention mechanisms to enhance the representation ability of features. However, after introducing various strategies or means, the detection accuracy is often not satisfactory, and the detection cost is also increased. Summary of the invention
[0004] The purpose of the present invention is to provide a urine sediment cell detection method based on convolutional neural network, which can effectively improve the accuracy of cell detection in urine sediment.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for detecting urine sediment cells based on a convolutional neural network, the method comprising:
[0007] The labeled urine sediment cell images are taken as sample images, and the sample images are preprocessed with data enhancement to form a training data set;
[0008] Input the sample images in the training data set into the backbone network CSPDarknet-53, and obtain the feature maps F of three different scales output by the dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53. 1 、F 2 、F 3 ;
[0009] The feature map F 1 、F 2 、F3 As a feature map They are respectively input into the feature fusion network NPANet for feature fusion processing to obtain the detection head P 1 , P 2 , P 3 ;
[0010] Take the detection head P 1 , P 2 After convolution of the classification branch and the regression branch, they are connected along the channel part, and then the connected feature map is stretched into two dimensions to obtain the stretched feature map F 11 、F 21 , the stretched feature map F 11 、F 21 Connect them to get the final feature map F, calculate the loss based on the feature map F and perform back propagation to update the gradient, and update the network parameters at the same time to complete a training;
[0011] If the training end condition is not met, continue to train using the training data set, otherwise save the latest weight file and end the training;
[0012] Load the pre-trained and saved weight file, and use the trained network to output the detection results for the urine sediment cell image to be detected.
[0013] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.
[0014] Preferably, the data enhancement preprocessing includes Mosaic data enhancement and MixUp data enhancement.
[0015] Preferably, the sample images in the training data set are first adjusted to a size of 640×640 and then input into the backbone network CSPDarknet-53.
[0016] Preferably, the sample images in the training data set are input into the backbone network CSPDarknet-53 based on the batch principle.
[0017] Preferably, the feature map F 1 、F 2 、F 3 As a feature map They are respectively input into the feature fusion network NPANet for feature fusion processing to obtain the detection head P 1 , P 2 , P 3 ,include:
[0018] The feature map F 1 、F 2 、F 3 As a feature map Will Directly input into the feature fusion network NPANet, first from top to bottom, after 1×1 convolution, upsampling and then with the feature map Perform concat splicing to get the feature map Continue to map the feature map After 1×1 convolution, upsampling and feature map Perform concat splicing to get the feature map The feature map As Direct output to get the detection head P 1 ; Then do bottom-up and cross-scale fusion, pass the bottom-level location information back to the shallow layer, After 3×3 convolution and the previous feature map Fusion splicing output Get the detection head P 2 ;Will After 3×3 convolution and the previous feature map Fusion and splicing to get the detection head P 3 .
[0019] Preferably, calculating the loss according to the feature map F includes calculating the classification loss, the target score loss and the border regression loss, the classification loss and the target score loss are BCELoss loss functions, the border regression loss is AIoULoss loss function, and the formula of the AIoULoss loss function is as follows:
[0020]
[0021] In the formula, IoU is the intersection-over-union ratio between the real box and the predicted box, A c A is the area of the minimum bounding rectangle of the real box and the predicted box and the difference between the real box and the predicted box. i is the area of the minimum bounding rectangle of the real box and the predicted box, w 1 is the length of the real frame, h 1 is the width of the real frame, w 2 is the length of the prediction box, h 2 is the width of the prediction box.
[0022] Preferably, when the trained network is used to output the detection result for the urine sediment cell image to be detected, the SimOTA positive and negative sample allocation strategy is used to screen the prediction box.
[0023] The urine sediment cell detection method based on convolutional neural network provided by the present invention improves the multi-scale fusion method in the existing YOLOX technical solution. Considering that the detection head with a large receptive field will introduce noise interference, which is not conducive to small target detection, it is not used for classification and regression tasks. At the same time, AIoULoss is used instead of IoULoss for the border regression loss, which can adaptively adjust the overlapping area and aspect ratio, and can effectively improve the accuracy of cell detection in urine sediment. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flow chart of the urine sediment cell detection method based on convolutional neural network of the present invention;
[0025] Figure 2 It is a schematic diagram of the structure of the cell detection network of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0028] In order to overcome the problem of low accuracy in detecting small objects in the prior art, this embodiment proposes a urine sediment cell detection method based on a convolutional neural network. The method of this embodiment mainly includes the following steps: first, the labeled urine sediment image is preprocessed with data enhancement, and then training is started in batches. For each batch, the image is extracted with the backbone network CSPDarknet-53 to obtain a feature map F 1 、F 2 、F 3 , and then the obtained feature map is used as Input, feature fusion is performed through NPANet to obtain the prediction head P 1 , P 2 , P 3 , and only the prediction head P 1 , P 2 Classification and regression are performed to obtain the predicted value, and then the loss is calculated with the true value of the image. After each batch of training, backpropagation is performed to update the gradient and the network parameters at the same time, and finally the training of the entire network is completed.
[0029] In one embodiment, Figure 1 As shown, a urine sediment cell detection method based on convolutional neural network is proposed, which includes the following steps:
[0030] Step S1: taking an annotated urine sediment cell image as a sample image, performing data enhancement preprocessing on the sample image to form a training data set.
[0031] In this embodiment, a urine sediment cell image with a urine sediment cell detection frame marked thereon is first obtained as a sample image, and then data enhancement is performed on the sample image to expand it. The data enhancement in this embodiment includes Mosaic data enhancement and MixUp data enhancement.
[0032] This embodiment performs Mosaic data enhancement and MixUp data enhancement on the sample images. Mosaic data enhancement, that is, taking out 4 sample images and splicing them by random cropping, random scaling, and random arrangement, has the advantage of enriching the background and small targets of the detected objects, and calculating the data of 4 images at one time during calculation, without a large overhead, and a GPU can achieve a relatively good effect. For MixUp data enhancement, that is, superimposing 2 sample images together, this can reduce the memory of wrong labels to enhance robustness.
[0033] This embodiment uses a training data set with high sample richness to train the cell detection network, such as Figure 2 As shown, the cell detection network includes a backbone network CSPDarknet-53, a feature fusion network NPANet, and a classification regression layer Head connected in sequence. The specific training process is shown in steps S2-S5.
[0034] Step S2: Input the sample images in the training data set into the backbone network CSPDarknet-53, and obtain the feature maps F of three different scales output by the dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53. 1 、F 2 、F 3 .
[0035] This embodiment uses CSPDarknet-53 as the backbone network for feature extraction. Compared with traditional networks such as Resnet-50, it ensures accuracy while maintaining lightweight. Pre-load the weights trained on MS COCO to facilitate faster and better convergence during training. Training is based on the batch principle. The batch size during training is 16 (that is, 16 pictures are processed in each batch, and the batch size can be adjusted). A total of 50 epochs are trained, including warm-up for the first 5 epochs, and data enhancement is turned off for the last 15 epochs. And stochastic gradient descent (SGD) is used for training, using a learning rate of lr×BatchSize / 64, an initial lr of 0.01, and cosine scheduling. The weight decay is 0.0005 and the SGD momentum is 0.9.
[0036] Since the original image size is 1920×1080, this embodiment scales the original image to 640×640 according to the proportion of the long side, and inputs the scaled image into the backbone network CSPDarknet-53. After extracting features through a series of convolution operations, the dark3 unit, dark4 unit, and dark5 unit are output, and three sizes of feature maps F of 256×80×80, 512×40×40, and 1024×20×20 are output successively. 1 、F 2 、F 3 The size of the feature map is determined by the backbone network CSPDarknet-53 and will not be described here.
[0037] Step S3: The feature map F 1 、F 2 、F 3 As a feature map They are respectively input into the feature fusion network NPANet for feature fusion processing to obtain the detection head P 1 , P 2 , P 3 .
[0038] In this embodiment, the feature map F 3 As Directly input into the feature fusion network NPANet, first from top to bottom, after 1×1 convolution, upsampling and then with the feature map F 2 As input Perform feature fusion and splicing to obtain feature maps Continue to map the feature map After 1×1 convolution, upsampling and feature map F 3 As input Perform feature fusion and splicing to obtain feature maps The feature map Directly output the feature map As the detection head P 1 Then do bottom-up fusion to pass the shallow position information back to the deep layer. After 3×3 convolution, it becomes The same size, so that they can be fused to obtain the feature map Output to get the detection head P 2 Similarly, After 3×3 convolution, it becomes The same size, so that they can be fused to obtain feature maps Output to get the detection head P 3 , but do not place the detection head P 3 Used for subsequent classification and regression operations.
[0039] Specifically, the feature map of 1024×20×20 Directly input into the top-down feature pyramid network NPANet, first through 1 × 1 convolution to make the number of channels become 512, and then after upsampling to 40 × 40 feature map and feature map (512×40×40) features are fused and spliced along the channel, and then the number of channels is changed to 512 through the CSP module to obtain (512×40×40). (512×40×40) is convolved with 1×1 to make the number of channels 256, and then upsampled to 80×80 feature map and then combined with feature map Perform feature fusion and splicing, and obtain it through the CSP module (256×80×80). We directly The output is the detection head P1. Then the bottom-up fusion is performed. The number of channels and size are directly converted to The same number of channels and length and width are fused and spliced and then obtained through the CSP module Will Output to get the detection head P 2 At this point, we have obtained two detection heads P 1 With P 2 , only using P 1 With P 2 To complete the subsequent classification and regression tasks.
[0040] It should be noted that the role of the CSP module mentioned above is to enhance the learning ability of CNN, deepen the network while maintaining lightness and accuracy, and on the other hand, reduce the computational bottleneck. NPANet feature fusion is based on PANet bidirectional fusion, removing the detection head P with a larger receptive field. 3, considering that the cells in the urine sediment dataset are generally small, P 3 The detection head will introduce noise interference, and removing the detection head P3 can reduce the missed detection rate.
[0041] Step S4: Take the detection head P 1 , P 2 After convolution of the classification branch and the regression branch, they are connected along the channel part, and then the connected feature map is stretched into two dimensions to obtain the stretched feature map F 11 、F 21 , the stretched feature map F 11 、F 21 The connections are made to obtain the final feature map F, the loss is calculated based on the feature map F and the gradient is updated by back propagation to complete a training.
[0042] The detection head P in this embodiment 1 , P 2 The decoupled head method is used to separate the classification branch and the regression branch. First, convolution is used in each detection head to convert the channel to 256, and then the channel part is connected, and the feature map obtained by the connection is stretched into two dimensions (along W×H) to obtain the stretched feature map F 11 、F 21 , and then the stretched feature map F 11 、F 21 Connect them to get the final feature map F, calculate the loss of each part and perform back propagation to update the gradient to complete the network training.
[0043] In this embodiment, after the convolution of the detection head classification branch and the regression branch, the channel part is connected, and the new feature maps generated are two tensors of size {W×H×[(cls+reg+obj)]×N}, where W×H is the feature map size, cls is the category classification, reg is the bounding box regression, including the predicted upper left corner point (x 1 ,y 1 ) and the lower right corner (x 2 ,y 2 ), obj is the target score prediction, N is the number of predicted anchor boxes, and in this embodiment, 1 is taken. Then W is multiplied by H, and the spatial dimension is stretched into two dimensions along W×H to obtain the feature map F 11 、F 21 Then put F 11 、F 21 The connection is performed to obtain the final feature map F. Finally, the classification loss, target score loss and border regression loss are calculated, and back propagation is performed to reduce the loss. At the same time, the network parameters are updated to make the network reach final convergence.
[0044] Specifically, F 11 、F21 After convolution of the classification branch and regression branch (classifier and regressor), each feature map generates 3 new feature maps F cls ∈{N×W×H×cls}, F reg ∈{N×W×H×4}, F obj ∈{N×W×H×1}, first connect along the channel part, and the new feature maps are two {N×W×H×[(cls+reg+obj)]} size tensors, W, H∈{40,80}. Then multiply W and H, stretch the spatial dimension into two dimensions to obtain two {N×(cls+reg+obj)×(W×H)} size tensors. Then multiply F along W*H 11 、F 21 The connection is performed to obtain the final feature map F∈{N×(cls+reg+obj)×8000}.
[0045] The prediction head of this embodiment adopts the decoupled head method. Considering that the classification task and the regression task focus on different areas, the classification task and the regression task are separated for convolution operations, which can achieve better detection results. The number of prediction boxes at each position is reduced from 3 to 1, and the anchor-free method is adopted to avoid the problem of imbalance between positive and negative samples.
[0046] Since the output feature value cannot be directly used for loss calculation, regression is required to obtain the actual prediction value. According to the following formula, classification loss, target score loss and bounding box regression loss are performed on the feature map F. The classification loss and target score loss are BCELoss loss functions, and the bounding box regression loss uses the newly designed AIoULoss loss function. The specific formula is as follows:
[0047] BCELoss=-(ylog(p(x))+(1-y)log(1-p(x)))
[0048]
[0049] It should be noted that the grid setting in this embodiment is on the final feature map, which is an abstract concept. The purpose is to facilitate the bounding box regression calculation. For the 40*40 and 80*80 feature maps, there are 40×40 and 80×80 grids respectively. Dividing the feature map into multiple grids is a relatively mature technology in the field and will not be repeated here.
[0050] 1) Calculate the classification loss and target score loss using the Binary CrossEntropy Loss function:
[0051] BCELoss=-(y log(p(x))+(1-y)log(1-p(x)))
[0052] Where y indicates whether it is a target, the value is 1 or 0, and p(x) is the predicted target score.
[0053] 2) Calculate the border regression loss, which is essentially to compare the predicted box with the real box. The AIoULoss loss function of this embodiment is improved on the basis of the IoULoss loss function. IoU (Intersection of Union) is the intersection of the predicted box and the real box. The formula of IoULoss is as follows:
[0054]
[0055] Where S1 is the ground-truth box, S2 is the predicted box, I(S1, S2) is the area of the intersection of the ground-truth box and the predicted box, and U(S1, S2) is the area of the ground-truth box and the predicted box. The lower the IoULoss value, the more accurate the prediction.
[0056] The existing IoULoss loss function has a drawback, that is, when the real box and the predicted box do not intersect, the relative position of the two cannot be measured. Therefore, this embodiment proposes an AIoULoss loss function, which considers the problem in a segmented form. First, when the two do not intersect, by finding the minimum bounding rectangle of the real box and the predicted box, Ac represents the area of the difference between the minimum bounding rectangle and the real box and the predicted box, A i represents the area of the minimum enclosing rectangle. This overcomes the defect of not being able to measure relative positions. Secondly, when the real box and the predicted box intersect, we consider the aspect ratio factor. The (w 1 ,h 1 ), (w 2 ,h 2 ) represent the length and width of the true box and the predicted box respectively. Considering the aspect ratio can achieve better regression effect and make the predicted box regression closer to the true box.
[0057] Thus, the loss between the predicted value and the true value is obtained. Before each batch ends, back propagation is performed to reduce the loss and update the network parameters. In this embodiment, the backbone network CSPDarknet-53, feature fusion network NPANet, classification branch and regression branch in steps S2, S3 and S4 are referred to as a cell detection network as a whole. The network parameters of the cell detection network are continuously updated during training to obtain a high-precision detection result in actual detection.
[0058] Step S5: If the training end condition is not met, continue training using the training data set; otherwise, save the latest weight file and end the training.
[0059] This embodiment starts the next batch of training after updating the network parameters until all batches of training data are trained, and finally obtains the trained weights, and all updated parameters are saved in the Outputs weight file.
[0060] Step S6: In the actual detection task, the pre-trained and saved weight file is loaded, and the trained network is used to output the detection result for the urine sediment cell image to be detected.
[0061] In this embodiment, the image to be detected is also scaled to 640×640 size and input into the network. The feature maps of Dark3 unit, Dark4 unit, and Dark5 unit are output through the CSPDarknet-53 backbone network. 1 、F 2 、F 3 As a feature map They are respectively input into the feature fusion network NPANet for feature fusion processing to obtain the detection head P 1 , P 2 , P 3 ; For the detection head P 1 , P 2 After classification and regression, the predicted value is obtained, including the category cls, the target score obj and the bounding box regression reg. The three are combined to draw the corresponding prediction box to obtain the final prediction result.
[0062] This embodiment also uses the SimOTA positive and negative sample allocation strategy to screen the prediction box in the actual detection task. First, the prediction box is preliminarily screened, and only those prediction boxes whose center points are within the groundtruth and within a square with a side length of 5 are retained. After the initial screening is completed, the border loss of the prediction box and the groundtruth is calculated, and the classification loss is calculated using the binary cross entropy, and the cost matrix is calculated:
[0063]
[0064] Represents the cost relationship between each true box and each feature point. The first k fixed prediction boxes with the smallest loss to the groundtruth are used as positive samples, and the rest are used as negative samples, thus avoiding additional hyperparameters.
[0065] The urine sediment cell detection method based on convolutional neural network in this embodiment improves the original YOLOX technical solution to obtain a new multi-scale fusion method (NPANet), and designs a better border regression loss function AIoULoss, which effectively improves the accuracy of cell detection in urine sediment.
[0066] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0067] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A method for detecting urine sediment cells based on convolutional neural network, characterized in that: The urine sediment cell detection method based on convolutional neural network comprises: The labeled urine sediment cell images are taken as sample images, and the sample images are preprocessed with data enhancement to form a training data set; Input the sample images in the training data set into the backbone network CSPDarknet-53 to obtain the feature maps of three different scales output by the dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53. ; The feature map As a feature map They are input into the feature fusion network NPANet for feature fusion processing to obtain the detection head ; Take the detection head After convolution of the classification branch and the regression branch, they are connected along the channel part, and then the connected feature map is stretched into two dimensions to obtain the stretched feature map. , the stretched feature map Connect to get the final feature map , according to the feature map Calculate the loss and perform back propagation to update the gradient, while updating the network parameters to complete a training session; If the training end condition is not met, continue to train using the training data set, otherwise save the latest weight file and end the training; Load the pre-trained and saved weight file, and use the trained network to output the detection results for the urine sediment cell image to be detected; Wherein, the feature map As a feature map They are input into the feature fusion network NPANet for feature fusion processing to obtain the detection head ,include: The feature map As a feature map ,Will Directly input into the feature fusion network NPANet, first from top to bottom, after 1×1 convolution, upsampling and then with the feature map Perform concat splicing to get the feature map ; Continue to map the feature map After 1×1 convolution, upsampling and feature map Perform concat splicing to get the feature map ; The feature map As Direct output to get the detection head ; Then do bottom-up and cross-scale fusion, pass the bottom-level location information back to the shallow layer, After 3×3 convolution and the previous feature map Fusion splicing output Get the detection head ;Will After 3×3 convolution and the previous feature map Fusion and splicing to get the detection head ; Among them, according to the feature map Calculating the loss includes calculating the classification loss, target score loss and bounding box regression loss. The classification loss and target score loss are Loss function, the bounding box regression loss is The loss function is The formula of the loss function is as follows: ; In the formula, is the intersection-over-union ratio of the real box and the predicted box, is the area of the minimum enclosing rectangle of the real box and the predicted box and the difference between the real box and the predicted box, is the area of the minimum enclosing rectangle of the real box and the predicted box, is the length of the real frame, is the width of the real frame, is the length of the prediction box, is the width of the prediction box.
2. The method for detecting urine sediment cells based on convolutional neural network according to claim 1, characterized in that: The data enhancement preprocessing includes Mosaic data enhancement and MixUp data enhancement.
3. The method for detecting urine sediment cells based on convolutional neural network according to claim 1, characterized in that: The sample images in the training data set are first resized to 640×640 and then input into the backbone network CSPDarknet-53.
4. The method for detecting urine sediment cells based on convolutional neural network according to claim 1, characterized in that: The sample images in the training data set are input into the backbone network CSPDarknet-53 based on the batch principle.
5. The method for detecting urine sediment cells based on convolutional neural network according to claim 1, characterized in that: When the trained network is used to output the detection result for the urine sediment cell image to be detected, the Positive and negative sample allocation strategies are used to filter prediction boxes.