Deep Supervision Based Density Regression Cell Counting Method and System
By adopting a deep supervision-based density regression method in the cell counting method, using fast connection and deep supervised learning strategies of non-adjacent layers, the accuracy problem of traditional methods in low contrast and complex contexts is solved, and higher cell counting accuracy and model performance are achieved.
Patent Information
- Application Number
- CN202210044331.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing cell counting methods are difficult to achieve sufficient counting accuracy when dealing with low image contrast, complex tissue background, and large differences in cell size and shape, and the local minimum value optimization problem of traditional neural network models affects the overall performance.
Deep supervision-based density regression cell counting method is adopted to connect convolutional neural networks with multi-scale features through fast connections of non-adjacent layers, integrate multi-scale image features, and provide direct supervision for the intermediate layer through deep supervision learning strategies to enhance the training process.
It improves the accuracy of cell counting work, avoids the risk of intermediate layer optimization falling into local minimum values, and enhances the overall performance of model performance.
Smart Images

Figure CN114549410B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a density regression cell counting method and system based on deep supervision. Background Art
[0002] Among the existing cell counting methods, they are mainly divided into two types, namely direct detection methods and density regression methods; in medical and biological research, it is necessary to accurately calculate the number of cells in microscopic images. In the existing image processing methods, although the accuracy has made obvious progress, due to differences in various image acquisition technology levels, low image contrast, complex tissue backgrounds, large differences in cell size and shape, and mutual blocking between cells in two-dimensional microscopic images, etc., designing an efficient automatic method with sufficient counting accuracy is still a challenging task.
[0003] The inventor found that the network layers of traditional neural network models are hierarchical structures, and the output of each layer only depends on the output of its directly adjacent layer; this limits the neural network model to making a more realistic density map for cell counting work; in addition, the training of the original neural network is based on a single loss measured at the final output layer, and all its intermediate layers are only optimized based on the gradients backpropagated from this single loss. The reduction of the gradient may cause the optimization of the intermediate layers to fall into local minima and endanger the performance of the overall network. Summary of the Invention
[0004] In order to solve the above problems, the present invention proposes a density regression cell counting method and system based on deep supervision. The present invention uses skip connections of non-adjacent layers to connect a convolutional neural network with multi-scale features. In this cascaded network architecture, multi-scale image features extracted by all layers along the downsampling path can be integrated into the inputs of the layers along the upsampling path to further improve the model performance; in addition, through a deep supervision learning strategy, direct supervision is provided for the intermediate layers of the designed neural network to enhance their training, thereby improving the accuracy of cell counting work.
[0005] To achieve the above object, the present invention is implemented through the following technical solutions:
[0006] In a first aspect, the present invention provides a density regression cell counting method based on deep supervision, including:
[0007] Obtain a cell image to be processed;
[0008] According to the cell image and a preset density regression cell counting model, obtain the number of cells in the cell image;
[0009] Among them, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; in the convolutional neural network, there are multiple first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, multiple third network blocks that are cascaded with the multiple first network blocks and used to restore the resolution of the cell image, and a fourth network block for obtaining a density map based on the image with the restored resolution; the third network block is set to perform upsampling; there is also a deep supervision module that provides direct supervision for the intermediate layer.
[0010] Further, after the second network block extracts representative features, it generates a first feature map; based on the first feature map, multiple third network blocks sequentially generate multiple different feature maps.
[0011] Further, the deep supervision module uses multiple convolutional neural networks. During the generation of the feature map, it generates multiple corresponding density maps, estimates the difference between the density map and the annotated image, and performs deep supervision on the backbone network framework.
[0012] Further, the training of the density regression cell counting model includes:
[0013] Input the cell image into multiple first network blocks with upsampling, convolution, and activation functions to extract low-dimensional features and generate low-dimensional feature maps;
[0014] Input the low-dimensional feature map into a second network block with convolution and activation functions to generate a first feature map, and input the information of the first feature map into an auxiliary convolutional network;
[0015] Input the first feature map into three third network blocks with upsampling, convolution, and activation functions to respectively generate a second feature map and a third feature map, and obtain an image with the restored resolution through the three third network blocks;
[0016] Train the parameters of the loss function, refer to the difference between the generated density map and the true annotation map, as well as the difference between the estimated density map and the true annotation map in each network block of the auxiliary convolutional network, and train the entire backbone network;
[0017] Input the image with the restored resolution into a fourth network layer with convolution and activation functions to obtain the density map corresponding to the input cell image.
[0018] Further, the first network block includes convolution, activation, and pooling functions; each pooling layer in the first network block performs a downsampling operation on the input feature map by only outputting the maximum value of each downsampling region in the feature map. The convolutional layer is associated with a set of learnable kernels and is used to extract local features from the output of its previous layer. The activation layer in each block is used to increase the non-linear characteristics of the network, without affecting the receptive field of the convolutional layer, setting the negative responses of the previous layer to zero and keeping the positive responses unchanged.
[0019] Further, the second network block includes convolution and activation functions.
[0020] Further, the third network block includes upsampling, convolution, and activation functions; the fourth network block includes convolution and activation functions.
[0021] In a second aspect, the present invention also provides a density regression cell counting system based on deep supervision, including:
[0022] A data acquisition module, configured to: acquire a cell image to be processed;
[0023] A cell counting module, configured to: obtain the number of cells in the cell image according to the cell image and a preset density regression cell counting model;
[0024] Wherein, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; the convolutional neural network includes a plurality of first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, a plurality of third network blocks for restoring the resolution of the cell image after being cascaded with the plurality of first network blocks, and a fourth network block for obtaining a density map according to the image with restored resolution; the third network block is set for upsampling; and it also includes a deep supervision module for providing direct supervision for intermediate layers.
[0025] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the density regression cell counting method based on deep supervision described in the first aspect are implemented.
[0026] In a fourth aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the density regression cell counting method based on deep supervision described in the first aspect are implemented.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. The present invention uses shortcut connections between non - adjacent layers to connect a convolutional neural network for multi - scale features. In such a cascaded network architecture, multi - scale image features extracted by all layers along the down - sampling path can be integrated into the inputs of layers along the up - sampling path to further improve the model performance. In addition, through a deep supervision learning strategy, direct supervision is provided for the intermediate layers of the designed neural network to enhance their training, thereby improving the accuracy of cell counting work.
[0029] 2. The present invention performs multi - level processing on images through a neural network model in a multi - level cascading mode. This processing mode can fuse multi - scale features, improve the granularity of the extracted features, facilitate the regression of density maps, and promote the learning of intermediate layers in the down - sampling path.
[0030] 3. Based on a deep - learning counting model, the present invention uses a method of auxiliary training of a convolutional neural network to provide direct and in - depth supervision for the learning of its intermediate layers to improve the performance of cell counting.
[0031] 4. The present invention sets up a neural network model with multiple cascades, which helps to process cell images of any size and uses the fully convolutional layer in the neural network model to estimate the density map of the cell image. If the cell types of interest are labeled in the training images, this method can also be applied to images containing multiple types of cells. This deep - supervision learning framework can make the training process focus on the cell types of interest and regard other types of cells as the background without including them in the result count. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings forming a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions thereof of this embodiment are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0033] Figure 1 is a flowchart of Embodiment 1 of the present invention;
[0034] Figure 2 is an image of a microscopic cell dataset set when training the neural network in Embodiment 1 of the present invention;
[0035] Figure 3 is a schematic diagram of the network structure in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present invention will be further described below in conjunction with the drawings and embodiments.
[0037] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0038] Example 1:
[0039] A density regression cell counting method based on deep supervision, comprising:
[0040] Obtaining a cell image to be processed;
[0041] Obtaining the number of cells in the cell image according to the cell image and a preset density regression cell counting model;
[0042] Wherein, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; the convolutional neural network includes a plurality of first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, a plurality of third network blocks for restoring the resolution of the cell image after being cascaded with the plurality of first network blocks, and a fourth network block for obtaining a density map according to the image with the restored resolution; the third network block is set to perform upsampling; and it further includes a deep supervision module for providing direct supervision for the intermediate layer.
[0043] As Figure 1 shown, the cell counting method in this embodiment includes the following steps:
[0044] S1. Extracting low-dimensional features by processing the image to obtain a low-dimensional feature map, inputting it into subsequent network layers, and continuing to extract high-dimensional features to obtain a feature map;
[0045] S2. While continuing to refine the features, gradually restoring the resolution of the feature map;
[0046] S3. Setting up a convolutional neural network for auxiliary training;
[0047] S4. Training the loss functions of the backbone network and the auxiliary convolutional neural network, and performing backpropagation, while receiving the supervision information fed back by the auxiliary convolutional neural network;
[0048] S5. Inputting the image with the restored resolution into the last network block to obtain a density map Y, and calculating the number of cells in the image by processing the density map through a density regression function
[0049] Specifically, in step S1, the cell images are input into the neural network. In this embodiment, the first network blocks are set to 3, which are the first three blocks in the neural network. Each first network block includes a convolutional layer, a ReLU layer, and a pooling layer, and is used to extract low-dimensional feature maps. In this embodiment, the second network block is set to 1, which is the fourth block in the neural network. The second network block includes a convolutional layer and a ReLU layer, and is used to further extract highly representative features. Each pooling layer in the 3 first network blocks performs a downsampling operation on the input feature map by only outputting the maximum value of each downsampling region in the feature map. The convolutional layer is associated with a set of learnable kernels and is used to extract local features from the output of its previous layer. The ReLU layer in each block is used to increase the non-linearity of the network without affecting the receptive field of the convolutional layer by setting the negative responses of the previous layer to zero while keeping the positive responses unchanged. Each pooling performs a downsampling operation on the input feature map by only outputting the maximum value of each downsampling region in the feature map. Therefore, as the spatial size of the input feature map decreases, multi-scale information features are gradually extracted.
[0050] In step S2, in this embodiment, the third network blocks are set to 3, which are located in the 5th to 7th blocks in the neural network, and are used to gradually restore the resolution of the feature map while refining the extracted feature map. Each third network block includes an upsampling layer, a convolution, and a ReLU layer, and performs an upsampling operation upward to gradually restore the resolution of the final estimated density map. This network design allows the feature extraction to be integrated into the density regression process.
[0051] In step S3, 3 auxiliary convolutional neural networks are used to provide direct supervision for the intermediate layer of the main learning framework network. The low-dimensional feature map is input into a network block with convolutional and activation functions, and the information of the first feature map is input into the first auxiliary network block of the auxiliary convolutional network. The feature map is input into three auxiliary network blocks with upsampling, convolutional, and activation functions. A second feature map is generated in the first auxiliary network area The information is input into the second auxiliary network block of the auxiliary convolutional network, and a third feature map is generated in the second auxiliary network block The information is input into the third auxiliary network block of the auxiliary convolutional network. Each auxiliary neural network contains two network blocks composed of a convolutional layer and a ReLU layer, and respectively estimates the low-resolution density map according to each input feature map. The difference between the estimated density map and the relevant ground truth is used to support the training of the backbone network model.
[0052] In step S4, the density regression function Θ in it is defined as Θ = (θ1, θ2, θ3, θ4); where θ1 represents the trainable parameters in the first 4 blocks, θ2 represents the parameters in the 5th block, θ3 represents the parameters in the 6th block, and θ4 represents the parameters in the last two blocks. is defined as the low-resolution density map output by each auxiliary neural network.
[0053] and are jointly trained by minimizing the combined loss function, and the formula of the loss function is as follows:
[0054]
[0055] where L(Θ) represents the error between the density map and the corresponding ground-truth annotation density map; represents the error measured between the low-resolution density map estimated by the k-th auxiliary convolutional neural network and the corresponding low-resolution ground-truth annotation density map, represents the parameter vector in the k-th auxiliary convolutional neural network; the parameter α k ∈[0, 1] controls the supervision intensity of the k-th auxiliary neural network; the parameter λ controls the intensity of the l2 loss penalty to reduce overfitting. The l2 loss is defined as the sum of the squared differences, which is expressed as:
[0056]
[0057] Calculate the error between the estimated density map generated by the backbone neural network and the ground-truth annotation density map, and represent it as the loss function and L(Θ) are defined as:
[0058]
[0059] where, Y b represents the full-size ground-truth annotation density map of the training data X b with respect to the training image B; is the low-resolution ground-truth annotation density map generated from Y b
[0060] The loss L cmb is numerically minimized by the momentum stochastic gradient descent method based on the equation, and the formula is:
[0061]
[0062] where, represents the updated parameter of Θ k , t represents the number of iterations; β represents the momentum parameter that controls the contribution of the previous iteration result, and η represents the learning rate that determines the parameter update speed.
[0063] After that, the network is trained using the backpropagation algorithm, and the formula is:
[0064]
[0065] In step S5, for a given two-dimensional microscopic image X ∈ R M×N including N c cells, the density map corresponding to X can be expressed as Y ∈ R M×N , and each value in Y represents the number of cells of the corresponding pixel in X. Let be the feature map extracted from X, and the density regression function can be defined as the mapping function from X to Y:
[0066]
[0067] The number of cells in X can then be calculated as follows:
[0068]
[0069] The key part of the density regression-based method is to learn from and the corresponding Θ
[0070] In the backbone network framework, can be simplified to F(X; Θ), because it can directly learn from X, so the number of cells can also be calculated using the following formula:
[0071]
[0072] where, [F(X; Θ)] i , j represents the estimated density of the pixel (i, j) in X.
[0073] Example 2:
[0074] This embodiment provides a density regression cell counting system based on deep supervision, including:
[0075] A data acquisition module, configured to: acquire a cell image to be processed;
[0076] A cell counting module, configured to: obtain the number of cells in the cell image according to the cell image and a preset density regression cell counting model;
[0077] Among them, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; the convolutional neural network includes a plurality of first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, a plurality of third network blocks for restoring the resolution of the cell image after being cascaded with the plurality of first network blocks, and a fourth network block for obtaining a density map based on the image with the restored resolution; the third network block is set to perform upsampling; and a deep supervision module for providing direct supervision to the intermediate layer is also included.
[0078] The working method of the system is the same as that of the density regression cell counting method based on deep supervision in Embodiment 1, which will not be elaborated here.
[0079] Embodiment 3:
[0080] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the density regression cell counting method based on deep supervision described in Embodiment 1 are implemented.
[0081] Embodiment 4:
[0082] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the density regression cell counting method based on deep supervision described in Embodiment 1 are implemented.
[0083] The above are only the preferred embodiments of this embodiment and are not used to limit this embodiment. For those skilled in the art, various changes and modifications can be made to this embodiment. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this embodiment shall be included within the protection scope of this embodiment.
Claims
1. A density regression cell counting method based on deep supervision, characterized in that, Including: Obtain a cell image to be processed; According to the cell image and a preset density regression cell counting model, obtain the number of cells in the cell image; Wherein, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; the convolutional neural network includes a plurality of first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, a plurality of third network blocks that are cascaded with the plurality of first network blocks and are used to restore the resolution of the cell image, and a fourth network block for obtaining a density map according to the image with restored resolution; the third network block is set to perform upsampling; it also includes a deep supervision module that provides direct supervision for the intermediate layer; The first network block is set to 3, which are the first three blocks in the neural network. The first network block includes convolution, activation, and pooling functions; each pooling layer in the first network block performs a downsampling operation on the input feature map by only outputting the maximum value of each downsampling area in the feature map. The convolutional layer is associated with a set of learnable kernels and is used to extract local features from the output of its previous layer. The activation layer in each block is used to increase the non-linear characteristics of the network, does not affect the receptive field of the convolutional layer, sets the negative response of the previous layer to zero, and keeps the positive response unchanged. The second network block is set to 1, which is the fourth block in the neural network. The second network block includes convolution and activation functions. The third network block is set to 3 and is located in the 5th to 7th blocks of the neural network. The third network block includes upsampling, convolution, and activation functions; the fourth network block includes convolution and activation functions; Three auxiliary convolutional neural networks are used to provide direct supervision for the intermediate layer of the main learning framework network. The low-dimensional feature map is input into a network block with convolutional and activation functions, and the information of the first feature map is input into the first auxiliary network block of the auxiliary convolutional network; the feature map is input into three auxiliary network blocks with upsampling, convolutional, and activation functions, and a second feature map is generated in the first auxiliary network area, and the information is input into the second auxiliary network block of the auxiliary convolutional network, and a third feature map is generated in the second auxiliary network block, and the information is input into the third auxiliary network block of the auxiliary convolutional network. Each auxiliary neural network contains two network blocks composed of convolutional layers and ReLU layers, and the low-resolution density map is estimated according to each input feature map respectively. The difference between the estimated density map and the relevant ground truth is used to support the training of the backbone network model; where X is a given two-dimensional microscopic image. Define the density regression function in as ; where represents the trainable parameters in the first 4 blocks, represents the parameters in the 5th block, represents the parameters in the 6th block, represents the parameters in the last two blocks; Define as the low-resolution density map output by each auxiliary neural network; and are jointly trained by minimizing a combined loss function, the formula of which is as follows: wherein, represents the error between the density map and the corresponding true annotation density map; represents measuring the error between the low-resolution density map estimated by the -th auxiliary convolutional neural network and the corresponding low-resolution true annotation density map, represents the parameter vector in the -th auxiliary convolutional neural network; the parameter controls the supervision intensity of the -th auxiliary neural network; the parameter controls the intensity of the loss penalty to reduce overfitting; wherein the loss is defined as the sum of the squared differences, expressed as in the above formula: Calculate the error between the estimated density map generated by the backbone neural network and the ground truth annotation density map, and represent it as a loss function and is defined as: Among them, represents the training data Regarding the training images Full-size ground truth annotation density maps; Represented as the low-resolution ground truth annotation density maps generated by ; Simplified to .
2. The density regression cell counting method based on deep supervision according to claim 1, characterized in that, After the second network block extracts representative features, it generates a first feature map; based on the first feature map, a plurality of third network blocks sequentially generate a plurality of different feature maps.
3. A density regression cell counting system based on deep supervision, characterized in that, Including: A data acquisition module, configured to: obtain a cell image to be processed; A cell counting module, configured to: according to the cell image and a preset density regression cell counting model, obtain the number of cells in the cell image; Wherein, the density regression cell counting model is a convolutional neural network that connects multi-scale features in a non-adjacent layer manner; the convolutional neural network includes a plurality of first network blocks for processing and extracting low-dimensional features, a second network block for extracting highly representative features, a plurality of third network blocks that are cascaded with the plurality of first network blocks and are used to restore the resolution of the cell image, and a fourth network block for obtaining a density map according to the image with restored resolution; the third network block is set to perform upsampling; it also includes a deep supervision module that provides direct supervision for the intermediate layer; The first network blocks are set to 3, which are the first three blocks in the neural network. The first network blocks include convolution, activation, and pooling functions. Each pooling layer in the first network blocks performs downsampling on the input feature map by only outputting the maximum value of each downsampling region in the feature map. The convolutional layer is associated with a set of learnable kernels and is used to extract local features from the output of its previous layer. The activation layer in each block is used to increase the non-linear characteristics of the network, without affecting the receptive field of the convolutional layer, setting the negative responses of the previous layer to zero and keeping the positive responses unchanged. The second network block is set to 1, which is the fourth block in the neural network. The second network block includes convolution and activation functions. The third network block is set to 3, located at the 5th to 7th blocks in the neural network. The third network block includes upsampling, convolution, and activation functions. The fourth network block includes convolution and activation functions. Three auxiliary convolutional neural networks are used to provide direct supervision for the intermediate layer of the main learning framework network. The low-dimensional feature map is input into a network block with convolutional and activation functions, and the information of the first feature map is input into the first auxiliary network block of the auxiliary convolutional network; the feature map is input into three auxiliary network blocks with upsampling, convolutional, and activation functions, and a second feature map is generated in the first auxiliary network area and the information is input into the second auxiliary network block of the auxiliary convolutional network, and a third feature map is generated in the second auxiliary network block and the information is input into the third auxiliary network block of the auxiliary convolutional network. Each auxiliary neural network contains two network blocks composed of convolutional layers and ReLU layers, and the low-resolution density map is estimated according to each input feature map respectively. The difference between the estimated density map and the relevant ground truth is used to support the training of the backbone network model; where X is a given two-dimensional microscopic image. Define the density regression function in as ; where represents the trainable parameters in the first 4 blocks, represents the parameters in the 5th block, represents the parameters in the 6th block, represents the parameters in the last two blocks; define as the low-resolution density map output by each auxiliary neural network; and are jointly trained by minimizing a combined loss function, the formula of which is as follows: wherein, represents the error between the density map and the corresponding true annotation density map; represents measuring the error between the low-resolution density map estimated by the th auxiliary convolutional neural network and the corresponding low-resolution true annotation density map, represents the th parameter vector in the auxiliary convolutional neural network; the parameter controls the supervision intensity of the th auxiliary neural network; the parameter controls the intensity of the loss penalty to reduce overfitting; wherein the loss is defined as the sum of squared differences, expressed as in the above formula: Calculate the error between the estimated density map generated by the backbone neural network and the true annotation density map, and represent it as a loss function and is defined as: Among them, represents the training data for the training images of the full-size ground truth annotation density map; is represented by the generated low-resolution ground truth annotation density map; is simplified to .
4. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, the steps of the density regression cell counting method based on deep supervision as described in any one of claims 1-2 are implemented.
5. An electronic device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the density regression cell counting method based on deep supervision as described in any one of claims 1-2 are implemented.
Citation Information
Patent Citations
Dense crowd counting method and device
CN109241895A
Small convolutional nuclear cell counting method and system based on deep convolutional neural network
CN110659718A
Method for fusing panchromatic image and multispectral image based on dense connection network
CN111223044A