System, Method and Storage Medium for Image Target Counting
Through the combination of integrated regression network and multi-task learning module, the problems of high annotation cost and poor robustness in image target counting are solved, and faster and more accurate target counting is achieved to adapt to more scenarios.
Patent Information
- Application Number
- CN202211579538.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-09
AI Technical Summary
The prior art has problems in image target counting with high annotation cost, poor robustness, and insufficient ability to adapt to dense scenes. In particular, the global regression method has poor robustness in remote sensing image counting, and the detection method has poor performance in dense scenes.
The integrated regression network is used to combine multi-task learning modules, and through feature learning subnets, multiple regression subnets and prediction modules, sorting loss, negative correlation learning and counting loss functions are used to reduce the computational complexity and improve generalization ability.
It achieves faster counting speed, higher counting accuracy and wider scope of application, reduces time cost and calculation complexity, avoids huge annotation costs, and improves counting efficiency and robustness.
Smart Images

Figure CN115797712B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image information processing, and particularly relates to a system, a method and a storage medium for counting image targets. Background Art
[0002] Counting targets in images, such as pedestrians, vehicles, buildings, etc., is an important and challenging task in the field of computer vision.
[0003] Currently, the methods for solving the target counting problem in the prior art are mainly divided into detection-based methods, density map estimation-based methods, and global regression-based methods:
[0004] 1) Detection-based methods: As the early used counting methods, they estimate the number of targets by detecting target instances. However, these methods can only perform well in scenarios where the object positions are sparse and at a relatively large scale, and perform poorly in dense and crowded scenarios, such as dense houses in remote sensing images; meanwhile, since the annotation of bounding boxes is required to identify objects for counting, the annotation cost is the highest among the three types of methods;
[0005] 2) Density map estimation-based methods: As the most mainstream counting method for the current image target counting problem, the density map estimation-based methods require the application of point-level annotations to generate the true density map; this method maps the image into a density map, which reflects the density distribution of the predicted objects in the entire image, and then obtains the count by summing the density map;
[0006] But this also requires a huge annotation cost. At the same time, the annotated point labels are not used to evaluate the counting performance, which means that the point labels are redundant to some extent;
[0007] 3) Global regression-based methods: The commonly used loss function for training the CNN network is the mean square error loss. In the highly non-linear regression counting task, it is difficult to optimize the model using this loss, especially the robustness to dense scenarios is also poor. Summary of the Invention
[0008] In order to overcome one or more defects and deficiencies existing in the prior art, the first object of the present invention is to provide a system for counting image targets, the second object is to provide a method for counting image targets, and the third object is to provide a storage medium for counting image targets, which is used to count the targets in the image with high precision.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions.
[0010] A system for counting image targets includes an integrated regression network and a multi-task learning module that are connected to each other;
[0011] The integrated regression network includes a feature learning sub-network, a regression sub-network, and a prediction module that are connected in sequence;
[0012] The feature learning sub-network outputs the features of the input image according to the input image, and divides the features into multiple feature subsets;
[0013] There are multiple regression sub-networks, and the number of regression sub-networks matches the number of feature subsets; the regression sub-network is used to output a predicted value according to the feature subset;
[0014] The prediction module is used to perform an average operation on the predicted values of multiple regression sub-networks, so as to obtain the final count predicted value of the target in the image;
[0015] The multi-task learning module is used to set the loss function for training the integrated regression network; the multi-task learning module is respectively connected to the regression sub-network and the prediction module;
[0016] The loss function in the multi-task learning module consists of a ranking loss function, a negative correlation learning loss function, and a counting loss function.
[0017] Preferably, it further includes a training database;
[0018] The training database is connected to the feature learning sub-network and is used to input the stored images into the feature learning sub-network.
[0019] Preferably, the prediction module is used to perform an average operation on the predicted values of multiple regression sub-networks and use the average value as the count predicted value of the target in the image.
[0020] Preferably, the loss function L total in the multi-task learning module, is composed of the ranking loss function L Rank , the negative correlation learning loss function L NCL , and the counting loss function L Count after being respectively assigned weight parameters and then accumulated, as shown in the following formula:
[0021] L total = L Rank + λL NCL + μL Count
[0022] Among them, the weight parameter of L Rank is 1, λ is the weight parameter of L NCL , and μ is the weight parameter of L Count .
[0023] Furthermore, the calculation of the ranking loss function L Rank is as shown in the following formula:
[0024]
[0025] Among them, NDCG(x i ,) is the approximate normalized discounted cumulative gain calculated according to the sorting of images. N is the number of regression sub-networks. max and min respectively correspond to the images with the largest and smallest label values in a set of images input during training. i represents the i-th image, and n represents the n-th regression sub-network.
[0026] Furthermore, the negative correlation learning loss function L NCL has the following calculation formula:
[0027]
[0028] Among them, M represents the number of a set of images input to the integrated regression network during training. n and m both represent the n-th regression sub-network. F n (i) and F m (i) respectively represent the predicted values of the n-th and m-th regression sub-networks on the i-th image. (i) represents the average value of the predicted values of the i-th image on all regression sub-networks.
[0029] Even further, the counting loss function L Count has the following calculation formula:
[0030]
[0031] Among them, d(i) represents the true count value of the i-th image.
[0032] A method for image target counting, including any one of the foregoing image target counting systems, comprising the following steps:
[0033] S1. According to the type of image for which target counting is to be performed, use the dataset of corresponding same-type images in the training database as the training set for image enhancement, and adjust the size of the image to match the input channel size of the feature learning sub-network in the integrated regression network;
[0034] S2. Input the enhanced training set images into the feature learning sub-network for feature extraction to form a respective feature set for each image;
[0035] S3. Divide the feature set into multiple feature subsets, and then input these multiple feature subsets into the corresponding multiple regression sub-networks for processing to obtain corresponding predicted values;
[0036] S4. Use the multi-task learning module to perform multi-task learning training on the outputs of multiple regression sub-networks based on the loss function until the entire integrated regression network is trained;
[0037] S5. Input the image for which target counting is to be performed into the trained integrated regression network. After passing through the feature learning sub-network and the regression sub-network in sequence, the final counting prediction value of the target in the image is obtained in the prediction module, and this final counting prediction value is used as the counting result of the target in the input image.
[0038] Preferably, the number of feature subsets and regression sub-networks is eight respectively.
[0039] A storage medium is used to store the computer program of the system for image target counting described in any one of the foregoing. It is characterized in that when the computer program is executed, the system for image target counting described in any one of the foregoing is correspondingly implemented.
[0040] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0041] The system of the present invention uses the method of multi-task learning, introduces a new loss function to form end-to-end training, and uses normalized discounted cumulative gain for regression counting, avoiding the problem of poor robustness of the existing global regression-based method for image, especially remote sensing image counting; in addition, by introducing negative correlation learning loss into the integrated regression network, good generalization ability is achieved, and the overfitting problem in the global regression-based method is alleviated; different from traditional ensemble learning, the system of the present invention uses the method of training a single model, reducing the time cost and computational complexity, and improving the counting performance of the entire system; compared with the method based on density map estimation, the system of the present invention avoids huge annotation costs and improves the counting efficiency; compared with the detection-based method, the system of the present invention can better adapt to the scenario with a higher density of counting targets and has a wider adaptation range; the method of the present invention realizes faster counting speed, higher counting accuracy and larger applicable range by using the system of the present invention for target counting; the storage medium of the present invention provides the implementation basis for the system of the present invention. Description of the Drawings
[0042] Figure 1 It is the structural framework diagram of one kind of image target counting system of the present invention;
[0043] Figure 2 is Figure 1 The schematic structural diagram of the feature learning sub-network in the integrated regression network;
[0044] Figure 3 is Figure 1 The schematic structural diagram of the regression sub-network in the integrated regression network;
[0045] Figure 4 It is the schematic general flow diagram of one kind of image target counting method of the present invention. Detailed Embodiments
[0046] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and their embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] Embodiment 1
[0048] As Figure 1 shown, a system for image target counting in this embodiment includes a training database, an integrated regression network, and a multi-task learning module that are connected in sequence.
[0049] The integrated regression network is composed of a feature learning sub-network, a plurality of regression sub-networks, and a prediction module. The feature learning sub-network is connected to the plurality of regression sub-networks, and the plurality of regression sub-networks are connected to the prediction module.
[0050] The feature learning sub-network uses the pre-trained convolutional neural network VGG16 on the ImageNet image library as the backbone and removes the last fully connected layer of VGG16. The feature learning sub-network is used to extract features from the input image to obtain features, and divide the features into multiple feature subsets according to the number of regression sub-networks.
[0051] As Figure 2 shown, the feature learning sub-network specifically includes thirteen convolutional layers Conv, four max pooling layers MaxPool, and two fully connected layers FC. Each convolutional layer Conv uses the Relu function as the activation function for non-linear transformation. The input of the feature learning sub-network is the image for target counting.
[0052] In this embodiment, the preferred connection relationship in the feature learning sub-network is as follows: convolutional layer (convolution kernel 3×3, number of channels 64), convolutional layer (convolution kernel 3×3, number of channels 64), max pooling layer (pooling kernel 2×2), convolutional layer (convolution kernel 3×3, number of channels 128), convolutional layer (convolution kernel 3×3, number of channels 128), max pooling layer (pooling kernel 2×2), convolutional layer (convolution kernel 3×3, number of channels 256), convolutional layer (convolution kernel 3×3, number of channels 256), convolutional layer (convolution kernel 3×3, number of channels 256), max pooling layer (pooling kernel 2×2), convolutional layer (convolution kernel 3×3, number of channels 512), convolutional layer (convolution kernel 3×3, number of channels 512), convolutional layer (convolution kernel 3×3, number of channels 512), max pooling layer (pooling kernel 2×2), convolutional layer (convolution kernel 3×3, number of channels 512), convolutional layer (convolution kernel 3×3, number of channels 512), convolutional layer (convolution kernel 3×3, number of channels 512), fully connected layer (4096 neurons), fully connected layer (4096 neurons), and finally the output of the last fully connected layer is respectively connected to the plurality of regression sub-networks.
[0053] In this embodiment, for the specific number of regression sub-networks, its optimal number setting is studied through experiments: different numbers of regression sub-networks are set, experiments are carried out on the RSOC_building dataset, and the optimal number is determined on the validation set. The experimental results are shown in Table 1 below:
[0054] Table 1
[0055] Number of regression sub-networks MAE RMSE 2 6.17 8.72 4 6.09 8.60 8 5.90 8.47 16 6.36 8.50 32 6.78 9.17
[0056] It can be seen from the experimental results in the above table that even if the number of regression sub-networks is small, the network performance can still be maintained at a good level. However, when the number increases to a certain extent, the performance will deteriorate. Because under the constraint of using the same number of parameters, increasing the number of regression sub-networks will make the input information passed to each sub-network less, which will lead to worse performance. Finally, the optimal number of regression sub-networks is determined to be eight. The output features of the feature learning sub-network in this embodiment are divided into eight feature subsets, and each feature subset is respectively used as the input of a regression sub-network and connected to the regression sub-network.
[0057] As Figure 3 shown, each regression sub-network is specifically composed of five fully connected layers FC. Among them, four fully connected layers all use the Relu function as the activation function for linear transformation, and the second-to-last fully connected layer uses the Sigmoid function as the activation function and then connects to the next layer. The specific connection relationship between the fully connected layers in each regression sub-network is as follows: fully connected layer (256 neurons), fully connected layer (128 neurons), fully connected layer (16 neurons), fully connected layer (1 neuron, Sigmoid activation function), fully connected layer (1 neuron, which is also the last linear transformation layer before output).
[0058] The prediction module is used to perform operations based on the outputs of multiple regression sub-networks to obtain the final counting result of the image to be target counted. In this embodiment, it is preferably that the operations performed by the prediction module on the outputs of multiple regression sub-networks are average calculations. After the eight regression sub-networks obtain the predicted values, the eight predicted values are input for average value calculation, and the obtained average value is used as the counting prediction value of the entire image.
[0059] Before the integrated regression network in this embodiment is put into practical application, it needs to be trained according to the loss function of the multi-task learning module. The training database is used to input images into the integrated regression network for training. The training database preferably used in this embodiment includes the following three datasets:
[0060] (1) RSOC_building dataset: A subset of buildings based on satellite images, which includes a total of 2,468 images, with 1,205 images in the training set and 1,263 images in the test set;
[0061] (2) VisDrone2019 People dataset: A crowd counting dataset based on drone images, originating from the VisDrone2019 dataset, which includes 2,392 training images, 329 validation images, and 626 test images;
[0062] (3) VisDrone2019 Vehicle dataset: A vehicle counting dataset based on drone images, including categories of cars, vans, trucks, and buses, which includes 3,953 training images, 364 validation images, and 986 test images.
[0063] In other embodiments, the training database may also contain various other types of datasets.
[0064] The multi-task learning module is used to provide a loss function when training the integrated regression network. The loss function in the multi-task learning module consists of a ranking loss function L Rank , a negative correlation learning loss function L NCL , and a counting loss function L Count These three parts.
[0065] The calculation of the ranking loss requires the introduction of Normalized Discounted Cumulative Gain (NDCG) and Discounted Cumulative Gain (DCG). The calculation derivation of the ranking loss function is as follows:
[0066] For the input images X = {x1,..., x n} for training, where x i ∈χ, i ∈ {1,..., n}, two intermediate parameters of relevance ψ rel and similarity ψ sim are constructed to assist in building the ranking;
[0067] Let the query image for ranking be q, and let the label value and predicted value of the query image q be y q and At this time, the set of retrieved images corresponds to x = {x1, x2,... x n}, let the retrieved image be x i , and let the label value and predicted value of the retrieved image x i be y i and The calculation formulas for the relevance ψ rel and similarity ψ sim are respectively:
[0068]
[0069]
[0070] Among them, ψ rel ∈ [0, 1], ψ sim ∈ [0, 1], y max and y min represent the maximum and minimum values of the label values in the entire dataset, respectively;
[0071] For a set of input images, only the images corresponding to the maximum and minimum values of the label values are taken as query images respectively; for the relevance ψ rel and similarity ψ sim , sort according to the relevance score, and two sorted sequences are obtained accordingly. At this time, the calculation formula of the approximate normalized discounted cumulative gain NDCG(x, q) is:
[0072]
[0073]
[0074] score(x i ) = T × ψ rel (i)
[0075] Among them, ψ(i) rank represents the ranking value of the i-th retrieved image in the sorted sequence, score represents the relevance score between the query image and the retrieved image, score is set in multiple levels, DCG(ψ) represents the discounted cumulative gain of the relevance ψ rel or similarity ψ sim , ψ rel (i) represents the relevance of the i-th retrieved image, and T is a hyperparameter;
[0076] The relevance ψ rel sorting corresponds to the ideal sorted sequence, and the similarity ψ sim sorting corresponds to the sorted sequence predicted by the integrated regression network;
[0077] To calculate the discounted cumulative gain, it is necessary to first clarify the ranking value of the retrieved image x i in the relevance ψ rel and similarity ψ sim sorted sequences. Since the normalized discounted cumulative gain is not differentiable and cannot be directly optimized using gradient descent, the ranking values corresponding to each retrieved image x sim in the similarity ψ i sorting are approximately represented by the sigmoid function, so as to realize gradient calculation;
[0078] For the retrieved image x i the relevance ψ rel ranking value ψ rel (i) rank and the similarity ψ sim ranking value ψ sim (i) rank are calculated as follows:
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] where i and j represent the i-th and j-th retrieved images respectively, and ψ rel (i) and ψ sim (i) represent the values of the j-th retrieved image in their respective corresponding sorting sequences, both are indicator functions, is approximate value, and α = 10 is set to control the approximation degree of the sigmoid function to the indicator function;
[0085] As can be seen from the above derivation, the multi-task learning module is connected to the output of the fully connected layer with the sigmoid function in each regression sub-network, and the sorting loss function L Rank during the training of the integrated regression network is calculated according to the output value. The calculation of the sorting loss function L Rank is shown as follows:
[0086]
[0087] where N is the number of regression sub-networks (preferably eight in this embodiment), and max and min correspond to the images with the largest and smallest label values in a group of input images respectively.
[0088] The negative correlation learning loss function L NCL as one of the loss functions of the multi-task learning module is used to interactively train all regression sub-networks, thereby improving the performance and generalization ability of the network; the calculation of the negative correlation learning loss function L NCL is shown as follows:
[0089]
[0090]
[0091] Among them, M represents the number of images in a group of input images during training, i represents the i-th input image, n and m both represent the n-th and m-th regression sub-networks respectively, and F n (i) and F m (i) respectively represent the predicted values of the n-th and m-th regression sub-networks on the i-th image, and (i) represents the average value of the predicted values of the i-th image on all regression sub-networks. It can be seen that the multi-task learning module is also related to the outputs of each regression sub-network and the prediction module respectively.
[0092] The counting loss function L Count , which is used to implement the end-to-end training method. The calculation of the counting loss function L Count is shown as follows:
[0093]
[0094] Among them, d(i) represents the true count value of the i-th image.
[0095] Combining the above-mentioned ranking loss function L Rank , the negative correlation learning loss function L NCL , and the counting loss function L Count , these three loss functions are respectively given different weight parameters and accumulated to obtain the loss function L total of the multi-task learning module as shown in the following formula:
[0096] L total = L Rank + λL NCL + μL Count
[0097] Among them, λ and μ are the weight parameters of the negative correlation learning loss function and the counting loss function respectively, and both are adjustable parameters.
[0098] In this embodiment, it is preferably to perform data augmentation on the RSOC_building dataset and the VisDrone2019 dataset in the aforementioned training database, and then perform training respectively. The mean absolute error (MAE) and root mean square error (RMSE) are used as two indicators to evaluate the performance of the network. When the results on the validation set reach stability, the training of the integrated regression network is terminated;
[0099] The stochastic gradient descent algorithm is used to determine the optimal hyperparameters T in the integrated regression network and the optimal values of λ and μ in the multi-task learning module in combination with the validation set in the dataset;
[0100] An experiment is carried out on the RSOC_building dataset. The experimental results of how to set the hyperparameter T are shown in Table 2 below:
[0101] Table 2
[0102]
[0103]
[0104] As can be seen from Table 2, as the value of T increases, the performance of the integrated regression network first increases and then decreases, indicating that a compromise value needs to be determined to measure the correlation score between images well. As can be seen from Table 2, this compromise value should be 30. Therefore, T = 30 is set in the embodiment;
[0105] Values of λ around 0.1 usually perform well. Values greater than 0.2 will cause training biases, while values less than 0.02 will make the performance of the regression sub-networks very similar, hardly bringing any empirical benefits. Therefore, λ = 0.1 is fixed and different values of μ are tried; first, μ is changed among {1.0, 0.5, 0.1, 0.05, 0.01}, and it is observed that the value of μ = 0.1 is the most effective. Then, more values are tried near 0.1 and experiments are carried out on the RSOC_building dataset. The experimental results for how to set λ and μ are shown in Table 3 below:
[0106] Table 3
[0107] λ μ MAE RMSE 0.1 1 9.87 13.71 0.1 0.5 6.56 9.66 0.1 0.1 5.68 7.74 0.1 0.05 6.50 8.16 0.1 0.01 8.25 11.17 0.1 0.16 6.15 8.53 0.1 0.14 6.04 8.16 0.1 0.12 5.96 7.98 0.1 0.08 6.07 8.19 0.1 0.06 6.18 7.93 0.1 0.04 6.65 8.50
[0108] As can be seen from Table 3, better network performance can be obtained when μ is around 0.1;
[0109] The experimental results of the preferred training in this embodiment are compared with the effects of other methods on the RSOC_building dataset as shown in Table 4 below:
[0110] Table 4
[0111] Method Year MAE RMSE MCNN 2016 13.65 16.56 CMTL 2017 12.78 15.99 CSRNet 2018 8.00 11.78 SANet 2018 29.01 32.96 SFCN 2019 8.94 12.87 SPN 2019 7.74 11.48 SCAR 2019 26.90 31.35 CAN 2019 9.12 13.38 SFANet 2019 8.18 11.75 BL 2019 11.51 15.96 ASPDNet 2021 7.59 10.66 MSCANet 2021 11.13 16.02 MCFA 2022 7.93 11.82 PSGCNet 2022 7.54 10.52 TransCrowd-Token 2022 8.88 12.48 TransCrowd-GAP 2022 8.58 12.51 This embodiment — 5.62 7.69
[0112] The effects are compared with those of other methods on the VisDrone2019 dataset as shown in Table 5 below:
[0113] Table 5
[0114]
[0115]
[0116] As can be seen from Table 4 and Table 5, the performance of this embodiment on the RSOC_building dataset has improved by 25.46% in MAE and 26.95% in RMSE compared to the existing best method. On the VisDrone2019People dataset, the performance has improved by 10.48% in MAE and 0.60% in RMSE compared to the existing best method. On the VisDrone2019Vehicle dataset, it has improved by 11.69% in MAE and 6.78% in RMSE.
[0117] Compared with the prior art, the beneficial effects of the image target counting system of this embodiment are as follows:
[0118] This embodiment uses multi-task learning, introduces a new loss function to form end-to-end training, and applies normalized discounted cumulative gain to regression counting, avoiding the problem of poor robustness of existing global regression-based methods for image counting, especially for remote sensing images. In addition, by introducing negatively correlated learning loss to the integrated regression network, good generalization ability is achieved, and the overfitting problem in global regression-based methods is alleviated. Different from traditional ensemble learning, this embodiment does not train multiple models but uses a single model training method, reducing the time cost and computational complexity and improving the counting performance of the entire system. Compared with density map estimation-based methods, this embodiment avoids huge annotation costs and improves counting efficiency. Compared with detection-based methods, this embodiment can better adapt to scenes with a higher density of counting targets and has a wider adaptation range.
[0119] Embodiment 2
[0120] As Figure 4 shown, an image target counting method of this embodiment is executed using the image target counting system in Embodiment 1, and includes the following steps:
[0121] S1. According to the type of image for which target counting is to be performed, use the dataset of corresponding same-type images in the training database as the training set for image enhancement, and adjust the size of the image to match the input channel size of the feature learning sub-network in the integrated regression network.
[0122] S2. Input the enhanced training set images into the feature learning sub-network for feature extraction to form respective feature sets corresponding to each image.
[0123] S3. Divide the feature set into eight feature subsets, and then input these eight feature subsets into eight regression sub-networks for processing to obtain corresponding eight predicted values.
[0124] S4. Use the multi-task learning module to process the outputs and other parameters of the eight regression sub-networks based on the loss function Ltotal Perform multi-task learning training until the entire integrated regression network is trained;
[0125] S5. Input the image to be subject to target counting into the trained integrated regression network. After passing through the feature learning sub-network and the regression sub-network in sequence, the final counting prediction value of the target in the image is obtained in the prediction module, and this final counting prediction value is used as the counting result of the target in the input image.
[0126] Compared with the prior art, the beneficial effects of the method for image target counting in this embodiment are as follows:
[0127] Compared with the prior art, it has a faster counting speed, higher counting accuracy, and a larger applicable range. The robustness of this embodiment is good, without serious overfitting problems, and without the need for a large number of annotations or running time.
[0128] Embodiment 3
[0129] This embodiment provides a storage medium.
[0130] The storage medium in this embodiment is provided in the computer system. In this embodiment, the storage medium is preferably a hard disk in the computer system. The hard disk is connected to the CPU of the computer system, and the hard disk is used to store the computer program for implementing the system for image target counting in Embodiment 1.
[0131] When the CPU of the computer system retrieves and executes this computer program from the storage medium, the system for image target counting in Embodiment 1 is established.
[0132] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An image target counting system, characterized in that, It includes an interconnected integrated regression network and a multi-task learning module; The integrated regression network includes a feature learning sub-network, a regression sub-network, and a prediction module connected in sequence; The feature learning sub-network outputs the features of the input image according to the input image, and divides the features into multiple feature subsets; There are multiple regression sub-networks, and the number of regression sub-networks matches the number of feature subsets; the regression sub-network is used to output a predicted value according to the feature subset; The prediction module is used to perform an average value operation on the predicted values of multiple regression sub-networks, so as to obtain the final count prediction value of the target in the image; The multi-task learning module is used to set the loss function for training the integrated regression network; the multi-task learning module is respectively connected to the regression sub-network and the prediction module; The loss function in the multi-task learning module consists of a ranking loss function, a negative correlation learning loss function, and a counting loss function; The loss function l described in the multi-task learning module total , is composed of the weighted sum of the ranking loss function l Rank , the negative correlation learning loss function l NCL , and the counting loss function L Count after assigning weight parameters respectively, as shown in the following formula: L total = L Rank + λL NCL + μL Count Among them, the weight parameter of L Rank is 1, λ is the weight parameter of L NCL , and μ is the weight parameter of L Count ; Sorting loss function L Rank is calculated as shown in the following equation: Among them, NDCG(x i , q) is the approximate normalized discounted cumulative gain calculated according to the ranking of images, x i represents the retrieved image, q represents the query image during ranking, N is the number of regression sub-networks, max and min respectively correspond to the images with the largest and smallest label values in a group of images input during training, i represents the i-th image, and n represents the n-th regression sub-network; Negative correlation learning loss function L NCL The calculation formula is shown as follows: Among them, M represents the number of a set of images input to the integrated regression network during training, both n and m represent the nth and mth regression subnetworks respectively, and F n (i), F m (i) respectively represent the predicted values of the nth and mth regression subnetworks on the ith image, and F(i) represents the average value of the predicted values of the ith image on all regression subnetworks; Counting loss function L Count The calculation formula is as follows: Among them, d(i) represents the true count value of the i-th image.
2. The system for image target counting according to claim 1, wherein It also includes a training database; The training database is connected to the feature learning sub-network and is used to input the stored images into the feature learning sub-network.
3. The system for image target counting according to claim 1, wherein The prediction module is used to perform an average value operation on the predicted values of multiple regression sub-networks, and use the average value as the count prediction value of the target in the image.
4. A method for counting image targets, comprising the system for counting image targets according to any one of claims 1-3, characterized in that, It includes the following steps: S1. According to the type of image for which target counting is to be performed, use the dataset of corresponding same-type images in the training database as the training set for image enhancement, and adjust the size of the image to match the input channel size of the feature learning sub-network in the integrated regression network; S2. Input the enhanced training set images into the feature learning sub-network for feature extraction to form respective feature sets corresponding to each image; S3. Divide the feature set into multiple feature subsets, and then input these multiple feature subsets into the corresponding multiple regression sub-networks for processing to obtain corresponding predicted values; S4. Use the multi-task learning module to perform multi-task learning training on the outputs of multiple regression sub-networks based on the loss function until the entire integrated regression network is trained; S5. Input the image for which target counting is to be performed into the trained integrated regression network, and after passing through the feature learning sub-network and the regression sub-network in sequence, obtain the final count prediction value of the target in the image in the prediction module, and use this final count prediction value as the count result of the target in the input image.
5. The method for counting image targets according to claim 4, wherein The number of feature subsets and regression sub-networks is eight respectively.
6. A storage medium for storing a computer program of the system for counting image targets according to any one of claims 1-3, characterized in that, When the computer program is executed, the system for image target counting described in any one of claims 1-3 is correspondingly implemented.