Method for crowd counting based on reinforcement learning and residual classification network
Through the combination of reinforcement learning and residual classification network, the redundant feature interference and occlusion problems in population counting are solved, and high-precision classification of population counting in complex scenarios is achieved.
Patent Information
- Application Number
- CN202211380112.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-05
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-11-05
AI Technical Summary
The prior art has problems with redundant feature interference and occlusion in crowd counting, resulting in low counting accuracy and poor performance in complex scenarios.
The method based on reinforcement learning and residual classification network is adopted to perform block classification through residual classification convolution neural network, and the classification results are adjusted in fine-grained manner in combination with reinforcement learning. The block classification loss function is used to accelerate the convergence of the network model and improve the classification accuracy.
It realizes a more accurate classification of the population number in complex scenarios, and improves the accuracy and accuracy of population counting.
Smart Images

Figure CN115761621B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to a method for crowd counting based on reinforcement learning and residual classification network. Background Art
[0002] Crowd counting, that is, Crowd Counting, is a key task in computer vision. Its purpose is to count the people in an image. In scenes with dense crowds such as tourist attractions and large gatherings, it is necessary to estimate the number of people in real time to prevent safety accidents such as crowd trampling. Currently, the main method adopted by the academic community for crowd counting problems is to regress the density map through a deep convolutional neural network (CNN), that is, to mark the people in the image to form a density map full of marked points, and the sum of the density map can obtain the number of people in the original image. However, due to problems such as interference caused by redundant features and the inability to mark when people are blocked, there is still a large room for improvement in the statistical accuracy of this method.
[0003] With the development of deep convolutional neural networks in recent years, some researchers have migrated the method of image segmentation to the crowd counting task and achieved certain results. However, the performance of the block classification method in complex scenes (such as the UCF-CC-50 dataset) is not ideal because in complex scenes, the image contains a large amount of redundant features and interference information, and using these features and information to perform block classification on the image will have a great impact on the final classification result. Summary of the Invention
[0004] The present invention proposes a method for crowd counting based on reinforcement learning and residual classification network, which can more accurately classify the number of people in the crowd and obtain the specific number of people in the crowd according to the classification result.
[0005] The present invention adopts the following technical solutions.
[0006] A method for crowd counting based on reinforcement learning and residual classification network includes the following steps:
[0007] Step S1: Define the number of people as multiple categories by using a preset function, and there is a corresponding relationship between the category and the range of the number of people;
[0008] Step S2: Use a camera to capture real-time dense crowd scenes prone to safety problems, input the original image into the residual classification convolutional neural network, and train the network with a block classification loss function until the network weights are stable, and obtain the feature map and block classification result belonging to the image, that is, the category map;
[0009] Step S3: Input the feature map of the image and the category map into the reinforcement learning evaluation network. The evaluation network makes precise adjustments to the category map based on the image features to obtain a more fine-grained category map.
[0010] Step S4: Use a preset function to map the category map of the image back to the number of people to obtain the counting map of the original image. The values in the counting map are accumulated to obtain the predicted number of people in the image monitored by the original camera.
[0011] Step S1 specifically includes the following steps;
[0012] Step S11: Determine which category the number of people count belongs to. It is necessary to judge in advance whether count is greater than the lower limit of the interval of the index-th class class index The lower limit of the interval of the index-th class is defined by the following formula:
[0013]
[0014] As can be seen from the above formula, the lower limit of the number of people in the first class is 0, the lower limit of the number of people in the second class is 0.13, and so on; if the number of people count is greater than the lower limit of the number of people in the index-th class, then proceed to step S12;
[0015] Step S12: Compare the number of people count with the upper limit of the interval of the index-th class class index The upper limit of the interval of the index-th class is defined by the following formula:
[0016]
[0017] As can be seen from the above formula, the upper limit of the number of people in the first class is 0.13, the upper limit of the number of people in the second class is 0.15, and so on;
[0018] Step S13: The upper and lower limits of the first class, the second class, the third class... and more classes can be defined through the above formula. Subsequently, use this as a standard to train the residual classification neural network so that it can classify the number of people.
[0019] Step S2 includes the following steps;
[0020] Step S21: The camera captures the crowded crowd scene in real time to obtain the crowd image image; input the original image image into the residual classification convolutional neural network;
[0021] Step S22: The backbone network part in the residual classification convolutional neural network extracts the feature map feature_map of the original image image image of the original image, and the residual block part extracts the residual attention map residual_map of the original image imageimage The residual attention map and the feature map are fused through a specific formula to obtain the final feature map final_feature_map of the original image image. image The formula is as follows:
[0022] fihal_feature_map image =
[0023] feature_map image + residual_map image × λ, where λ ∈ (0, 1]
[0024] Formula 3;
[0025] In the formula, λ belongs to the hyperparameter and can be adjusted according to the actual situation;
[0026] Since the original image image has undergone five downsampling operations in the backbone network part, each feature point feature_point in the feature map final_feature_map image stores the crowd information of each 32×32-sized patch image_patch in the corresponding original image image index ; Subsequently, the feature map is input into the classifier; index Step S23: The classifier formed by the 1×1 convolutional layer classifies each feature point feature_point in the feature map final_feature_map
[0027] to obtain the class map class_map image which stores the classification results of each 32×32-sized patch image_patch in the original image image index ; image Step S24: Use the patch classification loss function patch_class_loss to train the residual classification convolutional neural network until the network weights are stable, and obtain the optimal network weights. The patch classification loss function is as follows: index In the formula, class
[0028] represents the total number of classes,
[0029]
[0030] represents the number of times the index-th class class total appears in the class map, and image_patch represents the total number of 32×32-sized patches of the original image, index sum sum sum Represents the probability of the index-th class <class> obtained by the residual convolutional neural network, where γ belongs to hyperparameters and can be adjusted according to the actual situation. index The probability, where γ belongs to hyperparameters and can be adjusted according to the actual situation.
[0031] The specific steps of step S3 are as follows;
[0032] Step S31: Input the feature map final_feature_map image into the reinforcement learning evaluation network composed of three 1×1 convolutional layers. The evaluation network determines whether the class map class_map image needs to be adjusted and outputs an action map action_map image for adjusting the class map class_map imnage ;
[0033] Step S32: Add the action map action_map image and the class map class_map image to obtain the final class map final_class_map image , and the formula is as follows:
[0034] final_class_map image =action_map image +class_map image Formula Five.
[0035] The specific steps of step S4 are as follows;
[0036] Step S41: According to the formula, calculate the upper and lower limits of each class in the obtained class map final_class_map image , where i represents the abscissa of the class map and j represents the ordinate of the class map. The formula is as follows: (i,j)
[0037]
[0038]
[0039]
[0039] Step S42: After calculating the upper and lower limits of each class, calculate the specific number of people in each class according to the upper and lower limits. The formula is as follows:
[0040]
[0041] Step S43: Calculate the specific number of people count_map of each class in the class map (i,j) final_class_map(i,j) After that, a count map count_map of the original image image can be formed. image The count map stores the number of people results for each 32×32-sized patch image_patch in the original image image. index The number of people results;
[0042] Step S44: Sum the count map count_map image to obtain the total number of people count of the original image image image , and the summation formula is as follows:
[0043] count image = ∑ i=0 ∑ j=0 count_map (i,j) Formula IX.
[0044] The dense crowd scenes prone to security problems in step S1 include dense crowd scenes in tourist attractions and large-scale gathering scenes.
[0045] The present invention has the following beneficial effects compared with the prior art:
[0046] 1. The present invention improves the classification convolutional neural network, creates a residual attention module, strengthens the attention of the classification convolutional neural network to the crowd part in the image, and improves the classification effect of the classification convolutional neural network on the crowd;
[0047] 2. Based on the image segmentation idea, the present invention creates a block classification loss function on the basis of the cross-entropy loss function, accelerates the convergence of the network model, improves the sensitivity of the network model to categories, and thus obtains more accurate prediction results;
[0048] 3. Traditional methods only perform crowd counting based on convolutional neural networks. The present invention creatively uses reinforcement learning to finely adjust the results output by the convolutional neural network, makes up for the limitations of the convolutional neural network in classification prediction tasks, strengthens the classification ability of the network model, and obtains more accurate classification results than the convolutional neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0050] Attached Figure 1 is a flow diagram of the present invention. SPECIFIC EMBODIMENTS
[0051] As Figure 1 shown, a method for crowd counting based on reinforcement learning and residual classification network includes the following steps:
[0052] Step S1: Define the number of people into multiple categories using a preset function, where the categories have a corresponding relationship with the range of the number of people;
[0053] Step S2: The camera takes real-time pictures of crowded scenes where safety problems are likely to occur, and inputs the original images into the residual classification convolutional neural network. Train the network with a block classification loss function until the network weights are stable, and obtain the feature map and block classification result of the image, that is, the category map;
[0054] Step S3: Input the feature map and category map of the image into the reinforcement learning evaluation network. The evaluation network makes precise adjustments to the category map according to the image features to obtain a more fine-grained category map;
[0055] Step S4: Use a preset function to map the category map of the image back to the number of people to obtain the counting map of the original image. The values in the counting map are accumulated to obtain the predicted number of people in the image monitored by the original camera.
[0056] Step S1 specifically includes the following steps;
[0057] Step S11: Determine which category the number of people count belongs to. It is necessary to judge in advance whether count is greater than the lower limit of the interval of the index-th class class index The lower limit of the interval of the index-th class is defined by the following formula:
[0058]
[0059] As can be seen from the above formula, the lower limit of the number of people in the first class is 0, and the lower limit of the number of people in the second class is 0.13, and so on; if the number of people count is greater than the lower limit of the number of people in the index-th class, then continue to step S12;
[0060] Step S12: Compare the number of people count with the upper limit of the interval of the index-th class class index of the interval;
[0061] The upper limit of the interval of the index-th class is defined by the following formula:
[0062]
[0063] As can be seen from the above formula, the upper limit of the number of people in the first class is 0.13, and the upper limit of the number of people in the second class is 0.15, and so on;
[0064] Step S13: The upper and lower limits of the first class, the second class, the third class... and more classes can be defined through the above formula. Subsequently, use this as a standard to train the residual classification neural network so that it can classify the number of people.
[0065] The step S2 includes the following steps;
[0066] Step S21: The camera takes real-time pictures of the dense crowd scene to obtain the crowd image image; The original image image is input into the residual classification convolutional neural network;
[0067] Step S22: The backbone network part in the residual classification convolutional neural network extracts the feature map feature_map of the original image image image , and the residual block part extracts the residual attention map residual_map of the original image image image , and the residual attention map and the feature map are fused through a specific formula to obtain the final feature map final_feature_map of the original image image image . The formula is as follows:
[0068] final_feature_map image =
[0069] feature_map image +residual_map image ×λ λ∈(0, 1] Formula 3;
[0070] In the formula, λ belongs to the hyperparameter and can be adjusted according to the actual situation;
[0071] Since the original image image has undergone five downsampling operations in the backbone network part, each feature point feature_point in the feature map final_feature_map image stores the crowd information of each 32×32-sized patch image_patch in the corresponding original image image index ; Subsequently, the feature map is input into the classifier; index
[0072] Step S23: The classifier formed by the 1×1 convolutional layer classifies each feature point feature_point in the feature map final_feature_map image to obtain the class map class_map index , and this map stores the classification results of each 32×32-sized patch image_patch in the original image image image ; index
[0073] Step S24: Use the patch classification loss function patch_class_loss to train the residual classification convolutional neural network until the network weights are stable and the optimal network weights are obtained. The patch classification loss function is as follows:
[0074]
[0075] Where class total Represents the total number of categories, Represents the index-th class index The number of occurrences in the category graph, image_patch sum Represents the total number of 32×32 patches of the original image, Represents the index-th class class obtained by the residual convolutional neural network index The probability of γ is a hyperparameter and can be adjusted according to actual conditions.
[0076] The step S3 specifically includes the following steps:
[0077] Step S31: final_feature_map image Input to the reinforcement learning evaluation network composed of three layers of 1×1 convolutional layers, the evaluation network judges the category map class_map based on it image Whether adjustment is needed, output a class map for adjustment class_map image action_map image ;
[0078] Step S32: Set the action map action_map image With the class map class_map image Add together to get the final category map final_class_map image , the formula is as follows:
[0079] final_class_map image =action_map image +class_map image Formula five.
[0080] The step S4 comprises the following steps:
[0081] Step S41: According to the formula, the obtained category map final_class_map image Calculate the final_class_map for each category (i,j) The upper and lower limits of the category graph are as follows: i represents the horizontal coordinate of the category graph, and j represents the vertical coordinate of the category graph.
[0082]
[0083]
[0084] Step S42: After calculating the upper and lower limits of each category, calculate the specific number of people in each category according to the upper and lower limits. The formula is as follows:
[0085]
[0086] Step S43: After calculating the specific number of people count_map of each category final_class_map in the category graph (i,j) the count map of the original image image can be formed (i,j) after that, and the count map stores the number of people results of each 32×32-sized patch image_patch image in the original image image; index
[0087] Step S44: Sum the count map count_map image to obtain the total number of people count of the original image image image The summation formula is as follows:
[0088] count image = ∑ i=0 ∑ j=0 count_map (i,j) Formula IX.
[0089] The crowded people scenes prone to security problems in the said Step S1 include the crowded people scenes in tourist attractions and large-scale gathering scenes.
Claims
1. A method for crowd counting based on reinforcement learning and residual classification network, characterized in that: Including the following steps: Step S1: Define the number of people into multiple categories using a preset function, where the categories have a corresponding relationship with the range of the number of people; Step S2: The camera takes real-time pictures of crowded people scenes where safety problems are likely to occur, and inputs the original image into the residual classification convolutional neural network. Train the network with the patch classification loss function until the network weights are stable, and obtain the feature map and patch classification result of the image, that is, the category map; Step S3: Input the feature map and category map of the image into the reinforcement learning evaluation network. The evaluation network makes precise adjustments to the category map according to the image features to obtain a more fine-grained category map; Step S4: Use a preset function to map the category map of the image back to the number of people to obtain the counting map of the original image. The values in the counting map are accumulated to obtain the predicted number of people in the image monitored by the original camera; Step S2 includes the following steps; Step S21: The camera takes real-time pictures of crowded people scenes to obtain the crowd image image. Input the original image image into the residual classification convolutional neural network; Step S22: The backbone network part in the residual classification convolutional neural network extracts the feature map feature_map of the original image image image , and the residual block part extracts the residual attention map residual_map of the original image image image , and the residual attention map and the feature map are fused through a specific formula to obtain the final feature map final_feature_map of the original image image image ; The formula is as follows: final_feature_map image = feature_map image +residual_map image ×λ where λ ∈ (0, 1] Formula Three; In the formula, λ belongs to a hyperparameter and can be adjusted according to the actual situation; Since the original image undergoes five downsampling operations in the backbone network part, the feature map final_feature_map image each feature point feature_point index stores the crowd information of each 32×32-sized patch image_patch in the corresponding original image image; Subsequently, the feature map is input into the classifier; index Step S23: The classifier formed by the 1×1 convolutional layer classifies each feature point feature_point image in the feature map final_feature_map index to obtain the class map class_map image , which stores the classification results of each 32×32-sized patch image_patch index in the original image image; Step S24: Use the patch classification loss function patch_class_loss to train the residual classification convolutional neural network until the network weights are stable to obtain the optimal network weights. The patch classification loss function is as follows: Formula Four; where class total represents the total number of classes, represents the number of times the index-th class class index appears in the class graph, image_patch sum represents the total number of 32×32-sized patches of the original image, represents the probability of the index-th class class index obtained by the residual convolutional neural network, where γ belongs to hyperparameters and can be adjusted according to the actual situation; Step S4 includes the following steps; Step S41: According to the formula, for the obtained class map final_class_map image calculate the upper and lower limits of each class in it, where i represents the abscissa of the class map and j represents the ordinate of the class map; the formula is as follows: (i,j) The upper and lower limits of each class final_class_map, i represents the abscissa of the class map, and j represents the ordinate of the class map; the formula is as follows: Formula Six; Formula Seven; Step S42: After calculating the upper and lower limits of each category, calculate the specific number of people in each category according to the upper and lower limits. The formula is as follows: Step S43: Calculate the specific number of people count_map for each category final_class_map in the category map (i,j) After that, the count map count_map of the original image image can be formed (i,j) The count map stores the number of people results for each 32×32-sized patch image_patch in the original image image image ; index Step S44: Sum up the counting graph count_map image to obtain the total number of people count in the original image image image , and the summation formula is as follows:
2. The method for crowd counting based on reinforcement learning and residual classification network according to claim 1, characterized in that: The details of Step S1 Including the following steps; Step S11: Determine which category the number of people count belongs to. It is necessary to judge in advance whether count is greater than the lower limit of the interval of the index-th class class. index The lower limit of the interval of the index-th class is defined by the following formula: It can be seen from the above formula that the lower limit of the number of people in the first category is 0, the lower limit of the number of people in the second category is 0.13, and so on. If the number of people count is greater than the lower limit of the number of people in the index-th category, then continue with Step S12; Step S12: Compare the number of people count with the upper limit of the interval of the index index-th class class; The upper limit of the interval of the index-th class is defined by the following formula: It can be seen from the above formula that the upper limit of the number of people in the first category is 0.13, the upper limit of the number of people in the second category is 0.15, and so on; Step S13: The upper and lower limits of the first category, the second category, the third category... and more categories can be defined through the above formula. Subsequently, use this as a standard to train the residual classification neural network so that it can classify the number of people.
3. The method for crowd counting based on reinforcement learning and residual classification network according to claim 1, wherein: Step S3 specifically includes the following steps; Step S31: Input the feature map final_feature_map image into the reinforcement learning evaluation network composed of three 1×1 convolutional layers. The evaluation network determines whether the class map image needs to be adjusted, and outputs an action map image for adjusting the class map image ; Step S32: Add the action map action_map image to the class map class_map image to obtain the final class map final_class_map image , and the formula is as follows: final_class_map image = actio_map image + class_map image Formula Five.
4. The method for crowd counting based on reinforcement learning and residual classification network according to claim 1, characterized in that: The crowded people scenes where safety problems are likely to occur in Step S1 include crowded people scenes in tourist attractions and large-scale gathering scenes.
Citation Information
Patent Citations
Crowd counting method based on scene depth information
CN110059581A
Complex scene crowd counting method based on scene classification and multi-scale feature fusion
CN111783589A