DGCC-Net model-based crowd counting method

By adopting the population counting method based on the DGCC-Net model in population counting, differentiated attention to different density areas using density guide maps and multi-level supervision loss functions, the problems of uneven density distribution and insufficient fine-grained feature capture capabilities are solved, and the population counting effect with high accuracy and robustness is achieved.

CN120220065APending Publication Date: 2025-06-27SICHUAN CHUANJIAO ROAD & BRIDGE
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510313358.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems with uneven density distribution in population counting, and the existing network models have shortcomings in dealing with long-distance dependencies and global context information, especially in high-density crowd scenarios, which are difficult to effectively capture fine-grained feature information.

Method used

Using the population counting method based on the DGCC-Net model, differentiated attention to different density areas is achieved by generating a density guide map and introducing a multi-level supervised loss function, and the counting accuracy and robustness are improved. The method includes feature extraction network, local attention module, multi-level feature fusion network and prediction network, feature extraction is performed using Twins Transformer, and feature capture capabilities are enhanced through local attention module and density guidance module.

Benefits of technology

It significantly improves the feature extraction capability of the model in complex scenarios, allowing it to effectively capture fine-grained feature information in high-density crowd scenarios, and improves the accuracy and robustness of the overall counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220065A_ABST
    Figure CN120220065A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crowd counting, in particular to a crowd counting method based on a DGCC-Net model, and the method comprises the steps: obtaining a crowd counting image; a DGCC-Net model is constructed, and pre-training is carried out; extracting a multi-scale feature map from the crowd counting image by using a feature extraction network; a local attention module is used for extracting context information and detail information of a crowd region from a feature map on a high-resolution scale in the multi-scale feature map, and a high-resolution feature map is obtained; feature fusion of semantic and detail information is carried out on a feature map on a low-resolution scale and a high-resolution feature map in the multi-scale feature map by using a multi-level feature fusion network, and a density guide feature map is generated from the fused feature map through a density guide module; and predicting a crowd counting result according to the density guide feature map by using a prediction network. According to the invention, the counting precision and robustness are improved when crowd scenes with uneven density distribution are processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crowd counting, and particularly relates to a crowd counting method based on a DGCC-Net model. Background Art

[0002] In the field of crowd counting, traditional methods mainly rely on sensors and manual counting, and these methods have many deficiencies in practical applications. Sensor methods such as infrared, radar, and lidar devices can, to a certain extent, achieve the detection and counting of crowds, but are limited by factors such as the installation location of the devices and environmental interference, and are prone to large errors. The manual counting method is not only inefficient but also difficult to ensure accuracy in large crowds. Therefore, with the development of computer vision technology, more and more researchers have begun to use image processing and machine learning technologies for crowd counting.

[0003] Currently, computer vision-based crowd counting methods are mainly divided into three categories: regression-based methods, detection-based methods, and density map-based methods. The regression-based method extracts global or local features of an image, establishes a mapping relationship between the image features and the number of people, and directly outputs the number of people. However, this method cannot provide the spatial distribution information of the crowd, and the accuracy is often low when dealing with high-density crowds. The detection-based method uses object detection technology to achieve crowd counting by detecting the position and number of each human head, but in high-density crowd scenarios, the effect of object detection will be affected by occlusion and overlap problems, resulting in a decrease in detection accuracy. The density map-based method generates a density map, converts the number of people into density values in the image, and then calculates the total number of people by integrating the density map. Compared with the first two methods, the density map-based method performs better in high-density crowd scenarios, but there are still some technical challenges in dealing with the problem of uneven crowd density distribution.

[0004] Existing density map-based crowd counting methods mainly rely on the generation of crowd density maps and the design of network models. The generation of crowd density maps usually uses a Gaussian kernel function to expand the labeled points. However, due to the camera perspective effect and the difference in crowd distribution, a single-scale Gaussian kernel is difficult to adapt to different density regions, resulting in low-quality generation of density maps, which in turn affects the counting accuracy of the model. In addition, most existing network models are based on convolutional neural networks (CNNs), and although they have good performance in feature extraction, they have deficiencies in dealing with long-range dependence relationships and global context information. Especially in high-density crowd scenarios, it is difficult for the model to effectively capture fine-grained feature information.

[0005] In summary, the existing technologies have the following defects in the crowd counting method: First, the method for generating the crowd density map has poor performance in dealing with the problem of uneven density distribution; second, the existing network models have deficiencies in dealing with long-range dependence relationships and global context information, especially in high-density crowd scenarios, it is difficult to effectively capture fine-grained feature information. Summary of the Invention

[0006] In view of the above deficiencies in the existing technologies, the present invention provides a crowd counting method based on the DGCC-Net model. By generating a density guidance map and introducing a multi-level supervised loss function, it realizes differential attention to different density regions, and improves the counting accuracy and robustness in dealing with crowd scenarios with uneven density distribution.

[0007] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows: A crowd counting method based on the DGCC-Net model, comprising the following steps: Obtain a crowd counting image; Construct a DGCC-Net model and perform pre-training; the DGCC-Net model includes a feature extraction network, a local attention module, a multi-level feature fusion network, and a prediction network; Use the feature extraction network to extract multi-scale feature maps from the crowd counting image; Use the local attention module to extract context information and detailed information of the crowd region from the feature map at the high-resolution scale in the multi-scale feature map, and obtain a high-resolution feature map; Use the multi-level feature fusion network to perform feature fusion of semantic and detailed information on the feature map at the low-resolution scale and the high-resolution feature map in the multi-scale feature map, and generate a density guidance feature map by passing the fused feature map through a density guidance module; Use the prediction network to predict the crowd counting result according to the density guidance feature map.

[0008] Preferably, using the local attention module to extract context information and detailed information of the crowd region from the feature map at the high-resolution scale in the multi-scale feature map, and obtain a high-resolution feature map, includes: Use the multi-scale receptive field module included in the local attention module to extract multi-scale context information through parallel dilated convolutions with different receptive fields, and generate a feature map containing multi-scale context information; Use the rotation attention module included in the local attention module to extract detailed information of the crowd region from the feature map containing multi-scale context information by dynamically adjusting the feature map weights, and obtain a high-resolution feature map.

[0009] Preferably, the multi-scale receptive field module included in the local attention module extracts multi-scale context information through dilated convolutions with different receptive fields in parallel, generating a feature map containing multi-scale context information, including: Using the first dilated convolution, the second dilated convolution, and the third dilated convolution with three different dilation rates included in the multi-scale receptive field module, as well as the first convolutional layer, extract context information from the input feature map at different scales, generating four feature maps with different scales; Using the first global average pooling layer included in the multi-scale receptive field module to perform global average pooling on the input feature map, and restoring the feature map to the original resolution through the second convolutional layer and bilinear interpolation operation to obtain a global average pooling feature map; Using the first concatenation layer included in the multi-scale receptive field module to concatenate the four feature maps with different scales and the global average pooling feature map, generating a feature map containing multi-scale context information.

[0010] Preferably, the rotation attention module included in the local attention module extracts detailed information of the crowd area from the feature map containing multi-scale context information by dynamically adjusting the feature map weights, obtaining a high-resolution feature map, including: Using the rotation attention module to rotate the input feature map 90° counterclockwise along the height axis to obtain a first rotation tensor, then forming a first concatenated feature map through the concatenation operation of max pooling and average pooling, then performing a convolution operation through the third convolutional layer and obtaining a first attention weight through an activation function, and finally multiplying the first attention weight element-wise with the input feature map and then rotating it 90° clockwise along the height axis to obtain a first rotation attention feature map; Using the rotation attention module to rotate the input feature map 90° counterclockwise along the width axis to obtain a second rotation tensor, then forming a second concatenated feature map through the concatenation operation of max pooling and average pooling, then performing a convolution operation through the fourth convolutional layer and obtaining a second attention weight through an activation function, and finally multiplying the second attention weight element-wise with the input feature map and then rotating it 90° clockwise along the width axis to obtain a second rotation attention feature map; Using the rotation attention module to form a third concatenated feature map through the concatenation operation of max pooling and average pooling on the input feature map, then performing a convolution operation through the fifth convolutional layer and obtaining a third attention weight through an activation function, and finally multiplying the third attention weight element-wise with the input feature map to obtain a third rotation attention feature map; Using the rotation attention module to add the first rotation attention feature map, the second rotation attention feature map, and the third rotation attention feature map element-wise to obtain a high-resolution feature map.

[0011] Preferably, a local attention module is used to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map, and a high-resolution feature map is obtained. The method further includes: The high-resolution feature map is passed through a sixth convolutional layer to extract regional features and perform batch normalization operations, and then through a seventh convolutional layer for convolutional operations, and then through an activation function to generate feature map weights. Finally, the feature map weights are multiplied element-wise with the high-resolution feature map to obtain the output feature map of the local attention module.

[0012] Preferably, generating a density-guided feature map from the fused feature map by a density guidance module includes: The number of channels of the fused feature map is gradually reduced from 256 to 1, and the activation function is used to normalize the output feature values to the range of 0 to 1 to generate a density guidance map; After multiplying the density guidance map and the original feature map element-wise, the local and global attention mechanisms are fused through two layers of Twins-Transformer basic blocks to generate a density-guided feature map.

[0013] Preferably, when pre-training the DGCC-Net model, The head positions in the crowd counting image are labeled, and the crowd counting image is mapped into a point annotation map containing the number of people; The labeled points are processed by Gaussian filtering using a Gaussian function to generate a crowd density map label with a Gaussian distribution; According to the nearest neighbor distance, the current crowd density level and the size of the head area are determined to generate a density guidance map label.

[0014] Preferably, determining the current crowd density level and the size of the head area according to the nearest neighbor distance to generate a density guidance map label includes: For each labeled point, calculate the average distance between the nearest neighbor labeled points to estimate the local density degree; According to the local density degree of each head labeled point, assign corresponding density levels to each head labeled point in the image; According to the local density degree of each head labeled point, use a circular area to simulate the head area, form the corresponding density distribution on the density guidance map, and characterize the distribution characteristics of the crowd at different density levels by adjusting the radius of the circular area, and establish a matching relationship between the radius of the circular area and the local density level to generate a density guidance map label.

[0015] Preferably, the matching relationship between the radius of the circular area and the local density level is:

[0016] Wherein, sizei is the radius of the circular area of the i th personal head marker point, L i is the local density degree of the i th personal head marker point, min is the minimum value function, max is the maximum value function, e is the scientific notation symbol.

[0017] Preferably, when pre-training the DGCC-Net model, a multi-level supervised loss function is constructed to perform multi-level supervision on the three-scale branches and the final prediction branch, expressed as:

[0018]

[0019]

[0020]

[0021] Among them, is the multi-level supervised loss function, is the weight of the i th level, is the structure loss function of the i th level, is the balance coefficient, is the total variable loss function of the i th level, is the density-guided loss function of the i th level, is the mean of the pixel values in the ground truth label map, is the mean of the pixel values in the predicted density map, is the covariance between the pixel values of the ground truth label map and the predicted density map, is the variance of the pixel values in the ground truth label map, is the variance of the pixel values in the predicted density map, is a set constant, is the pixel value of the predicted density map at position (m,n), is the predicted pixel value, is the true label pixel value, N is the total number of pixels.

[0022] The present invention has the following beneficial effects: (1) The present invention uses Twins Transformer as the feature extraction network. Utilizing its powerful semantic information acquisition ability demonstrated globally, compared with traditional convolutional neural networks, it can better handle long-range dependencies and global context information, significantly enhancing the model's feature extraction ability in complex scenarios. It can still effectively capture fine-grained feature information in high-density crowd scenarios, improving the accuracy of overall counting.

[0023] (2) Aiming at the defect that existing crowd counting methods have poor effects in dealing with uneven density distribution problems, the present invention proposes a density guidance map generated based on the nearest neighbor distance. By introducing a density guidance module and using the density guidance map to pay differential attention to different density regions, the model can determine the crowding degree of the current crowd and the size of the human head region according to the nearest neighbor distance of the human head points, generating density guidance map labels, effectively solving the problem of uneven crowd density distribution and making the model more robust in dealing with scenarios with different density distributions.

[0024] (3) The present invention introduces a local attention module at the high-resolution scale, including a multi-scale receptive field module and a rotational attention module. The multi-scale receptive field module captures multi-scale context information through parallel dilated convolutions with different receptive fields, while the rotational attention module enhances the attention to the crowd region by dynamically adjusting the feature map weights, significantly enhancing the ability to capture detailed features. Through a multi-level feature fusion module, features at different levels are fused, and a dynamic learning weight mechanism is adopted to achieve effective fusion of semantic and detailed information, enabling the model to maintain a high counting accuracy when dealing with high-density and sparse crowd scenarios.

[0025] (4) The present invention constructs a multi-level supervised loss function, including density guidance loss, structural similarity loss, and total variation loss. The density guidance loss is used for the supervision of the density guidance map. The structural similarity loss selects the dense crowd region through a dense region mask to ensure the prediction quality of the dense region, while the total variation loss aims to reduce overly smooth predictions and enhance the model's adaptability to different density distributions. By introducing a multi-level supervision mechanism, the model can perform fine-grained supervision at different levels, improving the counting accuracy and robustness for scenarios with different density distributions. Brief Description of the Drawings

[0026] Figure 1 It is a schematic flow diagram of a crowd counting method based on the DGCC-Net model; Figure 2 It is a schematic diagram of the DGCC-Net model structure; Figure 3 It is a schematic diagram of the local attention module framework; Figure 4 It is a schematic diagram of the multi-level feature fusion module framework; Figure 5 It is a density-guided map effect display diagram; Figure 6 It is a visualization example diagram of the prediction result of DGCC-Net on the Part A dataset; Figure 7 It is a visualization example diagram of the prediction result of DGCC-Net on the UCF-QNRF dataset. Specific implementation manners

[0027] The following describes the specific implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0028] As Figure 1 shown, a crowd counting method based on the DGCC-Net model provided by an embodiment of the present invention includes the following steps S1 to S6: S1. Obtain a crowd counting image; S2. Construct a DGCC-Net model and perform pre-training; the DGCC-Net model includes a feature extraction network, a local attention module, a multi-level feature fusion network, and a prediction network; S3. Use the feature extraction network to extract multi-scale feature maps from the crowd counting image; S4. Use the local attention module to extract context information and detailed information of the crowd region from the feature map at the high-resolution scale in the multi-scale feature map to obtain a high-resolution feature map; S5. Use the multi-level feature fusion network to perform feature fusion of semantic and detailed information on the feature map at the low-resolution scale and the high-resolution feature map in the multi-scale feature map, and generate a density-guided feature map by passing the fused feature map through a density guidance module; S6. Use the prediction network to predict the crowd counting result according to the density-guided feature map.

[0029] In an optional embodiment of the present invention, as Figure 2 shown, in this embodiment, Twins Transformer is used as the feature extraction network to extract features from the input image to obtain feature maps of different scales, denoted as {F1, F2, F3}, where F1, F2, and F3 respectively represent feature maps at scales of 1 / 8, 1 / 16, and 1 / 32.

[0030] In an optional embodiment of the present invention, as Figure 3As shown, in this embodiment, a local attention module is introduced at the high-resolution scale to enhance the ability to capture detailed features.

[0031] In this embodiment, the local attention module is used to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map, and a high-resolution feature map is obtained, including: The multi-scale receptive field module included in the local attention module extracts multi-scale context information through parallel dilated convolutions with different receptive fields, and generates a feature map containing multi-scale context information; The rotation attention module included in the local attention module extracts detailed information of the crowd area from the feature map containing multi-scale context information by dynamically adjusting the feature map weights, and a high-resolution feature map is obtained.

[0032] In this embodiment, the multi-scale receptive field module included in the local attention module extracts multi-scale context information through parallel dilated convolutions with different receptive fields, and generates a feature map containing multi-scale context information, including: The first dilated convolution, the second dilated convolution, and the third dilated convolution with three different dilation rates and the first convolutional layer included in the multi-scale receptive field module extract context information from the input feature map at different scales, and generate four feature maps with different scales; The first global average pooling layer included in the multi-scale receptive field module performs a global average pooling operation on the input feature map, and restores the feature map to the original resolution through the second convolutional layer and the bilinear interpolation operation, and a global average pooling feature map is obtained; The first concatenation layer included in the multi-scale receptive field module concatenates the four feature maps with different scales and the global average pooling feature map, and generates a feature map containing multi-scale context information.

[0033] In this embodiment, the rotation attention module included in the local attention module extracts detailed information of the crowd area from the feature map containing multi-scale context information by dynamically adjusting the feature map weights, and a high-resolution feature map is obtained, including: The rotation attention module rotates the input feature map 90° counterclockwise along the height axis to obtain a first rotation tensor, then forms a first concatenated feature map through the concatenation operation of max pooling and average pooling, then performs a convolution operation through the third convolutional layer and obtains a first attention weight through an activation function, and finally multiplies the first attention weight element-wise with the input feature map and rotates it 90° clockwise along the height axis to obtain a first rotation attention feature map; The rotation attention module is used to rotate the input feature map counterclockwise by 90° along the width axis to obtain a second rotation tensor, and then a second concatenated feature map is formed through the concatenation operation of max pooling and average pooling. Then, a convolution operation is performed through a fourth convolutional layer and a second attention weight is obtained through an activation function. Finally, the second attention weight is multiplied element-wise with the input feature map and then rotated clockwise by 90° along the width axis to obtain a second rotation attention feature map; The rotation attention module is used to form a third concatenated feature map by the concatenation operation of max pooling and average pooling on the input feature map, and then a convolution operation is performed through a fifth convolutional layer and a third attention weight is obtained through an activation function. Finally, the third attention weight is multiplied element-wise with the input feature map to obtain a third rotation attention feature map; The rotation attention module is used to add the first rotation attention feature map, the second rotation attention feature map, and the third rotation attention feature map element-wise to obtain a high-resolution feature map.

[0034] In this embodiment, the local attention module is used to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map to obtain a high-resolution feature map, and it further includes: The high-resolution feature map is used to extract regional features through a sixth convolutional layer and perform a batch normalization operation, and then a convolution operation is performed through a seventh convolutional layer and a feature map weight is generated through an activation function. Finally, the feature map weight is multiplied element-wise with the high-resolution feature map to obtain the output feature map of the local attention module.

[0035] The multi-scale receptive field module in this embodiment captures multi-scale context information through parallel dilated convolutions with different receptive fields and generates a feature map :

[0036] Among them, is a dilated convolution function with a standard deviation of ; The rotation attention module in this embodiment enhances the attention to the crowd area by dynamically adjusting the feature map weight and generates a detailed feature map :

[0037] Among them, RAM is a rotation attention mechanism function; In the field of crowd counting, the core of the task is pixel-level prediction regression based on density maps. This process highly depends on the capture of detailed information, especially important when dealing with occlusion problems in densely distributed crowds. Although network deepening and applying downsampling can enhance the model's ability to obtain high-order semantic information, this process inevitably leads to the loss of detailed information. Therefore, maintaining the detailed information of high-resolution features is crucial for the crowd counting task. Considering that there is more noise in shallow features, the DGCC-Net model introduces a local attention module for high-resolution features at scales of 1 / 8 and 1 / 16 before feature fusion. This module aims to accurately locate the crowd area and retain its detailed information. The local attention module consists of a multi-scale receptive field module and a rotational attention module. Among them, the multi-scale receptive field module MRFM is responsible for capturing rich context information, while the rotational attention module RAM retains and enhances detailed features that may be occluded in dense scenes by performing attention calculations in three dimensions.

[0038] The multi-scale receptive field module MRFM of this embodiment includes dilated convolutions with three different dilation rates (6, 12, 18) and a standard convolutional layer. Through these four convolutional operations, context information is extracted at different scales, thereby generating four features with different scales. To enhance the regression ability of global context information, MRFM introduces a global average pooling layer based on the convolutional operations, and uses a 1×1 convolutional layer and bilinear interpolation to restore the features to the original resolution. Finally, the output of MRFM is obtained by adding the features.

[0039] In this embodiment, the rotational attention module RAM is introduced to improve the model's ability to capture detailed information. This module enhances the attention to the crowd area by dynamically adjusting the weights of the feature map, while strengthening the small features in the occluded area. Compared with the traditional spatial channel attention module, the triple attention mechanism realizes cross-dimensional feature interaction without dimensionality reduction, can effectively encode channel and spatial information, and has a lower computational cost. Therefore, the triple attention mechanism is introduced into the rotational attention module. RAM is first composed of a three-branch structure, which respectively captures the dependence relationship of the input tensor in the ( C , H ), ( C , W ), and ( H , W ) dimensions. Specifically, for the input tensor f a1 ∈R C×H×W , it is first rotated 90° counterclockwise along the H axis to obtain the rotated tensor f a1 '∈R W×H×C . The concatenation operation of max pooling and average pooling is used to form f a1'' ∈ R 2×H×C , after being processed by a 3×3 convolution and passing through the Sigmoid function to obtain attention weights W a1 ∈ R 1×H×C , and after element-wise multiplication with f a1 ' and then rotated clockwise, restored to the original dimension and output the features y a1 ∈ R C×H×W :

[0040]

[0041]

[0042] Among them, represents the Sigmoid activation function, and Conv represents the convolution operation. The second branch processes the input rotated along the W axis to produce the output feature y a2 . The third branch directly processes the original input and outputs the feature y a3 . The outputs of the three branches are added element-wise to form the final output of the three-branch structure y a , and the process is as follows:

[0043] Given that the crowd target usually occupies a continuous area, based on the three-branch structure, the present invention further proposes regional attention based on convolution with a specific receptive field to dynamically adjust the crowd detail features within the local area. For the output y a of the three-branch structure, a 5×5 convolution kernel is used to extract regional features, and then batch normalization is performed to stabilize the learning process and reduce internal covariate shift. The ReLU activation function is used to introduce non-linear processing. Then, a 1×1 convolution layer is used to reduce the dimension, and the Sigmoid function is used to generate feature weights. This weight is multiplied element-wise with the original feature map to obtain the output result of the local attention module.

[0044] In an alternative embodiment of the present invention, as Figure 4 shown, this embodiment fuses features at different levels through a multi-level feature fusion module and adopts a dynamic learning weight mechanism to achieve effective fusion of semantic and detail information.

[0045] This embodiment inputs the high-resolution feature maps F1', F2' and F3' into the multi-level feature fusion module, integrates all high-level features through a concatenation operation, and generates the fused feature :

[0046] Among them, is the dynamic learning weight, ; This embodiment adopts a multi-level supervision strategy, introducing additional density map supervision before feature fusion to improve the fine-grained control ability of the model.

[0047] This embodiment introduces a density guidance module DGM to implement supervised learning of the density guidance map, applying different degrees of attention to different density regions, so as to effectively identify population regions with different density distributions.

[0048] This embodiment generates a density guidance feature map by passing the fused feature map through the density guidance module, including: First, the density guidance map generation module (SEG) gradually reduces the number of channels of the fused feature map from 256 to 1 through three convolutional operations, and normalizes the output feature values to the range of 0 to 1 using an activation function to generate a density guidance map; Then, after multiplying the density guidance map element-wise with the original feature map, the local and global attention mechanisms are fused through two layers of Twins-Transformer basic blocks, and a density guidance feature map is generated through residual connection.

[0049] In an optional embodiment of the present invention, when pre-training the DGCC-Net model, the original images provided by the crowd counting dataset are preprocessed to generate crowd density map labels and density guidance map labels, including: Annotate the head positions in the crowd counting images, and map the crowd counting images into point annotation maps containing the number of people; Use the Gaussian function to perform Gaussian filtering on the annotated points to generate crowd density map labels with Gaussian distribution; Generate density guidance map labels according to the nearest neighbor distance to determine the current crowd density level and the size of the head region.

[0050] This embodiment annotates the head positions in the crowd counting images and maps the crowd counting images into point annotation maps containing the number of people, including: Annotate the head positions in the original images, select a single pixel point in the head region and mark it as 1 to form a point annotation map; Map the annotated points in the point annotation map into the original image containing the number of people:

[0051] Among them, H(x) represents the point annotation map, N is the total number of people in the image, is at the position The indicator function at

[0052] In this embodiment, a Gaussian function is used to perform Gaussian filtering on the labeled points to generate a crowd density map label with a Gaussian distribution, including: Performing a Gaussian filtering operation with a Gaussian kernel size of on each labeled point:

[0053] where \(F(x)\) is the crowd density map, is a Gaussian function with a standard deviation of ; where \(F(x)\) is the crowd density map, is a Gaussian function with a standard deviation of ; In this embodiment, the current crowd density level and the size of the human head region are determined according to the nearest neighbor distance to generate a density guidance map label, including: Calculating the average distance between the nearest neighbor marked points for each labeled point to estimate the local density degree :

[0054] where is the distance between the \(i\)-th i human head marked point and its \(j\)-th j nearest neighbor point, the correction coefficient is set to 0.1, m is set to 3; Based on the calculated local density degree , assign a corresponding density level to each human head marked point in the image, and discretize the calculated local density degree into three predefined density levels:

[0055] where is the density level of the \(i\)-th i human head marked point; Based on the local density degree of each human head marked point, use a circular region to simulate the human head region, thereby forming the corresponding density distribution on the density guidance map. By adjusting the radius of the circular region, the distribution characteristics of the crowd at different density levels are reflected. Referring to the setting of the standard deviation of the Gaussian function in the preprocessing of the crowd density map, establish a matching relationship between the region radius and the local density level:

[0056] where size i is the \(i\)-thi The radius of the circular area of the personal head marking point determines the distribution area of the crowd; L i For the i local density degree of the personal head marking point, min is the minimum value function, and max is the maximum value function. e is the scientific notation symbol.

[0057] In an alternative embodiment of the present invention, when pre-training the DGCC-Net model in this embodiment, a multi-level supervision loss function is constructed to perform multi-level supervision on the three-scale branches and the final prediction branch, realizing fine-grained control of different levels.

[0058] This embodiment supervises the density-guided supervision method: The mean squared error loss function is adopted for the supervision of the density guidance map:

[0059] Among them, is the density guidance loss function, is the predicted pixel value, is the true label pixel value, and N is the total number of pixels; This embodiment uses the structural similarity loss function to perform structural similarity supervision on the dense area:

[0060] Among them, is the mean of the pixel values in the ground truth label map, is the mean of the pixel values in the predicted density map, is the covariance between the pixel values of the ground truth label map and the predicted density map, is the variance of the pixel values in the ground truth label map, is the variance of the pixel values in the predicted density map, is a set constant.

[0061] This embodiment introduces the total variation loss function to reduce overly smooth predictions and improve the adaptability of the DGCC-Net model to different density distributions:

[0062] Among them, is the total variation loss function, is the pixel value of the predicted density map at position (i,j); This embodiment combines the multi-level supervision mechanism to construct the final multi-level supervision loss function, including density guidance loss, structural similarity loss, and total variation loss:

[0063]

[0064]

[0065]

[0066] Among them, is the multi-level supervision loss function, is the weight of the i th layer, is the structure loss function of the i th layer, is the balance coefficient, is the total variable loss function of the i th layer, is the density guidance loss function of the i th layer, is the mean of the pixel values in the ground truth label map, is the mean of the pixel values in the predicted density map, is the covariance between the pixel values of the ground truth label map and the predicted density map, is the variance of the pixel values in the ground truth label map, is the variance of the pixel values in the predicted density map, is a set constant, is the pixel value of the predicted density map at position (i, j), is the predicted pixel value, is the true label pixel value, and N is the total number of pixels.

[0067] In this embodiment, i takes 0, 1, 2, 3, representing the final prediction branch and the three hierarchical branches respectively. By introducing a multi-level supervision mechanism, fine-grained control of different levels is achieved, and the weight of the final prediction branch is set to 1, while the weights of the other three branches are set to 0.5. To ensure the balance between loss functions and the consistency of each loss in magnitude, the weight coefficient is set to 0.01.

[0068] In an alternative embodiment of the present invention, this embodiment generates a training dataset and uses the training dataset to train the network DGCC-Net model, including: Randomly cropping and horizontally flipping the original image for data augmentation, and the augmented image retains the original annotation information to form a training dataset; Using a geometric adaptive method based on the K-means algorithm to generate a crowd density ground truth label map, and creating a density guidance label map by the density guidance map generation algorithm generated in the previous step; For low-resolution images, the size of the random crop is set to 256×256, and for other crowd counting datasets, the size of the random crop is uniformly set to 512×512; In particular, for datasets with fewer images, a 5-fold cross-validation method is adopted for model training to improve the generalization ability and verification effect of the model.

[0069] In this embodiment, the AdamW optimizer is used to optimize the network parameters, including: Setting the learning rate and weight decay coefficient of the AdamW optimizer. The AdamW optimizer optimizes the network parameters through weight decay and momentum correction:

[0070] Among them, is the parameter value of the t-th iteration, is the learning rate, and are the momentum and acceleration terms respectively, is a small constant to prevent the denominator from being zero, is the weight decay coefficient; In this embodiment, the DGCC-Net model is trained, including: Inputting the enhanced training dataset into the DGCC-Net model for feature extraction, density map generation, and density guidance; Adopting a random crop operation to randomly crop the input image into image patches of a fixed size to increase the diversity of data and the robustness of the model; Adopting a horizontal flip operation to horizontally flip the input image with a probability of 50% to further increase the diversity of the training data; Calculating the loss by comparing the predicted density map and the real density map output by the model through the multi-level supervised loss function constructed above, and adjusting the model parameters; Repeating the above training process until the error of the model on the validation set reaches the preset threshold or the number of training iterations reaches the predetermined upper limit.

[0071] Next, the present invention analyzes the performance of a crowd counting method provided by the present invention based on the DGCC-Net model in combination with specific examples.

[0072] DGCC-Net is trained and tested in a hardware environment configured with 48GB Nvidia A40 graphics cards, using the Python 3.8 programming language and the PyTorch 1.12.1 deep learning framework, and accelerated based on CUDA 11.7. The feature extraction network of the model adopts the weights of the Twins-PCPVT-L model pre-trained on the ImageNet-1k dataset, and the remaining convolutional layers and fully connected layers use randomly initialized weights with a mean of 0 and a variance of 0.02. During the training process, the AdamW optimizer is used for learning, where the learning rate is set to 1×10-5, the weight decay is set to 1×10-4, and a dropout rate of 0.45 is used to optimize the Transformer feature extraction network. The batch size is set to 4 or 8 according to the characteristics of different datasets, and at the same time, random cropping and horizontal flipping with a random rate of 0.5 are used for data augmentation.

[0073] For the preprocessing part of the crowd counting dataset, the image size is uniformly adjusted to ensure that the longest side of each image is less than 1920 pixels and the shortest side is greater than 512 pixels. For the label data, a geometric adaptive method based on the K-means algorithm is used to generate the ground truth label map of crowd density, and the density guidance label map is created by the density guidance map generation algorithm proposed in the present invention. For the SHA dataset, since the image resolution of this dataset is generally low, the size of the random cropping is set to 256×256. For other crowd counting datasets, the size of the random cropping is uniformly set to 512×512. In particular, for the UCF-CC-50 dataset, considering the small number of images it contains, a 5-fold cross-validation method is adopted for model training to improve the generalization ability and verification effect of the model.

[0074] The DGCC-Net model was evaluated on the ShanghaiTech Part A dataset and compared with other leading crowd counting methods, as shown in Table 1. The core technical features of each algorithm include: MCNN uses a multi-column neural network to extract crowd features at different scales at the front end of the network; CSRNet uses parallel dilated convolution modules at the back end of the network to enhance the feature expression ability; BL and DM-Count respectively use Bayesian loss function and optimal transport loss function to reduce the generation error of pseudo crowd density maps; RAN generates a priority map containing prior information of crowd regions by introducing a feedback network with region awareness ability, improving the model's attention to crowd regions; SGANet combines a feature extraction network with Inception-v3 as the backbone and introduces an attention branch guided by a binary segmentation map to improve the extraction efficiency of crowd features; CrowdFormer uses a top-down visual perception mechanism and Transformer network architecture to regress crowd density maps and introduces a learnable density generation module; CFANet uses two types of coarse and fine-grained crowd attention maps to assist the model in learning different crowd regions; Gramformer introduces a graph-guided attention adjustment mechanism to achieve diverse attention to complementary information and adjusts input features according to centrality metrics; STEERER achieves precise fusion of low-resolution features to high-resolution features through selective inheritance learning method and performs selective supervision at different scales.

[0075] As shown in Table 1, DGCC-Net achieved the best counting results among all the listed algorithms. Compared with the second-place STEERER, the MAE and MSE were improved by 1.5% respectively. DGCC-Net adopted a more refined density-guided mask. Compared with SGANet using a binary segmentation mask, the MAE and MSE were improved by 6.7% and 15.3% respectively; compared with CFANet using two types of multi-level mask-assisted supervision, the MAE and MSE were reduced by 2.4 and 4.0 respectively, demonstrating the effectiveness of the density-guided module at the crowd attention level; compared with CrowdFormer using the Transformer architecture, the MAE and MSE were reduced by 3.2 and 11.8 respectively, verifying the efficiency of the model's multi-level feature fusion structure.

[0076] Table 1 Experimental Results on ShanghaiTech A Dataset

[0077] Figure 6Shows the visual prediction results of DGCC-Net on the ShanghaiTech Part A dataset. From left to right, they represent the original image, the annotated image, and the predicted image. Among them, GT represents the actual number of people, and Pred represents the predicted number of people. It can be seen from the figure that DGCC-Net can better focus on the crowd, realize the learning of different crowd density distributions of dense, sparse, and independent pedestrians, and can accurately count.

[0078] The DGCC-Net model was evaluated on the UCF-QNRF dataset and compared with other leading crowd counting methods, as shown in Table 2. Among them, P2PNet introduced the Hungarian matching algorithm into crowd counting to achieve end-to-end point annotation map prediction, providing a new solution for accurate counting and localization; CLTR combined the advantages of CNN and Transformer networks and used the Hungarian matching algorithm improved based on the nearest neighbor distance to achieve the regression of the point-level crowd prediction map; SASNet adopted a scale-adaptive feature layer selection strategy to select local patches from the most suitable feature layer for prediction, effectively using multi-scale feature representation; while S-DCNet (dcreg) introduced discrete constraint regression into the model, extended the regression problem from the continuous scale to discrete ranking, and achieved a more accurate counting effect.

[0079] The UCF-QNRF dataset is a dataset with dense crowds and large resolutions. The experimental results in Table 2 further prove the excellent performance of DGCC-Net in the crowd counting task. Compared with other algorithms, DGCC-Net achieved the best results in MAE and was also at a leading level in MSE. In particular, compared with the RAN model, the MAE of DGCC-Net decreased by 1.3; compared with the CLTR model that combined CNN and Transformer technologies, the MAE decreased by 3.7; compared with CFANet supervised by two fine-grained masks, DGCC-Net increased by 7.8% and 1.2% in MAE and MSE respectively. This result not only shows the advantages of DGCC-Net in dealing with high-density crowd scenes but also highlights the effectiveness of its density guidance and multi-level feature fusion strategy. Figure 7 Further shows the visual prediction effect of DGCC-Net on the UCF-QNRF dataset. It can be clearly seen from it that DGCC-Net accurately captures the crowd density and can accurately count in crowded scenes. This result confirms the leading counting level of DGCC-Net in the field of crowd counting, especially its ability to handle large-scale and high-density crowd scenes.

[0080] Table 2 Experimental results on the UCF-QNRF dataset

[0081] The present invention uses Twins Transformer as the feature extraction network. By leveraging its powerful semantic information acquisition ability demonstrated globally, compared with traditional convolutional neural networks, it can better handle long-range dependence relationships and global context information, significantly enhancing the model's feature extraction ability in complex scenarios. This enables it to effectively capture fine-grained feature information even in high-density crowd scenarios, improving the accuracy of overall counting.

[0082] Aiming at the defect that the existing crowd counting methods have poor performance in dealing with the problem of uneven density distribution, the present invention proposes a density guidance map generated based on the nearest neighbor distance. By introducing a density guidance module and using the density guidance map to pay differential attention to different density regions, the model can determine the crowd density and the size of the human head region according to the nearest neighbor distance of the human head points, generating a density guidance map label, effectively solving the problem of uneven crowd density distribution and making the model more robust in dealing with scenarios with different density distributions.

[0083] The present invention introduces a local attention module at the high-resolution scale, including a multi-scale receptive field module and a rotational attention module. The multi-scale receptive field module captures multi-scale context information through parallel dilated convolutions with different receptive fields, while the rotational attention module enhances the attention to the crowd region by dynamically adjusting the feature map weights, significantly enhancing the ability to capture detailed features. Through a multi-level feature fusion module, features at different levels are fused, and a dynamic learning weight mechanism is adopted to achieve effective fusion of semantic and detailed information, enabling the model to maintain a high counting accuracy when dealing with high-density and sparse crowd scenarios.

[0084] The present invention constructs a multi-level supervision loss function, including density guidance loss, structural similarity loss, and total variation loss. The density guidance loss is used for the supervision of the density guidance map. The structural similarity loss selects the dense crowd region through a dense region mask to ensure the prediction quality of the dense region, while the total variation loss aims to reduce overly smooth predictions and enhance the model's adaptability to different density distributions. By introducing a multi-level supervision mechanism, the model can perform fine supervision at different levels, improving the counting accuracy and robustness for scenarios with different density distributions.

[0085] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0088] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0089] Those of ordinary skill in the art will realize that the embodiments described here are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A crowd counting method based on the DGCC-Net model, characterized in that: The following steps are involved: Get crowd counting images; Constructing a DGCC-Net model and performing pre-training; the DGCC-Net model includes a feature extraction network, a local attention module, a multi-level feature fusion network and a prediction network; Use feature extraction network to extract multi-scale feature maps from crowd counting images; Use the local attention module to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map to obtain a high-resolution feature map; A multi-level feature fusion network is used to fuse the semantic and detail information of the feature maps at low resolution scales and high resolution feature maps in the multi-scale feature map, and the fused feature map is passed through a density guidance module to generate a density-guided feature map; The prediction network is used to predict the crowd counting results based on the density-guided feature map.

2. According to a crowd counting method based on the DGCC-Net model according to claim 1, it is characterized in that: The local attention module is used to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map to obtain a high-resolution feature map, including: The multi-scale receptive field module included in the local attention module is used to extract multi-scale context information through parallel dilated convolutions with different receptive fields to generate a feature map containing multi-scale context information. The rotation attention module contained in the local attention module is used to dynamically adjust the feature map weights to extract the detail information of the crowd area from the feature map containing multi-scale contextual information to obtain a high-resolution feature map.

3. According to a crowd counting method based on the DGCC-Net model according to claim 2, it is characterized in that: The multi-scale receptive field module included in the local attention module is used to extract multi-scale context information through parallel dilated convolutions with different receptive fields to generate a feature map containing multi-scale context information, including: The first atrous convolution, the second atrous convolution, and the third atrous convolution with three different atrous rates in the multi-scale receptive field module and the first convolution layer are used to extract context information at different scales from the input feature map to generate feature maps of four different scales. The first global average pooling layer included in the multi-scale receptive field module is used to perform a global average pooling operation on the input feature map, and the feature map is restored to the original resolution through the second convolution layer and bilinear interpolation operation to obtain a global average pooling feature map; The first concatenation layer contained in the multi-scale receptive field module is used to concatenate the feature maps of four different scales and the global average pooling feature map to generate a feature map containing multi-scale context information.

4. According to a crowd counting method based on the DGCC-Net model according to claim 2, it is characterized in that: The rotation attention module included in the local attention module is used to dynamically adjust the feature map weights to extract the detailed information of the crowd area from the feature map containing multi-scale context information, and obtain a high-resolution feature map, including: The input feature map is rotated 90° counterclockwise along the height axis by using the rotation attention module to obtain the first rotation tensor, and then the first concatenated feature map is formed by the splicing operation of maximum pooling and average pooling. Then, the convolution operation is performed through the third convolution layer and the first attention weight is obtained through the activation function. Finally, the first attention weight is element-wise multiplied with the input feature map and rotated 90° clockwise along the height axis to obtain the first rotated attention feature map. The input feature map is rotated 90° counterclockwise along the width axis by using the rotation attention module to obtain the second rotation tensor, and then the second concatenated feature map is formed by the splicing operation of maximum pooling and average pooling. Then, the convolution operation is performed through the fourth convolution layer and the second attention weight is obtained by the activation function. Finally, the second attention weight is element-wise multiplied with the input feature map and rotated 90° clockwise along the width axis to obtain the second rotated attention feature map. The rotation attention module is used to concatenate the input feature map through the maximum pooling and average pooling to form a third concatenated feature map, and then the convolution operation is performed through the fifth convolution layer and the third attention weight is obtained through the activation function. Finally, the third attention weight is element-wise multiplied with the input feature map to obtain the third rotated attention feature map; The rotational attention module is used to add the first rotational attention feature map, the second rotational attention feature map and the third rotational attention feature map by elements to obtain a high-resolution feature map.

5. The crowd counting method based on the DGCC-Net model according to claim 1, characterized in that: The local attention module is used to extract context information and detailed information of the crowd area from the feature map at the high-resolution scale in the multi-scale feature map to obtain a high-resolution feature map, which also includes: The high-resolution feature map is passed through the sixth convolutional layer to extract regional features and perform batch normalization. It is then convolved through the seventh convolutional layer and the feature map weights are generated through the activation function. Finally, the feature map weights are multiplied element-by-element with the high-resolution feature map to obtain the output feature map of the local attention module.

6. The crowd counting method based on the DGCC-Net model according to claim 1, characterized in that: The fusion feature map is passed through the density guidance module to generate a density-guided feature map including: The number of channels of the fusion feature map is gradually reduced from 256 to 1, and the output feature value is normalized to the range of 0 to 1 using the activation function to generate a density guidance map; After element-wise multiplication of the density-guided map and the original feature map, the local and global attention mechanisms are fused through two layers of Twins-Transformer basic blocks to generate a density-guided feature map.

7. The crowd counting method based on the DGCC-Net model according to claim 1, characterized in that: When pre-training the DGCC-Net model, Mark the head positions in the crowd counting image and map the crowd counting image into a point annotation map containing the number of people; Use the Gaussian function to perform Gaussian filtering on the marked points to generate crowd density map labels with Gaussian distribution; The current crowd density level and head area size are determined based on the nearest neighbor distance, and a density guidance map label is generated.

8. The crowd counting method based on the DGCC-Net model according to claim 7, characterized in that: Determine the current crowd density level and head area size based on the nearest neighbor distance, and generate density guide map labels, including: For each labeled point, calculate the average distance between the nearest neighboring labeled points and estimate the local density; Assign a corresponding density level to each head marker point in the image according to the local density of each head marker point; According to the local density of each head marker point, a circular area is used to simulate the head area to form the corresponding density distribution on the density guidance map. The distribution characteristics of the crowd at different density levels are characterized by adjusting the radius of the circular area, and the matching relationship between the circular area radius and the local density level is established to generate the density guidance map label.

9. A crowd counting method based on the DGCC-Net model according to claim 8, characterized in that: The matching relationship between the circular area radius and the local density level is: in, size i For the i The radius of the circular area of ​​the individual head marker points, L i For the i The local density of individual head markers, min is the minimum function, and max is the maximum function. e is the scientific notation symbol.

10. A crowd counting method based on the DGCC-Net model according to claim 1, characterized in that: When pre-training the DGCC-Net model, a multi-level supervision loss function is constructed to perform multi-level supervision on the three scale branches and the final prediction branch, which is expressed as: in, is the multi-level supervision loss function, For the i The weight of the level, For the i The hierarchical structure loss function, is the balance coefficient, For the i The total variable loss function of the layer, For the i The density-guided loss function of the layer, is the mean of the pixel values ​​in the true value label image, is the mean of the pixel values ​​in the predicted density map, is the covariance between the pixel values ​​of the true value label map and the predicted density map, is the variance of the pixel values ​​in the true label image, is the variance of pixel values ​​in the predicted density map, To set the constant, To predict the pixel value of the density map at position (m,n), To predict pixel values, is the true label pixel value, and N is the total number of pixels.

Citation Information

Cited By

  • Multi-dimensional dynamic perception and progressive focusing network model for crowd counting

    CN120976848A

  • Method for constructing multi-dimensional dynamic perception and progressive focusing network model for crowd counting

    CN120976848B

  • Fish school counting method based on multi-scale double-branch joint training network

    CN121545183A

  • Crowd density estimation method and system based on feature perception weighted contrast learning

    CN121617047A