A region complement-based population counting method and related device
By using multi-scale feature fusion and multiple deep supervision through a regional complementary aggregation network model, the problem of low accuracy in crowd counting is solved, and higher accuracy in crowd counting is achieved.
Patent Information
- Application Number
- CN202211384415.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-11-07
AI Technical Summary
Existing crowd counting techniques suffer from low accuracy, especially affected by scale variations and background noise interference.
A population counting method based on region complement aggregation is adopted. Through a pre-trained region complement aggregation network model, the correlation between multi-scale features and regions in the image is used to perform bidirectional iterative fusion and weighted complementary concatenation to generate a coarse density map. Then, the attention map is determined through multiple deep supervision, and finally a fine density map is output.
It improves the accuracy of crowd counting results, enhances the image quality of fine density maps, and effectively solves the problems of multi-scale and background noise interference.
Smart Images

Figure CN115731511B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology in computer vision, and in particular to a crowd counting method and related apparatus based on region complementary aggregation. Background Technology
[0002] With the increase in urban population and the advancement of urbanization, crowd counting has received widespread attention due to its important role in social security and intelligent transportation. The research content of crowd counting involves using computers to analyze a given image or video to determine the number of people present in the image or video, or to obtain the crowd density distribution within the image. Crowd counting tasks still face challenges such as scale variations and background noise interference, which contribute to the low accuracy of crowd counting.
[0003] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide a population counting method and related apparatus based on regional complementary aggregation, which addresses the shortcomings of the existing technology.
[0005] To address the aforementioned technical problems, a first aspect of this application provides a population counting method based on regional complementary aggregation, the method comprising:
[0006] Acquire images of dense crowds;
[0007] The dense crowd image is input into the feature module of a pre-trained region complementarity aggregation network model, and the feature module outputs several feature maps corresponding to the dense crowd image.
[0008] The aforementioned feature maps are input into the complementary iterative aggregation module of the region complementary aggregation network model, and a coarse density map is output through the complementary iterative aggregation module.
[0009] The aforementioned feature maps are input into the region localization module of the region complementary aggregation network model, and the attention map is output through the mutual region localization module.
[0010] The coarse density map and the attention map are input into the fusion module of the region complementary aggregation network model, and the fine density map is output through the fusion module.
[0011] The crowd count result corresponding to the dense crowd image is determined based on the fine density map.
[0012] The population counting method based on regional complementary aggregation, wherein the complementary iterative aggregation module includes two iterative fusion branches and a first fusion branch; the step of inputting the plurality of feature maps into the complementary iterative aggregation module of the regional complementary aggregation network model, and outputting a coarse density map through the complementary iterative aggregation module specifically includes:
[0013] The aforementioned feature maps are respectively input into two iterative fusion branches. One iterative fusion branch fuses the feature maps in a bottom-up order to obtain a first fusion map, and the other iterative fusion branch fuses the feature maps in a top-down order to obtain a second fusion map. The iterative fusion branch includes several cascaded weighted complementary cascaded modules.
[0014] The first fusion map and the second fusion map are input into the first fusion branch, and a coarse density map is output through the fusion unit.
[0015] The population counting method based on regional complementary aggregation, wherein the regional localization module includes two localization branches and a second fusion branch; the step of inputting the plurality of feature maps into the regional complementary aggregation network model and outputting an attention map through the regional localization module specifically includes:
[0016] The feature maps are input into two positioning branches respectively. The feature maps are fused in a bottom-up order through one positioning branch to obtain a first positioning map. The feature maps are fused in a top-down order through the other positioning branch to obtain a second positioning map. The positioning branch includes several cascaded weighted complementary cascaded modules.
[0017] The first positioning map and the second positioning map are input into the second fusion branch, and the attention map is output through the fusion unit.
[0018] The population counting method based on regional complementary aggregation includes a positioning branch comprising a plurality of positioning units cascaded sequentially. Each positioning unit includes a weighted complementary cascade module and a partial attention integration module. The weighted complementary cascade module is connected to the partial attention integration module. Both the weighted complementary cascade module and the partial attention integration module are connected to the partial attention integration module in the preceding positioning unit. The foremost partial attention integration module in the positioning branch used to determine the first positioning map is connected to the last partial attention integration module in the positioning branch used to determine the second positioning map.
[0019] The population counting method based on region complementary aggregation includes a weighted complementary cascade module comprising a first convolutional block, a first scaling layer, a second convolutional block, a second scaling layer, complementary units, a first multiplier, a second multiplier, an adder, a third convolutional block, a first connection layer, and a fourth convolutional block. The first convolutional block is connected to the first scaling layer, the second convolutional block is connected to the second scaling layer, both the first and second scaling layers are connected to complementary units, both the first scaling layer and the complementary unit are connected to the first multiplier, and both the second scaling layer and the complementary unit are connected to the second multiplier. The first and second multipliers are both connected to the adder, the adder is connected to the third convolutional block, the first scaling layer, the second scaling layer, and the third convolutional block are all connected to the first connecting layer, and the first connecting layer is connected to the fourth convolutional block. The complementary unit includes a first global average pooling layer, a second global average pooling layer, a second connecting layer, and an activation function layer. The first global average pooling layer is connected to the first scaling layer and the second connecting layer, the second global average pooling layer is connected to the second scaling layer and the second connecting layer, and the second connecting layer is connected to the activation function layer.
[0020] The crowd counting method based on region complementary aggregation includes a partial attention integration module comprising a fifth convolutional block, a third scaling layer, a sixth convolutional block, a fourth scaling layer, an activation function layer, and a fusion layer. The fifth convolutional block is connected to the third scaling layer, the third scaling layer is connected to the activation function layer, the sixth convolutional block is connected to the fourth scaling layer, and both the activation function layer and the fourth scaling layer are connected to the fusion layer.
[0021] The crowd counting method based on regional complementary aggregation includes a loss function used during the training of the regional complementary aggregation network model. This loss function includes a supervised loss term determined by the predicted density map and the ground truth value of the density map output by the regional complementary aggregation network model, a supervised loss term determined by the attention map output by the regional localization module and the ground truth value of the attention map, and a multi-depth supervised loss term determined by the partial attention maps output by the partial attention integration module in the regional localization module and the ground truth value of the attention maps. The attention map value is obtained by processing the ground truth value of the density map through a threshold function.
[0022] A second aspect of this application provides a population counting system based on regional complementary aggregation, the system comprising:
[0023] The acquisition module is used to acquire images of dense crowds;
[0024] The control module is used to input the dense crowd image into a pre-trained region complementary aggregation network model to determine a fine density map. Specifically, the process of determining the fine density map involves: outputting several feature maps corresponding to the dense crowd image through the feature module; inputting the several feature maps into the complementary iterative aggregation module of the region complementary aggregation network model to output a coarse density map; inputting the several feature maps into the region localization module of the region complementary aggregation network model to output an attention map; and inputting the coarse density map and the attention map into the fusion module of the region complementary aggregation network model to output a fine density map.
[0025] The determination module is used to determine the crowd count result corresponding to the dense crowd image based on the fine density map.
[0026] A third aspect of this application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the population counting method based on regional complementary aggregation as described above.
[0027] A fourth aspect of this application provides a terminal device, which includes: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;
[0028] The communication bus enables communication between the processor and the memory;
[0029] When the processor executes the computer-readable program, it implements the steps in the population counting method based on region complementary aggregation as described above.
[0030] Beneficial Effects: Compared with the prior art, this application provides a crowd counting method and related apparatus based on region complementary aggregation. The method includes: acquiring a dense crowd image; inputting the dense crowd image into the feature module of a pre-trained region complementary aggregation network model, and outputting several feature maps corresponding to the dense crowd image through the feature module; inputting the several feature maps into the complementary iterative aggregation module of the region complementary aggregation network model, and outputting a coarse density map through the complementary iterative aggregation module; inputting the several feature maps into the region localization module of the region complementary aggregation network model, and outputting an attention map through the region localization module; inputting the coarse density map and the attention map into the fusion module of the region complementary aggregation network model, and outputting a fine density map through the fusion module; and determining the crowd counting result corresponding to the dense crowd image based on the fine density map. This application generates a coarse density map through bidirectional iterative fusion and weighted complementary cascading of a complementary iterative aggregation module, and determines an attention map through multiple depth supervision using a region localization module. Then, a fine density map is determined based on the attention map and the coarse density map. In this way, multi-scale fusion through the complementary iterative aggregation module and multiple depth supervision through the region localization module effectively improves the image quality of the fine density map, thereby improving the accuracy of determining the crowd counting results based on the fine density map. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 A flowchart of the population counting method based on regional complementary aggregation provided in this application.
[0033] Figure 2 The flowchart illustrates the principle of the population counting method based on regional complementary aggregation provided in this application.
[0034] Figure 3 This is a schematic diagram of the complementary iterative aggregation module.
[0035] Figure 4 This is a schematic diagram of the structure of a weighted complementary cascaded module.
[0036] Figure 5 This is a schematic diagram of the structural principle of the regional positioning module.
[0037] Figure 6 This is a schematic diagram of the structural principle of a partial attention integration module.
[0038] Figure 7 The structural principle diagram of the population counting system based on regional complementary aggregation provided in this application.
[0039] Figure 8 A schematic diagram of the terminal device provided in this application. Detailed Implementation
[0040] This application provides a population counting method and related apparatus based on regional complementary aggregation. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0041] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0042] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0043] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.
[0044] The inventors discovered through research that with the increase in urban population and the advancement of urbanization, crowd counting has received widespread attention due to its important role in social security and intelligent transportation. The research content of crowd counting involves using computers to analyze a given image or video to determine the number of people present in the image or video, or to obtain the crowd density distribution within the image. Crowd counting tasks still face challenges such as scale variations and background noise interference, which are precisely the factors affecting the accuracy of crowd counting.
[0045] The primary cause of the multi-scale problem is the perspective effect. A key characteristic of the perspective effect is that objects closer to the camera appear larger, while those farther away appear smaller. Therefore, the scale of a pedestrian is related to their location within the image. When an image is input into a network, different layers exhibit varying sensitivities to scale features, resulting in inconsistent high-activation regions in their feature maps. These high-activation regions also demonstrate significant differences and complementarity.
[0046] Therefore, when addressing the low accuracy of crowd counting caused by multi-scale issues, the correlation between multi-scale features and regions in the image can be used as a starting point, based on the differences and complementarities of the regions containing the multi-scale features. Thus, in this embodiment, after acquiring a dense crowd image, the dense crowd image is input into the feature module of a pre-trained region complementarity aggregation network model, which outputs several feature maps corresponding to the dense crowd image. These feature maps are then input into the complementary iterative aggregation module of the region complementarity aggregation network model, which outputs a coarse density map. The feature maps are then input into the region localization module of the region complementarity aggregation network model, which outputs an attention map. The coarse density map and the attention map are input into the fusion module of the region complementarity aggregation network model, which outputs a fine density map. Based on the fine density map, the crowd counting result corresponding to the dense crowd image is determined. This application embodiment generates a coarse density map through bidirectional iterative fusion and weighted complementary cascading of a complementary iterative aggregation module, and determines an attention map through multiple depth supervision by a region localization module. Then, a fine density map is determined based on the attention and the coarse density map. In this way, multi-scale fusion through the complementary iterative aggregation module and multiple depth supervision through the region localization module effectively improve the image quality of the fine density map, thereby improving the accuracy of determining the crowd counting results based on the fine density map.
[0047] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.
[0048] This embodiment provides a population counting method based on regional complementary aggregation, such as... Figure 1 and Figure 2 As shown, the method includes:
[0049] S10. Obtain images of dense crowds.
[0050] Specifically, the dense crowd image can be obtained by taking pictures of a preset location using an image acquisition device, such as an image of a bus stop during the morning rush hour; it can also be obtained through the network (e.g., Baidu), sent by an external device, or be a surveillance image taken by a monitoring device. Furthermore, the dense crowd image carries several images of people, and the size of the image area occupied by each person in the several images can be different; in other words, the size of the people in the several images can be different.
[0051] S20. Input the dense crowd image into the feature module of the pre-trained region complementary aggregation network model, and output several feature maps corresponding to the dense crowd image through the feature module.
[0052] Specifically, the regional complementary aggregation network model is a pre-trained network model. The fine density map of dense crowd images can be determined through the regional complementary aggregation network model. The regional complementary aggregation network model takes the correlation between multi-scale features and regions in the image as the starting point, and determines the fine density map based on the differences and complementarity of the regions where the multi-scale features are located.
[0053] The region complementary aggregation network model includes a feature module, a complementary iterative aggregation module, a region localization module, and a fusion module. The feature module extracts feature maps, the complementary iterative aggregation module determines a coarse density map, the region localization module determines an attention map, and the fusion module fuses the coarse density map and the attention map to obtain a fine density map. Therefore, after acquiring a dense crowd image, the image can be input into the feature module of the region complementary aggregation network model. The feature module outputs several feature maps, each with a different image size; in other words, the feature maps are outputs from different network layers within the feature module.
[0054] In one implementation, the feature module can employ a VGG-16 network to extract several feature maps corresponding to the dense crowd image. These feature maps can include features from the last layer of the four stages in the first 13 layers of the VGG-16 network; that is, the feature maps consist of four feature maps, representing the outputs of Conv2-2, Conv3-3, Conv4-3, and Conv5-3 of the VGG-16 network. It's worth noting that in practical applications, the feature module can also use other network models, and the number of feature maps can vary, such as five or six feature maps.
[0055] S30. Input the aforementioned feature maps into the complementary iterative aggregation module of the regional complementary aggregation network model, and output a coarse density map through the complementary iterative aggregation module.
[0056] Specifically, the complementary iterative aggregation module takes several feature maps as input and outputs a coarse density map. The module employs a bidirectional iterative fusion approach, consisting of bottom-up and top-down iterations. This bidirectional iterative fusion method allows the complementary iterative aggregation module to simultaneously acquire shallow spatial information and deep semantic information, balancing scale features across different layers and maintaining a balance between shallow spatial and deep semantic information. This, in turn, improves the image quality of the subsequently generated fine density map.
[0057] In one implementation, the complementary iterative aggregation module includes two iterative fusion branches and a first fusion branch; the step of inputting the plurality of feature maps into the complementary iterative aggregation module of the region complementary aggregation network model, and outputting a coarse density map through the complementary iterative aggregation module specifically includes:
[0058] S31. Input the several feature maps into two iterative fusion branches respectively. The several feature maps are fused in a bottom-up order through one iterative fusion branch to obtain a first fusion map. The several feature maps are fused in a top-down order through the other iterative fusion branch to obtain a second fusion map. The iterative fusion branch includes several cascaded weighted complementary cascaded modules.
[0059] S32. Input the first fusion map and the second fusion map into the first fusion branch, and output a coarse density map through the fusion unit.
[0060] Specifically, both iterative fusion branches are connected to the first fusion branch. One of the two iterative fusion branches performs iterative fusion in a bottom-up order, while the other performs iterative fusion in a top-down order, thus achieving bidirectional iterative fusion. After bidirectional iterative fusion through the two iterative fusion branches, the first and second fusion maps from the bidirectional iterative fusion are fused through the first fusion branch to obtain a coarse density map.
[0061] The iterative fusion branch includes several cascaded weighted complementary modules. In any two adjacent cascaded weighted complementary modules, the output of the preceding module serves as the input of the following module. Each input module includes one feature map, with the first module containing two feature maps. Furthermore, the feature maps included in the inputs of each module are distinct. Therefore, the number of cascaded weighted complementary modules is one less than the number of feature maps. For example, if there are four feature maps, then there are three weighted complementary modules.
[0062] Furthermore, in the iterative fusion branch that performs iterative fusion in a bottom-up order, the feature map A included in the input of the weighted complementary cascade module located earlier in the cascade order is the lower-level feature map of the feature map B included in the input of the weighted complementary cascade module located later in the cascade order. In other words, the image size of feature map A is smaller than the image size of feature map B. Conversely, in the iterative fusion branch that performs iterative fusion in a top-down order, the feature map A included in the input of the weighted complementary cascade module located earlier in the cascade order is the top-level feature map of the feature map B included in the input of the weighted complementary cascade module located later in the cascade order. In other words, the image size of feature map A is larger than the image size of feature map B.
[0063] For example: Suppose several feature maps are Conv2-2, Conv3-3, Conv4-3, and Conv5-3 of VGG-16, denoted as f1, f2, f3, and f4 respectively. Figure 3As shown, in the iterative fusion branch that performs iterative fusion in a top-down and bottom-up order, the inputs of the first weighted complementary cascade module are f1 and f2, the inputs of the second weighted complementary cascade module are the output of the first weighted complementary cascade module and f3, and the inputs of the last weighted complementary cascade module are the output of the second weighted complementary cascade module and f4. In the iterative fusion branch that performs iterative fusion in a bottom-up order, the inputs of the first weighted complementary cascade module are f3 and f4, the inputs of the second weighted complementary cascade module are the output of the first weighted complementary cascade module and f2, and the inputs of the last weighted complementary cascade module are the output of the second weighted complementary cascade module and f1. Therefore, the process of determining the coarse density map can be expressed as:
[0064] F ↑ =M wcc (M wcc (M wcc (f1,f2),f3),f4)
[0065] F ↓ =M wcc (M wcc (M wcc (f4,f3),f2),f1)
[0066] F CIANet =Cat(Up(F) ↑ ),F ↓ )
[0067] Among them, f k ,k∈{1,2,3,4} represent Conv2-2, Conv3-3, Conv4-3 and Conv5-3 of feature module VGG-16 respectively; F ↑ F represents the output term of the iterative fusion branch that performs iterative fusion in a bottom-up order. ↓ This represents the output of the iterative fusion branch, which is iteratively fused in a top-down order; Cat represents cross-channel cascading, Up represents upsampling operation, and F... CIANet This represents a rough density map.
[0068] In one implementation, since the feature map resolutions of the output terms of the iterative fusion branch following a top-down order and the iterative fusion branch following a bottom-up order are different, when fusing the output terms from the two directions, the output terms of the iterative fusion branch following a top-down order with lower resolution are first upsampled using bilinear interpolation to be consistent with Conv2-2. By upsampling the output terms with lower resolution, the resolution of the density map can be guaranteed to maintain more detailed spatial information, and the spatial information loss caused by downsampling the output terms with higher resolution can be avoided, thereby improving the image quality of the generated fine density map.
[0069] like Figure 4 As shown, the weighted complementary cascaded module includes a first convolutional block, a first scaling layer, a second convolutional block, a second scaling layer, a complementary unit, a first multiplier, a second multiplier, an adder, a third convolutional block, a first connecting layer, and a fourth convolutional block. The first convolutional block is connected to the first scaling layer, the second convolutional block is connected to the second scaling layer, both the first and second scaling layers are connected to the complementary unit, both the first scaling layer and the complementary unit are connected to the first multiplier, both the second scaling layer and the complementary unit are connected to the second multiplier, both the first and second multipliers are connected to the adder, the adder is connected to the third convolutional block, the first scaling layer, the second scaling layer, and the third convolutional block are all connected to the first connecting layer, and the first connecting layer is connected to the fourth convolutional block. The complementary unit includes a first global average pooling layer, a second global average pooling layer, a second connection layer, and an activation function layer. The first global average pooling layer is connected to the first scaling layer and the second connection layer, the second global average pooling layer is connected to the second scaling layer and the second connection layer, and the second connection layer is connected to the activation function layer.
[0070] Furthermore, the structures of the first, second, third, and fourth convolutional blocks are all identical; here, we will use the first convolutional block as an example for explanation. Figure 4 As shown, the first convolutional block employs a uniform convolutional layer and an activation function layer. The convolutional layer adjusts the number of channels in the input terms so that the number of channels in the output terms of the convolutional layers in the first convolutional block is greater than the number of channels in the output terms of the convolutional layers in the second convolutional block. In one specific implementation, the kernel size of the convolutional layer is 1, and the number of kernels is 128, so that the number of channels in the output terms of the convolutional layer is 128. The activation function layer uses the Threshold function as the activation function, where the formula for the Threshold function is:
[0071]
[0072] Where T is a hyperparameter, and is assigned a value of 0 when the pixel value in the feature map is less than T.
[0073] In multi-scale feature fusion, since different layers of the network do not only predict the most sensitive regions, the feature map contains not only high-response regions predicted by that layer but also large areas of low-response regions. These low-response regions have poor prediction results, and therefore become redundant information during feature fusion, being iterated over by the network. This results in the final multi-scale features containing a lot of noise, failing to achieve the desired improvement in counting accuracy. Therefore, this embodiment uses the Threshold function as the activation function, which can suppress non-sensitive regions while retaining highly sensitive regions during feature fusion, thereby generating a high-quality multi-scale fused map.
[0074] Both the first scaling layer and the second scaling layer are used to adjust the resolution of the feature map. The resolution of the output item of the first scaling layer is the same as the resolution of the output item of the second scaling layer. Both the first scaling layer and the second scaling layer are configured with scaling operations, which include bilinear interpolation upsampling, adaptive pooling downsampling, or no operation.
[0075] Furthermore, to aggregate useful information from the two feature maps and suppress redundant information, the complementary unit obtains the fusion weights of the two feature maps input to the complementary press through Global Average Pooling (GAP) and an activation function layer. In one implementation, the activation function layer is configured with a softmax function, and the formula for calculating the fusion weights can be expressed as:
[0076]
[0077] Among them, w i σ(w) represents the global feature value extracted from a channel in a feature map after global average pooling. i ) represents the weights after corresponding reweighting of the same channel in two feature maps.
[0078] After performing the above operations on all channels corresponding to the two input feature maps through the complementary unit, a complete reweighted weight map of the two feature maps is obtained. Then, the weight map is multiplied by the output of the first scaling layer through the first multiplier to obtain the first target feature map. The weight map is then multiplied by the output of the second scaling layer through the second multiplier to obtain the second target feature map. Next, the first target feature map output from the first multiplier and the second target feature map output from the second multiplier are input into an adder, and pixel-by-pixel addition is performed to obtain a weighted feature map. Then, the weighted feature map is adjusted for the number of channels using a third convolutional block, and the outputs of the first scaling layer, the second scaling layer, and the third convolutional block are input into the first connection layer. The first connection layer concatenates these three components across channels. Finally, the concatenated feature map is input into a fourth convolutional block, which outputs the weighted complementary concatenated module (WCC module). The weighted complementary concatenated module, through the process of generating and redistributing weights, can effectively extract and fuse the real sensitive regions in the feature map.
[0079] S40. Input the aforementioned feature maps into the region localization module of the region complementary aggregation network model, and output the attention map through the mutual region localization module.
[0080] Specifically, the region localization module focuses solely on distinguishing between foreground and background, using an attention map to pinpoint the foreground location in dense crowd images, thus mitigating background interference. Understandably, this attention map is used to address background noise in dense crowd images, improving the image quality of the subsequently obtained fine-density map. Furthermore, the region localization module takes the differences and complementarities of regions containing multi-scale features as its starting point, determining the attention map through multiple depth supervisions. This facilitates the gradual refinement of the attention map generation, resulting in a more accurate attention map.
[0081] In one implementation, the region localization module includes two localization branches and a second fusion branch; the process of inputting the plurality of feature maps into the region complementary aggregation network model and outputting an attention map through the mutual region localization module specifically includes:
[0082] S41. Input the aforementioned feature maps into two positioning branches respectively. Merge the feature maps in a bottom-up order through one positioning branch to obtain a first positioning map. Merge the feature maps in a top-down order through the other positioning branch to obtain a second positioning map. The positioning branch includes several cascaded weighted complementary cascaded modules.
[0083] S42. Input the first positioning map and the second positioning map into the second fusion branch, and output the attention map through the fusion unit.
[0084] Specifically, the two positioning branches in the regional positioning module employ bidirectional iterative fusion. One positioning branch performs iterative fusion in a bottom-up order, while the other performs iterative fusion in a top-up order. Furthermore, to enhance the connectivity of multiple depth supervision processes within the regional positioning module, the two positioning branches can be connected in series. For example, the positioning branch performing iterative fusion in a bottom-up order can be connected to the positioning branch performing iterative fusion in a top-down order, or vice versa.
[0085] The localization branch includes several cascaded localization units. Each localization unit includes a weighted complementary cascaded module and a partial attention integration module. The weighted complementary cascaded module is connected to the partial attention integration module, and both the weighted complementary cascaded module and the partial attention integration module are connected to the partial attention integration module in the preceding localization unit. The foremost partial attention integration module in the localization branch used to determine the first localization map is connected to the last partial attention integration module in the localization branch used to determine the second localization map. The localization branch is equipped with a partial attention integration module, which supervises and corrects the weighted complementary cascaded module to improve its output and performs multiple depth supervisions, thereby improving the image quality of the subsequently obtained fine density map.
[0086] Furthermore, since the feature map resolutions of the output terms of the iterative fusion branches following the top-down order and the iterative fusion branches following the bottom-up order are different, when fusing the output terms from the two directions, the output terms of the iterative fusion branches following the top-down order with lower resolution are first upsampled using bilinear interpolation to be consistent with Conv2-2. By upsampling the output terms with lower resolution, the resolution of the density map can be guaranteed to maintain more detailed spatial information, which can avoid the loss of spatial information caused by downsampling the output terms with higher resolution, thereby improving the image quality of the generated fine density map.
[0087] For example: Suppose several feature maps are Conv2-2, Conv3-3, Conv4-3, and Conv5-3 of VGG-16, denoted as f1, f2, f3, and f4 respectively. Figure 5As shown, the localization branch includes three localization units. In the iterative localization branch where iterative fusion is performed in a top-down order, the inputs of the weighted complementary cascade module in the first localization unit are f1 and f2, and the inputs of the partial attention integration module are the outputs of the weighted complementary cascade module and the output of the partial attention integration module in the last localization unit in the iterative localization branch where iterative fusion is performed in a bottom-up order. The inputs of the weighted complementary cascade module in the second localization unit are the outputs of the partial attention integration module in the first localization unit and f3, and the inputs of the partial attention integration module are the outputs of the partial attention integration module in the first localization unit and the output of the weighted complementary cascade module in that localization unit. The inputs of the weighted complementary cascade module in the last localization unit are the outputs of the partial attention integration module in the second localization unit and f4, and the inputs of the partial attention integration module are the outputs of the partial attention integration module in the second localization unit and the output of the weighted complementary cascade module in that localization unit.
[0088] In the iterative localization branch that performs iterative fusion in a bottom-up order, the inputs of the weighted complementary cascade module in the first localization unit are f3 and f4, and the inputs of the partial attention integration module are the outputs of the weighted complementary cascade module and f4. The inputs of the weighted complementary cascade module in the second localization unit are the outputs of the partial attention integration module in the first localization unit and f2, and the inputs of the partial attention integration module are the outputs of the partial attention integration module in the first localization unit and the outputs of the weighted complementary cascade module in that localization unit. The inputs of the weighted complementary cascade module in the last localization unit are the outputs of the partial attention integration module in the second localization unit and f1, and the inputs of the partial attention integration module are the outputs of the partial attention integration module in the second localization unit and the outputs of the weighted complementary cascade module in that localization unit.
[0089] In one implementation, the model structure of the weighted complementary cascade module is the same as that of the weighted complementary cascade module in the complementary iterative aggregation module described above, and will not be repeated here. Figure 6 As shown, the partial attention integration module includes a fifth convolutional block, a third scaling layer, a sixth convolutional block, a fourth scaling layer, an activation function layer, and a fusion layer. The fifth convolutional block is connected to the third scaling layer, the third scaling layer is connected to the activation function layer, the sixth convolutional block is connected to the fourth scaling layer, and both the activation function layer and the fourth scaling layer are connected to the fusion layer.
[0090] The attention integration module performs deep supervision on the coarse attention map during the fusion process, refining the attention map generation process and thus effectively improving the quality of the generated attention map. The specific process is as follows:
[0091] F i ,F sup =M PAI (F i-1 M WCC (F i-1 ,f k ))
[0092] F sup =Sigmoid(Scale(F) i-1 ))
[0093] F i =F sup ⊙Scale(M WCC (F i-1 ,f k ))
[0094] Among them, F sup F represents a coarse attention map for multiple deep supervision. i The feature map representing the output or input of the PAI module, when i = {1, 2, 3, 4, 5, 6}, F i This represents the output of the PAI module. When i = 0, F i This represents the feature map extracted from the VGG-16 backbone network. M PAI and M WCC These represent the PAI and WCC modules, respectively. Sigmoid indicates a convolutional layer using the Sigmoid activation function. Scale indicates scaling operations such as upsampling, downsampling, and no change to ensure that the resulting attention map and density map have the same resolution. ⊙ indicates pixel-wise multiplication.
[0095] S50. Input the coarse density map and the attention map into the fusion module of the region complementary aggregation network model, and output the fine density map through the fusion module.
[0096] Specifically, the fusion module includes a first convolutional branch, a second convolutional branch, and a third convolutional branch. Both the first and second convolutional branches are connected to the third convolutional branch. The first convolutional branch includes a convolutional layer and an activation function layer; the second convolutional branch includes a convolutional layer; and the third convolutional branch includes a multiplier and a convolutional layer. The coarse density map output by the complementary iterative aggregation module is passed through the first convolutional branch and then input into the third convolutional branch. The attention map output by the region localization module is passed through the second convolutional branch and then input into the third convolutional branch. The third convolutional branch performs a dot product between the coarse density map and the attention map, and then passes the result through a convolutional layer to obtain a fine density map.
[0097] S60. Determine the crowd count result corresponding to the dense crowd image based on the fine density map.
[0098] Specifically, the crowd count result refers to the total number of people corresponding to the closely packed crowd image. After obtaining the fine density map, the total number of people corresponding to the dense crowd image can be obtained by statistically analyzing the human images in the fine density map.
[0099] In one implementation, since the regional complementary aggregation network model is equipped with a complementary iterative aggregation module and a regional localization module, multiple deep supervision is performed through the regional localization module. Thus, when training the regional complementary aggregation network model, the loss function can include a supervision loss term determined by the predicted density map and the ground truth of the density map output by the regional complementary aggregation network model, a supervision loss term determined by the attention map output by the regional localization module and the ground truth of the attention map, and multiple deep supervision loss terms determined by the partial attention maps output by the partial attention integration module in the regional localization module and the ground truth of the attention maps.
[0100] The attention map value is obtained by processing the ground truth value of the density map using a threshold function. A threshold T is determined, and when the probability value of an element in the density map is greater than T, the value of that element is set to 1; otherwise, it is set to 0. The resulting binary attention map will be used as the ground truth value for the attention map generated by the model, ensuring that the generated attention map only focuses on the distribution of the foreground crowd and not on the local density magnitude. This helps to prevent edge blurring caused by excessive smoothing of the Gaussian kernel.
[0101] The loss function can be expressed as:
[0102]
[0103] in, Representing the supervised loss of the density map, respectively Supervised loss of attention map Let α and λ represent the loss from multiple deep supervisions. iThe hyperparameter is used to balance multiple losses, where n represents the number of partial attention integration modules.
[0104] In one implementation, the supervised loss of the density map Mean squared loss can be used as the loss function, and its calculation formula is as follows:
[0105]
[0106] Among them, I n Let θ represent the nth input image, Θ represent the network parameters, and F(I) represent the input image. n ;Θ) represents the predicted density map output by the network. represents the ground truth of the density map, and N represents the size of the training dataset.
[0107] Both the loss from multiple deep supervision and the loss from detailed attention maps can be calculated using a pixel-wise binary classification cross-entropy loss function, with the following formula:
[0108]
[0109] Where H and W represent the height and width of the input image, D represents the predicted value at location (h,w). hw This represents the truth value at position (h, w).
[0110] In summary, this embodiment provides a crowd counting method based on region complementary aggregation. The method includes: acquiring a dense crowd image; inputting the dense crowd image into the feature module of a pre-trained region complementary aggregation network model, and outputting several feature maps corresponding to the dense crowd image through the feature module; inputting the several feature maps into the complementary iterative aggregation module of the region complementary aggregation network model, and outputting a coarse density map through the complementary iterative aggregation module; inputting the several feature maps into the region localization module of the region complementary aggregation network model, and outputting an attention map through the region localization module; inputting the coarse density map and the attention map into the fusion module of the region complementary aggregation network model, and outputting a fine density map through the fusion module; and determining the crowd counting result corresponding to the dense crowd image based on the fine density map. This application generates a coarse density map through bidirectional iterative fusion and weighted complementary cascading of a complementary iterative aggregation module, and determines an attention map through multiple depth supervision using a region localization module. Then, a fine density map is determined based on the attention map and the coarse density map. In this way, multi-scale fusion through the complementary iterative aggregation module and multiple depth supervision through the region localization module effectively improves the image quality of the fine density map, thereby improving the accuracy of determining the crowd counting results based on the fine density map.
[0111] Based on the above-described population counting method based on regional complementary aggregation, this embodiment provides a population counting system based on regional complementary aggregation, such as... Figure 7 The system includes:
[0112] Acquisition module 100 is used to acquire images of dense crowds;
[0113] The control module 200 is used to input the dense crowd image into a pre-trained region complementary aggregation network model to determine a fine density map. Specifically, the process of determining the fine density map involves: outputting several feature maps corresponding to the dense crowd image through the feature map module; inputting the several feature maps into the complementary iterative aggregation module of the region complementary aggregation network model to output a coarse density map; inputting the several feature maps into the region localization module of the region complementary aggregation network model to output an attention map; and inputting the coarse density map and the attention map into the fusion module of the region complementary aggregation network model to output a fine density map.
[0114] The determination module 300 is used to determine the crowd count result corresponding to the dense crowd image based on the fine density map.
[0115] Based on the above-described population counting method based on regional complementary aggregation, this embodiment provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the population counting method based on regional complementary aggregation as described in the above embodiment.
[0116] Based on the aforementioned population counting method based on regional complementary aggregation, this application also provides a terminal device, such as... Figure 8 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical instructions in the memory 22 to execute the methods described in the above embodiments.
[0117] Furthermore, the logical instructions in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0118] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.
[0119] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, as well as transient storage media.
[0120] Furthermore, the specific process of loading and executing multiple instruction processors in the aforementioned storage medium and terminal device has been described in detail in the above method, and will not be repeated here.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A population counting method based on regional complementary aggregation, characterized in that, The method includes: Acquire images of dense crowds; The dense crowd image is input into the feature module of a pre-trained region complementarity aggregation network model, and the feature module outputs several feature maps corresponding to the dense crowd image. The aforementioned feature maps are input into the complementary iterative aggregation module of the region complementary aggregation network model, and a coarse density map is output through the complementary iterative aggregation module. The complementary iterative aggregation module adopts a bidirectional iterative fusion method, namely bottom-up iterative fusion and top-down iterative fusion. The aforementioned feature maps are input into the region localization module of the region complementary aggregation network model, and the region localization module outputs an attention map. The coarse density map and the attention map are input into the fusion module of the region complementary aggregation network model, and the fine density map is output through the fusion module. The crowd count result corresponding to the dense crowd image is determined based on the fine density map; The region localization module includes two localization branches and a second fusion branch; the process of inputting the plurality of feature maps into the region complementary aggregation network model and outputting an attention map through the mutual region localization module specifically includes: The aforementioned feature maps are input into two localization branches. One localization branch fuses the feature maps in a bottom-up order to obtain a first localization map, while the other localization branch fuses them in a top-down order to obtain a second localization map. Each localization branch includes several cascaded localization units. Each localization unit includes a weighted complementary cascade module and a partial attention integration module. The weighted complementary cascade module is connected to the partial attention integration module. Both the weighted complementary cascade module and the partial attention integration module are connected to the partial attention integration module in the preceding localization unit. The first partial attention integration module in the localization branch used to determine the first localization map is connected to the last partial attention integration module in the localization branch used to determine the second localization map. The partial attention integration module supervises and corrects the weighted complementary cascade module. The weighted complementary cascade module extracts and fuses the true sensitive regions in the feature maps through a process of generating and redistributing weights. The first and second positioning maps are input into the second fusion branch, and the attention map is output through the fusion unit.
2. The population counting method based on regional complementary aggregation according to claim 1, characterized in that, The complementary iterative aggregation module includes two iterative fusion branches and a first fusion branch; the process of inputting the plurality of feature maps into the complementary iterative aggregation network model and outputting a coarse density map through the complementary iterative aggregation module specifically includes: The aforementioned feature maps are respectively input into two iterative fusion branches. One iterative fusion branch fuses the feature maps in a bottom-up order to obtain a first fusion map, and the other iterative fusion branch fuses the feature maps in a top-down order to obtain a second fusion map. The iterative fusion branch includes several cascaded weighted complementary cascaded modules. The first fusion map and the second fusion map are input into the first fusion branch, and a coarse density map is output through the fusion unit.
3. The population counting method based on regional complementary aggregation according to claim 1 or 2, characterized in that, The weighted complementary cascaded module includes a first convolutional block, a first scaling layer, a second convolutional block, a second scaling layer, complementary units, a first multiplier, a second multiplier, an adder, a third convolutional block, a first connection layer, and a fourth convolutional block. The first convolutional block is connected to the first scaling layer, the second convolutional block is connected to the second scaling layer, both the first and second scaling layers are connected to complementary units, both the first scaling layer and the complementary unit are connected to the first multiplier, both the second scaling layer and the complementary unit are connected to the second multiplier, and both the first and second multipliers are connected to the adder. The adder is connected to the third convolutional block, the first scaling layer, the second scaling layer, and the third convolutional block are all connected to the first connecting layer, and the first connecting layer is connected to the fourth convolutional block. The complementary unit includes a first global average pooling layer, a second global average pooling layer, a second connecting layer, and an activation function layer. The first global average pooling layer is connected to the first scaling layer and the second connecting layer, the second global average pooling layer is connected to the second scaling layer and the second connecting layer, and the second connecting layer is connected to the activation function layer.
4. The population counting method based on regional complementary aggregation according to claim 1, characterized in that, The attention integration module includes a fifth convolutional block, a third scaling layer, a sixth convolutional block, a fourth scaling layer, an activation function layer, and a fusion layer. The fifth convolutional block is connected to the third scaling layer, the third scaling layer is connected to the activation function layer, the sixth convolutional block is connected to the fourth scaling layer, and both the activation function layer and the fourth scaling layer are connected to the fusion layer.
5. The population counting method based on regional complementary aggregation according to claim 1, characterized in that, The loss function used in the training process of the region complementary aggregation network model includes a supervised loss term determined by the predicted density map and the ground truth value of the density map output by the region complementary aggregation network model, a supervised loss term determined by the attention map output by the region localization module and the ground truth value of the attention map, and a multi-depth supervised loss term determined by the partial attention maps output by the partial attention integration module in the region localization module and the ground truth value of the attention maps. The ground truth value of the attention map is obtained by processing the ground truth value of the density map through a threshold function.
6. A population counting system based on regional complementary aggregation, characterized in that, The system includes: The acquisition module is used to acquire images of dense crowds; The control module is used to input the dense crowd image into a pre-trained region complementary aggregation network model to determine a fine density map. Specifically, the process of determining the fine density map involves: outputting several feature maps corresponding to the dense crowd image through a feature module; inputting the several feature maps into a complementary iterative aggregation module of the region complementary aggregation network model to output a coarse density map; inputting the several feature maps into a region localization module of the region complementary aggregation network model to output an attention map; and inputting the coarse density map and the attention map into a fusion module of the region complementary aggregation network model to output a fine density map. The complementary iterative aggregation module employs a bidirectional iterative fusion method, consisting of bottom-up iterative fusion and top-down iterative fusion. The determination module is used to determine the crowd count result corresponding to the dense crowd image based on the fine density map; The region localization module includes two localization branches and a second fusion branch; the process of inputting the plurality of feature maps into the region complementary aggregation network model and outputting an attention map through the mutual region localization module specifically includes: The aforementioned feature maps are input into two localization branches. One localization branch fuses the feature maps in a bottom-up order to obtain a first localization map, while the other localization branch fuses them in a top-down order to obtain a second localization map. Each localization branch includes several cascaded localization units. Each localization unit includes a weighted complementary cascade module and a partial attention integration module. The weighted complementary cascade module is connected to the partial attention integration module. Both the weighted complementary cascade module and the partial attention integration module are connected to the partial attention integration module in the preceding localization unit. The first partial attention integration module in the localization branch used to determine the first localization map is connected to the last partial attention integration module in the localization branch used to determine the second localization map. The partial attention integration module supervises and corrects the weighted complementary cascade module. The weighted complementary cascade module extracts and fuses the true sensitive regions in the feature maps through a process of generating and redistributing weights. The first and second positioning maps are input into the second fusion branch, and the attention map is output through the fusion unit.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the population counting method based on regional complementary aggregation as described in any one of claims 1-5.
8. A terminal device, characterized in that, include: Processor, memory, and communication bus; the memory stores a computer-readable program that can be executed by the processor; The communication bus enables communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the population counting method based on regional complementary aggregation as described in any one of claims 1-5.