A crowd density estimation method and system based on efficient multi-semantic information aggregation
By adding semantic supervision and spatial information embedding strategies to the population density estimation method and combining the lightweight hollow space pyramid pooling structure, the shortcomings of the existing methods in view angle changes and occlusion processing are solved, and the accuracy and efficiency of density estimation are improved.
Patent Information
- Application Number
- CN202210600585.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The existing population density estimation method has shortcomings in dealing with the problem of viewing angle changes and occlusion, and the semantic gaps in high and low-level features have not been fully considered, resulting in room for improvement in accuracy.
The population density estimation method of efficient multi-semantic information aggregation is adopted, and by adding semantic supervision strategies and spatial information embedding strategies to the backbone network, the quality of low-level features is improved and the spatial representation ability of high-level features is strengthened. Combined with the lightweight hollow space pyramid pooling structure, multi-scale context information is captured through step-size convolution and feature fusion is performed to generate the final population density map.
It improves the accuracy and efficiency of crowd density estimation, reduces the complexity and running speed of the model, adapts to crowd density estimation in various scenarios, and the algorithm runs more stable.
Smart Images

Figure CN114943933B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crowd density calculation, and relates to a crowd density estimation method and system for efficient multi-semantic information aggregation. Background Art
[0002] With the rapid growth of my country's population and the acceleration of urbanization, large-scale crowd gatherings are increasing, which also brings various safety hazards such as stampede accidents caused by crowds and traffic dispatch pressure during rush hour. Therefore, crowd density estimation and counting has become an important research topic in the field of public safety.
[0003] At present, the main crowd counting method used in recent years is to regress the number of people by learning the mapping between the local features of the image and its corresponding density map. It can not only achieve the crowd counting task well, but also intuitively reflect the density distribution of the crowd through the density map, and has high application value. Among them, the deep learning-based method has a good performance in the field of crowd technology due to its excellent deep feature acquisition ability, but with the continuous increase in background complexity and perspective changes, the existing methods still have room for further improvement.
[0004] Information aggregation is an effective way to solve the problem of perspective change and occlusion. For example, by using the context space pyramid to aggregate information on the local and overall image, and introducing global context information into the feature map, the quality of density map generation can be effectively improved; however, this method cannot effectively deal with the problem of changes in the distribution of people and the loss of image feature information. The high- and low-level features of the backbone network are directly fused, and the channel attention module is used to optimize the feature fusion process. The hollow convolution is used to expand the receptive field and regress the density map; although the scale diversity problem caused by perspective is solved, the semantic gap between high- and low-level features is not considered during feature fusion, and there is still room for further improvement in accuracy. By encoding information about perspective changes from the "up-left-right-down" direction, deep global context information is captured through progressive aggregation, and scale relationship features of multi-dimensional perspectives are simultaneously extracted; although a high accuracy is achieved, the overly redundant network structure causes the model complexity to increase and the network operation speed to slow down. Summary of the invention
[0005] The purpose of the present invention is to solve the problems in the prior art and to provide a method and system for estimating crowd density by efficiently aggregating multi-semantic information.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] An efficient multi-semantic information aggregation crowd density estimation method includes the following steps:
[0008] S1: Input the crowd image to be detected into the constructed backbone network, combine the semantic supervision strategy, obtain the enhanced low-level features with some semantic information from the shallow layer of the network, introduce the spatial information embedding strategy in the deep layer of the backbone network, use bilinear interpolation upsampling for the high-level features in the backbone network, multiply the upsampled high-level features with the low-level features element by element, and obtain the enhanced high-level features;
[0009] S2: Construct a lightweight hollow spatial pyramid pooling structure. Through point-by-point convolution layers, the low-level features with semantic information obtained in S1 and the enhanced high-level features are respectively reduced in channel dimension. Stride convolution is used to capture the contextual information of low-level features and high-level features respectively, and multi-scale high-level features and multi-scale low-level features are obtained. The multi-scale high-level features and multi-scale low-level features are fused to obtain global multi-scale contextual features.
[0010] S3: Continue to upsample the global multi-scale context features through strided convolution to obtain the final crowd density map.
[0011] A further improvement of the present invention is:
[0012] In step S1, the backbone network is a VGG-19 backbone network, and the first 13 layers in the VGG-19 backbone network are selected to implement feature extraction.
[0013] The step S1 comprises the following steps:
[0014] S1.1: Input the crowd image to be detected into the layer network of the constructed backbone network, and output the low-level feature map F li And the high-level feature map F hi ;
[0015] S1.2 Low-level feature map F li Refine and reduce the feature mapping dimension; then reduce the number of parameters through global average pooling, form semantic boundary constraints, and generate low-level features F with partial semantic information li-1 ;
[0016] S1.3: The F in each lower layer network li-1 Fusion, to obtain low-level features F rich in semantic information Low level feature
[0017] S1.4: Introduce spatial information embedding strategy to embed high-level feature maps F hi Bilinear interpolation upsampling is used to scale the high-level channel size to the same dimension as the low-level channel size to obtain the upsampled feature map F hi-1 ;
[0018] S1.5: The low-level feature map F li With the upsampled feature map Fhi-1 Multiply element by element to obtain the enhanced high-level features F High-level feature .
[0019] In step S1.2, a 3×3 convolution and a 1×1 convolution are performed on the low-level feature map F. li To be refined.
[0020] The step S2 comprises the following steps:
[0021] S2.1: Through point-by-point convolutional layers, F High-level feature and F Low level feature Perform channel dimensionality reduction and implement channel information interaction;
[0022] S2.2: Through strided convolution, the dilation rate is used to enrich the receptive field of the feature map and capture contextual information;
[0023] S2.3: F after S2.2 processing High-level feature Feature map and F Low level feature The feature maps are fused to obtain multi-scale high-level features F H and multi-scale low-level features F L ;
[0024] S2.4: Fusion of multi-scale high-level features F H and multi-scale low-level features F L , and obtain the global multi-scale context feature F N .
[0025] In step S2.1, channel dimension reduction is performed through 4 point-by-point convolutional layers with a kernel size of 1.
[0026] In step S2.2, four expansion rates of different sizes are used to enrich the receptive field of the feature map.
[0027] An efficient multi-semantic information aggregation crowd density estimation system, including a feature extraction module, a feature fusion module and a global processing module;
[0028] The feature extraction module is used to input the crowd image to be detected into the constructed backbone network, combine the semantic supervision network, obtain the low-level features with semantic information, introduce the spatial information embedding strategy into the backbone network, use bilinear interpolation upsampling for the high-level features in the backbone network, multiply the upsampled high-level features by the low-level features element by element, and obtain the enhanced high-level features;
[0029] The feature fusion module is used to construct a lightweight hollow spatial pyramid pooling structure. Through point-by-point convolution layers, the low-level features with semantic information obtained in the feature extraction module and the enhanced high-level features are respectively reduced in channel dimension. The stride convolution is used to capture the contextual information of the low-level features and high-level features respectively, and multi-scale high-level features and multi-scale low-level features are obtained. The multi-scale high-level features and multi-scale low-level features are fused to obtain global multi-scale contextual features.
[0030] The global processing module is used to continue upsampling the global multi-scale context features through strided convolution to obtain the final crowd density map.
[0031] A terminal device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any method described in the present invention when executing the computer program.
[0032] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any method described in the present invention.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The present invention discloses a crowd density estimation method for efficient multi-semantic information aggregation, which adds a semantic supervision network to a backbone network to improve the quality of low-level features and alleviate the interference of background noise; combines with a spatial information embedding strategy to strengthen the spatial representation ability of high-level features, supplement the spatial information of high-level features, and alleviate the semantic loss of high-level features; obtains the final global multi-scale context features through stride convolution, which can reduce the number of parameters, overcome network redundancy, reduce the complexity of the model, speed up the operation of the network, and improve the efficiency of feature extraction; the method disclosed by the present invention can adapt to crowd density estimation in a variety of scenarios, the algorithm runs more stably, and the method can be used to realize crowd counting and density distribution detection in a variety of scenarios.
[0035] Furthermore, the VGG-19 backbone network disclosed in the present invention can better mine the deep features of images.
[0036] Furthermore, the present invention refines the low-level feature maps, strengthens the detailed expression of the low-level features, reduces the number of parameters through global average pooling, reduces the interference of noise, and improves the semantics of the low-level features; scales the high-level channel size to the same dimension as the low-level channel size, optimizes the feature fusion method, and avoids the loss of semantic information caused by direct fusion.
[0037] Furthermore, the present invention is to High-level feature and F Low level featureChannel dimension reduction is performed, channel information interaction is implemented, and more context information is captured at a lower computational cost through strided convolution, which enriches the receptive field of the feature map, improves the output quality of the feature map, and simplifies the calculation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 This is a network structure diagram of multi-semantic spatial information aggregation crowd density estimation of the present invention;
[0040] Figure 2 A multi-semantic feature extraction network diagram of the present invention;
[0041] Figure 3 Schematic diagram of the MSS and SE execution results of the present invention, (a is the input image; b is the second layer feature map; c is the 3×3 convolution execution result; d is the 1×1 convolution execution result; e is the 13th layer feature map; f is the bilinear interpolation upsampling result)
[0042] Figure 4 This is the global multi-scale context information aggregation network diagram of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0046] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0047] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0048] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0049] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0050] See also Figure 1 The present invention discloses a crowd density estimation method based on efficient multi-semantic information aggregation, comprising the following steps:
[0051] Firstly, a multivariate information extraction network is designed, using the VGG19 backbone network and combining the multi-layer semantic supervision strategy with the spatial information embedding strategy. The multi-layer semantic supervision strategy is used to encode low-level features to improve the semantic expression of low-level features, and spatial information embedding is used to enrich the spatial information representation of high-level features. The optimized low- and high-level features are the initial information.
[0052] Secondly, a multi-scale context information aggregation network is designed. The initial information obtained in step 1 is processed, and two lightweight dilated spatial pyramid pooling structures with strided convolutions are used to perform global multi-scale context information aggregation while alleviating model parameter redundancy to obtain the intermediate information of the crowd;
[0053] Finally, at the end of the network, stride convolution is used to upsample the features obtained in the previous step to obtain the final density map, which reduces the amount of calculation while ensuring accuracy.
[0054] The specific steps include:
[0055] Step 1: Multivariate Information Extraction
[0056] The embodiment of the present invention selects the VGG-19 network, which is similar to the VGG-16 network structure but has a deeper network layer, to obtain better initial features. Removing the fully connected layer of the VGG-19 network has little effect on the accuracy of crowd counting and can effectively reduce network parameters. Therefore, the embodiment of the present invention uses the VGG-19 network with the fully connected layer removed as the skeleton network to alleviate network redundancy while obtaining deep features. Low-level features contain more location and detail information, but have low semantics and more noise. To address this problem, the embodiment of the present invention proposes the following Figure 2 The multi-layer semantic supervision strategy MSS shown in Figure 2 processes low-level features and designs three semantic supervision modules (SS) at the 2nd, 4th and 6th layers of the VGG-19 backbone network. Figure 1 The schematic image in is taken as input, and its second layer feature map is taken as an example to illustrate the execution process of the SS module. The specific extraction process is as follows, see Figure 3 :
[0057] Step 1.1: Output the feature map F of the second layer of VGG19 l2 A 3×3 convolution and a 1×1 convolution are sent to refine the feature map output. The output feature map is as follows Figure 3 c and 3d show that the feature map dimension is reduced and the feature detail expression is enhanced;
[0058] Step 1.2: Reduce the number of parameters through a global average pooling, integrate global spatial information, form semantic boundary constraints, reduce noise interference, and generate high-quality low-level features F with partial semantic information l2-1 .
[0059] Step 1.3: Feature map F output from the fourth and sixth layers of VGG19 l4 and F l6 Repeat steps 1.1-1.2 to obtain the optimized low-level features F l4-1 、F l6-1 .
[0060] Step 1.4: Fusion F l2-1 、F l4-1 and F l6-1 Get low-level features F rich in semantic information Low level feature ;
[0061] Step 1.5: The feature map size of the VGG-19 network used in the embodiment of the present invention is only 1 / 16 of the input image at the 13th layer. Fusion of low-level features is one of the effective ways to supplement the spatial information of high-level features, but the direct fusion method will cause some semantic information loss due to the low overlap of the spatial resolution of the feature maps. Therefore, in order to enhance the spatial representation ability of high-level features, the following is proposed: Figure 2 The spatial information embedding strategy SE shown in the figure is used for the 13th layer feature F of VGG-19 h13 Use bilinear interpolation upsampling, such as Figure 3 As shown in f, the channel size is scaled to the same dimension as the 6th layer, and F is obtained h13-1 .
[0062] Step 1.6: The sixth layer feature F l6 With F h13-1 Multiply element by element to optimize the feature fusion method, supplement the spatial information for high-level features while alleviating the loss of semantic information caused by fusion, and obtain the enhanced high-level features F High-level feature ;
[0063] Step 2: Extraction of global multi-scale contextual information
[0064] See also Figure 4 , through two lightweight Simplify-atrous spatial pyramid pooling (S-ASPP) modules, the contextual information of low-level and high-level features at different scales is gradually captured and integrated step by step, and the global expression of features is enhanced under the premise of ensuring limited computational cost. For the convenience of description, the two S-ASPP modules are respectively denoted as S1-ASPP and S2-ASPP.
[0065] The embodiment of the present invention designs a lightweight atrous spatial pyramid pooling S-ASPP structure. Taking S1-ASPP as an example: first, through 4 point-by-point convolution layers with a kernel size of 1, channel dimension reduction is performed on the high-level features obtained by the multivariate information extraction network, and channel information interaction is performed; secondly, an Inception-like structure is adopted, and strided convolution is used to reduce model redundancy, and the feature map receptive field is enriched with expansion rates of 1, 6, 12, and 18 to capture more contextual information; finally, the processed feature map is fused to enhance the global expression of the feature.
[0066] The specific steps are:
[0067] Step 2.1: Through 4 point-wise convolutional layers with kernel size 1, High-level feature Perform channel dimensionality reduction and implement channel information interaction.
[0068] Step 2.2: Use an Inception-like structure and strided convolution to reduce model redundancy, enrich the feature map receptive field with expansion rates of 1, 6, 12, and 18 to capture more contextual information.
[0069] Step 2.3: Perform a fusion operation on the feature maps processed in steps 2.1-2.2 to enhance the global expression of the features and obtain multi-scale high-level features F H .
[0070] Step 2.4: F Low level feature Repeat the above steps 2.1-2.3 to obtain the multi-scale low-level features F L ;
[0071] Step 2.5: Fusion F H 、F L , and obtain the global multi-scale context feature F N .
[0072] Step 2.6: Use strided convolution to perform N Continue upsampling to get the final density map.
[0073] The method disclosed in the embodiment of the present invention maintains good stability for large-scale event scenes in different density areas. At the same time, the crowd counting results can be used to efficiently allocate resources for the venue. Moreover, the crowd distribution results through the density map can better reflect the corresponding queue neatness, providing assistance for the rehearsal of large-scale performances.
[0074] The embodiment of the present invention discloses a crowd density estimation system with efficient multi-semantic information aggregation.
[0075] It includes feature extraction module, feature fusion module and global processing module;
[0076] The feature extraction module is used to input the crowd image to be detected into the constructed backbone network, combine the semantic supervision network, obtain the low-level features with semantic information, introduce the spatial information embedding strategy into the backbone network, use bilinear interpolation upsampling for the high-level features in the backbone network, multiply the upsampled high-level features by the low-level features element by element, and obtain the enhanced high-level features;
[0077] The feature fusion module is used to construct a lightweight hollow spatial pyramid pooling structure. Through point-by-point convolution layers, the low-level features with semantic information obtained in the feature extraction module and the enhanced high-level features are respectively reduced in channel dimension. The stride convolution is used to capture the contextual information of the low-level features and high-level features respectively, and multi-scale high-level features and multi-scale low-level features are obtained. The multi-scale high-level features and multi-scale low-level features are fused to obtain global multi-scale contextual features.
[0078] Global processing module: The global multi-scale context features are further upsampled through strided convolution to obtain the final crowd density map.
[0079] A schematic diagram of a terminal device provided in an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0080] The computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to accomplish the present invention.
[0081] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0082] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0083] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0084] If the module / unit integrated in the terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0085] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An efficient crowd density estimation method based on multi-semantic information aggregation, It is characterized in that The following steps are involved: S1: Input the crowd image to be detected into the constructed backbone network, combine the semantic supervision strategy, obtain the enhanced low-level features with some semantic information from the shallow layer of the network, introduce the spatial information embedding strategy in the deep layer of the backbone network, use bilinear interpolation upsampling for the high-level features in the backbone network, multiply the upsampled high-level features with the low-level features element by element, and obtain the enhanced high-level features; S2: Construct a lightweight hollow spatial pyramid pooling structure. Through point-by-point convolution layers, the low-level features with semantic information obtained in S1 and the enhanced high-level features are respectively reduced in channel dimension. Stride convolution is used to capture the contextual information of low-level features and high-level features respectively, and multi-scale high-level features and multi-scale low-level features are obtained. The multi-scale high-level features and multi-scale low-level features are fused to obtain global multi-scale contextual features. S3: Continue to upsample the global multi-scale context features through stride convolution to obtain the final crowd density map; In step S1, the backbone network is a VGG-19 backbone network, and the first 13 layers in the VGG-19 backbone network are selected to implement feature extraction; The step S1 comprises the following steps: S1.1: Input the crowd image to be detected into the layer network of the constructed backbone network, and output the low-level feature map F li And the high-level feature map F hi ; S1.2 Low-level feature map F li Refine and reduce the feature mapping dimension; then reduce the number of parameters through global average pooling, form semantic boundary constraints, and generate low-level features F with partial semantic information li-1 ; S1.3: The F in each lower layer network li-1 Fusion, to obtain low-level features F rich in semantic information Lowlevelfeature S1.4: Introduce spatial information embedding strategy to embed high-level feature maps F hi Bilinear interpolation upsampling is used to scale the high-level channel size to the same dimension as the low-level channel size to obtain the upsampled feature map F hi-1 ; S1.5: The low-level feature map F li With the upsampled feature map F hi-1 Multiply element by element to obtain the enhanced high-level features F High-levelfeature .
2. According to claim 1, a method for estimating crowd density by efficient multi-semantic information aggregation, It is characterized in that In step S1.2, a 3×3 convolution and a 1×1 convolution are performed on the low-level feature map F. li To be refined.
3. The method for estimating crowd density by efficient multi-semantic information aggregation according to claim 1, It is characterized in that The step S2 comprises the following steps: S2.1: Through point-by-point convolutional layers, F High-levelfeature and F Lowlevelfeature Perform channel dimension reduction and implement channel information interaction; S2.2: Through strided convolution, the dilation rate is used to enrich the receptive field of the feature map and capture contextual information; S2.3: F after S2.2 processing High-levelfeature Feature map and F Lowlevelfeature The feature maps are fused to obtain multi-scale high-level features F H and multi-scale low-level features F L ; S2.4: Fusion of multi-scale high-level features F H and multi-scale low-level features F L , and obtain the global multi-scale context feature F N .
4. The method for estimating crowd density by efficient multi-semantic information aggregation according to claim 3, It is characterized in that In step S2.1, channel dimension reduction is performed through 4 point-by-point convolutional layers with a kernel size of 1.
5. The method for estimating crowd density by efficient multi-semantic information aggregation according to claim 3, It is characterized in that In step S2.2, four expansion rates of different sizes are used to enrich the receptive field of the feature map.
6. A crowd density estimation system based on efficient multi-semantic information aggregation according to the method of claim 1, It is characterized in that It includes feature extraction module, feature fusion module and global processing module; The feature extraction module is used to input the crowd image to be detected into the constructed backbone network, combine the semantic supervision network, obtain the low-level features with semantic information, introduce the spatial information embedding strategy into the backbone network, use bilinear interpolation upsampling for the high-level features in the backbone network, multiply the upsampled high-level features by the low-level features element by element, and obtain the enhanced high-level features; The feature fusion module is used to construct a lightweight hollow spatial pyramid pooling structure. Through point-by-point convolution layers, the low-level features with semantic information obtained in the feature extraction module and the enhanced high-level features are respectively reduced in channel dimension. The stride convolution is used to capture the contextual information of the low-level features and high-level features respectively, and multi-scale high-level features and multi-scale low-level features are obtained. The multi-scale high-level features and multi-scale low-level features are fused to obtain global multi-scale contextual features. The global processing module is used to continue upsampling the global multi-scale context features through strided convolution to obtain the final crowd density map.
7. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image semantic segmentation method and device based on codec
CN111292330A
Crowd density estimation method and device based on multi-feature information fusion, and storage medium
CN113743422A