A method, device, electronic device and storage medium for crowd counting in a dense scene
The method improves crowd counting accuracy in dense scenes by using a trained network for frequency domain enhancement and visual modeling to refine feature extraction and density estimation.
Patent Information
- Application Number
- CN202510421620.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In crowd-intensive scenarios, it is difficult for the prior art to extract sufficiently accurate pedestrian characteristics, resulting in low accuracy of crowd detection counts.
The grouped pyramid feature convolution, frequency domain conversion enhancement, visual information modeling and peak point filtering are used to process the detected population images by training a complete population counting network to obtain more complete pedestrian feature information.
The accuracy of crowd detection counting is improved, and the contours and features of pedestrians are captured through frequency domain conversion, combined with global and medium- and long-distance information modeling, reducing occlusion interference and achieving more accurate counting.
Smart Images

Figure CN119942464B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation, and particularly to a method, device, electronic device and storage medium for counting people in a dense scene. Background Art
[0002] The main goal of the crowd counting task is to support applications such as surveillance systems, intelligent transportation management, and retail store customer flow analysis, and to accurately estimate the number of people in a scene by analyzing images or videos. In some crowded scenes, such as in a subway station, many pedestrians are blocked by each other, which makes manual counting time-consuming and inaccurate. Therefore, it is particularly important to use deep learning technology to handle such problems. In the existing image recognition and counting methods, the density regression method based on point annotation is simple and efficient in the process of making a dataset. Only simple point annotations are required on the heads of the targets. In this way, the density distribution of the crowd can be effectively captured. The density regression method can better handle the counting problem in crowded areas by learning the mapping relationship between local features in the image and the crowd density map. This method usually uses a Gaussian kernel to simulate the corresponding position of each human head in the image and performs regularization processing to learn the non-linear mapping between local features and the target density map.
[0003] Currently, CNN (Convolutional Neural Networks) is applied to the crowd detection and counting of the density regression method due to its good accuracy. In crowd detection and counting, CNN can better extract the features of each pedestrian through its good local feature extraction ability and identify the pedestrians in the image. However, in crowded scenes, due to the mutual occlusion between pedestrians and the occlusion of some background buildings or objects, the features will interfere with each other. Existing CNN methods often have difficulty extracting sufficiently accurate pedestrian features in crowded scenes, resulting in low accuracy of crowd detection and counting.
[0004] Therefore, the existing technology has the problem that it is difficult to extract sufficiently accurate pedestrian features in crowded scenes, resulting in low accuracy of crowd detection and counting. Summary of the Invention
[0005] In view of this, it is necessary to provide a method, device, electronic device and storage medium for counting people in a dense scene to solve the technical problem in the existing technology that it is difficult to extract sufficiently accurate pedestrian features in crowded scenes, resulting in low accuracy of crowd detection and counting.
[0006] To solve the above technical problem, on the one hand, the present invention provides a method for counting people in a dense scene, including:
[0007] Input the obtained image of the population to be detected into the well-trained population counting network, perform grouped pyramid feature convolution on the image of the population to be detected to obtain preliminary convolution features, perform frequency domain transformation enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map;
[0008] Perform peak point filtering on the estimated density map to obtain the population count recognition output.
[0009] In a possible implementation, performing frequency domain transformation enhancement on the preliminary convolution features to obtain frequency domain enhanced features includes:
[0010] Perform two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function, and inverse Fourier transform on the preliminary convolution features in sequence to obtain frequency domain convolution features;
[0011] Fuse the preliminary convolution features and the frequency domain convolution features to obtain frequency domain enhanced features.
[0012] In a possible implementation, performing visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features includes:
[0013] Perform feature dimension transformation on the frequency domain enhanced features to obtain modal transformation features;
[0014] Perform global information visual transformation modeling on the modal transformation features to obtain the first modeling feature, and perform medium and long-range information visual sequence modeling on the modal transformation features to obtain the second modeling feature;
[0015] After fusing the first modeling feature and the second modeling feature, perform feature dimension transformation to obtain visually modeled enhanced features.
[0016] In a possible implementation, performing global information visual transformation modeling on the modal transformation features to obtain the first modeling feature includes:
[0017] Perform segmentation on the modal transformation features to obtain several segmented features;
[0018] Perform multi-head attention feature extraction and multi-layer perceptron feature extraction on the segmented features to obtain the first modeling feature.
[0019] In a possible implementation, performing medium and long-range information visual sequence modeling on the modal transformation features to obtain the second modeling feature includes:
[0020] Unfold the modal transformation features into a feature tensor to obtain an unfolded tensor;
[0021] Perform linear mapping, convolution operation, and bidirectional state space feature extraction on the unfolded tensor to obtain the second modeling feature.
[0022] In a possible implementation, peak point filtering is performed on the estimated density map to obtain the crowd counting recognition output, including:
[0023] Pooling operation, threshold segmentation, and non-maximum suppression are performed on the estimated density map to obtain the crowd counting recognition output.
[0024] In a possible implementation, a trained complete crowd counting network is obtained by training an initial crowd counting network. Training the initial crowd counting network includes:
[0025] Inputting the crowd counting training data into the initial crowd counting network to obtain a predicted density map;
[0026] Determining the parametric error decay loss of the initial crowd counting network according to the predicted density map and the corresponding true density map, and iteratively optimizing the initial crowd counting network according to the parametric error decay loss until the loss no longer decreases, to obtain a trained complete crowd counting network.
[0027] On the other hand, the present invention also provides a dense scene crowd counting device, including:
[0028] A feature extraction and prediction unit, configured to input the acquired crowd image to be detected into the trained complete crowd counting network, perform grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map;
[0029] A peak point filtering unit, configured to perform peak point filtering on the estimated density map to obtain the crowd counting recognition output.
[0030] On the other hand, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned dense scene crowd counting method is implemented.
[0031] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned dense scene crowd counting method is implemented.
[0032] The beneficial effects of the present invention are as follows: In the crowd counting method for dense scenarios provided by the present invention, first, the acquired crowd image to be detected is input into a well-trained crowd counting network. The crowd image to be detected is subjected to grouped pyramid feature convolution to obtain preliminary convolution features, the preliminary convolution features are enhanced through frequency domain conversion to obtain frequency domain enhanced features, the frequency domain enhanced features are subjected to visual information modeling to obtain visually modeled enhanced features, and the visually modeled enhanced features are subjected to density regression output to obtain an estimated density map. Then, peak point filtering is performed on the estimated density map to obtain the crowd counting recognition output. Through frequency domain conversion enhancement, the present invention can convert the image into the frequency domain to capture the contours and features of pedestrians, improving the accuracy of pedestrian features. Through visual information enhancement, visual information modeling is performed on pedestrian features at both the global information level and the medium- and long-distance information level to obtain more complete pedestrian feature information, effectively improving the accuracy of crowd detection and counting. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0034] Figure 1 It is a schematic flowchart of an embodiment of the crowd counting method for dense scenarios provided by the present invention;
[0035] Figure 2 It is a schematic flowchart of the frequency domain conversion enhancement in the embodiment of the present invention;
[0036] Figure 3 It is a schematic flowchart of the visual information modeling in the embodiment of the present invention;
[0037] Figure 4 It is a schematic flowchart of the global information visual conversion modeling in the embodiment of the present invention;
[0038] Figure 5 It is a schematic flowchart of the medium- and long-distance information visual sequence modeling in the embodiment of the present invention;
[0039] Figure 6 It is a schematic flowchart of training the initial crowd counting network in the embodiment of the present invention;
[0040] Figure 7 It is a schematic flowchart of an embodiment of the crowd counting device for dense scenarios provided by the present invention;
[0041] Figure 8 It is a schematic flowchart of an embodiment of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present invention.
[0043] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0044] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0045] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0046] The present invention provides a dense scene crowd counting method, device, electronic device, and storage medium, which will be described separately below.
[0047] Figure 1 It is a flowchart of an embodiment of the dense scene crowd counting method provided by the present invention. As Figure 1 shown, the dense scene crowd counting method includes:
[0048] S101. Input the obtained image of the crowd to be detected into a well-trained crowd counting network, perform multi-level feature convolution on the image of the crowd to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information enhancement on the frequency domain enhanced features to obtain visual enhanced features, and perform density regression output on the visual enhanced features to obtain an estimated density map;
[0049] S102. Filter the peak points of the estimated density map to obtain the crowd counting recognition output.
[0050] Compared with the prior art, the crowd counting method for dense scenes provided by the embodiments of the present invention first inputs the obtained crowd image to be detected into a well-trained crowd counting network, performs grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, performs frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, performs visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and performs density regression output on the visually modeled enhanced features to obtain an estimated density map; then filters the peak points of the estimated density map to obtain the crowd counting recognition output. Through frequency domain conversion enhancement, the present invention can convert the image to the frequency domain to capture the contours and features of pedestrians, improve the accuracy of pedestrian features, and perform visual information modeling on pedestrian features at the global information level and the medium- and long-distance information level through visual information enhancement respectively, obtaining more complete pedestrian feature information and effectively improving the accuracy of crowd detection and counting.
[0051] In some embodiments of the present invention, Figure 2 is a schematic flowchart of the frequency domain conversion enhancement of the embodiments of the present invention. As Figure 2 shown, performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features includes:
[0052] S201. Perform two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function, and inverse Fourier transform on the preliminary convolution features in sequence to obtain frequency domain convolution features;
[0053] S202. Fuse the preliminary convolution features and the frequency domain convolution features to obtain frequency domain enhanced features.
[0054] Specifically, in order to improve the accuracy of crowd counting, the embodiment proposes to perform feature enhancement through frequency domain conversion enhancement on the basis of CNN feature extraction. Frequency domain conversion enhancement processes image data in the frequency domain and the complex domain, extracts features by combining Fourier transform and deep learning techniques, and further processes these features in the spatial domain, thereby realizing frequency domain conversion enhancement. In this way, the model can be made to pay more attention to the texture detail part of the image, so as to extract more accurate and detailed pedestrian edge and contour features, thereby improving the accuracy of pedestrian detection. The specific process is as follows:
[0055] First, at the initial stage of feature extraction, the embodiment first obtains local shallow information under different receptive fields through grouped convolutional pyramids. Among them, each layer of the feature pyramid corresponds to a receptive field at a certain size, and the features obtained from each layer are merged after successive feature convolutions to obtain preliminary convolution features.
[0056] Then, in the frequency-domain conversion enhancement module, in the embodiment, the format of the input data is first adjusted, then the input data is subjected to two-dimensional Fourier transform and then complex convolution, and the features after convolution are processed by complex batch normalization and complex GELU activation function. Finally, the real and imaginary parts of the data are combined into a complex form and then inverse Fourier transform is performed to convert the data back from the frequency domain to the spatial domain. The formulas for the grouped pyramid feature convolution and the frequency-domain conversion enhancement process are expressed as:
[0057]
[0058]
[0059] Among them, is the pyramid convolution, is the original input image, are convolution kernels of different sizes, is the feature map obtained after convolution. is the frequency-domain conversion enhancement operation, is the feature map containing frequency-domain information.
[0060] Finally, the features after the frequency-domain conversion enhancement are further subjected to a convolution operation to complete the feature extraction of the CNN modality part, and after being fused with the preliminary convolution features in the early stage, frequency-domain enhanced features are obtained and sent to the next module.
[0061] In some embodiments of the present invention, Figure 3 is a schematic flow diagram of the visual information modeling in the embodiment of the present invention. As Figure 3 shown, visual information modeling is performed on the frequency-domain enhanced features to obtain visual modeling enhanced features, including:
[0062] S301. Perform feature dimension conversion on the frequency-domain enhanced features to obtain modality conversion features;
[0063] S302. Perform global information visual conversion modeling on the modality conversion features to obtain first modeling features, and perform medium- and long-range information visual sequence modeling on the modality conversion features to obtain second modeling features;
[0064] S303. After fusing the first modeling features and the second modeling features, perform feature dimension conversion to obtain visual modeling enhanced features.
[0065] Specifically, considering that although the traditional CNN network has strong feature extraction ability in local feature extraction, in the process of pedestrian feature extraction in dense scenes, due to the mutual occlusion between pedestrians and background objects, only local feature extraction will be interfered by complexity and it is difficult to extract accurate enough pedestrian features. Therefore, the embodiment combines CNN with ViT (Vision Transformer) and ViM (Vision Mamba), integrating the local feature extraction ability of CNN, the global feature capture ability of ViT, and the medium- and long-distance modeling advantage of ViM, effectively reducing the interference of complex features and obtaining more complete pedestrian feature information.
[0066] In visual information modeling, the embodiment first needs to perform feature dimension conversion on the data of the CNN modality with the ViT modality and the ViM modality, and then perform corresponding ViT modeling and ViM modeling respectively. The formula is expressed as:
[0067]
[0068]
[0069]
[0070]
[0071] Among them, represents the feature map dimension expansion operation, represent batch, channel, height, and width respectively. The method is used to change the shape of the tensor without changing the total number of its elements. The method is used to rearrange the dimensions of the tensor. It does not change the data in the tensor, but only changes the access order of the data. represents first changing the shape to , where 0 in means keeping the first dimension unchanged, and 2 and 1 mean moving the third dimension to the second position and moving the second dimension to the third position, that is, rearranging the dimensions to . is the corresponding feature map after ViT or ViM modeling. is the feature map density regression operation, which performs density regression output on the visually modeled enhanced features obtained by fusing ViT modeling and ViM modeling to obtain the finally output estimated density map
[0072] In some embodiments of the present invention, Figure 4 is the flow schematic diagram of the global information visual conversion modeling of the embodiment of the present invention, asFigure 4 As shown, global information visual transformation modeling is performed on the modality conversion feature to obtain a first modeling feature, including:
[0073] S401. The modality conversion feature is segmented to obtain a number of segmentation features;
[0074] S402. Multi-head attention feature extraction and multi-layer perceptron feature extraction are performed on the segmentation features to obtain a first modeling feature.
[0075] Specifically, in the crowd ViT module, the embodiment first segments the modality conversion feature into a number of segmentation feature patches, then performs multi-head attention feature extraction after a layer normalization, and merges the extracted features with the initially segmented features. Then, the features output by the multi-head attention are normalized again and then multi-layer perceptron feature extraction is performed, and the obtained features are merged with the features output by the multi-head attention to obtain the first modeling feature output by the crowd ViT module.
[0076] In some embodiments of the present invention, Figure 5 is a schematic flowchart of medium- and long-range information visual sequence modeling according to an embodiment of the present invention. As Figure 5 shown, medium- and long-range information visual sequence modeling is performed on the modality conversion feature to obtain a second modeling feature, including:
[0077] S501. The modality conversion feature is unfolded into a feature tensor to obtain an unfolded tensor;
[0078] S502. Linear mapping, convolution operation, and bidirectional state space feature extraction are performed on the unfolded tensor to obtain a second modeling feature.
[0079] Specifically, in the crowd ViM module, the embodiment first unfolds the modality conversion feature into a feature tensor to obtain an unfolded tensor, representing the high-dimensional tensor as a sum of simpler components. Then, the obtained unfolded tensor sequentially performs linear mapping, convolution operation, activation function, and bidirectional state space feature extraction to capture key features. Then, for the features of the two branches obtained by the bidirectional state space feature extraction, one of the branches is multiplied and fused with the features obtained after linear mapping and activation function with the unfolded tensor, and then concatenated with the features obtained from the other branch. Finally, after another linear mapping and concatenation with the initial unfolded tensor, tensor dimension adjustment is performed to obtain the second modeling feature output by the crowd ViM module.
[0080] In some embodiments of the present invention, peak point filtering is performed on the estimated density map to obtain the crowd counting recognition output, including:
[0081] Pooling operation, threshold segmentation, and non-maximum suppression are performed on the estimated density map to obtain the crowd counting recognition output.
[0082] Specifically, in the current density regression counting model, although the embodiments have made significant progress in density estimation through the above steps, the predictions output by the model still face adjustments in terms of localization accuracy. This is due to the dispersion of the targets in the predicted density map, which not only affects the accuracy of counting but also increases the complexity of subsequent processing. Therefore, the embodiments design a localization model to accurately identify and locate the targets in the image by filtering the peak points of the estimated density map, thereby improving the counting accuracy and reliability of the model.
[0083] Through steps such as grouped pyramids, frequency domain transformation enhancement, visual information modeling, and density regression output, the model can output a high-quality fine density estimation map. In the density estimation map, the regressed Gaussian blurred convex peaks are almost independent of each other. Each peak position can be regarded as the center position of the target. By accurately filtering out the accurate peak points, the accurate localization of pedestrian targets can be achieved.
[0084] In peak point filtering, the embodiments first apply a pooling operation to the generated estimated density map. This step not only helps to identify the initial peak points, that is, the possible center positions of pedestrian targets, but also makes the density map visually smoother by reducing local fluctuations, laying a foundation for subsequent processing steps. Next, the embodiments perform threshold segmentation on the pooled density map. By a preset threshold, the background outliers that do not belong to pedestrian targets can be effectively excluded, thereby achieving background suppression. Further purifying the useful signals in the image. Finally, in order to accurately locate the center point of the target, the embodiments adopt the maximum suppression algorithm, which can identify the local maximum points, that is, the center coordinate points of the target, from the density map after threshold processing.
[0085] In some embodiments of the present invention, the trained crowd counting network is obtained by training the initial crowd counting network. Figure 6 It is a schematic flowchart of training the initial crowd counting network according to an embodiment of the present invention. As Figure 6 shown, training the initial crowd counting network includes:
[0086] S601. Input the crowd counting training data into the initial crowd counting network to obtain a predicted density map;
[0087] S602. Determine the parametric error decay loss of the initial crowd counting network according to the predicted density map and the corresponding true density map, and iteratively optimize the initial crowd counting network according to the parametric error decay loss until the loss no longer decreases, to obtain the trained crowd counting network.
[0088] Specifically, in the crowd counting task of the density regression method, the loss function is mostly used between the predicted density map and the ground truth density map. To better complete the supervision task, the embodiment adopts PED-Loss (Parameterized Error Decay Loss) as the loss function during the model training process. The PED-Loss function is a loss function used in the regression task, which has a smooth characteristic and is robust to outliers. Its formula is:
[0089]
[0090] Where and represent the predicted value and the ground truth value of the th data point, respectively. is a weight parameter used to adjust the influence of the squared difference term. is another weight parameter used to adjust the influence of the exponential decay term. is the rate of exponential decay, which controls the term's decay speed. is the square of the difference between the predicted value and the ground truth value, which helps to emphasize larger errors. is the exponential decay term, which rapidly decreases as the error increases, helping to reduce the influence of extreme values.
[0091] In summary, for the crowd counting method in dense scenes provided by the present invention, first, the obtained crowd image to be detected is input into the trained crowd counting network, and the crowd image to be detected is subjected to grouped pyramid feature convolution to obtain preliminary convolution features. The preliminary convolution features are enhanced by frequency domain conversion to obtain frequency domain enhanced features. The frequency domain enhanced features are modeled for visual information to obtain visually modeled enhanced features. The visually modeled enhanced features are output by density regression to obtain an estimated density map. Then, peak point filtering is performed on the estimated density map to obtain the crowd counting recognition output. Through frequency domain conversion enhancement, the present invention can convert the image to the frequency domain to capture the contours and features of pedestrians, improving the accuracy of pedestrian features. Through visual information enhancement, visual information modeling is performed on pedestrian features at the global information level and the medium-long distance information level respectively to obtain more complete pedestrian feature information, effectively improving the accuracy of crowd detection and counting.
[0092] To better implement the crowd counting method in dense scenes in the embodiments of the present invention, correspondingly, based on the crowd counting method in dense scenes, as Figure 7 shown, the present invention also provides a crowd counting device for dense scenes. The crowd counting device 700 for dense scenes includes:
[0093] The feature extraction and prediction unit 701 is configured to input the acquired image of the population to be detected into a trained population counting network, perform grouped pyramid feature convolution on the image of the population to be detected to obtain preliminary convolution features, perform frequency domain transformation enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map.
[0094] The peak point filtering unit 702 is configured to perform peak point filtering on the estimated density map to obtain the population count recognition output.
[0095] The dense scene population counting device 700 provided in the above embodiment can implement the technical solutions described in the above embodiment of the dense scene population counting method. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above embodiment of the dense scene population counting method, which will not be elaborated here.
[0096] As Figure 8 shown, the present invention also correspondingly provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802, and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0097] In some embodiments, the processor 801 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is configured to run program codes stored in the memory 802 or process data, such as the dense scene population counting method in the present invention.
[0098] In some embodiments, the processor 801 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 801 may be local or remote. In some embodiments, the processor 801 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination of the above.
[0099] In some embodiments, the memory 802 may be an internal storage unit of the electronic device 800, such as the hard disk or memory of the electronic device 800. In some other embodiments, the memory 802 may also be an external storage device of the electronic device 800, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 800.
[0100] Furthermore, the memory 802 may also include both the internal storage unit of the electronic device 800 and external storage devices. The memory 802 is used to store the application software installed in the electronic device 800 and various types of data.
[0101] In some embodiments, the display 803 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 803 is used to display the information of the electronic device 800 and to display a visual user interface. The components 801 - 803 of the electronic device 800 communicate with each other through a system bus.
[0102] In one embodiment, when the processor 801 executes the dense scene crowd counting program in the memory 802, the following steps can be implemented:
[0103] Input the obtained crowd image to be detected into a trained crowd counting network, perform grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, perform frequency domain transformation enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map;
[0104] Perform peak point filtering on the estimated density map to obtain the crowd counting recognition output.
[0105] It should be understood that when the processor 801 executes the dense scene crowd counting program in the memory 802, in addition to the above functions, other functions can also be implemented. For specific details, please refer to the description of the corresponding method embodiments above.
[0106] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the dense scene crowd counting methods provided by the above method embodiments can be implemented.
[0107] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0108] The above has introduced in detail the method, device, electronic device and storage medium for crowd counting in dense scenarios provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for counting people in a dense scene, characterized in that, Including: Input the obtained image of the population to be detected into the trained population counting network, perform grouped pyramid feature convolution on the image of the population to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map; Perform peak point filtering on the estimated density map to obtain the population count recognition output; The performing visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features includes: Perform feature dimension conversion on the frequency domain enhanced features to obtain modality conversion features; Perform global information visual conversion modeling on the modality conversion features to obtain the first modeling feature, and perform medium and long-range information visual sequence modeling on the modality conversion features to obtain the second modeling feature; After fusing the first modeling feature and the second modeling feature, perform feature dimension conversion to obtain visually modeled enhanced features; The performing global information visual conversion modeling on the modality conversion features to obtain the first modeling feature includes: Perform segmentation on the modality conversion features to obtain a number of segmented features; Perform multi-head attention feature extraction and multi-layer perceptron feature extraction on the segmented features to obtain the first modeling feature; The performing medium and long-range information visual sequence modeling on the modality conversion features to obtain the second modeling feature includes: Unfold the modality conversion features into a feature tensor to obtain an unfolded tensor; Perform linear mapping, convolution operation, activation function, and bidirectional state space feature extraction on the unfolded tensor in sequence to capture key features; For the features of the two branches obtained by bidirectional state space feature extraction, multiply the features obtained by performing linear mapping and activation function on one branch with the unfolded tensor, then fuse them with the features obtained by the other branch, splice them, and after passing through a linear mapping, splice them with the initial unfolded tensor and then perform tensor dimension adjustment to obtain the second modeling feature.
2. The method for counting the number of people in a dense scene according to claim 1, wherein The performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features includes: Perform two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function, and inverse Fourier transform on the preliminary convolution features in sequence to obtain frequency domain convolution features; Fuse the preliminary convolution features and the frequency domain convolution features to obtain frequency domain enhanced features.
3. The method for counting the number of people in a dense scene according to claim 1, wherein, The performing peak point filtering on the estimated density map to obtain the population count recognition output includes: Perform pooling operation, threshold segmentation, and non-maximum suppression on the estimated density map to obtain the population count recognition output.
4. The method for counting the number of people in a dense scene according to claim 1, wherein The trained population counting network is obtained by training an initial population counting network. The training of the initial population counting network includes: Input population counting training data into the initial population counting network to obtain a predicted density map; Determine the parametric error decay loss of the initial population counting network according to the predicted density map and the corresponding true density map, and iteratively optimize the initial population counting network according to the parametric error decay loss until the loss no longer decreases, to obtain the trained population counting network.
5. An apparatus for counting people in a dense scene, characterized in that, Including: A feature extraction and prediction unit, which is configured to input the acquired image of the population to be detected into a trained population counting network, perform grouped pyramid feature convolution on the image of the population to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features, and perform density regression output on the visually modeled enhanced features to obtain an estimated density map; A peak point filtering unit, which is configured to perform peak point filtering on the estimated density map to obtain a population count recognition output; The performing visual information modeling on the frequency domain enhanced features to obtain visually modeled enhanced features includes: Performing feature dimension conversion on the frequency domain enhanced features to obtain modality conversion features; Performing global information visual conversion modeling on the modality conversion features to obtain a first modeled feature, and performing medium- and long-range information visual sequence modeling on the modality conversion features to obtain a second modeled feature; Fusing the first modeled feature and the second modeled feature and then performing feature dimension conversion to obtain visually modeled enhanced features; The performing global information visual conversion modeling on the modality conversion features to obtain a first modeled feature includes: Splitting the modality conversion features to obtain a number of split features; Performing multi-head attention feature extraction and multi-layer perceptron feature extraction on the split features to obtain a first modeled feature; The performing medium- and long-range information visual sequence modeling on the modality conversion features to obtain a second modeled feature includes: Unfolding the modality conversion features into a feature tensor to obtain an unfolded tensor; Successively performing linear mapping, convolution operation, activation function, and bidirectional state space feature extraction on the unfolded tensor to capture key features; For the features of the two branches obtained by the bidirectional state space feature extraction, multiply the features obtained by linearly mapping and applying an activation function to one of the branches and the unfolded tensor, fuse the result with the features obtained by the other branch, splice the result with the initial unfolded tensor after another linear mapping, and then perform tensor dimension adjustment to obtain a second modeled feature.
6. An electronic device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the dense scene population counting method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the dense scene population counting method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Dense crowd counting method combining high-resolution CNN (Convolutional Neural Network) and lightweight Transformer
CN116935316A
Anti-shielding head three-dimensional positioning method and system based on instance segmentation
CN119205924A
Image exposure adjustment method and device, equipment, medium and product
CN119583969A