Dense scene crowd counting method and device, electronic equipment and storage medium

By applying technical means such as grouped pyramid feature convolution, frequency domain conversion enhancement and visual information modeling in the crowd counting network, the problem of inaccurate pedestrian feature extraction in crowd-intensive scenarios is solved, and the accuracy of crowd detection counting is significantly improved.

CN119942464AActive Publication Date: 2025-05-06HUBEI UNIV OF TECH

Patent Information

Application Number
CN202510421620.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art is difficult to extract sufficiently accurate pedestrian characteristics in crowd-intensive scenarios, resulting in low accuracy of crowd detection counts.

Method used

By inputting the images of the people to be detected to train a complete population counting network, grouped pyramid feature convolution, frequency domain conversion enhancement, visual information modeling and density regression output are performed to obtain the estimated density map and peak point filtering to obtain the population count recognition output.

Benefits of technology

The accuracy of pedestrian characteristics is improved, and more complete pedestrian characteristic information is obtained through global and medium- and long-distance information modeling, which effectively improves the accuracy of crowd detection counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942464A_ABST
    Figure CN119942464A_ABST
Patent Text Reader

Abstract

The invention provides a crowd counting method and device for a dense scene, electronic equipment and a storage medium, and belongs to the field of intelligent traffic, and the method comprises the steps: inputting an obtained to-be-detected crowd image into a completely trained crowd counting network, carrying out the grouping type pyramid feature convolution of the to-be-detected crowd image, and obtaining a preliminary convolution feature, performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, performing visual information modeling on the frequency domain enhanced features to obtain visual modeling enhanced features, and performing density regression output on the visual modeling enhanced features to obtain an estimated density map; and performing peak point filtering on the estimated density map to obtain crowd counting recognition output. According to the method, the image is converted to the frequency domain through frequency domain conversion enhancement to capture the contour and features of the pedestrian, the pedestrian feature precision is improved, visual information modeling is carried out on the pedestrian features in global and medium-long distance levels through visual information enhancement, more complete feature information is obtained, and the crowd counting accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a method, device, electronic equipment and storage medium for counting people in dense scenes. Background Art

[0002] The main goal of crowd counting tasks is to support applications such as surveillance systems, intelligent traffic management, and retail passenger flow analysis by analyzing images or videos to accurately estimate the number of people in the scene. In some crowded scenes, such as subway stations, many pedestrians are blocked by each other, which makes manual counting time-consuming and inaccurate. Therefore, it is particularly important to use deep learning technology to deal with such problems. Among the existing image recognition and counting methods, the density regression method based on point annotation is simple and efficient in the process of data set preparation. It only requires simple point annotation on the head of the target. In this way, the density distribution of the crowd can be effectively captured. The density regression method can better handle the counting problem of densely populated areas by learning the mapping relationship between local features in the image and the crowd density map. This method usually uses a Gaussian kernel to simulate the corresponding position of each head in the image, and performs regularization to learn the nonlinear mapping between local features and the target density map.

[0003] At present, CNN (Convolutional Neural Networks, CNN) is used in crowd detection and counting based on density regression methods due to its good accuracy. In crowd detection and counting, CNN can better extract the features of each pedestrian and identify pedestrians in the image through its good local feature extraction ability. However, in densely populated scenes, pedestrians often block each other and are blocked by some background buildings or objects, and features interfere with each other. Existing CNN methods often find it difficult to extract sufficiently accurate pedestrian features in dense scenes, resulting in low accuracy of crowd detection and counting.

[0004] Therefore, the existing technology has the problem of difficulty in extracting sufficiently accurate pedestrian features in crowded scenes, resulting in low accuracy of crowd detection and counting. Summary of the invention

[0005] In view of this, it is necessary to provide a method, device, electronic device and storage medium for counting crowds in densely populated scenes to solve the technical problem in the prior art that it is difficult to extract sufficiently accurate pedestrian features in densely populated scenes, resulting in low accuracy of crowd detection and counting.

[0006] In order to solve the above technical problems, on the one hand, the present invention provides a method for counting people in a dense scene, comprising: The acquired crowd image to be detected is input into a well-trained crowd counting network, and the crowd image to be detected is subjected to grouped pyramid feature convolution to obtain preliminary convolution features, and the preliminary convolution features are enhanced by frequency domain conversion to obtain frequency domain enhanced features, and visual information modeling is performed on the frequency domain enhanced features to obtain visual modeling enhanced features, and density regression is performed on the visual modeling enhanced features to output an estimated density map; The estimated density map is peak-filtered to obtain the crowd counting recognition output.

[0007] In a possible implementation, the preliminary convolution features are enhanced by frequency domain conversion to obtain frequency domain enhanced features, including: The preliminary convolution features are sequentially subjected to two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function and inverse Fourier transform to obtain frequency domain convolution features; The preliminary convolution features and frequency domain convolution features are fused to obtain frequency domain enhanced features.

[0008] In a possible implementation, visual information modeling is performed on the frequency domain enhanced features to obtain visual modeling enhanced features, including: Perform feature dimension conversion on the frequency domain enhanced features to obtain modal conversion features; Performing global information visual conversion modeling on the modal conversion feature to obtain a first modeling feature, and performing medium- and long-range information visual sequence modeling on the modal conversion feature to obtain a second modeling feature; After fusing the first modeling feature and the second modeling feature, feature dimension conversion is performed to obtain a visual modeling enhancement feature.

[0009] In a possible implementation, performing global information visual conversion modeling on the modal conversion feature to obtain a first modeling feature includes: The modal conversion features are segmented to obtain a number of segmentation features; Multi-head attention feature extraction and multi-layer perceptron feature extraction are performed on the segmentation features to obtain the first modeling feature.

[0010] In a possible implementation, the modal conversion feature is subjected to medium- and long-range information visual sequence modeling to obtain a second modeling feature, including: Expand the feature tensor of the modal conversion feature to obtain an expanded tensor; The second modeling feature is obtained by performing linear mapping, convolution operation and bidirectional state space feature extraction on the unfolded tensor.

[0011] In a possible implementation, peak point filtering is performed on the estimated density map to obtain a crowd counting recognition output, including: Pooling, threshold segmentation and non-maximum suppression are performed on the estimated density map to obtain the crowd counting recognition output.

[0012] In one possible implementation, the trained crowd counting network is obtained by training an initial crowd counting network, and the training of the initial crowd counting network includes: Input the crowd counting training data into the initial crowd counting network to obtain the predicted density map; The parameterized error attenuation loss of the initial crowd counting network is determined according to the predicted density map and the corresponding true density map. The initial crowd counting network is iteratively optimized according to the parameterized error attenuation loss until the loss no longer decreases, thereby obtaining a fully trained crowd counting network.

[0013] On the other hand, the present invention also provides a device for counting people in a dense scene, comprising: A feature extraction and prediction unit is used to input the acquired crowd image to be detected into a well-trained crowd counting network, perform grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features, perform visual information modeling on the frequency domain enhanced features to obtain visual modeling enhanced features, and perform density regression on the visual modeling enhanced features to output an estimated density map; The peak point filtering unit is used to perform peak point filtering on the estimated density map to obtain a crowd counting recognition output.

[0014] On the other hand, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned method for counting crowds in dense scenes is implemented.

[0015] On the other hand, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned method for counting people in dense scenes is implemented.

[0016] The beneficial effects of the present invention are as follows: in the dense scene crowd counting method provided by the present invention, the acquired crowd image to be detected is first input into a well-trained crowd counting network, the crowd image to be detected is subjected to grouped pyramid feature convolution to obtain preliminary convolution features, the preliminary convolution features are subjected to frequency domain conversion enhancement to obtain frequency domain enhancement features, the frequency domain enhancement features are subjected to visual information modeling to obtain visual modeling enhancement features, the visual modeling enhancement features are subjected to density regression output to obtain an estimated density map; then the estimated density map is subjected to peak point filtering to obtain a crowd counting recognition output. The present invention can convert images to the frequency domain through frequency domain conversion enhancement to capture the contours and features of pedestrians, improve the accuracy of pedestrian features, and perform visual information modeling on pedestrian features at the global information level and the medium and long-distance information level through visual information enhancement to obtain more complete pedestrian feature information, thereby effectively improving the accuracy of crowd detection and counting. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A schematic diagram of a flow chart of an embodiment of a method for counting people in a dense scene provided by the present invention; Figure 2 A schematic diagram of a process of frequency domain conversion enhancement according to an embodiment of the present invention; Figure 3 A schematic diagram of a process of visual information modeling according to an embodiment of the present invention; Figure 4 A schematic diagram of a process for modeling global information visual conversion according to an embodiment of the present invention; Figure 5 A schematic diagram of a process for modeling a visual sequence of medium- and long-range information according to an embodiment of the present invention; Figure 6 A schematic diagram of a process of training an initial crowd counting network according to an embodiment of the present invention; Figure 7 A schematic diagram of a flow chart of an embodiment of a crowd counting device for densely populated scenes provided by the present invention; Figure 8 A schematic diagram of a flow chart of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0020] In the description of the embodiments of the present invention, unless otherwise specified, "multiple" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone.

[0021] The descriptions of "first" and "second" in the embodiments of the present invention are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined as "first" and "second" may explicitly or implicitly include at least one of the features.

[0022] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0023] The present invention provides a method, device, electronic device and storage medium for counting people in a dense scene, which are described below respectively.

[0024] Figure 1 The flowchart of an embodiment of the method for counting people in a dense scene provided by the present invention is as follows: Figure 1 As shown, the crowd counting method in dense scenes includes: S101, inputting the acquired crowd image to be detected into a well-trained crowd counting network, performing multi-level feature convolution on the crowd image to be detected to obtain preliminary convolution features, performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, performing visual information enhancement on the frequency domain enhancement features to obtain visual enhancement features, and performing density regression on the visual enhancement features to output an estimated density map; S102: Filter the estimated density map by peak points to obtain a crowd counting recognition output.

[0025] Compared with the prior art, the method for counting crowds in dense scenes provided by the embodiment of the present invention first inputs the acquired crowd image to be detected into a well-trained crowd counting network, performs grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, performs frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, performs visual information modeling on the frequency domain enhancement features to obtain visual modeling enhancement features, performs density regression output on the visual modeling enhancement features to obtain an estimated density map; then performs peak point filtering on the estimated density map to obtain a crowd counting recognition output. The present invention can convert images to the frequency domain through frequency domain conversion enhancement to capture the contours and features of pedestrians, improve the accuracy of pedestrian features, and perform visual information modeling on pedestrian features at the global information level and the medium and long-distance information level through visual information enhancement to obtain more complete pedestrian feature information, thereby effectively improving the accuracy of crowd detection and counting.

[0026] In some embodiments of the present invention, Figure 2 FIG. 1 is a flow chart of frequency domain conversion enhancement according to an embodiment of the present invention. Figure 2 As shown, the preliminary convolution features are transformed and enhanced in the frequency domain to obtain frequency domain enhanced features, including: S201, performing two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function and inverse Fourier transform on the preliminary convolution features in sequence to obtain frequency domain convolution features; S202, fusing the preliminary convolution features and the frequency domain convolution features to obtain frequency domain enhanced features.

[0027] Specifically, in order to improve the accuracy of crowd counting, the embodiment proposes feature enhancement through frequency domain conversion enhancement based on CNN feature extraction. Frequency domain conversion enhancement processes image data in the frequency domain and complex domain, extracts features by combining Fourier transform and deep learning technology, and further processes these features in the spatial domain, thereby achieving frequency domain conversion enhancement. In this way, the model can pay more attention to the texture details of the image, thereby extracting more accurate and detailed pedestrian edge and contour features, thereby improving the accuracy of pedestrian detection. The specific process is as follows: First, in the initial stage of feature extraction, the embodiment first obtains local shallow information under different perception fields of view by grouping convolution pyramids, wherein each layer of the feature pyramid corresponds to a perception field of view under a certain size, and after the level-by-level feature convolution, the features obtained at each layer are merged to obtain preliminary convolution features.

[0028] Then in the frequency domain conversion enhancement module, the embodiment first adjusts the format of the input data, then performs a two-dimensional Fourier transform on the input data and then performs a complex convolution, and performs complex batch normalization and complex GELU activation function processing on the convolved features, and finally combines the real and imaginary parts of the data into a complex form and performs an inverse Fourier transform to convert the data from the frequency domain back to the spatial domain. The formula for the grouped pyramid feature convolution and frequency domain conversion enhancement process is expressed as:

[0029]

[0030] in, is the pyramid convolution, is the original input image, are convolution kernels of different sizes, It is the feature map obtained after convolution. To enhance the frequency domain conversion operation, It is a feature map containing frequency domain information.

[0031] Finally, the features after frequency domain conversion and enhancement are subjected to convolution operation to complete the feature extraction of CNN modal part, and are fused with the previous preliminary convolution features to obtain frequency domain enhanced features and sent to the next module.

[0032] In some embodiments of the present invention, Figure 3FIG. 1 is a flow chart of visual information modeling according to an embodiment of the present invention. Figure 3 As shown, visual information modeling is performed on the frequency domain enhanced features to obtain visual modeling enhanced features, including: S301, performing feature dimension conversion on the frequency domain enhancement feature to obtain a modal conversion feature; S302, performing global information visual conversion modeling on the modal conversion feature to obtain a first modeling feature, and performing medium- and long-range information visual sequence modeling on the modal conversion feature to obtain a second modeling feature; S303: After fusing the first modeling feature and the second modeling feature, perform feature dimension conversion to obtain a visual modeling enhancement feature.

[0033] Specifically, although the traditional CNN network has strong feature extraction capabilities in local feature extraction, in the process of extracting features of pedestrians in dense scenes, due to the mutual occlusion between pedestrians and background objects, only local feature extraction will be subject to complex interference and it is difficult to extract sufficiently accurate pedestrian features. In this regard, the embodiment combines CNN with ViT (Vision Transformer) and ViM (Vision Mamba), combining the local feature extraction capability of CNN, the global feature capture capability of ViT and the medium and long distance modeling advantages of ViM, effectively reducing the interference of complex features and obtaining more complete pedestrian feature information.

[0034] In the visual information modeling, the embodiment first needs to convert the CNN modality data into the ViT modality and ViM modality feature dimensions, and then perform corresponding ViT modeling and ViM modeling respectively, and the formula is expressed as:

[0035]

[0036]

[0037]

[0038] in, Represents the feature map dimension expansion operation, Represents batch, channel, height and width respectively. method is used to change the shape of a tensor without changing the total number of its elements. The method is used to rearrange the dimensions of a tensor. It does not change the data in the tensor, but only changes the order in which the data is accessed. Indicates that first Reshape to , 0 means keeping the first dimension unchanged, 2 and 1 mean moving the third dimension to the second position and the second dimension to the third position, that is, rearranging the dimensions as . It is the feature map after the corresponding ViT or ViM modeling. It is a feature map density regression operation. The visual modeling enhanced features obtained by fusion of ViT modeling and ViM modeling are output through density regression to obtain the final output estimated density map. .

[0039] In some embodiments of the present invention, Figure 4 A schematic diagram of a flow chart of global information visual conversion modeling according to an embodiment of the present invention, such as Figure 4 As shown, the modal conversion feature is subjected to global information visual conversion modeling to obtain the first modeling feature, including: S401, segmenting the modal conversion feature to obtain a plurality of segmentation features; S402, performing multi-head attention feature extraction and multi-layer perceptron feature extraction on the segmentation features to obtain the first modeling features.

[0040] Specifically, in the crowd ViT module, the embodiment first divides the modal conversion feature into a number of segmented feature patches, then performs multi-head attention feature extraction after a layer of normalization, and merges the extracted features with the initial segmented features. Then, the features output by the multi-head attention are passed through a normalization layer and then multi-layer perceptron feature extraction is performed, and the obtained features are merged with the features output by the multi-head attention to obtain the first modeling feature output by the crowd ViT module.

[0041] In some embodiments of the present invention, Figure 5 FIG. 1 is a flow chart of modeling a visual sequence of medium and long-range information according to an embodiment of the present invention. Figure 5 As shown, the modal conversion feature is modeled by a medium- and long-distance information visual sequence to obtain a second modeling feature, including: S501, performing feature tensor expansion on the modal conversion feature to obtain an expanded tensor; S502, performing linear mapping, convolution operation and bidirectional state space feature extraction on the unfolded tensor to obtain a second modeling feature.

[0042] Specifically, in the crowd ViM module, the embodiment first expands the modal conversion feature into a feature tensor to obtain an expanded tensor, and represents the high-dimensional tensor as the sum of simpler components. The expanded tensor obtained is then subjected to linear mapping, convolution operation, activation function, and bidirectional state space feature extraction in sequence to capture key features. Then, for the features of the two branches obtained by the bidirectional state space feature extraction, the features obtained by linear mapping and activation function of one of the branches with the expanded tensor are multiplied and fused, and then spliced ​​with the features obtained from the other branch. Finally, after a linear mapping and splicing with the initial expanded tensor, the tensor dimension is adjusted to obtain the second modeling feature output by the crowd ViM module.

[0043] In some embodiments of the present invention, peak point filtering is performed on the estimated density map to obtain a crowd counting recognition output, including: Pooling, threshold segmentation and non-maximum suppression are performed on the estimated density map to obtain the crowd counting recognition output.

[0044] Specifically, in the current density regression counting model, although the embodiment has made significant progress in density estimation through the above steps, the prediction of the model output still faces adjustment in positioning accuracy. This is due to the dispersion of the targets in the predicted density map, which not only affects the accuracy of the count, but also increases the complexity of subsequent processing. Therefore, the embodiment designs a positioning model that accurately identifies and locates the targets in the image by filtering the peak points of the estimated density map, thereby improving the counting accuracy and reliability of the model.

[0045] Through grouped pyramid, frequency domain conversion enhancement, visual information modeling and density regression output, the model can output high-quality fine density estimation map. In the density estimation map, the regressed Gaussian fuzzy peaks are almost independent of each other. Each peak position can be regarded as the center position of the target. By accurately filtering out the accurate peak points, the pedestrian target can be accurately located.

[0046] In the peak point filtering, the embodiment first applies a pooling operation to the generated estimated density map. This step not only helps to identify the initial peak point, that is, the possible center position of the pedestrian target, but also makes the density map visually smoother by reducing local fluctuations, laying the foundation for subsequent processing steps. Next, the embodiment performs threshold segmentation on the pooled density map. Through the pre-set threshold, background anomalies that do not belong to pedestrian targets can be effectively excluded, thereby achieving background suppression. The useful signal in the image is further purified. Finally, in order to accurately locate the center point of the target, the embodiment adopts a maximum suppression algorithm, which can identify the local maximum point, that is, the center coordinate point of the target, from the density map that has been threshold processed.

[0047] In some embodiments of the present invention, the trained crowd counting network is obtained by training an initial crowd counting network. Figure 6 FIG. 1 is a flow chart of training an initial crowd counting network according to an embodiment of the present invention. Figure 6 As shown, train the initial crowd counting network, including: S601, inputting crowd counting training data into an initial crowd counting network to obtain a predicted density map; S602, determining a parameterized error attenuation loss of an initial crowd counting network according to the predicted density map and the corresponding true density map, iteratively optimizing the initial crowd counting network according to the parameterized error attenuation loss until the loss no longer decreases, and obtaining a fully trained crowd counting network.

[0048] Specifically, in the crowd counting task of the density regression method, the loss function is mostly used between the predicted density map and the real density map. In order to better complete the supervision task, the embodiment adopts PED-Loss (Parameterized ErrorDecay Loss) as the loss function in the model training process. The PED-Loss loss function is a loss function used in regression tasks. It has a smooth characteristic and is robust to outliers. Its formula is:

[0049] in, and Respectively represent The predicted and true values ​​for each data point. is a weight parameter used to adjust the influence of the squared error term. is another weight parameter used to adjust the impact of the exponential decay term. is the rate of exponential decay, which controls The decay rate of the term. is the square of the difference between the predicted value and the true value, which helps emphasize large errors. is an exponential decay term, which decreases rapidly as the error increases, which helps to reduce the impact of extreme values.

[0050] In summary, the dense scene crowd counting method provided by the present invention first inputs the acquired crowd image to be detected into a well-trained crowd counting network, performs grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, performs frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, performs visual information modeling on the frequency domain enhancement features to obtain visual modeling enhancement features, performs density regression output on the visual modeling enhancement features to obtain an estimated density map; then performs peak point filtering on the estimated density map to obtain a crowd counting recognition output. The present invention can convert images to the frequency domain through frequency domain conversion enhancement to capture the contours and features of pedestrians, improve the accuracy of pedestrian features, and perform visual information modeling on pedestrian features at the global information level and the medium and long-distance information level through visual information enhancement to obtain more complete pedestrian feature information, thereby effectively improving the accuracy of crowd detection and counting.

[0051] In order to better implement the crowd counting method in a dense scene in the embodiment of the present invention, based on the crowd counting method in a dense scene, correspondingly, Figure 7 As shown, the present invention also provides a crowd counting device for dense scene, and the crowd counting device 700 for dense scene includes: The feature extraction prediction unit 701 is used to input the acquired crowd image to be detected into a well-trained crowd counting network, perform grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, perform visual information modeling on the frequency domain enhancement features to obtain visual modeling enhancement features, and perform density regression on the visual modeling enhancement features to output an estimated density map; The peak point filtering unit 702 is used to perform peak point filtering on the estimated density map to obtain a crowd counting recognition output.

[0052] The dense scene crowd counting device 700 provided in the above embodiment can implement the technical solution described in the above dense scene crowd counting method embodiment. The specific implementation principles of the above modules or units can refer to the corresponding contents in the above dense scene crowd counting method embodiment, which will not be repeated here.

[0053] like Figure 8 As shown, the present invention also provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802 and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0054] In some embodiments, the processor 801 may be a central processing unit (CPU), a microprocessor or other data processing chip, which is used to run the program code or process data stored in the memory 802, such as the crowd counting method in a dense scene of the present invention.

[0055] In some embodiments, the processor 801 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 801 may be local or remote. In some embodiments, the processor 801 may be implemented in a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination thereof.

[0056] In some embodiments, the memory 802 may be an internal storage unit of the electronic device 800, such as a hard disk or memory of the electronic device 800. In other embodiments, the memory 802 may also be an external storage device of the electronic device 800, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 800.

[0057] Furthermore, the memory 802 may include both an internal storage unit of the electronic device 800 and an external storage device. The memory 802 is used to store application software installed in the electronic device 800 and various data.

[0058] In some embodiments, the display 803 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 803 is used to display information of the electronic device 800 and to display a visual user interface. The components 801-803 of the electronic device 800 communicate with each other via a system bus.

[0059] In one embodiment, when the processor 801 executes the crowd counting program in the memory 802, the following steps may be implemented: The acquired crowd image to be detected is input into a well-trained crowd counting network, and the crowd image to be detected is subjected to grouped pyramid feature convolution to obtain preliminary convolution features, and the preliminary convolution features are enhanced by frequency domain conversion to obtain frequency domain enhanced features, and visual information modeling is performed on the frequency domain enhanced features to obtain visual modeling enhanced features, and density regression is performed on the visual modeling enhanced features to output an estimated density map; The estimated density map is peak-filtered to obtain the crowd counting recognition output.

[0060] It should be understood that: when the processor 801 executes the dense scene crowd counting program in the memory 802, in addition to the above functions, other functions can also be implemented. For details, please refer to the description of the corresponding method embodiment above.

[0061] Accordingly, an embodiment of the present application also provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the program or instructions are executed by a processor, the steps or functions of the dense scene crowd counting method provided in the above-mentioned method embodiments can be implemented.

[0062] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0063] The above is a detailed introduction to the crowd counting method, device, electronic device and storage medium for dense scenes provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for counting people in a dense scene, characterized in that: include: Inputting the acquired crowd image to be detected into a well-trained crowd counting network, performing grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, performing visual information modeling on the frequency domain enhancement features to obtain visual modeling enhancement features, and performing density regression on the visual modeling enhancement features to obtain an estimated density map; The estimated density map is subjected to peak point filtering to obtain a crowd counting recognition output.

2. The method for counting people in dense scenes according to claim 1, characterized in that: The performing frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhanced features includes: The preliminary convolution features are sequentially subjected to two-dimensional Fourier transform, complex convolution, complex batch normalization, complex GELU activation function and inverse Fourier transform to obtain frequency domain convolution features; The preliminary convolution features and the frequency domain convolution features are fused to obtain frequency domain enhancement features.

3. The method for counting people in dense scenes according to claim 1, characterized in that: The performing visual information modeling on the frequency domain enhanced features to obtain visual modeling enhanced features includes: Performing feature dimension conversion on the frequency domain enhancement feature to obtain a modal conversion feature; Performing global information visual conversion modeling on the modal conversion feature to obtain a first modeling feature, and performing medium- and long-range information visual sequence modeling on the modal conversion feature to obtain a second modeling feature; After fusing the first modeling feature and the second modeling feature, feature dimension conversion is performed to obtain a visual modeling enhancement feature.

4. The method for counting people in dense scenes according to claim 3, characterized in that: The performing global information visual conversion modeling on the modal conversion feature to obtain a first modeling feature includes: Segmenting the modal conversion feature to obtain a plurality of segmentation features; Multi-head attention feature extraction and multi-layer perceptron feature extraction are performed on the segmented features to obtain the first modeling features.

5. The method for counting people in dense scenes according to claim 3, characterized in that: The step of performing medium- and long-range information visual sequence modeling on the modal conversion feature to obtain a second modeling feature includes: Expanding the modal conversion feature into a feature tensor to obtain an expanded tensor; The expanded tensor is subjected to linear mapping, convolution operation and bidirectional state space feature extraction to obtain a second modeling feature.

6. The method for counting people in dense scenes according to claim 1, characterized in that: The peak point filtering of the estimated density map to obtain a crowd counting recognition output includes: Pooling operations, threshold segmentation and non-maximum suppression are performed on the estimated density map to obtain a crowd counting recognition output.

7. The method for counting people in dense scenes according to claim 1, characterized in that: The fully trained crowd counting network is obtained by training an initial crowd counting network, wherein the trained initial crowd counting network includes: Input the crowd counting training data into the initial crowd counting network to obtain the predicted density map; The parameterized error attenuation loss of the initial crowd counting network is determined according to the predicted density map and the corresponding true density map, and the initial crowd counting network is iteratively optimized according to the parameterized error attenuation loss until the loss no longer decreases, thereby obtaining a fully trained crowd counting network.

8. A crowd counting device in a dense scene, characterized in that: include: A feature extraction and prediction unit is used to input the acquired crowd image to be detected into a well-trained crowd counting network, perform grouped pyramid feature convolution on the crowd image to be detected to obtain preliminary convolution features, perform frequency domain conversion enhancement on the preliminary convolution features to obtain frequency domain enhancement features, perform visual information modeling on the frequency domain enhancement features to obtain visual modeling enhancement features, and perform density regression on the visual modeling enhancement features to output an estimated density map; The peak point filtering unit is used to perform peak point filtering on the estimated density map to obtain a crowd counting recognition output.

9. An electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for counting people in dense scenes according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for counting people in dense scenes according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Dense crowd counting method combining high-resolution CNN (Convolutional Neural Network) and lightweight Transformer

    CN116935316A

  • Anti-shielding head three-dimensional positioning method and system based on instance segmentation

    CN119205924A

  • Multi-scale Mama transformation unmanned aerial vehicle visual gesture flight control method

    CN119206795A

  • Image exposure adjustment method and device, equipment, medium and product

    CN119583969A

  • Epileptic electroencephalogram recognition system based on hierarchical graph convolutional neural network, terminal, and storage medium

    WO2021226778A1

Cited By

  • Video crowd counting method based on time sequence interaction and global association network

    CN122176645A