Population Density Statistics Method, Device, Equipment and Medium Based on Detection and Segmentation
Through SwinTransformer object detection network and multi-scale feature fusion technology, the mesoscale compatibility and background interference problems of population density statistics are solved, and more accurate population density assessment is achieved.
Patent Information
- Application Number
- CN202210645973.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-06-08
AI Technical Summary
The prior art has problems such as inability to compatible with human information of different sizes, ignoring spatial information, being unable to directly judge density, inaccurate occlusion detection, and background interference impact in population density statistics, resulting in inaccurate assessment of population density.
Feature extraction is performed based on SwinTransformer object detection network, and population probability density graph prediction is performed by combining deformable convolution blocks, convolutional residual networks, multi-scale feature pyramids, CBAM attention mechanism and FAM model. Background interference is eliminated through multi-scale feature fusion and attention mechanism.
It improves the prediction effect of the population density statistical model, can be compatible with human information of different sizes, eliminates background interference, and improves the accuracy and robustness of the model.
Smart Images

Figure CN114898301B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data statistics, and more specifically to a method, device, equipment and medium for crowd density statistics based on detection and segmentation. Background Art
[0002] With the development of society, crowd counting or statistics is a research hotspot and difficulty in the current industrial and academic fields, and it has important application value in real life. There are the following several ways for crowd density statistics currently.
[0003] The first is to perform feature fusion of information with different receptive fields through dilated convolution and original convolution, as well as feature fusion of information with different receptive fields, and fuse different hierarchical semantic information of feature maps at different resolutions, so as to generate a crowd density map with higher quality. This patent uses a large number of dilated convolutions, which results in loss of continuous information in the picture features, has a great impact on density statistics, and ignores the spatial information that the crowd density is larger near and smaller far away.
[0004] The second is to use the AlexNet network to divide the crowd picture dataset into two categories: density and sparsity, and send them into the corresponding feature extraction network according to the differences in the density features of the two types of images. For density images, the method of attention mechanism is used for personnel density statistics, and for sparse crowd density, the method of dilated convolution is used for personnel density statistics. This patent first needs to analyze the personnel density of the picture to determine whether the personnel density in the picture is sparse or dense, and then perform density analysis on it. This patent cannot directly judge the crowd density, and needs to judge its density sparsity and density first before making a judgment, the network is bloated and cannot adapt to different levels of personnel density classification.
[0005] The third is a method for crowd density and quantity estimation based on a convolutional neural network. This method only uses a convolutional neural network without using a multi-scale method, and cannot be compatible with the feature information of targets of different scales, the model prediction effect is poor, and the recognition result is inaccurate.
[0006] The fourth is to detect the number of human heads through an object detection network and judge the crowd density according to the number of human heads. This method of judging the crowd density based on object detection of human heads will have missed detections for human body occlusion and human head occlusion, and at the same time, the detection recall rate for distant dense small targets is low, which will ultimately lead to inaccurate crowd density evaluation. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method, device, equipment and medium for crowd density statistics based on detection and segmentation, which can increase the supervision information of target detection, improve the prediction effect of the model, be compatible with human body information of different sizes, and facilitate the model to more easily eliminate background interference factors.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In the first aspect, a method for crowd density statistics based on detection and segmentation includes:
[0010] Obtain image data;
[0011] Process the image data to obtain a sample image;
[0012] Input the sample image into a crowd density statistics model for crowd probability prediction to obtain a crowd density map;
[0013] Calculate the number of people according to the crowd density map.
[0014] A further technical solution thereof is: for the step of inputting the sample image into a crowd density statistics model for crowd probability prediction to obtain a crowd density map, the processing method of the crowd density statistics model includes:
[0015] Input the sample image into a Swin Transformer model for object detection of human head data to obtain first detection features, second detection features, third detection features and fourth detection features;
[0016] Input the first detection feature, second detection feature, third detection feature and fourth detection feature into a deformable convolution block for processing to obtain first processing features, second processing features, third processing features and fourth processing features;
[0017] Perform transposed convolution processing on the fourth processing feature and then perform upsampling processing to obtain a fifth processing feature;
[0018] Perform anti-pooling upsampling processing on the third processing feature to obtain a sixth processing feature;
[0019] Perform convolution processing on the second processing feature to obtain a seventh processing feature;
[0020] Input the sample image into a convolutional residual network for processing to obtain an eighth processing feature;
[0021] Concatenate and merge the eighth processing feature, seventh processing feature, sixth processing feature, fifth processing feature and first processing feature to obtain a first merged feature;
[0022] Input the first merged feature into the PPM model for processing to obtain the ninth processed feature;
[0023] Input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into the CBAM attention mechanism model for processing to obtain the tenth processed feature, the eleventh processed feature, the twelfth processed feature, and the thirteenth processed feature;
[0024] Input the tenth processed feature and the ninth processed feature into the FAM model for processing to obtain the fourteenth processed feature;
[0025] Input the fourteenth processed feature and the eleventh processed feature into the FAM model for processing to obtain the fifteenth processed feature;
[0026] Input the fifteenth processed feature and the twelfth processed feature into the FAM model for processing to obtain the sixteenth processed feature;
[0027] Input the sixteenth processed feature into the upsampling block for three times of upsampling processing to obtain the first sampled feature, the second sampled feature, and the third sampled feature;
[0028] Normalize the third sampled feature through the sigmoid function to obtain the crowd density map.
[0029] Its further technical solution is: Input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into the deformable convolution block for processing to obtain the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature. The deformable convolution block is respectively composed of a deformable convolution, a relu activation function, and BatchNormaliization.
[0030] Its further technical solution is: Input the sample map into the convolutional residual network for processing to obtain the eighth processed feature. The convolutional residual network is respectively composed of a residual convolution and a mish activation function.
[0031] Its further technical solution is: The step of inputting the first merged feature into the PPM model for processing to obtain the ninth processed feature includes:
[0032] Perform downsampling pooling on the first merged feature through the multi-scale feature pyramid respectively to obtain a plurality of downsampling pooling features;
[0033] Perform convolutional upsampling processing on the plurality of downsampling pooling features respectively to obtain a plurality of convolutional upsampling features;
[0034] Merge the plurality of convolutional upsampling features and then perform convolutional processing to obtain the ninth processed feature.
[0035] Its further technical solution is as follows: input the first processing feature, the second processing feature, the third processing feature, and the fourth processing feature into the CBAM attention mechanism model for processing to obtain the tenth processing feature, the eleventh processing feature, the twelfth processing feature, and the thirteenth processing feature. The CBAM attention mechanism model is composed of a channel attention mechanism and a spatial attention mechanism.
[0036] Its further technical solution is as follows: input the sixteenth processing feature into the upsampling block for three times of upsampling processing to obtain the first sampling feature, the second sampling feature, and the third sampling feature. The upsampling block is composed of a transposed convolution and a relu activation function.
[0037] In a second aspect, a crowd density statistical device based on detection and segmentation includes an acquisition unit, a processing unit, a prediction unit, and a calculation unit;
[0038] The acquisition unit is used to acquire image data;
[0039] The processing unit is used to process the image data to obtain a sample map;
[0040] The prediction unit is used to input the sample map into the crowd density statistical model for crowd probability prediction to obtain a crowd density map;
[0041] The calculation unit is used to calculate the number of people according to the crowd density map.
[0042] In a third aspect, a computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the crowd density statistical method based on detection and segmentation as described above are implemented.
[0043] In a fourth aspect, a computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the steps of the crowd density statistical method based on detection and segmentation as described above.
[0044] The beneficial effects of the present invention compared with the prior art are as follows: The present invention extracts features based on the Swin Transformer object detection network and predicts the crowd probability density map based on the extracted features, increasing the supervision information of object detection and improving the prediction effect of the model. Through multi-scale feature fusion, it can well accommodate human body information of different sizes. By adding an attention mechanism, the model's understanding of crowd density is deepened, making it easier for the model to eliminate background interference factors.
[0045] The above description is only an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically described in detail as follows. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of the application scenario of the crowd density statistics method based on detection and segmentation provided by a specific embodiment of the present invention;
[0048] Figure 2 Flowchart of the crowd density statistics method based on detection and segmentation provided by a specific embodiment of the present invention;
[0049] Figure 3 Schematic block diagram of the device for the crowd density statistics method based on detection and segmentation provided by a specific embodiment of the present invention;
[0050] Figure 4 Schematic block diagram of a computer device provided by a specific embodiment of the present invention. Detailed Embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0052] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups.
[0053] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0054] It should be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0055] Please refer to Figure 1 and Figure 2 , Figure 1 FIG. is a schematic diagram of an application scenario of a crowd density statistical method based on detection and segmentation provided for a specific embodiment of the present invention; Figure 2 FIG. is a flowchart of a crowd density statistical method based on detection and segmentation provided for a specific embodiment of the present invention. The crowd density statistical method based on detection and segmentation is applied to a server and is executed by an application software installed in the server.
[0056] As Figure 2 shown, the crowd density statistical method based on detection and segmentation includes the following steps: S10-S40.
[0057] S10. Obtain image data.
[0058] In this embodiment, video data (i.e., image data) of the crowd in the subway car is collected by a monitoring device in the subway car. The monitoring device can be a common one on the market, and the present application does not make any limitations in this regard. In order to collect the video data of each car, monitoring devices can be installed in each car, and the video data collected by the monitoring devices installed in each car can be aggregated into the data background of the subway in a wired or wireless manner. By accessing the data background, the video data situation of each car can be queried.
[0059] S20. Process the image data to obtain a sample image.
[0060] In one embodiment, step S20 specifically includes the following steps: S201-S202.
[0061] S201. Segment the image data to obtain segmented image data.
[0062] In this embodiment, since the subway needs to stop at different stations, the boarding or alighting situation of each car basically changes after each stop at each station. Therefore, the image data can be segmented in the manner of each car corresponding to each station, and the segmented image data corresponding to each car at each station can be obtained.
[0063] S202. Select a frame of picture from the segmented image data as the sample image.
[0064] In this embodiment, since the segmented image data includes multiple frames of pictures, a frame of picture can be selected from the segmented image data as the sample image Iimage Perform head probability prediction.
[0065] S30. Input the sample image into the crowd density statistical model to perform crowd probability prediction, so as to obtain a crowd density map.
[0066] In one embodiment, the processing method of the crowd density statistical model specifically includes the following steps: S301 - S315.
[0067] S301. Input the sample image into the Swin Transformer model to perform object detection on the head data, so as to obtain the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature.
[0068] In this embodiment, input the sample image I image into the Swin Transformer model to perform object detection on the head data. There are a total of 4 stages in the Swin Transformer here, and thus the features ST1, ST2, ST3, and ST4 are obtained respectively.
[0069] The head data refers to the data of the human head in the image.
[0070] Swin Transformer is a hierarchical structure, similar to fpn, which extracts visual features at different levels to make it more suitable for tasks such as segmentation detection.
[0071] The overall structure of the Swin Transformer is similar to the hierarchical structure of convolution. The resolution becomes half at each layer, while the number of channels doubles. First, Patch Partition, which is the operation of equally dividing into small blocks in VIT; then it is divided into 4 stages, and each stage includes two parts, namely patch Merging (the first block is a linear layer) and the Swin Transformer Block. Patch Merging is an operation similar to pooling. Pooling will lose information, while patch Merging will not.
[0072] Feature extraction is performed based on the Swin Transformer object detection network, and crowd probability density map prediction is performed based on the extracted features, which increases the supervision information of object detection and improves the prediction effect of the model.
[0073] S302. Input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into the deformable convolution block for processing, so as to obtain the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature.
[0074] In this embodiment, the feature ST1, feature ST2, feature ST3, and feature ST4 are input into a deformable convolution block for processing to obtain feature T1, feature T2, feature T3, and feature T4.
[0075] The deformable convolution block is respectively composed of a deformable convolution, a relu activation function, and BatchNormaliization.
[0076] S303. Perform a transposed convolution process on the fourth processed feature and then perform an upsampling process to obtain a fifth processed feature.
[0077] In this embodiment, after performing a transposed convolution process on feature T4 and then performing an upsampling, feature C4 is obtained.
[0078] S304. Perform an anti-pooling upsampling process on the third processed feature to obtain a sixth processed feature.
[0079] In this embodiment, after performing an anti-pooling upsampling process on feature T3, feature C3 is obtained.
[0080] S305. Perform a convolution process on the second processed feature to obtain a seventh processed feature.
[0081] In this embodiment, for feature T2, a convolution process is performed to obtain feature C2.
[0082] S306. Input the sample image into a convolutional residual network for processing to obtain an eighth processed feature.
[0083] In this embodiment, the sample image I image is input into a convolutional residual network for processing to obtain feature CR1.
[0084] The convolutional residual network is respectively composed of a residual convolution and a mish activation function.
[0085] S307. Concatenate and merge the eighth processed feature, the seventh processed feature, the sixth processed feature, the fifth processed feature, and the first processed feature to obtain a first merged feature.
[0086] In this embodiment, feature CR1, feature C4, feature C3, feature C2, and feature T1 are concatenated and merged to obtain feature CT.
[0087] S308. Input the first merged feature into a PPM model for processing to obtain a ninth processed feature.
[0088] In this embodiment, feature CT is input into a PPM (Pyramid Pooling module) model for processing to obtain feature T0.
[0089] ppm is a special pooling model. Through pooling from more to less, the receptive field can be effectively increased, and the utilization efficiency of global information can be enhanced.
[0090] In one embodiment, step S308 specifically includes the following steps: S3081 - S3083.
[0091] S3081. Respectively perform downsampling pooling on the first merged feature through a multi - scale feature pyramid to obtain a plurality of downsampled pooling features.
[0092] S3082. Respectively perform convolutional upsampling processing on the plurality of downsampled pooling features to obtain a plurality of convolutional upsampling features.
[0093] S3083. After merging the plurality of convolutional upsampling features, perform convolutional processing to obtain the ninth processed feature.
[0094] In this embodiment, the feature CT is subjected to downsampling pooling through different - scale pyramids, then respectively subjected to convolutional upsampling and merging, and then convolutional processing to obtain the feature T0.
[0095] S309. Input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into the CBAM attention mechanism model for processing to obtain the tenth processed feature, the eleventh processed feature, the twelfth processed feature, and the thirteenth processed feature.
[0096] In this embodiment, the features T1, T2, T3, and T4 are input into the CBAM attention mechanism model to obtain the features TCM1, TCM2, TCM3, and TCM4.
[0097] The CBAM attention mechanism model is composed of a channel attention mechanism and a spatial attention mechanism.
[0098] By adding an attention mechanism, the model's understanding of crowd density is deepened, making it easier for the model to eliminate background interference factors.
[0099] S310. Input the tenth processed feature and the ninth processed feature into the FAM model for processing to obtain the fourteenth processed feature.
[0100] In this embodiment, the feature TCM1 and the feature T0 are input into the FAM (Flow Align Moudle) model to obtain the feature FAM2.
[0101] FAM is the semantic feature flow between adjacent stages of the model, which propagates semantic feature information to high - resolution spatial information, so that the features contain both semantic information and spatial information.
[0102] S311. Input the fourteenth processing feature and the eleventh processing feature into the FAM model for processing to obtain the fifteenth processing feature.
[0103] In this embodiment, the feature TCM2 and the feature FAM2 are input into the FAM model to obtain the feature FAM3.
[0104] S312. Input the fifteenth processing feature and the twelfth processing feature into the FAM model for processing to obtain the sixteenth processing feature.
[0105] In this embodiment, the feature TCM3 and the feature FAM3 are input into the FAM model to obtain the feature FAM4.
[0106] S314. Input the sixteenth processing feature into the upsampling block for three times of upsampling processing to obtain the first sampling feature, the second sampling feature, and the third sampling feature.
[0107] In this embodiment, the feature FAM4 is input into the upsampling block for upsampling. The upsampling block consists of a transposed convolution and a relu activation function. After 3 times of upsampling, the features F c1 , feature F c2 and feature F c3 are obtained. Among them, the size of the feature remains the same as the size of the sample image.
[0108] S315. Normalize the third sampling feature through the sigmoid function to obtain the crowd density map.
[0109] In this embodiment, the feature F c3 is normalized through the sigmoid function to obtain the output crowd density map F out . By setting each pixel value of the crowd density map F out between 0 and 1, it is the probability of a human head.
[0110] S40. Calculate the number of people according to the crowd density map.
[0111] In this embodiment, by accumulating and summing the probability density of each pixel on the crowd density map F out and taking the integer, the number of people in the picture is obtained.
[0112] In addition, the loss functions used in the crowd density statistical model are the object detection function, the crowd density segmentation function, and the total loss function. Among them, the object detection function performs object detection based on the features ST1, ST2, ST3, and ST4 output by the Swin Transformer. The loss here includes three losses, namely: the classification loss function, the regression loss function, and the giou loss function, that is, Loss od= Loss classification + Loss regression + Loss giou 。
[0113] The crowd density segmentation function uses the L2 loss function, and the loss function is as follows:
[0114]
[0115] B is the batch size, represents the crowd density annotation map, represents the crowd density prediction map. Resize F c2 , F c3 to the same size as the original image through interpolation resizing and use a 1x1 convolution to make its number of channels 1. After normalization by the sigmoid function, calculate the loss function of each with the true crowd density label map respectively. This is done to better accelerate the model training speed. At the same time, calculate the loss function of F out with the true crowd density label map to obtain Loss crowd density-1 , Loss crowd density-2 and Loss crowd density-3 respectively. The total crowd density loss is Loss crowd density-total = Loss crowd density-1 + Loss crowd density-2 + Loss crowd density-3 。
[0116] The total loss function is:
[0117] Loss total = αLoss od + βLoss crowd density ;
[0118] where α is 0.2 and β is 0.8.
[0119] The evaluation functions for the model are the MAE evaluation function and the RMSE evaluation function respectively.
[0120]
[0121]
[0122] where C i and respectively represent the true number of people, represents the predicted number of people.
[0123] The present invention extracts features based on the Swin Transformer object detection network, and predicts the crowd probability density map based on the extracted features, increasing the supervision information of object detection and improving the prediction effect of the model. Through multi-scale feature fusion, it can well accommodate human body information of different sizes. By adding an attention mechanism, the model's understanding of crowd density is deepened, making it easier for the model to eliminate background interference factors.
[0124] Figure 3 FIG. 4 is a schematic block diagram of a crowd density statistical device 100 based on detection and segmentation provided by an embodiment of the present invention. Corresponding to the above-mentioned crowd density statistical method based on detection and segmentation, a specific embodiment of the present invention further provides a crowd density statistical device 100 based on detection and segmentation. The crowd density statistical device 100 based on detection and segmentation includes units and modules for executing the above-mentioned crowd density statistical method based on detection and segmentation, and this device can be configured in a server.
[0125] As Figure 3 described, the crowd density statistical device 100 based on detection and segmentation includes an acquisition unit 110, a processing unit 120, a prediction unit 130, and a calculation unit 140.
[0126] The acquisition unit 110 is used to acquire image data.
[0127] In this embodiment, video data (i.e., image data) of the crowd in the subway car is collected through monitoring devices in the subway car. The monitoring devices can be common ones on the market, and the present application does not make any limitations in this regard. In order to collect the video data of each car, monitoring devices can be installed in each car, and the video data collected by the monitoring devices installed in each car can be aggregated into the subway data background through wired or wireless means. By accessing the data background, the video data situation of each car can be queried.
[0128] The processing unit 120 is used to process the image data to obtain a sample map.
[0129] In one embodiment, the processing unit 120 includes a splitting module and a selection module.
[0130] The splitting module is used to split the image data to obtain split image data.
[0131] In this embodiment, since the subway needs to stop at different stations, the boarding or alighting situation of each car basically changes after each stop at each station. Therefore, the image data can be split in the way of each car corresponding to each station, and the split image data of each car corresponding to each station can be obtained.
[0132] A selection module for selecting a frame of picture from the segmented image data as a sample picture.
[0133] In this embodiment, since the segmented image data includes multiple frames of pictures, a frame of picture can be selected from the segmented image data as sample picture I. image Perform head probability prediction.
[0134] A prediction unit 130 for inputting the sample picture into a crowd density statistical model to perform crowd probability prediction to obtain a crowd density map.
[0135] In one embodiment, the prediction unit 130 includes a detection module, a first processing module, a second processing module, a third processing module, a fourth processing module, a fifth processing module, a merging module, a sixth processing module, a seventh processing module, an eighth processing module, a ninth processing module, a tenth processing module, an eleventh processing module, and a twelfth processing module.
[0136] The detection module is used to input the sample picture into the SwinTransformer model to perform object detection on the head data to obtain a first detection feature, a second detection feature, a third detection feature, and a fourth detection feature.
[0137] In this embodiment, the sample picture I image is input into the SwinTransformer model to perform object detection on the head data. There are a total of 4 stages in the SwinTransformer here, and thus features ST1, ST2, ST3, and ST4 are obtained respectively.
[0138] The head data refers to the data of the human head in the picture.
[0139] SwinTransformerr is a hierarchical structure, similar to fpn, which extracts visual features at different levels to make it more suitable for tasks such as segmentation detection.
[0140] The overall structure of the Swin Transformer is similar to the hierarchical structure of convolution. The resolution becomes half at each layer, while the number of channels doubles. First, Patch Partition, which is the operation of equally dividing into small blocks in VIT; then it is divided into 4 stages, and each stage includes two parts, namely patch Merging (the first block is a linear layer) and the SwinTransformer Block. Patch Merging is an operation similar to pooling. Pooling will lose information, while patch Merging will not.
[0141] Feature extraction is performed based on the Swin Transformer object detection network, and crowd probability density map prediction is performed based on the extracted features, increasing the supervision information of object detection and improving the prediction effect of the model.
[0142] The first processing module is used to input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into a deformable convolutional block for processing to obtain the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature.
[0143] In this embodiment, the features ST1, ST2, ST3, and ST4 are input into a deformable convolutional block for processing to obtain the features T1, T2, T3, and T4.
[0144] The deformable convolutional block is respectively composed of a deformable convolution, a relu activation function, and Batch Normalization.
[0145] The second processing module is used to perform transposed convolution processing on the fourth processed feature and then perform upsampling processing to obtain the fifth processed feature.
[0146] In this embodiment, transposed convolution processing is performed on the feature T4 and then upsampling is performed to obtain the feature C4.
[0147] The third processing module is used to perform anti-pooling upsampling processing on the third processed feature to obtain the sixth processed feature.
[0148] In this embodiment, anti-pooling upsampling processing is performed on the feature T3 to obtain the feature C3.
[0149] The fourth processing module is used to perform convolution processing on the second processed feature to obtain the seventh processed feature.
[0150] In this embodiment, convolution processing is performed on the feature T2 to obtain the feature C2.
[0151] The fifth processing module is used to input the sample image into a convolutional residual network for processing to obtain the eighth processed feature.
[0152] In this embodiment, the sample image I image is input into a convolutional residual network for processing to obtain the feature CR1.
[0153] The convolutional residual network is respectively composed of a residual convolution and a mish activation function.
[0154] The merging module is used to perform concate merging on the eighth processed feature, the seventh processed feature, the sixth processed feature, the fifth processed feature, and the first processed feature to obtain the first merged feature.
[0155] In this embodiment, features CR1, C4, C3, C2, and T1 are concatenated to obtain feature CT.
[0156] The sixth processing module is configured to input the first merged feature into the PPM model for processing to obtain the ninth processed feature.
[0157] In this embodiment, feature CT is input into the PPM (Pyramid Pooling module) model for processing to obtain feature T0.
[0158] PPM is a special pooling model. Through pooling from more to less, the receptive field can be effectively increased, and the utilization efficiency of global information can be increased.
[0159] In one embodiment, the sixth processing module includes a first processing sub-module, a second processing sub-module, and a third processing sub-module.
[0160] The first processing sub-module is configured to perform downsampling pooling on the first merged feature through a multi-scale feature pyramid to obtain a plurality of downsampled pooling features.
[0161] The second processing sub-module is configured to perform convolutional upsampling processing on the plurality of downsampled pooling features respectively to obtain a plurality of convolutional upsampling features.
[0162] The third processing sub-module is configured to merge the plurality of convolutional upsampling features and then perform convolutional processing to obtain the ninth processed feature.
[0163] In this embodiment, feature CT is subjected to downsampling pooling through different-scale pyramids, then convolutional upsampling and merging are performed respectively, and then convolutional processing is performed to obtain feature T0.
[0164] The seventh processing module is configured to input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into the CBAM attention mechanism model for processing to obtain the tenth processed feature, the eleventh processed feature, the twelfth processed feature, and the thirteenth processed feature.
[0165] In this embodiment, features T1, T2, T3, and T4 are input into the CBAM attention mechanism model to obtain features TCM1, TCM2, TCM3, and TCM4.
[0166] The CBAM attention mechanism model is composed of a channel attention mechanism and a spatial attention mechanism.
[0167] By adding an attention mechanism, the model's understanding of crowd density is deepened, making it easier for the model to eliminate background interference factors.
[0168] The eighth processing module is used to input the tenth processing feature and the ninth processing feature into the FAM model for processing to obtain the fourteenth processing feature.
[0169] In this embodiment, the feature TCM1 and the feature T0 are input into the FAM (FlowAlign Moudle) model to obtain the feature FAM2.
[0170] FAM is the semantic feature flow of adjacent stages of the model, which propagates the semantic feature information to the high-resolution spatial information, so that the feature contains both semantic information and spatial information.
[0171] The ninth processing module is used to input the fourteenth processing feature and the eleventh processing feature into the FAM model for processing to obtain the fifteenth processing feature.
[0172] In this embodiment, the feature TCM2 and the feature FAM2 are input into the FAM model to obtain the feature FAM3.
[0173] The tenth processing module is used to input the fifteenth processing feature and the twelfth processing feature into the FAM model for processing to obtain the sixteenth processing feature.
[0174] In this embodiment, the feature TCM3 and the feature FAM3 are input into the FAM model to obtain the feature FAM4.
[0175] The eleventh processing module is used to input the sixteenth processing feature into the upsampling block for three times of upsampling processing to obtain the first sampling feature, the second sampling feature and the third sampling feature.
[0176] In this embodiment, the feature FAM4 is input into the upsampling block for upsampling. The upsampling block is composed of a transposed convolution and a relu activation function. After 3 times of upsampling, the features F c1 、feature F c2 and feature F c3 are obtained respectively. Among them, the size of the feature remains the same as the size of the sample graph.
[0177] The twelfth processing module is used to normalize the third sampling feature through the sigmoid function to obtain the crowd density map.
[0178] In this embodiment, the feature F c3 is normalized through the sigmoid function to obtain the output crowd density map F out . By setting that each pixel value of the crowd density map F out is between 0 and 1, it is the probability of a human head.
[0179] The calculation unit 140 is used to calculate the number of people according to the crowd density map.
[0180] In this embodiment, by summing up and rounding the probability density of each pixel on the crowd density map F out the number of people in the picture can be obtained.
[0181] In addition, the loss functions used in the crowd density statistical model are the object detection function, the crowd density segmentation function, and the total loss function. Among them, the object detection function performs object detection based on the features ST1, ST2, ST3, and ST4 output by Swin Transformer. The loss here includes three losses, namely: the classification loss function, the regression loss function, and the giou loss function, that is, Loss od = Loss classification + Loss regression + Loss giou .
[0182] The crowd density segmentation function uses the L2 loss function, and the loss function is as follows:
[0183]
[0184] B is the batch size, represents the crowd density annotation map, represents the crowd density prediction map. Let F c2 , F c3 Through interpolation and resize, this feature is resized to the same size as the original image and a 1x1 convolution is used to make its number of channels 1. After normalization by the sigmoid function, the loss functions with the true crowd density labels are calculated respectively. This is done to better accelerate the model training speed. At the same time, the loss functions of F out and the true crowd density label map are calculated to obtain Loss crowd density-1 , Loss crowd density-2 and Loss crowd density-3 respectively. The total crowd density loss is Loss crowd density-total = Loss crowd density-1 + Loss crowd density-2 + Loss crowd density-3 .
[0185] The total loss function is:
[0186] Loss total = αLoss od + βLoss crowd density ;
[0187] Among them, α is 0.2 and β is 0.8.
[0188] The evaluation functions for the model are the MAE evaluation function and the RMSE evaluation function respectively.
[0189]
[0190]
[0191] Among them, C i and respectively represent the actual population quantity, which represents the predicted population quantity.
[0192] The present invention extracts features based on the Swin Transformer object detection network, and predicts the crowd probability density map based on the extracted features, increasing the supervision information of object detection and improving the prediction effect of the model. Through multi-scale feature fusion, it can well accommodate human body information of different sizes. By adding an attention mechanism, the model's understanding of crowd density is deepened, facilitating the model to more easily eliminate background interference factors.
[0193] The above-mentioned crowd density statistics device based on detection and segmentation can be implemented in the form of a computer program, and this computer program can run on a computer device as Figure 4 shown.
[0194] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.
[0195] As Figure 4 shown, this computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned crowd density statistics method based on detection and segmentation.
[0196] The computer device 700 can be a terminal or a server. The computer device 700 includes a processor 720, a memory, and a network interface 750 connected through a system bus 710. Among them, the memory can include a non-volatile storage medium 730 and an internal memory 740.
[0197] The non-volatile storage medium 730 can store an operating system 731 and a computer program 732. When the computer program 732 is executed, it can cause the processor 720 to execute any crowd density statistics method based on detection and segmentation.
[0198] The processor 720 is used to provide computing and control capabilities to support the operation of the entire computer device 700.
[0199] The internal memory 740 provides an environment for the operation of the computer program 732 in the non-volatile storage medium 730. When the computer program 732 is executed by the processor 720, the processor 720 can be caused to execute any one of the crowd density statistical methods based on detection and segmentation.
[0200] The network interface 750 is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 700 to which the solution of this application is applied. Specifically, the computer device 700 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Among them, the processor 720 is used to run the program code stored in the memory to implement the following steps:
[0201] Obtain image data;
[0202] Process the image data to obtain a sample map;
[0203] Input the sample map into the crowd density statistical model for crowd probability prediction to obtain a crowd density map;
[0204] Calculate the number of people according to the crowd density map.
[0205] In one embodiment: for the step of inputting the sample map into the crowd density statistical model for crowd probability prediction to obtain a crowd density map, the processing method of the crowd density statistical model includes:
[0206] Input the sample map into the Swin Transformer model for object detection of human head data to obtain the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature;
[0207] Input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into the deformable convolution block for processing to obtain the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature;
[0208] Perform transposed convolution processing on the fourth processed feature and then perform upsampling processing to obtain the fifth processed feature;
[0209] Perform anti-pooling upsampling processing on the third processed feature to obtain the sixth processed feature;
[0210] Perform convolution processing on the second processed feature to obtain the seventh processed feature;
[0211] Input the sample image into the convolutional residual network for processing to obtain the eighth processed feature;
[0212] Concatenate the eighth processed feature, the seventh processed feature, the sixth processed feature, the fifth processed feature, and the first processed feature to obtain the first concatenated feature;
[0213] Input the first concatenated feature into the PPM model for processing to obtain the ninth processed feature;
[0214] Input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into the CBAM attention mechanism model for processing to obtain the tenth processed feature, the eleventh processed feature, the twelfth processed feature, and the thirteenth processed feature;
[0215] Input the tenth processed feature and the ninth processed feature into the FAM model for processing to obtain the fourteenth processed feature;
[0216] Input the fourteenth processed feature and the eleventh processed feature into the FAM model for processing to obtain the fifteenth processed feature;
[0217] Input the fifteenth processed feature and the twelfth processed feature into the FAM model for processing to obtain the sixteenth processed feature;
[0218] Input the sixteenth processed feature into the upsampling block for three times of upsampling processing to obtain the first sampled feature, the second sampled feature, and the third sampled feature;
[0219] Normalize the third sampled feature through the sigmoid function to obtain the crowd density map.
[0220] In one embodiment: Input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into the deformable convolutional block for processing to obtain the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature. The deformable convolutional block is respectively composed of a deformable convolution, a relu activation function, and BatchNormaliization.
[0221] In one embodiment: Input the sample image into the convolutional residual network for processing to obtain the eighth processed feature. The convolutional residual network is respectively composed of a residual convolution and a mish activation function.
[0222] In one embodiment: The step of inputting the first concatenated feature into the PPM model for processing to obtain the ninth processed feature includes:
[0223] Perform downsampling pooling on the first concatenated feature through the multi-scale feature pyramid respectively to obtain a plurality of downsampled pooling features;
[0224] Perform convolution upsampling processing on multiple downsampled pooling features respectively to obtain multiple convolution upsampled features;
[0225] Merge multiple convolution upsampled features and then perform convolution processing to obtain a ninth processed feature.
[0226] In one embodiment: input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into the CBAM attention mechanism model for processing to obtain the tenth processed feature, the eleventh processed feature, the twelfth processed feature, and the thirteenth processed feature. The CBAM attention mechanism model is composed of a channel attention mechanism and a spatial attention mechanism.
[0227] In one embodiment: input the sixteenth processed feature into an upsampling block for three times of upsampling processing to obtain a first sampling feature, a second sampling feature, and a third sampling feature. The upsampling block is composed of a transposed convolution and a relu activation function.
[0228] It should be understood that in the embodiments of the present application, the processor 720 may be a central processing unit (CPU), and this processor 720 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0229] Those skilled in the art can understand that Figure 4 the structure of the computer device 700 shown in
[0230] does not limit the computer device 700, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0231] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0232] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods, or units with the same function can be aggregated into one unit. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, or can be electrical, mechanical, or other forms of connection.
[0233] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0234] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0235] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs.
[0236] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for crowd density statistics based on detection and segmentation, characterized in that Including: Obtain image data; Process the image data to obtain a sample image; Input the sample image into a crowd density statistical model for crowd probability prediction to obtain a crowd density map; Calculate the number of people according to the crowd density map; The step of inputting the sample image into a crowd density statistical model for crowd probability prediction to obtain a crowd density map, and the processing method of the crowd density statistical model includes: Input the sample image into a Swin Transformer model for object detection of human head data to obtain a first detection feature, a second detection feature, a third detection feature, and a fourth detection feature; Input the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into a deformable convolutional block for processing to obtain a first processed feature, a second processed feature, a third processed feature, and a fourth processed feature; Perform transposed convolution processing on the fourth processed feature and then perform upsampling processing to obtain a fifth processed feature; Perform anti-pooling upsampling processing on the third processed feature to obtain a sixth processed feature; Perform convolutional processing on the second processed feature to obtain a seventh processed feature; Input the sample image into a convolutional residual network for processing to obtain an eighth processed feature; Concatenate and merge the eighth processed feature, the seventh processed feature, the sixth processed feature, the fifth processed feature, and the first processed feature to obtain a first merged feature; Input the first merged feature into a PPM model for processing to obtain a ninth processed feature; Input the first processed feature, the second processed feature, the third processed feature, and the fourth processed feature into a CBAM attention mechanism model for processing to obtain a tenth processed feature, an eleventh processed feature, a twelfth processed feature, and a thirteenth processed feature; Input the tenth processed feature and the ninth processed feature into a FAM model for processing to obtain a fourteenth processed feature; Input the fourteenth processed feature and the eleventh processed feature into a FAM model for processing to obtain a fifteenth processed feature; Input the fifteenth processed feature and the twelfth processed feature into a FAM model for processing to obtain a sixteenth processed feature; Input the sixteenth processed feature into an upsampling block for three times of upsampling processing to obtain a first sampled feature, a second sampled feature, and a third sampled feature; Normalize the third sampled feature through a sigmoid function to obtain a crowd density map; The step of inputting the first detection feature, the second detection feature, the third detection feature, and the fourth detection feature into a deformable convolutional block for processing to obtain a first processed feature, a second processed feature, a third processed feature, and a fourth processed feature, and the deformable convolutional block is respectively composed of a deformable convolution, a relu activation function, and BatchNormalization; The step of inputting the sample image into a convolutional residual network for processing to obtain an eighth processed feature, and the convolutional residual network is respectively composed of a residual convolution and a mish activation function.
2. The method for crowd density statistics based on detection and segmentation according to claim 1, wherein The step of inputting the first merged feature into a PPM model for processing to obtain a ninth processed feature includes: The first merged feature is respectively downsampled and pooled through a multi-scale feature pyramid to obtain multiple downsampled and pooled features; The multiple downsampled and pooled features are respectively subjected to convolutional upsampling processing to obtain multiple convolutional upsampling features; The multiple convolutional upsampling features are merged and then subjected to convolutional processing to obtain a ninth processed feature.
3. The method for crowd density statistics based on detection and segmentation according to claim 1, characterized in that, The first processed feature, the second processed feature, the third processed feature, and the fourth processed feature are input into the CBAM attention mechanism model for processing to obtain a tenth processed feature, an eleventh processed feature, a twelfth processed feature, and a thirteenth processed feature. The CBAM attention mechanism model is composed of a channel attention mechanism and a spatial attention mechanism.
4. The method for crowd density statistics based on detection and segmentation according to claim 1, wherein The sixteenth processed feature is input into an upsampling block for three times of upsampling processing to obtain a first sampled feature, a second sampled feature, and a third sampled feature. The upsampling block is composed of a transposed convolution and a relu activation function.
5. A crowd density statistical device based on detection and segmentation, when it runs, executes the crowd density statistical method based on detection and segmentation according to any one of claims 1-4, characterized in that, It includes an acquisition unit, a processing unit, a prediction unit, and a calculation unit; The acquisition unit is used to acquire image data; The processing unit is used to process the image data to obtain a sample map; The prediction unit is used to input the sample map into a crowd density statistical model for crowd probability prediction to obtain a crowd density map; The calculation unit is used to calculate the number of people according to the crowd density map.
6. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the crowd density statistical method according to any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the processor executes the steps of the crowd density statistical method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Crowd density estimation device and method and storage medium
CN113869285A
Satellite image small target detection method based on improved YOLOv5
CN114220015A