Crowd density estimation method and system oriented to traffic hub scene
By building a three-branch structure density grading population estimation network in the transportation hub scenario, the problem of inaccurate population density estimation is solved, accurate estimation and real-time early warning of population density are achieved, and the safety of transportation hubs is improved.
Patent Information
- Application Number
- CN202510300964.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-26
AI Technical Summary
The population density estimation results of the prior art in the transportation hub scenario are not accurate enough, the complex background is mistaken for the prospect and affects the calculation results, and the uneven distribution of population density leads to frequent safety accidents.
A three-branch structure density hierarchical population estimation network is built using a prospect segmentation method based on deep learning, including feature extraction network, prospect hierarchical network and scale factor extraction network. Combined with the density estimation network, predicted density maps are generated and real-time comparison and early warning are performed.
It improves the accuracy and safety of population density estimation, can promptly warn of high-density areas, reduce safety accidents, and improve the safety management level of transportation hubs.
Smart Images

Figure CN120544110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video image processing, and in particular to a crowd density estimation method and system for a transportation hub scenario. Background Art
[0002] As regional economic, political, and cultural centers, cities host diverse and frequent daily travel needs for urban residents, including commuting, shopping, and entertainment. However, challenges such as overcrowded public transportation, dense crowds in public spaces, and road traffic congestion, which hinder travel comfort, are becoming increasingly prominent.
[0003] Currently, many cities have introduced real-time travel indicators such as "congestion" and "traffic index." These indicators allow city residents to make informed travel decisions, making the implementation of measures such as crowd control and road traffic diversion more targeted. However, complex backgrounds in images are often mistaken by the network as foreground and included in the calculation. Furthermore, uneven density distribution in crowd areas can lead to inaccurate network calculations. Therefore, a crowd density estimation system for foreground segmentation in transportation hub scenarios is urgently needed. This system not only filters out background data but also classifies foreground areas into different levels of crowd density. Furthermore, it can provide early warnings for high-density areas, thereby reducing the occurrence of safety accidents and improving the intelligent level of transportation hub safety management. Summary of the Invention
[0004] Technical problems solved
[0005] In response to the above-mentioned shortcomings of the existing technology, the present invention provides a crowd density estimation method and system for transportation hub scenarios, which can effectively solve the problem of inaccurate calculation results of existing crowd density monitoring networks.
[0006] Technical Solution
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0008] In one aspect, the present invention provides a method for estimating crowd density in a transportation hub scenario, the method comprising the following steps:
[0009] Step S1, obtaining a large amount of crowd image data of different densities from surveillance videos of transportation hubs and establishing a crowd density dataset;
[0010] Step S2, building a density-graded crowd estimation network based on foreground segmentation, and inputting the crowd density dataset into the density-graded crowd estimation network for training;
[0011] The trained density-graded crowd estimation network is divided into two parts: the front-end and the back-end. The front-end adopts a three-branch structure, namely the feature extraction network, the foreground classification network and the scale factor extraction network, and the back-end is the density estimation network.
[0012] Step S3, inputting the crowd image to be processed into the feature extraction network to extract pedestrian head features to form an intermediate density map;
[0013] Input the crowd image into the scale factor extraction network to make a global crowd density level estimate, and assign corresponding scale factors according to different densities;
[0014] The foreground classification network is used to separate the crowd area and background of the input crowd image, and the crowd area is divided into local density levels to generate a graded foreground information map;
[0015] Step S4, generating a predicted density map using the density estimation network according to the obtained intermediate density map, scale factor, and graded foreground information map of the crowd image;
[0016] Step S5: Generate a predicted density map based on the above, and compare it with a preset crowd density threshold in real time and issue an early warning.
[0017] In one aspect, the present invention provides a crowd density estimation system for a transportation hub scenario, comprising:
[0018] The acquisition module is used to obtain the crowd image and crowd density data set to be processed and perform preliminary processing of the image data;
[0019] A model generation and execution module is used to build and train the density-graded crowd estimation network, and to calculate the input crowd image to be processed and output a predicted density map;
[0020] A counting and display module is used to estimate and count the number of people in the predicted density map and display the estimated results in real time;
[0021] The early warning module compares the crowd density value in the generated predicted density map with the set density threshold to determine whether it exceeds the safety threshold and initiates an automatic early warning.
[0022] Furthermore, in step S1, crowd image data is obtained through real-time monitoring videos of each transportation hub, and the collected crowd image data is preliminarily processed by denoising, enhancing and standardizing.
[0023] Furthermore, the feature extraction network consists of the first ten layers of VGG-16 cascaded with an upsampling layer and a convolution layer. The upsampling layer is used to amplify the feature map reduced by the pooling operation, and then perform a convolution operation through the convolution layer to fuse the upsampled local features.
[0024] Furthermore, the scale factor extraction network uses the first 10 layers of VGG16 and adds an average pooling layer after it to facilitate better evaluation of global head density. Then, four conv3×1 convolutional layers are added to change the output channel to 1. After the convolution operation, an adaptive average pooling operation is performed to quantize the feature map output by the network into a scale factor i. The scale factor value is then adjusted to between 0 and 2 through the following operations:
[0025] i=HardTan(i)+1
[0026] Among them, HardTan is an activation function with a value range of -1 to 1. When the crowd density of the image is large, the value of the scale factor i will be between 1 and 2, and when the crowd density is small, the value of the scale factor i will be between 0 and 1.
[0027] Furthermore, the steps of generating a graded foreground information graph through the foreground graded network are as follows:
[0028] Convolving the crowd image to be processed with a Gaussian function to convert it into a continuous density function to generate a true value density map;
[0029] Use a fixed-size sliding window to scan the image pixel by pixel, calculate the sum of the pixel values in the local area of each pixel, and obtain the neighborhood density map; where the area with a pixel value of 0 in the density map c is the background, and the area with a non-zero pixel value is the crowd area;
[0030] Find the maximum pixel value I max and the minimum non-zero pixel value I min , and set the average of these two values as the threshold;
[0031] After traversing the neighborhood density map pixel by pixel, non-zero pixels less than a threshold are divided into low-density areas, while pixels greater than or equal to the threshold are divided into high-density areas, and the hierarchical foreground information map is finally generated.
[0032] Furthermore, the step of inputting the crowd density dataset into the density-graded crowd estimation network for training includes:
[0033] The loss function is defined, and the network parameters are optimized through the back-propagation algorithm to train the U-Net network used to generate the hierarchical foreground information;
[0034] The loss function uses Euclidean loss to measure the error between the predicted density map and the true crowd density map, which is defined as shown in the following formula 1:
[0035]
[0036] where X i represents the crowd image of the i-th input network, F(X i ,Θ) is the predicted crowd density map, Θ represents a set of learnable parameters of the network, represents the true density map, L represents the loss between the true density map and the predicted density map, and N represents the number of crowd images used for model training.
[0037] Furthermore, in step S5, after obtaining and generating the predicted density map, the number of pedestrians in the crowd image can be estimated and the estimation effect can be evaluated. The specific method is as follows:
[0038] Each pixel value in the predicted density map represents the crowd density in the area, and the number of people is estimated by calculating the sum of all pixel values in the map;
[0039] To evaluate the estimation effect, the estimated results of the population size are compared with the actual data, and the performance of the model is quantified by calculating the mean absolute error (MAE) and mean square error (MSE), which are defined as shown in Formula 2 and Formula 3:
[0040]
[0041] Where N is the number of images, is the predicted value, is the true value. MAE determines the accuracy of the estimation results, while MSE indicates the robustness of the estimation results.
[0042] Furthermore, in step S5, a real-time density warning is performed based on a preset crowd density threshold. When the crowd density in the generated predicted density map is monitored to exceed a safety threshold, an alarm is automatically triggered or the management personnel is notified.
[0043] Beneficial effects
[0044] Compared with the known public technologies, the technical solution provided by the present invention has the following beneficial effects:
[0045] The present invention proposes a crowd density estimation method and system for foreground segmentation in transportation hub scenarios. By combining deep learning technology, a density-graded crowd estimation network for foreground segmentation is constructed. Based on the intermediate density map, scale factor and graded foreground information map of the input crowd image, the density estimation network is used to generate a predicted density map. Combined with AI video analysis technology, real-time monitoring and analysis of crowd flow, crowd density and crowd orientation in transportation hubs are carried out, and density warnings can be issued in a timely manner, thereby ensuring the safe and stable operation of transportation hubs. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0047] Figure 1 : A flow chart of a method for estimating crowd density in a transportation hub scenario according to an embodiment of the present invention;
[0048] Figure 2 : A density-graded crowd density estimation network based on foreground segmentation according to an embodiment of the present invention;
[0049] Figure 3 : The process of generating a crowd density graded foreground information graph according to an embodiment of the present invention;
[0050] Figure 4 : Visualization result diagram of predicted density and true density according to an embodiment of the present invention; DETAILED DESCRIPTION
[0051] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0052] The present invention will be further described below with reference to the embodiments.
[0053] Example:
[0054] On the one hand, the present invention provides a method for estimating crowd density in a transportation hub scenario, referring to Figure 1-2 , the method steps include:
[0055] Step S1, obtaining a large amount of crowd image data of different densities from surveillance videos of transportation hubs and establishing a crowd density dataset;
[0056] Step S2: Building a density-graded crowd estimation network based on foreground segmentation, and inputting the crowd density dataset into the density-graded crowd estimation network for training;
[0057] The trained density-graded crowd estimation network is divided into two parts: the front-end and the back-end. The front-end adopts a three-branch structure, namely the feature extraction network, the foreground classification network, and the scale factor extraction network, and the back-end is the density estimation network.
[0058] Step S3: Input the crowd image to be processed into a feature extraction network to extract pedestrian head features and form an intermediate density map;
[0059] The crowd image is input into the scale factor extraction network to make a global crowd density level estimate and assign corresponding scale factors according to the density.
[0060] The foreground classification network is used to separate the crowd area and background of the input crowd image, and the crowd area is divided into local density levels to generate a graded foreground information map;
[0061] Step S4, generating a predicted density map using a density estimation network based on the obtained intermediate density map, scale factor, and graded foreground information map of the crowd image;
[0062] Step S5: Generate a predicted density map and compare it with a preset crowd density threshold in real time to issue an early warning.
[0063] Specifically, in step S1, crowd image data is collected from real-time surveillance video at various transportation hubs. This data is then initially processed for denoising, enhancement, and standardization. Data is obtained from real-time surveillance video from transportation hubs such as station waiting rooms and airport concourses. The cameras used are network cameras with a 1080P resolution, and the video is transmitted to the system platform via the RTSP protocol. Using video analysis technology, real-time video image slices are obtained as crowd image data.
[0064] When the density-graded crowd estimation network based on foreground segmentation in this embodiment processes the input crowd image, the crowd image is first input into the three-branch structure of the front end, namely, the feature extraction network, the foreground classification network, and the scale factor extraction network. The specific processing method is as follows:
[0065] In step S31, the feature extraction network consists of the first ten layers of VGG-16, followed by an upsampling layer and a convolutional layer. The upsampling layer is used to amplify the feature map reduced by the pooling operation. Because the position information of objects is abstracted when the feature map is reduced by the pooling operation, resulting in a loss of spatial information, an upsampling layer is added to the end of the network to amplify the feature map from 1 / 8 the size of the original image to 1 / 4 of the original image. A convolution operation is then performed on the convolutional layer to fuse the upsampled local features. The crowd image is first input into the feature extraction network to extract pedestrian head features, forming an intermediate density map.
[0066] In step S32, the first 10 layers of the scale factor extraction network are the same as the first 10 layers of the feature extraction network. The first 10 layers of VGG16 are selected, and an average pooling layer is added after it to facilitate better evaluation of the global head density. Then, four conv3×1 convolutional layers are added to change the output channel to 1. After the convolution operation, an adaptive average pooling operation is performed to quantize the feature map output by the network into a scale factor i. Then, the scale factor value is adjusted to between 0 and 2 by the following operation:
[0067] i=HardTan(i)+1
[0068] Among them, HardTan is an activation function with a value range of -1 to 1. The role of the scale factor i is to make a global crowd density level estimate for the image input to the network; when the crowd density of the image is large, the value of the scale factor i will be between 1 and 2, and when the crowd density is small, the value of the scale factor i will be between 0 and 1.
[0069] Step S33, refer to Figure 3 , select the U-Net network as the foreground classification network branch of the entire model, separate the crowd area and background of the input image to eliminate the influence of complex background on the counting, and perform local density level division on the crowd area to generate a graded foreground information map. The steps are as follows:
[0070] Step S331 : for the crowd image a to be processed, which is of size M×N pixels, convolve it with a Gaussian function to convert it into a continuous density function, and generate a true value density map b;
[0071] In step S332, a 48×48 sliding window is used to scan image b pixel by pixel, and the sum of the pixel values of the local area of each pixel is calculated to obtain a neighborhood density map c; wherein, the area with a pixel value of 0 in the neighborhood density map c is the background, and the area with a non-zero pixel value is the crowd area.
[0072] And perform density classification on the foreground in c. After traversing the neighborhood density map pixel by pixel, find the maximum pixel value I max and the minimum non-zero pixel value Imin , and set the average of these two values as the threshold;
[0073] I threshold =(I max +I min ) / 2
[0074] Step S333, after traversing the neighborhood density map pixel by pixel, the non-zero pixels less than the threshold are divided into low-density areas, and the pixels greater than or equal to the threshold are divided into high-density areas. The final generated graded foreground information map is as follows: Figure 3 As shown in d.
[0075] In step S34, the density estimation network is used to generate the corresponding estimated density map. Dilated convolution is used to construct the density estimation network. Dilated convolution uses sparse kernels instead of pooling layers and convolution layers. In dilated convolution, if the dilation factor is r, the small kernel size of a k×k filter is expanded to (k+(k-1)(r-1))×(k+(k-1)(r-1)), which more flexibly aggregates multi-scale context information while maintaining the same resolution.
[0076] In step S2, the steps of inputting the obtained crowd density dataset into the density-graded crowd estimation network for training include:
[0077] Step S21, defining a loss function and optimizing network parameters through a back-propagation algorithm to train a U-Net network used to generate graded foreground information;
[0078] In step S22, the loss function is Euclidean loss to measure the error between the predicted density map and the true crowd density map, which is defined as shown in the following formula 1:
[0079]
[0080] where X i represents the crowd image of the i-th input network, F(X i ,Θ) is the predicted crowd density map, Θ represents a set of learnable parameters of the network, represents the true density map, L represents the loss between the true density map and the predicted density map, and N represents the number of crowd images used for model training.
[0081] Reference Figure 4 In step S5, after obtaining and generating the predicted density map, the number of pedestrians in the crowd image can be estimated and the estimation effect can be evaluated. The specific method is as follows:
[0082] S51, each pixel value in the predicted density map represents the crowd density of the area, and the number of people is estimated by calculating the sum of all pixel values in the map.
[0083] S52, in order to evaluate the estimation effect, the estimated results of the number of people are compared with the actual data, and the performance of the model is quantified by calculating the mean absolute error (MAE) and mean square error (MSE), as shown in the following formula 2 and formula 3:
[0084]
[0085]
[0086] Where N is the number of images, is the predicted value, is the true value. MAE determines the accuracy of the estimation results, while MSE indicates the robustness of the estimation results.
[0087] S53, a real-time density warning is performed according to a preset crowd density threshold. When the crowd density in the generated predicted density map is monitored to exceed a safety threshold, an alarm is automatically triggered or a management personnel is notified.
[0088] In one aspect, the present invention provides a crowd density estimation system for a transportation hub scenario, comprising:
[0089] The acquisition module is used to obtain the crowd image and crowd density data set to be processed and perform preliminary processing of the image data;
[0090] The model generation and execution module is used to build and train the density-graded crowd estimation network, calculate the input crowd image to be processed, and output the predicted density map;
[0091] The counting and display module is used to estimate and count the number of people in the predicted density map and display the estimated results in real time;
[0092] The early warning module compares the crowd density value in the generated predicted density map with the set density threshold to determine whether it exceeds the safety threshold and initiates an automatic early warning.
[0093] The present invention proposes a crowd density estimation method and system for foreground segmentation in transportation hub scenarios. By combining deep learning technology, a density-graded crowd estimation network for foreground segmentation is constructed. Based on the intermediate density map, scale factor and graded foreground information map of the input crowd image, the density estimation network is used to generate a predicted density map. Combined with AI video analysis technology, the crowd flow, crowd density and crowd orientation of transportation hubs are monitored and analyzed in real time, and density warnings can be issued in a timely manner, thereby ensuring the safe and stable operation of transportation hubs.
[0094] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A crowd density estimation method for a transportation hub scenario, characterized by: The method steps include: Step S1, obtaining a large amount of crowd image data of different densities from surveillance videos of transportation hubs and establishing a crowd density dataset; Step S2, building a density-graded crowd estimation network based on foreground segmentation, and inputting the crowd density dataset into the density-graded crowd estimation network for training; The trained density-graded crowd estimation network is divided into two parts: the front-end and the back-end. The front-end adopts a three-branch structure, namely the feature extraction network, the foreground classification network and the scale factor extraction network, and the back-end is the density estimation network. Step S3, inputting the crowd image to be processed into the feature extraction network to extract pedestrian head features to form an intermediate density map; Input the crowd image into the scale factor extraction network to make a global crowd density level estimate, and assign corresponding scale factors according to different densities; The foreground classification network is used to separate the crowd area and background of the input crowd image, and the crowd area is divided into local density levels to generate a graded foreground information map; Step S4, generating a predicted density map using the density estimation network according to the obtained intermediate density map, scale factor, and graded foreground information map of the crowd image; Step S5: Generate a predicted density map based on the above, and compare it with a preset crowd density threshold in real time and issue an early warning.
2. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: In step S1, crowd image data is obtained through real-time monitoring videos of various transportation hubs, and the collected crowd image data is preliminarily processed by denoising, enhancing and standardizing.
3. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: The feature extraction network consists of the first ten layers of VGG-16 cascaded with an upsampling layer and a convolution layer. The upsampling layer is used to amplify the feature map reduced by the pooling operation, and then a convolution operation is performed through the convolution layer to fuse the upsampled local features.
4. A crowd density estimation method for a transportation hub scenario according to claim 1 or 3, characterized in that: The scale factor extraction network uses the first 10 layers of VGG16 and adds an average pooling layer after it to facilitate better evaluation of global head density. It then adds four conv3×1 convolutional layers to change the output channel to 1. After the convolution operation, an adaptive average pooling operation is performed to quantize the feature map output by the network to the scale factor i. The scale factor value is then adjusted to between 0 and 2 through the following operations: i=HardTan(i)+1 Among them, HardTan is an activation function with a value range of -1 to 1. When the crowd density of the image is large, the value of the scale factor i will be between 1 and 2, and when the crowd density is small, the value of the scale factor i will be between 0 and 1.
5. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: The steps of generating a graded foreground information graph by the foreground classification network are as follows: Convolving the crowd image to be processed with a Gaussian function to convert it into a continuous density function to generate a true value density map; Use a fixed-size sliding window to scan the image pixel by pixel, calculate the sum of the pixel values in the local area of each pixel, and obtain the neighborhood density map; where the area with a pixel value of 0 in the density map c is the background, and the area with a non-zero pixel value is the crowd area; Find the maximum pixel value I max and the minimum non-zero pixel value I min , and set the average of these two values as the threshold; After traversing the neighborhood density map pixel by pixel, non-zero pixels less than a threshold are divided into low-density areas, while pixels greater than or equal to the threshold are divided into high-density areas, and the hierarchical foreground information map is finally generated.
6. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: The step of inputting the crowd density dataset into the density-graded crowd estimation network for training comprises: The loss function is defined, and the network parameters are optimized through the back-propagation algorithm to train the U-Net network used to generate the hierarchical foreground information; The loss function uses Euclidean loss to measure the error between the predicted density map and the true crowd density map, which is defined as follows: where X i represents the crowd image of the i-th input network, F(X i ,Θ) is the predicted crowd density map, Θ represents a set of learnable parameters of the network, represents the true density map, L represents the loss between the true density map and the predicted density map, and N represents the number of crowd images used for model training.
7. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: In step S5, after obtaining and generating the predicted density map, the number of pedestrians in the crowd image can be estimated and the estimation effect can be evaluated. The specific method is as follows: Each pixel value in the predicted density map represents the crowd density in the area, and the number of people is estimated by calculating the sum of all pixel values in the map; To evaluate the estimation effect, the estimated results of the population size are compared with the actual data, and the performance of the model is quantified by calculating the mean absolute error (MAE) and mean square error (MSE), which are defined as follows: Where N is the number of images, is the predicted value, is the true value. MAE determines the accuracy of the estimation results, while MSE indicates the robustness of the estimation results.
8. The method for estimating crowd density in a transportation hub scenario according to claim 1, characterized in that: In step S5, a real-time density warning is performed based on a preset crowd density threshold. When the crowd density in the generated predicted density map is monitored to exceed a safety threshold, an alarm is automatically triggered or the management personnel is notified.
9. A crowd density estimation system for a transportation hub scenario, used to implement the estimation method described in any one of claims 1 to 8, characterized in that: include: The acquisition module is used to obtain the crowd image and crowd density data set to be processed and perform preliminary processing of the image data; A model generation and execution module is used to build and train the density-graded crowd estimation network, and to calculate the input crowd image to be processed and output a predicted density map; A counting and display module is used to estimate and count the number of people in the predicted density map and display the estimated results in real time; The early warning module compares the crowd density value in the generated predicted density map with the set density threshold to determine whether it exceeds the safety threshold and initiates an automatic early warning.