Pedestrian Flow Statistics Method and System Based on Panoramic Images

Through panoramic image stitching and head detection technology, the intelligent problem of traffic statistics in tourist attractions has been solved, real-time traffic monitoring and information provision have been realized, and management efficiency and tourist experience have been improved.

CN114550077BActive Publication Date: 2025-07-25SOUTHEAST DIGITAL ECONOMY DEV INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210022221.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-10
Publication Date
2025-07-25
Estimated Expiration
2042-01-10

AI Technical Summary

Technical Problem

The traffic statistics system of existing tourist attractions has problems such as duplicate statistics, high labor costs, and the inability to intelligently obtain unmonitored area traffic, resulting in low management efficiency and tourists cannot obtain real-time scenic spot information, affecting the gaming experience.

Method used

The end-to-end flow statistics system is realized through camera image acquisition, panoramic stitching, head detection and visualization modules. The SIFT algorithm and the AdaBoost algorithm of Haar-like features are used for feature extraction and detection, and the panoramic image is generated and the number of people is displayed.

Benefits of technology

Remote monitoring of the number of people in the target area and real-time provision of density information, improving regional management efficiency, tourists can obtain intuitive attractions information, and improving the gaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550077B_ABST
    Figure CN114550077B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for counting the number of people based on panoramic images, which includes constructing a panoramic stitching module, a head detection module, and an interface for visualizing the panoramic view and statistical results. The specific method is as follows: 1) Collect camera images; 2) Extract the features of the images and perform matching; 3) Calculate the camera parameters using the matched features and correct them; 4) Use the corrected camera parameters to perform panoramic stitching and fusion on the output images of the cameras that can form a panorama, and the above steps can obtain a panoramic stitcher; 5) Train a head detection classification detector using the AdaBoost algorithm with Haar-like features; 6) Design a video display platform using Qt; 7) Deploy the panoramic and head detection modules, and input the camera parameters to achieve end-to-end counting of the number of people. The present invention can promote the intelligent management of tourist attractions and provide schedule reference information for travelers, and can also be extended to the queuing management of medical examinations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian flow statistics, and specifically, to a pedestrian flow statistics method and system based on panoramic images. Background Art

[0002] Smart tourism is a realistic manifestation of the booming development of the tourism industry. Among them, pedestrian flow statistics systems have been deployed in various scenic spots. These systems are used to address the risk of increased trampling accidents due to excessive pedestrian flow in tourist attractions and are also a requirement for epidemic prevention and standardized management.

[0003] Currently, the common passenger flow statistics systems in tourist attractions achieve management purposes by controlling the overall pedestrian flow. The common statistical method is to set up ticket inspection systems at the entrances and exits. The system counts the total volume and displays the data on an electronic display screen. Although this traditional electronic screen has been upgraded to a digital twin platform class, this visual large screen still displays the global data volume. Therefore, the congestion situation of internal scenic spots in tourist attractions depends on the self-awareness of tourists and the manual control of management personnel. Specifically, scenic spot management personnel count the pedestrian flow at each location through the feedback of fixed cameras at various points, and then make comprehensive judgments to obtain management decisions. However, the overlapping areas of these cameras are prone to cause duplicate statistics, increasing the manual statistical cost. It is difficult to intelligently obtain the pedestrian flow in unmonitored areas, so manual management is still required, which poses a challenge to the timeliness of solving accident risk problems. For tourists in scenic spots, travel itineraries are usually planned based on the travel experiences of others. It is difficult to adjust travel routes without obtaining information on the crowd flow at scenic spots. Therefore, the pleasure of travel obtained in this way is uncertain, and travel depends on scenic spot publicity and the travel experiences and feelings of others. Summary of the Invention

[0004] To solve the problems of the prior art, embodiments of the present invention provide a pedestrian flow statistics method and system based on panoramic images. The technical solutions are as follows:

[0005] On the one hand, a pedestrian flow statistics method based on panoramic images is provided, including:

[0006] Collect camera images and perform panoramic stitching and fusion on the camera images;

[0007] Detect the number of people's heads in the panoramic image;

[0008] Obtain the real-time pedestrian flow based on the detected heads.

[0009] Further, the step of collecting camera images and performing panoramic stitching and fusion on the camera images is specifically:

[0010] Step 1) Collect camera images, including deploying cameras on a unified horizontal line, reading camera streams, and setting image resolutions;

[0011] Step 2) Use the SIFT algorithm to extract features from each processed image, and then use the nearest neighbor method to calculate and complete the calculation of the best matching points;

[0012] Step 3) Calculate and correct the camera parameters using the matched features;

[0013] Step 4) Use the corrected camera parameters to perform panoramic stitching and fusion on the camera output pictures that can form a panorama.

[0014] Further, the specific content of step 3) is as follows:

[0015] Calculate the best estimated parameters of the camera according to the affine transformation and the best matching points, use the bundle adjustment method to correct the estimated camera parameters to reduce the feature distortion state in the panorama caused by the parameters. Then, perform the correction and format conversion of the camera parameters. According to the best adjacent matching standard, obtain the order of the images stitched into a panorama. According to this order, use the previously calculated camera parameters and the input images to extract the image corner features through affine transformation. Use the input image mask and the previous camera parameters to obtain the mask correction parameters through affine transformation. Perform exposure compensation on the input image transformation parameters, mask, and the calculated image corner points. Then, use the graph cut method to calculate the gaps between the matching features and save them in the corrected mask transformation parameters.

[0016] Further, the specific content of step 4) is as follows:

[0017] Calculate the current image transformation according to the previously calculated camera parameters, picture corner points, and the pictures taken from the real-time output stream. Calculate the mask transformation of the current image according to the camera parameters. Perform exposure compensation according to the corner points, real-time input images, and the newly obtained mask transformation. Then, perform dilation and size transformation on the mask transformation calculated in step 3) and merge the results with the new mask transformation in the current step to obtain the mask with optimized gaps. Finally, use the multi-frequency fusion algorithm to fuse the current real-time image transformation, the final optimized mask, and the image corner points to obtain the stitched panoramic image.

[0018] Further, the specific steps of using the multi-frequency fusion algorithm to fuse the current real-time image transformation, the final optimized mask, and the image corner points to obtain the stitched panoramic image are as follows:

[0019] 1) Calculate the Gaussian pyramid of the input image;

[0020] 2) Calculate the Laplacian pyramid of the input image;

[0021] 3) Fuse the Laplacian pyramids at the same level. For example, use simple linear fusion on both sides of the stitching seam;

[0022] 4) Expand the high-level Laplacian pyramids sequentially until they reach the same resolution as the source image;

[0023] 5) Stack the images of the cameras that make up the panorama in sequence, and the final output image is obtained.

[0024] Furthermore, the step of detecting the number of people's heads in the panoramic image is specifically as follows;

[0025] This step uses an AdaBoost cascade classifier with Haar-like features to implement a human head detector;

[0026] Specifically:

[0027] 1) Use the integral image to describe the global information of the image. The integral image at a specified coordinate of the image is the sum of all pixels in the upper left corner of this position. Calculate the pixel values of the image area in this way, and then obtain the global feature values of the image using the integral image;

[0028] 2) Construct a classification detector based on the obtained features, and obtain the optimal model by improving the recognition rate and reducing the false recognition rate. Here, the positioning loss uses the squared difference loss between the feature point position and the actual position.

[0029] On the other hand, a pedestrian flow statistics system based on panoramic images is provided, including:

[0030] A panoramic stitching module, which is used to collect camera images and perform panoramic stitching and fusion on the camera images;

[0031] A human head detection module, which is used to detect the number of people's heads in the panoramic image

[0032] A visualization panorama and statistics module, which is used to obtain the real-time pedestrian flow based on the detected people's heads.

[0033] Furthermore, the panoramic stitching module includes:

[0034] An image acquisition unit, which is used to collect multiple camera images of the target area. The cameras are deployed on the same plane and on the same horizontal line, and the camera streams are read through the vlc library and converted into images;

[0035] A feature extraction unit, which is used to extract the features of each processed image using the SIFT algorithm, and then use the nearest neighbor method to calculate and complete the calculation of the best matching points;

[0036] A parameter correction unit, which is used to calculate and correct the camera parameters using the matched features;

[0037] A panoramic stitching unit is used to perform panoramic stitching and fusion on the output images of cameras that can form a panorama by using the calibrated camera parameters.

[0038] Further, the human head detection module is implemented by using the AdaBoost cascade method, extracting features to train a classification model, and directly using the feature positions and squared error regression to obtain the detection frame.

[0039] The beneficial effects brought by the technical solution provided by the embodiments of the present invention are:

[0040] A method for counting the number of people based on panoramic images provided by the present invention uses this method to obtain an end-to-end people counting system. This system can remotely monitor the number of people in the target area and remotely provide density information through the camera voice information according to the statistical results and panoramic intuitive information, improving the efficiency of area management. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is the panoramic stitching flowchart of Embodiment 1 of the present invention;

[0043] Figure 2 is the integral image used to describe the global information of the image in Embodiment 1 of the present invention;

[0044] Figure 3 is the human head detection flowchart in Embodiment 1 of the present invention;

[0045] Figure 4 is the graphical interface for image display and number display in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.

[0047] The present invention provides a method for counting the number of people based on panoramic images. Refer to Figure 1 , including:

[0048] S1: Collect camera images, including deploying cameras on the same horizontal line, reading camera streams, and setting image resolutions.

[0049] In this embodiment, the camera brand in S1 is not limited to one type, and local videos can also be used. Based on the hardware level of the development environment, different resolutions can be set for the images.

[0050] Specifically, in S1, the vlc third-party library is used to read the real-time stream of the camera and convert it into image frames. The size of the obtained images is changed to 640*384, and other resolutions are also selected based on the hardware level of different development environments and the number of cameras.

[0051] S2: Use the SIFT algorithm to extract the features of each processed image, and then use the nearest neighbor method to complete the calculation of the best matching points.

[0052] Specifically, the nearest neighbor matching used in S2 is based on the matching of the original feature points.

[0053] In this embodiment, in S2, the SIFT algorithm, i.e., scale-invariant feature transform, is used for feature extraction, and the nearest neighbor algorithm is used to select the best matching features.

[0054] S3: Calculate and correct the camera parameters using the matched features;

[0055] Specifically, the best estimated parameters of the camera are calculated according to the affine transformation and the best matching points. The bundle adjustment method is used to correct the estimated camera parameters to reduce the feature distortion state in the panorama caused by the parameters. Then, the correction and format conversion of the camera parameters are carried out. In sequence, it includes using the previously calculated camera parameters and the input image to extract the corner features of the image through affine transformation, using the input image mask and the previous camera parameters to obtain the mask correction parameters through affine transformation, performing exposure compensation on the transformation parameters of the input image, the mask, and the calculated image corner points, and then using the graph cut method to calculate the gaps between the matching features and save them in the corrected mask transformation parameters.

[0056] In this embodiment, an affine transformation is used in S3 to obtain camera parameters. Specifically: 1) An affine transformation is used to calculate the internal parameter matrix, focal length matrix, and rotation transformation matrix of the camera from the best-matched features obtained in S2; 2) The bundle adjustment method is used to correct the camera parameters, that is, the above-mentioned internal parameter matrix, focal length matrix, and rotation matrix. The basic principle of bundle adjustment is to take multiple images of the object to be measured from different angles and positions. Then, all the feature points on the object, the object points, image points, and optical centers should fall on the corresponding light rays, that is, satisfy the collinearity relationship of imaging. The process of adjustment is to optimize to minimize the overall deviation of these collinearity relationships; 3) The original image is transformed using the calculated camera parameters and affine transformation, and the upper-left corner point features are extracted. Then, the mask of the original image is transformed in the same way to generate a transformed image, so that the feature direction of the transformed image is consistent with the feature direction of the matched image; 4) Exposure compensation is performed on the transformed image of the original image, its mask-transformed image, and corner points to reduce the difference in features of the stitched and synthesized image; 5) The graph cut method is used to smooth the gap between the mask-transformed images, that is, to smooth the boundary at the feature stitching when stitching the panorama, and eliminate the feature differences caused by inaccurate feature matching.

[0057] S4: Use the corrected camera parameters to perform panoramic stitching and fusion on the output pictures of the cameras that can form a panorama.

[0058] Specifically, calculate the current image transformation based on the previously calculated camera parameters, picture corner points, and the pictures taken from the real-time output stream. Calculate the mask transformation of the current image according to the camera parameters. Perform exposure compensation based on the corner points, real-time input image, and the newly obtained mask transformation. Then, perform dilation and size transformation on the mask transformation calculated in S3 and merge the result with the new mask transformation in the current step to obtain the mask with optimized gaps. Finally, use the multi-frequency fusion algorithm to fuse the current real-time image transformation, the final optimized mask, and the image corner points to obtain the stitched panoramic image.

[0059] Furthermore, the multi-band fusion steps in S4 are as follows:

[0060] 1) Calculate the Gaussian pyramid of the input image;

[0061] 2) Calculate the Laplacian pyramid of the input image;

[0062] 3) Fuse the Laplacian pyramids at the same level. For example, use simple linear fusion on both sides of the stitching seam;

[0063] 4) Expand the high-level Laplacian pyramids in sequence until they reach the same resolution as the source image;

[0064] 5) Stack the images of the cameras that form a panorama in sequence, and the final output image is obtained.

[0065] In this embodiment, S4 mainly covers the real-time panoramic stitching process. Here, first, new image transformation and mask transformation are calculated using the previous camera parameters and the real-time stream of the camera. Then, a new panoramic image is generated by integrating the previously calculated mask and the multi-band fusion method.

[0066] S5. This step is implemented using the AdaBoost cascade method. Feature extraction is used to train a classification model, and at the same time, the detection box is directly obtained using the feature position and the squared error regression.

[0067] Specifically, an AdaBoost cascade classifier using Haar-like features is used to implement a human head detector. Specifically:

[0068] 1) Refer to Figure 2 , the integral image is used to describe the global information of the image. The integral image of a specified coordinate of the image is the sum of all pixels in the upper left corner of this position. By analogy, the pixel values of the image area are calculated, and then the global feature values of the image are obtained using the integral image; refer to Figure 3 , the integral image is used to describe the global information of the image. The formula used is ii(x,y) = ∑ x'≤x,y'≤y i(x',y'), where i(x',y') is the pixel value at the original image (x',y'), and ii(x,y) is the integral image at (x,y). Then the pixel value of the area D in the figure = the pixel value of the area A + B + C + D + the pixel value of the area A - the pixel value of the area A + B - the pixel value of the area A + C, that is, the pixel value of the area D = ii(4) + ii(1) - ii(2) - ii(3). ii(1) represents the pixel value of the area A, ii(2) represents the pixel value of the area A + B, ii(3) represents the pixel value of the area A + C, and ii(4) represents the pixel value of the area A + B + C + D. By analogy, the pixel values of adjacent areas are calculated, and the difference between the pixel values of adjacent areas can be used to obtain the feature values by adding and subtracting the integral images of the rectangle endpoints;

[0069] 2) A classification detector is constructed based on the obtained features. The optimal model is obtained by improving the recognition rate and reducing the misrecognition rate. Here, the positioning loss uses the squared difference loss between the feature point position and the actual position.

[0070] In this embodiment, the Haar feature extraction method in S5 is selected because the feature calculation based on the Haar feature template is large. The integral image uses modularization to accelerate the process of extracting Haar features;

[0071] Furthermore, refer to Figure 3, including the whole process of head detection. The specific logic modules are as follows: 1) Collect images as data samples, covering various situations that may occur in actual applications as much as possible. Half of the samples are images containing head features at multiple angles as training positive samples, and the other half are images without head features as negative samples. The image size is normalized to a size that is easy to train. Then, mark all samples, including categories and target boxes; 2) Extract Haar features; 3) Generate weak classifiers. By using the position information of Haar features and performing statistics on the training samples, the corresponding feature parameters can be obtained; 4) Use the AdaBoost algorithm to select optimized weak classifiers. Here, basic binary classifiers are obtained; 5) Use the weak classifier set as the input. Under the constraints of the training detection rate and false positive rate, use the AdaBoost algorithm to select the optimal weak classifier; 6) Linearly combine the weak classifiers to obtain a strong classifier as the final detector.

[0072] S6: Implement a video display platform, which is implemented using qt and can achieve panoramic display and display of the total number of people in the panorama.

[0073] S7: Deploy the panoramic algorithm and the head detector to the platform implemented in S6.

[0074] Embodiment 2

[0075] The present invention provides a pedestrian flow statistics system based on panoramic images, including:

[0076] A panoramic stitching module for collecting camera images and performing panoramic stitching and fusion on the camera images.

[0077] A head detection module for detecting the number of heads in the panoramic image

[0078] A visual panoramic and statistics module for obtaining the real-time pedestrian flow based on the detected heads.

[0079] Further, the panoramic stitching module includes:

[0080] An image acquisition unit for collecting multiple camera images of the target area, where the cameras are deployed on the same plane and on the same horizontal line, reading the camera stream through the vlc library and converting it into an image;

[0081] A feature extraction unit for using the SIFT algorithm to extract the features of each processed image, and then using the nearest neighbor method to calculate the best matching points;

[0082] A parameter correction unit for calculating and correcting the camera parameters using the matched features;

[0083] A panoramic stitching unit for performing panoramic stitching and fusion on the output pictures of the cameras that can form a panorama using the corrected camera parameters.

[0084] Further, the human head detection module is implemented using the AdaBoost cascade method, extracts features to train a classification model, and directly uses the feature positions and squared error regression to obtain the detection boxes.

[0085] Further, the visualization panorama and statistics module uses video display to display the panorama and the total number of people in the panorama.

[0086] Specifically, it includes implementing a graphical interface for image display and number of people display, see Figure 4 , and the specific modules include parameter input, panorama display, and real-time data display, where the parameter input module can input camera parameters and local video addresses.

[0087] Deploy and merge the trained panorama fuser and human head detector on the interface, and finally obtain a panoramic human head detection system, which can be extended for use in crowd queuing and traffic control in various industries.

[0088] The beneficial effects brought by the technical solution provided in the embodiment of the present invention are:

[0089] A method for counting the number of people based on panoramic images provided by the present invention obtains an end-to-end people counting system using this method. This system can remotely monitor the number of people in the target area, and remotely provide density information through camera voice information according to the statistical results and panoramic intuitive information, improving the efficiency of area management.

[0090] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for counting the number of people based on panoramic images, characterized in that, Including: Collecting camera images and performing panoramic stitching and fusion on the camera images; Detecting the number of human heads in the panoramic image; Obtaining the real-time pedestrian flow based on the detected human heads; The steps of collecting camera images and performing panoramic stitching and fusion on the camera images are specifically as follows: Step 1) Collecting camera images, including deploying cameras on the same horizontal line, reading camera streams, and setting image resolutions; Step 2) Using the SIFT algorithm to extract the features of each processed image, and then using the nearest neighbor method to calculate and complete the calculation of the best matching points; Step 3) Calculating and correcting the camera parameters using the matched features; Step 4) Using the corrected camera parameters to perform panoramic stitching and fusion on the camera output pictures that can form a panorama; The specific content of step 3) is as follows: Calculating the best estimated camera parameters according to the affine transformation and the best matching points, using the bundle adjustment method to correct the estimated camera parameters to reduce the feature distortion state in the panorama caused by the parameters. Then, it is the correction and format conversion of the camera parameters. According to the image arrangement order determined based on the best adjacent matching standard for stitching into a panoramic image, using the previously calculated camera parameters and the input images to extract the image corner features through affine transformation, using the input image mask and the previous camera parameters to obtain the mask correction parameters through affine transformation, performing exposure compensation on the input image transformation parameters, the mask, and the calculated image corner points. Then, using the graph cut method to calculate the gaps between the matching features and save them in the corrected mask transformation parameters.

2. The method according to claim 1, wherein The specific content of step 4) is as follows: Calculating the current image transformation according to the previously calculated camera parameters, the picture corner points, and the pictures taken from the real-time output stream, calculating the mask transformation of the current image according to the camera parameters, performing exposure compensation according to the corner points, the real-time input image, and the newly obtained mask transformation. Then, performing dilation and size transformation on the mask transformation calculated in step 3) and merging the result with the new mask transformation in the current step to obtain the mask with optimized gaps. Finally, using the multi-frequency fusion algorithm to fuse the current real-time image transformation, the final optimized mask, and the image corner points to obtain the stitched panoramic image.

3. The method according to claim 2, wherein The steps of using the multi-frequency fusion algorithm to fuse the current real-time image transformation, the final optimized mask, and the image corner points to obtain the stitched panoramic image are specifically as follows: 1) Calculating the Gaussian pyramid of the input image; 2) Calculating the Laplacian pyramid of the input image; 3) Fusing the Laplacian pyramids at the same level; using linear fusion on both sides of the stitching seam; 4) Sequentially expanding the high-level Laplacian pyramids until they reach the same resolution as the source image; 5) Sequentially stacking the images of the cameras that form a panorama to obtain the final output image.

4. The method according to claim 3, wherein The steps of detecting the number of human heads in the panoramic image are specifically as follows; This step uses an AdaBoost cascade classifier with Haar-like features to implement a human head detector; Specifically: 1) Using the integral image to describe the global information of the image. The integral image of a specified coordinate in the image is the sum of all pixels in the upper left corner of this position. By analogy, calculate the pixel values of the image area, and then obtain the global feature values of the image using the integral image; 2) Construct a classification detector based on the obtained features, and obtain the optimal model by improving the recognition rate and reducing the false recognition rate. Here, the localization loss uses the squared difference loss between the feature point position and the actual position.

5. A pedestrian flow statistics system based on panoramic images, characterized in that, Including: A panoramic stitching module, a human head detection module, and a visualization panorama and statistics module for implementing the method according to any one of claims 1-4; wherein: The panoramic stitching module is used to collect camera images and perform panoramic stitching and fusion on the camera images. The human head detection module is used to detect the number of human heads in the panoramic image. The visualization panorama and statistics module is used to obtain the real-time pedestrian flow according to the detected human heads.

6. The system according to claim 5, characterized in that, Including: An image acquisition unit for collecting multiple camera images of the target area, where the cameras are deployed on the same plane and on the same horizontal line, and the camera streams are read through the vlc library and converted into images. A feature extraction unit for using the SIFT algorithm to extract the features of each processed image, and then using the nearest neighbor method to calculate and complete the calculation of the best matching points. A parameter correction unit for calculating and correcting the camera parameters using the matched features. A panoramic stitching unit for performing panoramic stitching and fusion on the camera output pictures that can form a panorama using the corrected camera parameters.

7. The system according to claim 6, characterized in that, The human head detection module is implemented using the AdaBoost cascade method, extracts features to train a classification model, and directly uses the feature position and squared error regression to obtain the detection frame.

Citation Information

Patent Citations

  • Visitor flow rate statistics method and device

    CN106650581A

  • Crack image splicing method

    CN111311492A