A bus passenger flow detection method based on RGB images

By using Kalman filter and YOLOX model to perform RGB image processing in bus passenger flow detection, the problems of high cost and low detection accuracy in the existing technology are solved, and high accuracy of bus passenger flow data acquisition is achieved.

CN115063450BActive Publication Date: 2025-06-10张维忠
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210648778.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-06-10
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

The existing bus passenger flow detection technology has problems such as high cost, easy equipment to be damaged, and low detection accuracy in crowded situations, and it is difficult to obtain high-accurate passenger flow data using color images.

Method used

The bus passenger flow detection method based on RGB images is used to predict the target position through the Kalman filter, and the prediction results are superimposed on the original image and input into the YOLOX model for detection. The Kuhn-Munkres algorithm is used to match to determine the bus passenger flow.

Benefits of technology

It realizes the use of color images to obtain high-accurate bus passenger flow data, which is low-cost and does not require expensive depth cameras. It can work in existing surveillance videos of buses, improves detection accuracy and ensures the accuracy of passenger flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063450B_ABST
    Figure CN115063450B_ABST
Patent Text Reader

Abstract

The present invention discloses a bus passenger flow detection method and system based on RGB images. The method includes: obtaining a video stream of passengers getting on and off the bus; the video stream includes multiple frames of images; for the Kth frame of image, predicting the target position in the Kth frame of image through the parameters of the Kalman filter that predicts the (K-1)th frame of image to obtain the prediction result of the Kalman filter; for the Kth frame of image, generating an image to be predicted based on the prediction result of the Kalman filter and the Kth frame of image; inputting the image to be predicted into the trained YOLOX model to obtain the detection result of the YOLOX model; using the Kuhn-Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model; determining the bus passenger flow according to the matching result. The present invention performs detection based on color images, reducing the cost. The present invention also inputs the result predicted by the Kalman filter and the original image into the YOLOX model for detection, improving the detection accuracy of the YOLOX model and ensuring the accuracy of the bus passenger flow volume.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bus passenger flow detection, and particularly to a bus passenger flow detection method and system based on RGB images. Background Art

[0002] Accurate passenger flow data is of great significance to intelligent transportation. Through accurate quantitative data, the number of passengers in each time period, on each bus route, and at each station can be understood. These data can help bus dispatching personnel conduct more timely and accurate dispatching, improve the user's riding experience, help bus companies conduct more scientific and reasonable network planning, make better use of public resources, and provide basic data for the planning and evaluation of the urban comprehensive transportation system.

[0003] In recent years, many scholars have conducted in-depth research on how to obtain accurate passenger flow data, but there are still certain limitations in practical applications. For example: some scholars use photoelectric switches and pressure sensors to obtain the number of passengers getting on and off. These methods are early works. The photoelectric switch cannot judge the situation of two people getting on the bus side by side, and the pressure sensor is easy to be damaged and has a high maintenance cost. Both methods cannot perform well in crowded situations and have basically been phased out; some scholars locate the approximate positions of passengers through information such as WIFI, Bluetooth, or MAC addresses of the electronic devices carried by passengers to obtain the number of passengers getting on and off. There are still many challenges in these technologies at present, including: some passengers carry zero or multiple electronic devices, are vulnerable to interference from electronic devices outside the bus, and many devices provide false dynamic MAC addresses, etc.; there are also some scholars who use video images, and some use depth images, but structured light cameras and time-of-flight cameras are costly, and the images of binocular cameras are vulnerable to textures, and the stability of accuracy cannot be guaranteed; there are also some scholars who use color images and use traditional digital image processing methods such as Hough transform circle detection to identify human heads, etc. to detect, but due to the complexity of the scene, this method cannot obtain a high accuracy rate.

[0004] With the rapid development of deep learning and the continuous improvement of the accuracy of object detection models, many networks deployed on mobile devices have emerged, making it possible to obtain passenger flow data with high accuracy using color images. However, how to use color images to obtain accurate bus passenger flow is still a major problem. Summary of the Invention

[0005] The purpose of the present invention is to provide a bus passenger flow detection method and system based on RGB images to obtain accurate bus passenger flow.

[0006] To achieve the above purpose, the present invention provides the following solutions:

[0007] A bus passenger flow detection method based on RGB images, the method includes:

[0008] Obtain the video stream of passengers getting on and off the bus; the video stream includes multiple frames of images;

[0009] For the Kth frame of image, predict the target position in the Kth frame of image through the parameters of the Kalman filter that predicts the (K - 1)th frame of image, and obtain the prediction result of the Kalman filter;

[0010] For the Kth frame of image, generate a to-be-predicted image based on the prediction result of the Kalman filter and the Kth frame of image;

[0011] Input the to-be-predicted image into the trained YOLOX model to obtain the detection result of the YOLOX model;

[0012] Use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model;

[0013] Determine the bus passenger flow according to the matching result.

[0014] Furthermore, for the first frame of image, directly detect the target position through the trained YOLOX model.

[0015] Furthermore, the YOLOX model includes a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer; the convolutional layer is a normal convolutional layer with a convolutional kernel size of 2 and a stride of 2.

[0016] Furthermore, for the Kth frame of image, generating a to-be-predicted image based on the prediction result of the Kalman filter and the Kth frame of image specifically includes:

[0017] Generate a target position mask from the prediction result of the Kalman filter;

[0018] Superimpose the target position mask and the Kth frame of image with different weights to generate a to-be-predicted image.

[0019] Furthermore, using the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model specifically includes:

[0020] Calculate the intersection over union (IoU) of the prediction result of the Kalman filter and the detection result of the YOLOX model;

[0021] Construct a cost matrix through the IoU;

[0022] Based on the cost matrix, use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model, and the matching result is the target trajectory.

[0023] Furthermore, determining the bus passenger flow according to the matching result specifically includes:

[0024] Set a counting area; the counting area is the area where passengers get on and off the bus.

[0025] Determine the passenger flow according to the positional relationship between the target trajectory and the counting area.

[0026] The present invention also provides a bus passenger flow detection system based on RGB images, and the system includes:

[0027] A video stream acquisition module, configured to acquire a video stream of passengers getting on and off the bus; the video stream includes multiple frames of images.

[0028] A prediction module, configured to, for the Kth frame of image, predict the target position in the Kth frame of image by predicting the parameters of the Kalman filter of the (K - 1)th frame of image, and obtain the prediction result of the Kalman filter.

[0029] An image generation module, configured to, for the Kth frame of image, generate a to-be-predicted image based on the prediction result of the Kalman filter and the Kth frame of image.

[0030] A detection module, configured to input the to-be-predicted image into a trained YOLOX model to obtain the detection result of the YOLOX model.

[0031] A matching module, configured to match the prediction result of the Kalman filter and the detection result of the YOLOX model by using the Kuhn - Munkres algorithm.

[0032] A determination module, configured to determine the bus passenger flow according to the matching result.

[0033] Further, the YOLOX model includes a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer; the convolutional layer is a common convolutional layer with a convolutional kernel size of 2 and a stride of 2.

[0034] Further, the image generation module includes:

[0035] A mask generation unit, configured to generate a target position mask from the prediction result of the Kalman filter.

[0036] An overlay unit, configured to overlay the target position mask and the Kth frame of image with different weights to generate a to-be-predicted image.

[0037] Further, the matching module includes:

[0038] A calculation unit, configured to calculate the intersection over union of the prediction result of the Kalman filter and the detection result of the YOLOX model.

[0039] A construction unit, configured to construct a cost matrix through the intersection over union.

[0040] A matching unit, which is used to match the prediction result of the Kalman filter and the detection result of the YOLOX model based on a cost matrix by using the Kuhn-Munkres algorithm, and the matching result is the target trajectory.

[0041] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0042] The bus passenger flow detection method based on RGB images proposed by the present invention obtains the video stream of passengers getting on and off the bus. For the Kth frame image, the target position in the Kth frame image is predicted by predicting the parameters of the Kalman filter of the (K - 1)th frame image, and the prediction result of the Kalman filter is obtained; for the Kth frame image, a to-be-predicted image is generated based on the prediction result of the Kalman filter and the Kth frame image; the to-be-predicted image is input into the trained YOLOX model to obtain the detection result of the YOLOX model; the Kuhn-Munkres algorithm is used to match the prediction result of the Kalman filter and the detection result of the YOLOX model; the bus passenger flow is determined according to the matching result. The present invention uses color images for passenger flow detection, with low cost, does not require expensive depth cameras, and can even use the existing monitoring videos of buses; the present invention inputs the result predicted by the Kalman filter and the original image into the YOLOX model for detection, improving the detection accuracy of the YOLOX model and ensuring the accuracy of bus passenger flow. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0044] Figure 1 It is a flowchart of the bus passenger flow detection method based on RGB images provided by the embodiments of the present invention;

[0045] Figure 2 It is a schematic diagram of generating a to-be-predicted image based on the prediction result of the Kalman filter and the Kth frame image provided by the embodiments of the present invention;

[0046] Figure 3 It is a schematic diagram of the original Focus structure provided by the embodiments of the present invention;

[0047] Figure 4 It is a schematic diagram of the YOLOX model provided by the embodiments of the present invention;

[0048] Figure 5 It is a comparison diagram of images before and after data augmentation provided by the embodiments of the present invention;

[0049] Figure 6 This is a schematic diagram showing the changes in precision and recall during the training process of the YOLOX model predicted by the fusion Kalman filter provided by the embodiments of the present invention. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] The purpose of the present invention is to provide a bus passenger flow detection method and system based on RGB images to obtain accurate bus passenger flow.

[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0053] As Figure 1 shown, a bus passenger flow detection method based on RGB images includes the following steps:

[0054] Step 101: Obtain the video stream of passengers getting on and off the bus; the video stream includes multiple frames of images.

[0055] Step 102: For the Kth frame of image, predict the target position in the Kth frame of image by predicting the parameters of the Kalman filter of the (K - 1)th frame of image to obtain the prediction result of the Kalman filter.

[0056] Among them, before obtaining the state estimation of the target, that is, for the first frame of image, directly detect the target position through the trained YOLOX model.

[0057] Among them, in a specific embodiment, step 2 includes: for the Kth frame of image in the video stream, the prior state estimation of the Kth frame of image can be calculated by using the posterior state estimation of the (K - 1)th frame of image, and the prior estimation covariance matrix of the Kth frame of image can be calculated by using the posterior estimation covariance matrix of the (K - 1)th frame of image. These two prior values are the prediction results of the Kalman filter, and the results include prediction box information. Let the time when the Kth frame of image is obtained be t, and the time when the (K - 1)th frame of image is obtained be t - 1, then the specific calculation formula is as follows:

[0058]

[0059] Among them is the prior state estimation of the Kth frame of image, is the prior covariance matrix estimate of the K-th frame image; is the posterior state estimate of the (K-1)-th frame image, P t-1 is the posterior covariance matrix estimate of the (K-1)-th frame image. It includes the center coordinates, area, aspect ratio of the prediction box, and the change rates of the center coordinates and area between frames. The specific values are obtained by updating the (K-1)-th frame image through the Kalman filter. F is the state transition matrix at the time of the K-th frame image, and its value is:

[0060]

[0061] B is the control matrix, and u is the control quantity, both of which take the value of 0; Q is the process noise matrix, which comes from the loss, overlap, and inaccuracy of the detection box. Since the initial velocity is unknown, a relatively large initial value is set for the covariance matrix, and more trust is placed on the measurement values in the early stage. The process noise comes from the uncertainty of motion, and its value is:

[0062]

[0063] Step 103: For the K-th frame image, generate the image to be predicted based on the prediction result of the Kalman filter and the K-th frame image. Specifically, it includes: generating a target position mask from the prediction result of the Kalman filter; superimposing the target position mask and the K-th frame image with different weights to generate the image to be predicted.

[0064] As Figure 2 shown, Figure 2 (a) is the target position mask, Figure 2 (b) is the K-th frame image, Figure 2 (c) is the generated image to be predicted. Based on the center coordinates, area, and aspect ratio of the prediction result of this Kalman filter obtain the coordinates and width and height of the target position mask Figure 2 (a), and superimpose this target position mask Figure 2 (a) and the K-th frame image Figure 2 (b) with different weights to generate the image to be predicted Figure 2 (c).

[0065] Step 104: Input the image to be predicted into the trained YOLOX model to obtain the detection result of the YOLOX model.

[0066] As Figure 3As shown, before the image enters the YOLOX backbone network, it will pass through the original Focus structure. This Focus structure performs interval pixel sampling on the rows and columns of the image to form a new image. Without losing information, it expands the number of channels by 4 times, which is equivalent to the downsampling function. Since the unconventional operation structure of Focus is not conducive to the deployment of the network on mobile devices, the YOLOX model in the present invention replaces the original Focus structure with a common convolutional layer with a convolutional kernel size of 2 and a stride of 2. The modified YOLOX model is as shown in Figure 4 shown, including a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer, where the convolutional layer is a common convolutional layer with a convolutional kernel size of 2 and a stride of 2.

[0067] Input the to-be-predicted image generated in step 3 into the trained YOLOX model to obtain the detection result of the YOLOX model.

[0068] Step 105: Use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model. Specifically, it includes:

[0069] Calculate the intersection over union (IoU) between the prediction result of the Kalman filter and the detection result of the YOLOX model;

[0070] Construct a cost matrix through the IoU;

[0071] Based on the cost matrix, use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model, and the matching result is the target trajectory.

[0072] In a specific embodiment, calculate the IoU between the prediction result (prediction box) of each Kalman filter and the detection result (measurement box) of each YOLOX model, and use 1 - IoU to construct the cost matrix. Based on this cost matrix, through the Kuhn - Munkres algorithm, the measurement boxes are assigned to the prediction boxes under the condition of the minimum total cost, and the matching result is the target trajectory. Its theoretical basis is: adding or subtracting a certain number to a certain row or a certain column of the cost matrix does not change the optimal assignment problem.

[0073] The specific process of the Kuhn - Munkres algorithm is as follows:

[0074] (1) Subtract the minimum value of each row from that row.

[0075] (2) Subtract the minimum value of each column from that column.

[0076] (3) Use the minimum number of horizontal and vertical lines to cover all zeros in the cost matrix. If n rows are needed, the optimal assignment is found and the algorithm ends; otherwise, execute step (4).

[0077] (4) Find the minimum element among the elements not covered by any line. Subtract this element from all the uncovered rows and add this element to all the covered columns. Repeat step (3).

[0078] In the actual operation process, since the number of prediction boxes and measurement boxes is not always equal, when the number of measurement boxes is greater than the number of prediction boxes, if the detection boxes that are not matched are detected continuously for 2 frames, a new trajectory is created; when the number of measurement boxes is less than the number of prediction boxes, if the missing measurement boxes do not appear again continuously for 2 frames, the trajectory is deleted.

[0079] Step 106: Determine the bus passenger flow according to the matching result. Specifically, it includes: setting a counting area; the counting area is the area where passengers get on and off the bus; determine the passenger flow according to the positional relationship between the target trajectory and the counting area.

[0080] In the actual operation process, first set the counting area, which can be set as a rectangular box area. Judge according to the specific scene of the acquired image. For example, when the target trajectory enters from below and exits from above the rectangular box area, it is considered getting off the bus, and vice versa for getting on the bus; the manifestation of the passenger flow is how many people get on and off at each station, and thus the bus passenger flow is determined.

[0081] As Figure 5 shown, in a specific embodiment, use the prediction result of the Kalman filter as a mask, and superimpose the mask on the original image as the training dataset of the YOLOX model. Since the installation height of the camera inside the bus is limited, there will be no situation where the target is too large or too small. Therefore, modify the random scaling ratio of the original image in the Mosaic data augmentation. At the same time, to improve the accuracy in the case of target occlusion, randomly crop and then splice the original image during the Mosaic data augmentation process. Finally, adjust the size of the training image to 416*416. Simulate the situations of strong light and insufficient light by random augmentation and reducing brightness, solve the lack of the dataset in specific scenarios, and improve the robustness of the network. The comparison of the images before and after data augmentation is as Figure 5 shown, Figure 5 (a) is to simulate a scene with strong sunlight in reality, such as noon in summer, Figure 5 (b) is to simulate a scene with insufficient light in reality, such as evening and night.

[0082] In the self-built dataset for YOLOX model training, the number of targets in 59% of the images does not exceed 3 people, resulting in a situation where the ratio of positive to negative samples is less than 1 / 4 when calculating the target loss. To alleviate the imbalance problem of positive and negative samples and easy and difficult samples, the present invention also replaces the original Sigmoid-BCELoss with FocalLoss, and the accuracy of the model after replacement has increased by 0.9%. In the experiment, the parameters γ and α are appropriately modified to achieve the optimal effect. The calculation formula of FocalLoss is:

[0083]

[0084] where y′ is the predicted value, and its value ranges from 0 to 1; α and γ are hyperparameters, γ is related to balancing positive and negative samples, and α is related to strengthening the learning of difficult samples.

[0085] In a specific embodiment, the cosine annealing decay strategy is used to dynamically adjust the learning rate, and the minimum learning rate is set to 0.05; the trained PTH format model is converted into the Open Neural Network Exchange (ONNX) format, and then the RKNN Toolkit is used to convert the ONNX format model into an RKNN format model that can be accelerated on the NPU. Finally, in the mobile terminal, the relevant interfaces in the RKNN library are used to load and use the RKNN format model, and finally the process of accelerating inference using the NPU is realized.

[0086] The applications of the three improvement methods are compared with the network training results of the original YOLOX-SORT algorithm on the self-built dataset, as shown in Table 1.

[0087] Among them, Method 1: Use the modified Mosaic data augmentation; Method 2: Replace Sigmoid-BCELoss with FocalLoss; Method 3: YOLOX integrated with Kalman prediction.

[0088] Table 1 Network training results

[0089]

[0090] The experimental data of the average accuracy of the original YOLOX-SORT algorithm and the three improvement methods in 150 rounds of training are as Figure 6 shown, where Figure 6 (a) is the precision rate in 150 rounds of training of the average accuracy, Figure 6 (b) is the recall rate in 150 rounds of training of the average accuracy.

[0091] The comparison of the MOT metrics between the original YOLOX-SORT algorithm and the bus passenger flow detection method based on RGB images proposed by the present invention is shown in Table 2.

[0092] Table 2 Network training results

[0093]

[0094] Comparing the experimental results deployed on the mobile side, as shown in Table 3. Among them, Experiment 1: Direct deployment; Experiment 2: Using NPU for accelerated inference; Experiment 3: Replacing Focus with ordinary convolution.

[0095] Table 3 Network Training Results

[0096]

[0097] Verified by the above experimental data, the accuracy rate of the bus passenger flow detection method proposed by the present invention for counting getting-on and getting-off behaviors can reach more than 94%, which can meet the actual needs.

[0098] The present invention also provides a bus passenger flow detection system based on RGB images, and the system includes:

[0099] A video stream acquisition module, configured to acquire a video stream of passengers getting on and off the bus; the video stream includes multiple frames of images;

[0100] A prediction module, configured to, for the Kth frame of image, predict the target position in the Kth frame of image through the parameters of the Kalman filter of the (K - 1)th frame of image to obtain the prediction result of the Kalman filter;

[0101] An image generation module, configured to, for the Kth frame of image, generate a to-be-predicted image based on the prediction result of the Kalman filter and the Kth frame of image;

[0102] A matching module, configured to use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model;

[0103] A determination module, configured to determine the bus passenger flow according to the matching result.

[0104] Among them, the image generation module specifically includes:

[0105] A mask generation unit, configured to generate a target position mask from the prediction result of the Kalman filter;

[0106] An overlay unit, configured to overlay the target position mask and the Kth frame of image with different weights to generate a to-be-predicted image.

[0107] A detection module, configured to input the to-be-predicted image into the trained YOLOX model to obtain the detection result of the YOLOX model;

[0108] Among them, the matching module specifically includes:

[0109] A calculation unit for calculating the intersection over union of the prediction result of the Kalman filter and the detection result of the YOLOX model;

[0110] A construction unit for constructing a cost matrix based on the intersection over union;

[0111] A matching unit for matching the prediction result of the Kalman filter and the detection result of the YOLOX model based on the cost matrix using the Kuhn-Munkres algorithm, and the matching result is the target trajectory.

[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0113] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A bus passenger flow detection method based on RGB images, characterized in that, this method includes: Obtain the video stream of passengers getting on and off the bus; the video stream includes multiple frames of images; For the Kth frame of image, predict the target position in the Kth frame of image through the parameters of the Kalman filter that predicts the (K - 1)th frame of image, and obtain the prediction result of the Kalman filter; specifically include: For the Kth frame of image in the video stream, use the posterior state estimate of the (K - 1)th frame of image to calculate the prior state estimate of the Kth frame of image, and use the posterior estimate covariance matrix of the (K - 1)th frame of image to calculate the prior estimate covariance matrix of the Kth frame of image. Let the time when the Kth frame of image is obtained be t, and the time when the (K - 1)th frame of image is obtained be t - 1, then the specific calculation formula is as follows: Among them is the prior state estimate of the K-th frame image, is the prior covariance matrix estimate of the K-th frame image; is the posterior state estimate of the (K - 1)-th frame image, P t-1 is the posterior covariance matrix estimate of the (K - 1)-th frame image; includes the center coordinates, area, aspect ratio of the prediction box, and the change rates of the center coordinates and area between frames. The specific values are obtained by updating the (K - 1)-th frame image through a Kalman filter. F is the state transition matrix at the time of the K-th frame image, and its value is: B is the control matrix, and u is the control quantity value, both are taken as 0; Q is the process noise matrix, which comes from the loss, overlap and inaccuracy of the detection frame, and the value is: For the Kth frame of image, generate the image to be predicted based on the prediction result of the Kalman filter and the Kth frame of image; specifically include: Generate a target position mask from the prediction result of the Kalman filter; Overlay the target position mask and the Kth frame of image with different weights to generate the image to be predicted; Input the image to be predicted into the trained YOLOX model to obtain the detection result of the YOLOX model; where the YOLOX model includes a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer; the convolutional layer is a normal convolutional layer with a convolutional kernel size of 2 and a stride of 2; Use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model; Determine the bus passenger flow according to the matching result.

2. The bus passenger flow detection method based on RGB images according to claim 1, characterized in that, For the first frame of image, directly detect the target position through the trained YOLOX model.

3. The bus passenger flow detection method based on RGB images according to claim 1, characterized in that, Use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model, specifically including: Calculate the intersection - over - union ratio of the prediction result of the Kalman filter and the detection result of the YOLOX model; Construct a cost matrix through the intersection - over - union ratio; Based on the cost matrix, use the Kuhn - Munkres algorithm to match the prediction result of the Kalman filter and the detection result of the YOLOX model, and the matching result is the target trajectory.

4. The bus passenger flow detection method based on RGB images according to claim 3, characterized in that, Determine the bus passenger flow according to the matching result, specifically including: Set a counting area; the counting area is the area where passengers get on and off the bus; Determine the passenger flow according to the positional relationship between the target trajectory and the counting area.

5. A bus passenger flow detection system based on RGB images, characterized in that, The system includes: A video stream acquisition module for acquiring the video stream of passengers getting on and off the bus; the video stream includes multiple frames of images; A prediction module, which is used to predict the target position in the K-th frame image by predicting the parameters of the Kalman filter of the (K - 1)-th frame image for the K-th frame image, and obtain the prediction result of the Kalman filter; specifically including: for the K-th frame image of the video stream, using the posterior state estimation of the (K - 1)-th frame image to calculate the prior state estimation of the K-th frame image, and using the posterior estimation covariance matrix of the (K - 1)-th frame image to calculate the prior estimation covariance matrix of the K-th frame image. Let the time when the K-th frame image is obtained be t, and the time when the (K - 1)-th frame image is obtained be t - 1, then the specific calculation formula is as follows: where is the prior state estimate of the K-th frame image, is the prior covariance matrix estimate of the K-th frame image; is the posterior state estimate of the (K - 1)-th frame image, and P t-1 is the posterior covariance matrix estimate of the (K - 1)-th frame image; includes the center coordinates, area, aspect ratio of the prediction box, and the change rates of the center coordinates and area between frames. The specific values are obtained by updating the (K - 1)-th frame image through a Kalman filter. F is the state transition matrix at the time of the K-th frame image, and its value is: B is the control matrix, and u is the value of the control quantity, both of which are taken as 0; Q is the process noise matrix, which comes from the loss, overlap and inaccuracy of the detection box, and the value is: An image generation module, which is used to generate a to-be-predicted image for the K-th frame image based on the prediction result of the Kalman filter and the K-th frame image; specifically including: Generating a target position mask from the prediction result of the Kalman filter; Superimposing the target position mask and the K-th frame image with different weights to generate a to-be-predicted image; A detection module, which is used to input the to-be-predicted image into the trained YOLOX model to obtain the detection result of the YOLOX model; where the YOLOX model includes a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer; the convolutional layer is a common convolutional layer with a convolutional kernel size of 2 and a stride of 2; A matching module, which is used to match the prediction result of the Kalman filter and the detection result of the YOLOX model by using the Kuhn-Munkres algorithm; A determination module, which is used to determine the bus passenger flow according to the matching result.

6. The bus passenger flow detection system based on RGB images according to claim 5, wherein, the YOLOX model includes a convolutional layer, a CSPDarknet layer, a PAFPN layer, and a HEAD layer; the convolutional layer is a common convolutional layer with a convolutional kernel size of 2 and a stride of 2.

7. The bus passenger flow detection system based on RGB images according to claim 5, wherein, the image generation module includes: A mask generation unit, which is used to generate a target position mask from the prediction result of the Kalman filter; A superimposing unit, which is used to superimpose the target position mask and the K-th frame image with different weights to generate a to-be-predicted image.

8. The bus passenger flow detection system based on RGB images according to claim 5, wherein, the matching module includes: A calculation unit, which is used to calculate the intersection over union of the prediction result of the Kalman filter and the detection result of the YOLOX model; A construction unit, which is used to construct a cost matrix through the intersection over union; A matching unit, which is used to match the prediction result of the Kalman filter and the detection result of the YOLOX model based on the cost matrix by using the Kuhn-Munkres algorithm, and the matching result is the target trajectory.

Citation Information

Patent Citations

  • Target association method, computer equipment and storage medium

    CN113139416A