Long video non-motor vehicle converse running proportion detection algorithm based on multi-model ensemble learning
Through multi-model ensemble learning and NeRF generation data set methods, the detection and prediction problems of non-motor vehicle retrograde behavior in surveillance video are solved, efficient and accurate estimation of reverse riding proportions is achieved, the challenges of large data volume, direction diversity and label acquisition difficulties are overcome, and the applicability and accuracy of the detection model are improved.
Patent Information
- Application Number
- CN202311392174.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-06
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to efficiently process a large amount of surveillance video data to achieve accurate monitoring and prediction of non-motor vehicle counter-travel behavior, especially in the problems of large amount of data, diverse cycling directions, uncertainty in non-motor vehicle shapes and difficulty in obtaining label data, resulting in insufficient accuracy of detection and prediction.
The multi-model ensemble learning method is adopted, combined with object detection and direction perception models, and the object rotation angle is predicted through K-means clustering and phase shift encoding. The synthetic data set is generated using NeRF for pre-training, and the Monte Carlo algorithm is used to sample video frames, the And strategy is used to fusion model output, and the reverse riding ratio estimation is performed in combination with K-means post-processing.
It improves the accuracy and efficiency of non-motor vehicle counter-travel behavior detection, reduces computing resource requirements, provides stable reverse riding proportion prediction, reduces errors and variances, and improves the applicability of the model in real scenarios.
Smart Images

Figure CN120279054A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning, ensemble learning, and computer vision, and in particular to a long video non-motor vehicle reverse ratio detection algorithm based on multi-model ensemble learning. Background Art
[0002] 1. Challenges of Reverse Cycling Behavior Reverse cycling behavior has become an important issue in urban road traffic management and road safety. This behavior not only endangers the safety of cyclists themselves but also poses a potential threat to other road users. Therefore, effective monitoring and prediction of reverse cycling behavior are crucial for urban traffic management. However, to achieve accurate prediction and monitoring, multiple technical challenges need to be overcome.
[0003] 2. Large Amount of Data Surveillance cameras usually capture videos at various sections of the city, resulting in a large amount of surveillance video data. These videos usually contain thousands to millions of frames of images, so a large amount of computing resources and time are required to process this data. In practical applications, processing such a large amount of data may lead to delays and performance degradation. Therefore, a method is needed to reduce the amount of data while maintaining the accuracy of prediction for more efficient monitoring and prediction of reverse cycling behavior.
[0004] 3. Diversity of Cycling Directions Surveillance cameras are installed at different positions and angles, and the reverse cycling behavior captured may have diversity in the cycling direction. For example, the reverse cycling direction in one video may be from top to bottom, while in another video, it may be completely opposite. To address this diversity, we adopted the K-means clustering algorithm to divide the predicted directions from a single camera into two groups. In this clustering process, small direction clusters are regarded as signs of incorrect cycling, while large direction clusters are regarded as signs of forward cycling. This method helps to handle different cycling direction situations to improve the accuracy of prediction.
[0005] 4. Uncertainty of Non-Motor Vehicle Shapes The shapes and sizes of non-motor vehicles may vary greatly in different situations. This uncertainty leads to inaccuracies in bounding boxes when detecting and tracking non-motor vehicles in videos. Especially in the case of reverse cycling, since the vehicle may appear at different angles and positions, traditional detection models may be more prone to making incorrect direction predictions. To solve this problem, we adopted a multi-model approach, which includes a detection model and a direction perception model. By combining these two models, we can better analyze and distinguish situations such as wrong-way driving, vehicle stationary, or moving backward due to intersections, thereby improving the accuracy of non-motor vehicle detection.
[0006] 5. Difficulty in obtaining labeled data To train monitoring and prediction models, a large amount of labeled data is usually required. However, obtaining non-motor vehicle data with orientation labels in the real world is a difficult and expensive task. In our method, we propose a "training - adjustment" architecture that combines the ability of NeRF (Neural Radiance Field) to generate viewpoints. In the pre-training stage, we use instant-ngp to construct a dataset that includes image data captured from real-world scenes. Next, we use instant-ngp to quickly train a virtual 3D model and render images from specific viewpoints. Through this method, we can generate various synthetic data with known orientation labels, thus solving the problem of obtaining labeled data and narrowing the domain gap between the test set and the fine-tuning set at the same time. Summary of the Invention
[0007] 1. Object tracking algorithm based on object detection model The multi-object tracking problem can be regarded as a data matching problem, that is, correlating the object detection results between adjacent frames in a video. A relatively successful and widely concerned method in this field is the "Simple Online and Real-time Tracking" (SORT) algorithm.
[0008] The SORT algorithm is an algorithm that integrates tracking technology into object detection technology specifically designed for the object tracking task. It uses a detector (such as Faster-RCNN) to locate the position of the object in each frame, and calculates the intersection over union (IoU) between the detection boxes of two frames to find the connection between the detection boxes of two frames to achieve the tracking effect.
[0009] In our research, we get inspiration from this algorithm and calculate the connection between adjacent frames in a similar way to obtain the object angle information.
[0010] 2. Object angle prediction Previous angle prediction problems generally occurred in the detection of rotating objects. For the problem of rotating object detection, the rotation of the object is often observed from an aerial view, that is, the object still maintains its original shape after rotation, and only the orientation of the image changes. For observations from other angles, the shape of the object and even the observed form of the entire object will change after rotation. Therefore, we propose a brand-new method for predicting the rotation angle of an object from a non-aerial view. We use a phase shift coding (PSC) to convert the periodic and discontinuous angle into a continuous vector representation when predicting the rotation angle of the object.
[0011] 3. Model integration Ensemble learning is a powerful technique that combines individual models to achieve better overall performance. One approach to ensemble learning is multi-model combination, where the fusion strategy plays a crucial role in ensemble learning as it determines how the outputs of these individual models are combined to obtain a more powerful result. In our work, we adopted the "AND strategy" as the fusion strategy. Description of the Drawings
[0012] Figure 1 This is a schematic diagram of the overall process of this algorithm.
[0013] Figure 2 This is a schematic diagram of the rendered image of non-motor vehicles.
[0014] Figure 3 This is a schematic diagram of the data obtained from the real world.
[0015] Figure 4 This is a schematic diagram of the evaluation data in the real scenario.
[0016] Figure 5 These are the loss function values for training with different methods.
[0017] Figure 6 This is a comparison of the performance between the ensemble model and each individual model.
[0018] Figure 7 This is a comparison between different training methods.
[0019] Figure 8 This is a comparison of the performance between methods using different expected intervals. Embodiments
[0020] A multi-model non-motor vehicle reverse detection algorithm based on deep learning includes the following steps: (1) Frame sampling; We use the Monte Carlo algorithm to sample frames from the video. Specifically, we define and hyperparameters. The random variable is independently generated according to the uniform distribution, indicating which moment of the video we want to include in the model.
[0021] Expected interval represents the average distance between adjacent random variables. We can represent the expected interval using the following formula:
[0022] On the other hand, the maximum deviation represents the maximum difference between any distance value and its expected value. We can represent the maximum deviation using the following formula:
[0023] where represents the th distance value, and its calculation formula is .
[0024] In this specific method, we extract the required moments from the video. The entire processed dataset is represented as , where and respectively represent the first frame image and the second frame image obtained from the moment . The entire inference process is based on the processed dataset .
[0025] (2) Angle prediction model based on object direction; Since the model aims to predict the direction of an object, which is a continuous and periodic variable, we apply a Phase-Shifting Coder (PSC) to convert the discontinuous degree system into a continuous -dimensional vector. In our work, we set to 3. The working principle of PSC is as follows: Encoding:
[0026] Decoding:
[0027] Specifically, we apply a pre-trained backbone network, here using ResNet-101, to generate an embedding vector for the image, and then apply a linear layer to convert the embedding vector into a vector .
[0028] During the training process, the label is encoded as . Then the loss is calculated as follows:
[0029] During the inference process, the vector is decoded into the final output .
[0030] In the direction prediction task, we define the metric as the distance between the predicted value and the label value. Since it is a cyclic number, its formula can be expressed as:
[0031]
[0032] Here, represents the label value, and represents the predicted value. and are both within the range, expressed in degrees. The error is calculated based on the difference between the maximum and minimum values of and . If this difference is less than or equal to 180 degrees, the error is equal to . Otherwise, if the difference is greater than 180 degrees, the error is calculated as . In the experimental section, we use this metric to evaluate the performance of the model.
[0033] To train our direction-based model, we adopted an effective pre-training and fine-tuning architecture. Considering the challenge of obtaining real-world data with precise direction labels and addressing the long-tail property problem, we generated synthetic data. These synthetic data are used as the pre-training data for the model and provide it with a beneficial bias towards direction information. By leveraging this approach, we aim to enhance the model's ability to understand and utilize direction cues in real scenarios.
[0034] First, we captured 360-degree videos using an ordinary camera in a real scenario. Then, we extracted 30 - 40 frames from the videos and utilized a trained COLMAP (Structure-from-Motion and Multi-View Stereo) system to estimate the position and pose parameters of the camera both inside and outside the captured scene. Using this information, we reconstructed a 3D model using instant-ngp.
[0035] Next, we utilized pre-designed camera poses to render images with corresponding direction labels. To simulate the real scenario, we rendered images at two different heights for each direction, as shown in Figure 2 . By default, we captured images at 10-degree intervals, generating 72 labeled images for each 3D model.
[0036] This process enables us to create a diverse dataset of labeled images, which serves as valuable pre-training data for our direction-based model. In total, we performed 12 3D models, and the total number of images for pre-training was 864.
[0037] After pre-training, the entire model is fine-tuned using real-world data.
[0038] (3) Ensemble learning method; We divide our multi - model ensemble learning into three parts: detection, direction prediction, and integration strategy. The details are as shown in the algorithm.
[0039]
[0040] At the beginning, a pair of consecutive frames are inferred by the detection model and can be represented as where \(D\) represents the detection function, represents the output bounding box. In this case, we use YOLOv5 as our selected detection model. is a function for calculating the intersection - over - union (IoU) matrix between two lists of bounding boxes. The expression represents a boolean matrix where when is true, is zero, and otherwise it is one. The value of is a predefined constant, set to 0.98. This is to exclude some stable vehicles from consideration. Then, we apply the Hungarian algorithm, represented as the function \(H\), to match the bipartite graph, where each side represents the bounding boxes of one of the two frames. The function \(H\) returns a series of matching index pairs, saved as For each pair of matches obtained by the Hungarian algorithm, we first calculate a detection - based direction by computing the direction of the center points of the bounding boxes:
[0041]
[0042] where represents the center point of the bounding box in the \(i\) - th frame.
[0043] Next, we apply our direction - aware model, represented as to extract direction information from the image - level data. First, we obtain a pair of bounding boxes from the matched by the Hungarian algorithm. Then, we crop the corresponding images from the original frames according to these bounding boxes. Each cropped image is then individually input into the trained direction - aware model to obtain the average direction prediction of the images captured in two consecutive frames.
[0044] The last step involves combining the outputs of the detection pipeline and the direction - aware model. In this step, we first verify and whether they satisfy the And strategy. Simply put, if we consider the output to be valid; otherwise, it is considered invalid. By default, we set Set to 50 to ensure and consistency. If the data is considered valid, we take their average as the final output.
[0045] Mathematically, it can be shown that the integration of two models using the And strategy can improve the final performance compared to using a single model. The proof is as follows.
[0046] Suppose the posterior probabilities of two models making mistakes are respectively and . Under ideal conditions, when one model makes a mistake and the other does not, the And strategy will identify this situation and consider the output invalid. In this case, the probability of being considered valid can be expressed as:
[0047] The integrated model will only make a mistake when both models make mistakes, and the probability of this happening can be written as:
[0048] However, we should only consider valid cases, so the actual probability of the integrated model making a mistake is:
[0049]
[0050] Next, we will prove that for all , there are and . It is generally considered that the error rate of each model is less than 0.5.
[0051] We take the derivative of with respect to (or ):
[0052]
[0053] Therefore, is a monotonically increasing function of (or ).
[0054] In the worst case, when , we have:
[0055] For any < 0.5, will be less than .
[0056] In this way, we finally proved that for all , and .
[0057] (4) K-Means post-processing; We use the K-means algorithm to process the list of directions and divide it into two groups. Specifically, if the distance between the two centers of the clustering group is less than 90, we can conclude that no reverse cycling occurs.
[0058] The working steps of the K-means algorithm are as follows: - Initialization: Randomly select two initial clustering centers as starting points.
[0059] - Assignment: For each direction, calculate its distance to the two clustering centers and assign it to the nearest center.
[0060] - Update: Calculate the new center of each cluster by computing the average of the directions assigned to that cluster.
[0061] - Repeat steps 2 and 3 until convergence: Iteratively execute the assignment and update steps until the clustering centers are stable and the assignments remain unchanged or meet the specified stopping criterion.
[0062] By iteratively assigning directions to clusters and updating the cluster centers, K-means effectively divides the directions into two groups. Based on our experience, we designate the majority group as the forward direction and consider the other group as an indication of reverse cycling.
[0063] Finally, we calculate the number of directions in each group and take the reverse cycling ratio as the output. This provides us with a quantitative measure to gauge the prevalence of reverse cycling in the dataset. Example
[0064] 1. Dataset Taking everything into consideration, we propose three datasets for different purposes. These datasets include a dataset for direction awareness, a detection training dataset, and a final validation dataset. Each dataset plays a specific role in our model development and evaluation process.
[0065] Direction-aware dataset: The Direction-aware dataset is specifically designed for training direction-aware models. The dataset consists of three subsets: pre-training set, fine-tuning set, and validation set. The pre-training set consists of synthetic images generated using instant-ngp. It includes 12 different 3D models, and for each model, we generate 72 labeled images taken from different heights, such as Figure 2 In total, the pre-training set contains 864 images. The fine-tuning set consists of real-world images that have been labeled and manually adjusted by the detection method, some of which are shown in Figure 3 As shown in Figure 2. The fine-tuning set contains 948 images and the validation set contains 132 images.
[0066] Detection Dataset: To train and validate the detection model, we constructed a training set and a validation set. There is only one category in this dataset, namely non-motorized vehicles. These datasets contain a total of 223 images taken under different conditions, covering 474 valid bounding boxes. The ratio between the training set and the validation set is 8:2, and the results of all our experiments are based on this ratio. Our experiments show that this amount of data is sufficient to effectively train a detection model.
[0067] Final validation dataset: To establish the evaluation metric, we randomly selected four 5-minute videos captured by road cameras at different locations. These videos serve as the basis for judging the performance of the method in terms of speed and accuracy level. By utilizing diverse video materials, such as Figure 4 As shown, we ensure that our approach is thoroughly evaluated in various real-world scenarios. This evaluation approach allows us to assess the effectiveness and reliability of our approach in accurately and efficiently performing the task of predicting the proportion of oncoming bicycles from road camera videos.
[0068] 2. Comparison with non-ensemble methods To demonstrate the advantages of our ensemble learning approach, we build two benchmarks: a single detection-based approach and a single direction-aware model-based approach. The detection-based approach is similar to our ensemble approach in that we take the direction (Odet) between the center points of two bounding boxes as the final output. On the other hand, the direction-aware model approach is different in that if we do not need the detection-based direction, we do not need to process both frames simultaneously. In this benchmark, we perform detection on only one frame and apply the direction-aware model to the cropped image. It should be noted that all model parameters and hyperparameters remain unchanged when comparing these methods. Specifically, we use YOLOv5-m as the detection model and a fine-tuned Resnet-101 plus a phase shift encoder as the direction-aware model. When extracting frames using the Monte Carlo algorithm, the expected interval Eg is set to 3 seconds. The interval time between two consecutive frames is set to 0.02 seconds.
[0069] As shown in Table 1, we list the number of predictions in the forward (F) and reverse (R) directions for the three methods. In our ensemble method, we achieved a mean error of 0.96%, which is considered acceptable for estimating the proportion of wrong-way bicycles in the video. In addition, the variance of the mean error is also small, indicating that the model is consistent in providing reliable predictions without making significant errors.
[0070] However, in both the detection-based approach and the direction-aware model approach, the mean error and variance are quite high. This means that the outputs of these methods are not accurate enough to provide acceptable final scale predictions. In contrast, the ensemble learning approach exploits the synergy between the two models. Considering that this task involves scale prediction, it proves effective to invalidate the outputs of the two models if they cannot support each other. By combining the strengths of both models, our ensemble learning approach improves the accuracy and reliability of the final prediction.
[0071] 3. Ablation Study of Direction Awareness Model like Figure 5 As shown in Figure 2, we compared four different training methods. Table 2 provides the specific numerical comparison of these four different training methods.
[0072] We designed a pre-train-fine-tune architecture to train the direction-aware model. To evaluate the impact of the pre-training dataset and this specific training architecture on the model performance, we performed an ablation study. This study allowed us to evaluate the effectiveness and importance of these components on the model performance.
[0073] Figure 5 A comparison of the change in validation error caused by different training methods is shown. The training methods are divided into the following categories: "Pure" means only real data is used for training.
[0074] "Pretrain" means training using only synthetic data.
[0075] "Mix" means training by combining real data and synthetic data.
[0076] "Finetune" denotes the process of fine-tuning on real data after a pre-training phase on synthetic data.
[0077] To facilitate subsequent analysis, we provide the specific numerical results of our method in Table 2.
[0078] 4. Monte Carlo Sampling Experiment The use of the Monte Carlo algorithm, i.e., the method of extracting frames, directly determines the amount of information that the subsequent model can receive. This algorithm plays a key role in our method. Therefore, we selected different hyperparameters, denoted as Eg, to examine their impact on the final result. In this context, Eg represents the expected time interval between two consecutive conditions (specifically, Vi and Vi+1).
[0079] To evaluate the performance of our method, we use two metrics to compare the final output: error and output range. The output range represents the confidence of the model in its prediction. Ideally, we hope to obtain a reliable model with a small output range.
[0080] We conducted this experiment on scenario 1 of the final validation dataset. As shown in Table 3, we compared the range of Eg values from 1.5 seconds to 9 seconds, and each experiment was repeated five times. It should be noted that the "error" value in Table 3 represents the average of five errors, rather than the error of the average ratio.
[0081] In our experiment, we observed that an increase in Eg led to fluctuations in the output and an increase in error, which was expected. The choice of Eg depends on achieving a balance between efficiency and accuracy. In our previous experiment, based on this trade-off, we chose Eg to be 3 seconds. However, it should be noted that the choice of Eg is greatly influenced by the video length. In our final validation dataset, the duration of each video is 5 minutes. If the video is longer, choosing a larger Eg value would be more flexible.
Claims
1. A long video non-motor vehicle reverse proportion detection algorithm based on multi-model ensemble learning, characterized in that, The method includes the following steps: (1) Frame sampling: using the Monte Carlo algorithm to sample frames from the video. (2) An angle prediction model based on the object direction, which uses a phase shift encoder; and a pre-training-fine-tuning architecture is used to train the model. (3) Ensemble learning method: an effective combination of the angle prediction model and the detection-based model to determine whether the non-motor vehicle is driving against the flow.
2. According to item (2) of the algorithm described in claim 1, the phase offset encoder converts the periodic angle value into a three-dimensional vector representation.
3. According to item (2) of the algorithm described in claim 1, the pre-training-fine-tuning architecture comprises the following steps: (1) shooting a 360-degree video using a regular camera; (2) estimating the position and attitude parameters of the camera using the COLMAP system; (3) reconstructing a three-dimensional model using instant-ngp; (4) rendering an image with a direction label using a pre-designed camera attitude; (5) training the model using the rendered image as pre-training data; and (6) fine-tuning the model using real-world data.
4. According to item (3) of the algorithm described in claim 1, the integration strategy is an "AND strategy", that is, the non-motor vehicle is considered to be traveling in the wrong direction only when the detection result and the direction prediction result are both judged to be traveling in the wrong direction.