Traffic target flow statistical algorithm based on content guidance module

By introducing content guidance module and ESIoU loss function into the traffic target traffic statistics algorithm, the problems of low quality and slow convergence speed of prediction boxes in complex traffic contexts are solved, and higher detection accuracy and faster convergence speed are achieved.

CN120220093APending Publication Date: 2025-06-27HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376632.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the context of complex road traffic, traffic targets are severely obstructed, resulting in the problems of low quality of the prediction box and slow convergence speed.

Method used

The traffic target flow statistics algorithm based on the content guidance module (CGM) is used to weight the feature maps at different levels by introducing a channel attention mechanism, and the adjustment and convergence of the prediction box are accelerated by combining the ESIoU loss function.

Benefits of technology

It effectively reduces the generation of low-quality prediction boxes, improves the accuracy and convergence speed of object detection, shortens training time, and improves the positioning accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220093A_ABST
    Figure CN120220093A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and discloses a traffic target flow statistical algorithm based on a content guidance module, which comprises the following steps: S1, preparation work: training through a target detection network containing the content guidance module, a training environment and parameters, adopting a cosine learning rate and a ReLU activation function, constructing a total loss function through weight combination, and calculating the total loss function; training is carried out through a target detection network containing a content guidance module, a training environment and parameters, the ability of the model to extract traffic target features under a complex background is enhanced, and feature maps of different levels are weighted by using a channel attention mechanism, so that the traffic target feature extraction efficiency is improved in a complex traffic scene with dense vehicles or mutual shielding. In addition, the ESIoU loss function designed by the invention integrates the central point distance, the aspect ratio and the newly added aspect difference penalty term, the rapid convergence of the prediction frame to the real frame is effectively guided, and the convergence speed in the model training process is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a traffic target flow statistical algorithm based on a content guidance module. Background Art

[0002] Traffic target flow statistics refers to installing flow statistical devices in traffic road areas, and by detecting and tracking targets (such as vehicles, pedestrians, bicycles, etc.) in traffic videos, images or sensor data, the flow distribution of the targets within a certain area or time range is statistically calculated. With the rapid development of deep learning and computer vision, multi-target detection and tracking methods based on video data have become the mainstream. On this basis, some new traffic target flow statistical systems under new intelligent monitoring are constantly emerging and are gradually expanding their usage scope. These systems have excellent performance, low cost, small and easy-to-install monitoring devices, can accurately distinguish moving vehicles and people, and can meet various types of requirements. Their core process can be summarized as: using a convolutional neural network to implement target detection, realizing multi-target trajectory tracking by associating detection results, using ReID or other technologies to improve tracking accuracy, and finally performing post-processing to obtain the flow statistical results.

[0003] Almost all existing detection models adopt FPN to fuse features at different levels, effectively utilize high-resolution shallow features (including spatial information) and low-resolution deep features (including semantic information), and improve the detection ability of multi-scale targets. However, in practice, the background features of road traffic are complex, and traffic target occlusion is serious. How to reduce low-quality prediction boxes and how to accelerate convergence have become a difficult problem. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a traffic target flow statistical algorithm based on a content guidance module, which solves the problems of low-quality prediction boxes and slow convergence speed caused by complex road traffic background features and serious traffic target occlusion.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A traffic target flow statistical algorithm based on a content guidance module includes the following steps: S1. Preparation work: Train through a target detection network, training environment and parameters including a content guidance module, adopt a cosine learning rate and a ReLU activation function, construct a total loss function with weight combinations, where the BBR loss adopts the ESIoU loss, the classification loss adopts the VFL loss, and the regression loss adopts the DFL loss. After training is completed, convert the obtained Pytorch model into an RKNN model and perform quantization processing on the RKNN model; S2. Data Acquisition and Encoding: Use a camera to collect traffic videos, transmit the video stream to the embedded platform through the RTSP protocol, and decode the video stream using the hardware decoding function or relevant software libraries to obtain frame-by-frame image data; S3. Image Preprocessing: Crop the collected images to a fixed size and set the counting and statistical areas; S4. Model Inference: Use the optimized FPN structure for feature fusion to improve the accuracy and robustness of object detection. Use the NPU of the embedded platform for multi-threaded inference to obtain the bounding box and object category label results; S5. Post-processing: Set a threshold based on the confidence score to filter out low-confidence objects and perform non-maximum suppression operations. Screen out the bounding box most likely to be the object by calculating the ESIoU value between the bounding boxes; S6. Multi-object Tracking: Input the detection results into the multi-object tracking algorithm to obtain the tracking results, and perform traffic statistics based on the tracking ID numbers; S7. Video Streaming and Result Echo: Re-encode the image frames containing the detection results into a video stream, push it to the receiving end through the RTSP protocol, and use a web player to display the detection results in real time.

[0006] Preferably, in step S1, the content guidance module takes the main feature map and the auxiliary feature map as the CGM input, fuses them through Concat, then performs global average pooling to generate a global feature vector. The first fully connected convolution reduces the number of channels from 2C to 2C / r, where r is the reduction factor. The second fully connected convolution then restores the number of channels to 2C through 1×1 convolution to generate the fused attention weights. Subsequently, Sigmoid activation is performed to normalize the weights, restricting them to the [0, 1] interval for channel weight adjustment. The obtained results are first weighted and adjusted with the original inputs respectively, then cross-fused once to extract deeper information, and finally Concat fusion is performed to obtain the output Y (with a size of (2C, H, W)).

[0007] Preferably, the CGM input calculation method is as follows: Wherein, is the main feature map, is the auxiliary feature map, is the globally feature vector after weighted adjustment.

[0008] Preferably, in step S1, some parameter settings in the training stage are as follows: epoch = 400. After 400 rounds of iteration, the model with the best performance in all iteration processes is selected as the model training result; batch size = 16; the SGD optimizer is adopted; the weight decay coefficient = 0.0005; the momentum hyperparameter = 0.937; the cosine learning rate is used to optimize the result, and the specific formula is: where is the learning rate at the t-th iteration, is the minimum learning rate, is the maximum learning rate, is the current iteration number, is the maximum iteration number.

[0009] Preferably, the constructed total loss function accelerates the adjustment and convergence of the prediction box through the width-height difference penalty terms of two rectangles, and the specific formula is: where w and h respectively represent the width and height of the predicted bounding box, respectively represent the width and height of the true bounding box, , respectively represent the width and height of the minimum closure, and the v reference formula is used to measure the difference in aspect ratio, indicating the center point distance penalty term, represents the aspect ratio penalty term, represents the width-height difference penalty term.

[0010] Preferably, in step S5, all bounding boxes are sorted according to the confidence of the bounding boxes. Starting from the bounding box with the highest confidence, calculate its ESIoU with other bounding boxes. The specific calculation method is: where represents the value of the loss function, represents the intersection over union between the predicted bounding box and the true bounding box, the penalty term in the loss function.

[0011] Preferably, in step S5, if the ESIoU of the bounding box with the highest confidence with several other bounding boxes is greater than the NMS threshold, the bounding box is deleted, and the remaining bounding boxes continue to be processed according to the above steps, and the remaining object detection results are output to step S6.

[0012] Preferably, in step S2, the RTSP stream is transmitted to the embedded platform through the TCP protocol, and the hardware decoding function and related software library are FFmpeg.

[0013] Preferably, in step S4, the optimized FPN structure includes a content guidance module, and the content guidance module is used for feature fusion at the connection of the FPN.

[0014] Preferably, in step S7, the video streaming and result echo are achieved by re-encoding the image frame containing the detection result into a video stream and pushing it to the specified receiver using the RTSP protocol.

[0015] Beneficial Effects The present invention provides a traffic target flow statistics algorithm based on a content guidance module. Compared with the prior art, it has the following beneficial effects: In the present invention, by introducing a content guidance module (CGM), the present invention enhances the model's ability to extract traffic target features in complex backgrounds, uses the channel attention mechanism to weight feature maps at different levels, and strengthens the information useful for target detection in the feature maps. Thus, in complex traffic scenarios with dense vehicles or mutual occlusion, the generation of low-quality prediction boxes is effectively reduced, and the accuracy of target detection is improved. In addition, the ESIoU loss function designed in the present invention combines the center point distance, aspect ratio, and newly added width-height difference penalty terms. These penalty terms jointly act on the bounding box regression in the model training process, effectively guiding the prediction box to quickly converge to the true box, accelerating the convergence speed in the model training process, shortening the training time, and improving the localization accuracy of the model. Brief Description of the Drawings

[0016] Figure 1 It is a flowchart of a traffic target flow statistics algorithm based on a content guidance module proposed by the present invention; Figure 2 It is a schematic diagram of the content guidance module CGM in a traffic target flow statistics algorithm based on a content guidance module proposed by the present invention; Figure 3 It is an actual effect diagram of the prediction box and the true box in a traffic target flow statistics algorithm based on a content guidance module proposed by the present invention; Figure 4 It is a prediction change diagram of the prediction box and the true box in a traffic target flow statistics algorithm based on a content guidance module proposed by the present invention; Figure 5 It is a change effect diagram of the prediction box and the true box in a traffic target flow statistics algorithm based on a content guidance module proposed by the present invention; Figure 6Simulation test diagram of a traffic target flow statistical algorithm based on a content guidance module proposed by the present invention; Figure 7 Convergence curve diagram of the localization loss when YOLOv8s+CGM+ESIoU and YOLOv8s are trained on the TD dataset; Figure 8 Schematic diagram of the center point distance and aspect ratio; Figure 9 Visualization effect diagram of the detection result after passing through CGM. Specific implementation manner

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Please refer to Figures 1 - 9 , the present invention provides two technical solutions, specifically including the following embodiments: Embodiment 1: A traffic target flow statistical algorithm based on a content guidance module includes the following steps: S1. Preparation work: Train through a target detection network, training environment, and parameters including a content guidance module, use a cosine learning rate and a ReLU activation function, construct a total loss function with a weight combination, where the BBR loss uses the ESIoU loss, the classification loss uses the VFL loss, and the regression loss uses the DFL loss. After training is completed, convert the obtained Pytorch model into an RKNN model and perform quantization processing on the RKNN model. In step S1, the content guidance module takes the main feature map and the auxiliary feature map as CGM inputs, fuses them through Concat, and then performs global average pooling to generate a global feature vector with a dimension of (2C, 1, 1). A single fully connected convolution reduces the number of channels from 2C to 2C / r, where r is the reduction factor. The second fully connected convolution then restores the number of channels to 2C through a 1×1 convolution to generate the fused attention weights, and then performs Sigmoid activation to normalize the weights and limit them in the [0, 1] interval for channel weight adjustment. The obtained results are first weighted and adjusted with the original inputs respectively, then cross-fused once to extract deeper information, and finally Concat-fused to obtain the output Y (with a size of (2C, H, W)). The calculation method of the CGM input is as follows: Among them, is the main feature map, It is an auxiliary feature map, It is the globally feature vector after weighted adjustment. Some parameter settings in the training stage are as follows: epoch = 400. After 400 rounds of iteration, the model with the best performance among all iterations is selected as the training result of the model; batch size = 16; the SGD optimizer is adopted; the weight decay coefficient = 0.0005; the momentum hyperparameter = 0.937; the cosine learning rate is used to optimize the result, and the specific formula is: where, is the learning rate at the t-th iteration, is the minimum learning rate, is the maximum learning rate, is the current iteration number, is the maximum iteration number. The total loss function is constructed to accelerate the adjustment and convergence of the prediction box through the width-height difference penalty terms of two rectangular boxes. The specific formula is:

[0019] where w and h respectively represent the width and height of the predicted bounding box, respectively represent the width and height of the ground truth bounding box, , respectively represent the width and height of the minimum closure. The v reference formula is used to measure the difference in aspect ratio and represents the center point distance penalty term, represents the aspect ratio penalty term, represents the width-height difference penalty term; S2. Data acquisition and encoding: A camera is used to collect traffic videos, and the video stream is transmitted to the embedded platform through the RTSP protocol. The video stream is decoded using the hardware decoding function or relevant software libraries to obtain frame-by-frame image data. The RTSP stream is transmitted to the embedded platform through the TCP protocol. The hardware decoding function and relevant software libraries are FFmpeg; S3. Image preprocessing: The collected images are cropped to a fixed size and the counting statistical area is set; S4. Model inference: The optimized FPN structure is used for feature fusion to improve the object detection accuracy and robustness. The NPU of the embedded platform is used for multi-threaded inference to obtain the bounding box and object category label results. The optimized FPN structure includes a content guidance module, and the content guidance module is used for feature fusion at the connection of the FPN; S5, post-processing: Set a threshold based on the confidence score to filter low-confidence targets, and perform non-maximum suppression operations. Filter out the bounding box that is most likely to be the target by calculating the ESIoU value between the bounding boxes. In step S5, sort all bounding boxes according to the confidence of the bounding box, starting from the bounding box with the highest confidence, and calculate its ESIoU with other bounding boxes. The specific calculation method is: in, represents the value of the loss function, represents the intersection-over-union ratio between the predicted bounding box and the true bounding box, The penalty term in the loss function is the highest confidence bounding box and the other bounding boxes. If the ESIoU is greater than the NMS threshold, the bounding box is deleted, and the remaining bounding boxes continue to be processed according to the above steps, and the retained target detection results are output to step S6. In a target detection task, multiple vehicle bounding boxes are detected, and their confidence levels are 0.9, 0.85, 0.7, 0.6, etc. from high to low. Starting from the bounding box with the highest confidence, calculate its ESIoU with other bounding boxes. If the ESIoU is greater than the NMS threshold (the default value is 0.25), it is deleted. The greater the overlap between the two bounding boxes, it means that they are likely to be detecting the same target; S6, multi-target tracking: input the detection results into the multi-target tracking algorithm to obtain the tracking results, and perform traffic statistics based on the number of tracked IDs; S7, video streaming and result echo: re-encode the image frames containing the detection results into video streams, push them to the receiving end through the RTSP protocol, and use a web player to display the detection results in real time. Video streaming and result echo re-encode the image frames containing the detection results into video streams, and push them to the specified receiving end using the RTSP protocol.

[0020] Embodiment 2: Based on Example 1, some parameter and rule settings in the training phase are as follows: the base learning rate lr0 = 0.01, and the cosine annealing algorithm is used to adjust the learning rate; the weight decay coefficient = 0.0005; the SGD optimizer is used; batch size = 32; the default value of epoch is 400. If the performance of the validation set does not improve for 100 consecutive rounds or reaches the number of iterations during the training process, the training stops, and the model with the best effect during the iteration process is selected as the model training result. The main data augmentation methods used are as follows: HSV space augmentation (the perturbation coefficients are 0.015, 0.7, and 0.4 respectively), horizontal flipping (50% probability), random translation (±10% ratio), multi-scale scaling (0.5∼2.0), Mosaic augmentation, RandAugment automatic augmentation, and random erasing (40% probability). In addition, in this chapter, the imgaug API is used to apply bad weather effects such as fog, frost, rain, and snow to some training images, enriching the data samples of non-clear and cloudy days in the dataset, which helps the model better adapt to detection tasks under various weather conditions and enhances robustness. To ensure fairness, the input size of the training images will be preprocessed to 640×640. The dataset used is the Traffic Dataset (TD dataset), which is composed of the open-source UA dataset and the BIT dataset mixed together.

[0021] The following are the results of the above experimental data: Method GFLOPs Params(M) FPS(bs = 8) mAP@0.5 <![CDATA[AP S > <![CDATA[AP M > <![CDATA[AP L > YOLOv8s + CIoU 28.4 10.6 139.3 64.8 15.3 42.6 57.3 YOLOv8s + CGM + CIoU 29.2 11.0 131.1 65.3 16.9 43.0 58.1 YOLOv8s + CGM + ESIoU 29.2 11.0 130.2 65.5 17.0 43.3 57.9 Ablation experiment of CGM on the TD dataset Loss / Method AP - Bic AP - Mot AP - Ped mAP - Other mAP@0.5 mAP@0.5:0.95 IoU(baseline) 44.5 51.7 57.8 71.0 61.2 39.4 GIoU 50.3(+13.0%) 54.6(+5.61%) 62.8(+8.65%) 74.1(+4.37%) 65.0(+6.12%) 42.7(+8.38%) DIoU 49.4(+11.0%) 56.0(+8.32%) 61.5(+6.40%) 73.9(+4.08%) 64.8(+5.88%) 42.8(+8.63%) CIoU 50.6(+13.7%) 55.3(+6.96%) 60.5(+4.67%) 74.0(+4.23%) 64.8(+5.88%) 42.6(+8.12%) Focal - CIoU 51.1(+12.3%) 54.4(+5.22%) 61.7(+7.09%) 73.7(+3.80%) 64.9(+6.05%) 42.8(+8.63%) ESIoU 51.5(+13.2%) 54.1(+4.64%) 63.7(+10.2%) 74.5(+4.93%) 65.5(+7.03%) 43.0(+9.14%) Comparative experiment of YOLOv8s using different loss functions on the TD dataset Among them, AP-Bic represents the AP value of Bicycle, AP-Mot represents the AP value of Motorcycle, AP-Ped represents the AP value of Pedestrain, and mAP-other represents the average AP value of Sedan, Bus, and Truck. It can be seen from the results shown in the table that the Focal-ESIoU loss performs better in processing the TD dataset, exceeding the CIoU loss, which is the most used and reliable loss function algorithm in the current YOLOv8 (mAP@0.5: +1.08%, mAP@0.5:@0.95: +0.94%). The Focal-ESIoU loss is for relatively small The detection performance of the target has been improved to a certain extent, further proving that the ESIoU loss can not only accelerate the convergence of the model, but also improve the model's ability to regress and locate the bounding box, thereby enhancing the overall performance of the target detector. At the same time, it also shows that the convergence speed of the model has been accelerated (Note: After adding CGM, the number of parameters and the amount of computation increase, but the convergence speed is faster).

[0022] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the application shall be included within the protection scope of the present application.

Claims

1. A traffic target flow statistics algorithm based on a content guidance module, characterized by: The following steps are involved: S1. Preparatory work: Train the target detection network and training environment and parameters including the content-guided module, use the cosine learning rate and ReLU activation function, and construct the total loss function with weight combination, where the BBR loss uses the ESIoU loss, the classification loss uses the VFL loss, and the regression loss uses the DFL loss. After the training is completed, the obtained Pytorch model is converted to the RKNN model, and the RKNN model is quantized; S2, data acquisition and encoding: use the camera to collect traffic video, transmit the video stream to the embedded platform through the RTSP protocol, and use the hardware decoding function or related software library to decode the video stream to obtain frame-by-frame image data; S3, image preprocessing: cropping the collected image to a fixed size and setting the counting and statistical area; S4, model reasoning: Use the optimized FPN structure for feature fusion to improve target detection accuracy and robustness, and use the embedded platform's NPU for multi-threaded reasoning to obtain bounding box and target category label results; S5, post-processing: set a threshold based on the confidence score to filter low-confidence targets, perform non-maximum suppression, and filter out the bounding box that is most likely to be the target by calculating the ESIoU value between the bounding boxes; S6, multi-target tracking: input the detection results into the multi-target tracking algorithm to obtain the tracking results, and perform traffic statistics based on the number of tracked IDs; S7, video streaming and result display: re-encode the image frames containing the detection results into a video stream, push it to the receiving end through the RTSP protocol, and use a web player to display the detection results in real time.

2. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In the step S1, the content guidance module uses the main feature map and the auxiliary feature map as CGM input, and generates a global feature vector after global average pooling after Concat fusion. A fully connected convolution reduces the number of channels from 2C to 2C / r, where r is the reduction factor. The second fully connected convolution restores the number of channels to 2C through 1×1 convolution to generate the fused attention weights, and then performs Sigmoid activation to normalize the weights and limit them to the [0, 1] interval for channel weight adjustment. The results are first weighted adjusted with the original input, and then cross-fused to extract deeper information. Finally, Concat fusion is performed to obtain the output Y.

3. The traffic target flow statistics algorithm based on the content guidance module according to claim 2 is characterized by: The CGM input calculation is as follows: in, is the main feature map, is the auxiliary feature map, is the weighted adjusted global eigenvector.

4. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In step S1, some parameters in the training phase are set as follows: epoch=400. After 400 iterations, the model with the best effect in all iterations is selected as the model training result. Batch size = 16; SGD optimizer is used; Weight decay coefficient = 0.0005; Momentum hyperparameter = 0.937; Use cosine learning rate to optimize the results. The specific formula is: in, is the learning rate at the tth iteration, is the minimum learning rate, is the maximum learning rate, is the current iteration number, is the maximum number of iterations.

5. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: The total loss function is constructed to accelerate the adjustment and convergence of the prediction frame through the penalty term of the width and height difference between the two rectangular frames. The specific formula is: Among them, w and h represent the width and height of the predicted bounding box respectively. Represent the width and height of the real bounding box respectively, , Respectively represent the width and height of the minimum closure, v refers to the formula, which is used to measure the difference in aspect ratio, indicating Center point distance penalty, represents the aspect ratio penalty, Represents the width-height difference penalty term.

6. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In step S5, all bounding boxes are sorted according to their confidences, starting from the bounding box with the highest confidence, and the ESIoU between it and other bounding boxes is calculated. The specific calculation method is: in, represents the value of the loss function, represents the intersection-over-union ratio between the predicted bounding box and the true bounding box, The penalty term in the loss function.

7. The traffic target flow statistics algorithm based on the content guidance module according to claim 6 is characterized by: In step S5, if the ESIoU of the bounding box with the highest confidence and the other bounding boxes is greater than the NMS threshold, the bounding box is deleted, and the remaining bounding boxes continue to be processed according to the above steps, and the retained target detection results are output to step S6.

8. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In the step S2, the RTSP stream is transmitted to the embedded platform via the TCP protocol, and the hardware decoding function and the related software library are FFmpeg.

9. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In the step S4, the optimized FPN structure includes a content guidance module, and the content guidance module is used to perform feature fusion at the connection of the FPN.

10. The traffic target flow statistics algorithm based on the content guidance module according to claim 1 is characterized by: In step S7, the video streaming and result display are performed by re-encoding the image frames containing the detection results into a video stream and pushing it to the designated receiving end using the RTSP protocol.