Dense pedestrian detection method, system and device for traffic scene

Through the improved RT-DETR network model, combined with feature fusion strategy and loss function, the problem of insufficient accuracy of dense pedestrian detection in traffic scenarios is solved, and efficient pedestrian detection in complex environments is achieved.

CN120259977APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510408392.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In traffic scenarios, intensive pedestrian detection tasks face problems such as large number of targets, large scale changes, and diverse pedestrian movement status. The existing detection methods lack detection accuracy in complex environments, especially in dense occlusion, which is difficult to effectively distinguish individuals, affecting detection performance.

Method used

The improved RT-DETR network model is adopted, combined with FasterNet Block, AIFI-HiLo module and BiFPN feature fusion strategy, and optimize the accuracy and robustness of object detection through bottom-up and top-down feature transmission, combined with MPDIoU loss function.

Benefits of technology

It improves the accuracy and robustness of pedestrian detection in complex backgrounds and dense occlusion environments, enhances the model's adaptability to targets at different scales, and improves the real-time and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259977A_ABST
    Figure CN120259977A_ABST
Patent Text Reader

Abstract

The invention relates to a dense pedestrian detection method for a traffic scene, and belongs to the technical field of machine vision, and the method comprises the following steps: S1, collecting pedestrian images of the traffic scene, and constructing a training and testing data set; s2, constructing a pedestrian detection network, fusing high and low frequency attention and a bidirectional feature fusion module, and introducing a target frame matching strategy; s3, preprocessing the image, and inputting the preprocessed image into a detection network training model; and S4, deploying the trained model to a traffic monitoring system to realize pedestrian detection. The method improves the detection precision and the calculation efficiency, still has the precise detection capability in a complex background and a dense shielding environment, and can be applied to an intelligent traffic monitoring and automatic driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and pedestrian detection, and relates to a pedestrian detection method, system and device for traffic scenarios. Background Art

[0002] In traffic scenarios, pedestrian detection is an important task in the fields of intelligent transportation, autonomous driving, and security monitoring, which can improve road safety and driving efficiency. However, due to the complex traffic environment, pedestrians are densely distributed and there are occlusions, making the detection task face great challenges. Especially in areas such as sidewalks, zebra crossings, and bus stops, pedestrians often appear in groups and are severely occluded from each other, affecting the accuracy of detection. In addition, the background in traffic scenarios is complex, and objects such as vehicles, billboards, and street lights may be similar to pedestrian features, resulting in false detections by detection algorithms. At the same time, affected by factors such as illumination changes and weather conditions (such as rainy and foggy days), the robustness of existing detection methods still needs to be improved.

[0003] In recent years, the rapid development of deep learning technology has greatly promoted the progress of pedestrian detection. Detection models based on convolutional neural networks and Transformers have improved the detection accuracy to a certain extent. Currently, mainstream detection methods can be divided into two-stage detection and one-stage detection. Two-stage detection relies on the region proposal mechanism, which helps to improve the detection accuracy, but has a high computational complexity and is difficult to meet the requirements of real-time applications. In contrast, one-stage detection adopts an end-to-end strategy and has a higher detection speed, but is prone to false detections and missed detections in dense pedestrian scenarios. In addition, although some lightweight models can adapt to resource-constrained devices, their detection accuracy in complex traffic environments is still insufficient and difficult to meet the actual application requirements.

[0004] The dense pedestrian detection task in traffic scenarios faces problems such as a large number of targets, large scale variations, and diverse pedestrian motion states, which further increases the detection difficulty. Although the Transformer structure has strong global modeling capabilities, there are still problems of feature aliasing and localization deviation in high-density scenarios, affecting the detection effect. At the same time, traditional bounding box detection methods are difficult to effectively distinguish individuals in highly occluded situations, affecting the final detection performance. Therefore, there is still room for optimization in dense pedestrian detection for traffic scenarios, and it is necessary to further improve the detection model to enhance its adaptability and detection accuracy in complex environments. Summary of the Invention

[0005] Based on this, the present invention provides a dense pedestrian detection method, system and device for traffic scenarios, which can improve the detection accuracy in the case of dense pedestrians, occlusions and complex backgrounds, while enhancing the model's adaptability to different scale targets and improving the robustness and real-time performance of detection.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] S1: Collect pedestrian images in the traffic scene and construct a pedestrian detection data set covering various environmental factors;

[0008] S2: Construct an improved RT-DETR network model for pedestrian detection;

[0009] S3: Input all images of the pedestrian data set at the traffic intersection into the constructed improved RT-DETR network for training.

[0010] S4: Use the trained pedestrian detection model for pedestrian detection in the actual traffic scene by means of the RT-DETR network.

[0011] Further, in step S1, first install the devices required for collecting traffic scene images, including monitoring devices and terminal computers, to ensure that the terminal computer can receive the traffic scene images transmitted by the monitoring devices; then annotate and screen the collected images.

[0012] The monitoring device is used to capture traffic scene images in real time and transmit pedestrian data at the traffic intersection to the terminal computer; the terminal computer is used to receive the traffic scene images transmitted by the camera and input them into the improved RT-DETR network for detection, so as to obtain the pedestrian distribution and status information in the traffic scene.

[0013] Further, the annotation and screening in step S1 include:

[0014] S11: Select the annotation tool LabelImg to annotate the images, select the pedestrian targets and annotate the categories to construct a complete information data set, and at the same time check and screen the annotation results to extract a high-quality training data set that meets the requirements of the dense occluded pedestrian detection task;

[0015] Step S2 specifically includes the following steps:

[0016] S21: Modify according to the FasterNet Block module in the FasterNet network structure and refer to reparameterization. The reparameterization consists of partial convolution, channel adjustment and residual connection layers to construct a new backbone network for extracting image features;

[0017] S22: Construct an AIFI-HiLo feature fusion module. The AIFI-HiLo module contains two branches: a low-frequency branch (Lo-Fi) and a high-frequency branch (Hi-Fi). The low-frequency branch is mainly responsible for capturing global information, while the high-frequency branch focuses on detailed information;

[0018] S23: In RT-DETR, a second feature fusion network, BiFPN, is constructed to fuse the S2, S3, S4, and S5 feature layers. It adopts bidirectional feature transmission from bottom to top and from top to bottom, and combines a weighted feature fusion strategy to adaptively adjust the contributions of features at different scales.

[0019] S24: A target box matching strategy is constructed, and the MPDIoU loss function is introduced to optimize the accuracy of object detection.

[0020] Furthermore, the backbone network constructed in step S21 includes the following steps:

[0021] S211: After the input image undergoes Embedding processing, initial feature extraction is performed through the Faster Rep Block to generate a feature map with a resolution of 1 / 4 and output the S2 feature layer. At this stage, the Faster Rep Block effectively fuses local spatial information with global features, and processes the input features through partial convolution, channel adjustment layers, and residual connection layers. Subsequently, the Merging layer performs a downsampling operation to further reduce the resolution of the feature map and provide input for subsequent deep feature extraction.

[0022] S212: Taking the S2 feature layer as the input, the Faster Rep Block is continued to perform deep feature extraction to generate a feature map with a resolution of 1 / 8 and output the S3 feature layer. At this stage, through deeper feature extraction, the network can capture richer semantic information and detailed features, further enhancing the ability to recognize targets. After that, downsampling is performed through the Merging layer to reduce the size of the feature map.

[0023] S213: Taking the S3 feature layer as the input, advanced features are further extracted through the Faster Rep Block to generate a feature map with a resolution of 1 / 16 and output the S4 feature layer. At this time, the network can extract more abstract high-level semantic information, especially in the case of complex backgrounds and dense occlusions, which helps to perform object detection more accurately. The subsequent Merging layer downsamples the feature map again.

[0024] S214: Taking the S4 feature layer as the input, the Faster Rep Block is used to perform the last layer of high-level semantic feature extraction to generate a feature map with a resolution of 1 / 32 and output the S5 feature layer. At this stage, the network can extract the most refined high-level semantic features, providing a more detailed feature representation for subsequent detection tasks.

[0025] Furthermore, the extracted features are processed through global pooling, point convolution, and fully connected layers and used as the input for feature fusion, providing feature input for subsequent BiFPN feature fusion.

[0026] Furthermore, the AIFI-HiLo module constructed in step S22 includes the following steps:

[0027] S221: The low-frequency branch extracts global information from the S5 feature, obtains global features using low-frequency convolutional kernels, adjusts the feature distribution through average pooling, and then calculates key region information using the global self-attention mechanism;

[0028] S222: The high-frequency branch extracts local detail information from the S5 feature, obtains local features using high-frequency convolutional kernels, and adjusts the feature distribution through normalization operations to extract key region information;

[0029] S223: The features of the high-frequency branch and the low-frequency branch are fused through residual connections to maintain the integrity of the feature information and avoid information loss. The fused features are used for subsequent object classification and detection tasks;

[0030] S23: Further fuse the features output by S2, S3, S4, and S5;

[0031] Furthermore, the BiFPN constructed in step S23 includes the following steps:

[0032] S231: During the multi-scale feature fusion process, first preprocess the features of S2, S3, S4, and S5. Each feature map undergoes a convolutional operation to ensure that its size and number of channels are consistent, facilitating subsequent feature fusion;

[0033] S232: Construct a BiFPN structure and adopt bottom-up feature transfer. By transferring low-resolution features to high-resolution layers, local information is strengthened, and at the same time, features are weighted in each layer to adjust the contribution of each scale feature to the final fusion result;

[0034] S233: In the top-down transfer, transfer high-resolution features to low-resolution layers to ensure that different scale features can fully interact, ensuring the integrity and diversity of the features. Through a weighted fusion strategy, combine the feature information of each scale to generate the final multi-scale fusion feature map;

[0035] Furthermore, the fused multi-scale features are input into the Decoder module of RT-DETR for object query and set prediction to complete the final pedestrian detection task;

[0036] Furthermore, the object box matching constructed in step S24 includes the following steps:

[0037] S241: In the Decoder module, the fused multi-scale features are used as input, and object localization and classification are performed by means of querying and aggregation. The model generates candidate boxes through object queries and makes predictions in combination with multi-scale features;

[0038] S242: The MPDIoU loss function is used to optimize the overlap between the predicted boxes and the ground truth boxes, which combines the traditional IoU and the geometric features of the boxes, and takes into account the distance information between the predicted boxes and the ground truth boxes;

[0039] S243: During the training process, the MPDIoU loss is jointly optimized with the conventional classification loss and regression loss, and various parameters of the model are optimized through backpropagation;

[0040] Step S3 specifically includes the following steps:

[0041] S31: Statistically analyze the traffic scene images and their annotation situations in the dataset, filter out objects with abnormally small or large annotation box sizes, and check for possible annotation errors to ensure data quality;

[0042] S32: Batch-normalize the annotation categories and uniformly modify them to "person" and "head" to improve annotation consistency.

[0043] S33: Normalize the dataset images to meet the input requirements of the RT-DETR network and input them into the network for feature extraction and fusion;

[0044] S34: Input the features extracted by the improved RT-DETR into the Decoder module for object set prediction to improve detection accuracy;

[0045] S35: In the Decoder module, output the final predicted result map, use logistic regression to filter and score the prior boxes, and annotate the final predicted results to complete the training of the pedestrian detection model;

[0046] On the other hand, the present invention provides a pedestrian detection device for an improved RT-DETR, including:

[0047] Input module: used to receive the images collected by the monitoring device and input them into the pre-trained improved RT-DETR pedestrian detection model;

[0048] Feature extraction module: The improved RT-DETR network extracts the features of the whole pedestrian and the head;

[0049] Feature fusion module: Adopt a multi-level feature fusion strategy to integrate pedestrian features of different scales;

[0050] Classification prediction module: Input the extracted overall pedestrian and head features into the Decoder module, generate prior boxes based on the set prediction method, and output the target category and scoring information, finally completing the pedestrian detection task. Brief Description of the Drawings

[0051] Figure 1 is the flowchart of the pedestrian detection method;

[0052] Figure 2 is the structural diagram of the improved RT-DETR overall network;

[0053] Figure 3 is the structural diagram of the FasterNet Rep backbone network;

[0054] Figure 4 is the structural diagram of the BiFPN module; Detailed Embodiments

[0055] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0056] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0057] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0058] In view of the background technology, in order to improve the accuracy of pedestrian detection, the present invention, asFigure 1 As shown in the figure, the present invention provides a pedestrian detection method for an improved RT-DETR network, including the following steps:

[0059] S1: Collect pedestrian images of traffic scenes and construct training and test data sets;

[0060] S2: Construct an improved RT-DETR network for pedestrian detection;

[0061] S3: Input the images of the traffic scenes into the improved RT-DETR network for training;

[0062] S4: Use the improved RT-DETR pedestrian detection model to detect pedestrians in actual traffic scenes.

[0063] In this embodiment, in order to better evaluate the detection effect of the improved RT-DETR pedestrian detection method of the present invention in traffic scenes, the publicly available dataset CrowdHuman is adopted. Since the CrowdHuman dataset already contains detailed annotation information, the manual annotation step is omitted.

[0064] In the above S1, constructing the test data set includes the following steps:

[0065] S11: Convert the original odgt format label file in the CrowdHuman dataset into the txt format required by RT-DETR to meet the input format requirements of model training.

[0066] S12: Adjust the label categories in the data set to person (pedestrian) and head (head);

[0067] S13: The training set contains 15,000 images, the validation set contains 4,370 images, and the test set contains 5,000 images.

[0068] Furthermore, in order to maintain the consistency of the training and inference processes, in this embodiment, the resolution of all input images is uniformly set to 640×640.

[0069] As Figure 3 shown, in the above S2, it includes the following steps:

[0070] S21: Input into the backbone network to extract features. The input image of 640×640×3 first passes through the Embedding layer, which contains three 3×3 convolutions and one 3×3 max pooling for preliminary feature extraction and dimensionality reduction. The first convolution uses a 3×3 convolution kernel with a stride of 2 to expand the number of channels to 32 and simultaneously reduce the size to 320×320×32. Subsequently, the two 3×3 convolutions keep the size unchanged, with the number of channels being 32 and 64 respectively, and finally output a feature map of 320×320×64. The max pooling layer uses a 3×3 pooling kernel with a stride of 2 and a padding of 1 to further reduce the size of the feature map to 160×160×64, providing input for the subsequent feature extraction of the backbone network and constructing a new backbone network for extracting image features.

[0071] S21: Modify according to the main module FasterNet Block in the FasterNet network structure, and refer to and combine the reparameterization idea to construct a new backbone network for extracting image features;

[0072] As Figure 3 shown, the main steps for constructing the backbone network in the above S21 are as follows:

[0073] S211: The size of the input feature map is 160×160×64. First, it passes through two Faster Rep Blocks for feature extraction. Each Block adopts the RepVGG structure, including 3×3 convolution, 1×1 convolution, and Identity branch. At this stage, the number of channels remains 64 and the size of the feature map remains unchanged; further, in the output feature map of each stage, 1×1 convolution is used to expand the number of channels, the Merging layer, and max pooling operations are used to reduce the spatial resolution and perform downsampling of the feature map;

[0074] S212: The size of the feature map output in the S2 stage is 160×160×64. First, it passes through a convolution with a stride of 2 to reduce the size of the feature map to 80×80 to reduce the computational amount and increase the receptive field. Subsequently, it passes through two Faster Rep Block modules, and the number of channels increases to 128. This stage is mainly used for extracting intermediate features to improve the model's recognition ability of pedestrian contours.

[0075] S213: The size of the feature map output in the S3 stage is 80×80×128. It passes through a convolution with a stride of 2 to further reduce it to 40×40 to enhance the model's perception ability of global information. Subsequently, it passes through two Faster Rep Blocks, and the number of channels increases to 256, enabling the model to learn richer semantic information.

[0076] S214: The size of the feature map output in stage S4 is 40×40×256. First, it goes through a convolution with a stride of 2, and the size shrinks to 20×20. At this time, the receptive field of the model is further expanded, and the number of channels increases to 512. This stage is mainly used to extract high-level features.

[0077] Furthermore, after the output of the fourth stage, global average pooling operation is performed, and global information is aggregated through a 1×1 point convolution layer and a fully connected layer, finally generating the output of the backbone network. This output represents the global features after multi-stage extraction.

[0078] The backbone network is as Figure 3 shown. The overall architecture of the FasterNet Rep backbone network is the same as that of the original backbone network ResNet18 of RT-DETR, and there are outputs in the last three stages. These three stages respectively contain features of different depths, which will be passed as inputs to the subsequent feature fusion structure for fusing features of different scales.

[0079] S22: In the improved RT-DETR, AIFI-HiLo is introduced as the first feature fusion module to specifically process the high-level features in stage S5. The feature map output in S5 first goes through a 1×1 convolution to reduce the number of channels from 2048 to 256, and then is input into the AIFI-HiLo module. The high-frequency branch extracts local details, the low-frequency branch extracts global information, and they are fused in the channel dimension. The processed features then go through a 1×1 convolution to align the number of channels, and are jointly input into BiFPN together with the features of S2, S3, and S4 to achieve multi-level feature fusion.

[0080] The AIFI feature fusion module constructed in S22 specifically includes the following steps:

[0081] S221: The size of the original feature map in stage S5 is 10×10×2048. First, it goes through a 1×1 convolution layer to reduce the number of channels from 2048 to 256 to reduce the computational load and adapt to the input requirements of the subsequent module. This operation is also equipped with batch normalization and ReLU activation function to enhance the stability and non-linear expression ability of the features. The processed feature map remains the same size and enters the AIFI-HiLo module for high-low frequency feature processing;

[0082] S222: After entering the AIFI-HiLo module, the features are first divided into a high-frequency branch and a low-frequency branch. The high-frequency branch directly acts on the original features using a 3×3 convolution to capture local detail information; the low-frequency branch reduces the resolution through a 4×4 convolution kernel and a mean pooling operation with a stride of 2 to extract global structural information. Subsequently, self-attention calculations are performed on the high-frequency and low-frequency features respectively. The high-frequency part performs attention operations within a local window, while the low-frequency part uses a linear transformation to construct a cross-region attention mechanism to enhance the perception ability of large-scale targets. Finally, the high- and low-frequency features are fused in the channel dimension to form a multi-scale feature map with stronger expressive ability;

[0083] S223: The feature map after being processed by AIFI-HiLo still maintains a size of 10×10×256. To adapt to the subsequent BiFPN module, channel alignment needs to be further performed. First, a 1×1 convolutional layer is used to keep the number of channels stable and ensure consistency with the feature dimensions in other stages (S2, S3, S4). Subsequently, this feature map and the features of S2, S3, and S4 are jointly input into BiFPN for multi-scale feature fusion;

[0084] S23: In RT-DETR, a second feature fusion network BiFPN is constructed, and an adaptive strategy is adopted for feature fusion. BiFPN consists of two parts: bottom-up and top-down convolutions.

[0085] As Figure 3 shown, the main steps for constructing BiFPN in S23 are as follows:

[0086] S231: In the bottom-up path, the feature map is first processed through five convolutional layers, which contain 1×1 and 3×3 convolutional kernels, and two upsampling layers are used to adjust the resolution of the feature map. The convolutional operations are used to enhance the expressive ability of the feature map, and the upsampling layers help improve the spatial resolution of the feature map to retain more detail information.

[0087] S232: In the top-down path, the feature map is processed through two convolutional layers, and each convolutional layer contains five convolutional operations, which are composed of 1×1 and 3×3 convolutional kernels. In this path, the feature map will also pass through two downsampling layers, which perform downsampling through 3×3 convolutional kernels to reduce the spatial resolution and highlight the more important global information.

[0088] S233: All the processed feature maps will be fused through an adaptive fusion module. In this process, first, the number of channels of the feature maps at different scales is adjusted, and then these feature maps are fused through an adaptive weighting mechanism. The adaptive weighting mechanism dynamically adjusts the contribution ratio of each feature map according to the importance of different feature maps to achieve more precise feature fusion.

[0089] Further, after passing through the Decoder, the improved RT-DETR network for pedestrian detection can be constructed.

[0090] S24: Through the feature fusion network of RT-DETR, three output feature layers are generated. Each feature contains four-dimensional information: batch N, feature grid, and the position information of the target (x, y offsets, width, height), confidence, and classification result. These features are input into the Decoder decoder module to complete the classification and detection of the target.

[0091] Further, the target box matching constructed in step S24 includes the following steps:

[0092] S241: In the Decoder module, the fused multi-scale features are used as input. Specifically, the three feature layers come from the feature grids of P5 / 32 (20×20), P4 / 16 (40×40), and P3 / 8 (80×80) respectively. Candidate boxes are generated through target queries, and the target is located and classified by combining these three layers of multi-scale features.

[0093] S242: The MPDIoU loss function is used to optimize the overlap between the predicted box and the ground truth box. This loss function combines the traditional IoU and the geometric features of the box, considering the distance information between the predicted box and the ground truth box to improve the accuracy of box matching.

[0094] S243: During the training process, the MPDIoU loss is jointly optimized with the conventional classification loss and regression loss. The parameters of the model are optimized through backpropagation to complete the construction of the pedestrian detection network.

[0095] Further, after passing through the Decoder, the pedestrian detection network can be constructed.

[0096] In S3, the following steps are included:

[0097] S31: Use the public dataset for training and the self-annotated dataset for testing. Statistically screen the images and annotations in the dataset for the construction site, filter out extremely small or large annotation boxes, check the existing annotation errors and modify and improve them;

[0098] S32: To ensure the consistency of testing, all annotation categories in the public dataset are uniformly modified to "person" and "head".

[0099] S33: Set the parameters required for training as follows: the number of training epochs is epoch = 200, the input batch size each time is batch_size = 4, the size of the training images is train_img_size = 640, the momentum is set to momentum = 0.9, the IoU threshold required for the loss function is Iou_threshold_loss = 0.5, and the weight decay rate is set to weight_decay = 0.0005. The learning rate is set to lr = 0.0001, and the number of threads is workers = 4,

[0100] Further, set the parameters required for testing: the size of the training images: train_img_size = 640, the test input batch size: batch_size = 4, the confidence threshold: conf_thresh = 0.005, the non-maximum suppression threshold: nms_thresh = 0.5, multi-scale validation: multi_scale_val = True;

[0101] Further, it also includes the settings of classes and anchors. The class settings are {"num": 2, "classes": ['person', 'head']}, the strides are: strides: [8, 16, 32], and the output scale is: anchors_pre_scale = 3,

[0102] S34: Input the data image normalized to the scale of 640×640×3 into a 3×3 convolutional layer. After passing through this convolutional layer, the image size becomes 320×320, and the number of channels is converted from 3 to 32. Then the image enters the improved RT-DETR network constructed by S2 for feature extraction and fusion.

[0103] After passing through the improved RT-DETR backbone network, three output feature maps are obtained. Each feature map contains four-dimensional information: the first dimension is the number of input images N, the second and third dimensions are grids of 20×20, 40×40, and 80×80 respectively, and the last dimension includes: the offsets of x and y, width and height, confidence, and classification results. Such an output can represent the positions of the prior boxes, and these features are input into the Decoder module for subsequent classification and prediction.

[0104] S35: Input the three scales obtained into the Decoder for decoding, and calculate the predicted bounding boxes, confidence values, and classes;

[0105] During the training process, the result of each forward propagation is evaluated by calculating the loss function, and the backpropagation and Adam optimization algorithm are used to minimize the loss function. After each epoch of training, the detection accuracy metric is calculated through evaluation on the validation set. At each cycle of training, the weights of the current model are saved. Finally, the model with the highest mAP accuracy on the validation set is used as the final detection model.

[0106] Further, when defining the loss function, MPDIoU is used to calculate the bounding box loss, and cross-entropy is used to calculate the loss of the object and the category. Among them, the input parameter information is the calculated bounding box and the dataset annotation box. After the calculation, the total loss value, the bounding box loss, the confidence loss, and the category loss value are obtained, and the model parameters are updated through backpropagation;

[0107] In S4, using the improved RT-DETR network, after multiple trainings, the pedestrian detection model with the highest accuracy is selected for pedestrian detection in the actual traffic scenario;

[0108] Further, in actual detection, information such as the detection model, the log path, the gpu / cpu setting, and the image path to be detected needs to be loaded, and multiple bounding boxes are predicted through the model;

[0109] Further, filtering is performed through set prediction. Among them, the bounding boxes need to be input, and the predefined bounding box threshold is used to obtain the final bounding boxes. The filtered bounding boxes and the relevant category information are directly drawn on the image to obtain the result;

[0110] The pedestrian detection device of the present invention includes:

[0111] An input module for inputting the image captured by the monitoring device into the pedestrian detection model of the improved RT-DETR that has been pre-trained;

[0112] A feature extraction module for extracting the human head and full-body features in the image using the improved RT-DETR network;

[0113] A classification prediction module for inputting the extracted human head and full-body features into the Decoder module, and performing set prediction through this module to obtain the predicted prior boxes, scores, and category information.

[0114] The present invention provides a pedestrian detection device for use in S1, including: a monitoring device, a computer.

[0115] The monitoring device can capture traffic scene images in real time and transmit the traffic scene images to the computer for facilitating the detection of pedestrian situations.

[0116] The computer can receive traffic scene images transmitted by monitoring devices, input the computer program proposed by the present invention, load a detection model to perform pedestrian detection, and then output detection results, such as the position and category information of pedestrians, etc.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A dense pedestrian detection method, system and device for traffic scenarios, comprising the following steps: S1: Collect pedestrian images in traffic scenarios and construct a pedestrian detection dataset; S2: Construct an improved RT-DETR pedestrian detection network for pedestrian detection; S3: Input the image dataset into the constructed improved RT-DETR pedestrian detection network for training; S4: Use the RT-DETR network to apply the trained pedestrian detection model to pedestrian detection in actual traffic scenarios. Further, in step S1, it specifically includes: setting the installation location and hardware conditions of the acquisition device, for example, camera performance requirements, installation location and shooting angle, etc., to meet the requirements of pedestrian detection; Step S2 includes the following steps: S21: Construct a FasterNet Rep backbone network, use Faster Rep Block for feature extraction, and combine partial convolution, channel adjustment and residual connection to improve feature extraction efficiency and information transfer ability; S22: Construct an AIFI feature fusion module on the S5 feature layer; S23: Construct a BiFPN to fuse multi-scale features of S2, S3, S4, and S5, and use an adaptive feature fusion strategy for feature adjustment; S24: The fused feature information is input into the Decoder of RT-DETR for object query and set prediction to complete the final pedestrian detection task. During this process, an MPDIoU loss function is introduced to optimize the matching accuracy and regression effect of the object box, which is specifically divided into the following steps; The backbone network constructed in step S21 includes the following steps: S211: Extract the core structure of the Faster Rep Block. The module consists of a partial convolution layer, a channel adjustment layer and a residual connection layer. Through the channel division strategy, some channels perform 3×3 convolution to extract local information, and other channels maintain the original features to reduce the computational amount; S212: Design a channel adjustment layer to integrate features using point convolution and non-linear activation functions. Point convolution adjusts the channel dimension, enhances the feature expression ability, and at the same time enhances the information interaction between channels, improving the network's adaptability to different scale targets. S213: Construct a BasicBlock_Faster_Block_Rep. The main branch consists of consecutive 3×3 convolution layers to extract deep features. Residual connection is used for feature fusion. Layers without downsampling are directly skipped and connected, and downsampling layers adjust the number of channels through pooling and point convolution to make the residuals match the main branch; S214: Stack BasicBlock_Faster_Block_Rep to construct the FasterNet Rep backbone network. At different network stages, the spatial resolution of the feature map gradually decreases, and the number of channels gradually increases to obtain rich semantic information. Finally, through global pooling, point convolution and fully connected layer processing, high-level features are formed to provide input for subsequent BiFPN adaptive fusion; Further, the extracted features are processed through global pooling, point convolution, and fully connected layers and used as the input for feature fusion, providing the input for the subsequent multi-scale BiFPN adaptive feature fusion. The AIFI-HiLo feature fusion module constructed in step S22 includes the following steps: S221: Before the Decoder module, the AIFI-HiLo module is used to adaptively fuse the input features of S5. The low-frequency branch extracts global information through global pooling and performs channel transformation, while the high-frequency branch extracts local detail information using local convolution operations. S222: Design an AIFI-HiLo feature processing structure, including a HiLo module and a feed-forward neural network. After receiving the input features, the HiLo module processes the high-frequency and low-frequency information respectively, and fuses the extracted global and local features through residual connections. The feed-forward network consists of two layers of point convolution. The first layer is used for dimensionality increase to enhance the feature expression ability, and the second layer is used for dimensionality reduction to make the features return to the original channel dimension. S223: Add a normalization layer and a Dropout mechanism to the AIFI-HiLo feature processing structure. The normalization layer is used to keep the feature distribution stable, and Dropout prevents overfitting. Adjust the attention mechanism to perform high-low frequency feature interaction. Finally, the fused features are adjusted through an activation function to adapt to subsequent detection tasks. The BiFPN feature fusion module constructed in step S23 includes the following steps: S231: Unify the channel dimension and feature input processing, construct a feature fusion module, receive the inputs from the feature layers of S2, S3, S4, and S5 scales, and perform channel transformation through point convolution to map the features of different scales to the same channel dimension. Adopt an adaptive fusion strategy to dynamically weight the multi-scale features and generate a fused feature representation. S232: Calculate the fusion weights and perform adaptive adjustment. During the adaptive fusion process, first concatenate the input features in the channel dimension and generate fusion weights through convolution operations. Subsequently, use softmax normalization to dynamically adjust the fusion weights of different scale features and adaptively allocate the contributions of each scale of information according to different scenario requirements. S233: Perform bidirectional feature flow and optimize the fusion strategy. Apply the fusion module to the BiFPN feature layer, and use the adaptive fusion mechanism to adjust the contributions of each scale of features during the bidirectional feature flow process. Strengthen the information flow through cross-layer connections, so that the fused features maintain the balance between global information and local details in a densely occluded scenario and improve the object detection ability. The target box matching strategy constructed in step S24 includes the following steps: S241: Calculate the basic bounding box loss. During the target box matching process, use the L1 loss to calculate the coordinate error between the predicted box and the ground truth box to ensure that the target box can gradually approach the ground truth box and improve the detection accuracy. S242: Calculate the MPDIoU loss. By constraining the matching process of the target bounding boxes through the Euclidean distance, make the target bounding boxes fit the real targets better and improve the detection accuracy. To further enhance the regression ability of low-IoU samples, combine Inner-IoU as a scale factor to dynamically adjust the loss calculation method, thereby improving the detection effect in occluded and dense scenes; S243: Optimize the loss weighting strategy. According to different stages of the target bounding box matching, adjust the loss weighting ratio to balance the learning processes of high-IoU and low-IoU samples, and reasonably allocate the weights of MPDIoU and other IoU variants, so that the model converges more stably in complex environments; The step S3 includes the following steps: S31: Conduct statistical analysis on the traffic scene images and their annotation situations in the dataset, screen out targets with abnormal annotation bounding box sizes, such as being too small or too large, and verify possible annotation errors to ensure the data quality; S32: Batch-standardize the annotation categories and uniformly modify them to "person" and "head"; S33: Normalize the dataset images to meet the input requirements of the RT-DETR network, and input them into the network for feature extraction and fusion; S34: Input the features extracted by the improved RT-DETR into the Decoder module for target set prediction to improve the detection accuracy; S35: In the Decoder module, predict and output the final result map, and use logistic regression to screen and score the prior bounding boxes, and annotate the final prediction results to complete the training of the pedestrian detection model.

2. The improved RT-DETR-based traffic scene pedestrian detection method according to claim 1, characterized in that: In step S1, first install the devices required to collect traffic scene images, including monitoring devices and terminal computers, and ensure that the terminal computer can receive the traffic scene images transmitted by the monitoring devices; then annotate and screen the collected images.

3. The improved RT-DETR-based traffic scene pedestrian detection method according to claim 2, characterized in that: The monitoring device is used to capture traffic scene images in real time and transmit the traffic scene images to the terminal computer; the terminal computer is used to receive the traffic scene images transmitted by the camera, input them into the improved RT-DETR pedestrian detection model for detection, so as to obtain the pedestrian detection results in the traffic scene.

4. The improved RT-DETR-based traffic scene pedestrian detection method according to claim 3, characterized in that: What is included in step S1 is as follows: S11: Select the annotation tool LabelImg to annotate the images, select pedestrian targets and annotate the categories to construct a complete information dataset, and at the same time check and screen the annotation results to extract a high-quality training dataset that meets the requirements of the dense occluded pedestrian detection task.

5. An improved RT-DETR pedestrian detection device for traffic scenes based on the method according to any one of claims 1-4, characterized in that: An input module, which is used to input the images captured by the monitoring device into the pedestrian detection model based on the RT-DETR network that has been pre-trained; A feature extraction module, which is used to extract the human head and full-body features by using the pedestrian detection model of the improved RT-DETR network; A classification and prediction module, which is used to input the human head and full-body features extracted from the pedestrian detection model of the improved RT-DETR network into the Decoder module, and obtain the predicted prior bounding boxes, scores and category information through set prediction.

Citation Information

Cited By

  • Multi-object tracking method based on global-local feature joint modeling

    CN121213616A

  • Communication big data-driven sewage plant peak clipping and valley filling low-carbon scheduling method

    CN121504031A

  • A communication big data driven peak load shifting and low carbon scheduling method for sewage treatment plants

    CN121504031B

  • Dense pedestrian detection method and system based on improved D-FINE-N network

    CN122493497A