Real-time human-vehicle detection method for aided driving
By introducing CPPA, DWRC and CG block modules into the human-vehicle detection network, the backbone network structure is improved, and the accuracy and real-time problems in multi-scale object detection and extreme environments in complex traffic scenarios are solved, and efficient human-vehicle detection results are achieved.
Patent Information
- Application Number
- CN202510498826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In complex traffic scenarios, existing human-vehicle detection algorithms are difficult to maintain high accuracy and real-time in multi-scale object detection and extreme environments, especially when dealing with extreme weather and complex lighting conditions.
An efficient detection network CDC-DETR is proposed. By introducing context pre-activated pooling attention mechanism (CPPA), expansion residual connection (DWRC) and context guidance module (CG block), the structure of backbone network and encoder is improved to improve multi-scale feature capture capability and detection accuracy.
It significantly improves the detection accuracy of small targets and the real-timeness of the model, enhances the adaptability to complex traffic scenarios and extreme environments, and is suitable for environmental perception tasks in autonomous driving systems.
Smart Images

Figure CN120032345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and target detection, and in particular to a real-time human-vehicle detection method for assisted driving. Background Art
[0002] In the context of the whole society attaching great importance to traffic safety, reducing the incidence of traffic accidents has become a key problem that needs to be solved urgently. The emergence of intelligent transportation systems and autonomous driving technology has opened up a new path to solve this problem. Among them, computer vision technology plays an extremely critical role, and human and vehicle detection technology is the core element to ensure road safety. In the field of autonomous driving, human and vehicle detection technology gives vehicles powerful environmental perception and obstacle recognition capabilities, assists in path planning and driving decisions, and greatly improves driving safety. Not only that, the technology is also integrated into the traffic management and information service system, from micro-driving to macro-traffic systems, to fully support the development of intelligent transportation.
[0003] Although human and vehicle detection technology has achieved remarkable results, in actual complex traffic scenarios, algorithms need to be able to accurately identify and locate various dynamic and static targets, such as vehicles, pedestrians, bicycles, traffic signs, and other obstacles that may affect traffic flow. These algorithms must not only be able to handle daily traffic scenarios, but also be able to cope with challenges in extreme weather, complex lighting conditions, and peak traffic hours. In addition to complex environmental factors that affect target feature recognition, human and vehicle target detection at different scales is also a thorny issue. While dealing with the challenges of multi-scale target detection, autonomous driving systems also place extremely high demands on the real-time and accuracy of detection. How to balance the relationship between the two has become the key to current research. In order to meet these challenges, target detection technology is moving towards a more intelligent and automated direction. Summary of the invention
[0004] In order to solve the problems existing in the above prior art, the present invention proposes a method for detecting people and vehicles in real road scenes and subway pedestrians, and an efficient detection network CDC-DETR with improved accuracy and speed compared with previous detection methods. Focusing on data consistency and dynamic adjustment strategies, it aims to improve the adaptability to different scenes and complex targets. Rethinking the model size and detection accuracy under occlusion, as well as the problems of false detection and missed detection of small targets with fewer pixels, improves them for human and vehicle detection applications in assisted driving systems.
[0005] A real-time human-vehicle detection method for assisted driving, the method comprising: step 1, acquiring a driving scene image based on a driver's visual area in real time through a vehicle-mounted image acquisition device, and performing preprocessing; wherein the acquired image data includes the annotation content of the detection target; Step 2: Build an RT-DETR model including a backbone network, an efficient hybrid encoder and a decoder, and make improvements. The improvements include: replacing the first three basic modules of the backbone network model with a context-guided module CGblock, replacing the fourth basic module with a dilated residual connection DWRC module, and adding a context pre-activation pooling attention mechanism CPPA module between the backbone network and the efficient hybrid encoder; Step 3: Based on the improved model, a real-time human and vehicle detection target detection model is obtained, and the target detection model is input into the training set data, and the position and type of the detection target in the picture are obtained from the output end of the model, and the position and type are compared with the content marked in step 1 to train the real-time human and vehicle detection target detection model; Step 4: Obtain the image to be identified, and input the image to be identified into the trained detection model to obtain the target type and location information in the image to be identified.
[0006] Preferably, the context pre-activation pooling attention mechanism CPPA module in step 2 includes: The contextual pre-activation pooling attention (CPAA) module introduces an average pooling layer before the activation function to improve feature extraction accuracy, enhance spatial adaptability, and optimize computational efficiency.
[0007] S21-1, set the input feature to , then global average pooling is applied to extract background features: ;in, represents the global average pooling operation, Feature compression is performed through 1×1 convolution to capture local area information; S21-2, use depthwise separable strip convolution to approximate standard depthwise large kernel convolution: ; in, represents the depth-wise separable convolution in the horizontal direction, kb is the convolution kernel size, which can change dynamically in different network layers; ; Indicates the network depth; S21-3. Generate weight mapping through attention mechanism: ; in, Represents the Sigmoid function, which maps attention to (0,1); Then apply the attention map to the original features to enhance the feature expression: ; And fused by 1×1 convolution and average pooling: ; ; in, It is a 1×1 convolution operation, which is used to fuse and compress feature information. represents the global average pooling operation; Based on this, the output characteristics of the backbone network are set as , then add CAA to its backend for enhancement: .
[0008] Preferably, the DWRC module in step 2 includes: S22-1, the DWRC module adopts a two-step strategy, regional residualization and semantic residualization; The regional residualization includes: generating regional feature maps of different scales using 3×3 convolution: ; Among them, X represents the input feature map, Represents a standard 3×3 convolution, BN is used for normalization, and the ReLU activation function ensures the sparsity of features; S22-2, the semantic residualization includes: applying a depth-separable convolution with a single dilation rate on the regional feature map obtained in the first step: ; in, It represents a depth-separable dilated convolution, whose dilation rate di is determined by the level of the network; Among them, each regional feature map only applies a convolution with a specific expansion rate to avoid computational redundancy. In this way, convolutions with different expansion rates can more effectively perform morphological filtering instead of directly extracting information on complex feature maps; S22-3, the regional feature maps obtained in the semantic residual stage will be fused: ; ; in, represents feature concatenation, Responsible for fusing information and adjusting the number of channels. This process ensures that multi-scale features can complement each other and ultimately enhance the original feature expression through residual connections. This method effectively improves the performance of the model in multi-scale scenarios while optimizing computational efficiency, making it suitable for real-time tasks.
[0009] Preferably, the above-mentioned regional feature maps are multi-scale fused through splicing operations and upsampling operations to ensure that the target information can be effectively integrated on feature maps of different scales, and target detection is performed through the decoder.
[0010] Preferably, the context guidance module CG block in step 2 comprises: a plurality of submodules to efficiently extract and fuse multi-scale features, including local features, surrounding context information and global context information; S23-1. Local feature extractor: Use standard 3×3 convolutional layer to obtain input features , focusing on capturing the details of the target in the image and learning local features: ; Among them, X is the input feature map, Represents local features; S23-2, surrounding context extractor: In order to introduce a larger receptive field, CGBlock uses dilated convolution to extract surrounding context information: ; Among them, r is the expansion rate, which is used to control the receptive field size of the convolution; S23-3, joint feature extractor: concatenate local features and surrounding context features, and activate them through BN and PReLU: ; Among them, [⋅,⋅] represents the concatenation in the channel dimension; S23-4, Global Context Extractor: In order to further optimize feature expression, CGBlock uses global average pooling to obtain global features: ; Adjust weights through the fully connected layer: ; in, Represents the Sigmoid function, which is used to generate attention weights; The global context information is then used to weight the channel attention to enhance useful features, suppress irrelevant features, and improve the model's ability to focus on key areas: ; Finally, the residual structure is adopted to make the information flow better: .
[0011] A real-time human and vehicle detection method for assisted driving ensures that latency is reduced and detection efficiency is improved when processing target detection tasks in real time on edge computing devices by streamlining, quantizing, and accelerating parallel computing of computational graphs. The feature graph is multi-scale fused through a splicing operation (cat operation) and an upsampling operation to ensure that target information can be effectively integrated on feature graphs of different scales, and target detection is performed through a decoder. The method can adapt to a variety of complex traffic scenarios, including highways, city streets, intersections, etc., and can operate stably in extreme weather conditions (such as rainy and snowy days) or low-light environments, and is suitable for environmental perception tasks in autonomous driving systems.
[0012] Compared with the prior art, the beneficial effects achieved by the present invention are: In the process of backbone network feature extraction, the contextual pre-activated pooling attention mechanism (CPPA), dilated residual connection (DWRC) and context-guided module (CG block) are introduced to effectively improve the ability to capture multi-scale features, enhance the detection accuracy of small targets, and improve the real-time performance and efficiency of the model, which is particularly suitable for target detection in complex traffic scenes. The CPPA module strengthens the contextual dependency between long-distance pixels, supplements the multi-scale local features, further improves the accuracy of target detection, and accelerates the reasoning process, so that the model can run efficiently on edge computing devices. The DWRC and CGblock modules are optimized in terms of multi-scale information capture and feature fusion, respectively, which improves the performance of the model, while reducing the computational burden and enhancing the robustness in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a schematic diagram of the overall structure of the model of the present invention; Figure 3 Schematic diagram of the dilated residual connection (DWRC) module of the present invention; Figure 4 The experimental results of all models in the present invention are based on the MMdetection framework and are trained and tested under unified experimental conditions; Figure 5 These are eight groups of ablation experiment result diagrams conducted based on RT-DETR in the present invention. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] See also Figure 1-Figure 5 , the present invention provides a technical solution: Embodiment 1: A real-time human-vehicle detection method for assisted driving, such as Figure 1 As shown, the method includes: Step 1: Data collection and processing; Through the on-board image acquisition device, the driving scene images based on the driver's visual area are acquired in real time and preprocessed.
[0016] Step 2: Construction of real-time human and vehicle detection target detection model; An RT-DETR model including a backbone network, an efficient hybrid encoder and a decoder is constructed, and then improvements are made. The improvements include replacing the first three basic modules of the backbone network model with a context-guided module CG block, replacing the fourth basic module with a dilated residual connection DWRC module, and adding a layer of context pre-activation pooling attention mechanism CPPA module between the backbone network and the efficient hybrid encoder.
[0017] Step 3: Input the training set data into the real-time human and vehicle detection target detection model constructed in step 2, obtain the position and type of the detection target in the image from the output of the model, compare it with the content marked in step 1, and train the real-time human and vehicle detection target detection model.
[0018] Step 4: Real-time human and vehicle detection: Input the image data to be identified into the model trained in step 3 to obtain the target type and location information in the image. In order to verify the effectiveness of CDC-DETR, a comparative experiment was conducted on the BDD100K dataset with other mainstream target detection algorithms, including SSD, Faster R-CNN, Mask-Rcnn, RetinaNet, RTMDet and YOLO series.
[0019] The BDD100K dataset contains 70,000 images, including images of various time periods and weather conditions, and the diversity of driving environments must be ensured. The dataset covers various common vehicle target information, pedestrian information, and traffic sign information under cloudy, rainy, snowy weather, large differences in lighting environments, many small targets, and complex occlusion conditions. 10,000 images are randomly selected and divided into three parts: training set, validation set, and test set. The training set is used to adjust the model parameters throughout the training process. The validation set is used to independently evaluate the training results and determine whether the model exhibits overfitting to ensure its generalization. The test set is used for the final evaluation. The dataset is divided into 8,000 training images, 1,000 validation images, and 1,000 test images.
[0020] To ensure the reliability and comparability of mean average precision (mAP) and other key indicators in horizontal comparison, all models are trained and tested under unified experimental conditions based on the MMdetection framework. The specific experimental results are as follows: Figure 4 ; Under the comprehensive consideration of multi-dimensional detection indicators, the CDC-DETR model shows significant advantages over all baseline algorithms. In terms of the core indicator of mean average precision (mAP), CDC-DETR is in a leading position under common intersection-over-union threshold settings, fully demonstrating its excellent prediction performance in target detection tasks. At the same time, in the face of different scenes, target scales and complex background interference, CDC-DETR's detection results maintain high consistency and accuracy, which strongly proves that the model has excellent generalization capabilities and can adapt to a variety of practical application scenarios.
[0021] In addition, in order to deeply evaluate the detection performance of the proposed algorithm and explore the effectiveness of each module enhancement, eight groups of ablation experiments were conducted based on RT-DETR. Each group of experiments used consistent hyperparameter settings and underwent the same training time. The experimental results are shown in Figure 2. Figure 5 As shown in the figure, A refers to the CPPA module, B refers to the DWRC module, and C refers to the ContextGuided module.
[0022] From the results, we can observe that the CDC-DETR model performs well in terms of recall, mAP@0.5 and mAP@0.5:0.95, and the number of parameters, which is significantly better than other model configurations. If any module is removed, the average accuracy of the model will decrease or the number of parameters will increase. After adding module A alone, the accuracy of the model increased by 1.5%, despite the increase in model complexity. When module B is introduced, although the accuracy is not significantly improved, in the experiment of combining modules BC, module B and module C show a significant synergistic effect. Specifically, module B plays a key role in improving model performance, especially when working with module C, it can effectively improve the detection accuracy and recall of the model. When the context-guided mechanism of module C globally optimizes the target features through semantic association modeling, it can significantly enhance the directional decision-making ability of module B during parameter recalibration. The integrated CDC-DETR architecture achieves a 3.4% accuracy gain while reducing the floating-point operations by 11% compared with the baseline model through optimization strategies such as sparse attention masking, which fully verifies the core principle of system synergy over local optimization in modular design.
[0023] Embodiment 2: Figure 2 This is a schematic diagram of the overall structure of the model described in the present invention, including a CDC-DETR model of a backbone network, an efficient hybrid encoder and a decoder. The improvements include: replacing the first three basic modules of the backbone network model with a context-guided module CG block, replacing the fourth basic module with a dilated residual connection DWRC module, and adding a layer of context pre-activation pooling attention mechanism CPPA module between the backbone network and the efficient hybrid encoder.
[0024] The contextual pre-activation pooling attention (CPAA) module introduces an average pooling layer before the activation function to improve feature extraction accuracy, enhance spatial adaptability, and optimize computational efficiency. The CPPA module first gives the input feature , first apply global average pooling to extract background features: ;in, Represents the global average pooling operation, which performs feature compression through 1×1 convolution to capture local area information.
[0025] Next, a depth-separable strip convolution is used to approximate the standard depth-large kernel convolution: ; in, It represents the depth-wise separable convolution in the horizontal direction. kb is the convolution kernel size, which can change dynamically in different network layers, such as: ; here, It is the network depth, which enables deep networks to use larger receptive fields.
[0026] Then, the weight map is generated through the attention mechanism: ; in, Represents the Sigmoid function, which maps attention to (0,1).
[0027] Then apply the attention map to the original features to enhance the feature expression: ; And fused by 1×1 convolution and average pooling: ; ; in: It is a 1×1 convolution operation, which is used to fuse and compress feature information. represents the global average pooling operation; Finally, assuming that the backbone network outputs the feature , add CAA to its backend for enhancement: .
[0028] The context-guided module CG Block consists of multiple sub-modules to efficiently extract and fuse multi-scale features, including local features, surrounding context information and global context information. Local Feature Extraction: It uses a standard 3×3 convolutional layer to obtain the input feature floc, focusing on capturing the details of the target in the image and learning local features: ; Here, X is the input feature map, representing the local features.
[0029] Surrounding Context Extraction: In order to introduce a larger receptive field, CGBlock uses dilated convolution to extract surrounding context information: ; Among them, r is the dilation rate, which is used to control the size of the receptive field of the convolution.
[0030] Joint Feature Extraction: concatenates local features with surrounding context features and activates them through BN and PReLU: ; in, Represents concatenation in the channel dimension.
[0031] Global Context Extraction: In order to further optimize feature expression, CGBlock uses global average pooling (GAP) to obtain global features: ; Then adjust the weights through the fully connected layer: ; here, Is the Sigmoid function used to generate attention weights.
[0032] The global context information is then used to weight the channel attention to enhance useful features, suppress irrelevant features, and improve the model's ability to focus on key areas: ; Finally, the residual structure is adopted to make the information flow better: .
[0033] Embodiment 3: Figure 3 This is a schematic diagram of the dilated residual connection (DWRC) module described in the present invention, which adopts a two-step strategy, Region Residualization (RR): 3×3 convolution is used to generate regional feature maps of different scales to make the feature expression more concise, so as to facilitate the subsequent efficient use of dilated convolution.
[0034] ; Among them, X is the input feature map, It is a standard 3×3 convolution, Batch Normalization (BN) is used for normalization, and the ReLU activation function ensures the sparsity of features.
[0035] Semantic Residualization (SR): A depth-wise separable convolution with a single dilation rate is applied to the regional feature map obtained in the first step to extract semantic information more efficiently.
[0036] ; in, represents the depth-wise separable dilated convolution, and its dilation rate Determined by the level of the network. Each regional feature map only applies a convolution with a specific dilation rate to avoid computational redundancy. In this way, convolutions with different dilation rates can more effectively perform morphological filtering rather than directly extracting information on complex feature maps. The feature maps obtained in the semantic residualization (SR) stage are fused: ; ; in, represents feature concatenation, Responsible for fusing information and adjusting the number of channels. This process ensures that multi-scale features can complement each other and ultimately enhance the original feature expression through residual connections. This method effectively improves the performance of the model in multi-scale scenarios while optimizing computational efficiency, making it suitable for real-time tasks.
[0037] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A real-time human-vehicle detection method for assisted driving, characterized in that: The method comprises: Step 1: Acquire a driving scene image based on the driver's visual area in real time through an on-board image acquisition device and perform preprocessing; wherein the acquired image data includes the annotation content of the detection target; Step 2: Build an RT-DETR model including a backbone network, an efficient hybrid encoder and a decoder, and make improvements. The improvements include: replacing the first three basic modules of the backbone network model with a context-guided module CG block, replacing the fourth basic module with a dilated residual connection DWRC module, and adding a context pre-activation pooling attention mechanism CPPA module between the backbone network and the efficient hybrid encoder; Step 3: Based on the improved model, a real-time human and vehicle detection target detection model is obtained, and the target detection model is input into the training set data, and the position and type of the detection target in the picture are obtained from the output end of the model, and the position and type are compared with the content marked in step 1 to train the real-time human and vehicle detection target detection model; Step 4: Obtain the image to be identified, and input the image to be identified into the trained detection model to obtain the target type and location information in the image to be identified.
2. A real-time human-vehicle detection method for assisted driving as claimed in claim 1, characterized in that: The context pre-activation pooling attention mechanism CPPA module in step 2 includes: S21-1, set the input feature to , then global average pooling is applied to extract background features: ;in, represents the global average pooling operation, Feature compression is performed through 1×1 convolution to capture local area information; S21-2, use depthwise separable strip convolution to approximate standard depthwise large kernel convolution: ;in, represents the depth-wise separable convolution in the horizontal direction, kb is the convolution kernel size, which can change dynamically in different network layers; ; Indicates the network depth; S21-3. Generate weight mapping through attention mechanism: ; in, Represents the Sigmoid function, which maps attention to (0,1); Then apply the attention map to the original features to enhance the feature expression: ; And fused by 1×1 convolution and average pooling: ; ; in, It is a 1×1 convolution operation, which is used to fuse and compress feature information. represents the global average pooling operation; Based on this, the output characteristics of the backbone network are set as , then add CAA to its backend for enhancement: 。 3. A real-time human-vehicle detection method for assisted driving as claimed in claim 1, characterized in that: The DWRC module in step 2 includes: S22-1, the DWRC module adopts a two-step strategy, regional residualization and semantic residualization; The regional residualization includes: generating regional feature maps of different scales using 3×3 convolution: ; Among them, X represents the input feature map, Represents a standard 3×3 convolution, BN is used for normalization, and the ReLU activation function ensures the sparsity of features; S22-2, the semantic residualization includes: applying a depth-separable convolution with a single dilation rate on the regional feature map obtained in the first step: ; in, represents a depth-wise separable dilated convolution, whose dilation rate Determined by the level of the network; Among them, each regional feature map only applies a convolution with a specific expansion rate to avoid computational redundancy; S22-3, the regional feature maps obtained in the semantic residualization stage will be fused: ; ; in, represents feature concatenation, Responsible for fusing information and adjusting the number of channels.
4. A real-time human-vehicle detection method for assisted driving as claimed in claim 1, characterized in that: The context guidance module CG block in step 2 includes: multiple sub-modules to efficiently extract and fuse multi-scale features, including local features, surrounding context information and global context information; S23-1. Local feature extractor: Use standard 3×3 convolutional layer to obtain input features , focusing on capturing the details of the target in the image and learning local features: ; Among them, X is the input feature map, Represents local features; S23-2, surrounding context extractor: In order to introduce a larger receptive field, CGBlock uses dilated convolution to extract surrounding context information: ; Among them, r is the expansion rate, which is used to control the receptive field size of the convolution; S23-3, joint feature extractor: concatenate local features and surrounding context features, and activate them through BN and PReLU: ; Among them, [⋅,⋅] represents the concatenation in the channel dimension; S23-4, Global Context Extractor: In order to further optimize feature expression, CGBlock uses global average pooling to obtain global features: ; Adjust weights through the fully connected layer: ; in, Represents the Sigmoid function, which is used to generate attention weights; The global context information is then used to weight the channel attention to enhance useful features, suppress irrelevant features, and improve the model's ability to focus on key areas: ; Finally, the residual structure is adopted to make the information flow better: .
5. A real-time human-vehicle detection method for assisted driving as claimed in claim 3, characterized in that: The regional feature map is multi-scale fused through a splicing operation and an upsampling operation to ensure that target information can be effectively integrated on feature maps of different scales, and target detection is performed through a decoder.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a real-time human-vehicle detection method for assisted driving as described in any one of claims 1-5 are implemented.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps in the real-time human and vehicle detection method for assisted driving as described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
RT-DETR infrared weak aircraft detection method and system
CN117974972A
Automatic sky bead image recognition method and system based on improved convolutional neural network
CN119418130A
Salient target detection method of intelligent auxiliary driving system
CN119559601A
Lightweight real-time target detection method and device, server and storage medium
CN119810428A