Vehicle pedestrian target detection method and device, computer equipment and storage medium

By adding inverted bottleneck UIB module and weighted bidirectional feature pyramid network to the RT-DETR model, the problem of high computational cost and difficult deployment in vehicle pedestrian target detection is solved, achieving more efficient detection performance and lower computational complexity.

CN119992511AActive Publication Date: 2025-05-13ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510048589.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing DETR model has high computational cost and difficult deployment in vehicle pedestrian target detection, resulting in poor detection performance in complex environments.

Method used

Improvements are made based on the RT-DETR model, and the inverted bottleneck UIB module and weighted bidirectional feature pyramid network are added, and the structure of the backbone network and hybrid encoder is adjusted to reduce the calculation complexity and parameter amount.

Benefits of technology

It significantly reduces the computational complexity and parameter quantity of the vehicle pedestrian object detection model, improves detection performance, especially in complex scenarios, and reduces inference time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992511A_ABST
    Figure CN119992511A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle pedestrian target detection method and device, computer equipment and a storage medium, and belongs to the technical field of computer vision. The method comprises the following steps: constructing a vehicle and pedestrian target detection model based on improved RT-DETR, including adding an inverted bottleneck UIB module in a backbone network of an original RT-DETR model and adjusting an internal feature fusion network structure of a hybrid encoder through a weighted bidirectional feature pyramid network in the hybrid encoder of the original RT-DETR model, and performing iterative training on the vehicle and pedestrian target detection model obtained through construction by using the obtained vehicle and pedestrian data set to obtain a trained model, detecting vehicles and pedestrians based on the trained model, and outputting a detection result. Therefore, on the basis of ensuring that the precision of the algorithm is not reduced, the method has the characteristics of shorter delay, smaller calculation complexity and smaller parameter quantity, and the capability of capturing vehicles and pedestrians by a target detection model in a complex scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a vehicle and pedestrian target detection method, device, computer equipment and storage medium. Background Art

[0002] During daily road driving, pedestrians and cyclists are likely to appear from behind vehicles on both sides of the road or in blind spots blocked by trees, causing vehicles and pedestrians to overlap and block each other. In a changing environment, existing target detection models still encounter difficulties in recognition when facing scenes with severe occlusion, which not only leads to increased missed detection and false detection rates, but also reduces detection accuracy, thus affecting traffic safety and making traffic accidents more likely to occur. Therefore, in the development of autonomous driving, how to detect vehicles and pedestrians in a timely manner and discover potential driving obstacles is a very important issue.

[0003] In the prior art, there are two typical architectures for target detectors, including CNN-based target detection models and Transformer-based target detection models. CNN-based target detection models have been widely used and studied due to their good accuracy and processing speed. Traditional CNN target detection models, such as the YOLO series, are widely used in real-time target detection fields, such as target tracking, autonomous driving, and video monitoring, because of their better robustness in processing complex scenes and simplicity in training and deployment. However, the two key steps of threshold screening and non-maximum suppression will reduce the robustness and detection speed of the target detection model. In order to further improve the detection performance, removing hand-crafted components has become an important trend in the field of target detection. The DETR model uses Transformer technology to redefine target detection as a set prediction problem, and adopts an end-to-end trainable encoder-decoder structure to replace the traditional region proposal-based method, eliminating many hand-crafted components, such as non-maximum suppression (NMS), without too many proposal generation and post-processing steps, greatly simplifying the target detection process and realizing end-to-end target detection.

[0004] Although the DETR model performs well in target detection tasks, its high computational cost limits the performance of the DETR model in complex environments. The high computational cost not only increases processing delays, but also places high demands on hardware resources, making the deployment of the DETR model more difficult. In this case, it is particularly important to design an improved DETR model to reduce computational costs and improve real-time and robustness. Summary of the invention

[0005] 1. Technical problems to be solved

[0006] In view of the problems in the prior art of high computational cost and difficult deployment when performing vehicle and pedestrian target detection through the DETR model, the present invention provides a vehicle and pedestrian target detection method, device, computer equipment and storage medium, which further reduce the computational complexity of the algorithm and the actual deployment cost of the target detection model, and improve the target detection model's ability to capture vehicles and pedestrians in complex scenarios.

[0007] 2. Technical solution

[0008] The purpose of the present invention is achieved through the following technical solutions.

[0009] A vehicle and pedestrian target detection method comprises the following steps:

[0010] Obtain and input a vehicle and pedestrian dataset, and preprocess the vehicle and pedestrian dataset;

[0011] Constructing a vehicle and pedestrian target detection model based on the improved RT-DETR; the construction of the vehicle and pedestrian target detection model based on the improved RT-DETR includes the following improvements on the basis of the original RT-DETR model:

[0012] 1) Add an inverted bottleneck UIB module to the original RT-DETR model backbone network;

[0013] 2) In the original RT-DETR model hybrid encoder, the internal feature fusion network structure of the hybrid encoder is adjusted through a weighted bidirectional feature pyramid network;

[0014] The preprocessed vehicle and pedestrian data set is used to iteratively train the constructed vehicle and pedestrian target detection model to obtain a trained vehicle and pedestrian target detection model;

[0015] Detect vehicles and pedestrians based on the trained vehicle and pedestrian target detection model and output the detection results.

[0016] As a further improvement of the present invention, the inverted bottleneck UIB module is composed of several UIB blocks; the UIB block includes a first DW convolution layer, a first optional PW layer, a second DW convolution layer and a second optional PW layer connected in sequence, and the first optional PW layer and the second optional PW layer present an inverted bottleneck structure.

[0017] As a further improvement of the present invention, the first DW convolution layer and the second DW convolution layer include a starting depth convolution layer and an intermediate depth convolution layer; the first optional PW layer and the second optional PW layer include an extended convolution layer and a projection convolution layer.

[0018] As a further improvement of the present invention, the starting deep convolution layer performs preliminary processing on the feature map input by the backbone network, the extended convolution layer uses 1×1 convolution to increase the dimension of the feature map after the preliminary processing, and increases the number of channels according to the expansion ratio to capture detail features through the intermediate deep convolution layer, and the projection convolution layer uses 1×1 convolution to reduce the feature image channels to the target number of channels to fuse the feature map information.

[0019] As a further improvement of the present invention, the hybrid encoder includes several convolutional layers, an attention-based intra-scale feature interaction module and a cross-scale feature fusion module; the cross-scale feature fusion module includes a Bi-FPN feature fusion module and a RepC3 feature processing module.

[0020] As a further improvement of the present invention, the Bi-FPN feature fusion module normalizes the weights of the feature maps input by the backbone network, traverses different feature maps, and assigns corresponding weights to them, and stacks the weighted feature maps to form a fused feature map.

[0021] As a further improvement of the present invention, in the Bi-FPN feature fusion module, fast normalization fusion is used as a feature fusion method, and weights are assigned to different feature maps input by the backbone network to determine the influence of different feature maps on the output results of the vehicle pedestrian target detection model. The weight calculation formula is:

[0022]

[0023] Among them, Output represents the output result, i represents the index of the input feature layer, j represents the set of feature layer indexes participating in the weighted calculation, and w i represents the weight of a specific feature layer i, w j Represents the weight of the feature layer in the feature layer set, I i Represents the feature map of a specific feature layer i, ∈ represents a small value set to 0.0001 to ensure the stability of the result.

[0024] A vehicle and pedestrian target detection device, comprising:

[0025] The data processing module obtains and inputs the vehicle and pedestrian data set, and pre-processes the vehicle and pedestrian data set;

[0026] The model building module builds a vehicle and pedestrian target detection model based on the improved RT-DETR; the building of the vehicle and pedestrian target detection model based on the improved RT-DETR includes the following improvements on the basis of the original RT-DETR model:

[0027] 1) Add an inverted bottleneck UIB module to the original RT-DETR model backbone network;

[0028] 2) In the original RT-DETR model hybrid encoder, the internal feature fusion network structure of the hybrid encoder is adjusted through a weighted bidirectional feature pyramid network;

[0029] The training module uses the preprocessed vehicle and pedestrian data set to iteratively train the constructed vehicle and pedestrian target detection model to obtain a trained vehicle and pedestrian target detection model;

[0030] The detection module detects vehicles and pedestrians based on the trained vehicle and pedestrian target detection model and outputs the detection results.

[0031] A computer device comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements any of the above-mentioned methods when executing the computer program.

[0032] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above methods is executed.

[0033] 3. Beneficial effects

[0034] Compared with the prior art, the advantages of the present invention are:

[0035] The vehicle-pedestrian target detection method, device, computer equipment and storage medium of the present invention are improved based on the RT-DETR model, and further lightweight processing is performed on it. An inverted bottleneck UIB module is added to the backbone network, and the flexibility of space and channel mixing is introduced to enhance the selection of receiving domains, thereby reducing the computational complexity and parameter quantity of the vehicle-pedestrian target detection model; at the same time, a weighted bidirectional feature pyramid network is used to adjust the network structure of the hybrid encoder, increase the balance of the importance of features at different levels of the vehicle-pedestrian target detection model, improve the efficiency of feature fusion, and improve the vehicle-pedestrian target detection model. The ability to capture small targets in overlapping situations. On the basis of ensuring that the accuracy of the algorithm is not reduced, the present invention has the characteristics of shorter delay, smaller computational complexity and parameter quantity, and significantly improves the lightweight performance of the vehicle-pedestrian target detection model algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a vehicle and pedestrian target detection method according to an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of a detection network structure according to an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of the UIB module structure according to an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram of the Bi-FPN feature fusion network structure of an embodiment of the present invention;

[0040] Figure 5 This is a training effect diagram of the vehicle and pedestrian target detection method according to an embodiment of the present invention;

[0041] Figure 6 This is a test effect diagram of the vehicle and pedestrian target detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0043] Example

[0044] like Figure 1 As shown, a vehicle-pedestrian target detection method provided by this embodiment includes the following steps: obtaining and inputting a vehicle-pedestrian data set, and preprocessing the vehicle-pedestrian data set; constructing a vehicle-pedestrian target detection model based on an improved RT-DETR; iteratively training the constructed vehicle-pedestrian target detection model using the preprocessed vehicle-pedestrian data set to obtain a trained vehicle-pedestrian target detection model; detecting vehicles and pedestrians based on the trained vehicle-pedestrian target detection model, and outputting the detection results. It is worth noting that in this embodiment, with respect to constructing a vehicle-pedestrian target detection model based on an improved RT-DETR, the following improvements are made on the basis of the original RT-DETR model: 1) adding an inverted bottleneck UIB module to the original RT-DETR model backbone network; 2) adjusting the internal feature fusion network structure of the hybrid encoder through a weighted bidirectional feature pyramid network in the hybrid encoder of the original RT-DETR model.

[0045] Specifically in this embodiment, first, publicly available image datasets of vehicles and pedestrians are screened and obtained from existing platforms such as roboflow and github, and the obtained image dataset contains a total of 10,000 road pictures. Then, the vehicle and pedestrian dataset is preprocessed. In this embodiment, the preprocessing includes annotating the image dataset, and dividing the annotated image dataset into a training set, a validation set, and a test set according to a certain ratio. The specific steps include: annotating the acquired image dataset through existing annotation tools, such as make sense, etc. Each set of annotations will obtain a picture and a .txt format label file, and the label class is divided into five categories: pedestrians, non-motor vehicles, buses, cars, and driving obstacles; the pictures and labels in the 10,000 image dataset are divided into 7,000 training sets, 2,000 validation sets, and 1,000 test sets, and the corresponding test, train, and valid folders are created, and subfolders images and labels are created under each folder to store pictures and labels respectively.

[0046] Furthermore, a vehicle and pedestrian target detection model based on improved RT-DETR is constructed. In this embodiment, a vehicle and pedestrian target detection model based on improved RT-DETR is constructed, including the following improvements on the basis of the original RT-DETR model: adding an inverted bottleneck UIB module to the original RT-DETR model backbone network, and adjusting the internal feature fusion network structure of the hybrid encoder through a weighted bidirectional feature pyramid network in the original RT-DETR model hybrid encoder. Thus, a vehicle and pedestrian target detection model based on improved RT-DETR is obtained.

[0047] like Figure 2 As shown, in this embodiment, the vehicle pedestrian target detection model based on the improved RT-DETR includes a backbone network, a hybrid encoder and a decoder.

[0048] For the backbone network, the backbone network includes a convolutional normalization layer, an inverted bottleneck UIB module, and a 1×1 convolutional layer. The convolutional normalization layer consists of a 2D convolutional block, a batch normalization module, and a RELU activation function. The inverted bottleneck UIB module consists of several UIB blocks. In this embodiment, the inverted bottleneck UIB module consists of six UIB blocks. Figure 3 As shown, each UIB block includes a first DW (DepthWise) convolution layer, a first optional PW (PointWise) layer, a second DW convolution layer and a second optional PW layer connected in sequence. In particular, in this embodiment, the first optional PW layer and the second optional PW layer are in an inverted bottleneck structure.

[0049] In this embodiment, the first DW convolution layer and the second DW convolution layer include a starting depth convolution layer and an intermediate depth convolution layer, and the first optional PW layer and the second optional PW layer include an extended convolution layer and a projection convolution layer. Furthermore, the starting depth convolution layer performs preliminary processing on the feature map input by the backbone network, and the extended convolution layer uses 1×1 convolution to perform dimensionality processing on the feature map after preliminary processing, and increases the number of channels according to the expansion ratio to capture detail features through the intermediate depth convolution layer. The projection convolution layer uses 1×1 convolution to reduce the feature image channel to the target channel number and fuse the feature map information.

[0050] Therefore, in this embodiment, the backbone network includes a convolutional normalization layer, a convolutional normalization layer, a convolutional normalization layer, an inverted bottleneck UIB module and an inverted bottleneck UIB module connected in sequence.

[0051] like Figure 2 As shown in the figure, the vehicle and pedestrian images input to the backbone network are first processed by the first convolution normalization layer, which uses a 3×3 convolution block, with 3 input channels and 32 output channels. The image features are input to the second convolution normalization layer, which has two convolution blocks. The first convolution block is 3×3, and the second convolution block is 1×1. The number of input and output channels is 32. The output of the second convolution normalization layer is input into the hybrid encoder as the first image feature. At the same time, the backbone network continues to process the image features. The third convolution normalization layer has two convolution blocks. The first convolution block is 3×3, the number of input channels is 32, and the number of output channels is 96. The second convolution block is 1×1, the number of input channels is 96, and the number of output channels is 64. The output of the third convolution normalization layer is input into the hybrid encoder as the second image feature. At the same time, the backbone network continues to process the image features. The image features are input into the first inverted bottleneck UIB module, which contains six UIB blocks. Each UIB block consists of depthwise separable convolution and 1×1 convolution. The image features are reduced in dimension by depthwise separable convolution, and then increased in dimension by convolution. The first UIB block has 64 input channels and 96 output channels, and the number of input and output channels of each of the remaining UIB blocks is 96. The output of the first inverted bottleneck UIB module is input into the hybrid encoder as the third image feature. At the same time, the backbone network continues to process the image features. The image features are input into the second inverted bottleneck UIB module. The structure of the second inverted bottleneck UIB module is the same as that of the first inverted bottleneck UIB module. The first UIB block has 96 input channels and 128 output channels, and the number of input and output channels of each of the remaining UIB blocks is 128. The output of the second inverted bottleneck UIB module is input into the hybrid encoder as the fourth image feature.

[0052] Therefore, this embodiment adds an inverted bottleneck UIB module to the backbone network, introduces the flexibility of space and channel mixing, enhances the selection of receiving fields, and reduces the computational complexity and parameter amount of the vehicle and pedestrian target detection model.

[0053] For the hybrid encoder, in this embodiment, if Figure 4 As shown, in the original RT-DETR model hybrid encoder, the internal feature fusion network structure of the hybrid encoder is adjusted by a weighted bidirectional feature pyramid network. In this embodiment, the hybrid encoder includes several convolutional layers, an attention-based intra-scale feature interaction module (AIFI) and a cross-scale feature fusion module (CCFM). The CCFM module includes a Bi-FPN feature fusion module and a RepC3 feature processing module.

[0054] It should be noted that the Bi-FPN feature fusion module normalizes the weights of the feature maps input by the backbone network, traverses different feature maps, assigns corresponding weights to them, and stacks the weighted feature maps to form a fused feature map. Specifically, the weighted feature maps are stacked to form a new tensor, and the stacking operation is performed according to dimension 0, that is, multiple feature maps are stacked along the batch direction, and all weighted stacked feature maps are summed on dimension 0 to obtain the final fused feature map for use by subsequent network modules.

[0055] In this embodiment, in the Bi-FPN feature fusion module, fast normalization fusion is used as the feature fusion method, and weights are assigned to different feature maps input by the backbone network to determine the impact of different feature maps on the output results of the vehicle pedestrian target detection model. The weight calculation formula is:

[0056]

[0057] Among them, Output represents the output result, i represents the index of the input feature layer, j represents the set of feature layer indexes participating in the weighted calculation, and w i represents the weight of a specific feature layer i, w j Represents the weight of the feature layer in the feature layer set, I i Represents the feature map of a specific feature layer i, ∈ represents a small value set to 0.0001 to ensure the stability of the result.

[0058] Therefore, the first image feature, the second image feature and the third image feature output by the backbone network are respectively input into the CCFM module of the hybrid encoder through 1×1 convolutional layers, and the fourth image feature is input into the AIFI module as a high-level feature through a 1×1 convolutional layer, which saves computing resources and avoids the problem of confusion between high-level features and low-level features. Finally, the fourth image feature is input into the CCFM module through the AIFI module.

[0059] In this embodiment, the four image features output by the backbone network are imported by the 1×1 convolution layer, fused with the high-level features via the Bi-FPN feature fusion module, and the weighted bidirectional feature pyramid network aggregates multi-scale features in a top-down manner. In this embodiment, taking the third image feature and the fourth image feature as an example, the feature transmission paths at different scale levels in the CCFM module are as follows:

[0060]

[0061] in, represents the input features of the third image feature layer, represents the input features of the fourth image feature layer, represents the intermediate features of the third image feature layer in the top-down path, represents the output features of the third image feature layer of the bottom-up path, represents the intermediate features of the fourth image feature layer in the top-down path, Represents the output features of the fourth image feature layer of the bottom-up path, Conv represents the convolution operation for feature processing, w1 represents the weight of the specific feature layer 1, w2 represents the weight of the specific feature layer 2, and Resize represents the upsampling or downsampling operation for resolution matching.

[0062] Therefore, in the hybrid encoder, after the features of different scale levels are obtained through the CCFM module, the features are output to the decoder module for decoding and then the fused features are output. In this embodiment, the network structure of the hybrid encoder is adjusted by using a weighted bidirectional feature pyramid network, the vehicle and pedestrian target detection model balances the importance of features at different levels, improves the efficiency of feature fusion, and improves the vehicle and pedestrian target detection model's ability to capture small targets in overlapping situations.

[0063] It should be noted that in this embodiment, an experimental platform was built on a desktop window system 11, the GPU used was NVDIAGeForce RTX 4070super, the deep learning framework was Pytorch 2.2.0 version, and in the training hyperparameter settings, the initial learning rate was set to 0.0001, the batch was set to 4, the learning rate momentum value was set to 0.9, the initial input vehicle and pedestrian image size was 640×640 pixels, the AdamW optimizer was used, the training rounds were set to 100, and the constructed vehicle and pedestrian target detection model was trained based on the training set to obtain the best.pt weight file. The vehicle and pedestrian target detection model in the training process was evaluated based on the validation set to obtain a trained vehicle and pedestrian target detection model. Figure 5 As shown, this is the training effect diagram of the trained vehicle and pedestrian target detection model.

[0064] Then, the best.pt weight file obtained by training is tested on the test set, and the average accuracy MAP, model inference time INFERENCE and other results obtained by detection are compared with the original RT-DETR target detection method. The comparison results are shown in Table 1 below.

[0065] Table 1

[0066]

[0067] As can be seen from Table 1 above, the MAP value of the vehicle-pedestrian target detection model based on the improved RT-DETR constructed in this embodiment is better than that of the original RT-DETR model. At the same time, the number of parameters is reduced by 65%, the computational complexity is reduced by 70%, and the inference time is reduced by 0.6ms, which significantly reflects the superior performance of the vehicle-pedestrian target detection model based on the improved RT-DETR constructed in this embodiment. Figure 6 As shown, the test set is input into the trained vehicle and pedestrian target detection model for detection, and the detection results are output.

[0068] This embodiment uses a vehicle-pedestrian target detection model based on an improved RT-DETR, and only needs to use a lightweight detector for vehicle-pedestrian target detection to achieve better results. Using a detector with a simpler structure in the vehicle-pedestrian target detection process can greatly reduce the resources required for calculation, improve the detection speed of the vehicle-pedestrian target detection model based on the improved RT-DETR, and reduce the video memory occupancy required in the detection process, thereby avoiding unnecessary waste and effectively utilizing limited computing resources, thereby realizing the practical application of the vehicle-pedestrian target detection model based on the improved RT-DETR in vehicle-pedestrian target detection, and has strong practicality and wide applicability.

[0069] This embodiment also provides a vehicle-pedestrian target detection device, including a data processing module, a model improvement module, a training module and a detection module. The data processing module is used to obtain and input a vehicle-pedestrian data set, and pre-process the vehicle-pedestrian data set. The model construction module constructs a vehicle-pedestrian target detection model based on the improved RT-DETR. The construction of the vehicle-pedestrian target detection model based on the improved RT-DETR includes the following improvements on the basis of the original RT-DETR model: 1) adding an inverted bottleneck UIB module to the original RT-DETR model backbone network; 2) adjusting the internal feature fusion network structure of the hybrid encoder through a weighted bidirectional feature pyramid network in the original RT-DETR model hybrid encoder. The training module uses the pre-processed vehicle-pedestrian data set to iteratively train the constructed vehicle-pedestrian target detection model to obtain a trained vehicle-pedestrian target detection model. The detection module detects vehicles and pedestrians based on the trained vehicle-pedestrian target detection model and outputs the detection results. The vehicle pedestrian target detection device provided in this embodiment can implement any of the vehicle pedestrian target detection methods, and the specific working process of a vehicle pedestrian target detection device can refer to the corresponding process in the vehicle pedestrian target detection method embodiment. The method and device provided in this embodiment can be implemented in other ways. For example, the device embodiment described above is only schematic; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual connection or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.

[0070] This embodiment also provides a computer device. A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the vehicle-pedestrian target detection method when executing the computer program.

[0071] This embodiment also provides a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a vehicle pedestrian target detection method described in this embodiment is executed. The computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, device or device; the program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0072] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited thereto, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all belong to the protection scope of the present invention. In addition, the word "including" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate the name, and do not indicate any specific order.

Claims

1. A vehicle and pedestrian target detection method, comprising the following steps: Obtain and input a vehicle and pedestrian dataset, and preprocess the vehicle and pedestrian dataset; Construct a vehicle and pedestrian target detection model based on improved RT-DETR; The vehicle and pedestrian target detection model based on the improved RT-DETR is constructed, including the following improvements on the basis of the original RT-DETR model: 1) Add an inverted bottleneck UIB module to the original RT-DETR model backbone network; 2) In the original RT-DETR model hybrid encoder, the internal feature fusion network structure of the hybrid encoder is adjusted through a weighted bidirectional feature pyramid network; The preprocessed vehicle and pedestrian data set is used to iteratively train the constructed vehicle and pedestrian target detection model to obtain a trained vehicle and pedestrian target detection model; Detect vehicles and pedestrians based on the trained vehicle and pedestrian target detection model and output the detection results.

2. A vehicle-pedestrian target detection method according to claim 1, characterized in that: The inverted bottleneck UIB module is composed of several UIB blocks; the UIB block includes a first DW convolution layer, a first optional PW layer, a second DW convolution layer and a second optional PW layer connected in sequence, and the first optional PW layer and the second optional PW layer present an inverted bottleneck structure.

3. A vehicle-pedestrian target detection method according to claim 2, characterized in that: The first DW convolution layer and the second DW convolution layer include a starting depth convolution layer and an intermediate depth convolution layer; the first optional PW layer and the second optional PW layer include an extended convolution layer and a projection convolution layer.

4. A vehicle-pedestrian target detection method according to claim 3, characterized in that: The starting deep convolution layer performs preliminary processing on the feature map input by the backbone network, the extended convolution layer uses 1×1 convolution to increase the dimension of the feature map after preliminary processing, and increases the number of channels according to the expansion ratio to capture detailed features through the intermediate deep convolution layer, and the projection convolution layer uses 1×1 convolution to reduce the feature image channels to the target number of channels to fuse the feature map information.

5. A vehicle-pedestrian target detection method according to claim 4, characterized in that: The hybrid encoder includes several convolutional layers, an attention-based intra-scale feature interaction module and a cross-scale feature fusion module; the cross-scale feature fusion module includes a Bi-FPN feature fusion module and a RepC3 feature processing module.

6. A vehicle-pedestrian target detection method according to claim 5, characterized in that: The Bi-FPN feature fusion module normalizes the weights of the feature maps input by the backbone network, traverses different feature maps, assigns corresponding weights to them, and stacks the weighted feature maps to form a fused feature map.

7. A vehicle-pedestrian target detection method according to claim 6, characterized in that: In the Bi-FPN feature fusion module, fast normalization fusion is used as the feature fusion method. By assigning weights to different feature maps input to the backbone network, the influence of different feature maps on the output results of the vehicle pedestrian target detection model is judged. The weight calculation formula is: Among them, Output represents the output result, i represents the index of the input feature layer, j represents the set of feature layer indexes participating in the weighted calculation, and w i represents the weight of a specific feature layer i, w j Represents the weight of the feature layer in the feature layer set, I i Represents the feature map of a specific feature layer i, ∈ represents a small value set to 0.0001 to ensure the stability of the result.

8. A vehicle and pedestrian target detection device, characterized in that: include: The data processing module obtains and inputs the vehicle and pedestrian data set, and pre-processes the vehicle and pedestrian data set; Model building module, building a vehicle and pedestrian target detection model based on improved RT-DETR; The vehicle and pedestrian target detection model based on the improved RT-DETR is constructed, including the following improvements on the basis of the original RT-DETR model: 1) Add an inverted bottleneck UIB module to the original RT-DETR model backbone network; 2) In the original RT-DETR model hybrid encoder, the internal feature fusion network structure of the hybrid encoder is adjusted through a weighted bidirectional feature pyramid network; The training module uses the preprocessed vehicle and pedestrian data set to iteratively train the constructed vehicle and pedestrian target detection model to obtain a trained vehicle and pedestrian target detection model; The detection module detects vehicles and pedestrians based on the trained vehicle and pedestrian target detection model and outputs the detection results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Strip steel camouflage surface defect detection network and method

    CN117392111A

  • Real-time traffic target detection method for low-illuminance scene

    CN119007149A

  • Traffic small target detection method, device and equipment for high-speed driving vehicle and medium

    CN119152463A

  • Target detection model training method and apparatus, map generation method and apparatus, and device

    WO2024037552A1