Traffic sign target detection method and device in extreme weather and storage medium

The modified RT-DETR model with Ortho attention and high-level feature pyramid network enhances traffic sign detection in extreme weather by improving channel selection and feature fusion, achieving superior precision and efficiency for small targets.

CN120318795APending Publication Date: 2025-07-15SHANGHAI UNIVERSITY OF ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510300394.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In extreme weather, the existing traffic sign detection methods have unstable detection accuracy, insufficient flexibility and adaptability of feature fusion methods, especially poor small target detection performance.

Method used

Using the improved RT-DETR model, the Ortho attention mechanism and high-level screening feature pyramid network are introduced. By screening important channel information and multi-scale features fusion, the small object detection accuracy is improved, and the extreme weather environment training model is simulated through data augmentation.

Benefits of technology

In extreme weather, the detection accuracy of small targets is improved, the amount of model parameters is reduced, the detection efficiency and accuracy are ensured, and different lighting and weather conditions are adapted to different lighting and weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318795A_ABST
    Figure CN120318795A_ABST
Patent Text Reader

Abstract

The invention relates to a traffic sign target detection method and device in extreme weather and a storage medium. The method comprises the steps of obtaining a to-be-detected traffic sign image, performing image enhancement, inputting the to-be-detected traffic sign image after image enhancement into a trained traffic sign target detection model based on a BH-RT-DETR model, and obtaining a traffic sign target detection result; wherein the BH-RT-DETR model comprises a backbone network, a neck network and a head decoder which are connected in sequence, the backbone network comprises a Basic BlockOrtho module, the neck network adopts a high-level screening feature pyramid network, the Basic BlockOrtho module is used for screening channel information playing a key role in small target detection, and the head decoder is used for decoding the channel information; the high-level screening feature pyramid network is used for screening high-level features and fusing low-level feature information, and the small target is a target whose area occupied in the traffic sign image is smaller than a preset area threshold. Compared with the prior art, the method has the advantages that the detection efficiency is ensured while the detection precision of the small target under the severe weather condition is improved, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic sign target detection, and in particular to a traffic sign target detection method, device and storage medium under extreme weather conditions. Background Technique

[0002] With the rapid development of artificial intelligence technology, autonomous driving, as a new field, has received extensive attention and in-depth research. In an autonomous driving system, the detection and recognition of traffic signs is a crucial component, which can provide real-time road environment information for intelligent vehicles and assist them in making safe and reasonable decisions. However, extreme weather conditions (such as rainy days and foggy days) will seriously affect the imaging quality of in-vehicle sensors, reduce the contrast and clarity of traffic signs, and thus affect the detection accuracy and robustness. Therefore, accurate detection and recognition of traffic signs under extreme weather conditions remains a huge challenge.

[0003] Traditional traffic sign detection methods mainly rely on manually designed color and shape features, which usually achieve good detection results under clear weather conditions. However, in complex weather environments (such as rainy days or foggy days), they often exhibit high false detection rates and missed detection rates. To address these problems, deep learning technology has been widely applied to the field of traffic sign detection in recent years. Deep learning models can automatically learn multi-scale features and diverse features of traffic signs and have the ability to adapt to different lighting and weather conditions, showing superior performance compared to traditional methods. Methods based on deep learning are mainly divided into one-stage detection algorithms and two-stage detection algorithms. Two-stage detection algorithms such as RCNN and Faster RCNN first generate a large number of candidate regions and then identify the candidate regions, so the detection speed is slow, the algorithm structure is complex, and it is difficult to meet the real-time requirements of traffic sign detection. In contrast, one-stage detection algorithms are more efficient and faster in detection, mainly including the SSD algorithm, CenterNet algorithm, RetinaNet algorithm, and YOLO series algorithms, etc.

[0004] In terms of traffic sign detection, the literature "Traffic Sign Detection Based on Lightweight RT-DETR" discloses a lightweight RT-DETR traffic sign detection method. Based on RT-DETR, it uses ShuffNetv2 to replace the original Resnet backbone. Although this improvement reduces the computational cost and the number of parameters and improves the detection speed, it does not consider the missed detection and misdetection phenomena of small targets and traffic signs in extreme weather conditions; the deep fusion network disclosed in the literature "Traffic Sign Recognition Under Adverse Weather Using a Deep Fusion Network with Attention Mechanisms" combines traditional convolutional neural networks and attention mechanisms to improve the traffic sign detection ability in rainy weather, but the detection accuracy for small targets is still low. Therefore, the current traffic sign detection mainly has the following problems: First, the detection accuracy under different weather conditions is unstable, and the flexibility and adaptability of the feature fusion method are insufficient; second, the performance of small target detection is poor.

[0005] Literature 1: Wang Z X, Lei X M and Zhou SH S, Traffic Sign Detection Based on Lightweight RT-DETR, 2024 IEEE 2nd International Conference on Sensors, Electronics and Computer Engineering (ICSECE), Jinzhou, China, 2024, pp. 1593-1597.

[0006] Literature 2: LIU M, XU P, YANG S, et al. Traffic Sign Recognition Under Adverse Weather Using a Deep Fusion Network with Attention Mechanisms[C]. Proceedings of the 2024 IEEE / CVF International Conference on Computer Vision (ICCV), Paris, France. IEEE, 2024: 2043-2050. Summary of the Invention

[0007] The object of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a traffic sign target detection method, device and storage medium under extreme weather conditions, which can improve the detection accuracy of small targets under adverse weather conditions while reducing the number of model parameters and ensuring the detection efficiency.

[0008] The object of the present invention can be achieved by the following technical solutions:

[0009] According to the first aspect of the present invention, a traffic sign target detection method under extreme weather based on an improved RT-DETR is provided, including the following steps: S1, obtaining a traffic sign image to be detected and performing image enhancement; S2, inputting the traffic sign image to be detected after image enhancement into a traffic sign target detection model based on the BH-RT-DETR model that has been trained to obtain a traffic sign target detection result; wherein, the BH-RT-DETR model includes a backbone network, a neck network and a head decoder connected in sequence, the backbone network includes a BasicBlock_Ortho module, the neck network adopts a high-level filtering feature pyramid network, the BasicBlock_Ortho module is used to filter the channel information that plays a key role in small target detection, the high-level filtering feature pyramid network is used to filter high-level features and fuse low-level feature information, and the small target is a target whose area occupied in the traffic sign image is smaller than a preset area threshold.

[0010] As a preferred technical solution, the construction and training process of the traffic sign target detection model includes: S21, constructing a BH-RT-DETR model; S22, obtaining a traffic sign data set and performing data enhancement; S23, configuring a deep learning environment and initializing the parameters of the BH-RT-DETR model; S24, inputting the traffic sign data set after data enhancement into the initialized BH-RT-DETR model for training and testing.

[0011] As a preferred technical solution, the operation process of the BasicBlock_Ortho module includes: initializing an orthogonal filter; compressing the pre-obtained input feature map through the orthogonal filter to obtain a compressed feature vector; generating an attention vector by passing the compressed feature vector through a preset activation function; performing channel-by-channel multiplication on the compressed feature vector and the original input feature to obtain a weighted output feature; adding the weighted output feature to the original input feature to obtain a final output feature.

[0012] As a preferred technical solution, the high-level screening feature pyramid network at least includes a feature selection module, a feature fusion module, and a channel attention module. The running process includes: in the feature selection module, feature maps of different scales from the backbone network are respectively subjected to feature weighting processing through corresponding channel attention modules to obtain a plurality of weighted feature maps; the plurality of weighted feature maps are adaptively fused with feature information of different levels through the feature fusion module.

[0013] As a preferred technical solution, in the feature selection module, the specific process of weighting the feature maps of different scales from the backbone network includes: respectively performing global max pooling and global average pooling processing, adding the processing results, calculating each channel weight through an activation function to obtain the current weighted feature map; reducing the dimension of the current weighted feature map through a convolutional layer to obtain a plurality of finally weighted feature maps.

[0014] As a preferred technical solution, the running process of the feature fusion module includes: sequentially performing transposed convolution and bilinear interpolation processing on the high-level features to obtain processed high-level features, and the processed high-level features match the dimensions of the low-level features; using the channel attention module to weight the processed high-level features and fuse them with the low-level features to obtain multi-scale output features.

[0015] As a preferred technical solution, in the S22, the process of data augmentation specifically includes: screening out traffic signs with a quantity greater than a preset quantity threshold from the traffic sign dataset, and performing rain, fog, and snow operations on the original dataset.

[0016] As a preferred technical solution, in the S24, the average precision mean, the computational amount of the model, and the number of parameters of the model are used as evaluation indicators during model testing.

[0017] According to the second aspect of the present invention, there is provided a traffic sign target detection device based on an improved RT-DETR in extreme weather, including a memory, a processor, and a program stored in the memory. When the processor executes the program, the method described above is implemented.

[0018] According to the third aspect of the present invention, there is provided a storage medium, on which a program is stored, and when the program is executed, the method described above is implemented.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] 1. The present invention uses the BH-RT-DETR model to achieve traffic sign object detection. Based on the original RT-DETR model, the Ortho attention mechanism is introduced into the backbone network to improve the BasicBlock module, resulting in the BasicBlock_Ortho module. This module uses orthogonal filters to reduce feature redundancy and screen out important channel information that plays a key role in small object detection, thereby improving the detection accuracy of small objects. At the same time, the BH-RT-DETR model uses a high-level screening feature pyramid network to replace the cross-scale context feature mixer in the original model. By screening high-level features and fusing low-level feature information, the detection accuracy of the model for low-contrast and blurred objects in extreme weather is improved;

[0021] 2. The present invention uses a data augmentation method to simulate extreme weather environments and trains the BH-RT-DETR model using the traffic sign dataset after data augmentation, which can improve the model's recognition ability for traffic signs in these environments. Compared with other classic one-stage object detection algorithms, the model proposed in the present invention has higher detection accuracy for small objects in extreme weather and also has better detection effects on traffic sign pictures outside the dataset than traditional algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic flowchart of the method provided in Embodiment 1 of the present invention;

[0023] Figure 2 It is a schematic diagram of the network structure of the Otrho attention mechanism in Embodiment 1 of the present invention;

[0024] Figure 3 It is a schematic diagram of the structure of the BasicBlock_Ortho module in Embodiment 1 of the present invention;

[0025] Figure 4 It is a schematic diagram of the network structure of the HS-FPN in Embodiment 1 of the present invention;

[0026] Figure 5 It is a schematic diagram of the network structure of the SFF in Embodiment 1 of the present invention;

[0027] Figure 6 It is a structural diagram of the BH-RT-DETR network in Embodiment 1 of the present invention;

[0028] Figure 7 It is a diagram of the types of traffic signs selected according to the TT 100K dataset in Embodiment 1 of the present invention;

[0029] Figure 8 It is an example of a traffic sign image after data augmentation in Embodiment 1 of the present invention;

[0030] Among them: part (a) is an example of a traffic sign image before data augmentation, and part (b) is an example of a traffic sign image after data augmentation;

[0031] Figure 9 It is a comparison chart of traffic sign detection results in Embodiment 1 of the present invention;

[0032] Among them: part (a) is the detection result of the RT-DETR model, part (b) is the detection result of the YOLOv5s model, part (c) is the detection result of the YOLOv7 model, (d) is the detection result of the YOLOv8m model, and (e) is the detection result of the BH-RT-DETR model in this embodiment. Detailed implementation manners

[0033] In the context of the present invention: an important channel refers to a channel that plays a key role in small target detection and low-contrast target detection in object detection; high-level features refer to abstract and semantic information extracted from the deep layers of the network (such as the overall shape and category of the object); low-level features refer to simple local features extracted from the first few layers of the network, usually involving detailed information (such as edges and textures); large targets refer to targets that occupy a large area and have a clear shape in the traffic sign image, such as speed limit signs, stop signs, etc.; small targets refer to targets that occupy a small area and are easily occluded in the traffic sign image, such as traffic signs at a long distance, small warning signs, etc. Optionally, "large area" and "small area" can be determined according to a preset area threshold.

[0034] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.

[0035] Embodiment:

[0036] As Figure 1 shown, this embodiment provides a traffic sign object detection method based on BH-RT-DETR, and the method specifically includes the following implementation steps:

[0037] Step S1, obtain the traffic sign image to be detected and perform image augmentation. Optionally, the traffic sign image to be detected is real-time data or historical data.

[0038] Step S2, input the traffic sign image to be detected after image augmentation into a traffic sign object detection model based on the trained BH-RT-DETR model to obtain a traffic sign object detection result;

[0039] Among them, the construction and training process of the traffic sign object detection model includes:

[0040] Step S21: Construct the BH-RT-DETR model.

[0041] The BH-RT-DETR model (i.e., the BasicBlock_Ortho and HS-FPN based RT-DETR model) provided in this embodiment includes a backbone network, a neck network, and a head decoder connected in sequence. Among them, the backbone network includes at least the BasicBlock_Ortho module, and the neck network adopts the High-level Screening-feature Pyramid Networks (HS-FPN).

[0042] Specifically, the BH-RT-DETR model is based on the original RT-DETR model. The Orthogonal Attention mechanism (i.e., the Ortho Attention mechanism) is introduced into the backbone network to obtain the BasicBlock_Ortho module. The orthogonal filter is used to reduce feature redundancy, screen important channel information, and improve the detection accuracy of small targets. In addition, HS-FPN is used to replace the cross-scale context feature mixer in the original model. By screening high-level features and fusing low-level feature information, the detection accuracy of the model for low-contrast and blurred targets in extreme weather is improved.

[0043] Figure 2 The network structure of the Ortho Attention mechanism in this embodiment is shown. Figure 3 The network structure of the improved BasicBlock module is shown. Specifically, in this embodiment, the Ortho Attention mechanism is fused with the BasicBlock module in the backbone network to form the improved BasicBlock_Ortho module. In the figure, BasicBlock represents the basic block, Conv represents convolution, and Ortho Attention represents orthogonal attention.

[0044] The improvement principle analysis of the BasicBlock_Ortho module is as follows:

[0045] First, a group of filters is randomly initialized where H×W represents the size of the filter. This group of filters is orthonormalized using Gram-Schmidt to obtain the orthogonal filter Orthogonalization ensures that the feature information captured by each filter during compression is independent of each other, which is beneficial to retaining more feature information in different dimensions.

[0046] Subsequently, the input feature map of size H×W×C is compressed through the orthogonal filter to obtain the compressed feature vector:

[0047]

[0048] The size of the compressed feature vector is 1×1×C.

[0049] Next, the feature vector F ortho (X) generates the attention vector A(X) through an activation function:

[0050] A(X) = σ(W2δ(W1F ortho (X)))

[0051] where σ is the Sigmoid activation function, δ is the ReLU function, and W1 and W2 are learnable weight matrices.

[0052] Then, the compressed feature vector is multiplied channel by channel with the original input feature X to obtain the weighted output feature X weighted :

[0053] X weighted = A(X) ⊙ X

[0054] where ⊙ represents channel-by-channel multiplication.

[0055] Finally, the weighted output feature is added to the original input feature X to obtain the final output feature Y:

[0056] Y = X weighted + X

[0057] After improving the BasicBlock module and integrating the Ortho attention mechanism into the BasicBlock module, the model can dynamically adjust the importance of each channel and enhance the feature extraction ability for small targets.

[0058] In this embodiment, HS-FPN is used to replace the cross-scale context feature mixer in the original model. Figure 4 The network structure of HS-FPN is shown. Specifically, HS-FPN mainly includes the following three modules: a feature selection module, a feature fusion module, and a channel attention module.

[0059] In the feature selection module, S3, S4, and S5 represent the feature maps of three different scales in the backbone network. The feature maps of each scale are processed by a channel attention (CA) module for feature weighting. The specific processing process of the CA module is as follows: the feature map is successively processed by global maximum pooling (MAX Pool) and global average pooling (Avg Pool), and then the processing results of the two are added. Then, the weights of each channel are calculated through the Sigmoid activation function, and the weighted feature map is processed by a convolutional layer for dimensionality reduction, and finally three weighted feature maps P3, P4, and P5 are obtained.

[0060] In the feature fusion module (i.e., the SFF module), the high-resolution feature map contains rich detailed information but less semantic information. On the contrary, the low-resolution feature map has rich semantic information but relatively rough target localization. To solve this problem, the SFF module in the HS-FPN network selectively fuses high-resolution and low-resolution feature maps, enabling the model to adaptively fuse feature information at different levels, thereby improving the target recognition accuracy. The network structure of SFF is as Figure 5 shown.

[0061] In Figure 5 , first, the high-level feature f high containing rich semantic information is processed through transposed convolution and bilinear interpolation to match its dimension with the low-level feature f low containing rich detailed information, resulting in the feature map f att . Secondly, the CA module is used to weight the processed high-level feature f att and fuse it with the low-level feature f low to form the multi-scale output feature f out . By combining multi-scale features through HS-FPN, effective fusion of multi-scale features is achieved, thereby improving the detection accuracy of the model for low-contrast and blurred targets under extreme weather conditions.

[0062] Combining the aforementioned improvements, Figure 6 shows one of the overall network structures of the BH-RT-DETR model. Specifically:

[0063] The BH-RT-DETR model includes a backbone network, a neck network, and a head decoder module connected in sequence. At the input end, the input image is first uniformly adjusted to 320×320 pixels.

[0064] The backbone network includes a convolutional layer and the BasicBlock_Ortho module, and the neck network is the high-level screening feature pyramid network (HS-FPN), which utilizes multi-scale feature information to further enhance the network's detection ability for targets of different scales. In the improved RT-DETR model, the BasicBlock_Ortho module can effectively enhance the discrimination ability between feature channels by introducing the Ortho attention mechanism, enabling the network to pay more attention to key target features while suppressing unimportant background interference. HS-FPN adaptively selects and fuses feature maps at different levels through the selective feature fusion module and the channel attention mechanism, thereby improving the accuracy of the network in detecting small targets and targets in complex backgrounds.

[0065] Step S22, obtain the traffic sign dataset and perform data augmentation.

[0066] In this step, the process of data augmentation specifically includes: screening traffic signs with a quantity greater than a preset quantity threshold from the traffic sign dataset to obtain a traffic sign detection dataset. Optionally, the traffic sign dataset uses the TT 100K traffic sign image dataset, selects the top 15 traffic sign pictures in terms of quantity to construct the dataset, and the selected traffic signs are as Figure 7 shown. Specifically, 15 types of traffic signs with a relatively large quantity are screened out from the traffic sign images as the dataset for data augmentation. There are a total of 9050 pictures. The imgaug library is used to perform data augmentation on the 9050 pictures. For example, iaa.Rain() is used to add raindrops, iaa.GaussianBlur() is used to add different degrees of blur effects to simulate haze, and iaa.Snow() is used to introduce snow noise to simulate snowy days. The enhanced images are as Figure 8 shown. Through data augmentation, a total of 13239 pictures are finally obtained, and they are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 for model training.

[0067] Step S23: Configure the deep learning environment and initialize the parameters of the BH-RT-DETR model.

[0068] In this embodiment, the operating system used in the experiment is Windows 11 64-bit, the processor is Intel(R) Core(TM) i9-13980H, the graphics card is NVIDIA GeForce RTX 4060 Laptop, the RAM size is 32G, the Python version used in the experiment is 3.9.7, the CUDA version is 11.6, the number of training epochs is set to 100, the batch size is 32, the Stochastic Gradient Descent (SGD) optimizer is adopted, and the initial learning rate is 1e-2.

[0069] Step S24: Input the traffic sign dataset after data augmentation into the initialized BH-RT-DETR model for training and testing.

[0070] Specifically, the training set divided from the traffic sign detection dataset in step S23 is input into the initialized BH-RT-DETR model for training. During the training process, the validation set is used for validation, and the test set is used to test the trained model.

[0071] This embodiment uses the mean average precision, the computational complexity of the model, and the number of parameters as evaluation indicators.

[0072] Among them, the calculation method of the mean average precision is:

[0073]

[0074] Where R is the recall rate, P is the precision, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, FP is the number of incorrectly predicted positive samples, AP is the area under the precision-recall curve, mAP is the mean of the APs for all detection classes, and N is the number of detection classes.

[0075] Using steps S21 - S24, a trained traffic sign object detection model based on BH - RT - DETR is obtained. On the basis of the original RT - DETR, this model combines the BasicBlock module of RT - DETR with the Ortho attention mechanism to obtain the BasicBlock_Ortho module. At the same time, the cross - scale context feature mixer in the original network is replaced with HS - FPN to improve the detection accuracy of traffic signs under extreme weather conditions, and the number of parameters of the BH - RT - DETR model is also lower than that of RT - DETR. Based on this, the real - time traffic sign image to be detected or the historical data of the traffic sign image to be detected is input into the trained improved model to obtain the corresponding traffic sign object detection result.

[0076] Exemplarily, when evaluating the trained BH - RT - DETR model, it can be found that the mAP of this method reaches 87.84% on the traffic sign detection dataset, the computational cost is 53.4 GFLOPs, and the number of model parameters is 18.22M, meeting the requirements of high precision and low number of parameters for traffic sign detection. Figure 9 The figure shows a comparison chart of the detection results obtained by using other existing methods and the method provided in this embodiment. Figure 9 As can be seen from the results, in the first group of pictures, the traffic signs are small in size, especially the targets at a long distance. YOLOv5s and YOLOv7 have false detection and missed detection phenomena in small target detection, while YOLOv8m and RT - DETR have improved this problem to a certain extent, but their confidence levels are still lower than that of BH - RT - DETR. In the second group of pictures, some targets are occluded, the background is relatively complex, and there is a lot of interference information. YOLOv5s, YOLOv7, and YOLOv8m all have missed detection phenomena, while the RT - DETR model has over - detection problems. In contrast, BH - RT - DETR can accurately identify the targets. In the third group of pictures, there are multiple traffic signs. YOLOv5s, YOLOv7, and YOLOv8m all have missed detections to varying degrees. Although RT - DETR can accurately identify all traffic signs, its confidence level is lower than that of BH - RT - DETR. The high detection accuracy and confidence level of BH - RT - DETR benefit from its adopted Ortho attention mechanism and HS - FPN feature fusion network.

[0077] Embodiment 2:

[0078] This embodiment provides a traffic sign target detection device based on improved RT-DETR in extreme weather, including a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the method in Embodiment 1. Exemplarily, the program in the memory includes an image acquisition module and a target detection module during operation. Among them, the image acquisition module is used to acquire the traffic sign image to be detected and perform image enhancement, and the target detection module is used to input the traffic sign image to be detected after image enhancement into the traffic sign target detection model based on the BH-RT-DETR model that has been trained to obtain the traffic sign target detection result. The basic structure and operation process of the BH-RT-DETR model are basically the same as the implementation process of the method in Embodiment 1, and will not be elaborated here.

[0079] Exemplarily, the processor in the device includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks. The processing unit executes the various methods and processes described above, such as one or more steps of the foregoing method. For example, in some embodiments, one or more steps of the foregoing method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the foregoing method can be executed. Alternatively, in other embodiments, the CPU can be configured to execute one or more steps of the foregoing method in any other appropriate manner (for example, by means of firmware). The functions described above can be at least partially executed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate array (FPGA), application specific integrated circuit (ASIC), application specific standard product (ASSP), system on chip (SOC), complex programmable logic device (CPLD), and so on.

[0080] Further, this embodiment also provides a storage medium, on which a program is stored, and when the program is executed, one or more steps of the method in Embodiment 1 are implemented. The storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined in the present invention, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0081] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A traffic sign target detection method based on improved RT-DETR under extreme weather, characterized in that, It includes the following steps: S1. Obtain the traffic sign image to be detected and perform image enhancement; S2. Input the traffic sign image to be detected after image enhancement into the traffic sign object detection model based on the BH-RT-DETR model that has been trained to obtain the traffic sign object detection result; Among them, the BH-RT-DETR model includes a backbone network, a neck network, and a head decoder connected in sequence. The backbone network includes a BasicBlock_Ortho module. The neck network adopts a high-level screening feature pyramid network. The BasicBlock_Ortho module is used to screen the channel information that plays a key role in small object detection. The high-level screening feature pyramid network is used to screen high-level features and fuse low-level feature information. The small object is an object whose area in the traffic sign image is smaller than the preset area threshold.

2. The traffic sign target detection method based on improved RT-DETR under extreme weather according to claim 1, wherein The construction and training process of the traffic sign object detection model includes: S21. Construct the BH-RT-DETR model; S22. Obtain the traffic sign data set and perform data augmentation; S23. Configure the deep learning environment and initialize the parameters of the BH-RT-DETR model; S24. Input the traffic sign data set after data augmentation into the initialized BH-RT-DETR model for training and testing.

3. The method for traffic sign target detection in extreme weather based on improved RT-DETR according to claim 1, characterized in that, The operation process of the BasicBlock_Ortho module includes: Initialize the orthogonal filter; Compress the pre-obtained input feature map through the orthogonal filter to obtain a compressed feature vector; Generate an attention vector by passing the compressed feature vector through a preset activation function; Perform channel-wise multiplication on the compressed feature vector and the original input feature to obtain a weighted output feature; Add the weighted output feature to the original input feature to obtain the final output feature.

4. The traffic sign target detection method based on improved RT-DETR under extreme weather according to claim 1, characterized in that The high-level screening feature pyramid network at least includes a feature selection module, a feature fusion module, and a channel attention module. The operation process includes: In the feature selection module, feature maps of different scales from the backbone network are respectively weighted through corresponding channel attention modules to obtain multiple weighted feature maps; The multiple weighted feature maps are adaptively fused through the feature fusion module for different-level feature information.

5. The method for traffic sign target detection in extreme weather based on improved RT-DETR according to claim 4, wherein, In the feature selection module, the specific process of weighting the feature maps of different scales from the backbone network includes: Perform global max pooling and global average pooling respectively, add the processing results, and calculate the weight of each channel through an activation function to obtain the current weighted feature map; Reduce the dimension of the current weighted feature map through a convolutional layer to obtain multiple finally weighted feature maps.

6. The traffic sign target detection method based on improved RT-DETR in extreme weather according to claim 4, characterized in that The operation process of the feature fusion module includes: Perform transposed convolution and bilinear interpolation on the high-level feature in sequence to obtain the processed high-level feature, and the processed high-level feature matches the dimension of the low-level feature; Weight the processed high-level feature by using the channel attention module and fuse it with the low-level feature to obtain a multi-scale output feature.

7. The method for traffic sign target detection in extreme weather based on improved RT-DETR according to claim 1, characterized in that, In the S22, the process of data augmentation specifically includes: screening out traffic signs with a quantity greater than a preset quantity threshold from the traffic sign dataset, and performing rain, fog, and snow operations on the original dataset.

8. The method for traffic sign target detection in extreme weather based on improved RT-DETR according to claim 1, wherein, In the S24, the mean average precision, the computational amount of the model, and the number of parameters are used as evaluation indicators during model testing.

9. A traffic sign target detection device based on improved RT-DETR in extreme weather, comprising a memory, a processor, and a program stored in the memory, characterized in that When the processor executes the program, the method described in any one of claims 1-8 is implemented.

10. A storage medium, on which a program is stored, characterized in that, When the program is executed, the method described in any one of claims 1-8 is implemented.