A traffic target detection method and system based on improved RT-DETR

CN122336263BActive Publication Date: 2026-10-09NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610770687.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-10-09
Estimated Expiration
2046-06-01

AI Technical Summary

Technical Problem

然而,在应对真实世界中高度动态、复杂多变的交通场景时,现有RT-DETR模型仍面临若干挑战:首先,交通目标尺度差异巨大,模型需要更强大的多尺度特征表征与融合能力;其次,密集车流、部分遮挡等情况要求模型具备更精细的局部语义理解和全局上下文推理能力;再者,光照变化、天气因素导致的图像质量下降,要求模型的前端特征提取网络具备更强的鲁棒性与动态感知能力

Benefits of technology

[0018] Compared with existing technologies, this invention has the following advantages: The traffic target detection method proposed in this invention, based on an improved RT-DETR, effectively enhances the environmental perception performance in complex traffic scenarios through targeted collaborative improvements to three model structures: First, the Dynamic Perception Enhancement Module (DPB) embedded in the backbone network, through a unique parallel branch and cascade enhancement design, integrates original shallow features, intermediate semantics, and deep abstract features, significantly enhancing the model's feature representation ability and robustness for multi-scale targets, especially small targets and severely occluded targets; Second, the original module is replaced with a dual-scale collaborative attention module (DSCAB), utilizing collaborative channel mapping branches and spatial attention branches, and introducing dilated convolution and overlapping window extraction mechanisms, enabling the model to simultaneously and efficiently model long-range channel dependencies and extremely fine local spatial context, greatly improving the quality of multi-feature interaction extraction; Third, the Global Context Convolution Module (GCC3) replaces some standard convolutions, and its unique gating branch mechanism realizes dynamic feature selection and channel fusion enhancement, strengthening the model's ability to capture key information in traffic images without introducing redundant computational burden. Therefore, while maintaining the real-time performance of the original model's inference, this invention has made a comprehensive breakthrough in the bottleneck of multi-scale target detection under complex autonomous driving conditions, and has extremely high accuracy and practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336263B_ABST
    Figure CN122336263B_ABST
Patent Text Reader

Abstract

The application discloses a traffic target detection method and device based on an improved RT-DETR and a medium, and belongs to the technical field of automatic driving and auxiliary driving. The method comprises the following steps: acquiring an initial driving video image and dividing the initial driving video image into a training image set and a to-be-detected image set; performing structural adjustment on an initial RT-DETR model to construct an improved RT-DETR model; the structural adjustment comprises embedding a dynamic perception enhancement module (DPB) in a backbone network to enhance multi-scale extraction capability, replacing an adaptive feature interaction module with a double-scale collaborative attention module (DSCAB) to realize channel and spatial attention collaboration, and replacing part of standard convolution blocks with a global context convolution module (GCC3) to realize dynamic feature enhancement; iteratively optimizing and training the improved RT-DETR model by using the training image set; and inputting the to-be-detected image set into the model to output a detection result. The application solves the problem of low detection precision of multi-scale, small targets and occluded targets in a complex traffic scene, and significantly improves the accuracy and robustness of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving and assisted driving technology, and in particular to a traffic target detection method and system based on an improved RT-DETR. Background Technology

[0002] In the wave of automotive intelligence, autonomous driving and assisted driving technologies are core to improving traffic safety and efficiency. Accurate and real-time perception of the traffic environment, especially the detection and identification of various targets such as vehicles, pedestrians, and traffic signs, is the foundation and prerequisite for realizing high-level autonomous driving functions.

[0003] Currently, deep learning-based object detection algorithms mainly fall into two major technical routes: two-stage detectors, represented by Faster R-CNN, which have high detection accuracy but slow inference speed; and single-stage detectors, represented by YOLO and SSD, which have significant advantages in speed, but their detection accuracy in complex scenes, especially their ability to identify small and occluded targets, often needs improvement. For autonomous driving systems that require high real-time performance, single-stage detectors are usually a more practical choice, but this also places higher demands on their accuracy and robustness.

[0004] In recent years, detection models based on the Transformer architecture, such as the DETR series, have received widespread attention due to their powerful global modeling capabilities and end-to-end detection processes. RT-DETR, as a representative improvement for real-time applications, has achieved a good balance between speed and accuracy. However, existing RT-DETR models still face several challenges when dealing with highly dynamic and complex traffic scenarios in the real world: First, the scale of traffic targets varies greatly, requiring models to have stronger multi-scale feature representation and fusion capabilities; second, dense traffic flow and partial occlusion require models to have more refined local semantic understanding and global contextual reasoning capabilities; third, image quality degradation caused by changes in lighting and weather factors requires the model's front-end feature extraction network to have stronger robustness and dynamic perception capabilities. Existing standard convolutional and attention mechanisms still have room for improvement in efficiently capturing features of such complex and dynamic traffic scenarios.

[0005] Therefore, how to optimize the existing RT-DETR model in a targeted manner to address the specific characteristics of traffic scenarios, and significantly improve its detection accuracy and robustness for multi-scale targets, especially small targets and occluded targets, under complex and harsh conditions while maintaining its real-time performance, has become a key technical problem that urgently needs to be solved in this field. Summary of the Invention

[0006] Based on this, this application proposes a traffic target detection method and system based on improved RT-DETR, aiming to improve the accuracy and efficiency of traffic target detection in the fields of autonomous driving and assisted driving technology.

[0007] In a first aspect, the present invention provides a traffic target detection method based on an improved RT-DETR, comprising the following steps: Acquire initial driving video images and divide the initial driving video images into a training image set and a test image set; An improved RT-DETR model is constructed; the improved RT-DETR model is obtained by structurally adjusting the existing RT-DETR model; the specific process of structural adjustment includes: embedding a dynamic perception enhancement module into the backbone network of the existing RT-DETR model; replacing the adaptive feature interaction module in the existing RT-DETR model with a dual-scale collaborative attention module; and replacing some standard convolutional blocks in the existing RT-DETR model with global context convolutional modules. The training image set is input into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; The set of images to be tested is input into the trained traffic target detection model for feature extraction and target detection, and the traffic target detection results are output.

[0008] As an optional implementation of the first aspect of this application, in the step of embedding a dynamic perception enhancement module into the backbone network of the existing RT-DETR model, the enhanced backbone features are output through the dynamic perception enhancement module. The specific process includes: acquiring a first input feature map transmitted by the backbone network; adjusting the number of channels in the first input feature map using a main channel adjustment layer with a first specific-size convolutional kernel to obtain a target feature map; uniformly dividing the target feature map along the channel dimension into a first feature subset and a second feature subset; retaining the first feature subset as a direct branch feature; inputting the second feature subset as a starting feature into a processing branch; the processing branch contains a preset number of high-frequency perception enhancement blocks that are sequentially connected; using each of the high-frequency perception enhancement blocks in the processing branch to sequentially perform deep feature extraction on the starting feature to obtain local intermediate features corresponding to each high-frequency perception enhancement block; concatenating the direct branch features, the starting feature, and each of the local intermediate features along the channel dimension to obtain multi-dimensional concatenated features; adjusting the number of channels in the multi-dimensional concatenated features using a main feature fusion layer with a second specific-size convolutional kernel to output the enhanced backbone features corresponding to the dynamic perception enhancement module.

[0009] As an optional implementation of the first aspect of this application, the process of sequentially performing deep feature extraction on the initial features using each of the high-frequency sensing enhancement blocks in the processing branch to obtain the local intermediate features corresponding to each of the high-frequency sensing enhancement blocks includes: obtaining the current input features entering the current high-frequency sensing enhancement block, and uniformly dividing the current input features into local feature parts and high-frequency feature parts along the channel dimension; inputting the local feature parts into the local feature extraction branch, processing them using a first convolutional layer with a third specific-size convolutional kernel and a first nonlinear activation function, and outputting the first local extracted features; inputting the high-frequency feature parts into the high-frequency enhancement branch, and sequentially passing them through a max pooling layer of a set size, a second convolutional layer with a fourth specific-size convolutional kernel, and a third high-frequency feature extraction layer. The first high-frequency enhanced feature is output by processing two nonlinear activation functions; the first local extracted feature and the first high-frequency enhanced feature are added element-wise to generate a first additive interactive feature, and the first local extracted feature and the first high-frequency enhanced feature are multiplied element-wise to generate a first multiplicative interactive feature; the first local extracted feature, the first high-frequency enhanced feature, the first additive interactive feature, and the first multiplicative interactive feature are concatenated along the channel dimension to obtain a composite interactive feature; the composite interactive feature is fused and mapped using a third convolutional layer with a fifth specific-size convolutional kernel, and the fused and mapped feature is added to the current input feature by residual addition to output the local intermediate feature corresponding to the current high-frequency perceptual enhancement block.

[0010] As an optional implementation of the first aspect of this application, in the step of replacing the adaptive feature interaction module in the existing RT-DETR model with a dual-scale collaborative attention module, the collaborative interaction features are output through the dual-scale collaborative attention module. Specifically, the process includes: obtaining a second input feature map that enters the dual-scale collaborative attention module; inputting the second input feature map to a channel attention branch for long-range channel dependency extraction, and outputting channel enhancement features; the channel attention branch sequentially includes a channel-mapped attention module and a first feedforward network module; inputting the channel enhancement features to a spatial attention branch for local spatial context extraction, and outputting the collaborative interaction features corresponding to the dual-scale collaborative attention module; the spatial attention branch sequentially includes a local spatial attention module and a second feedforward network module.

[0011] As an optional implementation of the first aspect of this application, the process of inputting the second input feature map to a channel attention branch for long-range channel dependency extraction and outputting channel-enhanced features includes: performing a first-layer normalization process on the second input feature map through the channel-mapped attention module, and mapping it into a channel query vector, a channel key vector, and a channel value vector using a first projection convolutional layer of size 1×1; inputting the channel query vector, the channel key vector, and the channel value vector into a preset dual-depth separable convolutional block for feature enhancement to obtain an enhanced query vector, an enhanced key vector, and an enhanced value vector; dividing the enhanced query vector, the enhanced key vector, and the enhanced value vector into several attention heads; and for each attention head, calculating its corresponding enhanced query vector and the channel key vector. The matrix multiplication of the enhancement key vectors is used, and channel attention weights are calculated in conjunction with learnable temperature parameters. The channel attention weights are then used to perform weighted aggregation calculations on the enhancement value vectors to obtain channel aggregation features. After feature reconstruction and processing by a second projection convolutional layer of size 1×1, these features are added to the second input feature map using a first residual, outputting a first residual feature. This first residual feature is then input to the first feedforward network module, where a first dimensionality reduction and expansion convolutional layer with an expansion rate of 2 and a size of 3×3 is used to extract the spatial context of the first residual feature. After passing through a third nonlinear activation function, a second dimensionality reduction and expansion convolutional layer with an expansion rate of 2 and a size of 3×3 is used for dimensionality reduction processing. The result is then added to the first residual feature using a second residual, outputting the channel enhancement feature corresponding to the channel attention branch.

[0012] As an optional implementation of the first aspect of this application, the process of inputting the channel enhancement features into the spatial attention branch for local spatial context extraction and outputting the collaborative interaction features corresponding to the dual-scale collaborative attention module includes: performing a second-layer normalization process on the channel enhancement features through the local spatial attention module, and mapping them into a spatial query vector, a spatial key vector, and a spatial value vector using a third projection convolutional layer of size 1×1; inputting the spatial query vector, the spatial key vector, and the spatial value vector into a preset dual-depth separable convolutional block for feature enhancement to obtain a reset query vector, a reset key vector, and a reset value vector; reshaping the reset query vector into a regular non-overlapping window format query feature; and comparing the features based on the set window size and overlap ratio. A sliding window operation is performed on the reset key vector and the reset value vector to extract sliding key features and sliding value features of the overlapping context region. The spatial attention weight between the regular non-overlapping window format query feature and the sliding key feature is calculated by combining the relative position encoding information and the fixed position bias matrix. The sliding value feature is aggregated using the spatial attention weight to obtain local aggregated features. The local aggregated features are projected using a fourth projection convolutional layer of size 1×1, and the third residual is added to the channel enhancement feature to output the second residual feature. The second residual feature is input to the second feedforward network module for dimensionality upscaling and downscaling extraction, and the result is added to the second residual feature to output the collaborative interaction feature corresponding to the spatial attention branch.

[0013] As an optional implementation of the first aspect of this application, the process of feature enhancement using the preset dual-depth separable convolutional block includes: obtaining the feature vector to be enhanced; extracting spatial features from the feature vector to be enhanced using a first depth-separable convolutional layer of size 5×5 to obtain basic spatial features; inputting the basic spatial features into a first linear transformation branch and a second normalization transformation branch set in parallel; processing the first transformation feature using a fifth projection convolutional layer of size 1×1 in the first linear transformation branch to output a first transformation feature, and processing the second transformation feature using a sixth projection convolutional layer of size 1×1 in the second normalization transformation branch to output a second transformation feature; activating the first transformation feature using a fourth nonlinear activation function, and then multiplying it element-wise with the second transformation feature to generate a feature fusion base term; performing channel fusion on the feature fusion base term using a seventh projection convolutional layer of size 1×1, and further extracting spatial context using a second depth-separable convolutional layer of size 5×5 to obtain depth context features; adding the depth context features to the feature vector to be enhanced using a fifth residual, and outputting the enhanced feature vector corresponding to the preset dual-depth separable convolutional block.

[0014] As an optional implementation of the first aspect of this application, in the step of replacing some standard convolutional blocks in the existing RT-DETR model with a global context convolutional module, the global context enhancement features are output through the global context convolutional module. Specifically, the process includes: obtaining a third input feature map entering the global context convolutional module; performing dimensionality transformation on the third input feature map using two parallel 1×1 channel dimensionality reduction convolutional layers, outputting shortcut connection features and depth transformation initial features; inputting the depth transformation initial features into a sequence of gated convolutional blocks containing the target number in sequence; for any gated convolutional block in the sequence: uniformly dividing the received block input features along the channel dimension into a first channel part and a second channel part; inputting the first channel part into the value extraction branch, sequentially extracting spatial features through a first pointwise convolutional layer of size 1×1 and a grouped depth convolutional layer of size 3×3. The system first extracts spatial features and outputs the first spatial features. The second channel portion is input to a gated branch and sequentially passed through a 1×1 second pointwise convolutional layer and a logistic activation layer to generate spatial attention weights. Element-wise multiplication and addition operations are performed on the first spatial features and the spatial attention weights simultaneously to generate second multiplicative interactive features and second additive interactive features. The first spatial features, the spatial attention weights, the second multiplicative interactive features, and the second additive interactive features are concatenated along the channel dimension, normalized, and then mapped using a 1×1 fusion projection convolutional layer to output the depth transformation stage features corresponding to the current gated convolutional block. The final state transformation features output by the last gated convolutional block in the sequence are obtained, and the final state transformation features are added element-wise with the shortcut connection features to form a residual connection, outputting the global context enhancement features corresponding to the global context convolutional module.

[0015] Secondly, embodiments of this application provide a traffic target detection system based on an improved RT-DETR, comprising: An image acquisition module is used to acquire initial driving video images and divide the initial driving video images into a training image set and a test image set; A model improvement module is used to construct an improved RT-DETR model; the improved RT-DETR model is obtained by structurally adjusting the initial RT-DETR model; the model improvement module includes: A module embedding unit is used to embed a dynamic perception enhancement module into the backbone network of the initial RT-DETR model; The first module replacement unit is used to replace the adaptive feature interaction module in the initial RT-DETR model with a dual-scale collaborative attention module. The second module replacement unit is used to replace some standard convolutional blocks in the initial RT-DETR model with global context convolutional modules. The model training module is used to input the training image set into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; The target detection module is used to input the set of images to be tested into the trained traffic target detection model for feature extraction and target detection, and output the traffic target detection results.

[0016] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the method described in the first aspect.

[0017] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0018] Compared with existing technologies, this invention has the following advantages: The traffic target detection method proposed in this invention, based on an improved RT-DETR, effectively enhances the environmental perception performance in complex traffic scenarios through targeted collaborative improvements to three model structures: First, the Dynamic Perception Enhancement Module (DPB) embedded in the backbone network, through a unique parallel branch and cascade enhancement design, integrates original shallow features, intermediate semantics, and deep abstract features, significantly enhancing the model's feature representation ability and robustness for multi-scale targets, especially small targets and severely occluded targets; Second, the original module is replaced with a dual-scale collaborative attention module (DSCAB), utilizing collaborative channel mapping branches and spatial attention branches, and introducing dilated convolution and overlapping window extraction mechanisms, enabling the model to simultaneously and efficiently model long-range channel dependencies and extremely fine local spatial context, greatly improving the quality of multi-feature interaction extraction; Third, the Global Context Convolution Module (GCC3) replaces some standard convolutions, and its unique gating branch mechanism realizes dynamic feature selection and channel fusion enhancement, strengthening the model's ability to capture key information in traffic images without introducing redundant computational burden. Therefore, while maintaining the real-time performance of the original model's inference, this invention has made a comprehensive breakthrough in the bottleneck of multi-scale target detection under complex autonomous driving conditions, and has extremely high accuracy and practical application value. Attached Figure Description

[0019] Figure 1 The flowchart is a traffic target detection method based on an improved RT-DETR proposed in the first embodiment of this application; Figure 2This is a structural diagram of an improved RT-DETR traffic target detection model proposed in the first embodiment of this application; Figure 3 This is a structural diagram of the Dynamic Perception Enhancement Module (DPB) proposed in the first embodiment of this application; Figure 4 This is a structural diagram of the high-frequency sensing enhancement block (HAPEB) in the dynamic sensing enhancement module (DPB) proposed in the first embodiment of this application; Figure 5 This is a structural diagram of the dual-scale collaborative attention module (DSCAB) proposed in the first embodiment of this application; Figure 6 This is a structural diagram of the channel-mapped attention module (DMDTA) in the dual-scale collaborative attention module (DSCAB) proposed in the first embodiment of this application; Figure 7 This is a structural diagram of the feedforward network module (DFFN) in the dual-scale collaborative attention module (DSCAB) proposed in the first embodiment of this application; Figure 8 This is a structural diagram of the Local Spatial Attention Module (DOCA) in the Dual-Scale Collaborative Attention Module (DSCAB) proposed in the first embodiment of this application; Figure 9 This is a structural diagram of the dual-depth separable convolutional block (DDB) in the dual-scale collaborative attention module (DSCAB) proposed in the first embodiment of this application; Figure 10 This is a structural diagram of the Global Context Convolutional Module (GCC3) proposed in the first embodiment of this application; Figure 11 This is a structural diagram of the gated convolutional block (GCB) in the Global Context Convolutional Module (GCC3) proposed in the first embodiment of this application; Figure 12 This is a schematic diagram of the structure of a traffic target detection system based on an improved RT-DETR proposed in the second embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0022] To illustrate the technical solution described in this application, specific embodiments are provided below.

[0023] Example 1 Please see Figure 1 This is a flowchart of a traffic target detection method based on an improved RT-DETR proposed in the first embodiment of this application. Figure 2 This is a structural diagram of an improved RT-DETR traffic target detection model proposed in the first embodiment of this application. The proposed method includes: S01: Acquire initial driving video images and divide the initial driving video images into a training image set and a test image set.

[0024] For example, this embodiment selects images from a publicly available driving video dataset. The selected images need to cover diverse driving scenarios, various weather conditions, and different time periods to ensure the representativeness and comprehensiveness of the data. Subsequently, the original label files in the dataset are uniformly converted to the YOLO format suitable for RT-DETR model training and evaluation. The entire image dataset is divided into three parts according to a 7:1:2 ratio: training image set, validation image set, and test image set. This division strategy aims to ensure the balance and independence of data distribution among the subsets, thereby effectively preventing overfitting during model training and providing a reliable basis for model performance evaluation.

[0025] In terms of data composition, the dataset covers six typical weather conditions: sunny, cloudy, overcast, rainy, snowy, and foggy, to enhance the model's adaptability to different weather environments. Regarding scene types, it encompasses six common driving scenarios: residential areas, highways, city streets, parking lots, gas stations, and tunnels, thereby improving the model's generalization performance in real-world, complex road conditions. In terms of time, the dataset includes three key time periods: dawn or dusk, daytime, and nighttime, with a focus on daytime and nighttime samples, further enriching the environmental diversity of the data.

[0026] To construct high-quality training samples, we performed rigorous data cleaning on the training image set, removing samples with annotation errors, blurry images, and low information value. The test image set remained independent, serving as an objective benchmark for the final evaluation of the model's generalization ability.

[0027] During the model training phase, to further enhance its generalization ability and robustness, we applied various data augmentation techniques to the training image set. These included random horizontal and vertical image flipping, random scaling and cropping, random color jittering across hue, saturation, and brightness dimensions, adding Gaussian noise, and simulating different weather visual effects such as raindrops and fog. These augmentation strategies effectively increased the diversity of the training data, realistically simulating complex image changes that might be encountered in the real world, thereby driving the model to learn more stable and general feature representations.

[0028] S02: Construct an improved RT-DETR model, which is obtained by structurally adjusting the existing RT-DETR model.

[0029] Specifically, S02 includes S021, S022, and S023, wherein: S021: Embed a dynamic perception enhancement module (DPB) into the backbone network of the existing RT-DETR model. Please see Figure 3 This is a structural diagram of the Dynamic Perceptual Enhancement Module (DPB) proposed in this application. The DPB module aims to enhance the backbone network's feature extraction capabilities for multi-scale targets, especially small targets and occluded targets. Its image processing includes: the input feature map first undergoes initial channel adjustment through a 1×1 convolutional layer. Subsequently, the feature map is uniformly divided into two parts: the first part, as the "direct branch," is completely preserved, maintaining the integrity of the original features; the second part, as the starting feature of the "processing branch," is input into a sequence of n cascaded High-Frequency Perceptual Enhancement Blocks (HAPEBs) for deep feature extraction and enhancement. After processing by n HAPEBs, the original features of the direct branch, the starting features of the processing branch, and the local intermediate features of each HAPEB—a total of n+2 feature maps—are concatenated along the channel dimension to obtain multi-dimensional concatenated features. This design integrates shallow details, mid-level semantics, and deep abstract features. Finally, a convolutional layer with a 1×1 kernel is used to adjust the number of channels in the concatenated high-dimensional features to the target output number, resulting in the final output of the DPB module—enhanced backbone features.

[0030] For details, please refer to Figure 4This is a structural diagram of the High-Frequency Perception Enhancement Block (HAPEB) in the Dynamic Perception Enhancement Module (DPB). The HAPEB employs a dual-branch collaborative enhancement design to achieve deep fusion of local feature extraction and high-frequency information enhancement. Specifically, the input features are first evenly divided into two halves along the channel dimension. The first half enters the "Local Feature Extraction Branch," which sequentially passes through a 3×3 convolutional layer and a GELU activation function to extract local spatial features. The second half enters the "High-Frequency Enhancement Branch," which sequentially passes through a 3×3 max-pooling convolutional layer with a 1×1 kernel and a GELU activation function to enhance high-frequency information such as edges and textures. Subsequently, the output of the Local Feature Extraction Branch and the output of the High-Frequency Enhancement Branch are element-wise added and multiplied to generate two complementary feature interaction forms. Next, the outputs of the two original branches, the addition result, and the multiplication result are concatenated along the channel dimension, fused and channel-adjusted through a 1×1 convolutional layer, and then added to the original input of the module through a residual connection to form the final output—the local intermediate feature. This design not only enhances the module's ability to represent multi-scale features, but also promotes information flow and training stability through residual connections.

[0031] S022: Replace the Adaptive Feature Interaction Module (AIFI) in the existing RT-DETR model with the Dual-Scale Cooperative Attention Module (DSCAB); Please see Figure 5 This is a structural diagram of the Dual-Scale Cooperative Attention Module (DSCAB) proposed in this application. The DSCAB aims to efficiently model long-range dependencies on fine local context through a collaborative channel and spatial attention mechanism. Its image processing is a sequential, residual-connected four-stage process: the input feature map first passes through a channel attention branch for long-range channel dependency extraction, outputting channel-enhanced features. This channel attention branch includes a channel-mapped-from-attention module (DMDTA) and a feedforward network module (DFFN); then, it passes through a spatial attention branch for local spatial context extraction, containing a Local Spatial Attention Module (DOCA) and another feedforward network module (DFFN); finally, it outputs the collaborative interaction features corresponding to the dual-scale cooperative attention module.

[0032] For details, please refer to Figure 6This is the Channel Mapped From Attention Module (DMDTA) within the Dual Scale Collaborative Attention Module (DSCAB). Input features are first normalized. Then, a 1×1 projective convolutional layer maps the features into query, key, and value vectors. Next, the query, key, and value vectors are enhanced together using a dual depthwise separable convolutional block (DDB). The enhanced features are then divided into multiple attention heads. The sharpness of the attention distribution is adjusted by multiplying the transpose of the query and key, combined with a learnable temperature parameter, to calculate the inter-channel attention weights. These weights are used for weighted aggregation of the value vectors. The aggregated features are then projected back to the original dimensions through a 1×1 convolutional layer. Finally, the projected output is added to the input features of the Channel Mapped From Attention Module (DMDTA) via a residual connection.

[0033] For details, please refer to Figure 7 This is the feedforward network module (DFFN) in DSCAB. The summed features are normalized layer by layer and then input into the DFFN. This module first performs feature dimensionality upscaling and spatial context extraction using a 3×3 convolutional layer with a dilation rate of 2. Then, it introduces non-linearity through the GELU activation function. Finally, it performs feature dimensionality reduction and output using another 3×3 convolutional layer with a dilation rate of 2. The output of the DFFN is also added to the module input through residual connections to complete the channel attention branch processing.

[0034] For details, please refer to Figure 8 This is the Local Spatial Attention Module (DOCA) in DSCAB. Input features are first normalized. Then, a 1×1 projective convolutional layer is used to map them to query, key, and value vectors, followed by feature enhancement through a Dual Depthwise Separable Convolutional Block (DDB). The query vector is reshaped into a regular, non-overlapping window. The key and value vectors are then extracted using a sliding window operation based on a set window size and overlap ratio, which captures contextual information across windows. Combining relative position encoding (providing the relative positional relationships between pixels) with a fixed position bias matrix (providing prior spatial structure information), spatial attention weights are calculated within each query window and between it and its overlapping context regions. These spatial attention weights are used to aggregate the sliding value features to obtain local aggregated features. The local aggregated features are projected through a 1×1 convolutional layer and added to the input features of the DOCA module via residual connections.

[0035] For details, please refer to Figure 9This is a Dual Depthwise Separable Convolutional Block (DDB) in DSCAB. First, the input features are initially extracted spatially using a 5×5 depthwise separable convolution. Then, the extracted features are simultaneously fed into two parallel 1×1 convolutional layers for transformation: the first branch performs linear mapping, and the second branch performs normalization. Next, the output of the first branch is activated by the ReLU6 activation function and then multiplied element-wise with the output of the second branch. The features are then channel-fused using a 1×1 convolutional layer and further extracted spatial context using a second 5×5 depthwise separable convolution. Finally, this output is added to the original input of the module through a residual connection to output the enhanced feature vector corresponding to the dual depthwise separable convolutional block, achieving effective reuse and enhancement of the original features. This design significantly improves the discriminative ability of features and the computational efficiency of the module by adaptively fusing multi-scale features through a gating mechanism while expanding the receptive field.

[0036] S023: Replace some of the standard convolutional blocks in the existing RT-DETR model with global context convolutional modules (GCC3). Please see Figure 10 This is a structural diagram of the Global Context Convolutional Module (GCC3) proposed in this application. The GCC3 aims to dynamically enhance key features through a gating mechanism, improving the model's ability to discriminate complex scenes. Its image processing includes: the input feature map is first processed through two parallel 1×1 convolutional layers. The output of one branch serves as a residual connection for the identity mapping, preserving the original feature flow. The other branch enters a sequence of n cascaded gated convolutional blocks (GCBs) for depth-sensing feature transformation and selection, outputting initial features of the depth transformation. After processing by n GCBs, the final transformed feature output by the last GCB of this branch is element-wise added to the features of the shortcut connection branch, forming a residual connection, resulting in the final output of the GCC3—the global context-enhanced feature.

[0037] For details, please refer to Figure 11This is a structural diagram of the Gated Convolutional Block (GCB) in the Global Context Convolutional Module (GCC3). The GCB employs a channel-splitting and gated collaborative enhancement design to achieve adaptive feature selection and fusion. The input features are first uniformly divided into two parts along the channel dimension. The first part enters the feature value branch, where spatial features are extracted sequentially through a 1×1 pointwise convolution and a 3×3 grouped depthwise convolution. The second part enters the feature gating branch, where spatial attention weights are generated sequentially through a 1×1 pointwise convolution and a sigmoid activation function. Subsequently, the outputs of both branches are simultaneously multiplied and summed element-wise to generate two complementary interactive features. Next, the outputs of the original two branches, the multiplication result, and the summation result are concatenated along the channel dimension and projected and fused through a 1×1 convolutional layer, achieving collaborative enhancement and selection of multi-dimensional features. This design not only achieves dynamic feature selection through a gating mechanism but also enhances the expressive power of features through various interactive methods.

[0038] S03: Input the training image set into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; S04: Input the set of images to be tested into the trained traffic target detection model for feature extraction and target detection, and output the traffic target detection results.

[0039] In summary, this application presents a traffic target detection method based on an improved RT-DETR, which synergistically enhances model performance through three core structural improvements: The Dynamic Perception Enhancement Module (DPB) embedded in the backbone network significantly enhances the model's feature representation ability and robustness for multi-scale, occluded, and small targets through parallel multi-level feature processing and a unique High-Frequency Perception Enhancement Block (HAPEB) design; the Dual-Scale Collaborative Attention Module (DSCAB), replacing the original Adaptive Feature Interaction Module (AIFI), efficiently models long-range dependencies and local details through collaborative channel and spatial attention mechanisms, utilizing dilated convolutions and overlapping windows, thus improving the quality of feature interaction and context awareness; and the introduced Global Context Convolution Module (GCC3) achieves dynamic feature selection through gated convolution mechanisms, strengthening the capture of key information without significantly increasing computational burden. These improvements enable the model to maintain the original real-time performance of RT-DETR while effectively improving detection accuracy and overall robustness for various targets such as vehicles, pedestrians, and traffic signs in complex autonomous driving scenarios, especially under conditions of small targets, dense occlusion, and poor lighting.

[0040] Example 2 Please see Figure 12The diagram shown is a schematic representation of a traffic target detection system based on an improved RT-DETR, as proposed in the second embodiment of this application. This system includes the following key modules: The image acquisition module 100 is used to acquire an initial driving video image and divide the initial driving video image into a training image set and a test image set. Model improvement module 200 is used to construct an improved RT-DETR model; the improved RT-DETR model is obtained by structurally adjusting the initial RT-DETR model; the model improvement module includes: The module embedding unit 210 is used to embed a dynamic perception enhancement module into the backbone network of the initial RT-DETR model; The first module replacement unit 220 is used to replace the adaptive feature interaction module in the initial RT-DETR model with a dual-scale collaborative attention module. The second module replacement unit 230 is used to replace some standard convolutional blocks in the initial RT-DETR model with global context convolutional modules. The model training module 300 is used to input the training image set into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; The target detection module 400 is used to input the set of images to be tested into the trained traffic target detection model for feature extraction and target detection, and output the traffic target detection results.

[0041] The traffic target detection system based on improved RT-DETR in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), etc. This application embodiment does not impose specific limitations.

[0042] The traffic target detection system based on the improved RT-DETR in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0043] The traffic target detection system based on an improved RT-DETR provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiment based on an improved RT-DETR traffic target detection method are not described in detail here to avoid repetition.

[0044] Optionally, embodiments of this application also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described embodiment of a traffic target detection method based on improved RT-DETR and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0045] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described embodiment of a traffic target detection method based on an improved RT-DETR, and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0046] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0047] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0048] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0049] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A traffic target detection method based on improved RT-DETR, characterized in that, Includes the following steps: Acquire initial driving video images and divide the initial driving video images into a training image set and a test image set; An improved RT-DETR model is constructed; the improved RT-DETR model is obtained by structurally adjusting the RT-DETR model. The specific process of the structural adjustment includes: The dynamic perception enhancement module is embedded in the backbone network of the RT-DETR model, specifically including: acquiring a first input feature map transmitted by the backbone network; adjusting the number of channels of the first input feature map using a main channel adjustment layer with a first specific-size convolutional kernel to obtain a target feature map; uniformly dividing the target feature map along the channel dimension into a first feature subset and a second feature subset; retaining the first feature subset as a direct branch feature; inputting the second feature subset as a starting feature into a processing branch; the processing branch contains a preset number of high-frequency perception enhancement blocks that are sequentially connected; using each of the high-frequency perception enhancement blocks in the processing branch to sequentially perform deep feature extraction on the starting feature to obtain local intermediate features corresponding to each of the high-frequency perception enhancement blocks; concatenating the direct branch feature, the starting feature, and each of the local intermediate features along the channel dimension to obtain a multi-dimensional concatenated feature; adjusting the number of channels of the multi-dimensional concatenated feature using a main feature fusion layer with a second specific-size convolutional kernel to output the enhanced backbone feature corresponding to the dynamic perception enhancement module. The adaptive feature interaction module in the RT-DETR model is replaced with a dual-scale collaborative attention module, specifically including: obtaining a second input feature map that enters the dual-scale collaborative attention module; inputting the second input feature map into a channel attention branch for long-range channel dependency extraction, and outputting channel-enhanced features; the channel attention branch sequentially includes a channel-mapped attention module and a first feedforward network module; inputting the channel-enhanced features into a spatial attention branch for local spatial context extraction, and outputting the collaborative interaction features corresponding to the dual-scale collaborative attention module; the spatial attention branch sequentially includes a local spatial attention module and a second feedforward network module. The process of replacing some standard convolutional blocks in the RT-DETR model with global context convolutional modules includes: obtaining the third input feature map entering the global context convolutional module; performing dimensionality transformation on the third input feature map using two parallel 1×1 channel dimensionality reduction convolutional layers, outputting shortcut connection features and depth transformation initial features; inputting the depth transformation initial features into a sequence of gated convolutional blocks containing the target number in sequence; for any gated convolutional block in the sequence: uniformly dividing the received block input features along the channel dimension into a first channel part and a second channel part; inputting the first channel part into the value extraction branch, sequentially extracting spatial features through a 1×1 first pointwise convolutional layer and a 3×3 grouped depth convolutional layer, outputting the first spatial extracted features; inputting the second channel part into... The gated branch sequentially generates spatial attention weights through a second pointwise convolutional layer and a logistic activation layer of size 1×1. Element-wise multiplication and element-wise addition operations are simultaneously performed on the first spatial extraction features and the spatial attention weights to generate second multiplication interaction features and second addition interaction features. The first spatial extraction features, the spatial attention weights, the second multiplication interaction features, and the second addition interaction features are concatenated along the channel dimension, and after layer normalization, a 1×1 fusion projection convolutional layer is used for mapping calculation to output the depth transformation stage features corresponding to the current gated convolutional block. The final state transformation features output by the last gated convolutional block in the sequence are obtained, and the final state transformation features are added element-wise with the shortcut connection features to form a residual connection, outputting the global context enhancement features corresponding to the global context convolutional module. The training image set is input into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; The set of images to be tested is input into the trained traffic target detection model for feature extraction and target detection, and the traffic target detection results are output.

2. The method according to claim 1, characterized in that, The process of sequentially extracting deep features from the initial features using each of the high-frequency sensing enhancement blocks in the processing branch to obtain the local intermediate features corresponding to each high-frequency sensing enhancement block includes: Obtain the current input features entering the current high-frequency sensing enhancement block, and uniformly divide the current input features into local feature parts and high-frequency feature parts along the channel dimension; The local feature portion is input into the local feature extraction branch, processed by the first convolutional layer with the third specific size convolutional kernel and the first nonlinear activation function, and the first local extracted feature is output. The high-frequency feature is input into the high-frequency enhancement branch and processed sequentially through a max pooling layer of a set size, a second convolutional layer with a fourth specific-size convolutional kernel, and a second nonlinear activation function to output the first high-frequency enhancement feature. The first local extracted feature and the first high-frequency enhanced feature are added element-wise to generate a first additive interactive feature, and the first local extracted feature and the first high-frequency enhanced feature are multiplied element-wise to generate a first multiplicative interactive feature. The first local extraction feature, the first high-frequency enhancement feature, the first additive interaction feature, and the first multiplicative interaction feature are concatenated along the channel dimension to obtain the composite interaction feature. The composite interactive features are fused and mapped using the third convolutional layer with a fifth specific-size convolutional kernel, and the fused and mapped features are added to the current input features by residual addition, and the local intermediate features corresponding to the current high-frequency perception enhancement block are output.

3. The method according to claim 1, characterized in that, The process of inputting the second input feature map into the channel attention branch for long-range channel dependency extraction and outputting channel-enhanced features includes: The second input feature map is normalized in the first layer by the channel mapping attention module, and then mapped into a channel query vector, a channel key vector and a channel value vector by a first projection convolutional layer of size 1×1. The channel query vector, the channel key vector, and the channel value vector are respectively input into a preset dual-depth separable convolutional block for feature enhancement to obtain enhanced query vector, enhanced key vector, and enhanced value vector; The enhanced query vector, the enhanced key vector, and the enhanced value vector are divided into several attention heads; for each attention head, the matrix product of the corresponding enhanced query vector and the enhanced key vector is calculated, and the channel attention weight is calculated in combination with the learnable temperature parameter. The channel aggregation feature is obtained by weighting and aggregating the enhanced value vector using the channel attention weights. After feature reconstruction and processing by a second projection convolutional layer of size 1×1, it is added to the second input feature map by the first residual to output the first residual feature. The first residual feature is input into the first feedforward network module. The first dimensionality reduction and dimensionality increase convolutional layer with an expansion rate of 2 and a size of 3×3 is used to extract the spatial context of the first residual feature. After passing through the third nonlinear activation function, the second dimensionality reduction and dimensionality increase convolutional layer with an expansion rate of 2 and a size of 3×3 is used for dimensionality reduction processing. The result is added to the first residual feature as a second residual, and the channel enhancement feature corresponding to the channel attention branch is output.

4. The method according to claim 1, characterized in that, The process of inputting the channel enhancement features into the spatial attention branch for local spatial context extraction and outputting the collaborative interaction features corresponding to the dual-scale collaborative attention module includes: The channel enhancement features are normalized in the second layer by the local spatial attention module, and mapped into spatial query vector, spatial key vector and spatial value vector by a third projection convolutional layer of size 1×1. The spatial query vector, the spatial key vector, and the spatial value vector are respectively input into a preset dual-depth separable convolutional block for feature enhancement to obtain a reset query vector, a reset key vector, and a reset value vector. The reset query vector is reshaped into a regular non-overlapping window format query feature; a sliding window operation is performed on the reset key vector and the reset value vector based on the set window size and overlap ratio to extract the sliding key feature and sliding value feature of the overlapping context region; By combining relative position encoding information and a fixed position bias matrix, the spatial attention weight between the regular non-overlapping window format query feature and the sliding key feature is calculated; the spatial attention weight is then used to aggregate the sliding value feature to obtain a local aggregated feature. The local aggregated features are projected using a fourth projection convolutional layer of size 1×1, and then added to the channel enhancement features using a third residual to output the second residual features. The second residual feature is input into the second feedforward network module for dimensionality upscaling and downscaling extraction. The result is added to the second residual feature as a fourth residual, and the collaborative interaction feature corresponding to the spatial attention branch is output.

5. The method according to claim 3 or 4, characterized in that, The process of feature enhancement using the preset dual-depth separable convolutional blocks includes: The feature vector to be enhanced is obtained, and the spatial features of the feature vector to be enhanced are extracted using a first depthwise separable convolutional layer with a size of 5×5 to obtain the basic spatial features. The basic spatial features are respectively input to the first linear transformation branch and the second normalized transformation branch set in parallel; in the first linear transformation branch, the first transformation feature is processed by the fifth projection convolution layer with a size of 1×1, and the second transformation feature is processed by the sixth projection convolution layer with a size of 1×1. After the first transformation feature is activated by the fourth nonlinear activation function, it is multiplied element-wise with the second transformation feature to generate the feature fusion basic term; Channel fusion is performed on the feature fusion base terms using a seventh projection convolutional layer with a size of 1×1, and spatial context is further extracted using a second depthwise separable convolutional layer with a size of 5×5 to obtain depth context features. The deep context features are added to the feature vector to be enhanced using the fifth residual, and the enhanced feature vector corresponding to the preset dual-depth separable convolutional block is output.

6. A traffic target detection system based on an improved RT-DETR, characterized in that, include: An image acquisition module is used to acquire initial driving video images and divide the initial driving video images into a training image set and a test image set; The model improvement module is used to construct an improved RT-DETR model; the improved RT-DETR model is obtained by structurally adjusting the RT-DETR model. The model improvement module includes: The module embedding unit is used to embed a dynamic perception enhancement module into the backbone network of the RT-DETR model. Specifically, it includes: acquiring a first input feature map transmitted by the backbone network; adjusting the number of channels in the first input feature map using a main channel adjustment layer with a first specific-size convolutional kernel to obtain a target feature map; uniformly dividing the target feature map along the channel dimension into a first feature subset and a second feature subset; retaining the first feature subset as a direct branch feature; inputting the second feature subset as a starting feature into a processing branch; the processing branch contains a preset number of high-frequency perception enhancement blocks connected in series; sequentially extracting deep features from the starting feature using each of the high-frequency perception enhancement blocks in the processing branch to obtain local intermediate features corresponding to each high-frequency perception enhancement block; concatenating the direct branch feature, the starting feature, and each of the local intermediate features along the channel dimension to obtain a multi-dimensional concatenated feature; adjusting the number of channels in the multi-dimensional concatenated feature using a main feature fusion layer with a second specific-size convolutional kernel to output the enhanced backbone feature corresponding to the dynamic perception enhancement module. The first module replacement unit is used to replace the adaptive feature interaction module in the RT-DETR model with a dual-scale collaborative attention module. Specifically, it includes: obtaining a second input feature map that enters the dual-scale collaborative attention module; inputting the second input feature map into a channel attention branch for long-range channel dependency extraction and outputting channel enhancement features; the channel attention branch sequentially includes a channel mapping attention module and a first feedforward network module; inputting the channel enhancement features into a spatial attention branch for local spatial context extraction and outputting the collaborative interaction features corresponding to the dual-scale collaborative attention module; the spatial attention branch sequentially includes a local spatial attention module and a second feedforward network module. The second module replacement unit is used to replace some standard convolutional blocks in the RT-DETR model with global context convolutional modules. The specific process includes: obtaining the third input feature map entering the global context convolutional module; performing dimensionality transformation on the third input feature map using two parallel 1×1 channel dimensionality reduction convolutional layers, outputting shortcut connection features and depth transformation initial features; inputting the depth transformation initial features into a sequence of gated convolutional blocks containing the target number in sequence; for any gated convolutional block in the sequence: uniformly dividing the received block input features along the channel dimension into a first channel part and a second channel part; inputting the first channel part into the value extraction branch, sequentially extracting spatial features through a first pointwise convolutional layer of size 1×1 and a grouped depth convolutional layer of size 3×3, outputting the first spatial extracted features; and inputting the second channel part into the value extraction branch. The channel input is fed into the gated branch, and spatial attention weights are generated sequentially through a second pointwise convolutional layer and a logistic function activation layer of size 1×1. The first spatial extraction feature and the spatial attention weights are simultaneously multiplied and added element-wise to generate a second multiplicative interaction feature and a second additive interaction feature. The first spatial extraction feature, the spatial attention weights, the second multiplicative interaction feature, and the second additive interaction feature are concatenated along the channel dimension, normalized, and then mapped using a 1×1 fusion projection convolutional layer to output the depth transformation stage feature corresponding to the current gated convolutional block. The final state transformation feature output by the last gated convolutional block in the sequence is obtained, and the final state transformation feature is added element-wise with the shortcut connection feature to form a residual connection, outputting the global context enhancement feature corresponding to the global context convolutional module. The model training module is used to input the training image set into the improved RT-DETR model for iterative optimization training to obtain the trained traffic target detection model; The target detection module is used to input the set of images to be tested into the trained traffic target detection model for feature extraction and target detection, and output the traffic target detection results.

7. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of a traffic target detection method based on an improved RT-DETR as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Traffic sign detection method based on improved RT-DETR

    CN117975418A

  • Unmanned aerial vehicle infrared small target detection method based on PGF-RTDETR

    CN121191040A