Lightweight Real-Time Object Detection Method, Apparatus, Server, and Storage Medium

By modifying the RT-DETR model's backbone network with lightweight downsampling and multi-scale dilated attention modules, the method addresses the computational resource constraints of RT-DETR, enabling real-time target detection on resource-limited devices with improved accuracy and reduced resource usage.

CN119810428BActive Publication Date: 2025-07-15TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294167.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-15
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The RT-DETR model has a large computing overhead and high hardware resource requirements, making it difficult to realize real-time object detection on intelligent security cameras with limited computing power, and image transmission consumes a lot of resources.

Method used

The backbone network of the RT-DETR model is improved, and the lightweight downsampling submodule and multi-scale expanded attention submodule are adopted. Feature extraction is performed through hollow convolution and self-attention mechanisms, reducing the amount of parameters and calculations, and improving detection accuracy.

Benefits of technology

Real-time object detection is realized on devices with limited computing power, reducing image transmission resources, extending the working time of the equipment, and improving the accuracy of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810428B_ABST
    Figure CN119810428B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight real-time object detection method, device, server and storage medium. The method includes: extracting shallow image features by using a basic convolution module of an RT-DETR model backbone network; extracting context features by using a lightweight downsampling sub-module of a multi-scale dilated convolution module and fusing them with local features; extracting multi-scale features by using a multi-scale dilated attention sub-module of the multi-scale dilated convolution module and fusing them to obtain deep image features; and obtaining an object detection result of image data by using a hybrid encoder and decoder. By improving the BasicBlock module with the lightweight downsampling sub-module and the multi-scale dilated attention sub-module, using dilated convolution to capture surrounding context semantic information without changing the convolution kernel size instead of the original method, and the sparsity of multi-scale dilated attention at different scales, the number of model parameters and the amount of computation are reduced, enabling the model to be deployed on devices with limited computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a lightweight real-time object detection method, device, server and storage medium. Background Art

[0002] In fields such as intelligent security and autonomous driving, it is often necessary to perform real-time detection of objects in images. Object detection models such as the YOLO series and the DETR series can extract features from images to detect specific objects in the images. RT-DETR (Real-Time DEtection TRansformer) is an object detection model based on the Transformer structure, adopting an encoder-decoder architecture, and using the self-attention mechanism to extract and associate object features. RT-DETR can better capture long-range dependencies in images, and in images with dense objects or complex backgrounds, it can show stronger detection effects than the YOLO series through adaptive feature capture.

[0003] The RT-DETR model has a large computational overhead and high requirements for hardware resources. To achieve real-time object detection, it often needs to be deployed in a data center with sufficient computing power. The data center uses the RT-DETR model to perform object detection on the received images and transmit the detection results back to the terminal device. However, the deployment of intelligent security systems often has requirements for low cost. Transmitting image data to the data center consumes a large amount of transmission resources, and the cost of establishing a data center with high computing power is high. The intelligent security camera, as the terminal of the security system, has limited computing power itself and cannot meet the real-time object detection computing power requirements of the RT-DETR model. Summary of the Invention

[0004] Embodiments of the present invention provide a lightweight real-time object detection method, device, server and storage medium to solve the technical problem that the RT-DETR model cannot run in an environment with limited computing power.

[0005] In a first aspect, embodiments of the present invention provide a lightweight real-time object detection method, including:

[0006] Input the collected image data into the RT-DETR model, where the backbone network of the RT-DETR model includes a basic convolution module and a multi-scale dilated convolution module, and the multi-scale dilated convolution module includes a lightweight downsampling sub-module and a multi-scale dilated attention sub-module;

[0007] Use the basic convolution module to perform basic feature extraction on the collected image data to obtain shallow image features;

[0008] The lightweight downsampling sub-module is used to extract context features from the shallow image features through dilated convolution, and fuse the context features with the local features extracted from the shallow image by the lightweight downsampling sub-module through standard convolution to obtain joint image features containing context;

[0009] The multi-scale dilated attention sub-module is used to extract features of different scales from the joint image features at different dilation rates respectively, and fuse the features of different scales to obtain deep image features;

[0010] The hybrid encoder of the RT-DETR model is used to fuse the deep image features, and the decoder of the RT-DETR model is used to identify target features to obtain the target detection result of the image data.

[0011] Further, the process of using the lightweight downsampling sub-module to extract context features from the shallow image features through dilated convolution, and fusing the context features with the local features extracted from the shallow image by the lightweight downsampling sub-module through standard convolution to obtain joint image features containing context includes:

[0012] The lightweight downsampling sub-module performs standard convolution processing on the shallow image features to obtain refined local features, and at the same time performs dilated convolution processing on the shallow image features to obtain surrounding context features;

[0013] The lightweight downsampling sub-module fuses the refined local features and the surrounding context features to obtain fused features;

[0014] The lightweight downsampling sub-module strengthens the key information of the fused features to obtain joint image features containing context.

[0015] Further, the process of the lightweight downsampling sub-module fusing the refined local features and the surrounding context features to obtain fused features includes:

[0016] The refined local features and the surrounding context features are concatenated to obtain concatenated features;

[0017] The concatenated features are batch-normalized and fused features are generated through an activation function.

[0018] Further, the process of the lightweight downsampling sub-module strengthening the key information of the fused features to obtain joint image features containing context includes:

[0019] The fused features are globally average-pooled to obtain global features;

[0020] A multi-layer perceptron is used to strengthen the key information of the global features to obtain joint image features containing context.

[0021] Further, the multi-scale dilated attention sub-module extracts features of different scales from the joint image features at different dilation rates respectively, and fuses the features of different scales to obtain deep image features, including:

[0022] The multi-scale dilated attention sub-module performs multi-scale SWDA operations on the joint image features at different dilation rates using multiple heads respectively to obtain head features for each head;

[0023] The multi-scale dilated attention sub-module aggregates the features of all heads using a linear layer to obtain deep image features.

[0024] Further, fusing the deep image features using the hybrid encoder of the RT-DETR model, and performing target feature recognition using the decoder of the RT-DETR model to obtain the target detection result of the image data, including:

[0025] The hybrid encoder performs multi-scale feature fusion on the deep image features respectively generated by the last three multi-scale dilated convolutional modules in the backbone network to obtain encoder features;

[0026] The decoder performs decoding iteration and prediction classification on the encoder features to obtain the target detection result of the image data.

[0027] Further, the multi-scale dilated convolutional module further includes a basic convolutional sub-module and an activation function sub-module.

[0028] In a second aspect, an embodiment of the present invention provides a lightweight real-time target detection device, including:

[0029] An image acquisition unit for acquiring image data and transmitting it to the RT-DETR model;

[0030] A basic convolutional unit for performing standard convolution on the acquired image data;

[0031] A multi-scale dilated convolutional unit for performing multi-scale dilated convolution on the shallow image features generated after standard convolution;

[0032] A hybrid encoder unit for fusing the deep image features generated after multi-scale dilated convolution;

[0033] A decoder unit for performing target detection after the feature fusion of the hybrid encoder.

[0034] In a third aspect, an embodiment of the present invention provides a server, including:

[0035] One or more processors;

[0036] A storage device for storing one or more programs,

[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the above lightweight real-time object detection method.

[0038] In a fourth aspect, an embodiment of the present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the above lightweight real-time object detection method when executed by a computer processor.

[0039] A lightweight real-time object detection method, device, server, and storage medium provided by an embodiment of the present invention. The method improves the BasicBlock module in the backbone network of the RT-DETR model, and uses a lightweight downsampling sub-module and a multi-scale dilated attention sub-module to replace the original basic convolution sub-module. The lightweight downsampling sub-module jointly extracts features through standard convolution and dilated convolution to obtain joint image features containing context, reduces the number of parameters and computational complexity of the model, and improves the accuracy of object detection by capturing context semantic information. The multi-scale dilated attention sub-module uses different dilation rates to extract features at different scales to capture multi-scale semantic information, effectively compensates for the detection performance of the lightweight downsampling sub-module in the face of complex scenes, small objects, and high-density objects, and uses the sparsity of the self-attention mechanism at different scales to further reduce computational redundancy while ensuring the detection performance of the model. The deep image features extracted by the lightweight downsampling sub-module and the multi-scale dilated attention sub-module contain the details of local features, surrounding context semantic information, and multi-scale semantic information, which can improve the accuracy of object detection while reducing the number of parameters and computational complexity of the model, enabling the RT-DETR model to be deployed on terminal devices or edge devices with limited computing power such as intelligent security cameras to achieve real-time object detection, and only transmitting the detection results to the data center without transmitting image or video data, greatly reducing the occupancy of transmission resources and computational energy consumption, and extending the working hours of the device. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0041] Figure 1 It is a flowchart of a lightweight real-time object detection method according to Embodiment 1 of the present invention;

[0042] Figure 2 It is a schematic diagram of the architecture of the RT-DETR model according to Embodiment 1 of the present invention;

[0043] Figure 3Flowchart of a lightweight real-time object detection method according to the second embodiment of the present invention;

[0044] Figure 4 Schematic structural diagram of the lightweight downsampling sub-module according to the second embodiment of the present invention;

[0045] Figure 5 Flowchart of a lightweight real-time object detection method according to the third embodiment of the present invention;

[0046] Figure 6 Schematic structural diagram of the multi-scale dilated attention sub-module according to the third embodiment of the present invention;

[0047] Figure 7 Schematic structural diagram of a lightweight real-time object detection device according to the fourth embodiment of the present invention;

[0048] Figure 8 Structural diagram of the server according to the fifth embodiment of the present invention. Detailed implementation manners

[0049] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0050] The RT-DETR (Real-Time Detection Transformer) model is a real-time object detection model based on the Transformer structure. By adopting an encoder-decoder structure and using the self-attention mechanism to extract and associate object features, it can better capture long-range dependencies in images, improve the detection accuracy of the model in complex scenarios, and obtain high object detection accuracy through adaptive feature capture in images with dense objects or complex backgrounds. It also has an end-to-end detection structure, simplifies the detection process, and reduces error accumulation to a certain extent. However, the RT-DETR model has a large number of parameters and a large amount of computation. To meet the real-time requirement of object detection, it has high requirements for hardware resources and computing resources. On some devices with limited computing resources or edge devices, it is difficult to meet the required computing resource requirements, thus making it difficult to achieve real-time object detection. Especially in the intelligent security system, to meet the convenience of system deployment and reduce costs, real-time object detection needs to be performed on the intelligent security cameras at the system terminal, and only the detection results are transmitted back to the data center, replacing the large amount of transmission resources occupied by traditional cameras when transmitting image / video data back to the data center. However, the computing power of intelligent security cameras is limited and it is difficult to meet the computing requirements for the RT-DETR model to achieve real-time object detection.

[0051] Example 1

[0052] Figure 1 The figure is a flowchart of a lightweight real-time object detection method according to Embodiment 1 of the present invention. In this embodiment, the module for extracting deep image features in the backbone network is improved, and the specific steps are as follows:

[0053] S101: Input the collected image data into the RT-DETR model. The backbone network of the RT-DETR model includes a basic convolution module and a multi-scale dilated convolution module. The multi-scale dilated convolution module includes a lightweight downsampling sub-module and a multi-scale dilated attention sub-module.

[0054] The RT-DETR model mainly extracts features of the input image through the backbone network (Backbone), as Figure 2 shown. The backbone network includes multiple layers of basic convolution modules ConvBN and multiple layers of multi-scale dilated convolution modules CCM. The basic convolution module ConvBN is used to extract features of the image input into the model using a standard convolution kernel. Each basic convolution module includes a standard convolution sub-module Conv, a batch normalization sub-module BN, and an activation function sub-module Relu. The multi-scale dilated convolution module CCM is used to extract features of the image using a standard convolution and multi-scale dilated convolutions. Each multi-scale dilated convolution module includes a basic convolution sub-module ConvBN, a lightweight downsampling sub-module CG, a multi-scale dilated attention sub-module MSDA, and an activation function sub-module Relu. By improving the BasicBlock module in the model backbone network to a multi-scale dilated convolution module CCM, the number of parameters and the amount of computation of the improved model are significantly reduced, and by paying attention to the context features through multi-scale dilated convolutions, the accuracy of object detection is ensured.

[0055] S102: Use the basic convolution module to perform basic feature extraction on the collected image data to obtain shallow image features.

[0056] The multiple layers of basic convolution modules ConvBN in the backbone network respectively perform multi-layer feature extraction on the image input into the model through convolution kernels of different sizes, gradually extracting features from low-level to high-level, and finally obtaining shallow image features. When extracting shallow image features through multiple layers of basic convolution modules, more attention is paid to the local features in the image.

[0057] S103: Use the lightweight downsampling sub-module to perform dilated convolution on the shallow image features to extract context features, and fuse the context features with the local features extracted by the lightweight downsampling sub-module through standard convolution on the shallow image to obtain joint image features containing context.

[0058] In the traditional RT-DETR model, the deep BasicBlock modules in the backbone network extract deep image features based on shallow image features and use them as the input for the hybrid encoder for feature fusion. The BasicBlock module contains the basic convolutional sub-module ConvBN and the activation function sub-module Relu. The lightweight downsampling sub-module CG (CG Block, Context Guided Block) is used to replace the ConvBN sub-module. The lightweight downsampling sub-module performs dilated convolution on the shallow image features to extract the context information of the local features, enhancing the scene understanding ability, improving the semantic segmentation accuracy during the deep feature extraction of the backbone network, and fusing the context features extracted by dilated convolution with the local features extracted by standard convolution at the same time. The fused joint image features contain the semantic information of the context. By using dilated convolution instead of standard convolution to extract the context around the local features, the number of parameters during feature extraction is reduced, and thus the computational complexity of the model is reduced.

[0059] S104. Use the multi-scale dilated attention sub-module to perform feature extraction on the joint image features at different scales with different dilation rates respectively, and fuse the features at different scales to obtain deep image features.

[0060] Using the multi-scale dilated attention sub-module MSDA as the subsequent processing of the lightweight downsampling sub-module is because the lightweight downsampling module reduces the computational complexity and the number of parameters of the model by introducing dilated convolution, but it is prone to insufficient detection accuracy of complex targets or small targets in high-density target scenarios. Using the self-attention mechanism of the multi-scale dilated attention sub-module to perform multi-scale SWDA with different dilation rates on the joint image features at different heads respectively, and aggregating the features with different dilation rates extracted by different heads through feature concatenation and linear layers to obtain deep image features containing multi-scale features. The features extracted by MSDA contain multi-scale semantic information. Especially in complex scenes, the information in the image can be captured more comprehensively through the semantic information at different scales. At the same time, using the sparsity of the self-attention mechanism at different scales reduces the computational redundancy while ensuring the detection performance.

[0061] S105. Use the hybrid encoder of the RT-DETR model to perform feature fusion on the deep image features, and use the decoder of the RT-DETR model to perform target feature recognition to obtain the target detection result of the image data.

[0062] The Efficient Hybrid Encoder receives the deep image features extracted from the deep CCM module of the backbone network, performs feature alignment and feature fusion. The fused features contain shallow local details and deep semantic information and are sent to the decoder. The decoder is composed of multiple stacked Transformer decoding layers, and through multi-layer iterative decoding of the features input to the decoder and classifying the detected targets, the final object detection result is obtained.

[0063] In this embodiment, by improving the BasicBlock module in the backbone network of the RT-DETR model, the original basic convolutional sub-module is replaced by a lightweight downsampling sub-module and a multi-scale dilated attention sub-module. The standard convolution and dilated convolution in the lightweight downsampling sub-module are jointly used for feature extraction. The dilated convolution can capture the surrounding context semantic information of the local features without changing the size of the convolution kernel, replacing the original way of enlarging the size of the standard convolution kernel. Then, the refined local features and surrounding context features extracted by the standard convolution and dilated convolution respectively are fused, and finally the joint image features containing context are obtained, reducing the number of model parameters and computational complexity. And by capturing the context semantic information, the accuracy of object detection is improved. Further, the multi-scale dilated attention sub-module uses different dilation rates to extract features at different scales to capture multi-scale semantic information, and uses the sparsity of the self-attention mechanism at different scales to further reduce computational redundancy while ensuring the detection performance of the model. The deep image features extracted by the lightweight downsampling sub-module and the multi-scale dilated attention sub-module contain the details of local features, the surrounding context semantic information and multi-scale semantic information, which can not only improve the accuracy of object detection through multiple semantic information, but also reduce the number of model parameters and computational complexity, enabling the RT-DETR model to be deployed on terminal devices or edge devices with limited computing power such as intelligent security cameras. The intelligent security camera can perform real-time object detection and only transmit the detection results to the data center, without transmitting image or video data, greatly reducing resource occupancy.

[0064] An alternative implementation of this embodiment is that the multi-scale dilated convolution module further includes a basic convolutional sub-module and an activation function sub-module.

[0065] When improving the RT-DETR model by using the CCM module instead of the BasicBlock module, the original BasicBlock module includes two basic convolution ConvBN sub-modules and one activation function Relu sub-module. The second ConvBN sub-module in the original BasicBlock module is replaced by CG-MSDA to form a ConvBN-CG-MSDA-Relu structure. Data is processed and transmitted in the ConvBN sub-module, CG sub-module, MSDA sub-module, and Relu sub-module in sequence after entering the CCM module. Finally, the CCM module outputs deep image features. The CG sub-module uses dilated convolution instead of standard convolution when extracting the surrounding context features of local features. Dilated convolution can expand the receptive field of the convolution kernel through the dilation rate, replacing the method of extracting the surrounding context features of local features by increasing the size of the convolution kernel, significantly reducing the number of model parameters and computational complexity. Then, the MSDA sub-module is used for multi-scale feature extraction. MSDA can adaptively extract features of different targets from a multi-scale perspective, improving the feature extraction accuracy in complex environments. The sparse performance of the self-attention mechanism of MSDA at different scales further reduces computational redundancy and the computational complexity of the model on the basis of maintaining the model performance. Through the combination of the two, the richness of context information can be guaranteed in complex environments, and the detection accuracy of small targets, high-density targets, and targets of different scales in complex backgrounds is improved.

[0066] Embodiment 2

[0067] Figure 3 The following is a flowchart of a lightweight real-time object detection method according to Embodiment 2 of the present invention. This embodiment is optimized based on the above embodiment. In this embodiment, the dilated convolution is used to extract the context features from the shallow image features by using the lightweight downsampling sub-module, and the context features are fused with the local features extracted by the standard convolution of the shallow image by the lightweight downsampling sub-module to obtain the joint image features containing the context. The specific optimization is as follows:

[0068] The lightweight downsampling sub-module performs standard convolution processing on the shallow image features to obtain refined local features, and at the same time performs dilated convolution processing on the shallow image features to obtain surrounding context features;

[0069] The lightweight downsampling sub-module fuses the refined local features and the surrounding context features to obtain fused features;

[0070] The lightweight downsampling sub-module strengthens the key information of the fused features to obtain the joint image features containing the context.

[0071] Correspondingly, the lightweight real-time object detection method provided in this embodiment specifically includes:

[0072] S201. Input the collected image data into the RT-DETR model. The backbone network of the RT-DETR model includes a basic convolutional module and a multi-scale dilated convolutional module. The multi-scale dilated convolutional module includes a lightweight downsampling sub-module and a multi-scale dilated attention sub-module.

[0073] S202. Use the basic convolutional module to perform basic feature extraction on the collected image data to obtain shallow image features.

[0074] S203. The lightweight downsampling sub-module performs standard convolutional processing on the shallow image features to obtain refined local features, and at the same time performs dilated convolutional processing on the shallow image features to obtain surrounding context features.

[0075] The lightweight downsampling sub-module CG Block includes a local feature extractor ( ), a surrounding context extractor ( ), a joint feature extractor ( ), and a global context extractor ( ), as shown in Figure 4 . First, the local feature extractor performs standard convolutional extraction on the input shallow image features to obtain detailed local features. At the same time, the surrounding context extractor parallelly performs dilated convolutional extraction on the input shallow image features to obtain surrounding context features with a wider receptive field. Dilated convolution can expand the receptive field of the convolutional kernel through the dilation rate, and no longer extract surrounding context information by expanding the size of the standard convolutional kernel, reducing the number of parameters and computational complexity.

[0076] S204. The lightweight downsampling sub-module performs feature fusion on the refined local features and the surrounding context features to obtain fused features;

[0077] The joint feature extractor receives the refined local features and the surrounding context features extracted in parallel by the local feature extractor and the surrounding context extractor, and generates fused features for the refined local features and the surrounding context features through feature concatenation, batch normalization, and activation functions. The generated fused features contain local detail information and context semantic information.

[0078] S205. The lightweight downsampling sub-module performs key information enhancement on the fused features to obtain joint image features containing context.

[0079] The global context extractor receives the fused features fused by the joint feature extractor, aggregates the global context features through global average pooling, and then performs key information enhancement and refinement through a multi-layer perceptron, improves the attention to key features through weighted processing, and uses rich context semantic information to improve the accuracy of feature recognition.

[0080] S206. Use the multi-scale dilation attention sub-module to perform feature extraction on the joint image features at different scales with different dilation rates, and fuse the features at different scales to obtain deep image features.

[0081] S207. Use the hybrid encoder of the RT-DETR model to perform feature fusion on the deep image features, and use the decoder of the RT-DETR model to perform target feature recognition to obtain the target detection result of the image data.

[0082] In this embodiment, the standard convolution of the local feature extractor is used to extract detailed local features, and at the same time, the dilated convolution of the surrounding context extractor is used in parallel to extract the surrounding context features. Then, feature fusion is performed by the joint feature extractor, and the key information is enhanced through weighted processing by the global context extractor. The receptive field is enlarged by the dilation rate of the dilated convolution, replacing the original method of enlarging the standard convolution kernel to enlarge the receptive field, reducing the number of model parameters and computational complexity, and improving the attention to key information through feature fusion and weighted processing, aggregating rich context semantic information, and improving the performance of the model.

[0083] Specifically, the lightweight downsampling sub-module fuses the refined local features and the surrounding context features to obtain fused features, including:

[0084] Perform feature concatenation on the refined local features and the surrounding context features to obtain concatenated features.

[0085] After the joint feature extractor receives the refined local features and the surrounding context features extracted in parallel by the local feature extractor and the surrounding context extractor, it first uses the Concat layer to perform feature concatenation on the refined local features and the surrounding context features. The concatenated features generated after concatenation contain local features and the context features around the local features. The context features are extracted by dilated convolution, and the size of the convolution kernel remains unchanged, replacing the method of obtaining context features by enlarging the convolution kernel, reducing the number of model parameters and computational complexity.

[0086] Perform batch normalization processing on the concatenated features, and generate fused features through an activation function.

[0087] Use batch normalization (BN) to process the concatenated features to improve the stability and generalization ability of the model. Then, use the parametric ReLU (PReLU) activation function to improve the expression ability of the model, enhance the non-linear expression ability of the model, and finally generate fused features containing local details and context semantics, ensuring the robustness and training efficiency of the model.

[0088] Specifically, the lightweight downsampling sub-module strengthens the key information of the fused features to obtain joint image features containing context, including:

[0089] Perform global average pooling on the fused features to obtain global features.

[0090] By average pooling, the fused features of all channels are aggregated, retaining the features between channels, reducing the number of model parameters, and lowering the model complexity. Smoothing the global features through mean calculation improves the robustness of the model.

[0091] Use a multi-layer perceptron to enhance the key information of the global features and obtain joint image features containing context.

[0092] The multi-layer perceptron consists of two fully connected layers FC. Use the two fully connected layers to weight the global features. By increasing the weight of the key information, the key information is enhanced and irrelevant information is suppressed. The fully connected layer can adaptively enhance the key information of different input images through weight learning, thereby improving the accuracy of the model.

[0093] Embodiment III

[0094] Figure 5 It is a flowchart of a lightweight real-time object detection method according to Embodiment III of the present invention. This embodiment is optimized based on the above embodiments. In this embodiment, the multi-scale dilated attention sub-module is used to perform feature extraction of different scales on the joint image features at different dilation rates, and the features of different scales are fused to obtain deep image features. The specific optimization is as follows:

[0095] The multi-scale dilated attention sub-module uses multiple heads to perform multi-scale SWDA operations on the joint image features at different dilation rates to obtain the head features of each head.

[0096] The multi-scale dilated attention sub-module uses a linear layer to aggregate the features of all heads to obtain deep image features.

[0097] Correspondingly, the lightweight real-time object detection method provided in this embodiment specifically includes:

[0098] S301, Input the collected image data into the RT-DETR model. The backbone network of the RT-DETR model includes a basic convolution module and a multi-scale dilated convolution module. The multi-scale dilated convolution module includes a lightweight downsampling sub-module and a multi-scale dilated attention sub-module.

[0099] S302, Use the basic convolution module to perform basic feature extraction on the collected image data to obtain shallow image features.

[0100] In S303, the lightweight downsampling sub-module is used to perform dilated convolution on the shallow image features to extract context features, and the context features are fused with the local features extracted by the lightweight downsampling sub-module through standard convolution on the shallow image to obtain joint image features containing context.

[0101] In S304, the multi-scale dilated attention sub-module uses multiple heads to perform multi-scale SWDA operations on the joint image features at different dilation rates respectively to obtain the head features of each head.

[0102] The multi-scale dilated attention sub-module first performs a linear projection on the input feature map, as Figure 6 shown, to obtain Query, Key, and Value. Then, different heads perform multi-scale SWDA (Sliding Window Dilated Attention) operations at different dilation rates respectively, finely capturing target features and semantic information at different scales, being able to adapt to different factors such as the size, shape, and background of the target, and adaptively processing different targets through the self-attention mechanism.

[0103] In S305, the multi-scale dilated attention sub-module uses a linear layer to perform feature aggregation on all head features to obtain deep image features.

[0104] The features output by all heads are concatenated, and then feature aggregation is performed by a linear layer. The finally output deep image features contain multi-scale semantic information, which can effectively improve the detection performance of the model.

[0105] In S306, the hybrid encoder performs multi-scale feature fusion on the deep image features respectively generated by the last three multi-scale dilated convolution modules in the backbone network to obtain encoder features.

[0106] In this embodiment, the multi-scale dilated attention sub-module performs different-scale SWDA at different heads using different dilation rates, and then performs feature aggregation on the outputs of all heads through a linear layer. The multi-scale self-attention mechanism is used to capture semantic information and target features at different scales, improving the detection accuracy of the model for small targets, high-density targets, and targets of different sizes in complex environments, effectively making up for the deficiencies of the lightweight downsampling sub-module in detecting complex targets, small targets, and high-density scenes. At the same time, the computational cost is further reduced through attention sparsification, and the detection performance of the model can still be guaranteed.

[0107] Optionally, the hybrid encoder of the RT-DETR model is used to perform feature fusion on the deep image features, and the decoder of the RT-DETR model is used to identify target features to obtain the target detection result of the image data, including:

[0108] The hybrid encoder performs multi-scale feature fusion using the deep image features respectively generated by the last three multi-scale dilated convolution modules in the backbone network to obtain encoder features.

[0109] In the backbone network of the RT-DETR model, the deep multi-scale dilated convolution modules are mainly used to extract deep semantic information, improve the model's ability to understand features and detection accuracy. Feature maps of different scales are respectively output to the hybrid encoder of the RT-DETR model through the multi-scale dilated convolution modules at three levels of p3, p4, and p5. As Figure 2 shown, the hybrid encoder first independently processes the feature maps of each scale through the attention-based intra-scale feature interaction module (AIFI) using the self-attention mechanism to enhance the expression ability of the features. Then, the cross-scale feature fusion module (CCFM) is used to fuse the features of different scales to generate encoder features and transmit them to the decoder.

[0110] The decoder decodes and iterates on the encoder features and performs prediction classification to obtain the object detection result of the image data.

[0111] After receiving the encoder features output by the hybrid encoder, the decoder performs initialization query selection (Uncertainty-Minimal Query Selection), then adds position embeddings to each query, and then performs decoding iteration by the multi-layer Transformer module. Each layer of the Transformer module contains a self-attention mechanism and a cross-attention mechanism, uses the prediction head to classify the object and determine the bounding box, and finally outputs the object detection result.

[0112] Embodiment 4

[0113] Figure 7 FIG. is a schematic structural diagram of a lightweight real-time object detection device according to Embodiment 4 of the present invention. In this embodiment, the lightweight real-time object detection device includes:

[0114] An image acquisition unit 810 for acquiring image data and transmitting it to the RT-DETR model;

[0115] A basic convolution unit 820 for performing standard convolution on the acquired image data;

[0116] A multi-scale dilated convolution unit 830 for performing multi-scale dilated convolution on the shallow image features generated after standard convolution;

[0117] A hybrid encoder unit 840 for performing feature fusion on the deep image features generated after multi-scale dilated convolution;

[0118] A decoder unit 850 for performing object detection after feature fusion of a hybrid encoder.

[0119] In this embodiment, by performing multi-scale dilated convolution on the shallow image features through a scale dilated convolution unit, it is possible to expand the receptive field using dilated convolution to extract surrounding context features without increasing the size of the standard convolution kernel, reducing the number of model parameters and the amount of computation. Then, multi-scale feature extraction is performed through dilated convolutions with different dilation rates, better aggregating the semantic information of the context, making object detection more accurate, enabling real-time object detection under limited computing power, and reducing resource occupancy.

[0120] Based on the above embodiments, the multi-scale dilated convolution unit further includes:

[0121] A lightweight downsampling subunit for jointly extracting joint image features containing context from shallow feature images through standard convolution and dilated convolution;

[0122] A multi-scale dilation attention sub-module for performing multi-scale feature extraction on the joint image features.

[0123] Based on the above embodiments, the lightweight downsampling subunit further includes:

[0124] A local feature extractor for extracting refined local features from shallow image features through standard convolution;

[0125] A surrounding context extractor for extracting surrounding context features from shallow image features through dilated convolution;

[0126] A joint feature extractor for fusing the refined local features and the surrounding context features;

[0127] A global context extractor for enhancing key information of the fused features.

[0128] The lightweight real-time object detection device provided by the embodiments of the present invention can execute the lightweight real-time object detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0129] Embodiment Five

[0130] Figure 8 It is a structural diagram of a server according to Embodiment Five of the present invention, Figure 8 showing a block diagram of an exemplary server 12 suitable for implementing the embodiments of the present invention. Figure 8 The shown server 12 is only an example and should not impose any limitation on the functions and scope of use of the embodiments of the present invention.

[0131] As shown Figure 8 in FIG. 1, server 12 is embodied in the form of a general-purpose computing device. The components of server 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 and the processing unit 16.

[0132] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, or a local bus using any of a variety of bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0133] Server 12 typically includes a variety of computer system readable media. Such media may be any available media that can be accessed by server 12, including both volatile and nonvolatile media, removable and non-removable media.

[0134] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Server 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 8 not shown, typically referred to as a "hard disk drive"). Although Figure 8 not shown in FIG. 1, a magnetic disk drive for reading and writing on a removable nonvolatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing on a removable nonvolatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) can be provided. In these instances, each drive can be connected to bus 18 by one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the various embodiments of the present invention.

[0135] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. The program modules 42 generally carry out the functions and / or methods of the described embodiments of the present invention.

[0136] Server 12 can also communicate with one or more external devices 14 (such as keyboards, pointing devices, monitors 24, etc.), and can also communicate with one or more devices that enable users to interact with the device / server / server 12, and / or communicate with any device that enables the server 12 to communicate with one or more other computing devices (such as network cards, modems, etc.). Such communication can be carried out through the input / output (I / O) interface 22. In addition, the server 12 can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the server 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the server 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0137] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the lightweight real-time object detection method provided by the embodiments of the present invention.

[0138] Embodiment Six

[0139] Embodiment Six of the present invention also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the lightweight real-time object detection method provided by the above embodiments when executed by a computer processor.

[0140] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0141] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0142] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0143] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., via an Internet service provider through the Internet).

[0144] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it may include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A lightweight real-time object detection method, characterized in that, Including: Input the collected image data into the RT-DETR model. The backbone network of the RT-DETR model includes a basic convolution module and a multi-scale dilated convolution module. The multi-scale dilated convolution module includes a lightweight downsampling sub-module and a multi-scale dilated attention sub-module; Use the basic convolution module to perform basic feature extraction on the collected image data to obtain shallow image features; Use the lightweight downsampling sub-module to perform dilated convolution on the shallow image features to extract context features, and fuse the context features with the local features extracted by the lightweight downsampling sub-module through standard convolution on the shallow image to obtain joint image features containing context; Use the multi-scale dilated attention sub-module to perform feature extraction at different scales on the joint image features at different dilation rates, and fuse the features at different scales to obtain deep image features; Use the hybrid encoder of the RT-DETR model to fuse the deep image features, and use the decoder of the RT-DETR model to perform target feature recognition to obtain the target detection result of the image data; The step of using the lightweight downsampling sub-module to perform dilated convolution on the shallow image features to extract context features, and fuse the context features with the local features extracted by the lightweight downsampling sub-module through standard convolution on the shallow image to obtain joint image features containing context includes: The lightweight downsampling sub-module performs standard convolution on the shallow image features to obtain refined local features, and at the same time performs dilated convolution on the shallow image features to obtain surrounding context features; The lightweight downsampling sub-module fuses the refined local features and the surrounding context features to obtain fused features; The lightweight downsampling sub-module strengthens the key information of the fused features to obtain joint image features containing context.

2. The method according to claim 1, characterized in that, The step that the lightweight downsampling sub-module fuses the refined local features and the surrounding context features to obtain fused features includes: Perform feature concatenation on the refined local features and the surrounding context features to obtain concatenated features; Perform batch normalization on the concatenated features and generate fused features through an activation function.

3. The method according to claim 1, characterized in that, The step that the lightweight downsampling sub-module strengthens the key information of the fused features to obtain joint image features containing context includes: Perform global average pooling on the fused features to obtain global features; Use a multi-layer perceptron to strengthen the key information of the global features to obtain joint image features containing context.

4. The method according to claim 1, wherein The step of using the multi-scale dilated attention sub-module to perform feature extraction at different scales on the joint image features at different dilation rates, and fuse the features at different scales to obtain deep image features includes: The multi-scale dilated attention sub-module uses multiple heads to perform multi-scale SWDA operations on the joint image features at different dilation rates to obtain head features for each head; The multi-scale dilated attention sub-module uses a linear layer to aggregate the features of all heads to obtain deep image features.

5. The method according to claim 1, characterized in that, Performing feature fusion on deep image features using the hybrid encoder of the RT-DETR model, and performing target feature recognition using the decoder of the RT-DETR model to obtain the target detection result of the image data, including: The hybrid encoder performs multi-scale feature fusion on the deep image features respectively generated by the last three multi-scale atrous convolution modules in the backbone network to obtain encoder features; The decoder performs decoding iteration and prediction classification on the encoder features to obtain the target detection result of the image data.

6. The method according to claim 1, characterized in that: The multi-scale atrous convolution module further includes a basic convolution sub-module and an activation function sub-module.

7. A lightweight real-time object detection device, characterized in that, Including: An image acquisition unit for acquiring image data and transmitting it to the RT-DETR model; A basic convolution unit for performing standard convolution on the acquired image data; A multi-scale atrous convolution unit for performing multi-scale atrous convolution on the shallow image features generated after standard convolution; A hybrid encoder unit for performing feature fusion on the deep image features generated after multi-scale atrous convolution; A decoder unit for performing target detection after the feature fusion of the hybrid encoder; The multi-scale atrous convolution unit includes: A lightweight downsampling sub-unit for jointly extracting joint image features containing context by performing standard convolution and atrous convolution on the shallow feature image features; A multi-scale dilated attention sub-unit for performing multi-scale feature extraction on the joint image features to generate deep image features; The lightweight downsampling sub-unit includes: A local feature extractor for performing standard convolution on the shallow image features to extract and refine local features; A surrounding context extractor for performing atrous convolution on the shallow image features to extract surrounding context features; A joint feature extractor for performing feature fusion on the refined local features and the surrounding context features to obtain fused features; A global context extractor for strengthening key information of the fused features to obtain joint image features containing context.

8. A server, characterized in that, The server includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the lightweight real-time target detection method as described in any one of claims 1-6.

9. A storage medium containing computer-executable instructions, the computer-executable instructions being used to execute the lightweight real-time target detection method as described in any one of claims 1-6 when executed by a computer processor.