Transform-based road end monocular 3D small object detection method
Through multi-stage depth estimation and multiple rounds of area of interest decoding methods, the depth prediction and object positioning capabilities of 3D small object detection on the road end monocular is enhanced, and the problem of insufficient detection accuracy in the existing technology is solved, thereby achieving higher detection accuracy and accuracy.
Patent Information
- Application Number
- CN202510410508.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing 3D small object detection method for road-end monocular 3D small object detection is difficult to effectively capture depth information and contextual features in complex traffic scenarios, resulting in insufficient accuracy of the model positioning and recognition of small objects, especially when detecting small objects.
Using the Transformer-based 3D small object detection method, multi-stage depth estimation and multiple rounds of region of interest refinement decoding, combined with multi-scale depth predictors and refined area decoders, feature extraction and area-level attention modeling are enhanced, and depth prediction accuracy and object positioning accuracy are improved.
It significantly improves the detection effect of small target objects in complex traffic scenarios, reduces background interference, and improves the model's positioning accuracy of small objects and the accuracy of bounding boxes.
Smart Images

Figure CN120259635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning object detection, and in particular to a method for road-side monocular 3D small object detection based on Transformer. Background Art
[0002] In the development process of autonomous driving technology, 3D small object detection has gradually become an important part of ensuring vehicle safety and decision-making accuracy. 3D small object detection is crucial for an autonomous driving system to understand a complex road environment. Especially when identifying obstacles such as pedestrians and bicycles, it can help the system perform more effective path planning and real-time obstacle avoidance. However, relying solely on in-vehicle sensors for object detection often faces problems such as limited field of view and occlusion, resulting in a decline in detection performance. To make up for the deficiencies of the in-vehicle perception system, vehicle-road collaborative technology has emerged. By interacting the road-side devices with in-vehicle sensors, the perception range of autonomous driving vehicles is significantly extended. Compared with in-vehicle cameras, road-side cameras are installed at a higher position, which can avoid the problem of visual blind spots caused by vehicle occlusion, ensure that the road conditions can be monitored from a wider angle, discover potential risk factors in advance, and gain more sufficient reaction time for autonomous driving vehicles.
[0003] However, most existing road-side methods rely on simple convolutional layers to predict depth-related features. Although this method based on simple convolutional operations can extract local features, due to its limited receptive field, it is more difficult to capture effective context information from a monocular perspective. This will lead to insufficient ability of the model to extract depth information, especially when detecting small objects, it is difficult to form sufficient feature representations, and thus unable to accurately identify the positions of small objects. Moreover, there are a large number of long-distance detection requirements in road-side perception, and the visual information of target objects is sparser under simple convolutional operations, further exacerbating the problem of insufficient features. On the other hand, existing original Transformers are only based on global attention modeling, relying too much on global information, resulting in the lack of the ability to gradually optimize complex backgrounds in the existing object modeling process, and unable to effectively exclude irrelevant background noise. At the same time, the complexity of road-side scenes makes it easy for monocular models to generate errors when locating small objects under single-dimensional features, resulting in insufficient ability of the model to locate bounding boxes, and increasing the risk of misidentifying small object detection. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose a method for road-side monocular 3D small object detection based on Transformer. Through feature enhancement based on multi-stage depth estimation and multi-round region of interest (RoI) refinement decoding, while improving the accuracy of single-view depth prediction, it strengthens object localization guided by regional-level attention modeling, thereby enhancing the detection effect of small target objects in complex traffic scenarios.
[0005] To achieve the above object, the technical solution provided by the present invention is: a roadside monocular 3D small object detection method based on Transformer, comprising the following steps:
[0006] S1: Obtain roadside image data, and preprocess the data to obtain an enhanced data set;
[0007] S2: Extract features from the image data in the enhanced data set to extract multi-scale image features;
[0008] S3: Design a pyramid depth predictor, take the image data of the enhanced data set as input, extract depth information of different receptive fields through multi-stage hybrid dilated convolutions, and finally aggregate the depth information of each stage to form multi-scale depth features;
[0009] S4: Input the multi-scale image features in step S2 and the multi-scale depth features in step S3 into the Transformer-based encoder respectively, encode the image data from two perspectives of depth and vision to form two types of features, namely visual encoding features and depth encoding features; subsequently, input these two types of features into the Transformer-based depth cross-decoder, and process them through the depth cross-attention layer and the self-attention layer to obtain the decoded features;
[0010] S5: Input the decoded features into the object detection network to obtain the initial object category label and the initial object bounding box coordinates;
[0011] S6: Design a refined region decoder to refine the initial object category label and the initial object bounding box coordinates in multiple rounds of loops; in each round, the refined region decoder extracts the region of interest features using the detection results of the previous round, gradually optimizes the bounding box and the category, and obtains the final object category label and the object bounding box coordinates.
[0012] Further, in step S1, the preprocessing includes rotating, flipping, translating, and scaling the data.
[0013] Further, in step S2, use the Resnet-50 neural network to extract features from the enhanced data set to form multi-scale image features x.
[0014] Further, in step S3, use a multi-stage pyramid depth predictor to extract depth features from the image data to obtain multi-scale depth features. The specific steps are as follows:
[0015] Input the image data into the pyramid depth predictor. In the first stage, perform downsampling and local feature extraction on the image through multiple convolutional layers to obtain downsampled features;
[0016] In the second, third, and fourth stages, the output features of the previous stage are used as the input features; in each stage, the input features are first downsampled and then passed into a hybrid dilated convolution module for feature extraction. The core of the hybrid dilated convolution module is to linearly aggregate multiple dilated convolutions with different dilation rates. For each dilated convolution, when a two-dimensional input feature f[k] is given, its corresponding dilated convolution output g[k] is defined as:
[0017]
[0018] where L is the length of the convolution kernel, h[l] is the weight parameter of the convolution kernel at the l-th position, and f[k + r·l] represents the skip sampling of the input feature through the dilation rate r, selecting the input feature corresponding to the l-th convolution kernel position;
[0019] When the input feature X passes through the hybrid dilated convolution module, the output feature is obtained
[0020]
[0021] where GELU(·) is the activation function that can effectively alleviate the vanishing gradient problem, BN(·) is the batch normalization layer, and Conv r (·) is a 3*3 dilated convolution with a dilation rate of r;
[0022] Before the features enter each stage, they are concatenated with the input of the pooled image data to reduce the loss of spatial information caused by the reduction of the feature size; after four-stage processing by the pyramid depth predictor, multi-scale depth features with multi-level information are formed.
[0023] Furthermore, the specific operation steps of step S4 are as follows:
[0024] The multi-scale image features and multi-scale depth features are respectively input into two Transformer-based encoders to obtain visual encoding features and depth encoding features through encoding;
[0025] The visual encoding features and depth encoding features are passed into a Transformer-based depth cross-decoder. Each layer of this depth cross-decoder includes a depth cross-attention layer, a self-attention layer, a visual cross-attention layer, and a feed-forward neural network layer. A set of learnable object queries q is used to capture object information from the depth cross-decoder, and the two types of feature information are seamlessly integrated to form decoded features;
[0026] In each layer, first, depth information is extracted from the depth-encoded features through the depth cross-attention layer. In this process, the object query and the depth-encoded features are linearly transformed respectively to generate the initial query Q q , depth key K D , and depth value V D . Subsequently, they are input into the depth cross-attention layer to obtain the depth query Q D :
[0027] Q D = Attn(Q q , K D , V D )
[0028] Subsequently, Q D is input into the self-attention layer for further interaction to avoid redundant predictions of the bounding boxes of the same object by each query; immediately afterwards, the output of the self-attention layer is passed to the visual cross-attention layer to obtain the visual information of the visual-encoded features. In this process, the visual-encoded features are linearly transformed respectively to generate the visual key K V and visual value V V . Subsequently, they are input into the visual cross-attention layer to obtain the generated visual query
[0029]
[0030] Finally, is passed into the feed-forward neural network layer to extract deeper feature information Q DV :
[0031]
[0032] The output Q DV of each round of the decoder is used as the input of the next round. Through multiple loops, the two types of feature information are seamlessly integrated and fully incorporated into the object query q
[0033] Furthermore, in step S6, the preliminary object category label O cls (0) and the preliminary object bounding box coordinates O box (0) are input into the refinement region decoder. Through multiple rounds of refinement, the final object category label and object bounding box coordinates are obtained. Among them, the specific operation steps in the i-th round of the refinement region decoder are as follows
[0034] Obtain the bounding box output O box (i*1) of the i*1-th round, and extract the regional visual feature R(i):
[0035] R(i) = f ext (x, RoI(Obox (i*1), α(i)))
[0036] In the formula, f ext (·) represents the feature extraction operation based on RoIAlign, x represents the multi-scale image features, RoI(·) represents the RoI calculation, and α(i) is the magnification parameter in the i-th round, which is used to adjust the magnification factor of the bounding box;
[0037] By performing cross-attention to model the relationship between the regional visual feature R(i) and the previous round of attention output M dec (i - 1). In this process, R(i) and M dec (i - 1) are respectively linearly transformed to generate the regional query Q R , key K M and value V M , and finally calculate the decoded regional feature M r (i):
[0038] M r (i) = Attn(Q R , K M , V M )
[0039] Subsequently, calculate the attention output M dec (i):
[0040]
[0041] In the formula, represents the feature concatenation operation;
[0042] In the refined regional decoder, for the decoding in the i-th round, detect the object according to the following formula:
[0043] O cls (i) = F cls (M dec (i))
[0044] O box (i) = F box (M dec (i) + O box (i - 1))
[0045] In the formula, O cls (i) and O box (i) respectively represent the object category label and the bounding box coordinates in the i-th round, and O box (i - 1) represents the bounding box coordinates in the (i - 1)-th round, which are obtained through the object detection network;
[0046] Finally, after multiple rounds of gradual optimization of the object bounding box and category, the final object category label and object bounding box coordinates are obtained.
[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0048] 1. The present invention designs a multi-scale pyramid depth predictor, which uses hybrid dilated convolutions to extract multi-scale information with different receptive fields and aggregates the depth information at each stage, effectively alleviating the problem of information loss in feature maps at different scales.
[0049] 2. The present invention designs a multi-round refined region decoder. Through the cyclic region decoding process, the detection process is guided to focus on the regions related to the bounding box, significantly reducing the interference of irrelevant backgrounds, thereby more accurately locating the position and boundary of small objects.
[0050] 3. The present invention enhances the detection effect of small target objects in complex traffic scenarios, is applicable to various road target detection tasks, has a wide application prospect, and is worthy of promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the overall framework diagram of the method of the present invention.
[0052] Figure 2 is the internal structure schematic diagram of the hybrid dilated convolution module.
[0053] Figure 3 is the effect diagram of the output target box of the refined region decoder. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0055] As Figures 1 to 3 shown, this embodiment discloses a road-side monocular 3D small object detection method based on Transformer, which enhances visual features through a more accurate depth pyramid predictor and a refined region decoder, and improves the detection ability for small objects such as pedestrians and bicycles. The specific situation is as follows:
[0056] 1) Pre-data processing, including performing contrast transformation, rotation, flipping, translation, scaling, and adding noise to the original data set to obtain an enhanced data set; among them, using contrast transformation to enhance the image contrast, and using rotation, flipping, translation, scaling, and adding noise to increase the physical feature diversity and complexity of the defective images, so as to obtain the enhanced data set as the training data set.
[0057] 2) Use RestNet-50 as the backbone network, which will enhance the given image I of the dataset cam ∈R H×W×3 , where R represents the set of real numbers, H and W represent the height and width of the image respectively, and convert it into multi-scale image features F = {f 1 / 8 , f 1 / 16 , f 1 / 32}, where f 1 / 8 , f 1 / 16 , f 1 / 32 represent the image features with downsampling ratios of 1 / 8, 1 / 16, and 1 / 32 respectively.
[0058] 3) Input the given picture I of the enhanced dataset cam ∈R H×W×3 into the pyramid depth predictor. The pyramid depth predictor first passes the image into the convolutional layer. The image is first downsampled by a 3×3 convolutional layer with a stride of 2 in the first stage, and then local feature extraction is performed by two 3×3 convolutions with a stride of 1, and finally a feature of size H / 2×W / 2×C1 is obtained, where C1 represents the number of channels. Before entering the second stage, the feature is concatenated with the pooled image to reduce the loss of spatial information caused by the reduction of the feature size. Then the concatenated feature is input into the downsampling layer with a 3×3 convolution to obtain a feature map of size H / 4×W / 4×C2 for the concatenation operation in the next stage, where C2 represents the number of channels. The downsampled feature is simultaneously input into the hybrid dilated convolution module to obtain multi-scale features, and is passed to the concatenation layer before the next stage to transfer multi-scale information from different receptive fields, promote the transfer of information between stages, and thus maintain the consistency of features. After four stages of processing, a multi-scale depth feature of size H / 16×W / 16×C4 with multi-level information is obtained, where C4 represents the number of channels.
[0059] The hybrid dilated convolution module consists of multiple dilated convolutions with different dilation rates to extract multi-scale local features. The dilated convolution can increase the receptive field without changing the size of the output feature map. However, directly using multiple identical dilated convolutions may cause a large number of image pixels not to participate in the calculation. Therefore, the present invention inserts several consecutive dilated convolutions with different dilation rates at each stage to cover the entire image and achieve the aggregation of multi-scale context information. For each dilated convolution, when a two-dimensional input feature f[k] is given, its corresponding dilated convolution output g[k] can be defined as:
[0060]
[0061] Wherein, L is the length of the convolutional kernel, h[l] is the weight parameter of the convolutional kernel at the l-th position, and f[k + r·l] represents performing skip sampling on the input feature through the dilation rate r and selecting the input feature corresponding to the l-th convolutional kernel position;
[0062] When the input feature X passes through the hybrid dilated convolution module, the output feature is obtained
[0063]
[0064] Wherein, GELU(·) is the activation function, which can effectively alleviate the problem of gradient vanishing, BN(·) is the batch normalization layer, and Conv r (·) is a 3*3 dilated convolution with a dilation rate of r.
[0065] 4) Input the multi-scale image features in step 2) and the multi-scale depth features in step 3) into the Transformer-based encoder respectively, encode the image data from both the depth and visual perspectives to form two types of features: visual encoding features and depth encoding features; subsequently, input these two types of features into the Transformer-based depth cross-decoder, and process them through the feature cross-attention layer and the self-attention layer to obtain the decoded features; specifically as follows:
[0066] Input the multi-scale image features and the multi-scale depth features into two Transformer-based encoders (which can be called the depth encoder and the visual encoder) respectively, and encode to obtain the visual encoding features and the depth encoding features;
[0067] Input the visual encoding features and the depth encoding features into the Transformer-based depth cross-decoder. Each layer of this depth cross-decoder includes a depth cross-attention layer, a self-attention layer, a visual cross-attention layer, and a feed-forward neural network layer. Use a set of learnable object queries q to capture object information from the depth cross-decoder, and seamlessly integrate the two types of feature information to form the decoded features;
[0068] In each layer, first extract the depth information from the depth encoding features through the depth cross-attention layer. In this process, the object query and the depth encoding features are respectively linearly transformed to generate the initial query Q q , the depth key K D and the depth value V D , and then input them into the depth cross-attention layer to obtain the depth query Q D :
[0069] Q D = Attn(Q q , K D , V D )
[0070] Subsequently, Q D is input into the self-attention layer for further interaction, avoiding redundant predictions of the bounding boxes of the same object by each query; immediately afterwards, the output of the self-attention layer is passed to the visual cross-attention layer to obtain the visual information of the visual encoding features. In this process, the visual encoding features are respectively linearly transformed to generate the visual key K V and the visual value V V , and then input into the visual cross-attention layer to obtain the generated visual query
[0071]
[0072] Finally, is passed into the feed-forward neural network layer to extract deeper feature information Q DV :
[0073]
[0074] Taking the output Q DV of each round of the decoder as the input of the next round, through multiple loops, the two types of feature information are seamlessly integrated and fully incorporated into the object query q.
[0075] 5) After decoding is completed, the decoded feature M dec (0) is input into the object detection network to obtain the preliminary object class label and the preliminary bounding box coordinates:
[0076] O cls (0) = Head cls (M dec (0))
[0077] O box (0) = Head box (M dec (0))
[0078] where, O cls (0) represents the preliminary object class label, O box (0) represents the preliminary bounding box coordinates of the object, and Head cls (·), Head box (·) respectively represent the class detector and the bounding box detector of the object detection network.
[0079] 6) Design a multi-round refined region decoder, which consists of multiple refined region sub-decoders. Each refined region sub-decoder uses the output of the target box in the previous round to select the region of interest, and then conducts attention modeling to obtain a more fine-grained result and improve the localization ability of the model's bounding box. At the beginning of each round of the refined region sub-decoder, it receives the output of the previous stage to achieve cyclic refinement. In each round, the bounding box generated in the previous stage is used to obtain the RoI (region of interest) to extract the corresponding region features. Then, according to the attention decoding output of the previous round, the region features are converted into a more refined attention decoding output.
[0080] Obtain the bounding box output O box (i - 1) of the (i - 1)-th round in the i-th round, and extract the regional visual feature R(i):
[0081] R(i) = f ext (x, RoI(O box (i - 1), α(i)))
[0082] In the formula, f ext (·) represents the feature extraction operation based on RoIAlign, x represents the multi-scale image features, RoI(·) represents the RoI calculation, and α(i) is the magnification parameter of the i-th round, which is used to adjust the magnification factor of the bounding box;
[0083] Then, cross-attention is performed to model the relationship between the regional visual feature and the attention output M dec (i - 1) of the previous round. In this process, R(i) and M dec (i - 1) are respectively linearly transformed to generate the query Q R , the key K M , and the value V M , and finally calculate the decoded region feature M r (i):
[0084] M r (i) = Attn(Q R , K M , V M )
[0085] Subsequently, calculate the attention output M dec (i) of this round:
[0086]
[0087] Generally speaking, in the refined region decoder, for the refined region sub-decoder of the i-th round, we detect the object according to the following formula:
[0088] O cls (i) = Fcls (M dec (i))
[0089] O box (i) = F box (M dec (i) + O box (i - 1))
[0090] Wherein, O cls (i) and O box (i) respectively represent the object category label and the bounding box coordinates in the i-th round, and O box (i - 1) represents the bounding box coordinates in the (i - 1)-th round, obtained through the object detection network. Since the detection results in the initial rounds may not be reliable enough, we tend to magnify the perspective of each bounding box more at the beginning of the loop for regional feature extraction to make full use of the context information and capture the target object more accurately. In the subsequent processing rounds, we gradually reduce the size of the extracted regional features to obtain more local details, so as to achieve more accurate detection.
[0091] Finally, after gradually optimizing the object bounding box and category through multiple rounds, the final object category label and bounding box coordinates are obtained.
[0092] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A Transformer-based roadside monocular 3D small object detection method, characterized in that, It includes the following steps: S1: Obtain roadside image data and preprocess the data to obtain an enhanced dataset; S2: Extract features from the image data in the enhanced dataset to extract multi-scale image features; S3: Design a pyramid depth predictor. Take the image data of the enhanced dataset as input, extract depth information of different receptive fields through multi-stage hybrid dilated convolutions, and finally aggregate the depth information of each stage to form multi-scale depth features; S4: Input the multi-scale image features in step S2 and the multi-scale depth features in step S3 into the Transformer-based encoder respectively to encode the image data from both the depth and visual perspectives to form two types of features: visual encoding features and depth encoding features; Subsequently, input these two types of features into the Transformer-based depth cross-decoder and process them through the depth cross-attention layer and self-attention layer to obtain the decoded features; S5: Input the decoded features into the object detection network to obtain the preliminary object category labels and the preliminary object bounding box coordinates; S6: Design a refined region decoder to refine the preliminary object category labels and the preliminary object bounding box coordinates in multiple rounds of loops; In each round, the refined region decoder extracts the region of interest features using the detection results of the previous round, gradually optimizes the bounding box and category, and obtains the final object category labels and the object bounding box coordinates.
2. The method for roadside monocular 3D small object detection based on Transformer according to claim 1, wherein, In step S1, the preprocessing includes rotating, flipping, translating, and scaling the data.
3. The method for road-end monocular 3D small object detection based on Transformer according to claim 1, characterized in that, In step S2, use the Resnet-50 neural network to extract features from the enhanced dataset to form multi-scale image features x.
4. The method for end-side monocular 3D small object detection based on Transformer according to claim 1, wherein In step S3, use a multi-stage pyramid depth predictor to extract depth features from the image data to obtain multi-scale depth features. The specific steps are as follows: Input the image data into the pyramid depth predictor. In the first stage, downsample the image and extract local features through multiple convolutional layers to obtain the downsampled features; In the second, third, and fourth stages, take the output features of the previous stage as the input features; In each stage, first perform downsampling on the input features, and then pass them into the hybrid dilated convolution module for feature extraction. The core of the hybrid dilated convolution module is to linearly aggregate multiple dilated convolutions with different dilation rates. For each dilated convolution, when a two-dimensional input feature f[k] is given, its corresponding dilated convolution output g[k] is defined as: In the formula, L is the length of the convolution kernel, h[l] is the weight parameter of the convolution kernel at the l-th position, and f[k + r·l] represents the skip sampling of the input feature through the dilation rate r to select the input feature corresponding to the l-th convolution kernel position; When the input feature X passes through the hybrid dilated convolution module, the output feature is obtained Wherein, GELU(·) is an activation function that can effectively alleviate the vanishing gradient problem, BN(·) is a batch normalization layer, and Conv r (·) is a 3×3 dilated convolution with a dilation rate of r; Before the features enter each stage, they will be concatenated with the input of the image data after pooling to reduce the loss of spatial information caused by the reduction of the feature size; After four-stage processing by the pyramid depth predictor, multi-scale depth features with multi-level information are formed.
5. The method for road-end monocular 3D small object detection based on Transformer according to claim 1, wherein The specific operation steps of step S4 are as follows: The multi-scale image features and multi-scale depth features are respectively input into two Transformer-based encoders, and visual encoding features and depth encoding features are obtained through encoding; The visual encoding features and depth encoding features are fed into a Transformer-based depth cross-decoder. Each layer of this depth cross-decoder includes a depth cross-attention layer, a self-attention layer, a visual cross-attention layer, and a feed-forward neural network layer. A set of learnable object queries q is used to capture object information from the depth cross-decoder, and the two types of feature information are seamlessly integrated to form decoded features; In each layer, first, depth information is extracted from the depth-encoded features through the depth cross-attention layer. In this process, the object query and the depth-encoded features are linearly transformed respectively to generate the initial query Q q , the depth key K D and the depth value V D , and then they are input into the depth cross-attention layer to obtain the depth query Q D : Q D = Attn(Q q , K D , V D ) Then, Q D It is input into the self-attention layer and further interacts to avoid redundant predictions of the bounding box of the same object for each query; then the output of the self-attention layer is The visual information is passed to the visual cross-attention layer to obtain the visual encoding features. In this process, the visual encoding features are linearly transformed to generate the visual key K V and visual value V V , which is then fed into the visual cross-attention layer to generate the visual query Finally, is fed into the feedforward neural network layer to extract deeper feature information Q DV : Take the output Q of each round of the decoder DV as the input for the next round. Through multiple loops, the two types of feature information are seamlessly integrated and fully incorporated into the object query q.
6. The method for roadside monocular 3D small object detection based on Transformer according to claim 1, wherein In step S6, the initial object category label O cls (0) and the initial object bounding box coordinates O box (0) are input into the refinement region decoder, and through multiple rounds of refinement, the final object category label and object bounding box coordinates are obtained. Among them, the specific operation steps in the i-th round of the refinement region decoder are as follows: Obtain the bounding box output O of the (i - 1)-th round box (i - 1), extract the regional visual feature R(i): R(i) = f ext (x, RoI(O box (i - 1), α(i))) where f ext (·) represents the feature extraction operation based on RoIAlign, x represents the multi-scale image features, RoI(·) represents the RoI calculation, and α(i) is the magnification parameter in the i-th round, which is used to adjust the magnification factor of the bounding box; Model the relationship between the regional visual feature R(i) and the previous round of attention output M dec (i - 1) through cross-attention. In this process, R(i) and M dec (i - 1) are linearly transformed respectively to generate the regional query Q R , the key K M and the value V M , and finally calculate the decoded regional feature M r (i) of the current round: M r (i) = Attn(Q R , K M , V M ) Subsequently, the attention output M for this round is calculated dec (i): In the formula, represents a feature splicing operation; In the refinement region decoder, for the decoding in the i-th round, objects are detected according to the following formula: O cls (i) = F cls (M dec (i)) O box (i) = F box (M dec (i) + O box (i - 1)) where O cls (i) and O box (i) respectively represent the object class label and the bounding box coordinates in the i-th round, and O box (i - 1) represents the bounding box coordinates in the (i - 1)-th round, which are obtained through the object detection network; Finally, the object bounding boxes and categories are gradually optimized through multiple rounds to obtain the final object category labels and object bounding box coordinates.