An object detection method and device based on an adaptive decoder

By introducing adaptive decoder and AdaMixer model into the object detection framework, using 3D sampling and adaptive hybrid technology, the existing object detection framework relies on human prior knowledge and query-based detection framework insufficient performance, achieving efficient, fast and accurate object detection results.

CN114612716BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210227694.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-06-13
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

The existing object detection framework requires manual modules that rely on human prior knowledge to adjust the parameters in detail, and the query-based Transformer object detection framework has problems such as limited spatial resolution, poor detection performance of small objects, and slow convergence speed.

Method used

A target detection method based on an adaptive decoder is designed, and the AdaMixer model is constructed. Multi-scale features of the picture are sampled through the 3D sampling feature space. Based on the query mechanism, the sampling point position and feature decoding are adaptively adjusted through the encoder and the decoder, and combined with FFN to complete the query enhancement to realize the detection of the query position.

Benefits of technology

It achieves better target detection effect, without the need for additional feature encoder outside the feature extraction network, can directly, efficiently, quickly and accurately generate the enclosure frames and categories of the target object, and has faster convergence speed and stronger adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612716B_ABST
    Figure CN114612716B_ABST
Patent Text Reader

Abstract

An object detection method and device based on an adaptive decoder, which constructs an object detection model AdaMixer, including a network configuration stage, a training stage, and a testing stage. Different-sized feature maps obtained in cooperation with a backbone network are combined into a 3D feature space, where efficient feature sampling is performed, and the sampled features are enhanced by adaptively coordinating the spatial information and position information of the query amount, thereby realizing the object detection task. The present invention effectively utilizes the information in the query amount through an adaptive module for different picture query amounts, avoids redundant network structures, saves computational amounts, and enables the network to converge quickly and stably. The sampling of the 3D feature space is introduced to efficiently encode the position information and semantic information, which can better cooperate with the adaptive module to flexibly, efficiently, quickly, and accurately complete the object detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer software, relates to the technology of sequential action detection, and specifically provides a target detection method and device based on an adaptive decoder. Background Art

[0002] Object detection has always been a fundamental but difficult task in the field of computer vision. The goal of this task is to find the positions of different objects in an image and classify them. So far, the two main paradigms of object detection are divided into two types:

[0003] The first paradigm is a dense object detector, which is based on the idea of a sliding window. It assumes that objects in an image may appear densely and evenly at any spatial position in the image. In the era of deep learning, object detectors based on this assumption can well cover objects that may be targets. There are many well-known works based on this dense prior assumption, but their disadvantage is that they need to densely generate multi-scale anchor boxes to exhaustively generate proposal regions in the feature map or directly classify and locate objects, which consumes a lot of computing resources and is prone to redundant detection results.

[0004] The second paradigm is a query-based object detector, mainly represented by the recently emerged DETR. It formalizes the object detection problem as a direct set prediction problem. First, it needs to use an encoder and a decoder based on the Transformer structure to generate many boxes to predict the positions of targets, and then perform bipartite graph matching based on these predicted boxes and the ground truth boxes. Although the performance of this paradigm is comparable to the baseline of Faster R-CNN, it still has problems such as limited spatial resolution, insufficient small object detection ability, and slow convergence speed. Moreover, these query-based detectors all require additional feature encoders outside the feature extraction network, and the computational overhead brought by these feature encoders is huge. There are also some works that use some of the dense assumptions in the first paradigm to solve some of the problems here, but at the same time introduce some problems in the first paradigm. These query-based detectors cannot be applied to practical applications due to various problems mentioned above for the time being. Summary of the Invention

[0005] The problems to be solved by the present invention are as follows: Many existing object detection frameworks need to rely on many manual modules based on human prior knowledge and require fine-tuning of parameters; moreover, the newly emerged query-based Transformer object detection framework also has problems such as limited spatial resolution, poor small object detection performance, and slow convergence speed, and does not fully utilize the query information, so the object detection effect needs to be improved.

[0006] The technical solution of the present invention is as follows: A target detection method based on an adaptive decoder constructs a target detection model AdaMixer, samples multi-scale features of an image according to a 3D sampling feature space, and adaptively adjusts the positions of sampling points and feature decoding based on the query mechanism through an encoder and a decoder according to the spatial position information and semantic content information of the query, and then cooperates with the FFN to complete the enhancement of the query to realize the detection of the query position. The implementation of the target detection network includes a 3D feature generation stage, a network configuration stage, a training stage, and a testing stage:

[0007] 1) 3D feature generation stage: Use a backbone network to extract features from training example images. For each input image, based on the feature maps with different lengths, widths, and numbers of channels output by different stages of the backbone network, a 3D feature space is obtained for subsequent sampling processing therein;

[0008] 2) Network configuration stage: Based on the initial query configuration and the decoder, establish the target detection model AdaMixer, including the following configurations:

[0009] 2.1) Initial query configuration: Initialize the encoding for the input feature map to generate N initial query quantities. The query quantities include an initial semantic vector q 0 , and the initial position vector (x, y, z, r) of the corresponding query sampling point. (x, y, z) are the coordinates of the sampling point in the 3D feature space, and r is the base-2 logarithm of the aspect ratio of the feature map. The initial semantic vector q 0 is randomly sampled from the standard normal distribution N(0, 1), and the initial position vector (x, y, z, r) is set to cover the entire feature map;

[0010] 2.2) Decoder: The input of the decoder is the query quantity encoded in 2.1), and the output is the query quantity of the same format after being optimized by the decoder. The decoder contains the following modules for adaptively using the semantic information and position information in the query:

[0011] 2.2.1) Multi-head self-attention module: Input N query quantities into a multi-head self-attention module, attach the position information in the form of sine to the semantic vector, and add the intersection over foreground ratio IoF as a bias to the weight of the attention. After passing through this multi-head self-attention module, an enhanced semantic vector output q is obtained;

[0012] 2.2.2) 3D sampling module: Use q obtained in 2.2.1), after a linear layer transformation, use the semantic information of q to obtain a set of displacements of P in sampling points, and combine with the corresponding position vector quadruple in the query quantity to obtain P inFor a 3D feature space sampling point coordinate P, first perform two-point interpolation in the (x, y) feature space according to the sampling point coordinate P to obtain a feature matrix without considering the z-axis weight, and then perform Gaussian weight interpolation on the z-axis to obtain the complete sampled feature matrix X;

[0013] 2.2.3) Adaptive mixing module: Decode the feature matrix X adaptively, and perform adaptive mixing on the sampled feature matrix X in two steps, namely adaptive semantic channel mixing and adaptive spatial information mixing. In adaptive semantic channel mixing, use a dynamic weight matrix based on q to enhance the channel semantics of the feature matrix X on the feature channels; in adaptive spatial information mixing, use a dynamic weight matrix based on q to enhance the spatial information of the feature matrix X on the spatial information; finally, obtain the feature matrix X' with enhanced information;

[0014] 2.2.4) FFN module: For the updated feature matrix X' output in 2.2.3), combine its position vector to update the semantic vector and position vector in the overall query, that is, flatten the feature matrix X' obtained in 2.2.3), pass it through a group of FFNs, convert its number of channels to the same number of channels as q, and obtain the updated semantic vector q' in the query; according to the updated semantic vector q', pass it through another group of FFNs to obtain a group of updated position vectors (x', y', z', r') in the query;

[0015] 2.3) After obtaining the final query semantic vector q' and position vectors (x', y', z', r'), send q' into an FFN classification network to obtain the classification result, and translate (x', y', z', r') into the coordinates of the bounding box to obtain the result of the bounding box;

[0016] 3) Training stage: Train the configured network model with training data, use the combination of focal loss, L1 loss and GIoU loss as the loss function, use the AdamW optimizer, and update the network parameters through the backpropagation algorithm. Continuously repeat steps 1) and 2) until the number of iterations is reached;

[0017] 4) Training stage: Input the image features of the data to be tested into the trained AdaMixer model. According to the method in 2.3), obtain the final target classification result and the target bounding box position, verify the effect of the trained AdaMixer model, and use the AdaMixer model that achieves the required target detection effect as the final obtained target detection model for target detection.

[0018] The present invention also provides an object detection device based on an adaptive decoder, which has a computer-readable storage medium. A computer program is configured in the computer storage medium. The computer program is used to implement the above-mentioned object detection model AdaMixer, and when the computer program is executed, the above-mentioned object tracking method is implemented.

[0019] The present invention proposes a new decoding method, which can better utilize the spatial position information and semantic content information of queries in a query-based Transformer object detection framework. The designed decoder is more adaptive to queries, can achieve better results, and does not require an additional feature encoder outside the feature extraction network, and can directly, efficiently, quickly, and accurately generate the bounding box of the target object and its category.

[0020] The present invention has the following advantages compared with the prior art

[0021] The present invention proposes a simple and accurate detector with the ability to adapt to different query amounts for different pictures. Without manual modules such as anchor boxes, dense matching, and non-maximum suppression that rely on human prior knowledge, it can fully utilize the spatial information and semantic information in the query amount, is easy to debug, and has a faster convergence speed during training, as Figure 4 shown.

[0022] The decoder proposed by the present invention can adaptively sample features based on the query amount, and dynamically decode features by feature channel mixing and spatial information mixing. After experiments, it is found that the feature information and spatial information are very helpful for the accuracy of the results. The detection model designed by the present invention effectively utilizes the semantic and position information contained in the query amount, helps to improve the decoder's understanding of semantic information and spatial information, and has stronger adaptability.

[0023] The present invention proposes 3D feature sampling, which can effectively encode feature maps with different numbers of feature channels and integrate more effective information into the query amount. Based on this 3D feature sampling strategy, the present invention can obtain multi-scale feature information only by using the features of different channels output by a backbone network, and can adapt to the scale transformation of different objects, which helps to simplify the network and consume less computing resources without any additional network output heads.

[0024] The present invention has the characteristics of high adaptability, efficiency, speed, and accuracy in object detection tasks. Compared with existing methods, the present invention has better performance in mainstream data sets and practical applications. Brief Description of the Drawings

[0025] Figure 1 is the system framework diagram used in the present invention.

[0026] Figure 2 It is a schematic diagram of the 3D feature space of the present invention.

[0027] Figure 3 It is a schematic diagram of the adaptive feature mixing module of the present invention.

[0028] Figure 4 It is a diagram showing the convergence speed of the present invention.

[0029] Figure 5 It is a diagram showing the results of the 1x training strategy of the present invention.

[0030] Figure 6 It is a diagram showing the results of the 3x training strategy of the present invention.

[0031] Figure 7 It is a schematic diagram of the overall process of the present invention. Detailed implementation manners

[0032] The present invention is a query-based adaptive object detection method, which constructs an AdaMixer model, combines self-adaptive 3D sampling technology and self-adaptive channel and spatial mixing technology, can improve the performance of the decoder in the query-based object detector, and adaptively adjusts features through the attention mechanism to achieve a high-performance and high-efficiency image object detector. In the method of the present invention, the implementation of the AdaMixer model includes a 3D feature generation stage, a network configuration stage, a training stage, and a testing stage, as Figure 7 shown, and the specific description is as follows.

[0033] 1) 3D feature generation stage: Using the feature maps with different lengths, widths, and numbers of channels output at different stages of the backbone network ResNet-50 based on the input image, a 3D feature space is defined. First, each feature map is transformed to the same number of channels d feat , preferably d feat = 256, and each feature Figure 1 map is given a label j. For the feature map j, its z-axis coordinate can be calculated as:

[0034] z j = log 2 (s j / s base )

[0035] where s j is the downsampling strip width of the j-th feature map, and s base is the width of the basic downsampling strip, which is a preset amount for image downsampling. In the present invention, for the second to fifth stages (C 2 ~C 5The feature maps are linearly transformed to stretch them to the same length and width and unify the number of channels to d. feat , and stacked in a three-dimensional space, as Figure 2 , to obtain the 3D feature space of the present invention.

[0036] 2) In the network configuration stage, based on the encoder and decoder, a target detection model, i.e., the AdaMixer model of the present invention, is established, as Figure 1 shown. The model includes the following configurations:

[0037] 2.1) Initial query generation configuration: The present invention first configures the encoder to initialize the encoding of the input feature maps to obtain N initial query quantities for the initialization of the subsequent framework learning, specifically as follows:

[0038] 1. Query definition:

[0039] In the present invention, in order to enable the encoder to achieve the expected effect, a query quantity containing semantic information and position information is customized. Each query quantity contains a semantic vector with a vector dimension of d q , and a position vector (x, y, z, r) corresponding to the query sampling point, where (x, y, z) are the coordinates in the 3D feature space of the sampling point. r is the base-2 logarithm of the aspect ratio of the feature map. The main reason for such a definition is to facilitate encoding and conversion with the bounding box of the target.

[0040] The position vector in the query quantity and the bounding box can be mutually converted. The specific conversion relationship is:

[0041] x B = s base · x, y B = s base · y

[0042] w B = s base · 2 z-r , h B = s base · 2 z+r

[0043] That is, a position vector (x, y, z, r) can be converted into the information of a bounding box (x B , y B , w B , h B ). (x B , y B ) corresponds to the center point of the bounding box, and (w B , h B ) corresponds to the width and height of the bounding box. s baseis the width of the basic image downsampling strip, which is determined according to the scaling factor of the largest feature map. In the example of the present invention, s base is taken as 4.

[0044] 2. Initial query configuration:

[0045] Initialize the encoding for the input feature map to generate N initial query quantities. The semantic vector q in each query 0 is initialized as a d q -dimensional tensor, where the value of each dimension is randomly sampled from the standard normal distribution N(0,1). Assuming the length of the input image is H and the width is W, the position vector of each query sampling point is initialized as a quadruple that can be transformed into a bounding box covering the entire image which is a 4-dimensional vector. The initial query quantity only depends on the width and height of the image. The initial query quantity does not come from the encoded image but is related to the width and height of the encoded image.

[0046] 2.2) Decoder: The input at this stage is the query quantity in the format generated in 2.1), and the output is the query quantity in the same format after being optimized by the decoder. The network is deepened by stacking M layers in advance to achieve better results. In the example of the present invention, we take M = 6 and stack 6 layers of decoders. For the input query quantity, the decoder of the present invention first sends it into a multi-head self-attention module to obtain an enhanced semantic query quantity Then, after using this q to pass through a linear layer transformation, P in sampling point sets can be obtained Using this sampling point set, sampling can be performed in the 3D feature space defined in the present invention. The feature space is divided into g groups, and the number of channels in each group of feature spaces after grouping is d feat / g. P in points are sampled in each group, and a feature matrix can be obtained The feature matrix of each group is where C = d feat / g. For adaptive mixing, first perform adaptive semantic channel mixing to obtain the feature output, and then perform adaptive spatial information mixing to obtain Finally, use this output to send it into the FFN network to update the semantic vector and position vector of the query, and obtain the query quantity q' and position vector (x', y', z', r') in the same format as the input but after adaptive optimization. The specific implementation of the decoder of the present invention is as follows.

[0047] 1. Multi-head self-attention module:

[0048] For the input N query volumes, they are fed into a multi-head self-attention module. In order to combine the position information and semantic information of the query volumes in this module, sine-form position information is used and appended to the semantic vectors. Additionally, the Intersection over foreground (IoF) is added as a bias to the attention weights, so that the relationships between the queries can be explicitly taken into account. The specific form is as follows:

[0049] Q = K = V = Linear(q + PosEmbedding(x, y, z, r))

[0050]

[0051] where each element in the IoF information B ∈ = 10 -7 , where exp(B ij ) = 1 means that bounding box j completely encloses bounding box i, and exp(B ij ) = ∈ means that bounding box j and bounding box i do not intersect at all. respectively represent the query (query), key (key), and value (value) in the self-attention mechanism. PosEmbedding() is a sine encoding of the query volume position quadruple, which can make different positions have different encodings, and the dimension of the output vector is N×d q . Specifically as follows:

[0052]

[0053]

[0054] α is a learnable scalar for each attention head. Finally, an enhanced semantic vector output can be obtained through this multi-head self-attention module

[0055] 2.3D Sampling Module:

[0056] a) Adaptive Sampling in the 3D Feature Space:

[0057] Using q obtained from the multi-head self-attention module, after a linear layer transformation, the semantic information of q can be used to obtain P in sets of sampling points:

[0058]

[0059] Combined with the position vector quadruple corresponding to the query volume in the input, the following P in 3D feature space sampling point coordinates P can be obtained:

[0060]

[0061] It can be found that when {Δx i , Δy i ∈[-0.5, 0.5]}, the sampling points are located within the bounding box. However, for {Δx i , Δy i} obtained by the present invention, there is no restriction on the range. That is, when necessary, according to different query contents, the sampling points can be adaptively adjusted outside the bounding box.

[0062] After obtaining the coordinates P of the sampling points, first perform bilinear interpolation in the feature space of (x, y):

[0063]

[0064] Obtain a planar feature matrix, and then perform Gaussian weight interpolation on the z-axis. The weight for the j-th feature map is:

[0065]

[0066]

[0067] where τ z is a coefficient on the z-axis. In the example of the present invention, τ z = 2. After such interpolation, the obtained output is a sampled feature matrix X, and its dimension is

[0068] b) Group sampling strategy:

[0069] For the diversity of sampling points, the present invention also adopts a group sampling strategy during sampling. Similar to the multi-head attention mechanism, the d feat in the 3D feature space is divided into g groups, and the number of channels in each group of the 3D feature space then becomes d feat / g. Then, for each group of sampled feature matrices where C = d feat / g, perform independent 3D feature space adaptive sampling. Finally, the dimension of the obtained feature matrix X becomes

[0070] 3. Adaptive mixing module:

[0071] After obtaining the sampled feature matrix, it can be adaptively decoded. Based on the group sampling strategy in 2, the present invention decodes each group of feature matrices in two steps, as Figure 2 perform adaptive mixing in sequence according to the semantic information in its channels and the spatial information in the query.

[0072] a) Adaptive Semantic Channel Mixing ACM:

[0073] For the given feature matrix where C = d feat / g, this module uses a q-based dynamic weight matrix to perform channel semantic enhancement on the sampled feature matrix x in the feature channels:

[0074]

[0075] ACM(x) = ReLU(LayerNorm(xM c ))

[0076] where the output is the feature output after mixing on the semantic channels. According to Figure 3 , the linear layer Linear c () is independent between different groups, but the dynamic weight matrix M c is shared among different sampling points in the 3D feature space. The LayerNorm() layer is used on all output channels.

[0077] b) Adaptive Spatial Information Mixing (ASM):

[0078] For the given feature matrix where C = d feat / g, this module uses a q-based dynamic weight matrix to perform spatial information enhancement on the sampled feature matrix x in the spatial information:

[0079]

[0080] ASM(x) = ReLU(LayerNorm(x T M s ))

[0081] where the output is the feature output after mixing on the spatial information. Here, P out is the number of outputs after spatial mixing, which can be adjusted manually. In the examples of this invention, after testing, when P in = 32, P out = 128, the performance can reach the best state. And because of the transpose operation on x, the dynamic weight matrix M s is shared among different channels.

[0082] c) Overall Dynamic Mixing Module:

[0083] As Figure 3, the overall dynamic mixing module first performs an adaptive semantic channel operation (ACM) on the sampled feature matrix , and then performs an adaptive spatial information mixing operation (ASM) on its output , finally obtaining the output with enhanced information

[0084] 4. FFN Module:

[0085] a) Update of the semantic vector q of the query volume:

[0086] After passing through the overall dynamic mixing module in (3), the output The outputs of g groups are concatenated and flattened to After passing through an FFN, its number of channels is converted to and then added to the original query semantic vector q, thus obtaining the updated query semantic vector q'.

[0087] b) Update of the position vector (x, y, z, r) of the query volume:

[0088] According to the updated semantic vector q', after passing through a group of FFNs, a group of updated position vectors (x', y', z', r') in the query can be obtained

[0089] {(Δx i , Δy i , Δz i )} = FFN(q')

[0090]

[0091] If the feature matrix X is not grouped, the adaptive mixing module directly performs adaptive decoding on the feature matrix X, performs 3D feature space adaptive sampling to obtain the feature matrix X', flattens the feature matrix X', passes through a group of FFNs, converts its number of channels to the same number of channels as q, obtains the updated query semantic vector q', and obtains the update of the position vector.

[0092] 3) Training Phase:

[0093] 1. Definition of the Loss Function:

[0094] The loss function of the present invention mainly consists of three parts, the focal loss, the L1 bounding box loss, and the matching loss composed of the GIoU loss. The following separately introduces these three loss functions:

[0095]

[0096]

[0097]

[0098] In the example of the present invention, L is taken focal of λ cls = 2, p t is the confidence score for classification t, and x of L L1bbox of is, and λ of L GIoU is giou = 2, and A c is the smallest bounding box that encloses multiple bounding boxes participating in the operation at the same time. The combination of these three loss functions forms the loss function L of the present invention. The supervision of the loss function L will be used in each stage of the decoder.

[0099] 2. Training strategy: In this example, focal loss, L1 bbox loss, and GIoU loss are used as loss functions, and the AdamW optimizer is used. The decay rate of the optimizer is set to 0.0001. The batch size BatchSize is set to 16, that is, 16 samples are taken from the training set for training each time, and the initial learning rate is 2.5×10 -5 . There are two training strategies in total, which are specifically as follows:

[0100] a) 1x training strategy:

[0101] The total number of training rounds is set to 12 rounds, and the short side of the training image is adjusted to 800 for input. The standard data augmentation operations for this strategy only include random horizontal flipping. Here, for a fair comparison with some popular detectors (such as FCOS and Cascade R-CNN), the present invention only allocates 100 learnable object queries, and divides the learning rate by 10 at the 8th and 11th rounds. Training on eight V100 GPUs takes about 9 hours.

[0102] b) 3x training strategy:

[0103] The total number of training rounds is set to 36 rounds. Since some popular query-based object detectors generally train more rounds on data and use some cropping or multi-scale data augmentation. Here, for a fair comparison with these detectors, the same data augmentation as theirs is used, and training is performed with 3 times the number of training rounds. The present invention will generate 300 object queries under this training strategy, and divide the learning rate by 10 at the 24th and 33rd rounds. Training on eight V100 GPUs takes about 29 hours.

[0104] 4) Testing phase

[0105] The processing of the test set input data is the same as that of the training data. The short side of the input image is scaled to 800, and the ResNet50 network is used for feature extraction. The test metrics used are AP, AP 50 , AP 75 , Ap s , AP m , AP l , which are a series of metrics representing the accuracy of object detection. AP 50 refers to the average precision when the intersection over union (IoU) between the ground truth bounding box and the predicted bounding box of an object is greater than 0.5, indicating that the prediction is accurate; AP 75 refers to the average precision when the intersection over union (IoU) between the ground truth bounding box and the predicted bounding box of an object is greater than 0.75, indicating that the prediction is accurate; AP refers to the average precision measured at 10 thresholds where the IoU threshold between the ground truth bounding box and the predicted bounding box of an object ranges from 0.5 to 0.95 with an interval of 0.05 for each threshold, and it is also the most important metric. Ap s , AP m , AP l respectively refer to the average precision when the objects to be detected are small objects, medium objects, and large objects. In the COCO dataset, when using the 1x training strategy, compared with other object detection frameworks that also use ResNet50 as the backbone network, the best results can be obtained in all metrics, as shown in Figure 5 ; when using the 3x training strategy, even better results are achieved in all metrics except AP l . When switching the backbone network of the present invention to ResNeXt-101-DCN and Swin-S, even better results are achieved in all of the above metrics compared to the current object detection frameworks using the same backbone network. Examples of this dataset are shown in Figure 6 .

[0106] The present invention focuses on a framework capable of adapting based on the query volume, without modules such as anchor boxes, dense matching, and non-maximum suppression. For the feature extraction module, the present invention uses an encoding form of 3D feature sampling, enabling the use of the spatial position information of the query volume to sample feature maps of different resolutions simultaneously. In view of the current situation that other existing query-based object detection frameworks cannot fully utilize the information in the query volume, the present invention proposes an adaptive semantic feature mixing module and an adaptive spatial information mixing module, which can be used in combination. To address the problem of large network computational complexity, the network designed by the present invention is relatively small in scale, only including a decoder for initial query generation and a decoder module capable of adapting to the query volume information, achieving excellent results while reducing the network scale and computational complexity. AdaMixer first applies the adaptive mixing module to the object detection module, designing a simple and neat framework, removing the manually designed modules and the dense matching paradigm; AdaMixer proposes 3D feature sampling, adaptive semantic feature mixing, and adaptive spatial information mixing to solve the problem that existing query-based frameworks cannot fully utilize query information; AdaMixer achieves state-of-the-art results on all metrics of the MSCOCO minival dataset.

Claims

1. An object detection method based on an adaptive decoder, characterized in that an object detection model AdaMixer is constructed, multi-scale features of an image are sampled according to a 3D sampling feature space, and based on a query mechanism, the positions of sampling points and the decoding of features are adaptively adjusted according to the spatial position information and semantic content information of the query through an encoder and a decoder, and then the enhancement of the query is completed in cooperation with an FFN to realize the detection of the query position. The implementation of the object detection network includes a 3D feature generation stage, a network configuration stage, a training stage, and a testing stage: 1) 3D feature generation stage: Use a backbone network to extract features from training example images. For each input image, based on the feature maps with different lengths, widths, and numbers of channels output by different stages of the backbone network, a 3D feature space is obtained for subsequent sampling processing therein; 2) Network configuration stage, based on an initial query configuration and a decoder, an object detection model AdaMixer is established, including the following configurations: 2.1) Initial query configuration: Initialize the encoding for the input feature map to generate N initial query quantities. The query quantities include the initial semantic vector q 0 , and the initial position vectors (x, y, z, r) corresponding to the query sampling points. (x, y, z) are the coordinates of the sampling points in the 3D feature space, and r is the base-2 logarithm of the aspect ratio of the feature map. The initial semantic vector q 0 is randomly sampled from the standard normal distribution N(0, 1), and the initial position vectors (x, y, z, r) are set to cover the entire feature map; 2.2) Decoder: The input of the decoder is the query quantity generated by encoding in 2.1), and the output is the query quantity of the same format after being optimized by the decoder. The decoder contains the following modules for adaptively using the semantic information and position information in the query: 2.2.1) Multi-head self-attention module: Input N query quantities into a multi-head self-attention module, attach the position information in the form of sine to the semantic vector, and add the intersection over foreground ratio IoF as a bias to the weight of the attention. After passing through this multi-head self-attention module, an enhanced semantic vector output q is obtained; 2.2.2) 3D Sampling Module: Using q obtained in 2.2.1), after a linear layer transformation, the semantic information of q is utilized to obtain P in a set of displacement of sampling points, combined with the corresponding position vector quadruple in the query volume, to obtain P in coordinates P of 3D feature space sampling points. According to the sampling point coordinates P, two-point interpolation is first performed in the (x, y) feature space to obtain a feature matrix without considering the z-axis weight, and then Gaussian weight interpolation is performed on the z-axis to further obtain the complete sampled feature matrix X; 2.2.3) Adaptive mixing module: Decode the feature matrix X adaptively, and perform adaptive mixing on the sampled feature matrix X in two steps, namely adaptive semantic channel mixing and adaptive spatial information mixing. In adaptive semantic channel mixing, use a dynamic weight matrix based on q to perform channel semantic enhancement on the feature matrix X in the feature channels; In adaptive spatial information mixing, use a dynamic weight matrix based on q to perform spatial information enhancement on the feature matrix X in the spatial information; finally, an information-enhanced feature matrix X' is obtained; 2.2.4) FFN module: For the updated feature matrix X' output in 2.2.3), combined with its position vector, update the semantic vector and position vector in the overall query, that is, flatten the feature matrix X' obtained in 2.2.3), pass through a group of FFNs, convert its number of channels to the same number of channels as q, and obtain the updated semantic vector q' in the query; according to the updated semantic vector q', pass through another group of FFNs to obtain a group of updated position vectors (x', y', z', r') in the query; 2.3) After obtaining the final query semantic vector q' and position vectors (x', y', z', r'), send q' into an FFN classification network to obtain a classification result, and translate (x', y', z', r') into the coordinates of a bounding box to obtain the result of the bounding box; 3) Training stage: Train the configured network model using training data. Use the combination of focal loss, L1 loss, and GIoU loss as the loss function, and use the AdamW optimizer. Update the network parameters through the backpropagation algorithm. Continuously repeat steps 1) and 2) until the number of iterations is reached; 4) Training stage: Input the image features of the data to be tested into the trained AdaMixer model. According to the method in 2.3), obtain the final target classification result and the position of the target bounding box, verify the effect of the trained AdaMixer model, and use the AdaMixer model that achieves the required object detection effect as the final obtained object detection model for object detection.

2. An object detection method based on an adaptive decoder according to claim 1, characterized in that The backbone network in step 1) is ResNet. For the feature maps of the second to fifth stages of the backbone network ResNet, after linear transformation, the number of channels is unified to d feat , stacked in a three-dimensional space to obtain a 3D feature space.

3. An object detection method based on an adaptive decoder according to claim 1, characterized in that The decoder first feeds the query volume into a multi-head self-attention module to obtain an enhanced semantic query volume After transforming q through a linear layer, P is obtained in a set of sampling points Sample using the set of sampling points in the 3D feature space. First, divide the feature space into g groups, and the number of channels in each group of the feature space after grouping is d feat / g, d feat is the number of channels of the feature space, and sample P in points in each group to obtain a feature matrix The feature matrix of each group is where C = d feat / g, and perform adaptive mixing. First, perform adaptive semantic channel mixing to obtain an enhanced output of semantic features Then perform adaptive spatial information mixing to obtain an enhanced output of spatial information Finally, send this output into the FFN network to update the semantic vector and position vector of the query, and obtain the query volume q' and position vector (x', y', z', r') that have the same input format and are adaptively optimized 4. An object detection method based on an adaptive decoder according to claim 1, characterized by stacking multiple layers of decoders to deepen the network. Each layer of the decoder includes a multi-head self-attention module, a 3D sampling module, an adaptive mixing module, and an FFN module.

5. An object detection device based on an adaptive decoder, characterized in that it has a computer-readable storage medium, and a computer program is configured in the computer-readable storage medium. When the computer program is executed by a processor, it implements the object detection method according to any one of claims 1-4.