Object detection method based on dynamic weights and hierarchical query selection strategies

JP7927367B1Active Publication Date: 2026-10-01HANGZHOU DIANZI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2026098440
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2026-01-07
Filing Date
2026-06-12
Publication Date
2026-10-01
Estimated Expiration
2046-06-12

Smart Images

  • Figure 0007927367000001_ABST
    Figure 0007927367000001_ABST
Patent Text Reader

Abstract

The method includes the steps of: extracting a multiscale feature map and obtaining its shape information and mask information; performing convolution, flattening, and dimensional rearrangement operations on the multiscale feature map, inputting it into a Transformer encoder in combination with a position code to obtain an enhanced image feature sequence and its mask; optimizing the image feature sequence, performing initial predictions using a classification head and a detection head to obtain category scores and position information, and generating a global mask; calculating the dynamic weights of the feature maps in each layer; assigning the number of queries for the feature maps in each layer based on the dynamic weights; selecting local queries and projection queries from the upper layers layer by layer from top to bottom; generating reference points and position embedding vectors using the selected queries; and outputting the final object detection result. [Effect] Improves the accuracy of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the fields of computer vision and object detection technology, and more specifically to an object detection method based on dynamic weights and a hierarchical query selection strategy. [Background technology]

[0002] Object detection is an important research direction in the field of artificial intelligence, aiming to identify and classify objects in images. In recent years, the development of object detection technology has gone through three stages: conventional, two-stage deep learning, and one-stage deep learning. In the conventional detector stage, object detection algorithms mainly rely on manual feature construction. However, due to problems such as limited feature representation capabilities, complex detection processes, and low computational efficiency, it is difficult to handle complex scenarios.

[0003] With the advent of deep convolutional neural networks, two-stage methods, such as RCNN, SPPNET, and Fast RCNN, became mainstream. The basic idea behind these methods is to first generate candidate regions, and then perform classification and bounding box regression for each region. However, this "propose candidate regions first, classify later" approach leads to computational redundancy and process complexity, resulting in slower detection speeds and increased process complexity. To simplify the process, one-stage detectors treat detection as a unified regression problem, directly predicting bounding boxes and class probabilities on the image. Methods such as YOLO, SSD, and RetinaNet significantly improve detection speed. However, both conventional two-stage methods and earlier one-stage models rely at their core on predefined anchor frames and complex, manually designed components. To solve this problem, the DETR model was proposed, which was the first to introduce the concepts of Transformer architecture and ensemble prediction into the field of object detection, creating a true end-to-end object detection framework that does not require anchor boxes or post-processing NMS. In early DETR-like models, the decoder query input to the Transformer decoder was a static embedding that did not capture any encoder features. This method resulted in slow model convergence and low accuracy because the information generated by the encoder after image processing was discarded. Therefore, Deformable DETR proposes a two-stage method to enhance the decoder query by selecting the top K encoder features from the last encoder layer as preferred options based on class scores.

[0004] Based on this idea, Chinese Patent Application Publication No. 114663915 discloses an image-based human-object interaction localization method and system based on a Transformer model. The method includes the steps of: obtaining a target image and a descriptive phrase, where the descriptive phrase is used to describe the human-object interaction relationship in the target image, and the target image includes a human-object interaction scene; performing feature extraction on the target image and the descriptive phrase, respectively, to obtain image features and linguistic features; and inputting the image features and linguistic features into a pre-trained human-object interaction localization model, thereby outputting all human and object localization boxes in the target image that match the descriptive phrase, thereby realizing human-object interaction localization.

[0005] However, selecting features based solely on the order of class scores can lead to situations where certain feature maps are selected many times, while others are selected infrequently or not at all. As a result, some feature information is ignored, and the benefits of multiscale features are not fully utilized. Therefore, there is an urgent need for a method that can more effectively utilize multiscale features so that models can better understand image information. [Overview of the project] [Problems that the invention aims to solve]

[0006] The objective of this invention is to overcome the shortcomings of the prior art and propose an object detection method based on dynamic weighting and a hierarchical query selection strategy. The weights of each feature layer are calculated by analyzing the results and confidence scores after an initial screening query, the number of queries for each layer is then dynamically assigned, and finally, the high-quality queries selected in the feature maps of each layer are projected downwards, improving the image information of the high-quality queries in the other feature maps, thereby improving the accuracy of object detection. [Means for solving the problem]

[0007] To achieve the above objectives, the technical solutions specifically employed by the present invention are as follows.

[0008] An object detection method based on dynamic weights and a hierarchical query selection strategy, Step 1 involves preprocessing the input image, extracting multiscale feature maps via a backbone network, and obtaining the shape and mask information of the feature maps in each layer. Step 2 involves performing convolution, flattening, and dimensional rearrangement operations on the multiscale feature map, inputting it into a Transformer encoder in combination with a position code, and obtaining an enhanced image feature sequence and its mask. Step 3 involves optimizing the aforementioned image feature sequence, performing initial predictions using a classification head and a detection head to obtain class scores and location information, and generating a global mask by non-maximum suppression processing. Step 4 involves calculating the dynamic weights of the feature maps in each layer based on the shape of the feature map, the image feature sequence, the class score, and the global mask. Step 5 involves assigning the number of queries for each layer's feature map based on the aforementioned dynamic weights, wherein each layer's query includes two parts: a local query and a projected query for the upper layer. Step 6 involves selecting local queries and projected queries for higher layers from top to bottom, layer by layer, wherein the projected queries for higher layers are obtained by projecting the queries selected in the previous layer onto the feature map of the current layer. Step 7 involves using the selected query as a content query and combining it with its corresponding location information to generate a reference point and a location embedding vector. The process includes step 8, which inputs content queries, reference points, and position embedding vectors into a Transformer decoder, and outputs the final object detection result using a classification head and a detection head.

[0009] Preferably, in step 1, the backbone network is ResNet-50, and the output multiscale feature map includes feature maps of four different scales.

[0010] Preferably, in step 2, the position code is calculated using a sine function and a cosine function. The formula is as follows:

number

[0011] Preferably, the Transformer encoder includes six encoder layers, each encoder layer including a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation.

[0012] Preferably, in step 3, the method for optimizing the image feature sequence is as follows: Each feature map is traversed sequentially to create a 2D grid coordinate system covering the pixel location of each feature map in the current layer. These coordinate axes are integrated into a tensor of shape H, W, 2, representing the coordinates of each pixel, where H and W represent the height and width of the feature map. Based on the resulting mask, the effective width and height of each image in the current layer are calculated and used as scaling factors. The grid coordinates are normalized to the effective region, and the width and height of the ground truth box where the grid coordinate values ​​normalized to the effective region are located are scaled based on a hierarchical index. The grid coordinates are combined with the width and height of the bounding box to which the grid coordinate values ​​are located to obtain a tensor of shape N, H*W, 4. For each of the four elements, one quadruple (cx,cy,w,h) is constructed, and each quadruple represents one candidate proposal init_proposal. Valid prediction proposals are selected based on whether each proposal value falls within the range of (0.01,0.99). Finally, a logarithmic transformation is performed on the selected proposals, and this transformation maps the coordinates from the [0,1] space to the real space to obtain the initial proposal result output_proposal, whose formula is as follows:

number

[0013] Preferably, in step 4, the method for calculating the dynamic weights is as follows: First, the effective feature points and their class scores are extracted from the feature map of each layer. The L2 norm of each effective feature point is calculated. The importance score is obtained by multiplying the class score by the L2 norm. The initial weights are obtained by averaging the importance scores of all effective feature points in the feature map of each layer. The initial weights of all layers are normalized, and finally, the final weights of the feature map of each layer are obtained.

[0014] Preferably, in step 5, the number of queries for the feature map of each layer is obtained by multiplying the weight of the layer by the total number of queries, the feature map of the top layer only includes local queries, and the remaining layers include local queries and projection queries from upper layers.

[0015] Preferably, in step 6, the generation of the projection queries of the upper layers comprises: obtaining the center coordinates of the queries selected in the previous layer; calculating a scaling coefficient based on the aspect ratio of the feature maps of adjacent layers, and projecting the center coordinates to the current layer; finding 2×2 neighboring feature points with the closest projection coordinates in the feature map of the current layer; selecting valid feature points as projection queries of upper layers based on masks and class scores.

[0016] Preferably, in step 7, the reference points are projected by sinusoidal position encoding and multi-layer perceptron to obtain position embedding vectors.

[0017] Preferably, in step 8, the Transformer decoder comprises 6 decoder layers, each layer includes a self-attention mechanism, a multi-scale deformable attention mechanism, a feed-forward network, a dropout layer, and a layer normalization operation, and the final prediction result is matched by the Hungarian algorithm and ground truth. Effects of the Invention

[0018] The present invention has the following features and beneficial effects.

[0019] This method improves the query selection strategy and proposes a dynamic weighting module that assigns weights of different magnitudes based on the importance of feature maps in each layer, obtains the number of selected candidate points for each feature map based on these weights, selects feature points in each layer's feature map based on the class score, and the selected feature points consist of two parts: one is the feature point of the current layer, and the other is the feature point projected from the previous layer to the current layer. By projecting the feature points, the model can be input with relevant information from other feature maps, thereby enabling the model to better capture the features of objects and improve its detection and recognition capabilities. The present invention introduces dynamic weighting, selects candidate points layer by layer, and projects them to adjacent layers, enabling the model to better understand object information in an image. [Brief explanation of the drawing]

[0020] [Figure 1] Figure 1 is a framework diagram of the object detection method based on the dynamic weight and hierarchical query selection strategy of the present invention. [Figure 2] Figure 2 is a flowchart showing the steps of the object detection method based on the dynamic weight and hierarchical query selection strategy of the present invention. [Figure 3] Figure 3 is a structural diagram of the query cross-layer projection in the object detection method based on the dynamic weight and hierarchical query selection strategy of the present invention. [Figure 4] Figure 4 is a framework diagram of the dynamic weight calculation module in the object detection method based on dynamic weights and a hierarchical query selection strategy of the present invention. [Modes for carrying out the invention]

[0021] The present invention will be described in detail below, along with specific examples. The following examples will be helpful to those skilled in the art in gaining a deeper understanding of the present invention, but will not limit the present invention in any way. The examples and features of the present invention may be combined in any way that does not contradict each other.

[0022] This invention proposes an object detection method based on dynamic weights and a hierarchical query selection strategy. The model constructed by this method is the Level-DETR image object detection model. Level-DETR is an improved version of DDQ-DETR, which mainly includes an image feature extraction module, a Transformer-based encoder, an NMS-based query filtering model, dense query auxiliary queries, an NMS-based Transformer decoder, multiple prediction heads, and a denoising training strategy. After the original image features are extracted by the backbone network, they are input to the Transformer encoder to enhance the features, then the query is obtained using the NMS-based query filtering model to be input to the NMS-based Transformer decoder, and finally, the classification and detection results of the object are obtained based on the classification head and detection head.

[0023] The Level-DETR model proposed in this method improves upon the NMS-based query filtering model, enabling query selection based on feature maps. By projecting the selected queries to the next layer, supplementary information is added to the selected queries, achieving a query enhancement effect.

[0024] This embodiment proposes an object detection method based on dynamic weights and a hierarchical query selection strategy, and the specific steps are as follows, as shown in Figures 1 and 2.

[0025] Step 1: Preprocess the input image, extract multiscale feature maps via the backbone network, and obtain the shape and mask information of the feature maps in each layer.

[0026] In this embodiment, the input images are first preprocessed to scale different images within the same batch to the same size and fill in smaller images within the same batch. Then, features are extracted from the preprocessed images using a ResNet-50 backbone network to obtain a multiscale feature map, and the shape of the feature map of each layer and the mask information of the multiscale feature map are calculated.

[0027] The specific implementation process is as follows: First, the images are scaled to a shape with a short side of 800 and a long side of less than 1333. Different images within the same batch are scaled to the same shape, smaller images within the same batch are filled in, and the images are horizontally flipped with a probability of 0.5. Next, the preprocessed images are input into a ResNet-50 backbone network, which employs a residual connection mechanism to extract image features stepwise through multiple layers of convolution and residual modules, and finally a multiscale image feature tensor Image Feat Obtain information about masks. Image Feat ∈[res2,res3,res4,res5] Here, res2 to res5 represent feature maps of the original image at four different scales.

[0028] Specifically, multiscale features are extracted via a backbone network, and the shape of each feature map is (N, H × W, C), where N represents the batch size, C represents the number of channels, and H and W represent the width and height of the feature map, respectively.

[0029] Furthermore, two main processes are performed on the extracted multiscale features. First, we sequentially traverse each feature map to obtain the shape of each feature map, and store all shapes in spatial shape variables. Next, each feature map is joined along the second dimension to obtain an image feature tensor, the dimension of which is (N, Σ i lvl H i, ×W i ,C)

[0030] Step 1 primarily yields three variables: the image feature tensor (used for calculations in the Transformer), the spatial shape (used to divide the feature map where the features are located), and the mask (used to distinguish between valid and invalid regions).

[0031] Furthermore, the method for obtaining the shape and mask information of the feature maps of each layer is as follows: Based on the width and height of the highest resolution feature map obtained from the backbone network, the feature maps are sequentially downsampled, reducing their width and height to half of their original values, until the width and height of all feature maps are finally obtained.

[0032] In this embodiment, through Step 1, Image feature tensor used in calculations in Transformer, The spatial shape used to divide the feature map in which the features are located. Please understand that three variables are obtained: the mask information used to distinguish between the active and inactive regions.

[0033] Step 2: The multiscale feature map is subjected to convolution, flattening, and dimensional rearrangement operations, and combined with the position code, it is input to a Transformer encoder to obtain an enhanced image feature sequence and its mask.

[0034] In this embodiment, the image feature tensor of the multiscale feature map obtained in step 1 is subjected to convolution, flattening, and dimensional rearrangement operations. Simultaneously, position coding is performed on the input multiscale feature map. The flattened image feature sequence and position code are input to a Transformer encoder to obtain an encoded and enhanced image feature sequence. Based on the mask information obtained in step 1, a mask of the image feature sequence is obtained.

[0035] The specific implementation process is as follows: Step 2-1: The image feature tensor of the multiscale feature map is mapped by a convolutional layer with a kernel size of 1x1 and a stride of 1, reducing the channel dimension to 256 to match the input dimension of the Transformer. Next, the width and height of the image are flattened to one dimension to obtain the sequence length. Finally, the dimensions are rearranged based on the sequence length and feature dimension to obtain an image feature sequence containing all feature information. Step 2-2: Position coding is performed on each scale feature in the image feature tensor of the multiscale feature map. Position information is coded using sine and cosine waves of different frequencies. The embedding vector for each position consists of a series of sine and cosine values, and the frequencies of these values ​​increase or decrease according to a certain rule. Finally, the position code Pos is output. The formula for calculating the position code is as follows:

number

number

[0036] Step 3: The image feature sequence is optimized, initial predictions are made using the classification head and detection head to obtain class scores and position information, and a global mask is generated by non-maximum value suppression processing.

[0037] In this embodiment, the image feature sequence obtained in step 2 is predicted, corresponding location information and classification score are obtained, non-maximum NMS processing is performed on all feature points, mask information is added to feature points whose similarity exceeds a threshold, and finally, global mask information for the entire image feature sequence is obtained.

[0038] The specific implementation process is as follows: Step 3-1: This model method is a two-stage detection method. All image feature sequences obtained in Step 2 are processed by a classification head and a detection head to generate an initial proposal for object detection for each image feature sequence. Specifically, each feature map is traversed sequentially to create a two-dimensional grid coordinate system covering each pixel position of the feature map in the current layer. These coordinate axes are integrated into a tensor of shape (H, W, 2) representing the coordinates of each pixel, where H and W represent the height and width of the feature map. Based on the mask obtained in Step 2-3, the effective width and height of each image in the current layer are calculated and used as scaling factors to normalize the grid coordinates to the effective region. Note that because the sizes of images within the same batch are non-uniform, when filling smaller images, the region to be filled is the invalid region, and the effective region is the region that does not need to be filled. The width and height of the ground truth box where the normalized grid coordinate values ​​are located within the effective region are scaled based on a hierarchical index, reflecting a hierarchical perceptual characteristic where higher-level features correspond to larger objects and lower-level features correspond to smaller objects, where the hierarchical index is the hierarchical number of the multiscale feature. The scaling method for the hierarchical index is to raise the index number to the power of 2 and multiply the result by 0.05 to obtain the scaled width and height. In this embodiment, the initial values ​​of the width and height of the ground truth box are 1. Next, the scaled grid coordinates are combined with the width and height of the ground truth box where the grid coordinate values ​​are located to obtain a tensor of shape (N, H*W, 4). Please understand that the scaled width and height shape is (N, H*W, 2), where N represents the batch size, H and W represent the width and height of the current feature layer, the third dimension represents the width and height of the grid, the coordinate shape is (N, H*W, 2), the third dimension represents the center coordinates, and these two are combined according to the third dimension to obtain the initial proposal. One proposal (cx, cy, w, h) is represented for each of the four elements, where the four elements respectively represent the x-axis of the center coordinate, the y-axis, the width, and the height. Each quadruple represents one candidate proposal init_proposal, and effective prediction proposals are screened based on whether each proposal value is within the range of (0.01, 0.99). Finally, logarithmic transformation is performed on the screened proposals, through which coordinates are mapped from the [0,1] space to the real number space, so as to obtain the initial proposal result output_proposal. This transformation improves the numerical stability of the proposals and facilitates optimization.

Math

[0039] Step 4: Based on the shape of the feature map, the image feature sequence, the class score, and the global mask, the dynamic weights of the feature maps in each layer are calculated.

[0040] In this embodiment, based on the shape of the feature map obtained in step 1, the image feature sequence obtained in step 2, and the class score and mask obtained in step 3, the number of feature points held in the feature map of each layer and the mask information of the feature map of each layer are first obtained according to the mask. Next, the corresponding feature vectors are obtained from the image feature sequence via the mask information, and the L2 norm of each feature vector is calculated. Subsequently, importance scores are calculated based on the class score and L2 norm of the feature points. Finally, the weights of all feature maps are normalized by using the average importance score of the valid feature points in the feature map of each layer as the weight of that layer.

[0041] As shown in Figure 4, the specific implementation process is as follows: Step 4-1: First, based on the shape of the feature map, obtain the boundary range of the feature map for each layer. Then, from the image feature sequence, class score, and mask, obtain the feature sequence, class score, and mask corresponding to the feature map of that layer, based on the boundary range. Step 4-2: Calculate the L2 norm of the feature sequence of the feature map of that layer. Based on the effective feature sequence of the resulting feature map, calculate the L2 norm of each effective feature point, square each element, sum all the squared values, and finally take the square root of the sum. Step 4-3: The L2 norm of the obtained feature point is multiplied by the corresponding class score to obtain the importance score of that feature point. Next, the average importance of the feature map in this layer is calculated and used as the initial weight of that layer. Finally, the initial weights of all feature maps are normalized to obtain the final weight of the feature map in each layer. The formula for calculating the weight of the feature map is as follows:

number

number

[0042] Step 5, assigning the number of queries for each layer's feature map based on the dynamic weights, wherein each layer's query includes two parts: a local query and a projected query for the upper layer.

[0043] Specifically, as shown in Figure 3, in this embodiment, based on the shape of the feature map obtained in step 1 and the image feature sequence obtained in step 2, the image feature sequence of the feature map of each layer is obtained according to the shape of the feature map, the start index and end index of each feature map are obtained based on the boundary range of the feature map of each layer obtained in step 4-1, and the number of elements K to be assigned to each layer is determined based on the weights of the feature map and the total number of elements K in the model configuration obtained in step 4-3. level The query for each layer's feature map is calculated as follows: local and Query upper It is divided into two parts, and since the feature map of the top layer does not contain the features of the higher layers, the feature map of this layer is a query localIt is composed of the following, and the feature maps of the lower three layers are divided into two parts, and these two parts equally divide the number of feature maps in each layer. K level =weight × K Here, `weight` represents the weight of the feature map, and `K` represents the number of queries input to the decoder.

[0044] Step 6 involves selecting local queries and projected queries for higher layers from top to bottom, layer by layer, wherein the projected queries for higher layers are obtained by projecting the queries selected in the previous layer onto the feature map of the current layer.

[0045] Specifically, in this embodiment, the assignment configuration of the feature maps for each layer is obtained through step 5. Selecting from top to bottom, the number of selected elements from each part of the feature map of each layer is taken from the configuration, and for the four-layer feature map, the corresponding number of features are Queryed according to the class score generated by the classification head. local Selected as such, the feature maps of the lower three layers are Query upper Select and perform a projection calculation on the feature map of the previous layer. First, calculate the center coordinates of the Query selected in the previous layer. Next, calculate a scaling factor according to the aspect ratio of adjacent feature maps. Next, project the center coordinates of the upper layer onto the current layer to find the nearest feature point in the current feature map. Next, generate a 2x2 neighborhood. Once the projection operation is complete, according to the mask information and the classification score generated by the classification head, apply the corresponding number of features to the valid region. upper Select it as such.

[0046] The specific implementation process is as follows: Step 6-1: Based on the assignment configuration of the feature maps for each layer obtained in Step 5, select from top to bottom for all feature maps and take the number of selected parts from the two parts of the feature map from the assignment configuration. Based on the class scores of the feature vectors obtained in Step 6-2 and Step 3, the masks of the feature maps obtained in Step 4, the image feature sequences of the feature maps of each layer obtained in Step 5, and the number selected from each part of the feature maps of each layer obtained in Step 6-1, first sort the feature maps of this layer in descending order according to the masks and class scores, select the index of the first corresponding number, and then extract the corresponding feature vectors from the image feature sequences and query local It forms. Step 6-3: Perform a Query projection operation on the feature maps of the lower three layers. Based on the positional information of the feature points obtained in step (3) and the Query selected in the previous layer, first obtain the center coordinates of the prediction box based on the positional information. Next, multiply these center coordinates by the width and height of the current feature map to obtain the true center coordinates. Then, calculate the aspect ratio of the two layer feature maps based on the shape of the feature map obtained in step (2) to obtain scaling factors for width and height. Next, multiply the x-axis coordinate of the center coordinates by the scaling factor for width and the y-axis coordinates by the scaling factor for height to obtain new center coordinates. Step 6-4: First, generate a 2x2 neighborhood offset offset ∈ [[0,0],[0,1],[1,0],[1,1]], add the offset to the center coordinates to obtain the projected center coordinates, and based on the starting index of each feature map obtained in step (5), add the result of multiplying the y-axis coordinate of the projected center coordinates by the width of the current feature map, add the x-axis coordinate of the projected center coordinates to obtain the final projected index. Step 6-5: After obtaining all index sets, deduplication is performed on duplicate indices, invalid indices are removed using a mask, and finally, class scores are extracted based on the index, sorted in descending order, and the number to select from the configuration is determined. If the remaining number is less than the number to be selected, all remaining indices are selected, and the corresponding feature vectors are extracted from the image feature sequence and query... upper This forms the following. Finally, these two parts are combined as the query for the current layer. Step 6-6: If the total number of selections at the end is less than a predetermined number K, select and complete the remaining indices according to the class score.

[0047] Step 7: The selected query is used as a content query and combined with its corresponding location information to generate reference points and location embedding vectors.

[0048] Specifically, the Query selected in step 6 is used as the content query, reference points are generated based on the location information obtained in step 3, and these reference points are embedded by generating a periodic sinusoidal position.

[0049] The specific implementation process is as follows: The query obtained in step 6 is used as the content query input to the Transformer decoder. According to the index of the content query, the corresponding ground truth box representation is extracted from the location information to generate a reference point. This reference point is used as the prior information for the geometric coordinates of the Transformer decoder. The reference point is encoded with a sinusoidal position and projected by the MLP to obtain the corresponding position embedding vector.

[0050] Step 8: The content query, reference point, and position embedding vector are input to the Transformer decoder, and the classification head and detection head output the final object detection result.

[0051] The content query obtained in step 6 and the reference point and position embedding vector obtained in step 7 are input to the Transformer decoder for calculation. After prediction by the classification head and detection head, the coordinate boxes and class information of all objects are obtained, and then the final prediction results are obtained by matching using the Hungarian algorithm.

[0052] The specific implementation process is as follows: The content queries obtained in step 6, the reference points and location embedding vectors obtained in step 7 are input to a Transformer decoder consisting of six decoder layers for computation. Inside each decoder layer are a self-attention mechanism, a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation. Here, the self-attention mechanism models the global relationships between all queries and enables information interaction between different objects. The multiscale deformable attention mechanism uses the sampling positions of the reference points at different feature scales to adaptively aggregate attention across multilayer feature maps, allowing the model to pay attention to both large and small objects simultaneously. By iterating through the six decoder layers layer by layer, each query vector is gradually brought into focus on a specific object region, ultimately yielding a query feature representation. These feature representations are then processed by a classification head and a bounding box detection head to output class predictions and ground truth box coordinates corresponding to each query. Finally, the model uses the Hungarian method to establish a one-to-one correspondence between the predicted results and ground truth, obtaining the final detection result and achieving end-to-end object detection.

[0053] The object detection method based on dynamic weights and a hierarchical query selection strategy provided in this embodiment was implemented in an experimental environment using Python 3.8 (Ubuntu 18.04) and CUDA 11.3. The hardware environment consisted of an RTX 4080 Super and an 11th Gen Intel(R) Core(TM) i7-11700. The AdamW optimizer was used with a BASE_LR of 0.0002, a batch size of 2, and a training batch size of 12. In this experiment, ResNet-50 was used as the backbone network, with 6 decoder and encoder layers, and the number of queries was set to 300.

[0054] In this example, 7393 images were randomly selected from the MS COCO dataset for training, and 313 images were used for validation. The extracted data was guaranteed to contain 80 classes. The MS COCO dataset is a large dataset that can be used for image detection and is classified into 80 object classes (pedestrians, cars, elephants, etc.). (Table 1) Comparison of the overall AP representation of the present invention and conventional models in object detection tasks [Table 1]

[0055] As can be seen from Table 1, compared to the original DDQ-DETR model, the present invention uses a method of assigning queries hierarchically, resulting in improvements to all metrics except AP@75 and APm when the projection area is 2x2, and improvements to all metrics except APm when the projection area is 4x4.

[0056] Experimental results show that the DETR object detection model, which assigns queries hierarchically, improves the overall detection accuracy of the model.

[0057] The basic principles, main features, and advantages of the present invention have been shown and described above. Those skilled in the art will understand that the present invention is not limited by the above embodiments, and that what is described in the above embodiments and specification is merely a preferred example of the present invention, not limiting the invention, and that various modifications and improvements can be made to the present invention without departing from the spirit and scope of the invention, and that all such modifications and improvements fall within the claimed scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and equivalents.

[0058] (Note) (Note 1) An object detection method based on dynamic weights and a hierarchical query selection strategy, Step 1 involves preprocessing the input image, extracting multiscale feature maps via a backbone network, and obtaining image feature tensors, shape, and mask information for each layer's feature map. Step 2 involves performing convolution, flattening, and dimensional rearrangement operations on the image feature tensor of the multiscale feature map, inputting it into a Transformer encoder in combination with a position code, and obtaining an enhanced image feature sequence and its mask. Step 3 involves optimizing the aforementioned image feature sequence, performing initial predictions using a classification head and a detection head to obtain category scores and location information, and generating a global mask by non-maximum value suppression processing. Step 4 involves calculating the dynamic weights of the feature maps in each layer based on the shape of the feature map, the image feature sequence, the category score, and the global mask. Step 5 involves assigning the number of queries for each layer's feature map based on the aforementioned dynamic weights, wherein each layer's query includes two parts: a local query and a projected query for the upper layer. Step 6 involves selecting local queries and projected queries for higher layers from top to bottom, layer by layer, wherein the projected queries for higher layers are obtained by projecting the queries selected in the previous layer onto the feature map of the current layer. Step 7 involves using the selected query as a content query and combining it with its corresponding location information to generate a reference point and a location embedding vector. An object detection method comprising step 8, inputting a content query, a reference point, and a position embedding vector into a Transformer decoder, and outputting a final object detection result using a classification head and a detection head.

[0059] (Note 2) The method according to Appendix 1, characterized in that, in step 1, the backbone network is ResNet-50, and the output multiscale feature map includes feature maps of four different scales.

[0060] (Note 3) The method according to Appendix 1, characterized in that, in step 2, the position code is calculated using a sine function and a cosine function.

[0061] (Note 4) The method according to Appendix 1, wherein the Transformer encoder comprises six encoder layers, each encoder layer comprising a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation.

[0062] (Note 5) In step 3, the method for optimizing the image feature sequence is as follows: Each feature map is traversed sequentially to create a 2D grid coordinate system covering the pixel position of each feature map in the current layer. These coordinate axes are integrated into a tensor of shape H,W,2 representing the coordinates of each pixel, where H and W represent the height and width of the feature map. Based on the resulting mask, the effective width and height of each image in the current layer are calculated and used as scaling factors. The grid coordinates are normalized to the effective region, and the width and height of the ground truth box where the grid coordinate values ​​normalized to the effective region are located are calculated based on a hierarchical index. Next, the grid coordinates are scaled, and then the scaled grid coordinates are combined with the width and height of the ground truth box where the grid coordinate values ​​are located to obtain a tensor of shape N,H*W,4. For each of the four elements, a quadruple is constructed, represented by the center coordinate x-axis, y-axis, width, and height. Each quadruple represents a candidate proposal, and valid prediction proposals are selected based on whether each proposal value falls within a predetermined threshold range. Finally, a logarithmic transformation is performed on the selected proposals, and this transformation maps the coordinates from the [0,1] space to the real space to obtain the initial proposal results. Finally, the method according to Appendix 1, characterized in that invalid proposed values ​​among the initial proposed results are set to null values ​​according to the obtained mask, invalid regions of the image feature sequence are filled with 0, and a new feature sequence is obtained.

[0063] (Note 6) In step 4, the method for calculating the dynamic weights is as follows: The method described in Appendix 1, characterized by first extracting the effective feature points and their category scores from the feature map of each layer, calculating the L2 norm for each effective feature point, multiplying the category score by the L2 norm to obtain an importance score, averaging the importance scores of all effective feature points in the feature map of each layer to obtain initial weights, normalizing the initial weights of all layers, and finally obtaining the final weights of the feature map of each layer.

[0064] (Note 7) The method according to Appendix 1, wherein in step 5, the number of queries in the feature map of each layer is obtained by multiplying the weight of that layer by the total number of queries, the feature map of the top layer contains only local queries, and the remaining layers contain local queries and projected queries from the upper layers.

[0065] (Note 8) In step 6, the generation of the projection query for the upper layer is performed as follows: Steps include obtaining the center coordinates of the query selected in the previous layer, The steps include: calculating a scaling factor based on the aspect ratio of the feature maps of adjacent layers and projecting the center coordinates onto the current layer; The steps include finding the closest 2x2 neighboring feature points with projected coordinates within the feature map of the current layer, The method according to Appendix 1, characterized by comprising the step of selecting valid feature points as projection queries for the upper layer based on a mask and category score.

[0066] (Note 9) The method according to Appendix 1, characterized in that, in step 7, sinusoidal position coding and projection by a multilayer perceptron are performed on the reference point to obtain a position embedding vector.

[0067] (Note 10) The method according to Appendix 1, wherein in step 8, the Transformer decoder comprises six decoder layers, each layer comprising a self-attention mechanism, a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation, and the final prediction results are matched by the Hungarian method and ground truth.

Claims

1. An object detection method based on dynamic weights and a hierarchical query selection strategy, Step 1 involves preprocessing the input image, extracting multiscale feature maps via a backbone network, and obtaining image feature tensors, shape, and mask information for each layer's feature map. Step 2 involves performing convolution, flattening, and dimensional rearrangement operations on the image feature tensor of the multiscale feature map, inputting it into a Transformer encoder in combination with a position code, and obtaining an enhanced image feature sequence and its mask. Step 3 involves optimizing the aforementioned image feature sequence, performing initial predictions using a classification head and a detection head to obtain category scores and location information, and generating a global mask by non-maximum value suppression processing. Step 4 involves calculating the dynamic weights of the feature maps in each layer based on the shape of the feature map, the image feature sequence, the category score, and the global mask. Step 5 involves assigning the number of queries for each layer's feature map based on the aforementioned dynamic weights, wherein each layer's query includes two parts: a local query and a projected query for the upper layer. Step 6 involves selecting local queries and projected queries from the upper layer, layer by layer, where the projected queries from the upper layer are obtained by projecting the queries selected in the previous layer onto the feature map of the current layer. Step 7 involves using the selected query as a content query and combining it with its corresponding location information to generate a reference point and a location embedding vector. Step 8 includes inputting content queries, reference points, and position embedding vectors into the Transformer decoder, and outputting the final object detection result using a classification head and a detection head. In step 3, the method for optimizing the image feature sequence is as follows: Each feature map is traversed sequentially to create a two-dimensional grid coordinate system covering the pixel position of each feature map in the current layer. These coordinate axes are integrated into a tensor of shape H, W, 2, representing the coordinates of each pixel, where H and W represent the height and width of the feature map. Based on the resulting mask, the effective width and height of each image in the current layer are calculated and used as scaling factors. The grid coordinates are normalized to the effective region, and the width and height of the ground truth box where the grid coordinate values ​​normalized to the effective region are located are calculated based on a hierarchical index. Next, the grid coordinates are scaled, and then the scaled grid coordinates are combined with the width and height of the ground truth box where the grid coordinate values ​​are located to obtain a tensor of shape N, H*W, 4. For each of the four elements, a quadruple is constructed, represented by the center coordinate x-axis, y-axis, width, and height. Each quadruple represents a candidate proposal, and valid prediction proposals are selected based on whether each proposal value falls within a predetermined threshold range. Finally, a logarithmic transformation is performed on the selected proposals, and this transformation maps the coordinates from the [0,1] space to the real space to obtain the initial proposal results. Finally, an object detection method characterized by setting invalid proposed values ​​among the initial proposed results to null values ​​according to the obtained mask, filling the invalid regions of the image feature sequence with zeros, and obtaining a new feature sequence.

2. The method according to claim 1, characterized in that, in step 1, the backbone network is ResNet-50, and the output multiscale feature map includes feature maps of four different scales.

3. The method according to claim 1, characterized in that in step 2, the position code is calculated using a sine function and a cosine function.

4. The method according to claim 1, wherein the Transformer encoder comprises six encoder layers, each encoder layer comprising a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation.

5. In step 4, the method for calculating the dynamic weights is as follows: The method according to claim 1, characterized by first extracting effective feature points and their category scores from the feature map of each layer, calculating the L2 norm for each effective feature point, multiplying the category score by the L2 norm to obtain an importance score, averaging the importance scores of all effective feature points in the feature map of each layer to obtain initial weights, normalizing the initial weights of all layers, and finally obtaining the final weights of the feature map of each layer.

6. The method according to claim 1, wherein in step 5, the number of queries in the feature map of each layer is obtained by multiplying the weight of that layer by the total number of queries, the feature map of the top layer includes only local queries, and the remaining layers include local queries and projected queries from the upper layers.

7. In step 6, the generation of the projection query for the upper layer is performed as follows: Steps include obtaining the center coordinates of the query selected in the previous layer, The steps include: calculating a scaling factor based on the aspect ratio of the feature maps of adjacent layers and projecting the center coordinates onto the current layer; The steps include finding the closest 2x2 neighboring feature points with projected coordinates within the feature map of the current layer, The method according to claim 1, further comprising the step of selecting valid feature points as projection queries for a higher layer based on a mask and category score.

8. The method according to claim 1, characterized in that in step 7, sinusoidal position coding and projection by a multilayer perceptron are performed on the reference point to obtain a position embedding vector.

9. The method according to claim 1, wherein in step 8, the Transformer decoder comprises six decoder layers, each layer comprising a self-attention mechanism, a multiscale deformable attention mechanism, a feedforward network, a dropout layer, and a layer normalization operation, and the final prediction results are matched by the Hungarian method and ground truth.

Citation Information

Patent Citations

  • Image retrieval method based on multi-scale feature extraction

    CN108764246A

  • Image human-object interaction positioning method and system based on Transform model

    CN114663915A

  • Method and device for retrieving picture

    JP1996137908A

  • Pre-training framework for neural networks

    JP2023007419A