Transform-based multi-view image target detection method
By introducing a hybrid encoder and feature query module into the multi-view image object detection method, combined with the multi-view consistency loss function, the problems of insufficient robustness of view angle changes and difficulty in matching targets in the prior art are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510242469.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-16
AI Technical Summary
The existing multi-view image object detection methods are not robust enough when dealing with viewing angle changes, and it is difficult to correctly associate and match the same target, especially when the target object is blocked.
Using a multi-view image object detection method based on Transformer, multi-scale features are extracted through the backbone network, and a hybrid encoder and feature query module are introduced into the multi-view image object detection module to generate an image feature sequence with more semantic information. At the same time, the detection results are optimized using the multi-view consistency loss function.
A richer and more detailed image feature sequence is achieved, more attention is paid to high-scale features, and the robustness and accuracy of detection is improved, especially in the case of viewing angle changes and target occlusion.
Smart Images

Figure CN120014246A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing methods, and in particular to a Transformer-based multi-view image target detection method. Background Art
[0002] Existing multi-view image target detection methods are mainly divided into Transformer-based methods and multi-view Figure 1 The Transformer-based method uses self-attention to capture the global dependency between images of different viewpoints, adopts multi-head attention to learn features of different scales, improves detection accuracy by fusing information from different viewpoints, processes image data of all viewpoints in parallel, and adds position encoding to the Transformer to retain the spatial information of the image. The encoder is used to process the input image, and the decoder generates the detection result. However, in multi-view image object detection, the Transformer model may be sensitive to viewpoint changes, and additional design is required to ensure the robustness of the model to different viewpoints.
[0003] Based on multi-view Figure 1 The consistency method takes advantage of the fact that the target information observed from different viewpoints should be consistent. The core idea of this method is that if an object is detected in multiple viewpoints, then these detection results should be consistent with each other. This method performs target detection in each viewpoint separately, finds the features of the same object in different viewpoints through feature matching, and the consistency check is responsible for verifying whether the attributes of the object in different viewpoints are consistent. Finally, the consistent detection results are combined to optimize the position and attributes of the object. However, this method still has limitations in how to correctly associate and match the same object in different viewpoints, especially when the target is occluded due to the change of viewpoint. Summary of the invention
[0004] The technical problem to be solved by the present invention is how to provide a multi-view image target detection method that can obtain richer and more detailed image feature sequences, pay more attention to high-scale features, and better perform detection.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: a multi-view image target detection method based on Transformer, comprising the following steps: First, the input left camera image and the image from the right camera The multi-scale features of the image are extracted through the backbone network, and the parameters are shared between the two backbone networks; Subsequently, the extracted image multi-scale features are input into the multi-view image object detection module for processing. The view image object detection module consists of two parts: a hybrid encoder module and a feature query module, which perform encoding operations in sequence to generate an image feature sequence, and a feature query module to filter the sequence to generate a new image feature sequence with more semantic information. Finally, the Transformer-based decoder and detection head are used to perform decoding operations and final predictions to obtain image detection results under multiple perspectives.
[0006] The beneficial effects of adopting the above technical solution are: the method described in the present application is based on multi-view fusion in the input stage, and adopts left and right camera images for input, giving full play to the advantage that RGB images can provide rich semantic information, and obtaining a richer and more detailed image feature sequence; on the basis of the traditional encoder-decoder, a new hybrid encoder structure is introduced to remove low-scale features and pay more attention to high-scale features. This structure can better capture the connection between concepts and entities in the image, thereby facilitating the detection and recognition of objects in the image by subsequent modules; at the same time, a multi-view consistency loss function is used to match and optimize each query point, and better detection is carried out by jointly considering predictions under multiple different versions of perspectives and learning viewpoint equivariance. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0008] Figure 1 is a processing flow chart of the method described in an embodiment of the present invention; Figure 2 is a diagram of a backbone network structure in the method described in an embodiment of the present invention; Figure 3 is a structural diagram of a multi-view image target detection module in the method described in an embodiment of the present invention; Figure 4 is a structural diagram of a cross-scale feature fusion module in the method described in an embodiment of the present invention; Figure 5 is a structural diagram of a fusion module in the method described in an embodiment of the present invention; Figure 6 It is a structural diagram of a decoder and a detection head in the method described in an embodiment of the present invention. DETAILED DESCRIPTION
[0009] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0010] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0011] Overall, such as Figure 1 As shown, the embodiment of the present invention discloses a multi-view image target detection method based on Transformer, comprising the following steps: First, the input left camera image and the image from the right camera The multi-scale features of the image are extracted through the backbone network. and Parameter sharing between Subsequently, the extracted image multi-scale features are input into the multi-view image object detection module for processing. The view image object detection module consists of two parts: a hybrid encoder module and a feature query module, which perform encoding operations in sequence to generate an image feature sequence, and a feature query module to filter the sequence to generate a new image feature sequence with more semantic information. Finally, the Transformer-based decoder and detection head are used to perform decoding operations and final predictions to obtain image detection results under multiple perspectives.
[0012] 1) Backbone network: The overall structure of the network includes a series of convolutional layers and pooling layers, which are used to extract image features. The structure diagram is as follows Figure 2 The figure shows the operation of the left image P1. The operation process of the right image P2 is the same. The backbone network includes multiple convolution kernels and ReLU activation functions. The 13 convolution layers are divided into five stages, respectively. , , , , Indicates that different numbers of The convolution kernel performs convolution operations on the input image to complete feature extraction from local to global. The feature maps extracted at different stages have different dimensions and channels, which are called multi-scale features. As the convolution operation continues to deepen, the richer the semantic information included in the extracted features, the more conducive to image recognition tasks, and are called high-level features. This structure enables the backbone network to capture rich visual features and achieve relatively outstanding results in a variety of image recognition tasks.
[0013] 2) Multi-view image object detection module The multi-scale features obtained by the backbone network are input into the multi-view image object detection module. First, the hybrid encoder module mainly optimizes the image features. Secondly, the IOU-aware query selection method is used in the feature query module to select a fixed number of image features from the encoder output sequence. The Transformer decoder is responsible for continuously updating the input query through repeated operations of multiple decoding layers. Finally, the predicted detection box is output through the detection head. The architecture of this module is shown in the figure below. Figure 3 shown.
[0014] Its input is the high-level multi-scale features extracted from the backbone network, which are input into an efficient hybrid encoder for encoding processing. The hybrid encoder module gradually improves the model communication accuracy by decoupling the multi-scale feature interaction into two-step operations of intra-scale interaction and cross-scale fusion, while significantly reducing the computational cost, and obtains an optimized image feature sequence. This sequence is screened in the feature query module and input into the Transformer-based decoder for decoding operations. Finally, the detection result is obtained through the detection head, and a detection box is generated in the image.
[0015] 2-1) Hybrid encoder module In the encoding process, high-level multi-scale features often have richer semantic features. If self-attention operations are applied to high-level features, the connection between concepts and entities in the image can be better captured, thereby facilitating the detection and recognition of objects in the image by subsequent modules. At the same time, due to the lack of semantic concepts and the risk of duplication and confusion with high-level features, the intra-scale interaction of low-level features is unnecessary. In summary, the method described in this application only utilizes the last three stages of the backbone ( ) as the input of the hybrid encoder. First, the internal scale feature interaction module only processes the last stage in the backbone network. The output, The feature map is first flattened in the module through the Flatten operation, and then input into the self-attention mechanism for calculation. Then, the features processed by the self-attention mechanism are reshaped into Compared with the previous method of first using a single-scale Transformer encoder to complete the intra-scale interaction and then using a Panet-like structure to perform cross-scale fusion, the cross-scale feature fusion module proposed in this application is further adjusted on the basis of the original Transformer encoder. Its main structure is as follows Figure 4 Show: After being processed by the internal scale feature interaction module The feature map will be combined with the feature maps of the third and fourth stages of the backbone network The SiLU activation function is an improved version of Sigmoid and ReLU. SiLU has the characteristics of no upper bound, lower bound, smoothness, and non-monotonicity. It has been verified that the effect of this function on deep models is better than ReLU. Multiple fusion blocks are added to this module. The fusion block mainly includes multiple convolutional layers. Its main function is to add features from different stages of the input, and its function is to fuse them to generate a new feature. The main structure of the fusion block is as follows: Figure 5 As shown: The main computational operations of the hybrid encoder proposed in this application for multi-scale image features can be expressed by the following formula: Among them, Q, K, and V represent the vectors obtained by linear transformation of the input vector in the Transformer. The Flatten operation flattens the multi-dimensional feature map into one dimension for self-attention calculation; reshape represents the reshaping operation, which aims to restore the feature shape to Attention represents multi-head self-attention operation; It represents the unified image feature obtained by fusing multi-scale features together.
[0016] 2-2) Feature Query Module This application adds an IOU-based feature query module to the proposed network, and the image feature sequence obtained by hybrid encoder encoding is input into this module. The main function of this module is to constrain the network model. The main constraint part is concentrated in the training process. Its purpose is to select features with higher IOU scores and generate corresponding higher classification scores for them. Correspondingly, for features with lower IOU scores, lower classification scores are generated. In this way, when the training model selects the top-ranked features obtained by the encoder according to the classification scores, it can ensure that these features have both high classification scores and high IOU scores. Finally, after screening, we get , the specific operations are as follows: in, and g represent the predicted value and the true value respectively, c and b represent the category and the detection box respectively, , ; This query method introduces the IoU score into the objective function of the classification branch, thereby achieving consistency constraints on the classification and positioning of positive samples.
[0017] 3) Decoder and detection head The new image feature sequence obtained by the feature query module is input into the decoder together with the conditional query. The decoder includes multiple Transformer layers. In each Transformer layer, self-attention and cross-attention operations are performed. After a series of self-attention and cross-attention layers, the conditional query is updated to include aggregated information from multi-view images. At the same time, for the conditional query, the method described in the present application applies a multi-view consistency loss function to it, which takes into account multiple predictions from different query views of the same query point to ensure their geometric consistency in different views.
[0018] Subsequently, the query updated by the decoder will be input into the detection head, which includes 3 fully connected layers. The fully connected layer learns the high-level features extracted from the RoI feature map and classifies the RoI based on these features to determine which category it belongs to in the dataset (people, cars, bicycles). In addition to classification, the fully connected layer is also responsible for fine-tuning the bounding box position of each RoI. This step is to locate the object more accurately. At the same time, the fully connected layer outputs a set of numerical values representing the four coordinates of the adjusted bounding box (x and y coordinates of the upper left and lower right corners), making the detection result more accurate. Finally, the detection result is output. The architecture diagram of the decoder and detection head is shown in the figure. Figure 6 As shown: 3-1) Geometry Encoder Module The geometry encoder in this module is mainly used to process object queries, which are based on 2D query points. and query view T v = [ q ¯ v , t v ] Build.
[0019] First, initialize a set of 2D query points in the global coordinate system These points are usually randomly initialized and learned and updated via backpropagation during training. Subsequently, for each 2D query point , define a query view T v = [ q ¯ v , t v ] Each query view includes two parameters, is a quaternion representing the view direction, is the translation vector representing the view position. At the same time, the query view consists of two parts, namely the default global view and the additionally generated virtual view. The main purpose of generating virtual views is to enhance the robustness of the model to viewpoint changes. These virtual views are generated by sampling Euler angles from a uniform distribution. and translation vector To obtain, for the Euler angle and translation vector, define them to be uniformly distributed within a certain range, which is generally [ 0 , 2 π ] , randomly sample two values from the uniform distribution defined above, corresponding to the rotation angles around the x and y axes. These values are expressed as , , and then convert the Euler angles to the equivalent quaternion , this process is completed through the quaternion rotation formula, the conversion process is as follows: The converted quaternion Combined with the sampled translation vector, a query view is constructed . Then, each 2D query point need to be transformed into This operation is done by inverse transformation To achieve this, namely: in is the transformed query point. Finally, the transformed query point Combined with query view to form geometry query set [ cv j , q ¯ v , t v ] The geometric query set will be Fourier transformed and MLP operated to generate conditional query , the formula is as follows: q v j = MLPdec ( γ [ cv j , q ¯ v , t v ] ) in, represents the Fourier transform, represents a multilayer perceptron used to decode queries. During training, each query point Combined with each query view, a set of conditional queries is generated. These conditional queries and the new image feature sequence output by the feature query module Together, they are input into the decoder for self-attention and cross-attention calculations to extract aggregate information from image features of different views. In this way, the network framework proposed in this application uses virtual perspectives to learn perspective consistency during training, which helps the model better understand the scene structure in the image.
[0020] 3-2) Multi-view consistency loss function Based on multi-view object detection, the consistency in multi-view geometric space is an important part to consider. Assuming that query viewpoints, then the bounding box prediction from a single query point There will be corresponding versions. Although these bounding boxes are in different coordinate systems, they actually correspond to the same true value target. According to multi-view geometry, observing the same target, even from different coordinate systems, is geometrically constrained, and the only difference is the transformation matrix. Based on this idea, this application applies a multi-view consistency loss function in the decoder and detection head, jointly considering multi-view predictions from the same query point and targets from all query views. The main calculation steps are as follows: use Indicates the number of virtual views generated in the geometry encoder, plus the original default view, for a total of During the training process, the generated virtual view will be different from the default query view. Figure 1 First, we need to ensure that the query points To achieve this goal, a collection can be created by concatenating the predictions from different query views. , that is, the set of predictions for different query views Get in touch: X ^ j = [ x ^ j 0 , x ^ j 1 ,..., x ^ j V ] This collection contains all The prediction results of the view angles are: Indicates that for The set of prediction boxes for query points, Indicates that in the default query view Next for the The prediction box information of the query point, Indicates that the network The prediction information for the query point under each query view, and then, based on the label data provided in the dataset, we can get all the ground truth values The bounding boxes also form a set at each query view : (g represents the bounding box of the ground truth) : G m = [ g m 0 , g m 1 ,..., g m v ] in, represents the ground truth bounding box in the default view, Indicated in The ground truth bounding box obtained from each query view.
[0021] Next, Hungarian matching is performed using the cost function to determine { }and{ The optimal allocation between , the cost function includes the classification score and regression loss , the specific formula is as follows: in, is the ground truth label, is the weighted L1 loss function, which is defined by the following formula: Among them, the prediction results of each virtual view and the ground-truth bounding box The differences between them are weighted and summed, and the predicted box set is finally calculated. and the ground truth set The weighted regression loss Once the optimal allocation is determined Next, we will calculate the multi-view consistency loss on the above set : For the prediction box set and the ground truth set , By classification loss and regression loss Composition, multiply them by weights and And finally sum them up to get . This constraint function uses focal loss as the classification loss and the same form of regression loss as in matching.
[0022] In general, the role of this loss function is to better perform detection by jointly considering the predictions under multiple different versions of the perspective for each query point during the matching and optimization process and by learning viewpoint equivariance.
[0023] This application uses a hybrid encoder to fuse high-level multi-scale features of the image and a multi-view consistency loss function to deal with the problem that the same object has deviations from the ground truth value in two versions of the view.
Claims
1. A multi-view image object detection method based on Transformer, characterized in that The steps include: First, the input left camera image and the image from the right camera The multi-scale features of the image are extracted through the backbone network, and the parameters are shared between the backbone networks; Subsequently, the extracted image multi-scale features are input into the multi-view image object detection module for processing. The multi-view image object detection module consists of two parts: a hybrid encoder module and a feature query module, which perform encoding operations in sequence to generate an image feature sequence, and a feature query module to filter the sequence to generate a new image feature sequence with more semantic information. Finally, the Transformer-based decoder and detection head are used to perform decoding operations and final predictions to obtain image detection results under multiple perspectives.
2. The Transformer-based multi-view image object detection method according to claim 1, characterized in that: The multi-scale features obtained through the backbone network are input into the multi-view image object detection module for processing.
3. The Transformer-based multi-view image object detection method according to claim 1, characterized in that: The backbone network includes multiple convolution kernels and ReLU activation functions. The 13 convolution layers are divided into five stages. , , , , Indicates that different numbers of The convolution kernel performs a convolution operation on the input image to complete feature extraction from local to global. The feature maps extracted at different stages have different dimensions and numbers of channels, which are called multi-scale features.
4. The Transformer-based multi-view image object detection method according to claim 1, characterized in that: The multi-view image object detection module includes a hybrid encoder module and a feature query module. The hybrid encoder module is used to optimize image features. Its input is the high-level multi-scale features extracted from the backbone network, which are input into the hybrid encoder for encoding. The hybrid encoder module decouples the multi-scale feature interaction into two steps: intra-scale interaction and cross-scale fusion, to obtain an optimized image feature sequence. The feature query module uses an IOU-aware query selection method to select a fixed number of image features from the hybrid encoder output sequence.
5. The Transformer-based multi-view image object detection method according to claim 4, characterized in that: The hybrid encoder module includes a cross-scale feature fusion module, which utilizes the backbone network The output features of the three stages are used as input. First, the internal scale feature interaction module only processes the last stage in the backbone network. The output, The feature map is first flattened in the module through the Flatten operation, and then input into the self-attention mechanism for calculation. Then, the features processed by the self-attention mechanism are reshaped into , and then the Panet structure is used to carry out cross-scale fusion; after being processed by the internal scale feature interaction module Feature map, and feature map of the third and fourth stages of the backbone network They are input together into the cross-scale feature fusion module; multiple fusion blocks are added to the module, and the fusion block includes multiple convolutional layers, whose main function is to add operations to the features of different stages input, and its function is to fuse them to generate a new feature.
6. The Transformer-based multi-view image object detection method according to claim 5, characterized in that: The calculation operation of the hybrid encoder for multi-scale features is expressed by the following formula: ; ; ; Among them, Q, K, and V represent the vectors obtained by linear transformation of the input vector in Transformer. The Flatten operation flattens the multi-dimensional feature map into one dimension. Reshape represents the reshaping operation, which is used to restore the feature shape to Attention represents multi-head self-attention operation. It represents the unified image feature obtained by fusing multi-scale features together.
7. The Transformer-based multi-view image object detection method according to claim 4, characterized in that: The feature query module is an IOU-based feature query module, and the image feature sequence is obtained by encoding with a hybrid encoder. It is input into this module to constrain the network model. The constraint part is concentrated in the training process to select features with higher IOU scores. The specific operations are as follows: ; in, and g represent the predicted value and the true value respectively, c and b represent the category and the detection box respectively, , .
8. The Transformer-based multi-view image object detection method according to claim 1, characterized in that: The new image feature sequence obtained after processing by the multi-view image object detection module is input into the decoder together with the conditional query. The decoder includes multiple Transformer layers. In each Transformer layer, self-attention and cross-attention operations are performed. After a series of self-attention and cross-attention layers, the conditional query is updated, including the aggregated information from the multi-view image. At the same time, for the conditional query, a multi-view consistency loss function is used. Subsequently, the query updated by the decoder is input into the detection head, which includes 3 fully connected layers. The fully connected layer learns the high-level features extracted from the RoI feature map, and classifies the RoI according to these features to determine which category it belongs to in the data set. The fully connected layer fine-tunes the position of each bounding box. At the same time, the fully connected layer outputs a set of values representing the four coordinates of the adjusted bounding box, and finally outputs the detection result.
9. The Transformer-based multi-view image object detection method according to claim 8, characterized in that: The input of the conditional query includes the processing result of the geometric encoder and the multi-view consistency function; The geometry encoder is used to process object queries based on 2D query points. and query view Build; First, initialize a set of 2D query points in the global coordinate system , these points are randomly initialized and learned and updated through back-propagation during the training process; Then, for each 2D query point , define a query view Each query view includes two parameters, a quaternion representing the view direction, The translation vector representing the view position; at the same time, the query view consists of two parts, namely the default global view and additionally generated virtual views ; The generated virtual view is used to enhance the robustness of the model to viewpoint changes; the virtual view is generated by sampling Euler angles from a uniform distribution and translation vectors To obtain, for the Euler angle and translation vector, define it in Uniformly distributed in the range, two values are randomly sampled from the uniform distribution defined above, corresponding to the rotation angles around the x and y axes respectively; these values are expressed as , , and then convert the Euler angles to the equivalent quaternion , this process is completed through the quaternion rotation formula, the conversion process is as follows: ; The converted quaternion Combined with the sampled translation vector, construct a virtual view ; Then, each 2D query point is transformed into In the virtual view defined; this operation is done by inverse transformation To achieve this, namely: ; in, is the transformed query point; finally, the transformed query point Combined with query view to form geometry query set ; The geometric query set undergoes Fourier transform and MLP operations to generate conditional queries , the formula is as follows: ; in, represents the Fourier transform, represents a multilayer perceptron used to decode queries. During training, each query point Combined with each query view, a set of conditional queries is generated; these conditional queries and the new image feature sequence output by the feature query module Together, they are input into the decoder for self-attention and cross-attention calculation to extract aggregate information from image features of different views.
10. The Transformer-based multi-view image object detection method according to claim 9, characterized in that: The multi-view consistency function is constructed by the following steps: Based on multi-view object detection, given V query perspectives, the bounding box prediction from a single query point There will be V versions; use Indicates the number of virtual views generated in the geometry encoder, plus the original default view, a total of During the training process, the generated virtual view will be used for training together with the default query view; first, we need to ensure that the images from the same query point in these viewpoints The detection boxes of each version are assigned to the same ground truth object; to achieve this goal, a collection is created by concatenating the predictions from different query views. , aggregate predictions from different query views Get in touch: ; This collection contains all The prediction results of the view angles are: Indicates that for The set of prediction boxes for query points, Indicates that in the default query view Next for the The prediction box information of the query point, represents the network's prediction information for the query point under the V+1th query view. Then, based on the label data provided in the dataset, we can get all the ground truth values The bounding boxes also form a set at each query view : ; in, represents the ground truth bounding box in the default view, Indicated in The ground truth bounding box obtained in each query view; Next, Hungarian matching is performed using the cost function to determine { }and{ The optimal allocation between , the cost function includes the classification score and regression loss , the specific formula is as follows: ; in, is the ground truth label, is the weighted L1 loss function, which is defined by the following formula: ; Among them, the prediction results of each virtual view and the ground-truth bounding box The differences between them are weighted and summed, and the predicted box set is finally calculated. and the ground truth set The weighted regression loss , when the optimal allocation is determined After that, calculate the multi-view consistency loss on the above set : ; For the prediction box set and the ground truth set , Including classification loss and regression loss , the classification loss and regression loss Multiply by the weights and And finally summed up, the constraint function uses focal loss as the classification loss and the regression loss in the same form as in matching.
Citation Information
Cited By
Vehicle body part detection method and device
CN121095934A