An object detection method based on attention mechanism and contrastive learning loss function
By introducing attention mechanism and comparative learning loss function into the object detection algorithm, combined with rotation data enhancement, the problem that the single-stage object detection algorithm cannot guarantee feature invariance and isodenality at the same time is solved, achieving higher object detection accuracy and faster training convergence.
Patent Information
- Application Number
- CN202211053660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The single-stage object detection algorithm cannot guarantee the invariance and isotropy of the features simultaneously in the processing of image features, resulting in the inability to further improve the object detection accuracy.
The object detection method based on attention mechanism and contrast learning loss function is adopted. By introducing spatial and channel attention mechanisms into the network, three encoders are designed, combining rotation data augmentation and comparison learning, and adding classification comparison loss and regression comparison loss, the network shows rotation invariance in the classification branch and rotation in the regression branch.
The network's classification and regression ability of the target is improved, the target detection accuracy is improved, and the training convergence process is accelerated while ensuring accuracy.
Smart Images

Figure CN115424004B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence and computer vision, and particularly relates to an object detection method based on attention and contrast learning. Background Art
[0002] Object detection is an algorithm that can output the categories and positions of several objects in an image. Different from image classification algorithms, it involves not only classification but also regression tasks. With the continuous development of deep learning, the accuracy of detection algorithms based on neural networks is much higher than that of traditional detection algorithms based on manually extracted features, and the generalization ability of neural networks is stronger. Many object detection algorithms based on deep learning can be classified into two categories, namely one-stage object detection algorithms and two-stage object detection algorithms. The difference between the two lies in whether they need to generate candidate boxes in advance. One-stage object detection algorithms are widely used because they do not need to generate candidate boxes in advance, have a more concise network architecture, a simpler training process, and a faster inference speed.
[0003] The attention mechanism was first applied to the field of natural language processing. However, due to its excellent performance, it was soon applied to the field of computer vision. For example, DETR based on a multi-head attention network introduced the attention mechanism into the object detection task and used the Hungarian algorithm for end-to-end training and inference, which has research value. Therefore, many variants of DETR have been proposed by researchers. For example, Conditional DETR that introduces conditional spatial queries speeds up network convergence, UP-DETR that adds unsupervised contrast learning improves object detection accuracy, and DeformableDETR that uses a method of focusing on sparse spatial localization greatly reduces the time required for network convergence while ensuring accuracy. In the network design of the present invention, three encoders are designed by using spatial and channel attention mechanisms in appropriate network layers, so that the network finally outputs the category and position prediction results based on different types of Anchors. Compared with DETR, this network replaces the decoder with an encoder of the same scale, and this encoder can strengthen the relationship between the classification branch and the localization branch, making these two different types of tasks more unified in the prediction results and achieving faster convergence while ensuring accuracy.
[0004] In the Internet environment, there is a vast amount of data, but this data cannot be directly used to train network models because there are no corresponding labels. If you want to obtain data labels, it often requires manual annotation, and the time and financial costs brought by this are extremely high. To break this dilemma, contrast learning emerged and is widely used in the field of unsupervised training. Without labels, it is only necessary to compare the data to train the network, and good results can be achieved by simply fine-tuning on a small amount of data sets containing labels.
[0005] Since contrastive learning is mainly used in the pre-training task of image classification, while contrastive learning is rarely used in object detection tasks. However, objects of the same class are comparable, and this aspect is often overlooked in existing research, which will lead to the inability to improve the classification ability and regression ability of the network simultaneously. This is disadvantageous for the entire object detection task. The present invention changes this unfavorable factor by enhancing the contrast between four rotated images. Specifically, the present invention adopts a contrastive form in the training process, enhances the images through four rotations, mines the comparable parts in object detection, and adds a classification contrast loss and a regression contrast loss, so that the classification branch exhibits rotational invariance, and at the same time makes the localization branch exhibit rotational equivariance. By differentiating the characteristics of the two branches, the network better conforms to the object detection task, thereby improving the detection accuracy. Summary of the Invention
[0006] The purpose of the present invention is to provide an object detection method based on an attention mechanism and a contrastive learning loss function, aiming to enhance the network's classification ability for objects and regression ability for object positions simultaneously, and solve the problem that the object detection accuracy cannot be further improved because the single-stage object detection algorithm cannot ensure the invariance and equivariance of features simultaneously in the processing of image features.
[0007] The technical solution adopted by the present invention is: an object detection method based on an attention mechanism and a contrastive learning loss function, comprising the following steps:
[0008] Step 1: Augment the original images in the dataset images, perform rotation operations of 0 degrees, 90 degrees, 180 degrees, and 270 degrees counterclockwise on each original image respectively, generating four views with different rotation angles. Every four homologous images form a group. Among them, the images rotated by 0 degrees constitute the original image dataset, and the images with other rotation angles constitute the rotated image dataset rotation images, and then proceed to step 2.
[0009] Step 2: Construct an initial online network for Encoder-only Contrast-based Object Detection with Transformer (ECODT). This network consists of three parts, namely the Transformer Encoder Backbone (TEB), the Transformer Encoder Neck (TEN), and the Transformer Encoder Head (TEH). The TEB extracts features at different scales of the image, the TEN aligns and fuses the image features, and the TEH decouples the prediction network and makes predictions. Construct the initial ECODT target network in the same way, and then go to Step 3.
[0010] Step 3: Define the update method of the initial ECODT online network as backpropagation of the total loss gradient. Through this method, the trained ECODT online network is obtained. Define the update method of the initial ECODT target network as the EMA method. Through this method, the trained ECODT target network is obtained, and then go to Step 4.
[0011] Step 4: Input the original image dataset into the initial ECODT target network, and input the rotated image dataset into the initial ECODT target network. Both networks will finally output four high-level semantic features, namely class features, feature point coordinates, class Logits, and bounding boxes. Align the class features and feature point coordinates output by the two networks respectively to obtain the target contrast group for calculating the contrast loss, and then go to Step 5;
[0012] Step 5: Constrain the class features using supervised contrast learning to minimize the feature differences of the same-class objects, and calculate the supervised class contrast loss; align the feature point coordinates, and then make the positions of the same object in different views satisfy the rotation affine transformation relationship, and then calculate the MSE loss; calculate the Focal Loss of the class Logits; calculate the GIOU loss of the bounding boxes, and then go to Step 6.
[0013] Step 6: Perform weighted summation on the supervised class contrast loss, MSE loss, Focal Loss, and GIOU loss to calculate the total loss value L, and then go to Step 7.
[0014] Step 7: In the training stage, perform gradient backpropagation on the initial ECODT online network to update the network parameters of the initial ECODT online network, obtaining the trained ECODT online network; select the online network with the highest score on the validation set as the final ECODT online network. In the detection stage, input the detection image into the final ECODT online network to obtain class Logits and bounding boxes, and use the class Logits to perform non-maximum suppression (NMS) on the bounding boxes to obtain the final prediction results and complete the detection.
[0015] Compared with the prior art, the significant advantages of the present invention are as follows:
[0016] (1) By combining the form of rotation data augmentation and contrastive learning, the present invention enhances the robustness of the network, enabling the network to exhibit rotational invariance in the classification branch and rotational equivariance in the regression, that is, differentiating the task attributes of the classification and localization branches.
[0017] (2) By combining the attention mechanism and designing three encoders, the present invention enhances the consistency of the classification and localization branches in perceiving the same target, improving the training convergence speed while ensuring accuracy, which is beneficial to the training, transformation, and promotion of the Transformer network.
[0018] (3) The loss function of the present invention has portability. When rotation augmentation is not performed, the same-class contrast loss can also be added. This contrast is between all the same-class targets within a batch, and this supervised class contrast can be conveniently added to the classification branches of other classification tasks or detection tasks. Description of the Drawings
[0019] Figure 1 is the flowchart of the method of the present invention.
[0020] Figure 2 is the overall structural schematic diagram of the present invention for enhancing the object detection ability through contrastive learning.
[0021] Figure 3 is the overall model network structure diagram of the present invention.
[0022] Figure 4 is the structural schematic diagram of the network Backbone of the present invention.
[0023] Figure 5 is the structural schematic diagram of the network Neck of the present invention.
[0024] Figure 6 is the structural schematic diagram of the network Head of the present invention. Detailed Embodiments
[0025] The present invention will be further described in detail below with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0026] Combined with Figures 1 to 6 , the object detection method based on the attention mechanism and the contrast learning loss function of the present invention includes the following steps:
[0027] Step 1: Augment the original images in the dataset images. Rotate each original image counterclockwise by 0 degrees, 90 degrees, 180 degrees, and 270 degrees respectively to generate four views with different rotation angles. Every four homologous images form a group. Among them, the images rotated by 0 degrees constitute the original image dataset, and the images with other rotation angles constitute the rotation image dataset rotation images.
[0028] Step 2: Construct an initial Encoder-only Contrast-based Object Detection with Transformer (abbreviated as ECODT) online network. This network consists of three parts, namely Transformer Encoder Backbone (abbreviated as TEB), Transformer Encoder Neck (abbreviated as TEN), and Transformer Encoder Head (abbreviated as TEH). TEB extracts different scale features of the image, TEN aligns and fuses the image features, and TEH decouples the prediction network and makes predictions. Construct the initial ECODT target network in the same way. The specific steps are as follows:
[0029] Step 2.1: Combined with Figure 3 , the structure of ECODT can be summarized as follows:
[0030] Input each original image in the dataset images into TEB, where represents the real number space, H represents the image height, W represents the image width, C represents the number of image channels, and extract four scale features, which are The corresponding scale is Among them, represents the feature at 4 times downsampling, represents the feature at 8 times downsampling, represents the feature at 16 times downsampling, represents the feature at 32 times downsampling, C 1 represents the number of channels at 4 times downsampling, C 2 represents the number of channels at 8 times downsampling, C3 Indicates the number of channels during 16x downsampling, C 4 Indicates the number of channels during 32x downsampling. The features of these four scales are input into TEN for feature alignment and multi-scale feature fusion to obtain feature f ten , which is input into TEH, and finally four types of high-level semantic features are obtained, namely class feature (Preclsfeat), class Logits (Pre cls), feature point coordinates (Pre boxfeat), and bounding box (Pre box).
[0031] Step 2.2, Combine Figure 4 , The structure and data transmission process of TEB can be summarized as follows:
[0032] After the original image image is successively input into the linear layer and patch embedding layer, the feature Then After successively outputting the patch merging layer, spatial-wise attention layer, channel-wise attention layer, and up sample layer, the obtained feature is added to the original , and the result is passed into the channel-wise attention layer to obtain the feature of the first scale Then is input into the down sample layer to obtain the feature While is transmitted along a network path similar to , and the feature of the second scale can be obtained and the feature for continuing forward propagation in the network Similarly, the feature of the third scale and the feature can also be obtained in a similar way. Finally, the feature can be used to obtain the feature of the fourth scale The feature of the i-th scale obtained by TEB will be successively input into TEN.
[0033] Step 2.3, Combine Figure 5 , The structure and data transmission process of TEN can be summarized as follows:
[0034] After inputting into the down sample layer, it is combined with Add them up, then input them into the Channel-wise Attention layer and the Down sample layer successively, and then combine with Add them up, then input them into the Channel-wise Attention layer and the Down sample layer, and then combine with Add them up, finally input them into a Channel-wise Attention layer, and finally obtain the feature f ten , each downsampling operation aligns two features to be added Each channel attention layer assigns higher weights to the key information in the fused feature map, and the final fused feature map f ten will be passed into the TEH; this process can be described as:
[0035]
[0036] where N CA represents the channel attention layer, and N DS represents the downsampling layer.
[0037] Step 2.4, combine Figure 6 , the structure and data transmission process of the TEH can be summarized as follows:
[0038] First, input the fused feature map f ten obtained from Step 2.3 into several Groupchannel-wise Attention layers to align the target category and bounding box in the feature domain, and then pass this aggregated feature into three branches respectively. The first branch directly outputs the category feature, denoted as Pre_cls_feat. The second branch passes through a Convolution layer and outputs the category Logits, denoted as Pre_cls. The third branch first passes through several Channel-wise Attention layers and then is divided into two sub-branches. One of them directly outputs to obtain the feature point coordinates, denoted as Pre_box_feat, and the other passes through a Convolution layer and outputs the bounding box, denoted as Pre_box. This process can be described as:
[0039]
[0040]
[0041] Pre_cls = N C (Pre_cls_feat),
[0042]
[0043] Pres_box = N C (Pre_box_feat)
[0044] where N GCA represents the grouped channel attention layer of the m-th layer, N C represents the convolutional layer, N CA represents the channel attention layer; represents the features obtained from the first split, and the class features are The class Logits are The feature point coordinates are The bounding box is na represents the number of Anchors, and nc represents the number of classes.
[0045] Step 2.5, the structure and data transmission process of the initial ECODT target network are as follows:
[0046] The structure of the initial ECODT target network is the same as that of the initial ECODT online network, but the input data is different. The input of the structure of the initial ECODT target network is the rotation image dataset rotation_images.
[0047] Step 3, define the update method of the initial ECODT online network as backpropagation of the total loss gradient. Through this method, the trained ECODT online network is obtained. Define the update method of the initial ECODT target network as the EMA method. Through this method, the trained ECODT target network is obtained. Combine Figure 2 The following elaborates on the specific process:
[0048] Step 3.1, the initial ECODT online network is used to process the original image, and parameter updates are performed according to backpropagation of the gradient to obtain the trained ECODT online network:
[0049] P a = N online (a) = TEH θ (TEN θ (TEB θ (a)))
[0050] where a represents the original image, N online represents the initial or trained ECODT online network, including network parameters TEB θ 、TEN θ and TEH θ , P aRepresent four features of the a image after passing through the initial or trained ECODT online network, namely Pre_cls_feat a , Pre_cls a , Pre_box_feat a and Pre_box a .
[0051] Step 3.2: The update method of the initial ECODT target network ECODT is to use the online network to calculate the exponential moving average (EMA) to obtain the trained ECODT target network, which is used to process rotated images:
[0052] P x = N target (x) = TEH ξ (TEW ξ (TEB ξ (x))), x ∈ {a, b, c},
[0053]
[0054] where b, c, d respectively represent the enhanced images rotated counterclockwise by 90, 180, and 270 degrees, N target represents the initial or trained ECODT target network, including network parameters TEB ξ , TEN ξ and TEH ξ , P b , P c , P d represents the four features obtained for each of the three rotated images after passing through the target network; nb represents the number of batches in one iteration, and τ represents the momentum coefficient.
[0055] Step 4: Input the original image dataset into the initial ECODT target network, and input the rotated image dataset into the initial ECODT target network. Both networks will finally output four high-level semantic features, namely class features, feature point coordinates, class Logits, and bounding boxes. Align the class features and feature point coordinates output by the two networks respectively to obtain the target comparison group for calculating the comparison loss. The specific steps are as follows:
[0056] Step 4.1: Rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 90 degrees 90 degrees clockwise in the channel dimension, rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 180 degrees 180 degrees clockwise in the channel dimension, and rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 270 degrees 270 degrees clockwise in the channel dimension. The clockwise rotation operation on these feature point coordinates is to align with the original image for supervised class comparison. After alignment, the channels at the same position of the feature are responsible for predicting the same target. Due to the Anchor mechanism, each channel at a position can predict na different scales of the same target. When performing class feature comparison, only the class feature corresponding to the Anchor that best matches the Ground truth is selected for comparison.
[0057] Step 4.2: Similar to the way of aligning with the class feature, it is also necessary to perform corresponding clockwise rotation alignment operations on the feature point coordinates of the rotated image. However, after rotation alignment, affine alignment is also required. That is, first regard each 512-channel length as 256 coordinate points, that is, 16×16 coordinate pairs. These coordinate pairs are regarded as the feature point coordinates of each grid point, and clockwise rotation affine transformation is performed on these feature points. The formula for affine transformation is as follows:
[0058]
[0059]
[0060] where (x,y) represents the predicted feature point coordinates, θ is the clockwise rotation angle, and (x rot ,x rot ) represents the feature point coordinates of the aligned original image.
[0061] Step 5: Calculate four network losses: Use supervised contrast learning to constrain the class feature to make the features of the same-class targets as similar as possible, and calculate the supervised class loss; Align the feature point coordinates, and then make the positions of the same target in different views satisfy the rotation affine transformation relationship, and then calculate the MSE loss; Calculate the Focal Loss for the class Logits; Calculate the GIOU loss of the bounding box. The specific steps are as follows:
[0062] Step 5.1: Perform supervised class comparison on the target, which can be divided into two categories according to whether four types of rotation augmentations are performed. The first category is without rotation augmentation, then the same-class targets of different source images within a batch can be used for comparison. The advantage is easy transplantation, and the disadvantage is that if the batch setting is too small and there are no same-class coordinates within the batch, no class feature comparison will be performed.
[0063]
[0064]
[0065] where L SCC represents the class contrast loss, N is the number of targets, is the class contrast loss of the same-class targets in the original image, is the total number of targets with label y i , τ is the temperature coefficient, z i and z j are the feature vectors of different targets.
[0066] The second category is to perform rotation augmentation, then the class contrast loss of the same-class targets in the homologous images can be calculated. The advantage of this method is that there are at least four same-class targets in each batch, unless there are no targets in the original image. The disadvantage is that it requires more computing resources. The corresponding class feature contrast formula is as follows:
[0067]
[0068]
[0069] where is the class contrast loss between the same-class targets in the homologous images. This construction method is equivalent to splicing four rotated images into a large image, and then comparing the same-class targets in this large image to ensure the inevitability of the existence of the same-class targets in the image.
[0070] Step 5.2: Compare the feature point coordinates of the aligned targets to make the network more sensitive to the position transformation of the same target. The feature point comparison formula is as follows:
[0071]
[0072] where w, h, and c are the lengths of the three dimensions of the feature point coordinates, Y and Y' are the feature point coordinates of two different homologous images, and n is the number of combinations of two different homologous images, which refers to .
[0073] Step 6: Perform weighted summation on the supervised class contrast loss, MSE loss, FocalLoss, and GIOU loss to calculate the total loss value L. The specific steps are as follows:
[0074] L = ω 1 L cls + ω 2 L box + ω 3 L Scc + ω 4 LMSE
[0075] where L cls is the classification branch loss of TEH, and L box is the bounding box loss of the regression branch of TEH, and ω 1 , ω 2 , ω 3 , ω 4 are the weights of the four types of losses respectively, which are used to scale different losses to the same scale in order to balance each loss.
[0076] Step 7: In the training stage, perform gradient backpropagation on the initial ECODT online network to update the network parameters of the initial ECODT online network, and obtain the trained ECODT online network; select the online network with the highest score on the validation set as the final ECODT online network. In the detection stage, input the detection image into the final ECODT online network to obtain class Logits and bounding boxes, and use the class Logits to perform non-maximum suppression (NMS) on the bounding boxes to obtain the final prediction result and complete the detection.
[0077] The series of detailed descriptions listed in the present invention are only specific descriptions of the feasible implementation manners of the present invention, and they are not used to limit the protection scope of the present invention. Any equivalent implementation manners or changes made without departing from the spirit of the inventive art should be included in the protection scope of the present invention.
Claims
1. An object detection method based on an attention mechanism and a contrast learning loss function, characterized in that, the method comprises the following steps: Step 1, augment the original images in the dataset images, perform rotation operations on the original images one by one by rotating counterclockwise by 0 degrees, 90 degrees, 180 degrees, and 270 degrees respectively, and generate four views with different rotation angles correspondingly. Every four homologous images form a group. Among them, the images rotated by 0 degrees constitute the original image dataset, and the images with other rotation angles constitute the rotation image dataset rotation images, and then go to Step 2; Step 2, construct an initial ECODT online network. The initial ECODT online network includes TEB, TEN, and TEH. TEB extracts different scale features of the image, TEN aligns and fuses the image features, and TEH decouples the prediction network and makes predictions. Construct an initial ECODT target network in the same way, and then go to Step 3; Step 3, define the update method of the initial ECODT online network as backpropagation of the total loss gradient, and obtain the trained ECODT online network in this way. Define the update method of the initial ECODT target network as the EMA method, and obtain the trained ECODT target network in this way, and then go to Step 4; Step 4, input the original image dataset into the initial ECODT target network, and input the rotation image dataset rotation images into the initial ECODT target network. Both networks will finally output four high-level semantic features, namely category features, feature point coordinates, category Logits, and bounding boxes. Align the category features and feature point coordinates output by the two networks respectively to obtain a target comparison group for calculating the contrast loss, and then go to Step 5; Step 5, constrain the category features in a supervised contrast learning manner to minimize the feature differences of the same-class objects, and calculate the supervised category contrast loss; align the feature point coordinates, and then make the positions of the same object in different views satisfy the rotation affine transformation relationship, and then calculate the MSE loss; calculate the Focal Loss of the category Logits; calculate the GIOU loss of the bounding boxes, and then go to Step 6; Step 6, perform weighted summation on the supervised category contrast loss, MSE loss, Focal Loss, and GIOU loss to calculate the total loss value L, and then go to Step 7; Step 7, in the training stage, perform backpropagation of the gradient on the initial ECODT online network to update the network parameters of the initial ECODT online network to obtain the trained ECODT online network; take the online network with the highest score on the validation set as the final ECODT online network. In the detection stage, input the detection image into the final ECODT online network to obtain the category Logits and the bounding boxes, and use the category Logits to perform non-maximum suppression on the bounding boxes, abbreviated as NMS, to obtain the final prediction result and complete the detection.
2. The object detection method based on an attention mechanism and a contrast learning loss function according to claim 1, characterized in that: In step 2, an initial ECODT online network is constructed. The initial ECODT target network includes TEB, TEN, and TEH. TEB extracts features of different scales of the image, TEN aligns and fuses the image features, and TEH decouples the prediction network and makes predictions. The initial ECODT target network is constructed in the same way. The specific steps are as follows: Step 2.1, The structure and data transmission process of the initial ECODT online network are as follows: Each original image in the dataset images is input into TEB, where represents the real number space, H represents the image height, W represents the image width, C represents the number of image channels, and four scales of features are extracted, which are The corresponding scales are Among them, represents the feature at 4x downsampling, represents the feature at 8x downsampling, represents the feature at 16x downsampling, represents the feature at 32x downsampling, C 1 represents the number of channels at 4x downsampling, C 2 represents the number of channels at 8x downsampling, C 3 represents the number of channels at 16x downsampling, C 4 represents the number of channels at 32x downsampling. The features of these four scales are input into TEN for feature alignment and multi-scale feature fusion to obtain the feature f ten , which is input into TEH, and finally four types of high-level semantic features are obtained, namely class feature, class Logits, feature point coordinates, and bounding box; Step 2.2, The structure and data transmission process of TEB are as follows: After the original image is input into the linear layer and the block encoding layer successively, features are obtained. Then After being output from the block fusion layer, the spatial attention layer, the channel attention layer, and the upsampling layer successively, the obtained features are added to the original ones, and the result is then passed into the channel attention layer to obtain the features of the first scale Then is input into the downsampling layer to obtain features And is transmitted along a network path similar to that of to obtain the features of the second scale and the features for continuing the forward propagation Similarly, the features of the third scale and the features can also be obtained in a similar way. Finally, the features of the fourth scale are obtained from the features The features of the i-th scale obtained by TEB where i ∈ {1, 2, 3, 4}, will be successively passed into TEN; Step 2.3, The structure and data transmission process of TEN are as follows: After inputting into the downsampling layer, it is added to , and then successively input into the channel attention layer and the downsampling layer, and then added to . Next, it is input into the channel attention layer and the downsampling layer, and then added to . Then, it is input into the channel attention layer and the downsampling layer, and then added to . Finally, it is input into a channel attention layer to finally obtain the feature f ten . Each downsampling operation aligns two features to be added i ∈ {1, 2, 3, 4}; each channel attention layer assigns higher weights to the key information in the fused feature map, and the final fused feature map f ten will be passed into the TEH; this process is described as: Where N CA represents the channel attention layer, and N DS represents the downsampling layer; Step 2.4, The structure and data transmission process of TEH are as follows: The fused feature map f obtained from step 2.3 ten is fed into several grouped channel attention layers to align the object categories and bounding boxes in the feature domain. Then, this aggregated feature is fed into three branches respectively. The first branch directly outputs the category feature, denoted as Pre_cls_feat. The second branch passes through a convolutional layer and outputs the category Logits, denoted as Pre_cls. The third branch first passes through several channel attention layers and is then divided into two sub-branches. One sub-branch directly outputs the feature point coordinates, denoted as Pre_box_feat, and the other sub-branch passes through a convolutional layer and outputs the bounding box, denoted as Pre_box. This process is described as follows: Pre_cls = N C (Pre_cls_feat), Pres_box = N C (Pre_box_feat) Among them, N GCA represents the grouped channel attention layer of the m-th layer, N C represents the convolutional layer, N CA represents the channel attention layer; represents the features obtained from the first shunt, and the category features are The category Logits are The feature point coordinates are The bounding box is na represents the number of Anchors, and nc represents the number of categories; Step 2.5, The structure and data transmission process of the initial ECODT target network are as follows: The structure of the initial ECODT target network is the same as that of the initial ECODT online network, but the input data is different. The input of the structure of the initial ECODT target network is the rotation images dataset.
3. The object detection method based on the attention mechanism and the contrast learning loss function according to claim 2, characterized in that: In step 3, the update method of the initial ECODT online network is defined as the backpropagation of the total loss gradient. Through this method, the trained ECODT online network is obtained. The update method of the initial ECODT target network is defined as the EMA method. Through this method, the trained ECODT target network is obtained. Specifically as follows: Step 3.1, The initial ECODT online network is used to process the original image, and parameter updates are performed according to the backpropagation of the gradient to obtain the trained ECODT online network: P a = N online (a) = TEH θ (TEN θ (TEB θ (a))) Among them, a represents the original image, N online represents the initial or trained ECODT online network, including network parameters TEB θ , TEN θ and TEH θ , P a represents four features of the a image after passing through the initial or trained ECODT online network, namely Pre_cls_feat a , Pre_cls a , Pre_box_feat a and Pre_box a ; Step 3.2, The update method of the initial ECODT target network is to calculate the exponential moving average EMA using the online network to obtain the trained ECODT target network, and this target network is used to process the rotation image: P x = N target (x) = TEH ξ (TEN ξ (TEB ξ (x))), x ∈ {b, c, d}, where b, c, and d respectively represent the enhanced images rotated counterclockwise by 90, 180, and 270 degrees, and N target represents the initial or trained ECODT target network, including network parameters TEB ξ , TEN ξ and TEH ξ , P b , P c , P d represent the four features obtained for each of the three rotated images after passing through the target network; nb represents the number of batches in one iteration, and τ represents the momentum coefficient.
4. The object detection method based on the attention mechanism and the contrast learning loss function according to claim 3, characterized in that: The original image dataset is input into the initial ECODT target network, and the rotation images dataset is input into the initial ECODT target network. Both networks will finally output four high-level semantic features, namely class features, feature point coordinates, class Logits, and bounding boxes. Alignment operations are performed on the class features and feature point coordinates output by the two networks respectively to obtain the target comparison group for calculating the contrast loss. The specific steps are as follows: Step 4.1: Rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 90 degrees clockwise in the channel dimension, rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 180 degrees clockwise by 180 degrees in the channel dimension, and rotate the class feature Pre_cls_feat obtained from the image rotated counterclockwise by 270 degrees clockwise by 270 degrees in the channel dimension. The clockwise rotation operation on these feature point coordinates is to align with the original image for supervised class comparison. After alignment, the channels at the same position of the feature are responsible for predicting the same target. Due to the Anchor mechanism, each channel at a position can predict na different scales of the same target. When performing class feature comparison, only the class feature corresponding to the Anchor that best matches the Groundtruth is selected for comparison. Step 4.2: Similar to the way of aligning class features, it is also necessary to perform corresponding clockwise rotation alignment operations on the feature point coordinates of the rotated image. However, after rotation alignment, affine alignment is also required. That is, first consider each 512-channel length as 256 coordinate points, that is, 16×16 coordinate pairs. These coordinate pairs are regarded as the feature point coordinates of each grid point, and clockwise rotation affine transformation is performed on these feature points. The formula for affine transformation is as follows: where (x, y) represents the predicted coordinates of the feature points, θ is the clockwise rotation angle, and (x rot , y rot ) represents the coordinates of the feature points of the aligned original image.
5. The object detection method based on the attention mechanism and the contrast learning loss function according to claim 4, characterized in that: Four network losses are calculated. The class features are constrained by the method of supervised contrast learning to make the features of the same class of objects as similar as possible, and the supervised class contrast loss is calculated. The feature point coordinates are aligned, and then the positions of the same target in different views are made to satisfy the rotation affine transformation relationship, and then the MSE loss is calculated. Focal Loss is calculated for the class Logits. The GIOU loss of the bounding box is calculated. The specific steps are as follows: Step 5.1: Perform supervised class comparison on the target. It can be divided into two categories according to whether four types of rotation augmentations are performed. The first category is without rotation augmentation, and the same-class objects of different source images within a batch can be used for comparison. The advantage is that it is easy to transplant, and the disadvantage is that if the batch size is set too small and there are no same-class coordinates within the batch, class feature comparison will not be performed. where L SCC represents the class contrast loss, N is the number of targets, is the in-class target contrast loss within the original image, is the total number of targets with label y i , τ is the temperature coefficient, z i and z j are the feature vectors of different targets; The second category is with rotation augmentation, and the class comparison loss of the same-class objects in the homologous images can be calculated. The advantage of this method is that it can ensure that there are at least four same-class objects in each batch, unless there are no objects in the original image. The disadvantage is that it requires more computing resources. The corresponding class feature comparison formula is as follows: Among them is the contrast loss of the same type of targets between homologous images. This construction method is equivalent to splicing four rotated images into a large image, and then comparing the same type of targets in this large image to ensure the inevitability of the existence of the same type of targets in the image; Step 5.2: Compare the feature point coordinates of the aligned targets to make the network more sensitive to the position transformation of the same target. The feature point comparison formula is as follows: where w, h, and c are the lengths of the three dimensions of the feature point coordinates, Y and Y' are the feature point coordinates of two different homologous images, and n is the number of combinations of two different homologous images, which refers to C in the present invention 2 4 。 6. The object detection method based on the attention mechanism and the contrast learning loss function according to claim 5, characterized in that: Perform a weighted sum of the supervised class contrast loss, MSE loss, Focal Loss, and GIOU loss to calculate the total loss value L. The specific steps are as follows: L = ω 1 L cls + ω 2 L box + ω 3 L SCC + ω 4 L MSE where L cls is the classification branch loss of TEH, and L box is the bounding box loss of the regression branch of TEH, and ω 1 、ω 2 、ω 3 、ω 4 are the weights of the four types of losses respectively, used to scale different losses to the same scale in order to balance each loss.
Citation Information
Patent Citations
Zero sample image target detection method and device based on deep learning
CN113255829A
Pedestrian cross-mirror re-identification method and system based on background and attitude normalization
CN114120363A