A training method, annotation method, device and electronic device for an inspection target annotation model

By using paired image samples and common attention module technology in robot inspection tasks, the problem of insufficient image labeling accuracy in key scenes in data centers is solved, and a higher automatic labeling accuracy of patrol targets is achieved.

CN116229464BActive Publication Date: 2025-06-03CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310167075.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-06-03
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

In robot inspection tasks for key scenes in data centers, there are many types of objects in the image. The existing weak supervision detection methods are prone to incomplete labeling results and label errors, which cannot meet the labeling accuracy requirements of image training data.

Method used

By constructing a paired image sample dataset, two OICR networks shared by weight parameters are input, the common attention module is used to mine the same and different features in the image features, calculate the total loss of the model and adjust the weight parameters, and extract the inspection target annotation model to improve the annotation accuracy.

Benefits of technology

It significantly improves the accuracy of automatic labeling of inspection targets, reduces the impact of the wide variety of objects in the image on category recognition, and alleviates the local optimal labeling results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229464B_ABST
    Figure CN116229464B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, annotation method, device and electronic device for an inspection target annotation model. The training method includes: constructing a paired image sample data set including at least one type of the same category of objects, inputting the sample pairs into two weakly supervised detection networks OICR with shared weight parameters respectively, and using the feature convolutional layer in the OICR network to extract the first image feature and the second image feature; proposing a common attention module to mine the same features and different features between the first image feature and the second image feature; obtaining the preliminary annotation results of the targets in the two images output by the two OICR networks; calculating the total loss of the model based on the preliminary annotation results, the same features and the different features; adjusting the weight parameters in the total model based on the total loss of the model, and then extracting the OICR network from the total model as the inspection target annotation model. The technical solution provided by the present invention improves the accuracy of annotating inspection targets in robot inspection images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot target detection, and particularly to a patrol target annotation model training, annotation method, device and electronic device. Background Art

[0002] In the robot patrol task for key scenarios of the data center, an important task is to detect targets. A prerequisite for achieving accurate target detection is to use a large number of high-quality annotated images as training samples. The annotation accuracy of the target bounding box directly determines the accuracy of the detection model. However, at present, the annotation of industrial images is mainly completed manually, and it is very time-consuming and laborious to annotate the images accurately to the border level. In this case, automatic annotation technologies based on strong supervised learning have emerged. However, the models of strong supervised learning still need a large number of image samples accurately annotated to the border level for training, and the workload of generating such samples is still very large. In recent years, with the development of weak supervision technologies, the weak supervision detection network only needs to be trained with image-level annotation data to output annotation results. Since the image-level annotation data only requires the image category and does not need to annotate the border of the actual position of the object, the annotation efficiency has been greatly improved.

[0003] In 2016, Hakan Bilen et al. proposed a network structure WSDDN in the paper "Weakly Supervised Deep Detection Networks", which regarded image object detection as Multiple Instance Learning (MIL). This network used a pre-trained CNN as a feature descriptor, and then combined with a region recommendation algorithm to form a two-stream network for screening out categories and bounding boxes with high target scores. In 2017, Tang Peng et al. proposed a new network structure OICR in the paper "Multiple Instance Detection Network with Online Instance Classifier Refinement". This network was improved based on the WSDDN network, and proposed the idea of candidate region clustering. For the classification layer of MIL, multiple classifiers were deployed. Except for the first classifier trained with image-level annotations, the other classifiers used the bounding boxes and categories output by the previous classifier as the labels of the next classifier for training. Through this refinement process, the categories of the predicted bounding boxes were corrected, avoiding the output bounding boxes only focusing on a certain most characteristic part of the target and falling into local optimality, and improving the accuracy of object detection. In 2019, Zeng et al. proposed to improve OICR by adding a bounding box regressor at the end of OICR in the paper "WSOD2: Learning Bottom-up and Top-down Objectness Distillation for Weakly-supervised Object Detection", and adjusted the target bounding boxes output by the MIL-refinement network in it, enabling the network to automatically analyze the probability that the selected target is a complete object, and further improving the accuracy of the bounding boxes.

[0004] However, in the robot inspection task in the key scenarios of the data center, there are a wide variety of objects in the images collected by the robot, and some objects are similar and difficult to distinguish, which increases the difficulty of the neural network in learning features. The images labeled by existing weakly supervised detection methods are prone to problems such as incomplete labeling results and incorrect labeling labels, and cannot meet the labeling accuracy requirements of image training data. Summary of the Invention

[0005] In view of this, the embodiments of the present invention provide an inspection target annotation model training, annotation method, device and electronic device, thereby improving the accuracy of annotating inspection targets in robot inspection images.

[0006] According to a first aspect, an embodiment of the present invention provides a method for training an inspection target annotation model. The method includes: constructing a paired image sample data set including at least one same-category object, and separately inputting the paired image samples into two OICR networks with shared weight parameters, and using the feature convolutional layers in the two OICR networks to extract a first image feature and a second image feature respectively; using a preset common attention module to mine the same features and different features between the first image feature and the second image feature; respectively outputting preliminary annotation results of the targets in the two image samples through the two OICR networks; calculating a total model loss based on the preliminary annotation results, the same features and the different features, where the total model loss is the loss of a total model composed of two OICR networks with shared weight parameters and the common attention module; adjusting the weight parameters of the total model based on the total model loss, and extracting the OICR network from the total model as the inspection target annotation model.

[0007] Optionally, using a preset common attention module to mine the same features in the first image feature and the second image feature includes: classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature belonging to each category and the scores of each position in the second image feature belonging to each category; calculating a similarity index between all positions of the first image feature and all positions of the second image feature pairwise based on the scores of each position in the first image feature and the second image feature belonging to each category; extracting the same features from the first image feature and the second image feature respectively according to the similarity index.

[0008] Optionally, classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature and the second image feature belonging to each category includes: sequentially inputting the current image feature into a category-aware convolutional layer and a 1×1 convolutional layer to generate a category-aware activation map, where the current image feature is the first image feature or the second image feature; performing global average pooling on the category-aware activation map to obtain a category score vector of the current image feature, and the category score vector is used to represent the scores of each position in the current image feature belonging to each category.

[0009] Optionally, calculating a similarity index between all positions of the first image feature and all positions of the second image feature based on the scores of each position of the first image feature and the second image feature belonging to each category includes: flattening the first image feature and the second image feature into two-dimensional matrices respectively, and corresponding the scores of each position of the first image feature and the second image feature belonging to each category to their respective flattened two-dimensional matrices; calculating a similarity matrix using the two-dimensional matrix of the first image feature and the two-dimensional matrix of the second image feature, and each element in the similarity matrix is used to represent the similarity index between all positions of the first image feature and all positions of the second image feature pairwise.

[0010] Optionally, extracting the same feature from the first image feature and the second image feature respectively according to the similarity index includes: performing column normalization and row normalization on the similarity matrix respectively to obtain a first attention map matching the dimension of the flattened first image feature and a second attention map matching the dimension of the flattened second image feature; multiplying the first attention map and the flattened first image feature element by element to obtain a first common attention feature, and the first common attention feature is used to represent the object corresponding to the common category label of the paired images on the first image; multiplying the second attention map and the flattened second image feature element by element to obtain a second common attention feature, and the second common attention feature is used to represent the object corresponding to the common category label of the paired images on the second image; processing the first common attention feature and the second common attention feature respectively through a category-aware convolutional layer and a 1×1 convolutional layer, and using the processed first common attention feature and second common attention feature as the same feature.

[0011] Optionally, mining different features in the first image feature and the second image feature by using a co-attention module includes: respectively inputting the first co-attention feature and the second co-attention feature into a sigmod function for processing; calculating the difference between 1 and the processed first co-attention feature to obtain a first contrast attention map, and calculating the difference between 1 and the processed second co-attention feature to obtain a second contrast attention map; multiplying the first contrast attention map and the first image feature element by element to obtain a first contrast co-attention feature, and multiplying the second contrast attention map and the second image feature element by element to obtain a second contrast co-attention feature, where the first contrast co-attention feature is used to represent the objects unique to the first image, and the second contrast co-attention feature is used to represent the objects unique to the second image; respectively processing the first contrast co-attention feature and the second contrast co-attention feature through a category-aware convolutional layer and a 1×1 convolutional layer, and using the processed first contrast co-attention feature and second contrast co-attention feature as the different features.

[0012] Optionally, calculating the total loss of the model based on the preliminary annotation result, the same features, and the different features includes: calculating the total loss of the model according to the following loss function relationship

[0013]

[0014]

[0015]

[0016]

[0017]

[0018] where α and β represent weights, represents the loss function corresponding to the OICR network sharing two weight parameters, created based on the preliminary annotation result, represents the loss function of the co-attention module, I m and I n respectively represent the m-th image and the n-th image in the image samples, and I m and I n are a pair of images including at least one same-class object, is a loss function created according to the scores of each position in the image features of the m-th and n-th images belonging to each category, is a loss function created by mining the same features, is the loss function created by mining the different features, l m and l n respectively represent the image-level labels of the m-th image and the n-th image, s m and s n respectively represent the class score vectors of the m-th image and the n-th image, represents the common score vector of the m-th image containing common-class objects calculated based on the same features of the m-th image, represents the common score vector of the n-th image containing common-class objects calculated based on the same features of the n-th image, l m ∩l n represents the common-class label, represents the unique score vector of the m-th image containing unique objects calculated based on the different features of the m-th image, represents the unique score vector of the n-th image containing unique objects calculated based on the different features of the n-th image, l m \l n and l n \l m respectively represent the unique-class labels of the m-th image and the n-th image, and σ represents the cross-entropy operation.

[0019] According to a second aspect, an embodiment of the present invention provides an inspection target annotation method, the method including: obtaining an image to be annotated; inputting the image to be annotated into the inspection target annotation model pre-trained by any optional implementation manner of the first aspect, and outputting the annotation result of the inspection target annotation model for the image to be annotated.

[0020] According to a third aspect, an embodiment of the present invention provides an inspection target annotation model training device, the device including: a bounding box online generation module, configured to construct a paired image sample data set including at least one type of common-class objects, and separately input the paired image samples into two OICR networks with shared weight parameters, and use the feature convolutional layers in the two OICR networks to extract first image features and second image features respectively; and further configured to output preliminary annotation results of the targets in the two image samples through the two OICR networks respectively; a common attention module, configured to mine the same features and different features between the first image features and the second image features; a loss calculation module, configured to calculate the total model loss based on the preliminary annotation results, the same features and the different features, where the total model loss is the loss of the total model composed of the two OICR networks with shared weight parameters and the common attention module together; a parameter adjustment module, configured to adjust the weight parameters of the total model based on the total model loss, and extract the OICR network in the total model as the inspection target annotation model.

[0021] According to a fourth aspect, an embodiment of the present invention provides an inspection target annotation device, where the device includes: a data acquisition module, configured to acquire an image to be annotated; and a target annotation module, configured to input the image to be annotated into an inspection target annotation model pre-trained by any optional implementation manner of the first aspect, and output an annotation result of the inspection target annotation model for the image to be annotated.

[0022] According to a fifth aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method described in the first aspect, the second aspect, or any optional implementation manner of the first aspect and the second aspect.

[0023] According to a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method described in the first aspect or any optional implementation manner of the first aspect.

[0024] The technical solution provided by this application has the following advantages:

[0025] The technical solution provided by this application automatically annotates an image to be annotated through a pre-trained inspection target annotation model, and outputs an annotation result of the inspection target annotation model for the image to be annotated. Among them, the training method of the inspection target annotation model provided by the embodiment of the present invention is as follows: First, a pair of image samples are respectively input into two OICR networks with shared weight parameters, and preliminary annotation results of the targets in the two images are respectively output. It should be noted that at least one object of the same category is included in the pair of image samples. Then, a first image feature and a second image feature extracted by a feature convolutional layer in the two OICR networks are respectively obtained, and the same features and different features between the first image feature and the second image feature are mined through a common attention module. After that, the total loss of the model is calculated based on the preliminary annotation result, the same feature, and the different feature. Finally, the weight parameters of the total model composed of the OICR network and the common attention module are adjusted based on the total loss of the model. After multiple rounds of training, the OICR network with adjusted weight parameters is extracted from the total model as the inspection target annotation model. Through the inspection target annotation model trained by the embodiment of the present invention, on the basis of the original OICR network, it can drive the classifier to mine the similarity information of objects with the same semantics and the difference information of objects with different semantics, improve the ability of the network to mine the potential similarity and difference between objects across images, reduce the impact of a large variety of objects in the image on category recognition, alleviate the situation of local optimality of the annotation result, and significantly improve the accuracy of automatic annotation of inspection targets through the annotation model provided by the embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as limiting the present invention in any way. In the drawings:

[0027] Figure 1 Fig. shows a schematic diagram of the steps of a patrol target annotation model training method in an embodiment of the present invention;

[0028] Figure 2 Fig. shows a schematic diagram of the model structure of a patrol target annotation model training method in an embodiment of the present invention;

[0029] Figure 3 Fig. shows a schematic diagram of the model structure of OICR in an embodiment of the present invention;

[0030] Figure 4 Fig. shows a schematic diagram of the model structure of the co-attention module in an embodiment of the present invention;

[0031] Figure 5 Fig. shows a schematic diagram of the process of extracting co-attention features in an embodiment of the present invention;

[0032] Figure 6 Fig. shows a schematic diagram of the steps of a patrol target annotation method in an embodiment of the present invention;

[0033] Figure 7 Fig. shows a schematic diagram of the structure of a patrol target annotation model training device in an embodiment of the present invention;

[0034] Figure 8 Fig. shows a schematic diagram of the structure of a patrol target annotation device in an embodiment of the present invention;

[0035] Figure 9 Fig. shows a schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Please refer to Figure 1 and Figure 2 , in one embodiment, a patrol target annotation model training method specifically includes the following steps:

[0038] Step S1. Construct a paired image sample dataset including at least one type of the same object, separately input the paired image samples into two OICR networks with shared weight parameters, and use the feature convolutional layers in the two OICR networks to extract the first image feature and the second image feature respectively.

[0039] Step S2. Use a preset co-attention module to mine the same features and different features between the first image feature and the second image feature.

[0040] Step S3. Output the preliminary annotation results of the targets in the two image samples through the two OICR networks respectively.

[0041] Specifically, the inspection target annotation model training method proposed in the embodiment of the present invention includes two modules. One module is a bounding box online generation module composed of two OICR networks with shared weight parameters, and the other module is a co-attention module for enhancing the performance of the OICR network. In this embodiment, the bounding box online generation module and the co-attention module are regarded as a whole, and the total model is called a weakly supervised target annotation model based on the cross-image co-attention mechanism (COCA). A relevant loss function is created for the total model for training. After the training is completed, the co-attention module is removed, and one OICR network in the trained bounding box online generation module is used for target annotation. During the training process, in order to pay more attention to the similar information of the same object in different images and the difference information of different objects, so that the training process of the OICR network can match the co-attention module proposed in this embodiment, the structure of the bounding box online generation module designed in this embodiment is composed of two OICR networks with shared weight parameters. During training, the paired image samples need to be separately input into the two OICR networks with shared weight parameters, and the annotation results of the targets in the two images are output respectively. According to the output annotation results, a loss function for this part of the OICR network can be created, and the model is trained according to the loss value.

[0042] To facilitate the understanding of the solution provided in the embodiment of the present invention, the introduction of the OICR network structure is as follows:

[0043] As Figure 2 and Figure 3 shown, taking Figure 2Taking the OICR network on the road as an example, the image sample extracts image features through the feature convolution layer; at the same time, the image sample extracts the regions of interest (i.e., the borders of interest) through the Selective Search algorithm; combining the image features and the regions of interest, the region-of-interest features of each region of interest are obtained; the region-of-interest features are normalized to a unified size through spatial pyramid pooling, and then the features are integrated through two fully connected layers. After that, the MIL-refinement module transforms the integrated features into high dimensions (e.g., K + 1 dimensions) and inputs them into multiple classification branches (e.g., K + 1 branches). The first classification branch includes two fully connected layer branches along the category direction and the candidate region direction. The two branches along the category direction and the candidate region direction respectively determine the position of each target border (the candidate region detected to contain the target) and the category score of each target border (candidate region) from the regions of interest through two softmax functions (the category score represents the probability of classifying a certain border as a certain type of object, e.g., the probability of classifying the current border as a person or a horse); summing and pooling the category scores of each border can calculate the overall category score of the current image sample. For example, the category score of the current image sample is a matrix: [0.1, 0, 0.2, 0.9, 0], and it can be judged that the current image sample has the highest probability of containing the 4th type of object; according to the overall category score of the image sample and the image-level label of the current image sample, the first classification branch can be trained. Then, the border coordinates and border categories output by the first classification branch are used as the pseudo-labels of the next classification branch. The next classification branch classifies the integrated features along the category direction and adjusts the weight parameters of the current classification branch with the pseudo-labels output by the previous classification branch as the standard. After the refinement processing of multiple classification branches, the MIL-refinement module outputs more accurate annotation results.

[0044] Among them, the total loss function for the MIL-refinement module to train each classification branch is:

[0045]

[0046] In the formula, is the loss function of the first classification branch of the MIL-refinement module, is the loss function of the classification branch used for refinement in the MIL-refinement module. and are both prior arts and will not be elaborated in this application.

[0047] The output of the MIL-refinement module will enter the online regressor. On the one hand, based on the output of the MIL-refinement module, the online regressor also includes a classification branch for predicting the discrete probability distribution of each candidate region. The created loss function is as follows:

[0048]

[0049] where R is the number of candidate regions, w r is the weight, is the classification result of the prediction target, is the discrete probability distribution for predicting each candidate region.

[0050] On the other hand, the online regressor predicts the offsets of the positions and sizes of each candidate region through the regression branch. For example, the offset of the k-th candidate region is The classical smooth-L 1 can be used to calculate the loss function of this branch:

[0051]

[0052] In the formula, represents the predicted bounding box offset by the online regressor, and λ r represents the regression loss weight, represents the offset of the r-th predicted bounding box with respect to the standard box, represents the offset of the pseudo-bounding box output by the MIL-refinement module with respect to the standard box.

[0053] Combining the classification branch loss function and the regression branch loss function of the online regressor, the total loss function of the online regressor is expressed as:

[0054]

[0055] where represents the classification loss, represents the regression loss, and λ is used to control the balance between these two parts of the loss.

[0056] Specifically, for the two-way OICR network of the bounding box online generation module, the feature convolutional layer of each OICR extracts image features, which are respectively called the first image feature F m and the second image feature F n . Then, in this embodiment, through the co-attention module, from the first image feature F m and the second image feature F nExtract the same features that can represent the same type of object between the two feature maps, and extract the different features of the unique objects between the two feature maps. And create a loss function for adjusting the weight parameters of the OICR network based on the same features and different features.

[0057] Step S4. Calculate the total model loss based on the preliminary annotation results, the same features and different features. The total model loss is the loss of the total model composed of the OICR network and the co-attention module sharing two weight parameters.

[0058] Step S5. Adjust the weight parameters of the total model based on the total model loss, and extract the OICR network from the total model as the inspection target annotation model.

[0059] Specifically, in order to train the OICR network and the co-attention module in the above steps S1 to S3, it is necessary to create the loss functions of each sub-network in the OICR network and the co-attention module respectively. In this embodiment, the loss functions for training the above-mentioned sub-networks are fused together to obtain the total model loss function. Use a large number of paired image samples for multiple rounds of training, adjust the weight parameters of the total model composed of the OICR network and the co-attention module according to the total model loss output in each round of training, and extract the OICR network from the total model as the inspection target annotation model after training. Therefore, the model training method provided by the embodiment of the present invention can improve the ability of the model to mine the potential similarities and differences between objects across images, and alleviate the situation of local optimality of the annotation results. The accuracy of automatic annotation of inspection targets is significantly improved by the annotation model provided by the embodiment of the present invention.

[0060] Specifically, as Figure 4 shown, in one embodiment, the above step S3 specifically includes the following steps:

[0061] Step 1: Classify the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature belonging to each category and the scores of each position in the second image feature belonging to each category.

[0062] Specifically, the co-attention module proposed in this embodiment has a total of 3 branches, namely the classification branch, the co-attention branch and the contrast co-attention branch. The processing flow of Step 1 belongs to the classification branch, and the processing flows of the following Step 2 and Step 3 belong to the co-attention branch. The processing process of the contrast co-attention branch will be described in subsequent embodiments.

[0063] The classification branch is used to classify the first image feature F m and the second image feature F nClassify each position in it to prepare for mining objects of the same and different categories in two paired images. The position refers to the pixel position in the image feature. For example, the first image feature F m has a dimension of C*H*W, indicating that the feature convolution layer obtains C images through C convolution channels. The number of pixels in each image is H*W, and each pixel represents a position. Converting it into the form of a 3D matrix gives the first image feature F with a dimension of C*H*W m .

[0064] Specifically, in this embodiment, the classification branch of the co-attention module is different from the fully connected layer classification branch in the above OICR network. A new classification branch structure is adopted to increase the training intensity of the overall network. It is specifically implemented in the following ways:

[0065] 1. Input the current image feature into the class-aware convolutional layer and the 1×1 convolutional layer in sequence to generate a class-aware activation map. The current image feature is the first image feature or the second image feature

[0066] 2. Perform global average pooling on the class-aware activation map to obtain the class score vector of the current image feature. The class score vector is used to represent the scores of each position in the current image feature belonging to each category

[0067] Specifically, assume that the robot patrol target training data set is where I n represents the nth training image, and the image-level label is represented by , indicating that the image contains the kth category of objects. Conversely K represents the number of object categories, and n' represents that the data set contains a total of n' image samples. Select at least an image pair {I m , I n} that contains the same object content from the data set, that is, l m ∩l n ≠0. Input I m into one path of OICR, and input I n into another path of OICR. After a series of feature convolutional layers to generate corresponding feature maps where C, H, and W respectively correspond to the number of channels, height, and width of the feature map; then, F m and F n are sequentially input into the class-aware convolutional layer with K dimensions and the 1×1 convolutional layer, thereby generating class-aware activation maps S m and S n :

[0068]

[0069]

[0070] Category-aware Convolutional Layer It is a layer of convolution after the feature convolutional layer, used to further deepen the semantic information in the first image feature map or the second image feature map, and deeply extract the category features in the map. After passing through the category-aware convolutional layer and the image features after the 1×1 convolutional layer, while deepening the semantic information, reduce the feature dimension. Next, the class-aware activation maps S m and S n After the global average pooling operation (Global Average Pooling, GAP), the category detection values at each position of the first image feature and the second image feature are normalized to obtain the category score vectors of I m respectively I n respectively

[0071]

[0072]

[0073] The values in the category score vector are used to represent the scores of each position in the image feature belonging to each category.

[0074] Specifically, in this embodiment, the classification branch of the co-attention module uses the following sigmoid cross-entropy (CE) loss function as the training objective to train the above-mentioned category-aware convolutional layer, 1×1 convolutional layer, and OICR network feature convolutional layer. And the training process will also affect the weight parameters of other parts of the OICR network.

[0075]

[0076] In the formula is the loss function of the classification branch of the co-attention module, which is created according to the scores of each position in the image feature of the m-th image belonging to each category and the scores of each position in the image feature of the n-th image belonging to each category. σ represents the cross-entropy operation, l m and l n respectively represent the image-level label of the m-th image and the image-level label of the n-th image, s m and s n respectively represent the category score vector of the m-th image and the category score vector of the n-th image. If the category detection in the m-th image and the n-th image is more accurate, then s m and s n are respectively closer to l m and l n respectively. The smaller the performance, the more it is used to guide the adjustment of the weight parameters of the co-attention module and the online bounding box generation module.

[0077] Step 2: Calculate the similarity index between every two positions of all positions of the first image feature and all positions of the second image feature based on the scores of each position of the first image feature and the second image feature belonging to each category.

[0078] Step 3: Extract the same features from the first image feature and the second image feature respectively according to the similarity index.

[0079] Specifically, in the embodiment of the present invention, the score vector calculated in Step 1 is used to perform similarity analysis between every two positions of all positions of the first image feature F m and all positions of the second image feature F n to obtain the similarity index between every two positions. For example, if a certain position of the first image feature F m and a certain position of the second image feature F n contain objects and are the same type of objects, then the scores of these two positions belonging to the same category should be similar, so the similarity is relatively high. If a certain position of the first image feature F m and a certain position of the second image feature F n do not contain objects, or the contained objects are not the same type of objects, then the score values of these two positions belonging to the same category should differ greatly, so the similarity is relatively low. Based on this, the similarity between each position of the two image features is characterized by the similarity index, and the features with relatively high similarity are extracted from the first image feature F m and the second image feature F n as the same features. This can drive the classifier to mine the similarity information of objects with the same semantics and improve the ability of the annotation model to mine the potential similarity between objects across images.

[0080] Specifically, in the embodiment of the present invention, the above Step 2 specifically includes the following steps:

[0081] Step 4: Flatten the first image feature and the second image feature into two-dimensional matrices respectively, and map the scores of each position of the first image feature and the second image feature belonging to each category to their respective flattened two-dimensional matrices.

[0082] Specifically, for the three-dimensional first image feature F m and the second image feature F n , it is difficult to directly analyze the similarity between each position of the two features. To accurately calculate the first image feature F m and the second image feature F nRegarding the similarity between each position, the embodiments of the present invention adopt the idea of converting three-dimensional features into two-dimensional matrices and performing calculations in the two-dimensional space. As Figure 5 shown, first, one of the two image features with dimensions C*H*W is converted into a two-dimensional matrix with dimensions C*HW, and the other is converted into a two-dimensional matrix with dimensions HW*C. The method of converting a three-dimensional matrix into a two-dimensional matrix is a prior art and will not be elaborated here. Specifically, the reshape function in Python can be referred to. For example, in this embodiment, the first image feature F m is flattened into a two-dimensional matrix with dimensions HW*C, and the second image feature F n is flattened into a two-dimensional matrix with dimensions C*HW. The scores of each position of the first image feature and the second image feature belonging to each category are correspondingly mapped to each position of the two-dimensional matrix in sequence.

[0083] Step Five: Calculate a similarity matrix using the two-dimensional matrix of the first image feature and the two-dimensional matrix of the second image feature. Each element in the similarity matrix is used to represent the similarity index between every two positions of all positions of the first image feature and all positions of the second image feature.

[0084] Specifically, through a two-dimensional similarity matrix algorithm, use the two-dimensional matrix of the first image feature and the two-dimensional matrix of the second image feature to calculate the similarity matrix of the two features, as shown in the following formula.

[0085]

[0086] It should be noted that in the formula, and are in the form of two-dimensional matrices after being flattened. is a learnable matrix. The similarity matrix P stores the similarity indexes of all position pairs in F m and F n , that is, the value at the position (i,j) in P is the similarity between the i-th position in F th and the j-th position in F m . n Specifically, in this embodiment, the above Step Three specifically includes the following steps:

[0087] Step Six: Perform column normalization and row normalization on the similarity matrix respectively to obtain a first attention map that matches the dimension of the flattened first image feature, and obtain a second attention map that matches the dimension of the flattened second image feature.

[0088] Step Seven: Multiply the first attention map and the flattened first image feature element by element to obtain a first co-attention feature, which is used to represent the object corresponding to the common category label of the paired images on the first image.

[0089] ​

[0090] Step eight: Multiply the second attention map and the flattened second image features element by element to obtain the second co-attention feature, which is used to represent the object corresponding to the common class label of the paired images on the second image.

[0091] Step nine: Process the first co-attention feature and the second co-attention feature through a class-aware convolutional layer and a 1×1 convolutional layer respectively, and use the processed first co-attention feature and second co-attention feature as the same feature.

[0092] Specifically, in this embodiment, the similarity matrix calculated in the above step five is normalized column by column and row by row. On the one hand, it is used to convert the elements in the similarity matrix to the range of 0 to 1. On the other hand, it is used to make the elements of the similarity matrix match the elements in the two-dimensional matrices of the first image features and the second image features, so as to facilitate matrix multiplication. For example, in this embodiment, the first image feature F m is flattened into a two-dimensional matrix of HW*C, and the second image feature F n is flattened into a two-dimensional matrix of C*HW. Therefore, the first attention map A m (with dimension HW*HW) obtained by normalizing the similarity matrix column by column can match the elements of the flattened two-dimensional matrix of HW*C of the first image feature F m . The first attention map A m and the flattened first image feature F m are multiplied element by element to obtain the first co-attention feature with dimension HW*C. Similarly, the second attention map A n (with dimension HW*HW) obtained by normalizing the similarity matrix row by row can match the elements of the flattened two-dimensional matrix of C*HW of the second image feature F n , and multiplying them gives the second co-attention feature of HW*C.

[0093] The formulaic representation of the above process is as follows:

[0094] A m = softmax(P) ∈ [0,1] HW×HW

[0095]

[0096] where softmax(·) represents normalization along the column direction.

[0097] Next, multiply the attention map and the flattened two-dimensional matrix of the image features to obtain the first co-attention feature and the second co-attention feature

[0098]

[0099]

[0100] wherein represents the image I m and I n The corresponding object on F of the common label m is represents the object corresponding to the common label on F n The symbol ∩ represents the intersection. This is because the essence of co-attention is to calculate the similarity between all pairs of positions between I m and I n For example: if an element at a certain position in the first attention map A m is closer to 1, the feature of the first image feature F m at this position is multiplied by a number close to 1 and retained. If an element at a certain position in the first attention map A m is closer to 0, the feature of the first image feature F m at this position is multiplied by a number close to 0 and deleted. Thus, only the common object semantics in the two paired images will be retained in the two co-attention features.

[0101] After that, similar to the classification branch, the first co-attention feature and the second co-attention feature are input into the class-aware convolutional layer to generate the class-aware activation map and to obtain the same features between the two images:

[0102]

[0103]

[0104] Finally, using the GAP operation, the common score vector for predicting the common class can be obtained:

[0105]

[0106]

[0107] wherein, the common score vector of the first image represents the probability that the first image contains the common class object, and the common score vector of the second image represents the probability that the second image contains the common class object. Then, the common score vectors of the first image and the second image are used as the same features mined from the two paired images in the embodiments of the present invention.

[0108] The co-attention branch proposed in the embodiments of the present invention makes full use of the context information of all training data, which helps the classifier associate the labels with the corresponding object regions and understand the relationships between the various object parts, thereby mining the intrinsic correlation information of different morphological objects with the same semantics and helping to improve the detection performance of the model.

[0109] Based on the above same features, in order to train a co-owned object classifier based on co-attention, the known co-class label l m ∩l n (For example, the image-level annotation indicates that the object "horse" is included in both paired images) is used as the supervision information, and the loss function expression for mining the same features created in the co-attention branch is:

[0110]

[0111] In the formula, σ represents the cross-entropy operation. This loss function can train the aforementioned perception convolutional layer, 1×1 convolutional layer, and OICR network feature convolutional layer, and the training process also affects the weight parameters of other parts of the OICR network. If the weight parameters of the aforementioned category perception convolutional layer 1×1 convolutional layer and each item in the OICR network are adjusted better, then the loss function of the co-attention branch has a smaller value.

[0112] Specifically, as Figure 4 shown, in this embodiment, the above step S2 further includes the following steps:

[0113] Step ten: Input the first co-attention feature and the second co-attention feature into the sigmod function for processing respectively.

[0114] Specifically, based on the co-attention features calculated in the above steps seven and eight, input the first co-attention feature and the second co-attention feature into the sigmod function respectively, and convert the size of each element in the feature to the range of 0 to 1.

[0115]

[0116]

[0117] In the formula, and are the first co-attention feature and the second co-attention feature processed by the sigmod function respectively, is the parameter matrix implemented by 1×1 convolution.

[0118] Step Eleven: Calculate the difference between unity 1 and the processed first co-attention feature to obtain the first contrast attention map, and calculate the difference between unity 1 and the processed second co-attention feature to obtain the second contrast attention map.

[0119] Specifically, calculate the difference between unity 1 and and calculate the difference between unity 1 and , which is equivalent to obtaining a contrast attention map opposite to the co-attention within the range of 0 to 1. The formula is as follows: Contrast co-attention and can be obtained through the following formula:

[0120]

[0121]

[0122] Among them, the first contrast attention map focuses on the non-common object region in the first image I m , that is, the category objects unique to the first image. The second contrast attention map focuses on the non-common object region in the second image I n , that is, the category objects unique to the second image.

[0123] Step Twelve: Multiply the first contrast attention map element-wise with the first image feature to obtain the first contrast co-attention feature, and multiply the second contrast attention map element-wise with the second image feature to obtain the second contrast co-attention feature. Among them, the first contrast co-attention feature is used to represent the objects unique to the first image, and the second contrast co-attention feature is used to represent the objects unique to the second image.

[0124] Specifically, similar to the principle in Step Eight, multiply the first contrast attention map with the first image feature F m element-wise to obtain the first contrast co-attention feature and multiply the second contrast attention map with the second image feature F n element-wise to obtain the second contrast co-attention feature For example: If the element at a certain position in the first contrast attention map is close to 1, it means that the object category of the first image feature F m at this position is unique to it, and the value close to 1 is retained through multiplication; if the element at a certain position in the first contrast attention map is close to 0, it means that the object category of the first image feature F mThe object category at this position is not unique to it. Most likely, it is a category shared by the paired images and is deleted by multiplying a value close to 0.

[0125]

[0126]

[0127] Step 13: Process the first contrast common attention feature and the second contrast common attention feature through a category-aware convolutional layer and a 1×1 convolutional layer respectively, and use the processed first contrast common attention feature and the second contrast common attention feature as different features.

[0128] Next, perform the same operations as in the common attention branch, that is, and are input into the category-aware convolutional layer and the 1×1 convolutional layer to generate the category-aware activation maps and Obtain the different features between the two images:

[0129]

[0130]

[0131] Finally, through the GAP operation, obtain the predicted category scores of the different objects and

[0132]

[0133]

[0134] Among them, the unique score vector of the first image represents the probability of the occurrence of the unique category objects in the first image, and the unique score vector of the second image represents the probability of the occurrence of the unique category objects in the second image. Then, the unique score vectors of the first image and the second image are used as the same features mined from the two paired images in the embodiment of the present invention.

[0135] Similarly, in order to train a unique object classifier based on contrast common attention, obtain the unique category labels of the two images: that is, the unique object category l m in the first image I m \l n and the unique object category l n in the second image I n \l m

[0136] lm \l n = l m -l m ∩l n

[0137] l n \l m = l n -l m ∩l n

[0138] Then, use the known unique class label l m \l n and l n \l m (For example, the known image-level labels, the first image contains "person", the second image does not contain "person", the second image contains "car", and the first image does not contain "car") as the supervision information, the loss function expression for mining different features created by the contrastive co-attention branch is:

[0139]

[0140] In the formula, σ represents the cross-entropy operation. This loss function can train the aforementioned perceptual convolutional layer, 1×1 convolutional layer, and OICR network feature convolutional layer, and the training process also affects the weight parameters of other parts of the OICR network. If the weight parameters of the aforementioned class-aware convolutional layer 1×1 convolutional layer, and each item in the OICR network are adjusted better, then the loss function of the contrastive co-attention branch has a smaller value. From the above calculation process, it can be seen that the contrastive co-attention branch filters out objects with the same semantics, forcing the classifier to better focus on different semantic objects in the image, learning the object differences between paired images, so as to mine the intrinsic difference information of different semantic objects.

[0141] Specifically, in this embodiment, the processing process of the above step S4 is as follows:

[0142] Specifically, based on the processing of the above steps one to thirteen, the embodiment of the present invention integrates all the loss functions involved in the above-mentioned bounding box online generation module and co-attention module, so as to train the overall network composed of the bounding box online generation module and co-attention module through the total loss function of the model.

[0143]

[0144]

[0145]

[0146]

[0147]

[0148] In the formula, the total model loss is α and β represent weights, which are used to balance the losses of the bounding box online generation module and the co-attention module.

[0149] represents the loss function of the bounding box online generation module, specifically including those in the above embodiments and is the original loss function of the OICR network created based on the annotation results, and are both prior arts and will not be elaborated in this embodiment.

[0150] represents the loss function of the co-attention module, I m and I n respectively represent the m-th image and the n-th image in the image sample, and I m and I n are a pair of images that include at least one object of the same category, is a loss function created according to the scores of each position belonging to each category in the image features of the m-th and n-th images, is a loss function created by mining the same features, is a loss function created by mining different features, l m and l n respectively represent the image-level labels of the m-th image and the n-th image, s m and s n respectively represent the category score vectors of the m-th image and the n-th image, represents the common score vector of the m-th image containing common category objects calculated based on the same features of the m-th image, represents the common score vector of the n-th image containing common category objects calculated based on the same features of the n-th image, l m ∩ l n represents the common category label, represents the unique score vector of the m-th image containing unique objects calculated based on the different features of the m-th image, represents the unique score vector of the n-th image containing unique objects calculated based on the different features of the n-th image, l m \l n and l n \l mrespectively represent the unique class labels of the m-th image and the n-th image, and σ represents the cross-entropy operation.

[0151] Loss function for the co-attention module For the principle explanation, please refer to the relevant detailed descriptions in the above steps 1 to 13, which will not be elaborated here.

[0152] The overall network composed of the bounding box online generation module and the co-attention module is trained through the total loss function of the model created in this embodiment. If the weight parameters in the bounding box online generation module and the co-attention module are adjusted better, the total loss of the model will be The smaller the value, which is used as the guiding direction for adjusting the weight parameters of the bounding box online generation module and the co-attention module. After the training is completed, the co-attention module is deleted, and the OICR network is output. The obtained OICR network has a stronger ability to recognize different category objects, significantly improving the annotation accuracy of the OICR network for images with more object categories.

[0153] As Figure 6 shown, in one embodiment, an inspection target annotation method specifically includes the following steps:

[0154] Step S101: Obtain the image to be annotated;

[0155] Step S102: Input the image to be annotated into the inspection target annotation model pre-trained in the above training method embodiment, and output the annotation result of the inspection target annotation model for the image to be annotated, making the annotation result more accurate and the error rate lower.

[0156] Through the above steps, the technical solution provided by this application automatically annotates the image to be annotated through a pre-trained inspection target annotation model, and outputs the annotation result of the inspection target annotation model for the image to be annotated. Among them, the training method of the inspection target annotation model provided by the embodiments of the present invention is as follows: First, the paired image samples are respectively input into two OICR networks with shared weight parameters, and the preliminary annotation results of the targets in the two images are respectively output. It should be noted that at least one object of the same category is included in the paired image samples. Then, the first image feature and the second image feature extracted by the feature convolutional layers in the two OICR networks are respectively obtained, and the same features and different features between the first image feature and the second image feature are mined through the common attention module. After that, the total model loss is calculated based on the preliminary annotation results, the same features and different features. Finally, the weight parameters of the total model composed of the OICR network and the common attention module are adjusted based on the total model loss. After multiple rounds of training, the OICR network with adjusted weight parameters is extracted from the total model as the inspection target annotation model. Through the inspection target annotation model trained by the embodiments of the present invention, on the basis of the original OICR network, it can drive the classifier to mine the similarity information of objects with the same semantics and the difference information of objects with different semantics, improve the ability of the network to mine the potential similarity and difference between objects across images, reduce the impact of the variety of objects in the image on category recognition, alleviate the situation of local optimality of the annotation results, and significantly improve the accuracy of automatic annotation of inspection targets through the annotation model provided by the embodiments of the present invention.

[0157] As Figure 7 shown, this embodiment also provides a training device for an inspection target annotation model. The device includes:

[0158] The bounding box online generation module 1 is used to construct a paired image sample data set including at least one object of the same category, separately input the paired image samples into two OICR networks with shared weight parameters, and use the feature convolutional layers in the two OICR networks to extract the first image feature and the second image feature respectively; it is also used to output the preliminary annotation results of the targets in the two image samples through the two OICR networks respectively; for the detailed content, please refer to the relevant descriptions of steps S1-S2 in the above method embodiment, and will not be elaborated here.

[0159] The common attention module 2 is used to mine the same features and different features between the first image feature and the second image feature. For the detailed content, please refer to the relevant description of step S3 in the above method embodiment, and will not be elaborated here.

[0160] The loss calculation module 3 is used to calculate the total loss of the model based on the preliminary annotation results, the same features, and different features. The total loss of the model is the loss of the total model composed of the OICR network and the co-attention module shared by two weight parameters. For the detailed content, refer to the relevant description of step S4 in the above method embodiment, and details will not be repeated here.

[0161] The parameter adjustment module 4 is used to adjust the weight parameters of the total model based on the total loss of the model, and extract the OICR network in the total model as the inspection target annotation model. For the detailed content, refer to the relevant description of step S5 in the above method embodiment, and details will not be repeated here.

[0162] The inspection target annotation model training device provided by the embodiment of the present invention is used to execute the inspection target annotation model training method provided by the above embodiment. Its implementation manner and principle are the same. For the detailed content, refer to the relevant description of the above method embodiment, and details will not be repeated.

[0163] As Figure 8 shown, this embodiment also provides an inspection target annotation device, and the device includes:

[0164] The data acquisition module 101 is used to acquire the image to be annotated. For the detailed content, refer to the relevant description of step S101 in the above method embodiment, and details will not be repeated here.

[0165] The target annotation module 102 is used to input the image to be annotated into the inspection target annotation model pre-trained by the above training device, and output the annotation result of the inspection target annotation model for the image to be annotated. For the detailed content, refer to the relevant description of step S102 in the above method embodiment, and details will not be repeated here.

[0166] The inspection target annotation device provided by the embodiment of the present invention is used to execute the inspection target annotation method provided by the above embodiment. Its implementation manner and principle are the same. For the detailed content, refer to the relevant description of the above method embodiment, and details will not be repeated.

[0167] Through the collaborative cooperation of the above-mentioned various components, the technical solution provided by this application automatically annotates the image to be annotated through a pre-trained inspection target annotation model, and outputs the annotation result of the inspection target annotation model for the image to be annotated. Among them, the training method of the inspection target annotation model provided by the embodiments of the present invention is as follows: First, a pair of image samples are respectively input into two OICR networks with shared weight parameters, and the preliminary annotation results of the targets in the two images are respectively output. It should be noted that at least one object of the same category is included in the pair of image samples. Then, the first image feature and the second image feature extracted by the feature convolutional layer in the two OICR networks are respectively obtained, and the common attention module is used to mine the same features and different features between the first image feature and the second image feature. After that, the total model loss is calculated based on the preliminary annotation result, the same feature and the different feature. Finally, the weight parameters of the total model composed of the OICR network and the common attention module are adjusted based on the model total loss. After multiple rounds of training, the OICR network with adjusted weight parameters is extracted from the total model as the inspection target annotation model. Through the inspection target annotation model trained by the embodiments of the present invention, on the basis of the original OICR network, it can drive the classifier to mine the similarity information of objects with the same semantics and the difference information of objects with different semantics, improve the ability of the network to mine the potential similarity and difference between objects across images, reduce the impact of the variety of objects in the image on category recognition, alleviate the situation of local optimality of the annotation result, and significantly improve the accuracy of automatic annotation of inspection targets through the annotation model provided by the embodiments of the present invention.

[0168] Figure 9 An electronic device according to an embodiment of the present invention is shown. The device includes a processor 901 and a memory 902, and can be connected through a bus or other means. Figure 9 Taking the connection through the bus as an example.

[0169] The processor 901 can be a central processing unit (CPU). The processor 901 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.

[0170] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 902, that is, implements the methods in the above method embodiments.

[0171] The memory 902 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor 901, etc. In addition, the memory 902 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely disposed relative to the processor 901, and these remote memories can be connected to the processor 901 through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0172] One or more modules are stored in the memory 902 and, when executed by the processor 901, implement the methods in the above method embodiments.

[0173] For the specific details of the above electronic device, reference can be made to the corresponding relevant descriptions and effects in the above method embodiments for understanding, and details are not described herein again.

[0174] Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The implemented program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0175] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for training an inspection target annotation model, characterized in that, the method includes: Construct a paired image sample data set including at least one same category of objects, and input the paired image samples into two OICR networks with shared weight parameters respectively. Use the feature convolutional layers in the two OICR networks to extract the first image feature and the second image feature respectively; Use a preset common attention module to mine the same features and different features between the first image feature and the second image feature; Output the preliminary annotation results of the targets in the two image samples through the two OICR networks respectively; Calculate the total model loss based on the preliminary annotation results, the same features and the different features. The total model loss is the loss of the total model composed of two OICR networks with shared weight parameters and the common attention module; Adjust the weight parameters of the total model based on the total model loss, and extract the OICR network from the total model as the inspection target annotation model; Using a preset common attention module to mine the same features in the first image feature and the second image feature includes: classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature belonging to each category and the scores of each position in the second image feature belonging to each category; calculating the similarity index between all positions of the first image feature and all positions of the second image feature pairwise based on the scores of each position in the first image feature and the second image feature belonging to each category; extracting the same features from the first image feature and the second image feature respectively according to the similarity index; The classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature and the second image feature belonging to each category includes: sequentially inputting the current image feature into a category-aware convolutional layer and a 1×1 convolutional layer to generate a category-aware activation map, where the current image feature is the first image feature or the second image feature; performing global average pooling on the category-aware activation map to obtain the category score vector of the current image feature, and the category score vector is used to represent the scores of each position in the current image feature belonging to each category; The calculating the similarity index between all positions of the first image feature and all positions of the second image feature pairwise based on the scores of each position in the first image feature and the second image feature belonging to each category includes: flattening the first image feature and the second image feature into two-dimensional matrices respectively, and corresponding the scores of each position in the first image feature and the second image feature belonging to each category to their respective flattened two-dimensional matrices; calculating a similarity matrix using the two-dimensional matrix of the first image feature and the two-dimensional matrix of the second image feature, and each element in the similarity matrix is used to represent the similarity index between all positions of the first image feature and all positions of the second image feature pairwise; Extracting the same features from the first image feature and the second image feature respectively according to the similarity index includes: performing column normalization and row normalization on the similarity matrix respectively to obtain a first attention map matching the dimension of the flattened first image feature, and obtaining a second attention map matching the dimension of the flattened second image feature; multiplying the first attention map and the flattened first image feature element by element to obtain a first common attention feature, where the first common attention feature is used to represent the object corresponding to the common class label of the paired images on the first image; multiplying the second attention map and the flattened second image feature element by element to obtain a second common attention feature, where the second common attention feature is used to represent the object corresponding to the common class label of the paired images on the second image; processing the first common attention feature and the second common attention feature respectively through a class-aware convolutional layer and a 1×1 convolutional layer, and using the processed first common attention feature and second common attention feature as the same features; Mining different features in the first image feature and the second image feature by using a common attention module includes: respectively inputting the first common attention feature and the second common attention feature into a sigmod function for processing; calculating the difference between 1 and the processed first common attention feature to obtain a first contrast attention map, and calculating the difference between 1 and the processed second common attention feature to obtain a second contrast attention map; multiplying the first contrast attention map and the first image feature element by element to obtain a first contrast common attention feature, and multiplying the second contrast attention map and the second image feature element by element to obtain a second contrast common attention feature, where the first contrast common attention feature is used to represent the unique object on the first image, and the second contrast common attention feature is used to represent the unique object on the second image; processing the first contrast common attention feature and the second contrast common attention feature respectively through a class-aware convolutional layer and a 1×1 convolutional layer, and using the processed first contrast common attention feature and second contrast common attention feature as the different features; Calculating the total loss of the model based on the preliminary annotation result, the same features and the different features includes: Calculate the total loss of the model according to the following loss function relationship Wherein, α and β represent weights, represents the loss function corresponding to the OICR network shared by two weight parameters, created based on the preliminary annotation results, represents the loss function of the co-attention module, I m and I n respectively represent the m-th image and the n-th image in the image sample, and I m and I n are a pair of images including at least one same-category object, is the loss function created according to the scores of each position in the image features of the m-th and n-th images belonging to each category, is the loss function created by mining the same features, is the loss function created by mining the different features, l m and l n respectively represent the image-level labels of the m-th image and the n-th image, s m and s n respectively represent the category score vectors of the m-th image and the n-th image, represents the common score vector of the m-th image containing common-category objects calculated based on the same features of the m-th image, represents the common score vector of the n-th image containing common-category objects calculated based on the same features of the n-th image, l m ∩l n represents the common-category label, represents the unique score vector of the m-th image containing unique objects calculated based on the different features of the m-th image, represents the unique score vector of the n-th image containing unique objects calculated based on the different features of the n-th image, l m \l n and l n \l m respectively represent the unique-category labels of the m-th image and the n-th image, and σ represents the cross-entropy operation.

2. An inspection target annotation method, characterized in that, the method includes: Obtaining an image to be annotated; Inputting the image to be annotated into an inspection target annotation model pre-trained by the method according to claim 1, and outputting the annotation result of the inspection target annotation model for the image to be annotated.

3. An inspection target annotation model training device, characterized in that, the device includes: A bounding box online generation module is used to construct a paired image sample data set including at least one type of the same category of objects, and separately input the paired image samples into two OICR networks with shared weight parameters, and use the feature convolutional layers in the two OICR networks to extract the first image feature and the second image feature respectively; it is also used to output the preliminary annotation results of the targets in the two image samples through the two OICR networks respectively. The co-attention module is used to mine the same features and different features between the first image feature and the second image feature; mining the same features in the first image feature and the second image feature by using the preset co-attention module includes: classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature belonging to each category and the scores of each position in the second image feature belonging to each category; calculating the similarity index between every two positions of all positions of the first image feature and all positions of the second image feature based on the scores of each position in the first image feature and the second image feature belonging to each category; extracting the same features from the first image feature and the second image feature respectively according to the similarity index; the classifying the first image feature and the second image feature respectively to obtain the scores of each position in the first image feature and the second image feature belonging to each category includes: sequentially inputting the current image feature into the category-aware convolutional layer and the 1×1 convolutional layer to generate a category-aware activation map, where the current image feature is the first image feature or the second image feature; performing global average pooling on the category-aware activation map to obtain the category score vector of the current image feature, and the category score vector is used to represent the scores of each position in the current image feature belonging to each category; the calculating the similarity index between every two positions of all positions of the first image feature and all positions of the second image feature based on the scores of each position in the first image feature and the second image feature belonging to each category includes: flattening the first image feature and the second image feature into two-dimensional matrices respectively, and corresponding the scores of each position in the first image feature and the second image feature belonging to each category to their respective flattened two-dimensional matrices; calculating a similarity matrix by using the two-dimensional matrix of the first image feature and the two-dimensional matrix of the second image feature, and each element in the similarity matrix is used to represent the similarity index between every two positions of all positions of the first image feature and all positions of the second image feature; the extracting the same features from the first image feature and the second image feature respectively according to the similarity index includes: performing column normalization and row normalization on the similarity matrix respectively to obtain a first attention map matching the dimension of the flattened first image feature and a second attention map matching the dimension of the flattened second image feature; multiplying the first attention map and the flattened first image feature element by element to obtain a first co-attention feature, and the first co-attention feature is used to represent the object corresponding to the common category label of the paired images on the first image; multiplying the second attention map and the flattened second image feature element by element to obtain a second co-attention feature, and the second co-attention feature is used to represent the object corresponding to the common category label of the paired images on the second image;Process the first co-attention feature and the second co-attention feature through a class-aware convolutional layer and a 1×1 convolutional layer respectively, and use the processed first co-attention feature and second co-attention feature as the same feature; use the co-attention module to mine different features in the first image feature and the second image feature, including: respectively input the first co-attention feature and the second co-attention feature into the sigmod function for processing; calculate the difference between 1 and the processed first co-attention feature to obtain a first contrast attention map, and calculate the difference between 1 and the processed second co-attention feature to obtain a second contrast attention map; multiply the first contrast attention map and the first image feature element by element to obtain a first contrast co-attention feature, and multiply the second contrast attention map and the second image feature element by element to obtain a second contrast co-attention feature, where the first contrast co-attention feature is used to represent the objects unique to the first image, and the second contrast co-attention feature is used to represent the objects unique to the second image; process the first contrast co-attention feature and the second contrast co-attention feature through a class-aware convolutional layer and a 1×1 convolutional layer respectively, and use the processed first contrast co-attention feature and second contrast co-attention feature as the different features; A loss calculation module, configured to calculate a total model loss based on the preliminary annotation result, the same features, and the different features, where the total model loss is the loss of a total model composed of an OICR network with two shared weight parameters and the co-attention module; calculate the total model loss according to the following loss function relationship Wherein, α and β represent weights, represents the loss function corresponding to the OICR network in which two weight parameters are shared, and is created based on the preliminary annotation result, represents the loss function of the co-attention module, I m and I n respectively represent the m-th image and the n-th image in the image sample, and I m and I n are a pair of images including at least one same-class object, is the loss function created according to the scores of each position belonging to each category in the image features of the m-th and n-th images, is the loss function created by mining the same features, is the loss function created by mining the different features, l m and l n respectively represent the image-level labels of the m-th image and the n-th image, s m and s n respectively represent the category score vectors of the m-th image and the n-th image, represents the common score vector of the m-th image containing common-class objects calculated based on the same features of the m-th image, represents the common score vector of the n-th image containing common-class objects calculated based on the same features of the n-th image, l m ∩l n represents the common-class label, represents the unique score vector of the m-th image containing unique objects calculated based on the different features of the m-th image, represents the unique score vector of the n-th image containing unique objects calculated based on the different features of the n-th image, l m \l n and l n \l m respectively represent the unique-class labels of the m-th image and the n-th image, and σ represents the cross-entropy operation; A parameter adjustment module is used to adjust the weight parameters of the total model based on the total loss of the model, and extract the OICR network in the total model as the inspection target annotation model.

4. An inspection target annotation device Characterized in that The device includes: A data acquisition module is used to acquire the image to be annotated. A target annotation module is used to input the image to be annotated into the inspection target annotation model pre-trained by the method according to claim 1, and output the annotation result of the inspection target annotation model for the image to be annotated.

5. An electronic device Characterized in that It includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1-2.

6. A computer-readable storage medium Characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1-2.