A self-supervised pre-training object detection method, system, device and storage medium

By pasting patch blocks of the foreground map in the background image and performing multi-scale feature encoding, combined with Transformer's self-attention module and data enhancement technology, the insufficient positioning and classification in the self-supervised pre-training agent task is solved, and the performance of target detection is improved.

CN116012658BActive Publication Date: 2025-07-11XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310112547.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-07-11
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

The existing self-supervised pre-training agent tasks have shortcomings in object detection positioning and classification, especially the accuracy of traditional algorithms limits the model's positioning and classification capabilities.

Method used

Proposals are extracted using Selective-search algorithm, and the pre-training process is optimized by pasting patch blocks of the foreground map in the background image and multi-scale feature encoding, combining Transformer's self-attention module and data enhancement technology.

Benefits of technology

The positioning and classification capabilities of object detection are improved, the robustness of feature learning is enhanced, and the positioning and classification performance in the pre-training process is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012658B_ABST
    Figure CN116012658B_ABST
Patent Text Reader

Abstract

The present invention discloses a self-supervised pre-training object detection method, system, device and storage medium, which includes extracting proposals from a given input picture, selecting the top 30 proposals as patch blocks to be pasted; pasting the obtained patch blocks into the selected background picture to obtain a synthetic picture, providing accurate position annotations for pre-training, extracting the color RGB values of the downstream objects to be detected, randomly selecting an area in the pasted patch block and changing its color to the color corresponding to the extracted color RGB values, optimizing the classification ability in pre-training object detection, respectively extracting the features of the synthetic picture and the multi-scale features of the pasted patch block, and encoding the multi-scale features of the patch block into object queries; the object queries are learned based on the extracted features of the synthetic picture, and the learned object queries are used to predict the categories and bounding boxes, obtaining a set of predictions, and matching the set of predictions with the set of true annotations. The present invention solves the problem of the deficiencies of pre-training proxy tasks in object detection positioning and classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a self-supervised pre-training object detection method, system, device and storage medium. Background Art

[0002] Object detection is an important task in computer vision, including the localization and classification of objects in images, and is a more complex task than image classification. Object detection technology has extensive applications in industries such as security, transportation, smart agriculture, and healthcare. Therefore, the research on deep learning object detection technology is very meaningful.

[0003] Generally, in order to achieve good detection effects, a large amount of labeled data is required to train the model. However, the cost of labeled data is very high and the accuracy is also affected by labeling errors. Unsupervised tasks use data that does not require labeling, but do not fully utilize the features of the data. Recently, people have focused on self-supervised learning. Self-supervised learning refers to using the data itself to provide weak supervision according to the designed proxy tasks, without the need to label the data for pre-training. Usually, pre-training is carried out on large datasets such as ImageNet or COCO to obtain valuable feature representations for downstream tasks. In the field of image processing, common proxy tasks include jigsaw puzzles, cutout (randomly removing a certain part of the image and using the remaining part to predict the removed part), color completion (using the grayscale image of the image as input and predicting the real color image), etc. For the object detection task, constructing a suitable proxy task enables the model to acquire certain localization and classification capabilities during the pre-training stage, enhancing its representation performance in downstream tasks.

[0004] Regarding the design and implementation of proxy tasks, most previous work focused on the backbone network part of the pre-trained model. The recent UP-DETR (also a pre-training method based on DETR) and DETReg (also a pre-training method constructed based on DETR) are proxy tasks designed based on the DETR model architecture, achieving end-to-end pre-training of the entire model. The proxy task of UP-DETR is to let the model learn to detect random patch blocks in the image during the pre-training stage, but the random patch blocks do not represent the actual objects appearing in the image. The proxy task of DETReg uses the traditional Selective-search algorithm to extract proposals as objects to be detected, promoting the model to recognize objects. However, due to the accuracy limitations of traditional algorithms, the localization ability and classification ability of DETReg are restricted. Summary of the Invention

[0005] The object of the present invention is to solve the problem of the deficiencies of self-supervised pre-training proxy tasks in object detection localization and classification in the prior art, and to provide a self-supervised pre-training object detection method, system, device and storage medium.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A self-supervised pre-training object detection method includes the following steps:

[0008] S1: Given an input image, extract proposals from the given input image, and select the top 30 proposals as patch blocks to be pasted.

[0009] S2: Select an image from the dataset as the background image, paste the patch blocks obtained in S1 into the background image to obtain a synthesized image, extract the RGB values of the colors of the downstream objects to be detected, and randomly select an area in the pasted patch blocks and change its color to the color corresponding to the extracted RGB values.

[0010] S3: Extract the features of the synthesized image and the multi-scale features of the patch blocks pasted in the synthesized image respectively, and encode the multi-scale features of the patch blocks into object queries.

[0011] S4: The object queries are learned based on the features of the extracted synthesized image, and the learned object queries are used to predict the categories and bounding boxes to obtain a set of predictions, and the set of predictions is matched with the set of true annotations.

[0012] A further improvement of the present invention lies in:

[0013] The step S1 includes the following steps:

[0014] Extract proposals through the Selective-search algorithm. Specifically, segment the given input image to obtain a series of regions, calculate the similarity of different regions according to the set loss function for merging, and select the top 30 proposals according to the similarity from high to low.

[0015] The step S2 further includes the following steps:

[0016] Horizontally flip the obtained patch blocks, change the brightness, change the contrast, change the saturation, and change the hue.

[0017] The step S3 includes the following steps:

[0018] Extract the features of the synthetic image and the multi-scale features of the patch blocks pasted in the synthetic image through the ResNet50 backbone network.

[0019] In the step S3, the process of encoding the pasted patch blocks into object queries includes:

[0020] Select several patch blocks from each synthetic image for encoding

[0021] Extract the features of each patch block through the backbone network to obtain several patch feature maps of different scales;

[0022] Perform pooling processing on the obtained patch feature maps;

[0023] Define corresponding linear layers based on the obtained patch feature maps of different scales to obtain object queries obtained by converting patch features of different scales;

[0024] According to the different size ratios of the multi-scale input features extracted from the synthetic image, repeat the operation on the object queries obtained by encoding each patch at different scale features.

[0025] In the step S4, the object query and the features of the synthetic image are input into the Transformer for learning, and an attention mask designed for the object query is introduced in the self-attention module of the decoder of the Transformer.

[0026] In the step S4, the process of predicting the category and bounding box of the learned object query includes:

[0027] Perform prediction through the prediction head, and the prediction head includes f box 、f cat and f rec ;

[0028] f box is used to predict the bounding boxes, f cat predicts whether it is a pasted patch, f rec is used to reconstruct the object descriptors, where then v1……v k represents the object query related to the image calculated by the Transformer.

[0029] A self-supervised pre-training object detection system, comprising a patch block extraction module, a composite graph construction module, a multi-scale feature transformation module, and a matching module;

[0030] The patch block extraction module is used to, given an input image, extract proposals from the given input image and select the first 30 proposals as patch blocks to be pasted;

[0031] The composite graph construction module is used to select an image from a dataset as the background image, paste the patch blocks obtained in S1 into the background image to obtain a composite graph, extract the RGB values of the colors of the downstream targets to be detected, and randomly select an area in the pasted patch blocks and change its color to the color corresponding to the extracted RGB values;

[0032] The multi-scale feature transformation module is used to respectively extract the features of the composite graph and the multi-scale features of the patch blocks pasted in the composite graph, and encode the multi-scale features of the patch blocks into object queries;

[0033] The matching module is used to learn the object queries based on the extracted composite graph features, predict the categories and bounding boxes of the learned object queries to obtain a set of predictions, and match the set of predictions with the set of true annotations.

[0034] A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of any method of the present invention are implemented.

[0035] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any method of the present invention are implemented.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] The present invention discloses a self-supervised pre-training object detection method. First, the pictures are synthesized. Specifically, more characteristic patch blocks are selected from the given input pictures, and the patch blocks are pasted and synthesized with the selected background pictures to form synthesized pictures, providing accurate position annotations for pre-training, optimizing the positioning problem in the pre-training object detection process, improving the positioning ability. At the same time, based on the RGB values of the colors of the downstream detection targets, the pixel values of some of the pasted patch blocks are changed to add downstream color noise, optimizing the classification ability in the pre-training object detection. The multi-scale features of the patch blocks are transformed into object queries for learning, so that there are different numbers of object queries to learn for features of different scales, and the features of different scales can be more fully mined to improve the detection performance. And the bounding box prediction is performed on the learned object queries to ensure the feature consistency between the predicted patch and the real patch.

[0038] Furthermore, in the present invention, the patch blocks obtained by extracting proposals through the Selective-search algorithm can learn better feature expressions for downstream tasks, making feature learning more valuable.

[0039] Furthermore, in the present invention, the obtained patch blocks are adjusted in terms of brightness, contrast, saturation, and hue to enhance the robustness of feature learning and better cooperate with the object detection tasks of downstream datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is the overall flowchart of the pre-training stage and the fine-tuning training stage of the present invention;

[0042] Figure 2 is the overall framework diagram of the CP-DETR self-supervised pre-training object detection model of the present invention;

[0043] Figure 3 is the multi-scale feature structure diagram based on the input of the Deformable-DETR architecture;

[0044] Figure 4 is the schematic diagram of the ratio expression of the input features and object queries of the present invention;

[0045] Figure 5 This is a schematic diagram of the attention mask added to the decoder in the present invention. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0047] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0048] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0049] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0050] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0051] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly defined and limited, if the terms "set", "install", "connected", "connected" appear, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0052] The following further describes the present invention in detail with reference to the drawings:

[0053] See Figure 1 , the embodiments of the present invention disclose a self-supervised pre-training object detection method. This method is a self-supervised pre-training object detection method based on DETR. This method includes a pre-training stage and a fine-tuning training stage. The pre-training stage includes a self-supervised prior task, a feature extraction module, a Transformer module, a prediction head module, and a Hungarian matching module.

[0054] The pre-training stage is a training process based on a self-supervised prior task on an unlabeled large dataset, and its purpose is to enable the model to initially have certain detection capabilities;

[0055] The fine-tuning stage is a supervised fine-tuning training on a relatively small dataset based on the model weights saved in the pre-training stage, referring to downstream tasks including object detection.

[0056] The self-supervised prior task based on DETR is a task of pasting the patches cropped from the foreground image using a cropping strategy into the background image to construct a synthetic image and letting the model learn to detect the pasted patches in the synthetic image. The purpose is to initially train the detection capabilities of the model;

[0057] The feature extraction module is to use the ResNet50 backbone network to extract the features of the input synthetic image for subsequent calculations;

[0058] The Transformer calculation module includes an encoder Encoder and a decoder Decoder. In the encoding stage, the input image features are encoded, and in the decoding stage, object query (a set of learnable position encodings in the Transformer) is used to learn in the encoded features;

[0059] The prediction head module is to use a multi-layer perceptron MLP to predict the category and bounding box of the object query after being calculated by the Transformer;

[0060] The Hungarian matching module matches the predicted set with the given set of labels and assigns them according to the principle of the minimum cost, and the cost rule is determined by a set of loss functions.

[0061] Specifically, it includes the following steps:

[0062] Pre-training stage:

[0063] Step 1: For a given dataset, construct a synthetic graph based on cut and paste

[0064] To solve the object detection localization problem:

[0065] To provide more accurate weak supervision of the bounding box for self-supervised pre-training, a synthetic graph method is used to construct the input image. The overall architecture is as Figure 2 , by selecting some regions from the current foreground image for cropping and then pasting them into the background image to construct a synthetic graph.

[0066] To achieve simplicity and not introduce additional datasets, in the specific implementation, the background image is a randomly selected image from the current dataset, and the pasting is randomly selected at a certain position within the appropriate width and height range according to the current patch size and the size of the background image.

[0067] To solve the object detection classification problem:

[0068] For the region to be cropped selected from the foreground image, inspired by DETReg, the present invention adopts the Selective-search algorithm to extract proposals, which is an unsupervised extraction method.

[0069] A series of regions are obtained through image segmentation, and then the similarity of different regions is calculated according to the set loss function for merging. Finally, the selected proposals are also obtained in descending order of similarity.

[0070] Furthermore, the present invention proves through experiments that using Selective-search has better effects than random cropping. For the localization task, it is not important which patch to detect. Any patch in the synthetic graph is regarded as the "object" to be detected. However, if the patch selected is the background region, good feature expressions are not learned for the downstream tasks. Therefore, using Selective-search to select more meaningful patches makes the features more valuable for learning.

[0071] Furthermore, in the embodiments of the present invention, in order to enhance the robustness of feature learning, for the pasted patches, the present invention has some data augmentation schemes:

[0072] Including horizontal flipping and using the colorjitter method provided in pytorch to change brightness, contrast, saturation, and hue.

[0073] Furthermore, in the embodiments of the present invention, in order to better adapt to the object detection task of the downstream dataset, sample the color pixel values of the object to be detected, and randomly select some regions in the patch to modify their pixel values. That is, in the pre-training stage, let the model learn to detect the patch (i.e., object) with these pixel values, which helps the detection of the downstream task.

[0074] Step 2: Use the resnet50 backbone network to extract features from the input image

[0075] The ResNet50 adopted by the present invention has a total of 50 layers and can be divided into five parts according to the feature scale. The convolutional kernel sizes and numbers used in each part are different, and feature maps with different scales and numbers of channels are output respectively to obtain C2, C3, C4, and C5. The existing DETR uses the last layer C5 output for encoding and decoding. Since the feature scale used is small, the detection accuracy of small targets in the final result is not good. The CP-DETR in the embodiments of the present invention is based on Deformable-DETR, which adopts multi-scale features on the basis of DETR. By using 1×1 convolution on C3, C4, and C5 and additionally using 3×3 convolution on C5, 4 scales of features with 256 dimensions in channel number are obtained. See Figure 3 。

[0076] Step 3: Construct a Transformer object query adapted to multi-scale input features and the corresponding attention mask design based on the pasted patch

[0077] After the Transformer structure flattens and inputs the image features, it is encoded by the Encoder of the encoder and then input into the Decoder decoder.

[0078] An important component of the decoder is the object query, which is used as the query of the cross-attention module of the decoder to search in the feature memory output by the encoder. The memory is interpreted as the intermediate result or intermediate feature calculated by the Transformer's Encoder encoder for the synthesized graph features.

[0079] Furthermore, in subsequent processing, after the decoder outputs, the prediction of the category and the bounding box is also based on the object query.

[0080] Furthermore, DETReg also adapts self-supervised prior tasks based on Deformable-DETR. During the pre-training and fine-tuning phases, the object queries are learnable encodings of size [300, batch_size, 256], implemented through nn.Embedding() in PyTorch. UP-DETR is adapted based on DETR, that is, the input feature is a single-scale feature, and it converts the feature of the last scale extracted by the backbone network from the patch into an object query.

[0081] The present invention combines multi-scale features and multi-scale object queries adapted for multi-scale features, and the specific implementation is as follows:

[0082] For the multi-scale input features extracted by the backbone network from the synthetic graph, before encoding into the Transformer, a flattening operation is performed, that is, the features of each scale are flattened into [x, 256], where x represents the width × height of the feature map. Then, the flattened feature maps of all scales are concatenated in the first dimension, and the size ratios of the feature maps of different scales are from large to small as In the embodiment of the present invention, the different-scale feature maps obtained in this step are actually 4 scales. However, since the ratio of the last scale is small and its feature map and the feature map of the previous scale are obtained from the same layer of the backbone network feature map, the sum of their ratios is approximately calculated according to the ratio of the previous scale Calculated.

[0083] Furthermore, the design of the object query also uses this ratio, that is, more queries are used to focus on large feature maps, and fewer queries are used to focus on small feature maps. See Figure 4 。

[0084] Furthermore, the process of converting the multi-scale features of the patch blocks into object queries is as follows:

[0085] For each synthetic image, 10 patch blocks are extracted to generate object queries.

[0086] Furthermore, the extracted patches are passed through the backbone network to extract 3-scale patch feature maps. The obtained patch feature maps are pooled through the AdaptiveAvgPool2d(1, 1) average pooling layer. For these 3-scale patch feature maps, 3 patch2query branches are defined, that is, 3 linear layers, with the input dimension being the dimension of different scales and the output dimension being 256 dimensions.

[0087] Further, the patch feature maps of three scales are used to obtain three groups of object queries obtained by converting the patch features of different scales through three patch2query branches.

[0088] Further, for each predicted object query, it is repeated a certain number of times through the repeat_interleave operation of pytorch, that is, multiple object queries are responsible for one patch of this scale.

[0089] In the embodiment of the present invention, 10 patches are used, each patch has features of three scales, and each scale feature obtains an object query through the patch2query branch. Now they are all repeated, that is, multiple object queries correspond to the features of one scale of one patch.

[0090] Further, based on the splicing ratio of the synthetic graph feature map, the repetition times of the object queries constructed by the patch features of different scales also follow the ratio. For the synthetic graph feature map of the last scale, since its feature scale is the smallest, each object query is repeated 3 times. In the embodiment of the present invention, the number of repetitions can be set according to the situation; among them, the object queries corresponding to the penultimate feature map are repeated 3*4, and the object queries corresponding to the largest feature map are repeated 3*4*4 times. Then, according to the splicing order of the input features, the object queries constructed by the features of different scales are spliced in the first dimension.

[0091] Since the present invention adopts a one-to-many strategy, that is, for each object query obtained by the feature mapping of each scale, a repetition operation needs to be performed, and multiple object queries are responsible for one patch feature of this scale. Therefore, for two levels of consideration, one is that there is no need for interaction between the object queries responsible for the features of different scales, and the other is that there is no need for interaction between the object queries responsible for different patches of the features of the same scale. Further, an attention mask is introduced into the Self-attention module of the decoder of the Transformer:

[0092] The attention mask is a matrix with the same length and width as the number of object queries. Based on the 10 patch blocks extracted previously, for the smallest-scale features, the number of queries is 10 * 3, and the number of queries for the first two layers is 10 * 3 * 4 and 10 * 3 * 4 * 4 respectively. So the total number of queries is 630. For a matrix of shape [630, 630], masks are set for features of different scales. The first 480 rows (the number of object queries corresponding to the largest scale, 10 * 3 * 4 * 4) and columns are the masks for the largest scale, and every 48 rows and columns (a total of 480 rows and columns generated by 10 patches, each is 48 rows and columns) form a group; the other scales are designed with reference to this, with 12 rows and columns and 3 rows and columns as a group respectively. See Figure 5 , since the matrix is too large, the schematic diagram disclosed in the embodiments of the present invention is only a simple example with a repetition number of 1 and only one patch.

[0093] Step 4: Perform class and bounding box predictions on the output of the Transformer, and use the Hungarian bipartite matching algorithm to perform the optimal allocation calculation for the set. The specific work process is as follows:

[0094] In the pre-training stage, the class prediction is a binary classification, that is, whether it is a pasted patch block; in the fine-tuning stage, the class prediction is the number of classes in the fine-tuning dataset;

[0095] Formally describe the pre-training process of CP-DETR. Assume that we paste M patch blocks, then for the synthetic graph, M bounding boxes b i and object descriptors z i are generated, i ∈ {1, …, M}, which are the encodings extracted by Swav from the pasted patches. Define y i =(b i , z i ). After CP-DETR is pre-trained, its N outputs are aligned with y, where N corresponds to the number of object queries.

[0096] Furthermore, the process of performing class and bounding box predictions on the learned object queries includes:

[0097] CP-DETR has three prediction heads, f box responsible for predicting bounding boxes, f cat predicting whether it is a pasted patch, and f rec responsible for reconstructing object descriptors. Then

[0098] where, v1……v k represents an object query related to the image calculated by the Transformer, and the object query is the output of the last layer decoder of the Transformer.

[0099] In DETR, N is larger than M, so y is expanded into N tuples, and a label c is assigned to each box in y i ∈{0, 1} indicates whether it is a pasted patch or an expanded proposal.

[0100] Furthermore, y and are matched through Hungarian matching, that is, the permutation σ that minimizes the matching cost between y and is found:

[0101]

[0102] In the formula, L match is the cost matrix of bipartite matching, and ∑N is the set of all permutations of {1…n};

[0103] The loss is defined as:

[0104]

[0105] λ f takes the value of 2 in the experiment, and λ b and λ r both take the value of 1 in the experiment;

[0106] L box is based on the L1 loss and the GIOU loss, as shown in formula (3):

[0107]

[0108] λ iou takes the value of 2 in the experiment, takes the value of 5 in the experiment, and L iou is as shown in formula (4):

[0109]

[0110] In formula (4),. represents the area, and the result box itself is represented by the intersection and union of the predicted box coordinates and the ground truth box coordinates. The area of the intersection or union is obtained through b σ(i) and calculated by the minimum / maximum value of a linear function, so that the loss performs well for stochastic gradients, refers to including b σ(i) and the maximum bounding box of, and the area of B is also calculated by the minimum / maximum value of a linear function of its boundary coordinates.

[0111] L class Adopts Cross Entropy Loss, and its definition is shown in formula (5):

[0112]

[0113] L rec Uses L1 loss, and its definition is shown in formula (6):

[0114]

[0115] In the embodiments of the present invention, proposals can be interpreted as candidate clusters or candidate boxes extracted by an algorithm; object query can be interpreted as a set of learnable position encodings in a Transformer; CP-DETR can be interpreted as a proxy task based on DETR, which is the abbreviation of crop-paste DETR.

[0116] The fine-tuning stage includes

[0117] Step 1, loading the weights trained in the pre-training stage to initialize the entire model;

[0118] Step 2, modifying the class prediction branch of the detection head part and removing the branches added for pre-training;

[0119] Step 3, performing fine-tuning training on a supervised dataset.

[0120] The embodiments of the present invention also disclose a self-supervised pre-training object detection system, including a patch block extraction module, a synthetic graph composition module, a multi-scale feature transformation module, and a matching module;

[0121] The patch block extraction module is used to, given an input picture, extract proposals from the given input picture and select the first 30 proposals as the patch blocks to be pasted;

[0122] The synthetic graph composition module is used to select a picture from the dataset as the background picture, paste the patch blocks obtained in S1 into the background picture to obtain a synthetic graph, extract the color RGB values of the downstream objects to be detected, and randomly select an area in the pasted patch block and change its color to the corresponding color;

[0123] A multi-scale feature transformation module is used to extract the features of the synthesized image and the multi-scale features of the pasted patch blocks respectively, and encode the multi-scale features of the patch blocks into object queries;

[0124] A matching module is used for the object queries to learn based on the extracted features of the synthesized image, and perform category and bounding box predictions on the learned object queries to obtain a set of predictions, and match the set of predictions with the set of true annotations.

[0125] A schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned various method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned various device embodiments are implemented.

[0126] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.

[0127] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0128] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0129] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the terminal device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory.

[0130] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0131] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A self-supervised pre-training object detection method, characterized in that It includes the following steps: S1: Given an input image, extract proposals from the given input image, and select the top 30 proposals as the patch blocks to be pasted; S2: Select an image from the dataset as the background image, paste the patch blocks obtained in S1 into the background image to obtain a synthesized image, extract the RGB values of the colors of the downstream targets to be detected, and randomly select an area in the pasted patch block and change its color to the color corresponding to the extracted RGB values; S3: Extract the features of the synthesized image and the multi-scale features of the patch blocks pasted in the synthesized image respectively, and encode the multi-scale features of the patch blocks into object queries; S4: The object queries are learned based on the features of the extracted synthesized image, and the learned object queries are used to predict the categories and bounding boxes to obtain a set of predictions, and the set of predictions is matched with the set of true annotations; In S4, the object queries and the features of the synthesized image are input into the Transformer for learning, and an attention mask designed for the object queries is introduced into the self-attention module in the decoder of the Transformer.

2. The self-supervised pre-training object detection method according to claim 1, wherein Step S1 includes the following steps: Extract proposals through the Selective-search algorithm. Specifically, segment the given input image to obtain a series of regions, calculate the similarity of different regions according to the set loss function for merging, and select the top 30 proposals according to the similarity from high to low.

3. The self-supervised pre-training object detection method according to claim 1, characterized in that, Step S2 also includes the following steps: Horizontally flip, change the brightness, change the contrast, change the saturation, and change the hue of the obtained patch blocks.

4. A self-supervised pre-training object detection method according to claim 1, characterized in that Step S3 includes the following steps: Extract the features of the synthesized image and the multi-scale features of the patch blocks pasted in the synthesized image through the resnet50 backbone network.

5. A self-supervised pre-training object detection method according to claim 1, characterized in that In step S3, the process of encoding the pasted patch blocks into object queries includes: Select several patch blocks from each synthesized image for encoding Extract features of each patch block through the backbone network to obtain several patch feature maps of different scales; Perform pooling processing on the obtained patch feature maps; Define corresponding linear layers based on the obtained patch feature maps of different scales to obtain object queries obtained by converting the patch features of different scales; According to the different size ratios of the multi-scale input features extracted from the synthesized image, repeat the operation of the object queries obtained by encoding each patch at different scale features.

6. A self-supervised pre-training object detection method according to claim 1, characterized in that In step S4, the process of predicting the categories and bounding boxes of the learned object queries includes: Make predictions through a prediction head, where the prediction head includes , and ; for predicting bounding boxes, predicting whether the patch is pasted, for reconstructing object descriptors, where , , , then , , represents an object query related to the image calculated by the Transformer.

7. A self-supervised pre-training object detection system for the method according to claim 1, characterized in that It includes a patch block extraction module, a synthesized image composition module, a multi-scale feature transformation module, and a matching module; The patch extraction module is used to extract proposals from the given input image for a given input image, and select the top 30 proposals as the patches to be pasted; The composite image construction module is used to select an image from the dataset as the background image, paste the patches obtained from the patch extraction module into the background image to obtain a composite image, extract the RGB values of the colors of the downstream targets to be detected, and randomly select an area in the pasted patches and change its color to the color corresponding to the extracted RGB values; The multi-scale feature transformation module is used to extract the features of the composite image and the multi-scale features of the patches pasted in the composite image respectively, and encode the multi-scale features of the patches into object queries; The matching module is used to learn the object queries based on the features of the extracted composite image, predict the categories and bounding boxes of the learned object queries to obtain a set of predictions, and match the set of predictions with the set of true annotations.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Training method of multi-label classification model and multi-label classification method of image

    CN114004992A

  • Modal transformation method based on deep learning

    WO2023005186A1