A self-supervised pre-training method based on a region screening module and multi-level comparison
By introducing technology based on region screening module and multi-level comparison in the self-supervised pre-training method, the self-supervised pre-training problem on multi-instance data sets is solved, and higher model accuracy and feature extraction performance are achieved.
Patent Information
- Application Number
- CN202210018471.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-01-07
AI Technical Summary
The existing self-supervised pre-training methods are difficult to effectively train on multi-instance datasets, especially in object detection tasks, and local instance information in the dataset cannot be effectively utilized.
The self-supervised pre-training method based on the region filtering module and multi-level comparison is adopted to filter out the local instance information block diagram through image RGB information entropy and LC significance detection values, and three hierarchical comparison loss functions are set, and the screened instance information is maximized.
It significantly improves the model accuracy on multi-instance datasets, can more effectively utilize local instance information in the dataset, and improves the performance of feature extractors.
Smart Images

Figure CN114387454B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a self-supervised pre-training method based on a region screening module and multi-level contrast. Background Art
[0002] The self-supervised pre-training method refers to using auxiliary tasks to mine its own supervision information from data without manual annotation. In this way, effective supervision information is constructed to pre-train a deep learning neural network model to learn a valuable feature extractor for downstream tasks such as image recognition, object detection, and semantic segmentation. Since manual data annotation in large-scale datasets is both expensive and time-consuming, the self-supervised method using the unlabeled training method has very important research value in the field of computer vision and has received more and more extensive attention.
[0003] Currently, most of the auxiliary tasks of the self-supervised pre-training method adopt contrastive learning. It maximizes the similarity between two different augmented images of the same image through a contrastive loss function, thereby learning the feature representation of the model in the dataset. Through contrastive learning, existing self-supervised learning methods have made effective progress in single-instance datasets such as ImageNet. ImageNet is an instance-centered dataset that not only contains only a single instance, but also maximally crops the background regions irrelevant to the instance. The key step of the self-supervised contrastive learning method is to maximize the similarity of different augmented representations in the same instance, which largely depends on the proportion of instance information in the entire image. Therefore, the current contrastive learning research is easy to achieve high accuracy on ImageNet. However, for multi-instance datasets commonly used in object detection, such as COCO and PASCAL VOC, the research on self-supervised pre-training has been difficult to make effective progress because this type of dataset not only does not center on instance information, but also does not crop a large amount of irrelevant background noise in the image.
[0004] Most of the existing self-supervised pre-training methods only perform contrastive learning on the global information of image augmentation. In order to improve the performance of the dense feature extractor through multi-scale prediction, some methods introduce the contrast of combining global and local information into the pre-training. However, they cannot guarantee that the local information of the contrast effectively contains instance information. This shows that the existing self-supervised methods are difficult to apply to the pre-training task of multi-instance datasets. Therefore, a self-supervised pre-training method based on a region screening module and multi-level contrast is urgently needed. Summary of the Invention
[0005] Objective of the Invention: To solve the problems existing in the prior art and realize the effective training of a self-supervised model based on a multi-instance dataset. The present invention provides a self-supervised pre-training method based on a region screening module and multi-level contrast. It can effectively screen out the local instance information block diagrams in the dataset and sets three levels of contrast loss functions, namely global, local, and "global-local", to maximize the utilization of the screened instance information, thereby effectively improving the accuracy of the pre-trained feature extractor.
[0006] Technical Solution: A self-supervised pre-training method based on a region screening module and multi-level contrast, characterized by comprising the following steps:
[0007] Step 1): Establish an initial deep learning neural network;
[0008] Step 2): Input the unlabeled training input data into the neural network and screen out the local block diagrams containing instance features based on the unsupervised data;
[0009] Step 3): Training step, train the deep learning neural network model based on the local block diagrams screened out from the unsupervised data through the loss function of multi-level contrast learning.
[0010] As an improvement of the present invention, in Step 2), enhanced diagrams of the dataset pictures are generated. For a given RGB picture x of the dataset, two enhanced diagrams v and v' of it are randomly generated. The generation methods of the enhanced diagrams include five methods: random size cropping, color jittering, random gray value transformation, gray image conversion, and random horizontal flipping;
[0011] After the two RGB enhanced diagrams of the picture are generated, they are divided into a plurality of neatly arranged block diagrams {P 1 , P 2 , …, P n} in a grid-like manner, where n represents the total number of block diagrams included in the enhanced diagram;
[0012] The image RGB information entropy is used to preliminarily screen the instance block diagrams. For a block diagram P of an enhanced diagram, it is divided into [P r : P g : P b according to the three different color channels of RGB. The calculation method of the image information entropy of the red channel P r is as follows:
[0013]
[0014] where p(r i ) represents the proportion of the pixel value i in the red channel P r , and the range of the pixel value is [0, 255]. The green channel P g and the blue channel Pb Image information entropy and The calculation method is similar to that of the above-mentioned red channel P r Next, calculate the total RGB information entropy H of the block diagram P , and the calculation method of the total RGB information entropy of the image is expressed as:
[0015]
[0016] In the entire enhanced image v, all the block diagrams {P 1 , P 2 , …, P n} divided by grid are sorted from high to low according to the image RGB information entropy H P , and the top k H block diagrams with high image information entropy are selected for further screening in Step 5;
[0017] Since the image RGB information entropy can only calculate the local information of the block diagrams in the enhanced image, the LC saliency detection value screening method for measuring global information is used to perform secondary screening on the block diagrams. In the enhanced image, the distance between a pixel and the pixels where other gray values are located in the image is used to measure the probability that the pixel belongs to the instance area. Assume that I k represents a pixel point in the enhanced image, then the saliency value of I k is calculated as follows:
[0018]
[0019] Among them, g(I k ) represents the gray value of the pixel I k , f n represents the occurrence frequency of the gray value n in the entire enhanced image, Dist(.) represents the Euclidean distance between two gray values. For an enhanced image v, it is converted into a grayscale image V g , and for all pixel points I k , calculate its saliency value in V g ;
[0020] Calculate the saliency value S P of the block diagram according to the saliency value of the pixel point, and its calculation method is expressed as:
[0021]
[0022] For the block diagrams screened in Step 4, sort them from high to low according to the saliency value S P of the block diagram, and further screen out the top k P with the highest saliency value S H) Small pieces, as the screening results of the instance area screening module, for the enhanced graph v, all the block graphs screened out are denoted as P(v);
[0023] Since the generation method of the enhanced graph includes a random size cropping method that can change the position features, it cannot be guaranteed that the positions of the block graphs screened out in step seven are corresponding and consistent in two enhanced graphs v and v' of an image. Therefore, the mutual information measurement method is used to match the block graphs screened in the two enhanced graphs v and v'. To accurately use them in the contrastive learning in step ten and step twelve, given two block graphs X and Y, the calculation method of their mutual information value M(X,Y) is as follows:
[0024] H(X,Y) = -∑ a,b p XY (a,b)log 2 p XY (a,b)
[0025] M(X,Y) = H X +H Y -H(X,Y)
[0026] where p XY (a,b) represents the joint probability distribution of two pixel values a and b in the two block graphs X and Y. Next, for a block graph, the one with the highest mutual information value in the other enhanced graph is selected as the matching block graph;
[0027] Calculate the global and local features for contrastive learning. The deep learning model of the present invention sequentially includes a backbone network f θ and two MLP heads. The backbone network selects the Resnet50 residual network. The MLP head includes a linear layer, a normalization operator, and a ReLu layer. For two enhanced graphs v and v' of an image in the multi-instance dataset and their block graphs P(v) and P(v') screened out in step seven, they are first put into the backbone network f θ for calculation, and the corresponding embedded feature vectors are output. Among them, the global feature vector I y , I y′ is obtained through the entire enhanced graph. The block graph is regarded as the local feature vector, denoted as P y , P y′ . After passing through the first MLP head, the corresponding projected features I z , I z′ and P z , P z′ are obtained. For the first enhanced graph v, its projected features also need to be input into the calculation of the second MLP head to obtain the predicted features I q and P q ;
[0028] As an improvement of the present invention, in step 3): Multilevel contrast learning is adopted to maximize the utilization of the instance information of the block diagrams screened in step seven. First, local contrast learning is carried out. For a block diagram in the enhanced graph v, its predicted feature is denoted as P q , in step eight, the matching block diagram of this block diagram from the enhanced graph v' is calculated, and the projected feature of this matching block diagram is denoted as P z′ . In order to enhance the feature similarity between the above-mentioned matching block diagrams, a local contrast loss function L local is set up, and its calculation method is expressed as follows:
[0029]
[0030] where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors;
[0031] Next, global contrast learning is carried out. For two enhanced graphs v and v' of an image in the dataset, the predicted feature of the enhanced graph v is set as I q , and the projected feature of the enhanced graph v' is denoted as I z′ , then the calculation method of the global contrast loss function L global is:
[0032]
[0033] where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors.
[0034] Since the positions of the local instance block diagrams are also very important potential information in downstream tasks, the present invention sets up a brand-new "global-local" contrast loss function which comprehensively applies the global and local feature representations and the position encoding of the local block diagrams to contrast learning. This position encoding is set as POS, representing the positioning information of a selected local instance block diagram in the entire enhanced graph. For an enhanced graph with a pixel size of 224×224, POS is set as a one-dimensional vector with all initial values of 0 and a length of 408. Assuming that the position coordinates of the pixel point at the upper left corner of a certain block diagram in the enhanced graph are [a, b], the setting method of its position encoding POC is to set the element values of vector subscripts a and 224 + b to 1. For an enhanced graph v, its comprehensive contrast learning connection representation C q is set up, and its calculation method is expressed as follows:
[0035] C q = cat(I q , P q , POS P,v )
[0036] where cat(·) represents the vector concatenation operation, and POS P,vRepresents the position encoding of block diagram P in the enhanced graph v, and the comprehensive contrast learning connection representation C for the corresponding enhanced graph v'. z′ , with a similar calculation method. Next, calculate the "global-local" contrast loss function:
[0037]
[0038] Step Thirteen: Next, set the total loss function, which is expressed as the combination of the above three-level contrast loss functions of global, local, and "global-local":
[0039]
[0040] where α, β, γ represent the weights for balancing these three contrast loss functions;
[0041] Apply the model with region screening module and multi-level contrast in Step Thirteen to the multi-instance dataset for unsupervised pre-training, then train the linear classification layer of the model according to the corresponding downstream task, and finally use the test set of the dataset to test the accuracy rate.
[0042] Furthermore, the method for generating the enhanced graph of the picture in Step One further includes random padding and affine transformation.
[0043] Furthermore, when using random size cropping to enhance the picture in Step One, set the cropping range to [0.08, 0.75].
[0044] Furthermore, set the grid segmentation of the block diagram in Step Two to the standard of dividing the pictures in the actual dataset into 32×32 blocks.
[0045] Furthermore, the averaging operation of the information entropy in Step Three is based on RGB images. If it is a dataset based on grayscale images, directly calculate the information entropy of the grayscale values.
[0046] Furthermore, the setting of the number k of block diagrams screened in each enhancement in Step Seven is determined according to the average information content of the instance regions in the actual dataset, and is set to 24 in the pre-training of the COCO2017 dataset.
[0047] Furthermore, the backbone network f in Step Nine θ is set to Resnet18, Resnet34, and Transformer feature extraction network.
[0048] Furthermore, the downstream tasks in Step Fourteen include image classification and object detection.
[0049] Beneficial effects: The present invention provides a self-supervised pre-training method based on a region screening module and multi-level contrast. When pre-training in the environment of a multi-instance dataset, a higher model accuracy can be obtained compared with the existing methods. This method can screen out local instance information block graphs in the dataset and set a multi-level contrast loss function to make the most of the screened instance information, effectively solving the problem that it is difficult to obtain and utilize effective instance information in contrastive learning based on multi-instance datasets. Description of the Drawings
[0050] Figure 1 It is the flowchart of the method of the present invention;
[0051] Figure 2 It is the screening result graph after instance region screening of the PASCAL VOC 2007 dataset pictures of the present invention; Detailed Embodiment
[0052] The present invention will be further described in detail below in conjunction with the drawings and the detailed embodiment:
[0053] This embodiment provides a self-supervised pre-training method based on a region screening module and multi-level contrast, and applies it to image classification and object detection in the PASCAL VOC and COCO datasets.
[0054] The flow of this method is as Figure 1 shown:
[0055] Step 1: Generate enhanced graphs of the dataset pictures. For a given RGB picture x of the dataset, randomly generate two of its enhanced graphs v and v′. The generation methods of the enhanced graphs include five methods: random size cropping, color jittering, random grayscale value transformation, grayscale image conversion, and random horizontal flipping.
[0056] Step 2: After the two RGB enhanced graphs of the picture are generated, divide them into a plurality of neatly arranged block graphs {P 1 , P 2 , …, P n} in a grid-like manner, where n represents the total number of block graphs included in the enhanced graph.
[0057] Step 3: Use the image RGB information entropy to initially screen the instance block graphs. For a block graph P of an enhanced graph, it is divided into [P r : P g : P b according to the three different color channels of RGB. The calculation method of the image information entropy of the red channel P r is as follows:
[0058]
[0059] where p(r i ) represents the proportion of pixel value i in the red channel P r . The range of pixel values is [0, 255]. The green channel P g and the blue channel P b Image information entropy and The calculation methods are similar to those of the above-mentioned red channel P r . Next, calculate the total RGB information entropy H P of the block diagram. The calculation method of the total RGB information entropy of the image is expressed as:
[0060]
[0061] Step 4: In the entire enhanced image v, sort all the grid-divided block diagrams {P 1 , P 2 , …, P n} in descending order according to the size of the image RGB information entropy H P , and select the top k H block diagrams with high image information entropy to enter the further screening in Step 5;
[0062] Step 5: Since the image RGB information entropy can only calculate the local information of the block diagrams in the enhanced image, the LC saliency detection value screening method for measuring global information is used to perform a secondary screening on the block diagrams. In the enhanced image, the distance between a pixel and the pixels of other gray values in the image is used to measure the probability that the pixel belongs to the instance area. Assume I k represents a pixel point in the enhanced image, then the saliency value of I k is calculated as follows:
[0063]
[0064] where g(I k ) represents the gray value of pixel I k , f n represents the occurrence frequency of gray value n in the entire enhanced image, Dist(.) represents the Euclidean distance between two gray values. For an enhanced image v, it is converted into a grayscale image V g . For all pixel points I k , calculate its saliency value in V g ;
[0065] Step 6: Calculate the saliency value S P of the block diagram according to the saliency value of the pixel point. Its calculation method is expressed as:
[0066]
[0067] Step Seven: For the block diagrams screened in Step Four, according to the significant value S of the block diagrams P sort them from high to low, and further screen out the top k (k < k P ) small blocks with the highest significant value S as the screening results of the instance area screening module. For the enhanced graph v, all the screened block diagrams are denoted as P(v); H ) small blocks, as the screening results of the instance area screening module. For the enhanced graph v, all the screened block diagrams are denoted as P(v);
[0068] Figure 2 Figure 11 shows the results of instance area screening for an image in the PASCAL VOC 2007 dataset by the method in Step Seven. It can be seen that in the absence of labels, the areas screened by the present invention very likely contain the most crucial instance information in the image. Therefore, it is effective to use the block diagrams screened by this module as local instance information for contrast learning.
[0069] Step Eight: Since the generation method of the enhanced graph includes the random size cropping method that can change the position features, it cannot be guaranteed that the positions of the block diagrams screened in Step Seven in two enhanced graphs v and v′ of an image are corresponding and consistent. Therefore, the method of mutual information measurement is used to match the block diagrams screened in the two enhanced graphs v and v′ so as to accurately use them in the contrast learning in Step Ten and Step Twelve. Given two block diagrams X and Y, the calculation method of their mutual information value M(X,Y) is as follows:
[0070] H(X,Y) = -∑ a,b p XY (a,b)log 2 p XY (a,b)
[0071] M(X,Y) = H X +H Y -H(X,Y)
[0072] where p XY (a,b) represents the joint probability distribution of two pixel values a and b in the two block diagrams X and Y. Next, for a block diagram, select the one with the highest mutual information value in the other enhanced graph as the matching block diagram;
[0073] Step Nine: Calculate the global and local features for contrast learning. The deep learning model of the present invention sequentially includes a backbone network f θ and two MLP heads. The backbone network selects the Resnet50 residual network, and the MLP head includes a linear layer, a normalization operator, and a ReLu layer. For two enhanced graphs v and v′ of an image in the multi-instance dataset and their block diagrams P(v) and P(v′) screened in Step Seven, first put them into the backbone network f θPerform calculations in [the medium] to output their corresponding embedded feature vectors, where the global feature vector I is obtained through the entire enhanced graph y , I y′ , the block graph is regarded as a local feature vector, denoted as P y , P y′ , after they pass through the first MLP head, the corresponding projected features I z , I z′ and P z , P z′ , for the first enhanced graph v, it is also necessary to input its projected features into the calculation of the second MLP head to obtain the predicted features I q and P q ;
[0074] Step Ten: Adopt multi-level contrastive learning to make the most of the instance information of the block graphs screened in Step Seven. First, perform local contrastive learning. For a block graph in the enhanced graph v, its predicted feature is denoted as P q , in Step Eight, the matching block graph of this block graph from the enhanced graph v' is calculated, and the projected feature of this matching block graph is denoted as P z′ , in order to improve the feature similarity between the above-mentioned matching block graphs, a local contrastive loss function L local is set up, and its calculation method is as follows:
[0075]
[0076] where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors;
[0077] Step Eleven: Next, perform global contrastive learning. For two enhanced graphs v and v' of an image in the dataset, set the predicted feature of the enhanced graph v as I q , and the projected feature of the enhanced graph v' is denoted as I z′ , then the calculation method of the global contrastive loss function L global is:
[0078]
[0079] where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors.
[0080] Step Twelve: Since the positions of the local instance block graphs are also very important potential information in downstream tasks, the present invention sets up a brand-new "global-local" contrastive loss function It comprehensively applies global and local feature representations and the position encoding of local block diagrams to contrastive learning. This position encoding is set as POS, representing the positioning information of a selected local instance block diagram in the entire augmented diagram. For an augmented diagram with a pixel size of 224×224, POS is set as a one-dimensional vector with initial values all being 0, and its length is 408. Assuming the position coordinates of the pixel point at the upper left corner of a certain block diagram in the augmented diagram are [a, b], the setting method of its position encoding POC is to set the element values of vector subscripts a and 224 + b to 1. For an augmented diagram v, its comprehensive contrastive learning connection representation C is established. q , and its calculation method is shown as follows:
[0081] C q = cat(I q , P q , POS P,v )
[0082] where cat(·) represents the concatenation operation of vectors, and POS P,v represents the position encoding of block diagram P in augmented diagram v. For the comprehensive contrastive learning connection representation C z′ of the corresponding augmented diagram v′, it has a similar calculation method. Next, calculate the "global - local" contrast loss function:
[0083]
[0084] Step 13: Next, set the total loss function, which is expressed as the combination of the above three - level contrast loss functions of global, local, and "global - local":
[0085]
[0086] where α, β, γ represent the weights for balancing these three contrast loss functions;
[0087] Step 14: Apply the model based on the region screening module and multi - level contrast in Step 13 to a multi - instance dataset for unsupervised pre - training, then train the linear classification layer of the model according to the corresponding downstream tasks, and finally use the test set of the dataset to test the accuracy rate.
[0088] In this example, self-supervised pre-training was first performed based on COCO2017, and then tested on the test set of PASCAL VOC 2007. The model proposed by the present invention achieved a top-1 image classification accuracy of 86.2% and object detection results of AP: 52.9, AP50: 79.5, and AP75: 58.0. Compared with previous self-supervised pre-training methods, this is a highly competitive result. In addition, pre-training was also performed based on the training and validation sets of the multi-instance dataset PASCAL VOC2007+2012, and an image classification accuracy of 66.1% could be obtained on the test set of PASCAL VOC 2007, showing a certain improvement compared with the results of previous self-supervised methods on the same dataset. These results strongly demonstrate that the self-supervised model proposed by the present invention based on the region screening module and multi-level contrast has excellent effects on the pre-training of multi-instance datasets.
Claims
1. A self-supervised pre-training method based on a region screening module and multi-level contrast, characterized in that: The method includes the following steps: Step 1): Establish an initial deep learning neural network; Step 2): Input the unlabeled training input data into the neural network, and based on unsupervised data, screen out the local block graphs containing instance features therein; Step 3): Training step, through the loss function of multi-level contrast learning, train the deep learning neural network model based on the local block graphs screened out by unsupervised data. The specific content of Step 2) includes: After the two RGB enhanced images of the picture are generated, they are segmented into multiple neatly arranged block images {P 1 , P 2 , …, P n} in a grid-like manner, where n represents the total number of block images contained in the enhanced image; Use the image RGB information entropy to preliminarily screen the instance block diagram. For a block diagram P of an enhanced image, it is divided into [P r :P g :P b according to the three different color channels of RGB. The calculation method of the image information entropy of the red channel P r is as follows: where p(r i ) represents the proportion of pixel value i in the red channel P r . The range of pixel values is [0, 255]. Next, calculate the total RGB information entropy H P of the block diagram. The calculation method of the total RGB information entropy of the image is expressed as: In the entire enhanced graph v, all the block graphs {P 1 , P 2 , …, P n} obtained by grid segmentation are sorted in descending order according to the image RGB information entropy H P in terms of magnitude, and the top k H block graphs with high image information entropy are selected; Since the RGB information entropy of the image can only calculate the local information of the block diagram in the enhanced image, the LC saliency detection value screening method that measures the global information is used to perform a secondary screening on the block diagram. In the enhanced image, the distance between a pixel and the pixels where other gray values are located in the image is used to measure the probability that the pixel belongs to the instance area. Assume I k represents a pixel point in the enhanced image, then I k The calculation method of the saliency value is as follows: where g(I k ) represents the gray value of pixel I k , f n represents the occurrence frequency of gray value n in the entire enhanced image, Dist(.) represents the Euclidean distance between two gray values. For an enhanced image v, it is converted into a grayscale image V g . For all pixel points I k , calculate its saliency value in V g ; Calculate the saliency value S of the block diagram according to the saliency value of the pixel points P , and its calculation method is expressed as: According to the significant value S of the block diagram P Sort from high to low, and further screen out the significant value S P The top k (k < k H ) small blocks, as the screening result of the instance area screening module. For the enhanced graph v, all the block diagrams screened out are represented as P(v); Calculate the global and local features for contrastive learning. The deep learning model sequentially includes a backbone network f θ and two MLP heads. The backbone network selects the ResNet50, and the MLP head includes a linear layer, a normalization operator, and a ReLU layer. For two augmented graphs v and v′ of an image in a multi-instance dataset and the selected patch graphs P(v) and P(v′), they are first put into the backbone network f θ for calculation, and the corresponding embedded feature vectors are output. Among them, the global feature vector I is obtained through the entire augmented graph y , I y′ , and the patch graph is regarded as the local feature vector, denoted as P y , P y′ . After passing through the first MLP head, the corresponding projected features I z , I z′ and P z , P z′ are obtained. For the first augmented graph v, its projected features also need to be input into the calculation of the second MLP head to obtain the predicted features I q and P q . Step 3) specifically includes: Adopt multi-level contrastive learning to maximize the utilization of the instance information of the screened block graphs. First, perform local contrastive learning. For a block graph in the enhanced graph v, its predicted feature is denoted as P q , calculate the matching block graph of this block graph from the enhanced graph v′, and the projected feature of this matching block graph is denoted as P z′ , in order to improve the feature similarity between the above-mentioned matching block graphs, set up a local contrastive loss function L local , and its calculation method is shown as follows: where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors; Perform global contrastive learning. For two augmented images v and v′ of an image in the dataset, set the predicted feature of augmented image v as I q , and denote the projected feature of augmented image v′ as I z′ . Then the global contrastive loss function L global is calculated as follows: where ‖·‖ represents the L-2 norm function, and <·> represents the inner product of vectors, Comprehensively apply global and local feature representations and the position encoding of local block diagrams to contrastive learning. This position encoding is set as POS, which represents the positioning information of a selected local instance block diagram in the entire augmented diagram. For an augmented diagram with a pixel size of 224×224, POS is set as a one-dimensional vector with all initial values of 0, and its length is 408. Assume that the position coordinates of the pixel point at the upper left corner of a certain block diagram in the augmented diagram are [a, b], and the setting method of its position encoding POS is to set the element values of vector subscripts a and 224 + b to 1. For an augmented diagram v, establish its comprehensive contrastive learning connection representation C q , and its calculation method is expressed as follows: C q = cat(I q ,P q ,POS P,v ) where cat(·) represents the concatenation operation of vectors, and POS P,v represents the position encoding of the block graph P in the enhanced graph v. For the comprehensive contrastive learning connection representation C of the corresponding enhanced graph v′ z′ , which has a similar calculation method. Next, calculate the "global-local" contrastive loss function: Set the total loss function, which is expressed as the combination of the above three-level contrast loss functions of global, local, and "global-local": where α, β, and γ represent the weights for balancing these three contrast loss functions.
2. The self-supervised pre-training method based on a region screening module and multi-level contrast according to claim 1, characterized in that The picture enhancement map generation method also includes random size cropping, color jittering, random grayscale value transformation, grayscale image conversion, random horizontal flipping, random padding, and affine transformation.
3. The self-supervised pre-training method based on a region screening module and multi-level contrast according to claim 2, characterized in that When using random size cropping enhancement for pictures, set the cropping range to [0.08, 0.75].
4. The self-supervised pre-training method based on a region screening module and multi-level contrast according to claim 1, characterized in that The block graph grid segmentation is set to the standard of dividing the pictures in the actual dataset into 32×32 blocks.
5. The self-supervised pre-training method based on a region screening module and multi-level contrast according to claim 1, characterized in that The averaging operation of the information entropy is based on RGB images. If it is a dataset based on grayscale images, directly calculate the information entropy of the grayscale values.
6. The self-supervised pre-training method based on a region screening module and multi-level contrast according to claim 1, characterized in that The setting of the number k of blocks for screening block graphs in each enhancement is determined according to the average information content of the instance regions in the actual dataset, and is set to 24 in the pre-training of the COCO2017 dataset.
Citation Information
Patent Citations
Language identification method of scene text image in combination with global and local information
CN110334705A
Earth observation image semantic segmentation method based on self-supervised learning
CN112308860A