Image self-supervised representation learning method and device based on mask pre-training
By constructing two encoders to generate visible blocks and mask block marks, and combining the image marks of the VAE model for fit training, the problems of insufficient semantic information feedback and insufficient representation capabilities in image representation learning in the prior art are solved, and stronger image feature learning and reasoning capabilities are achieved.
Patent Information
- Application Number
- CN202310778826.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-06-28
AI Technical Summary
The existing masked image modeling method based on Token Reconstruction has problems such as insufficient semantic information feedback and insufficient representation ability in image representation learning.
By constructing two encoders to generate marks of visible blocks and mask blocks, combining the image marks of the entire image generated by the VAE model, fitting training is performed to realize image self-supervised representation learning based on mask pre-training.
The visible block representation and mask block representation are effectively fused, which improves image feature learning and capture capabilities, enhances the overall learning and reasoning capabilities of the model, and solves the semantic complementarity and enhancement problems of image representation learning.
Smart Images

Figure CN116894994B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image representation learning, and in particular to a method and device for image self-supervised representation learning based on mask pre-training. Background Art
[0002] As a new image representation learning method, the mask pre-training method has stronger data representation than the autoencoder. Compared with the clustering method that relies heavily on data augmentation strategies, a large number of negative samples and low scalability, mask image modeling is more in line with the label-free and independence required by unsupervised learning, which can facilitate the realization of image representation learning in the pre-training stage.
[0003] In the prior art, image representation learning based on image mask pre-training usually adopts a mask image modeling method based on Token Reconstruction, that is, the token of the visible block is obtained through a Variational AutoEncoder (VAE), and then the token of the mask block is inferred. Specifically, the VAE goal is usually to map a complete 224×224 image to a 14×14 token matrix, and each matrix element corresponds to a 16×16 image block. Although the traditional mask image modeling method based on Token Reconstruction can learn high-level semantic information of the image, since the matrix mapped by VAE is used as the pre-training target for mask image modeling, it will inevitably limit the completeness of image representation learning. For example, it is difficult to fully represent the high-level semantic information of a 224×224 image using only a 14×14 token matrix; and it is necessary to rely on an additional VAE to provide pre-training targets for the model. The token matrix generated by VAE is the feedback of the entire image semantics. During the training process, incorrect tokens may be generated and there may be insufficient feedback of semantic information, resulting in insufficient ability to generate image representation. Summary of the invention
[0004] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a method and device for self-supervised image representation learning based on mask pre-training, which can compensate for and enhance image representation, and improve the image's high-level semantic representation ability as well as learning and reasoning ability.
[0005] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0006] A method for learning image self-supervised representation based on mask pre-training, comprising:
[0007] S01. Obtain an image data sample set and perform preprocessing operations to obtain a processed image data sample set;
[0008] S02. Cut each image data sample in the processed image data sample set into image blocks evenly and without overlap, select a portion of the cut image blocks as visible blocks and the rest as mask blocks;
[0009] S03. Inputting the visible blocks into two different encoders respectively, wherein the first encoder is used to generate a visible block mark, and the second encoder is used to infer a mask block mark according to the visible block content;
[0010] S04. Inputting the original image data sample into the VAE model to generate an image tag corresponding to the entire image, wherein the image tag includes a visible block tag and a mask block tag;
[0011] S05. Perform fitting training on the visible block tokens (tokens) generated by the two encoders, the mask block tokens, and the image tokens generated by the VAE model to obtain a trained model;
[0012] S06. Use the trained model to learn the high-level semantic representation corresponding to each image block in the target image.
[0013] Furthermore, the step S03 includes:
[0014] S301. Convert the visible block into a two-dimensional feature matrix, and expand the feature dimension of each image block through a linear mapping layer to obtain visible block features;
[0015] S302. Initialize two additional cls tokens and concatenate them with the visible block features to obtain two feature matrices;
[0016] S303. Add corresponding position information to the two feature matrices respectively, and use them as input features of the first encoder and the second encoder respectively;
[0017] S304. The first encoder of the two encoders generates a label corresponding to the visible block, and the second encoder infers a label corresponding to the mask block.
[0018] Furthermore, the step S05 further includes matching the mask block mark inferred by the second encoder with the mask block position, so as to construct a mark corresponding to the mask block according to the matching result.
[0019] Furthermore, the step of matching the mask block mark inferred by the second encoder with the mask block position to construct a mark corresponding to the mask block according to the matching result includes:
[0020] The image tags generated by the VAE model are combined with the position information of the corresponding mask blocks one by one to construct a token space.
[0021] Two sets A are constructed by the image tags generated by the VAE model and the mask block tags inferred by the second encoder. u and
[0022] A u and The sets are sorted so that the mean square error (MSE) between the two sets is minimized;
[0023] According to the sorted The element value of the combination corresponds to A u The elements in the set search and match the position information of the corresponding mask mark in the mark space;
[0024] The found position information of the mask mark is used as the position information of the mask block mark inferred by the second encoder to sort the mask block marks, and construct the mark corresponding to the mask block.
[0025] Furthermore, the step S05 also includes constructing a model pre-training objective function, the steps including:
[0026] S501. Calculate the visible block marker loss according to the visible block marker generated by the first encoder and the visible block marker generated by the VAE model, and calculate the mask block loss according to the mask block marker generated by the second encoder and the mask block marker generated by the VAE model;
[0027] S502. Construct a model pre-training objective function according to the visible block labeling loss and the mask block loss.
[0028] Furthermore, the model pre-training objective function constructed is:
[0029]
[0030] in, represents the visible block labeling loss, Y v represents the visible block marker generated by the first encoder, Represents the visible block marker generated by the VAE model, l c () represents the cross entropy function, represents the mask block loss, Y u represents the mask block tag generated by the second encoder, It represents the mark of the mask block generated by the VAE model after sorting according to the position of the mask mark, and λ is the proportion of the mask block semantic information in the entire image semantic information.
[0031] Furthermore, in step S06, by inputting the target image into the trained model, a high-level semantic representation corresponding to each image block in the entire image is obtained, wherein the image block representation generated by the first encoder is directly used as the corresponding image block representation, and the mask block representation inferred by the second encoder is matched with the position information of the mask block representation to infer the final mask block representation.
[0032] Furthermore, the specific steps of step S06 include:
[0033] S601. Cut the target image into image blocks evenly and without overlap, and divide them into a first set A and a second set B according to a specified ratio:
[0034] S602. Input the first set A as visible blocks and the second set B as mask blocks into the trained model to obtain a first feature representation e corresponding to the visible block set. A and the first mark And infer the second feature representation e corresponding to the mask block set A→B and the second marker And the second set B is used as the visible block and the first set A is used as the mask block to input into the trained model, so as to obtain the third feature representation e corresponding to the visible block set. B and the third marker And the fourth feature representation e is inferred as a set of mask blocks B→A and the fourth mark
[0035] S603. The third mark and the second marker Matching is performed based on the third tag The position information of the first feature represents A→B Sort by the first tag The position information of the fourth feature represents B→A Sort by
[0036] S604. Represent the sorted first features as e A→B , the fourth feature representation after sorting is e B→A With the first feature representation e A Combine them to get a high-level semantic representation of the target image.
[0037] Furthermore, the two encoders respectively include a Transformer-Encoder module, a linear layer and a Sofftmax layer, the Transformer-Encoder module extracts high-level semantic features of visible blocks through an Add&Norm layer, a Feed Forward layer and a multi-head attention layer, the high-level semantic features include visible block representation and mask block representation, the Add&Norm layer includes Add and Layer Normalization layers, the Add layer is used to implement residual connection, the Layer Normalization layer is used to normalize the hidden layer in the neural network to a standard normal distribution, and the FeedForward layer is a fully connected layer comprising two layers, wherein the activation function of the first layer is ReLU, and the second layer does not use an activation function; the two encoders are based on Generate the corresponding markup, where represents the visible block flag generated by the first encoder or the mask block flag generated by the second encoder, er represents the output of the Transformer-Encoder module, Softmax represents the Softmax layer, and W r e r +b r represents the linear layer, W r is the feature matrix to be trained; b r is the feature bias to be trained.
[0038] A computer device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.
[0039] Compared with the prior art, the advantages of the present invention are as follows: the present invention constructs two encoders to generate labels of visible blocks and mask blocks respectively, the first encoder learns the semantic information of the visible blocks, and the second encoder infers the semantic information of the mask blocks according to the content of the visible blocks, and combines the image labels of the entire image generated by the VAE model to jointly construct a token matrix for fitting training, thereby realizing image self-supervised representation learning based on mask pre-training, and the obtained image semantic representation effectively integrates the visible block representation and the mask block representation, which can improve the model's image feature learning and capture capabilities, and because the second encoder can realize reasoning, it can also make up for and enhance the image representation, thereby effectively enhancing the overall learning and reasoning capabilities of the model, and during the training process, the visible block representation and the mask block representation can promote and complement each other, greatly enhancing the high-level semantic representation capability of the image, and realizing semantic complementarity and enhancement of image representation learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1It is a schematic diagram of the implementation flow of the image self-supervised representation learning method based on mask pre-training in this embodiment.
[0041] Figure 2 Schematic diagram of the network architecture for realizing semantic complementation and enhanced image self-supervised representation learning in this embodiment.
[0042] Figure 3 Schematic diagram of the structure of the encoder used in this embodiment.
[0043] Figure 4 Schematic diagram of the network structure of the VAE module used in this embodiment. DETAILED DESCRIPTION
[0044] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0045] As shown in the disclosure of the present invention, unless the context clearly indicates an exception, the words "a", "an", "a kind" and / or "the" do not specifically refer to the singular, but may also include the plural. The words "first", "second" and similar words used in the disclosure of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, the words "include" or "comprise" and the like mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0046] The present invention constructs two encoders to generate tokens of visible blocks and mask blocks respectively. The first encoder learns the semantic information of the visible blocks, and the second encoder infers the semantic information of the mask blocks according to the content of the visible blocks. Combined with the image tags of the entire image generated by the VAE model, a token matrix is jointly constructed for fitting training, which can effectively improve the image feature learning and capture capabilities of the model. Moreover, since the second encoder can realize reasoning, it can also compensate and enhance the image representation, thereby effectively enhancing the overall learning and reasoning capabilities of the model. In addition, they can promote and compensate each other during the training process, effectively enhancing the high-level semantic representation capabilities of the image.
[0047] like Figure 1 , 2 As shown, the steps of the image self-supervised representation learning method based on mask pre-training in this embodiment include:
[0048] Step S01. Data acquisition: acquiring an image data sample set and performing a preprocessing operation to obtain a processed image data sample set.
[0049] The original image dataset ImageNet-1K is obtained, and each data sample in the original image dataset ImageNet-1K is preprocessed such as rotated, cropped, and colored to expand the sample size of the dataset ImageNet-1K, and then converted into a uniform size (such as 224×224) to obtain the processed image data sample set ImageNet-1K-1.
[0050] S02. Image block segmentation: cut each image data sample in the processed image data sample set into image blocks evenly and without overlap, and select part of the cut image blocks as visible blocks and the rest as mask blocks.
[0051] After step S01, each sample X of the ImageNet-1K-1 dataset is cut into patches of the same size (e.g., 16×16) and the same number of patches evenly and without overlap, and a set is randomly selected as the visible patch X according to a certain ratio. v , the remaining image blocks are used as mask blocks X u , we know that X=X v ∪X u The above extraction ratio may be specifically 50%, and of course it may also be configured and selected according to actual needs.
[0052] S03. Label generation: The visible blocks are input into two different encoders respectively, wherein the first encoder is used to generate visible block labels, and the second encoder is used to infer mask block labels according to the visible block content.
[0053] S301. Set the visible block x i ∈X v Convert it into a two-dimensional feature matrix and use a linear mapping layer to expand the feature dimension of each image block to obtain the visible block feature x′ i =x i E;
[0054] S302. Initialize two additional cls token x cla , and are respectively related to the visible block features x′ i Concatenate and obtain two feature matrices [x cla , x′ 1 , …, x′ i ,…,x′ I ];
[0055] S303. Respectively convert the two feature matrices [x cls , x′ 1 ,…,x′ i ,…,x′ I ] plus the corresponding position information, and used as the input features of the first encoder and the second encoder respectively, that is, e0 =[x cla , x′ 1 ,…,x′ i ,…,x′ I ]+E pos ;
[0056] S304. The e obtained in step S303 is 0 Through the first encoder f θ Generate a token corresponding to a visible block Right now and through the second encoder Infer the token corresponding to the mask block Right now
[0057] In a specific application embodiment, the visible block X v After being converted into a two-dimensional feature matrix, the RGB three channels are expanded to 1024 dimensions through a linear layer, and then the initialized additional cls token is added to the 0 position of the input feature and concatenated with the visible block feature to form a feature matrix [x cla , x′ 1 ,…,x′ i ,…,x′ I ].
[0058] In this embodiment, the two encoders adopt the same structure, such as Figure 3 As shown, they include a Transformer-Encoder module, a linear layer and a Softmax layer respectively, wherein the Transformer-Encoder module extracts high-level semantic features of the visible blocks through the Add&Norm layer, the Feed Forward layer and the multi-head attention layer, and the high-level semantic features include visible block representation and mask block representation, the Add&Norm layer includes the Add and Layer Normalization layers, the Add layer is a residual connection layer, the Layer Normalization layer is used to normalize the hidden layer in the neural network to a standard normal distribution, and the Feed Forward layer is a fully connected layer containing two layers, wherein the activation function of the first layer is ReLU, and the second layer does not use an activation function.
[0059] In a specific application embodiment, the specific steps of forward propagation of the Transformer-Encoder module to generate visible block representation include:
[0060] (1) e l As the input of the l+1th layer Transformer-Encoder, high-level features are extracted through the multi-head attention layer and the Add&Norm layer:
[0061] e′ l =LayerNorm(e l +MultiHeadAttention(e l )) (1)
[0062] Among them, e l is the input of the l-th layer Transformer-Encoder module, e′ l is the potential feature extracted from the lth layer, MultiHeadAttention(·) is the multi-head attention layer, LayerNorm(·) is the LayerNormalization layer in the Add&Norm layer used to normalize the hidden layer in the neural network to a standard normal distribution, e l +MultiHeadAttention(e l ) corresponds to the Add layer in the Add&Norm layer. Add is a residual connection. The multi-head attention layer function MultiHeadAttention(e l ) multiple self-attention functions Self-Attention(e l )’s output features are concatenated together.
[0063] Each self-attention function Self-Attention(e l ) specifically include:
[0064] Q l =e l *W l,Q (2)
[0065] K l =e l *W l,K (3)
[0066] V l =e l *W l,V (4)
[0067]
[0068] Among them, W l,Q , W l,K , W l,V are the three feature matrices to be trained in the lth Encoder module, used to train the feature e l Perform linear mapping operation, d is the matrix Q l and K l The number of columns, different self-attention functions Self-Attention(e l) corresponds to different linear mapping matrices W l,Q , W l,K , W l,V , to indicate that the model judges the degree of association between image features from different perspectives.
[0069] (2) The feature e′ extracted from the lth layer i Input to the Feed Forward layer and Add&Norm layer to extract semantic features:
[0070] e l+1 =LayerNorm(e′ l +FeedForward(e′ l )) (6)
[0071] Among them, FeedForward(·) is a two-layer fully connected layer. The activation function of the first layer is ReLU, and the second layer does not use an activation function. It is expressed by the following formula:
[0072] max(0,e′ l W l1 +b l1 )W l2 +b l2 (7)
[0073] Among them, W l1 and W l2 is the feature matrix to be trained; b l1 and b l2 is the feature bias to be trained.
[0074] (3) Repeat steps (1) to (3) for a total of Num times, where Num is a set number of times, and finally obtain high-level semantic information about each visible block, including visible block representation and mask block representation. Num can be set according to actual needs, for example, it can be set to 12 times.
[0075] (4) In this embodiment, the specific process of the linear layer and the Softmax layer can be expressed as follows:
[0076]
[0077] in, represents the first encoder f θ Generated visible block markers Or the second encoder f θ Generated mask block markers e r represents the feature output of the last layer of the Transformer-Encoder module, Softmax represents the Softmax layer, and W r er +b r represents the linear layer, W r is the feature matrix to be trained; b r is the feature bias to be trained. That is, the two encoders follow Generate the corresponding markup
[0078] In this embodiment, the first encoder is mainly used to learn the semantic information of the visible block, and the second encoder is mainly used to infer the semantic information of the mask block based on the visible block information. Image features that are difficult to learn or capture can be effectively learned and captured, thereby improving the learning and capture capabilities of image features. Furthermore, since the second encoder has reasoning capabilities, it can enhance image representation. The combination of the two encoders can effectively enhance the overall learning and reasoning capabilities of the model.
[0079] S04. Image label generation: Input the original image data sample into the VAE model to generate the image label corresponding to the entire image. The image label includes a visible block label and a mask block label.
[0080] Input the original image into the trained VAE model, such as Figure 4 As shown, the VAE model obtains the token matrix Y corresponding to the entire image through the tag head. v , Y u}; where Y v is the token corresponding to the visible block, Y u is the token corresponding to the mask block. The generated token matrix is specifically a 14×14 token matrix. Each matrix element represents the high-level semantic representation of the 16×16 image block at the corresponding position, and each matrix element value is a number between [1, 8192]. The decoder in the VAE model uses the 14×14 token matrix to restore the entire image. In this embodiment, the decoder in the VAE model does not participate in model training.
[0081] S05. Fitting training: Fit the visible block labels, mask block labels generated by the two encoders and the image labels generated by the VAE model to obtain the training result output.
[0082] The first encoder f θ The generated image tokens correspond one-to-one to the visible blocks, but the second encoder Since the inferred image token has no position information to match it with the mask block, the learned semantic representation cannot be matched with the corresponding mask block. Inferred mask block labels Match the mask block position to construct the tag corresponding to the mask block according to the matching result. Specifically, a mask tag matching header is constructed to be used for the second encoder Inferred mask block labels Matching is performed with the mask block position to construct a mark corresponding to the mask block according to the matching result. The specific straight line steps of the mask mark matching head include:
[0083] Step (1) Label the image generated by the VAE model as Y u Combined with the position information of the corresponding mask blocks one by one to construct a marking space;
[0084] Step (2) The images generated by the VAE model are labeled Y u , Second Encoder Inferred mask block labels Construct two sets A u and
[0085] Step (3) constructs A u and The sets are sorted so that the mean square error between the two sets is minimized;
[0086] Step (4) according to the sorted The element value of the set corresponds to A u The elements in the collection look up and match the corresponding masked marked position information in the mark space;
[0087] Step (5) uses the position information of the mask mark found as the second encoder The inferred position information of the mask block mark is used to mark the mask block Sort them, and finally construct the mark corresponding to the mask block by mask block mark-sorted mark ~.
[0088] This embodiment also includes constructing a model pre-training objective function, the steps include:
[0089] Calculate the visible block marker loss according to the visible block marker generated by the first encoder and the visible block marker generated by the VAE model, and calculate the mask block loss according to the mask block marker generated by the second encoder and the mask block marker generated by the VAE model;
[0090] The model pre-training objective function is constructed based on the visible block labeling loss and mask block loss.
[0091] The model pre-training objective function constructed in this embodiment is specifically:
[0092]
[0093] in, represents the visible block labeling loss, Y v represents the visible block marker generated by the first encoder, Represents the visible block marker generated by the VAE model, l c () represents the cross-entropy function, that is, the loss between tokens is calculated by the cross-entropy function. represents the mask block loss, Y u represents the mask block tag generated by the second encoder, It represents the mark of the mask block generated by the VAE model after sorting according to the position of the mask mark, and λ is the proportion of the mask block semantic information in the entire image semantic information. The specific value of λ can be 1, which means that all image blocks contribute equally to the semantic information of the entire image.
[0094] This embodiment constructs a model pre-training objective function according to the visible block labeling loss and the mask block labeling loss, and performs training according to the objective function during the pre-training process, so that the visible block labeling and the mask block labeling can promote and complement each other during the training process, thereby effectively enhancing the high-level semantic representation capability of the image.
[0095] S06. Use the trained model to learn the high-level semantic representation corresponding to each image block in the target image.
[0096] After the training is completed in step S05, the network parameters of the two encoders are retained, and the high-level semantic representation corresponding to each image block in the entire image is obtained by inputting the target image into the trained model. θ The image block representation generated by the Transformer-Encoder module directly corresponds to the representation of the image block, and the second encoder It is necessary to obtain the position information of each mask block representation by inputting the tokens generated by the two encoders through the mask block, that is, to match the mask block representation inferred by the second encoder with the position information of the mask block representation. The matching process is similar to the pre-training stage, so as to further infer the corresponding mask block representation. After steps S01 to S05, the two encoders in the trained model can generate the tokens of the visible blocks and mask blocks. Combined with the matching of the mask block representation and the corresponding position information, the high-level semantic representation corresponding to each image block in the entire image can be obtained.
[0097] S601. Cut the target image into image blocks evenly and without overlap, and divide them into a first set A and a second set B according to a specified ratio.
[0098] For example, the target image is converted into a size of 224×224, cut into patches of the same size of 16×16 and the same number of patches evenly and without overlap, and randomly divided into two sets A and B at a ratio of 50%.
[0099] S602. Input the first set A as the visible block and the second set B as the mask block into the trained model to obtain the first feature representation e corresponding to the visible block set A. A and the first mark And infer the second feature representation e corresponding to the mask block set B A→B and the second marker And the second set B is used as the visible block and the first set A is used as the mask block to input into the trained model, and the third feature representation e corresponding to the visible block set B is obtained. B and the third marker And the fourth feature representation e as the mask block set A is inferred B→A and the fourth mark
[0100] S603. Mark the third and the second marker Matching (specifically, it can be implemented by using the mask tag matching head constructed in the above step S03) to match the third tag The position information of the first feature represents e A→B Sort by first tag The position information of the fourth feature represents e B→A Sort by
[0101] S604. Represent the sorted first features as e A→B , the fourth feature representation after sorting is e B→A With the first characteristic representation e A The high-level semantic representation of the target image is obtained by combining (addition or concatenation). Since the obtained image semantic representation effectively integrates the visible block representation and the mask block representation, it can effectively improve the high-level semantic representation ability of the image and solve the problem of insufficient image representation generated by the traditional mask image modeling method based on Token Reconstruction.
[0102] This embodiment can further configure a new linear layer initialized after the Transformer-Encoder module to perform further linear probe or fine probe fine-tuning operations on the high-level semantic information of the target image when performing downstream tasks, so as to fit the downstream task target. The image representations learned through the above-mentioned architecture of the present invention can promote and complement each other in downstream tasks, thereby realizing image self-supervised representation learning with semantic complementarity and enhancement functions.
[0103] The present invention can realize image representation learning and can be applied to scenes such as image target classification and target detection, and can also be applied to multiple image scenes such as face recognition and target tracking, etc.
[0104] This embodiment also provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.
[0105] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A self-supervised image representation learning method based on mask pre-training, It is characterized in that the steps include: S01. Obtain an image data sample set and perform preprocessing operations to obtain a processed image data sample set; S02. Cut each image data sample in the processed image data sample set into image blocks evenly and without overlap, select a portion of the cut image blocks as visible blocks and the rest as mask blocks; S03. Inputting the visible blocks into two different encoders respectively, wherein the first encoder is used to generate a visible block mark, and the second encoder is used to infer a mask block mark according to the visible block content; S04. Inputting the original image data sample into the VAE model to generate an image tag corresponding to the entire image, wherein the image tag includes a visible block tag and a mask block tag; S05. Perform fitting training on the visible block marks, the mask block marks, and the image marks generated by the two encoders to obtain a trained model; S06. Use the trained model to learn the high-level semantic representation corresponding to each image block in the target image.
2. The image self-supervised representation learning method based on mask pre-training according to claim 1, It is characterized in that The step S03 comprises: S301. Convert the visible block into a two-dimensional feature matrix, and expand the feature dimension of each image block through a linear mapping layer to obtain visible block features; S302. Initialize two additional cls tokens and concatenate them with the visible block features to obtain two feature matrices; S303. Add corresponding position information to the two feature matrices respectively, and use them as input features of the first encoder and the second encoder respectively; S304. The first encoder of the two encoders generates a label corresponding to the visible block, and the second encoder infers a label corresponding to the mask block.
3. The image self-supervised representation learning method based on mask pre-training according to claim 1, It is characterized in that The step S05 further includes matching the mask block mark inferred by the second encoder with the mask block position, so as to construct a mark corresponding to the mask block according to the matching result.
4. The image self-supervised representation learning method based on mask pre-training according to claim 3, It is characterized in that The step of matching the mask block mark inferred by the second encoder with the mask block position to construct a mark corresponding to the mask block according to the matching result includes: The image tags generated by the VAE model are combined with the position information of the corresponding mask blocks one by one to construct a tag space; Two sets A are constructed respectively from the image tags generated by the VAE model and the mask block tags inferred by the second encoder u and A u and The sets are sorted so that the mean square error between the two sets is minimized; According to the sorted The element values of the set correspond to A u The elements in the set search for and match the position information of the corresponding mask marker in the said marker space; The found position information of the mask mark is used as the position information of the mask block mark inferred by the second encoder to sort the mask block marks, and construct the mark corresponding to the mask block.
5. The image self-supervised representation learning method based on mask pre-training according to claim 1, It is characterized in that The step S05 also includes constructing a model pre-training objective function, the steps including: Calculate the visible block marker loss according to the visible block marker generated by the first encoder and the visible block marker generated by the VAE model, and calculate the mask block loss according to the mask block marker generated by the second encoder and the mask block marker generated by the VAE model; A model pre-training objective function is constructed according to the visible block labeling loss and the mask block loss.
6. The image self-supervised representation learning method based on mask pre-training according to claim 5, It is characterized in that The model pre-training objective function constructed is: in, represents the visible block labeling loss, Y v represents the visible block marker generated by the first encoder, Represents the visible block marker generated by the VAE model, l c () represents the cross entropy function, represents the mask block loss, Y u represents the mask block tag generated by the second encoder, It represents the mark of the mask block generated by the VAE model after sorting according to the position of the mask mark, and λ is the proportion of the mask block semantic information in the entire image semantic information.
7. The image self-supervised representation learning method based on mask pre-training according to any one of claims 1 to 6, It is characterized in that In step S06, the target image is input into the trained model to obtain a high-level semantic representation corresponding to each image block in the entire image, wherein the image block representation generated by the first encoder is directly used as the corresponding image block representation, and the mask block representation inferred by the second encoder is matched with the position information of the mask block representation to infer the final mask block representation.
8. The image self-supervised representation learning method based on mask pre-training according to claim 7, It is characterized in that The specific steps of step S06 include: S601. Cut the target image into image blocks evenly and without overlap, and divide them into a first set A and a second set B according to a specified ratio; S602. Input the first set A as visible blocks and the second set B as mask blocks into the trained model to obtain a first feature representation e corresponding to the visible block set. A and the first mark And infer the second feature representation e corresponding to the mask block set A→B and the second marker And the second set B is used as the visible block and the first set A is used as the mask block to input into the trained model, so as to obtain the third feature representation e corresponding to the visible block set. B and the third marker And the fourth feature representation e is inferred as a set of mask blocks B→A and the fourth mark S603. The third mark and the second marker Matching is performed based on the third tag The position information of the first feature represents A→B Sort by the first tag The position information of the fourth feature represents B→A Sort by S604. Respectively combine the sorted first feature representation e A→B , the sorted fourth feature representation e B→A with the first feature representation e A to obtain the high-level semantic representation of the target image.
9. The image self-supervised representation learning method based on mask pre-training according to any one of claims 1 to 6, It is characterized in that The two encoders respectively include a Transformer-Encoder module, a linear layer and a Softmax layer. The Transformer-Encoder module extracts high-level semantic features of visible blocks through an Add&Norm layer, a Feed Forward layer and a multi-head attention layer. The high-level semantic features include visible block representation and mask block representation. The Add&Norm layer includes an Add and a Layer Normalization layer. The Add layer is used to implement residual connection. The LayerNormalization layer is used to normalize the hidden layer in the neural network to a standard normal distribution. The Feed Forward layer is a fully connected layer containing two layers, wherein the activation function of the first layer is ReLU, and the second layer does not use an activation function. The two encoders are based on Generate the corresponding markup, where represents the visible block mark generated by the first encoder or the mask block mark generated by the second encoder, e r represents the output of the Transformer-Encoder module, Softmax represents the Softmax layer, and w r e r +b r represents the linear layer, W r is the feature matrix to be trained; b r is the feature bias to be trained.
10. A computer device comprising a processor and a memory, wherein the memory is used to store a computer program. It is characterized in that The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Article recommendation method based on self-supervised variational auto-encoder
CN113868517A
Anatomy-aware motion estimation
US20210397886A1