Self-supervised training method and device of image encoder and terminal equipment
Patent Information
- Application Number
- CN202211667982.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-23
AI Technical Summary
[0004]本申请实施例提供的图像编码器的自监督训练方法、装置及终端设备,可以解决任务迁移能力差,影响了下游任务的准确度的问题
[0012]可以理解的是,上述第二方面至第五方面的有益效果可以参见上述第一方面中的相关描述,在此不再赘述。
Smart Images

Figure CN116188604B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of neural network technology, and in particular relates to a self-supervised training method, apparatus, terminal device and computer-readable storage medium for an image encoder. Background Technology
[0002] In recent years, deep neural networks have made remarkable progress in the field of computer vision. Algorithms based on deep neural networks have significantly outperformed traditional vision algorithms in tasks such as image recognition, object detection, and image segmentation. The core of these algorithms lies in training a high-performance neural network model using massive amounts of training data. However, manually labeling training data is expensive. Therefore, how to utilize virtually unlimited and inexpensive unlabeled data for network training is a problem of great concern to both academia and industry.
[0003] In related technologies, self-supervised learning of image encoders can be achieved through image transformations. For example, by performing a projection transformation on an image and using the original and transformed images as inputs, a Siamese network can be used to predict the projection transformation relationship between the two input images. However, self-supervised learning methods based on image transformations differ significantly from the tasks in practical applications, resulting in poor task transferability and affecting the accuracy of downstream tasks. Summary of the Invention
[0004] The self-supervised training method, apparatus, and terminal device for image encoders provided in this application can solve the problem of poor task transferability, which affects the accuracy of downstream tasks.
[0005] In a first aspect, embodiments of this application provide a self-supervised training method for an image encoder, including:
[0006] Acquire sample images; perform random cropping on the sample images to generate a first local image and a second local image of the sample images; input the first local image and the second local image into a reference image reconstruction model for cross-reconstruction to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image; determine the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result; update the network parameters of the reference image reconstruction model based on the loss value, and continue training using the updated reference image reconstruction model until the loss value of the updated reference image reconstruction model is less than the loss value threshold, and then determine the encoder of the updated reference image reconstruction model as the trained image feature encoder.
[0007] Secondly, embodiments of this application provide a self-supervised training apparatus for an image encoder, comprising:
[0008] The system includes a sample acquisition module for acquiring sample images; a cropping module for randomly cropping the sample images to generate a first local image and a second local image; an image reconstruction module for cross-reconstructing the first local image and the second local image into a reference image reconstruction model to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image; a loss analysis module for determining the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result; and an encoder determination module for updating the network parameters of the reference image reconstruction model based on the loss value, and continuing training with the updated reference image reconstruction model until the loss value of the updated reference image reconstruction model is less than a loss value threshold, at which point the encoder of the updated reference image reconstruction model is determined as the trained image feature encoder.
[0009] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the self-supervised training method of the image encoder described in any one of the first aspects above.
[0010] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the self-supervised training method for the image encoder described in any one of the first aspects.
[0011] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the self-supervised training method of the image encoder described in any of the first aspects.
[0012] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0013] The beneficial effects of this application embodiment compared with the prior art are as follows: By acquiring sample images and then randomly cropping them to generate a first local image and a second local image, the first local image and the second local image are then input into a reference image reconstruction model for cross-reconstruction to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. Based on the differences between the first local image and the first reconstruction result, and the differences between the second local image and the second reconstruction result, the loss value of the reference image reconstruction model is determined. The network parameters of the reference image reconstruction model are updated based on the loss value, and the updated reference image reconstruction model is used for further training until the loss value of the updated reference image reconstruction model is less than a loss value threshold. At this point, the encoder of the updated reference image reconstruction model is determined as the trained image feature encoder. Therefore, by reconstructing the second reconstruction result corresponding to the second local image from the first local image, and reconstructing the first reconstruction result corresponding to the first local image from the second local image, the contextual information of the image is utilized more efficiently to express image features more richly, thereby improving task transfer capability and increasing the accuracy of downstream tasks. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating a self-supervised training method for an image encoder provided in an embodiment of this application;
[0016] Figure 2 This is a schematic diagram of the structure of a reference image reconstruction model provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram of the structure of a reference image reconstruction model provided in another embodiment of this application;
[0018] Figure 4 This is a schematic diagram of the structure of a reference image reconstruction model provided in another embodiment of this application;
[0019] Figure 5 This is a schematic diagram showing the positional relationship of image blocks in local image m1 and local image m2 provided in another embodiment of this application;
[0020] Figure 6This is a schematic diagram of the structure of a self-supervised training device for an image encoder provided in an embodiment of this application;
[0021] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0026] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0028] It should be understood that the sequence number of each step in this embodiment does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.
[0029] In related technologies, self-supervised learning of image encoders can be achieved through image transformations. For example, by performing a projection transformation on an image and using the original and transformed images as inputs, a Siamese network can be used to predict the projection transformation relationship between the two input images. However, self-supervised learning methods based on image transformations differ significantly from the tasks in practical applications, resulting in poor task transferability and affecting the accuracy of downstream tasks.
[0030] This application provides a self-supervised training method for an image encoder. It acquires sample images and then randomly crops them to generate a first local image and a second local image. These two local images are then input into a reference image reconstruction model for cross-reconstruction, generating a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. Based on the differences between the first local image and the first reconstruction result, and the differences between the second local image and the second reconstruction result, the loss value of the reference image reconstruction model is determined. The network parameters of the reference image reconstruction model are updated based on the loss value, and training continues using the updated reference image reconstruction model until its loss value is less than a threshold. At this point, the encoder of the updated reference image reconstruction model is considered the successfully trained image feature encoder. Therefore, by reconstructing the second reconstruction result corresponding to the second local image from the first local image, and reconstructing the first reconstruction result corresponding to the first local image from the second local image, the method more efficiently utilizes the contextual information of the image, providing a richer expression of image features, thereby improving task transfer capability and increasing the accuracy of downstream tasks.
[0031] The self-supervised training method for image encoders provided in this application can be executed by a terminal device or a server. To illustrate the technical solution of this application, this method uses a terminal device as an example to explain the implementation process of the self-supervised training method for image encoders provided in the embodiments of this application.
[0032] In one embodiment, refer to Figure 1 This provides a self-supervised training method for image encoders, as an example rather than a limitation, such as... Figure 1 As shown, the self-supervised training method for this image encoder may include the following steps:
[0033] Step 101: Obtain the sample image.
[0034] The sample images can be images used to train neural network models, such as images used to train image encoders.
[0035] It should be understood that the acquired sample images can be multiple or a single image. Multiple sample images can be acquired simultaneously or sequentially as needed for training. The sample images can be stored in a database.
[0036] It should be understood that sample images can be generated in any way.
[0037] In one possible embodiment, the images can be retrieved from a database that stores sample images.
[0038] It should be understood that this database can be any database that is accessible during the training process.
[0039] Step 102: Perform random cropping on the sample image to generate a first local image and a second local image of the sample image.
[0040] The first local image can be an image of any region of the sample image.
[0041] The second local image can be an image of any region of the sample image.
[0042] In one possible embodiment, taking the random cropping of a sample image A as an example, the sample image A is randomly cropped according to preset cropping parameters to generate a first local image a1 and a second local image a2. The first local image a1 and the second local image a2 can be images of any region of the sample image A. The image content of the first local image a1 can completely overlap with the image content of the second local image a2, or they can not overlap, or they can partially overlap.
[0043] The preset cropping parameters can be used to limit the cropping ratio of random cropping. These preset cropping parameters can be set according to actual training needs, such as 0.1, 0.2, 0.3, or 0.4, etc.
[0044] In one example, taking the preset cropping parameter set to 0.2 as an example, when the preset cropping parameter is set to 0.2, the range of cropping ratios that can be selected during random cropping is 0.2-1.0. During random cropping, a cropping ratio is randomly selected from 0.2-1.0 to determine the cropping size.
[0045] In one example, taking the sample image A as an example with the preset cropping parameter set to 0.2, if a cropping ratio of 0.5 is randomly selected within the range of 0.2-1.0, then 0.5 times the area of sample image A will be used as the cropping area to crop out the corresponding local image a1. Then, if a cropping ratio of 0.3 is randomly selected, then 0.3 times the area of sample image A will be used as the cropping area to crop out the corresponding local image a2.
[0046] It should be understood that the cropping ratios selected randomly in two separate trials can be the same or different.
[0047] Furthermore, the image resolution of the cropped image can be scaled to make the first local image and the second local image have the same image resolution. That is, in one possible embodiment, the sample image is randomly cropped according to preset cropping parameters to obtain a first initial local image and a second initial local image; the image resolution of the first initial local image and the second initial local image is scaled according to a preset image resolution to generate a first local image and a second local image with the same image resolution.
[0048] The preset image resolution can be set as a multiple of the number of channels of the convolutional layer mapped according to the encoder's dimensions. For example, if the number of channels of the convolutional layer is 14, it can be 224×224 or 448×448, etc.
[0049] Step 103: Input the first local image and the second local image into the reference image reconstruction model for cross-reconstruction to generate the first reconstruction result corresponding to the first local image and the second reconstruction result corresponding to the second local image.
[0050] The first local image can be divided into multiple image blocks.
[0051] The second local image can be divided into multiple image blocks.
[0052] It should be understood that one sample image corresponds to one first local image and one second local image, and the number of image patches in the first local image and the second local image can be the same.
[0053] The first reconstruction result can be an image generated after reconstruction based on the second feature.
[0054] The second reconstruction result can be an image generated based on the first feature.
[0055] The reference image reconstruction model can be a reference model used for self-supervised training of the image encoder.
[0056] Among them, such as Figure 2 The schematic diagram of the reference image reconstruction model shown can include an encoder unit and a decoder unit. The encoder unit can include at least one encoder, and the decoder unit can include at least one decoder.
[0057] In one possible embodiment, such as Figure 3 As shown, the encoder unit includes an encoder, and the decoder unit includes a decoder.
[0058] In one possible embodiment, such as Figure 4 As shown, the encoder unit includes two encoders: a first encoder and a second encoder, and the decoder unit includes two decoders: a first decoder and a second decoder. The first encoder and the first decoder form the first branch of the reference image reconstruction model, and the second encoder and the second decoder form the second branch of the reference image reconstruction model.
[0059] The encoder can be an encoder network used for image encoding.
[0060] The decoder can be a decoder network used for image decoding.
[0061] Furthermore, the first local image and the second local image can be divided into multiple image blocks, and after preprocessing, the encoder and decoder units of the reference image reconstruction model are cross-reconstructed to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. That is, in one possible embodiment, step 103 includes:
[0062] The first local image and the second local image are each divided into multiple image blocks. Preprocessing is performed on these image blocks to generate visible image blocks corresponding to the first and second local images. Each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image. Based on the positional relationship between the first and second local images, first relative position encoding information of the first local image relative to the second local image and second relative position encoding information of the second local image relative to the first local image are generated. Based on the first relative position encoding information and the second feature, a first feature to be decoded is generated. Based on the second relative position encoding information and the first feature, a second feature to be decoded is generated. The first feature to be decoded is input into the decoder unit for decoding to generate a first reconstruction result corresponding to the first local image. The second feature to be decoded is input into the decoder unit for decoding to generate a second reconstruction result corresponding to the second local image.
[0063] Furthermore, since the image coding network is a Vision Transformers (ViT) network structure used for image input, it is necessary to divide the first local image and the second local image into multiple image blocks of the same size. That is, in one possible embodiment, taking the scenario of dividing a local image into image blocks as an example, the local image is divided into image blocks based on the image resolution of the local image and the number of channels of the convolutional layer of the encoder's dimension mapping, thereby generating the image blocks corresponding to the local image.
[0064] The number of channels in the convolutional layer mapped by the encoder's dimension can determine the pixel size of the local image.
[0065] It should be understood that, based on the pixel size of a local image, a local image is divided into multiple image blocks, and the image resolution of the local image divided by the pixel size of the local image equals the number of image blocks.
[0066] In one example, taking a local image with a resolution of 224×224 and a pixel count of 16×16, the 224×224 local image can be divided into 14×14 image blocks.
[0067] It should be noted that both the first and second local images need to be divided into image blocks separately, and the division method is the same, so it will not be described again.
[0068] Furthermore, preprocessing of the multiple image patches corresponding to the first local image and the second local image can include masking. That is, preprocessing the multiple image patches corresponding to the first local image and the second local image to generate each visible image patch corresponding to the first local image and the second local image includes:
[0069] According to the preset mask ratio, masking is performed on some image blocks in the first local image and the second local image respectively to generate each visible image block corresponding to the first local image and the second local image.
[0070] Among them, the masked image block can be an image block in a local image that has been masked, and the visible image block can be an image block in a local image that has not been masked.
[0071] The mask, in this context, can be a selected image, graphic, or object used to occlude (all or part) of the image being processed, thereby controlling the area or process of image processing. The image can be a first partial image or a second partial image.
[0072] Furthermore, the first local image and the second local image can be divided into multiple image blocks, and then, according to a preset mask ratio, some image blocks can be masked to generate visible image blocks corresponding to the first local image and the second local image.
[0073] The preset mask ratio can be the proportion of the number of masked image blocks to the total number of image blocks in the local image. This preset mask ratio can be set according to the actual situation, such as setting a mask ratio of 80%, 75%, or 70%, etc.
[0074] In this context, a masked image block can be an image block whose original pixel information has been erased. A visible image block can be an image block that has not undergone masking and retains its original pixel information.
[0075] In one example, taking the masking of a local image with a preset masking ratio of 75% as an example, when masking the local image, 75% of the image blocks corresponding to the local image are randomly masked. The pixel information of the 75% of the image blocks is erased, and the 75% of the image blocks are masked image blocks; the pixel information of the remaining 25% of the image blocks is retained, and the remaining 25% of the image blocks are visible image blocks.
[0076] It should be understood that this 75% image patch can be randomly selected from all image patches in the local image. That is, when masking different local images, the locations of the image patches whose pixel information is erased are not necessarily the same.
[0077] The first feature may be a feature that includes a visible image patch in the first local image.
[0078] The second feature may be a feature that includes visible image patches in the second local image.
[0079] Furthermore, image blocks of the first local image and the second local image can be obtained and encoded with position information. The image blocks of the first local image and the second local image, along with their corresponding encoded position information, are then input into the encoder unit for encoding processing to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image.
[0080] In one possible embodiment, taking a first local image and a second local image as examples, before inputting each visible image patch and its corresponding first position encoding information corresponding to the first local image into the encoder unit, the visible image patches corresponding to the first local image are standardized and then input into a convolutional layer to map them to a fixed-dimensional token, generating visible tokens corresponding to each visible image patch. The visible tokens corresponding to each visible image patch and their corresponding first position encoding information are then input into the encoder unit for encoding processing to generate a first feature corresponding to the first local image. Similarly, before inputting each visible image patch and its corresponding second position encoding information corresponding to the second local image into the encoder unit, the visible image patches corresponding to the second local image are standardized and then input into a convolutional layer to map them to a fixed-dimensional token, generating visible tokens corresponding to each visible image patch. The visible tokens corresponding to each visible image patch and their corresponding second position encoding information are then input into the encoder unit for encoding processing to generate a second feature corresponding to the second local image.
[0081] Further, in the case where the encoder unit includes an encoder, in one possible embodiment, each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are respectively input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image, including:
[0082] The first local image is obtained by acquiring each visible image block corresponding to the first local image and the first position encoding information of the first local image, wherein the first position encoding information includes the position encoding information of each image block in the first local image; the visible image blocks corresponding to the first local image and the corresponding first position encoding information are input into the encoder for encoding processing to generate a first feature corresponding to the first local image; the second local image is obtained by acquiring each visible image block corresponding to the second local image and the second position encoding information of the second local image, wherein the second position encoding information includes the position encoding information of each image block in the second local image; the visible image blocks corresponding to the second local image and the corresponding second position encoding information are input into the encoder for encoding processing to generate a second feature corresponding to the second local image.
[0083] It should be understood that when the encoder unit includes a single encoder, the encoder needs to generate a first feature corresponding to a first local image and a second feature corresponding to a second local image. The generation of the first feature and the second feature can be performed sequentially. Generating the first feature and the second feature of the first local image separately using a single encoder reduces memory usage.
[0084] Furthermore, in the case where the encoder unit includes two encoders, in one possible embodiment, each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are respectively input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image, including:
[0085] The first local image is obtained by acquiring each visible image block corresponding to the first local image and the first position encoding information of the first local image, wherein the first position encoding information includes the position encoding information of each image block in the first local image; the visible image blocks corresponding to the first local image and the corresponding first position encoding information are input into the first encoder for encoding processing to generate the first feature corresponding to the first local image; the second local image is obtained by acquiring each visible image block corresponding to the second local image and the second position encoding information of the second local image, wherein the second position encoding information includes the position encoding information of each image block in the second local image; the visible image blocks corresponding to the second local image and the corresponding second position encoding information are input into the second encoder for encoding processing to generate the second feature corresponding to the second local image.
[0086] It should be understood that when the encoder unit includes two encoders, the first encoder generates the first feature corresponding to the first local image, and the second encoder generates the second feature corresponding to the second local image. The first feature corresponding to the first local image and the second feature corresponding to the second local image can be generated simultaneously, which can improve the generation efficiency.
[0087] The first position encoding information may represent the position information of each image block of the first local image in the sequence.
[0088] The second position encoding information may represent the position information of each image block of the second local image in the sequence.
[0089] The first relative position encoding information may be the relative position encoding of each image block of the first local image relative to each image block of the second local image.
[0090] The second relative position encoding information may be the relative position encoding of each image block of the second local image relative to each image block of the first local image.
[0091] In one possible embodiment, based on the positional relationship between the first local image and the second local image, first relative position encoding information of the first local image relative to the second local image and second relative position encoding information of the second local image relative to the first local image are generated, including:
[0092] Based on the relative position and size ratio between the first local image and the second local image, and the position encoding information of the second local image, the first relative position encoding information of the first local image relative to the second local image is determined; based on the relative position and size ratio between the second local image and the first local image, and the position encoding information of the first local image, the second relative position encoding information of the second local image relative to the first local image is determined.
[0093] It should be understood that, with reference to Figure 2 The first feature to be decoded, generated using the second feature, serves as the feature for generating the first reconstruction result corresponding to the first local image. This needs to be combined with relative interpolation positional encoding for decoding to generate the first reconstruction result corresponding to the first local image. Since the first and second local images both originate from different regions of the same sample image, the image blocks in each image implicitly possess a relative positional relationship. For example, the wolf region in the lower left corner of the first local image corresponds to the wolf region in the upper right corner of the second local image. The positional and scale information of the same semantic block differs between the first and second local images. Because the input to the decoder unit, i.e., the second feature, comes entirely from the second local image, and the image block to be reconstructed is the image block of the first local image, and the image blocks in the first local image differ from those in the second local image in both position and scale, a relative positional encoding is required, derived from the positional and scale relationships of the two different sets of tokens.
[0094] In one example Figure 5The diagram shows the positional relationship of image blocks in local images m1 and m2. It is assumed that local images m1 and m2 are obtained by randomly cropping from sample image M. After scaling, local images m1 and m2 are divided into 4×4 image blocks. It is assumed that the area corresponding to local image m1 in sample image M is 16×16 pixels, i.e., width = 16pp, height = 16pp. The area corresponding to local image m2 in sample image M is 8×8 pixels. The ratio of the width (Scale W) and height (Scale H) of local image m2 to local image m1 is 0.5, i.e., Scale H = 0.5, Scale W = 0.5. Local image m2 is located to the lower right of local image m1, meaning that local image m2 has a translation of 6 pixels in both the width and height directions relative to local image m1, i.e., Translation X = 6pp, Translation Y = 6pp.
[0095] The position encoding of each image patch in local image m1 is as follows: Figure 5 As shown, i.e. (0, 0), (0, 1), ..., (3, 3), the position encoding of each image block in local image m2 relative to local image m1 is (1.25, 1.25), (1.25, 1.75), ..., (3.25, 3.25). Figure 5 The example illustrates the calculation method for the relative position encoding of the top-left image block on the second branch (i.e., the black image block (0, 0) in the first row and first column of local image m2). In summary, the formula for calculating the relative position encoding of the black image block in the first row and first column of local image m2 is (0 + Translation X / Height*num_X + 0.5*(Scale H-1), 0 + Translation Y / Width*num_Y + 0.5*(Scale W-1), where Translation X is the translation scale in the width direction, Height is the height of local image m1, num_X is the number of image block columns in local image m1, Scale H is the ratio of the height of local image m2 to that of local image m1, Translation Y is the translation scale in the length direction, Width is the width of local image m1, num_Y is the number of image block rows in local image m1, and Scale W is the ratio of the width of local image m2 to that of local image m1.
[0096] Correspondingly, the relative position encoding calculation formula for other image blocks (i.e., positions (i, j)) of local image m2 is (i*Scale H+Translation X / Height*num_X+0.5*(Scale H-1), j*Scale W+Translation Y / Width*num_Y+0.5*(Scale W-1), where i is the index of the image block in the height direction of local image m2, and j is the index of the image block in the width direction of local image m2.
[0097] The first feature to be decoded can be a feature sequence that includes the second feature.
[0098] The second feature to be decoded can be a feature sequence that includes the first feature.
[0099] Furthermore, the first mask token corresponding to the first local image can be concatenated with the second feature to generate a first feature sequence, and first relative position encoding information can be added to the first feature sequence to generate a first feature to be decoded. That is, in one possible embodiment, the first feature to be decoded is generated based on the first relative position encoding information and the second feature, including:
[0100] Based on the number of image blocks included in the first local image and the feature dimension of the second feature, a first mask token corresponding to the first local image is randomly generated; the first mask token corresponding to the first local image is concatenated with the second feature to generate a first feature sequence; and a first feature to be decoded is generated based on the first relative position encoding information and the first feature sequence.
[0101] The process of generating a first feature to be decoded based on the first relative position encoding information and the first feature sequence can be achieved by concatenating the first relative position encoding information and the first feature sequence, by superimposing the first relative position encoding information and the first feature sequence, or by using information fusion technology to fuse the first relative position encoding information and the first feature sequence to generate the first feature to be decoded.
[0102] The first mask token corresponding to the first local image can be a token without pixel information corresponding to the first local image. The number of first mask tokens is the same as the number of image blocks in the first local image. Assuming that the number of image blocks in the first local image is 14×14, then the number of first mask tokens should also be 14×14, for a total of 196 first mask tokens.
[0103] In one possible embodiment, taking the generation of a first mask token corresponding to a first local image as an example, a mask token with random initialization and the same feature dimension as the second feature is copied to generate a first mask token with the same number of image blocks as the first local image.
[0104] The number of copies is the same as the number of image blocks in the first local image. Assuming the first local image has 14×14 image blocks, then it is copied 196 times.
[0105] Furthermore, a first feature to be decoded can be generated based on the first relative position encoding information, the first feature sequence, and the position encoding information of the first local image. That is, in one possible embodiment, generating the first feature to be decoded based on the first relative position encoding information and the first feature sequence includes: generating the first feature to be decoded based on the first relative position encoding information, the first feature sequence, and the position encoding information of the first local image.
[0106] The first feature to be decoded is generated based on the first relative position encoding information, the first feature sequence, and the position encoding information of the first local image. This can be achieved by concatenating the first relative position encoding information, the position encoding information of the first local image, and the first feature sequence; by superimposing the first relative position encoding information, the position encoding information of the first local image, and the first feature sequence; or by using information fusion technology to fuse the first relative position encoding information, the position encoding information of the first local image, and the first feature sequence to generate the first feature to be decoded.
[0107] It should be understood that incorporating the positional encoding information of the first local image into the first feature to be decoded can further improve the accuracy of reconstruction.
[0108] Furthermore, the second mask token corresponding to the second local image can be concatenated with the first feature to generate a second feature sequence, and second relative position encoding information can be added to the generated second feature sequence to generate a second feature to be decoded. That is, in one possible embodiment, the second feature to be decoded is generated based on the second relative position encoding information and the first feature, including:
[0109] Based on the number of image blocks included in the second local image and the feature dimension of the first feature, a second mask token corresponding to the second local image is randomly generated; the second mask token corresponding to the second local image is concatenated with the first feature to generate a second feature sequence; and a second feature to be decoded is generated based on the second relative position encoding information and the second feature sequence.
[0110] The process of generating a second feature to be decoded based on the second relative position encoding information and the second feature sequence can be achieved by concatenating the second relative position encoding information and the second feature sequence, superimposing the second relative position encoding information and the second feature sequence, or by using information fusion technology to fuse the second relative position encoding information and the second feature sequence to generate the second feature to be decoded.
[0111] The second mask token corresponding to the second local image can be a token without pixel information corresponding to the second local image. The number of these second mask tokens is the same as the number of image blocks in the second local image. Assuming the second local image has 14×14 image blocks, the number of second mask tokens should also be 14×14, for a total of 196 second mask tokens.
[0112] In one possible embodiment, taking the generation of a second mask token corresponding to a second local image as an example, a mask token with random initialization and the same feature dimension as the first feature is copied to generate a second mask token with the same number of image blocks as the second local image.
[0113] The number of copies is the same as the number of image blocks in the second local image. Assuming the second local image has 14×14 image blocks, then it is copied 196 times.
[0114] Specifically, concatenating the first mask token corresponding to the first local image and the corresponding second feature can be achieved by adding the first mask token corresponding to the first local image to the sequence of the second feature.
[0115] Furthermore, a second feature to be decoded can be generated based on the second relative position encoding information, the second feature sequence, and the position encoding information of the second local image. That is, in one possible embodiment, generating a second feature to be decoded based on the second relative position encoding information and the second feature sequence includes: generating a second feature to be decoded based on the second relative position encoding information, the second feature sequence, and the position encoding information of the second local image.
[0116] Specifically, generating the second feature to be decoded based on the second relative position encoding information and the second feature sequence can be achieved by concatenating the second relative position encoding information, the position encoding information of the second local image, and the second feature sequence; it can also be achieved by superimposing the second relative position encoding information, the position encoding information of the second local image, and the second feature sequence; or it can be achieved by using information fusion technology to fuse the second relative position encoding information, the position encoding information of the second local image, and the second feature sequence to generate the second feature to be decoded.
[0117] It should be understood that incorporating the positional encoding information of the second local image into the second feature to be decoded can further improve the accuracy of reconstruction.
[0118] Furthermore, in the case where the decoder unit includes a decoder, in one possible embodiment, the first feature to be decoded is input into the decoder for decoding processing to generate a first reconstruction result corresponding to a first local image; the second feature to be decoded is input into the decoder for decoding processing to generate a second reconstruction result corresponding to a second local image. By generating the first reconstruction result corresponding to the first local image and the second reconstruction result corresponding to the second local image through the decoder respectively, the memory usage can be reduced.
[0119] Furthermore, in the case where the decoder unit includes two decoders, in one possible embodiment, the first feature to be decoded is input into the first decoder for decoding processing to generate a first reconstruction result corresponding to the first local image; the second feature to be decoded is input into the second decoder for decoding processing to generate a second reconstruction result corresponding to the second local image. By generating the first reconstruction result corresponding to the first local image and the second reconstruction result corresponding to the second local image through the first decoder and the second decoder respectively, the generation efficiency can be improved.
[0120] It should be understood that the information source of the first reconstruction result comes entirely from the second local image. The reason why pixel restoration can be successfully performed is that the second local images are all from different pixel regions of the same input sample image and have a certain semantic correlation.
[0121] It should be understood that this method utilizes the pixel context relationship across images. The first feature is input into the decoder unit of the reference image reconstruction model for decoding to generate a second reconstruction result corresponding to the second local image. The second feature is input into the decoder unit of the reference image reconstruction model for decoding to generate a first reconstruction result corresponding to the first local image. The information source of the second reconstruction result is entirely from the first local image, and the information source of the first reconstruction result is entirely from the second local image. Successful pixel restoration is achieved because the input samples all come from different pixel regions of the same input image, possessing a certain semantic correlation. Pre-training can be performed on image data without manual labels. Compared to fully supervised pre-training methods, this application does not rely on data labels, saving significant manpower and resources. Compared to other self-supervised algorithms, this application has a stronger ability to learn effective features and achieves higher accuracy. Compared to Masked Auto-Encoder (MAE), this application achieves an accuracy improvement of 0.5-1.0 on the ImageNet dataset using the same network structure as MAE (ViT-S, ViT-B, etc.).
[0122] It should be understood that the information source of the second reconstruction result comes entirely from the first local image. The reason why pixel restoration can be successfully performed is that the first local images are all from different pixel regions of the same input sample image and have a certain semantic correlation.
[0123] It should be understood that, with reference to Figure 2 The second feature to be decoded, generated using the first feature, serves as the feature for generating the second reconstruction result corresponding to the second local image. This needs to be combined with relative interpolation positional encoding for decoding to generate the second reconstruction result corresponding to the second local image. Since the first and second local images both originate from different regions of the same sample image, the image patches in each image implicitly contain relative positional relationships. For example, the wolf region in the lower left corner of the first local image corresponds to the wolf region in the upper right corner of the second local image. The positional and scale information of the same semantic block differs between the first and second local images. Because the input to the decoder unit, i.e., the second feature, comes entirely from the second local image, and the image patch to be reconstructed is the image patch of the second local image, and the image patch of the second local image differs from the image patch of the first local image in both position and scale, a relative positional encoding is required, derived from the positional and scale relationships of the two different sets of tokens.
[0124] It should be understood that the method for determining the second relative position encoding information is the same as the method for determining the first relative position encoding, and will not be repeated here.
[0125] Step 104: Determine the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result.
[0126] The loss value of the reference image reconstruction model can be used to measure the quality of the pre-trained image encoder, determine the performance of the image encoder, and then find optimization directions.
[0127] In one possible implementation, the loss function used during training is the Mean Square Error Loss (MSE LOSS).
[0128] The structure of the reference image reconstruction model is as follows Figure 4 As shown, in the case where the first encoder and the first decoder form the first branch of the reference image reconstruction model, and the second encoder and the second decoder form the second branch of the reference image reconstruction model, in one possible embodiment, the reference image reconstruction model is a Siamese Network, and the network parameters of the first encoder and the first decoder of the first branch of the reference image reconstruction model are shared with those of the second encoder and the second decoder.
[0129] Furthermore, when calculating the loss function in each training iteration, the loss corresponding to the input sample image can be weighted. That is, step 104 includes:
[0130] The translation weight of the sample image is determined based on the difference in translation position between the first local image and the second local image in the sample image; the scaling weight of the sample image is determined based on the cropping ratio of the first local image and the second local image in the sample image; the weight of the sample image is determined based on the translation weight and the scaling weight; and the loss value of the reference image reconstruction model is determined based on the difference between the first local image and the first reconstruction result, the difference between the second local image and the second reconstruction result, and the weight of the sample image.
[0131] It should be understood that since the first and second local images of the sample images input to the encoder unit are obtained by random cropping, their positions and scales are random. It is very likely that the first and second local images will be very different. In the early stage of network training, reconstructing from the first and second local images will lead to network degradation. Therefore, the loss corresponding to the input sample images can be weighted by the weights corresponding to the sample images. This way, the first and second local images with large differences are given low weights in the loss function, while the first and second local images with high similarity are given high weights. This can help improve the stability and accuracy of training.
[0132] The translational position difference between the first local image and the second local image in the corresponding sample image can be determined by taking the pixel at the top left corner of the sample image as the origin of the coordinate system, and the average of the normalized distance difference in the X direction and the normalized distance difference in the Y direction between the target pixel in the first local image and the corresponding target pixel in the second local image.
[0133] In one possible embodiment, the formula for calculating the translation weight corresponding to the sample image is:
[0134]
[0135] Among them, f T Let ΔTx be the translation weight corresponding to the sample image, ΔTy be the normalized distance difference in the X direction between the target pixel in the first local image and the corresponding target pixel in the second local image, and ΔTy be the normalized distance difference in the Y direction between the target pixel in the first local image and the corresponding target pixel in the second local image. τ T τ is a translation weighting coefficient that is greater than zero. TThe hyperparameters are determined experimentally. exp is an exponential function that is monotonically increasing. When the input is 0, the result is 1.
[0136] The cropping ratio of the first local image and the second local image in the corresponding sample image can be determined by taking the pixel at the top left corner of the sample image as the origin of the coordinate system and the average of the normalized difference between the cropping scales of the first local image and the second local image in the X direction and the normalized difference in the cropping scales in the Y direction.
[0137] In one possible embodiment, the scaling weights corresponding to the sample images are calculated using the following formula:
[0138]
[0139] Among them, f S ΔS represents the scaling weights corresponding to the sample images. x It is the difference between the first and second local images after cropping scale normalization in the X direction, ΔS y It is the difference between the first local image and the second local image after cropping scale normalization in the Y direction. S The scaling weight coefficient is greater than zero and is a hyperparameter determined experimentally.
[0140] In one possible embodiment, the formula for calculating the weights corresponding to the sample images is:
[0141] f = f T · s
[0142] Where f is the weight corresponding to the sample image.
[0143] It should be understood that if the translation positions and cropping scales of the first and second local images are completely identical, meaning they are exactly the same sample, then the weight corresponding to the sample image reaches its maximum, i.e., 1.0, which has a significant impact on the loss function. However, if the translation positions of the first and second local images differ greatly (i.e., ...), the weight will be significantly lower. (Very large) and the scaling scale varies greatly ( If the sample image weight is very large, the corresponding weight will be very small, and its impact on the loss function will be negligible.
[0144] Step 105: Update the network parameters of the reference image reconstruction model according to the loss value, and continue training using the updated reference image reconstruction model until the loss value of the updated reference image reconstruction model is less than the loss value threshold. Then, determine the encoder of the updated reference image reconstruction model as the trained image feature encoder.
[0145] The network parameters of the reference image reconstruction model include the network parameters of the encoder unit and the decoder unit.
[0146] Furthermore, the structure of the reference image reconstruction model is as follows: Figure 4 As shown, in the case where the first branch of the reference image reconstruction model is formed by the first encoder and the first decoder, and the second branch of the reference image reconstruction model is formed by the second encoder and the second decoder, in one possible embodiment, the network model structures of the two branches are completely identical. During the network initialization phase before training begins, the parameters of the two branches are the same after initialization. During training, in each training iteration cycle, the two input images are processed through their respective branches, and the final loss function is calculated. The gradients of the network parameters of the two branches are obtained according to the backpropagation algorithm, and the mean of the gradients of the network parameters of the two branches is calculated. Then, the mean gradient is used as the gradient of the network parameters of the two branches. The network parameters of the first and second branches are updated according to the gradients, and the next training iteration cycle is entered. The training is repeated until the loss value of the updated reference image reconstruction model is less than the loss value threshold. Then, the encoder of the updated reference image reconstruction model is determined as the image feature encoder that has been trained.
[0147] The trained image feature encoder can be used as a feature extractor, adapted to the corresponding classifier related to the task, and fine-tuned to make it applicable to downstream tasks such as image classification, detection and segmentation.
[0148] In one possible embodiment, the feature extractor can function as a feature extraction module in image classification. This feature extractor and classifier together form an image classification model. The image to be classified is input into the feature extractor for feature extraction, generating a feature image. This feature image is then input into the classifier for classification, yielding the classification result for the image. Because this feature extractor reconstructs a second reconstruction result corresponding to a second local image from a first local image, and a first reconstruction result corresponding to a first local image from a second local image, it more efficiently utilizes the contextual information of the image, providing a richer expression of image features. The extracted features are more representative, thereby improving classification accuracy.
[0149] In an image classification scenario, the network structure of a feature extractor can be used as the backbone to construct a classification network. For example, a fully connected layer with an n-dimensional output can be added to the last layer of the feature extractor, where n is the number of classification categories. The network parameters of the feature extractor are used to initialize the backbone of the classification network. After obtaining the classification network, it is trained using the corresponding image classification task data. The trained classification network can then be used for image classification. Each image to be classified is input into the classification network, and the classification results of each image are output. For example, this classification network is used to classify images of dogs and cats. When an animal image D is input into the classification network, the feature extractor of the classification network extracts features from the animal image D, extracting the feature image of the animal image D. The fully connected layer determines whether the animal in image D is a cat or a dog based on the feature image of the animal image D.
[0150] It should be understood that, compared to an uninitialized classification network, the classification network constructed using the feature extractor obtained through training in this application can achieve higher accuracy by initializing the network parameters, thereby resulting in more accurate classification results.
[0151] In an object detection scenario, the network structure of a feature extractor can be used as the backbone network to construct a detection network. For example, a two-stage object detection network, Faster R-CNN, adds a region proposal network and a detection head (including classification and regression branches) to the last layer of the feature extractor. The network parameters of the feature extractor are used to initialize the backbone network of the detection network. The detection network is trained using an object detection dataset. After training, the detection network can be used for object detection. The image to be detected is then input into the detection network to output the object detection result.
[0152] It should be understood that, compared to an uninitialized detection network, the detection network constructed using this feature extractor, when initialized with the network parameters of this feature extractor, can achieve higher accuracy, and thus the output detection results will also be more accurate.
[0153] The self-supervised training method for the image encoder described above acquires sample images and then randomly crops them to generate a first local image and a second local image. These first and second local images are then input into the first and second branches of a reference image reconstruction model for cross-reconstruction, generating a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. Based on the differences between the first local image and the first reconstruction result, and between the second local image and the second reconstruction result, the loss value of the reference image reconstruction model is determined. The network parameters of the reference image reconstruction model are updated based on this loss value, and training continues using the updated model until its loss value is less than a threshold. At this point, the encoder of the updated model is considered the successfully trained image feature encoder. Thus, by reconstructing the second reconstruction result corresponding to the second local image from the first local image, and reconstructing the first reconstruction result corresponding to the first local image from the second local image, the method more efficiently utilizes the contextual information of the image to express image features more richly, thereby improving task transfer capability and increasing the accuracy of downstream tasks.
[0154] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0155] In one embodiment, such as Figure 6 As shown, a self-supervised training device for an image encoder is provided. The device can be a software module or a hardware module, or a combination of both, as part of a computer device. The device specifically includes: a sample acquisition module 410, a cropping module 420, an image reconstruction module 430, a loss analysis module 440, and an encoder determination module 450.
[0156] The sample acquisition module 410 is used to acquire sample images.
[0157] The cropping module 420 is used to perform random cropping processing on the sample image to generate a first partial image and a second partial image of the sample image.
[0158] The image reconstruction module 430 is used to cross-reconstruct the first local image and the second local image with the corresponding input reference image reconstruction model to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image.
[0159] The loss analysis module 440 is used to determine the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result.
[0160] The encoder determination module 450 is used to update the network parameters of the reference image reconstruction model according to the loss value, and continue training using the updated reference image reconstruction model until the loss value of the updated reference image reconstruction model is less than the loss value threshold, and then determine the encoder of the updated reference image reconstruction model as the trained image feature encoder.
[0161] In one embodiment, the reference image reconstruction model includes an encoder unit and a decoder unit, and the image reconstruction module 430 is further configured to:
[0162] The first local image and the second local image are each divided into multiple image blocks. Preprocessing is performed on these image blocks to generate visible image blocks corresponding to the first and second local images. Each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image. Based on the positional relationship between the first and second local images, first relative position encoding information of the first local image relative to the second local image and second relative position encoding information of the second local image relative to the first local image are generated. Based on the first relative position encoding information and the second feature, a first feature to be decoded is generated. Based on the second relative position encoding information and the first feature, a second feature to be decoded is generated. The first feature to be decoded is input into the decoder unit for decoding to generate a first reconstruction result corresponding to the first local image. The second feature to be decoded is input into the decoder unit for decoding to generate a second reconstruction result corresponding to the second local image.
[0163] In one embodiment, the image reconstruction module 430 is further configured to: determine first relative position encoding information of the first local image relative to the second local image based on the relative position and size ratio between the first local image and the second local image and the position encoding information of the second local image; and determine second relative position encoding information of the second local image relative to the first local image based on the relative position and size ratio between the second local image and the first local image and the position encoding information of the first local image.
[0164] In one embodiment, the image reconstruction module 430 is further configured to: randomly generate a first mask token corresponding to the first local image based on the number of image blocks included in the first local image and the feature dimension of the second feature; concatenate the first mask token corresponding to the first local image with the second feature to generate a first feature sequence; and generate a first feature to be decoded based on the first relative position encoding information and the first feature sequence.
[0165] In one embodiment, the image reconstruction module 430 is further configured to: generate a first feature to be decoded based on the first relative position encoding information, the first feature sequence, and the position encoding information of the first local image.
[0166] In one embodiment, the image reconstruction module 430 is further configured to: randomly generate a second mask token corresponding to the second local image based on the number of image blocks included in the second local image and the feature dimension of the first feature; concatenate the second mask token corresponding to the second local image with the first feature to generate a second feature sequence; and generate a second feature to be decoded based on the second relative position encoding information and the second feature sequence.
[0167] In one embodiment, the image reconstruction module 430 is further configured to: generate a second feature to be decoded based on the second relative position encoding information, the second feature sequence, and the position encoding information of the second local image.
[0168] In one embodiment, the loss analysis module 440 is further configured to: determine the translation weight of the sample image based on the difference in translation position between the first local image and the second local image in the sample image; determine the scaling weight of the sample image based on the cropping ratio of the first local image and the second local image in the sample image; determine the weight of the sample image based on the translation weight and the scaling weight; and determine the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, the difference between the second local image and the second reconstruction result, and the weight of the sample image.
[0169] The self-supervised training device for the aforementioned image encoder acquires sample images and then randomly crops them to generate a first local image and a second local image. These two local images are then input into a reference image reconstruction model for cross-reconstruction, generating a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. Based on the differences between the first local image and the first reconstruction result, and between the second local image and the second reconstruction result, the loss value of the reference image reconstruction model is determined. The network parameters of the reference image reconstruction model are updated based on this loss value, and training continues using the updated model until its loss value is less than a threshold. At this point, the encoder of the updated reference image reconstruction model is considered the successfully trained image feature encoder. Thus, by reconstructing the second reconstruction result corresponding to the second local image from the first local image, and reconstructing the first reconstruction result corresponding to the first local image from the second local image, the device more efficiently utilizes the contextual information of the image to express image features more richly, thereby improving task transfer capability and increasing the accuracy of downstream tasks.
[0170] Specific limitations regarding the self-supervised training device for image encoders can be found in the limitations of the self-supervised training method for image encoders described above, and will not be repeated here. Each module in the aforementioned self-supervised training device for image encoders can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0171] Figure 7 This is a schematic diagram of the structure of a terminal device provided in one embodiment of this application. Figure 7 As shown, the terminal device 500 of this embodiment includes: at least one processor 510 ( Figure 7 (Only one is shown) a processor, a memory 520, and a computer program 521 stored in the memory 520 and capable of running on at least one processor 510. When the processor 510 executes the computer program 521, it implements the steps in the above-described self-supervised training method embodiment for the image encoder.
[0172] Terminal device 500 can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. This terminal device may include, but is not limited to, a processor 510 and a memory 520. Those skilled in the art will understand that... Figure 7This is merely an example of terminal device 500 and does not constitute a limitation on terminal device 500. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0173] The processor 510 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0174] In some embodiments, memory 520 may be an internal storage unit of terminal device 500, such as a hard disk or memory of terminal device 500. In other embodiments, memory 520 may be an external storage device of terminal device 500, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on terminal device 500. Furthermore, memory 520 may include both internal and external storage units of terminal device 500. Memory 520 is used to store operating system, applications, boot loader, data, and other programs, such as program code for computer programs. Memory 520 may also be used to temporarily store data that has been output or will be output.
[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0176] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0177] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0178] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0179] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0180] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0181] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0182] The implementation of all or part of the processes in the methods of the above embodiments can also be accomplished by a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the various method embodiments described above.
[0183] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A self-supervised training method for an image encoder, characterized in that, include: Acquire sample images; The sample image is randomly cropped to generate a first partial image and a second partial image of the sample image; The first local image and the second local image are input into a reference image reconstruction model for cross-reconstruction to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. The loss value of the reference image reconstruction model is determined based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result. The network parameters of the reference image reconstruction model are updated according to the loss value, and the updated reference image reconstruction model is used to continue training until the loss value of the updated reference image reconstruction model is less than the loss value threshold. Then the encoder of the updated reference image reconstruction model is determined as the trained image feature encoder. The reference image reconstruction model includes an encoder unit and a decoder unit. The step of inputting the first local image and the second local image into the reference image reconstruction model for cross-reconstruction to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image includes: The first local image and the second local image are each divided into multiple image blocks; Preprocessing is performed on multiple image blocks corresponding to the first local image and the second local image to generate each visible image block corresponding to the first local image and the second local image; Each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are respectively input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image; Based on the positional relationship between the first local image and the second local image, a first relative position encoding information of the first local image relative to the second local image and a second relative position encoding information of the second local image relative to the first local image are generated. Based on the first relative position encoding information and the second feature, a first feature to be decoded is generated; Based on the second relative position encoding information and the first feature, a second feature to be decoded is generated; The first feature to be decoded is input into the decoder unit for decoding processing to generate the first reconstruction result corresponding to the first local image; The second feature to be decoded is input into the decoder unit for decoding processing to generate the second reconstruction result corresponding to the second local image.
2. The method as described in claim 1, characterized in that, The step of generating first relative position encoding information of the first local image relative to the second local image and second relative position encoding information of the second local image relative to the first local image based on the positional relationship between the first local image and the second local image includes: Based on the relative position and size ratio between the first local image and the second local image, and the position encoding information of the second local image, the first relative position encoding information of the first local image relative to the second local image is determined; Based on the relative position and size ratio between the second local image and the first local image, and the position encoding information of the first local image, the second relative position encoding information of the second local image relative to the first local image is determined.
3. The method as described in claim 1, characterized in that, The step of generating a first feature to be decoded based on the first relative position encoding information and the second feature includes: Based on the number of image blocks included in the first local image and the feature dimension of the second feature, a first mask token corresponding to the first local image is randomly generated. The first mask token corresponding to the first local image is concatenated with the second feature to generate a first feature sequence; Based on the first relative position encoding information and the first feature sequence, a first feature to be decoded is generated.
4. The method as described in claim 3, characterized in that, The step of generating the first feature to be decoded based on the first relative position encoding information and the first feature sequence includes: A first feature to be decoded is generated based on the first relative position encoding information, the first feature sequence, and the position encoding information of the first local image.
5. The method as described in claim 1, characterized in that, The step of generating a second feature to be decoded based on the second relative position encoding information and the first feature includes: Based on the number of image blocks included in the second local image and the feature dimension of the first feature, a second mask token corresponding to the second local image is randomly generated. The second mask token corresponding to the second local image is concatenated with the first feature to generate a second feature sequence; A second feature to be decoded is generated based on the second relative position encoding information and the second feature sequence.
6. The method as described in claim 5, characterized in that, The step of generating the second feature to be decoded based on the second relative position encoding information and the second feature sequence includes: A second feature to be decoded is generated based on the second relative position encoding information, the second feature sequence, and the position encoding information of the second local image.
7. The method according to any one of claims 1-6, characterized in that, The step of determining the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result, includes: The translation weight of the sample image is determined based on the difference in translation position between the first local image and the second local image in the sample image; The scaling weight of the sample image is determined based on the cropping ratio of the first local image and the second local image in the sample image; The weights of the sample images are determined based on the translation weights and the scaling weights. The loss value of the reference image reconstruction model is determined based on the difference between the first local image and the first reconstruction result, the difference between the second local image and the second reconstruction result, and the weight of the sample image.
8. A self-supervised training device for an image encoder, characterized in that, include: The sample acquisition module is used to acquire sample images. The cropping module is used to randomly crop the sample image to generate a first partial image and a second partial image of the sample image; The image reconstruction module is used to cross-reconstruct the first local image and the second local image by correspondingly inputting them into the reference image reconstruction model, so as to generate a first reconstruction result corresponding to the first local image and a second reconstruction result corresponding to the second local image. The loss analysis module is used to determine the loss value of the reference image reconstruction model based on the difference between the first local image and the first reconstruction result, and the difference between the second local image and the second reconstruction result. The encoder determination module is used to update the network parameters of the reference image reconstruction model according to the loss value, and continue training using the updated reference image reconstruction model until the loss value of the updated reference image reconstruction model is less than the loss value threshold, and then determine the encoder of the updated reference image reconstruction model as the trained image feature encoder. The reference image reconstruction model includes an encoder unit and a decoder unit, and the image reconstruction module is specifically used for: The first local image and the second local image are each divided into multiple image blocks; Preprocessing is performed on multiple image blocks corresponding to the first local image and the second local image to generate each visible image block corresponding to the first local image and the second local image; Each visible image block corresponding to the first local image and each visible image block corresponding to the second local image are respectively input into the encoder unit to generate a first feature corresponding to the first local image and a second feature corresponding to the second local image; Based on the positional relationship between the first local image and the second local image, a first relative position encoding information of the first local image relative to the second local image and a second relative position encoding information of the second local image relative to the first local image are generated. Based on the first relative position encoding information and the second feature, a first feature to be decoded is generated; Based on the second relative position encoding information and the first feature, a second feature to be decoded is generated; The first feature to be decoded is input into the decoder unit for decoding processing to generate the first reconstruction result corresponding to the first local image; The second feature to be decoded is input into the decoder unit for decoding processing to generate the second reconstruction result corresponding to the second local image.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Model training method and device, electronic equipment and storage medium
CN114926338A