Model Training Method and Apparatus, Electronic Device, and Storage Medium

By stitching two sample images into mixed images and training the target encoder with image reconstruction model, the problem of low training efficiency and poor universality caused by image patches covered by special symbols in existing MIM tasks is solved, and a more efficient and general model training effect is achieved.

CN114926338BActive Publication Date: 2025-05-30SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210583591.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-05-30
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

In model training, existing MIM tasks use a large number of special symbols to fill the image patch masked by masks, resulting in inconsistency of inputs that have potential negative impacts on model training and consume a large amount of computing resources, reducing training efficiency and model universality.

Method used

A model training method is proposed, by stitching image blocks in two sample images into mixed images and using the image reconstruction encoder and decoder in the model, training the target encoder to learn visual representations and generate reconstructed images.

Benefits of technology

This method reduces the negative impact of input inconsistency on model training, reduces the waste of computing resources, improves model training efficiency and performance, and has high versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926338B_ABST
    Figure CN114926338B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, apparatus, electronic device, and storage medium. The method includes: obtaining a mixed image, which is an image formed by splicing image patches in two sample images; encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image; decoding the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images; training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder. Embodiments of the present disclosure can improve the overall model training efficiency and the performance of the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a model training method, an apparatus, an electronic device, and a storage medium. Background Art

[0002] In the related art, a Masked Image Modeling (MIM) task based on self-supervised learning of "visual representations (or visual expressions, image representations)" is proposed. Good "visual representations", that is, good encoded features, can provide information important for the task and ignore information irrelevant to the task. In the MIM task, first, the original image is segmented into non-overlapping image patches, then a random mask is used to cover some of the image patches, and the covered part of the image patches is filled with special symbols to obtain an image to be processed. Then, the encoder in the image reconstruction model is used to process the image to be processed to obtain an implicit "visual representation", and a lightweight decoder is used to generate a reconstructed image based on the "visual representation". Then, the relative mean square error between the reconstructed image and the original image is used as the reconstruction loss to train the encoder and the decoder. The encoder after multiple trainings can be used in the network model for downstream visual tasks.

[0003] In the current MIM task, although progress has been made in the model training of self-supervised learning of "visual representations", in order to make the task difficult enough, it is inevitable to use a large number of special symbols to fill the covered part of the image patches by the mask. However, these meaningless special symbols do not exist in real images. This difference in input will have a potential negative impact on model training, and a large amount of computing resources will be consumed on a large number of meaningless special symbols (that is, artificial inputs). This training method has a long training time, low training efficiency, and poor generality. Summary of the Invention

[0004] The present disclosure proposes a technical solution for model training.

[0005] According to an aspect of the present disclosure, there is provided a model training method, including: obtaining a mixed image, where the mixed image is an image formed by splicing image blocks in two sample images; encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image; decoding the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images; and training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder.

[0006] In a possible implementation, the encoder includes N sub-encoders, each sub-encoder including a multi-head attention mechanism layer, where N is a positive integer; wherein, encoding the mixed image through the encoder in a preset image reconstruction model to obtain the target feature map of the mixed image includes: determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder, and determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder; encoding the mixed image through the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain the target feature map of the mixed image; wherein, the attention mask is used to indicate the multi-head attention between the features of the same sample image calculated by the multi-head attention mechanism layer, and the attention window is used to indicate the multi-head attention between the features within the same attention window calculated by the multi-head attention mechanism layer.

[0007] In a possible implementation, determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention mask adopted by the first sub-encoder among the N sub-encoders according to the mask graph used when splicing the mixed image; downsampling the attention mask adopted by the (n - 1)-th sub-encoder according to the scale of the feature map encoded by the n-th sub-encoder among the N sub-encoders to obtain the attention mask adopted by the n-th sub-encoder, where 2 ≤ n ≤ N.

[0008] In a possible implementation, determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder according to the preset window size for the multi-head attention mechanism layer in each sub-encoder; wherein, the attention window includes at least one of an attention window for calculating global multi-head attention and an attention window for calculating local multi-head attention in a segmented manner.

[0009] In a possible implementation, encoding the mixed image by the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain the target feature map of the mixed image includes: converting the mixed image into an input vector of a specified dimension; encoding the input vector by a first sub-encoder according to the attention mask and attention window adopted in the first sub-encoder to obtain a first output feature map; downsampling the (n-1)th output feature map to obtain the (n-1)th input feature map with reduced resolution and increased number of channels; encoding the (n-1)th input feature map by the nth sub-encoder according to the attention mask and attention window adopted in the nth sub-encoder to obtain the nth output feature map, where 2 ≤ n ≤ N; and using the Nth output feature map encoded by the Nth sub-encoder as the target feature map.

[0010] In a possible implementation, converting the mixed image into an input vector of a specified dimension includes: performing channel unfolding and linear transformation on multiple image patches stitched into the mixed image to obtain a sequence vector; and embedding the position encoding vector corresponding to the mixed image into the sequence vector to obtain the input vector, where the position encoding vector is used to indicate the position information of each of the multiple image patches in the two sample images, and the position encoding vectors adopted for the image patches in different sample images are different.

[0011] In a possible implementation, decoding the target feature map by a decoder in the image reconstruction model to obtain two decoded reconstructed images includes: disassembling the target feature map into two sub-feature maps according to the attention mask adopted during encoding of the target feature map; and decoding the two sub-feature maps by the decoder to obtain two decoded reconstructed images.

[0012] In a possible implementation, the two sample images include a first sample image and a second sample image, and obtaining the mixed image includes: determining a first image patch extracted from the first sample image according to the first sample image and a preset mask map; determining a second image patch extracted from the second sample image according to the second sample image and the inverse mask map corresponding to the mask map; and stitching the first image patch and the second image patch to obtain the mixed image.

[0013] In a possible implementation, the two sample images include a first sample image and a second sample image, and the two reconstructed images include a first reconstructed image corresponding to the first sample image and a second reconstructed image corresponding to the second sample image. Among them, training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder includes: determining a first image difference between the first reconstructed image and the first sample image, and a second image difference between the second reconstructed image and the second sample image; determining the loss according to the first image difference and the second image difference, and training the image reconstruction model according to the loss to obtain a trained target encoder.

[0014] In a possible implementation, determining the first image difference between the first reconstructed image and the first sample image, and the second image difference between the second reconstructed image and the second sample image includes: determining the first image difference between the first reconstructed image and the first sample image according to a reverse mask graph used when extracting image patches from the second sample image; determining the second image difference between the second reconstructed image and the second sample image according to a mask graph used when extracting image patches from the first sample image; where the mask graph and the reverse mask graph are opposite to each other.

[0015] In a possible implementation, the target encoder is applied to a network model for a downstream task, and the downstream task includes at least one of object detection, image inpainting, image segmentation, and image classification.

[0016] According to one aspect of the present disclosure, a model training device is provided, including: an acquisition module for acquiring a mixed image, which is an image formed by splicing image patches in two sample images; an encoding module for encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image; a decoding module for decoding the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images; a training module for training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder.

[0017] In a possible implementation, the encoder includes N sub-encoders, each sub-encoder including a multi-head attention mechanism layer, where N is a positive integer; wherein, the encoding module includes: a determination sub-module, configured to determine the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder, and determine the attention window adopted by the multi-head attention mechanism layer in each sub-encoder; an encoding sub-module, configured to encode the mixed image through the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain a target feature map of the mixed image; wherein, the attention mask is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image, and the attention window is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features within the same attention window.

[0018] In a possible implementation, determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention mask adopted by the first sub-encoder among the N sub-encoders according to the mask map used when splicing the mixed image; downsampling the attention mask adopted by the (n - 1)-th sub-encoder according to the scale of the feature map encoded by the n-th sub-encoder among the N sub-encoders to obtain the attention mask adopted by the n-th sub-encoder, where 2 ≤ n ≤ N.

[0019] In a possible implementation, determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder according to the window size preset for the multi-head attention mechanism layer in each sub-encoder; wherein, the attention window includes at least one of an attention window for calculating global multi-head attention and an attention window for calculating local multi-head attention in a block manner.

[0020] In a possible implementation, encoding the mixed image by the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain the target feature map of the mixed image includes: converting the mixed image into an input vector of a specified dimension; encoding the input vector by the first sub-encoder according to the attention mask and attention window adopted in the first sub-encoder to obtain a first output feature map; downsampling the (n-1)th output feature map to obtain the (n-1)th input feature map with reduced resolution and increased number of channels; encoding the (n-1)th input feature map by the nth sub-encoder according to the attention mask and attention window adopted in the nth sub-encoder to obtain the nth output feature map, where 2 ≤ n ≤ N; and using the Nth output feature map encoded by the Nth sub-encoder as the target feature map.

[0021] In a possible implementation, converting the mixed image into an input vector of a specified dimension includes: performing channel unfolding and linear transformation on the multiple image patches that are stitched into the mixed image to obtain a sequence vector; and embedding the position encoding vector corresponding to the mixed image into the sequence vector to obtain the input vector, where the position encoding vector is used to indicate the position information of each of the multiple image patches in the two sample images, and the position encoding vectors adopted for the image patches in different sample images are different.

[0022] In a possible implementation, the decoding module includes: a disassembling sub-module, configured to disassemble the target feature map into two sub-feature maps according to the attention mask adopted during encoding of the target feature map; and a decoding sub-module, configured to decode the two sub-feature maps using the decoder to obtain two decoded reconstructed images.

[0023] In a possible implementation, the two sample images include a first sample image and a second sample image, and the obtaining module includes: a first extraction sub-module, configured to determine a first image patch extracted from the first sample image according to the first sample image and a preset mask map; a second extraction sub-module, configured to determine a second image patch extracted from the second sample image according to the second sample image and the inverse mask map corresponding to the mask map; and a stitching sub-module, configured to stitch the first image patch and the second image patch to obtain the mixed image.

[0024] In a possible implementation, the two sample images include a first sample image and a second sample image, and the two reconstructed images include a first reconstructed image corresponding to the first sample image and a second reconstructed image corresponding to the second sample image. Among them, the training module includes: a difference determination sub-module, configured to determine a first image difference between the first reconstructed image and the first sample image, and a second image difference between the second reconstructed image and the second sample image; a training sub-module, configured to determine the loss according to the first image difference and the second image difference, and train the image reconstruction model according to the loss to obtain a trained target encoder.

[0025] In a possible implementation, determining the first image difference between the first reconstructed image and the first sample image, and the second image difference between the second reconstructed image and the second sample image includes: determining the first image difference between the first reconstructed image and the first sample image according to a reverse mask graph used when extracting image patches from the second sample image; determining the second image difference between the second reconstructed image and the second sample image according to a mask graph used when extracting image patches from the first sample image; where the mask graph and the reverse mask graph are opposite to each other.

[0026] In a possible implementation, the target encoder is applied to a network model of a downstream task, and the downstream task includes at least one of object detection, image completion, image segmentation, and image classification.

[0027] According to one aspect of the present disclosure, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.

[0028] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0029] In the embodiments of the present disclosure, the encoder is used to extract the target feature map from the mixed image, which is equivalent to learning the "visual representation" in the mixed image. Then, the decoder is used to predict the reconstructed image based on the target feature map, which is equivalent to predicting the partial image patches in one sample image that are covered by the other sample image. Among them, since the mixed image is obtained by splicing the image patches in two sample images, that is, the mixed image input to the encoder comes from real sample images. Compared with using meaningless special symbols to fill the partial image patches covered by the mask, it can reduce the potential negative impact on model training caused by input inconsistency, while reducing the waste of computing resources, improving the overall model training efficiency and the performance of the trained model, and having high generality.

[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure.

[0032] Figure 1 The flowchart showing the model training method according to the embodiments of the present disclosure.

[0033] Figure 2 The schematic diagram showing the model structure of a Transformer block according to the embodiments of the present disclosure.

[0034] Figure 3 The schematic diagram showing the framework of an image reconstruction model according to the embodiments of the present disclosure.

[0035] Figure 4 The block diagram showing the model training apparatus according to the embodiments of the present disclosure.

[0036] Figure 5 The block diagram showing an electronic device 1900 according to the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0038] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" should not necessarily be construed as superior to or better than other embodiments.

[0039] As used herein, the term "and / or" is merely a description of the associated relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0040] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.

[0041] It is known that self-supervised learning is a type of unsupervised learning, mainly aiming to learn a general visual representation (i.e., feature expression) for downstream tasks. Self-supervised learning can be understood as pre-training a model to a preliminary form, that is, the pre-training process of the model; after pre-training the model to a certain extent, the model can be further trained to a complete form according to different labeled datasets of different downstream tasks, which can improve the overall training efficiency of the model and obtain better model performance.

[0042] The encoder in the embodiments of the present disclosure can be understood as a model that can extract general visual representations in images. However, as described above, in the masked image modeling (MIM) task used for pre-training the encoder currently, a large number of special symbols are used to fill the image patches (or image patches) covered by the mask, which will have a potential negative impact on model training, and the encoder will consume a large amount of computing resources on a large number of meaningless special symbols (i.e., artificial inputs), resulting in a long model training time, low training efficiency, and poor generality.

[0043] Based on the above problems, an embodiment of the present disclosure proposes a model training method, which can also be referred to as a hybrid masked image modeling method for efficiently learning visual representations. This method trains an encoder to learn visual representations by using a hybrid image formed by splicing and combining two random sample images from a training set, and uses a decoder to reconstruct the two sample images. From the perspective of a sample image, instead of replacing the mask symbols used to mask some image patches in the image with special symbols, this method uses the replacement mask symbols of another sample image, which can reduce the potential negative impact on model training caused by input inconsistency, while reducing the waste of computing resources, improving the overall model training efficiency and model performance, and having high generality.

[0044] In the model training method of the embodiment of the present disclosure, an encoder-decoder design is adopted, that is, the image reconstruction model includes an encoder and a decoder. Among them, the encoder can adopt the model structure of a hierarchical Vision Transformer (ViT). The encoder processes the hybrid image to obtain the visual representation hidden in the hybrid image, and the decoder reconstructs the two sample images (that is, generates two reconstructed images) based on the hidden visual representation; then, the image difference between the two reconstructed sample images and the original two sample images is used as the loss to train the encoder and the decoder. After multiple trainings, the decoder can be discarded, and the trained target encoder is retained for use in downstream tasks.

[0045] Figure 1 The flowchart showing the model training method according to the embodiment of the present disclosure. The model training method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by the processor calling the computer-readable instructions stored in the memory, or the method can be executed by the server. As Figure 1 shown, the model training method includes:

[0046] In step S11, a hybrid image is obtained.

[0047] Among them, the hybrid image is an image formed by splicing image patches in two sample images, and the sample images can be randomly selected images from an image set.

[0048] In a possible implementation, two sample images can be divided into multiple image patches according to a preset division size, and then some image patches can be randomly selected from the two sample images. The image patches randomly selected from the two sample images respectively (for example, 50% of the image patches can be randomly selected from the two sample images respectively) can be spliced and combined into a mixed image. It should be understood that the division size, that is, the size of the image patches, can be custom-set. For example, the division size can be set to 4×4 in terms of length×width. The embodiments of the present disclosure do not limit this.

[0049] In step S12, the encoder in the preset image reconstruction model is used to encode the mixed image to obtain the target feature map of the mixed image.

[0050] In a possible implementation, the encoder can, for example, adopt the model structure of a hierarchical Vision Transformer (ViT), or can also adopt the model structure of a hierarchical Vision Transformer based on moving windows (SwinTransformer). The embodiments of the present disclosure do not limit this. The encoder adopting the above model structure can extract better visual representations in the image, which is beneficial to better implementing various downstream tasks. It should be understood that the embodiments of the present disclosure do not limit the specific model structure of the encoder.

[0051] It can be known that the encoder adopting the above model structure usually includes multiple sub-encoders. There is a downsampling layer between every two sub-encoders. The downsampling layer is used to reduce the resolution of the input feature map (that is, the feature map input to the sub-encoder) by a factor of two and increase the number of channels of the input feature map by a factor of two. Each sub-encoder includes at least one group of Transformer blocks, and each group of Transformer blocks includes two Transformer blocks. Figure 2 The schematic diagram of the model structure of a Transformer block according to an embodiment of the present disclosure is shown as Figure 2 shown. Each Transformer block includes a multi-head attention (MSA) mechanism layer, a multi-layer perceptron (MLP, or multi-layer perceptron) layer, and a layer normalization (LN) layer. Residual connections are used between the MSA layer and the NLP layer, and an LN layer is connected between the MSA and the MLP.

[0052] Among them, the above encoder encodes the mixed image to obtain the target feature map of the mixed image. It can be understood that the mixed image is input into the above encoder, and after passing through multiple sub-encoders of the encoder, the target feature map is output. It should be understood that the specific encoding process and specific model structure of the encoder in the embodiments of the present disclosure are not limited.

[0053] In step S13, the decoder in the image reconstruction model decodes the target feature map to obtain two decoded reconstructed images.

[0054] It should be understood that those skilled in the art can customize the model structure of the decoder. For example, it can be designed as a decoder composed of multiple Transformer blocks, and the embodiments of the present disclosure are not limited thereto. The decoder can decode the target feature map to obtain two decoded reconstructed images, that is, the decoder can reconstruct two original sample images using the target feature map.

[0055] In step S14, according to the loss between the two reconstructed images and the two sample images, the image reconstruction model is trained to obtain the trained target encoder.

[0056] In a possible implementation manner, the loss between the two reconstructed images and the two sample images can be determined according to the image difference between the two reconstructed images and the two sample images. Among them, the image difference between the entire image area of each reconstructed image and the corresponding sample image can be determined, and the image difference between the local image area filled by the other sample image in each reconstructed image and the corresponding sample image can also be determined. The embodiments of the present disclosure are not limited thereto.

[0057] Among them, the image difference can adopt error functions such as mean square error and absolute error, and can also adopt distance functions such as L1 distance and L2 distance. In a possible implementation manner, determining the loss between the two reconstructed images and the two sample images according to the image difference between the two reconstructed images and the two sample images can include, for example: taking the sum of the two image differences between the two reconstructed images and their respective corresponding sample images as the loss between the two reconstructed images and the two sample images; of course, other calculation methods can also be used to determine the loss between the two reconstructed images and the two sample images, and the embodiments of the present disclosure are not limited thereto.

[0058] Among them, training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain the trained target encoder may include: adjusting the model parameters of the encoder and the model parameters of the decoder according to the loss between the two reconstructed images and the two sample images. It can be understood that during the model training process, the encoder and the decoder can be trained simultaneously, that is, when adjusting the model parameters of the encoder, the model parameters of the decoder can also be adjusted simultaneously. After multiple trainings, the decoder can be discarded, and the trained target encoder can be applied to downstream tasks.

[0059] It should be understood that model training usually includes multiple iterative trainings. Then, multiple mixed images can be used to perform the above steps S11 to S14 multiple times to iteratively train the image reconstruction model multiple times until the training end index is reached, and the target encoder in the trained image reconstruction model is obtained. The training end index can include, for example, loss convergence, loss set to 0, the number of iterations reaching the specified number of training times, etc. The embodiments of the present disclosure do not limit this.

[0060] In a possible implementation manner, the target encoder can be applied to the network model of the downstream task. The downstream task can include at least one of object detection, image completion, image segmentation, and image classification. Of course, it can also be applied to any other type of downstream task in this field. The embodiments of the present disclosure do not limit the type of the downstream task.

[0061] For example, when the target encoder is applied to the detection model of object detection, such as Faster RCNN (Faster Region Convolutional Neural Networks), the trained target encoder can be used as the feature extraction module in the Faster RCNN to extract the feature map of the image to be detected, that is, to extract the general image representation in the image to be detected. Then, the candidate anchor box generation module in the Faster RCNN can generate a large number of candidate anchor boxes based on the feature map. The candidate anchor box classification module in the Faster RCNN can classify the candidate anchor boxes to obtain the target anchor box containing the object to be detected. Then, the size and / or position of the target anchor box can be adjusted to obtain the detection box indicating the area where the object to be detected is located in the image to be detected, realizing object detection.

[0062] For another example, when the target encoder is applied to the classification model of image classification, such as a binary classification model, etc., the target encoder can be used as the feature extractor in the binary classification model to extract the feature map of the image to be classified, that is, to extract the general image representation in the image to be classified. Then, through the two classifiers in the binary classification model, classification is performed based on the feature map respectively to obtain the classification result of the image to be classified, realizing image classification.

[0063] Among them, when the above-mentioned target encoder is adopted in the network model of any downstream task, the network model adopting the above-mentioned target encoder can be further trained based on different labeled sample data adopted in different downstream tasks, and a fully formed network model in the downstream task can be obtained. In this way, the training efficiency of the network model in the entire downstream task can be improved, and the model performance of the trained network model can be improved.

[0064] In the embodiments of the present disclosure, by extracting the target feature map in the mixed image through the encoder, it is equivalent to learning the "visual representation" in the mixed image, and then using the decoder to respectively predict the reconstructed image based on the target feature map, which is equivalent to predicting the partial image patches in the two sample images that are covered by the other sample image. Among them, since the mixed image is obtained by splicing the image patches in the two sample images, that is, the mixed image input to the encoder comes from real sample images. Compared with using meaningless special symbols to fill the partial image patches covered by the mask, it can reduce the potential negative impact on model training caused by input inconsistency, while reducing the waste of computing resources, improving the overall model training efficiency and the model performance after training, and having high generality.

[0065] As described above, the mixed image is an image formed by splicing the image patches of two sample images. In a possible implementation manner, the two sample images include a first sample image and a second sample image. In step S11, obtaining the mixed image includes:

[0066] Determining a first image patch extracted from the first sample image according to the first sample image and a preset mask map; determining a second image patch extracted from the second sample image according to the second sample image and the inverse mask map corresponding to the mask map; splicing the first image patch and the second image patch to obtain the mixed image. In this way, a mixed image formed by splicing the image patches of two sample images can be effectively obtained.

[0067] Among them, the mask map can be a mask map composed of binary masks. The user can design the mask map according to the size of the image patches to be divided and the number of image patches to be extracted from the first sample image, or the number of image patches to be masked in the first sample image. The embodiments of the present disclosure do not limit this.

[0068] Among them, the reverse mask graph can be a mask graph composed of binary masks that are the opposite of the binary masks in the mask graph. For example, assuming the mask graph is represented as M, the reverse mask graph is represented as 1 - M, that is, the sum of the two binary masks at the same position in the mask graph and the reverse mask graph is 1. It should be understood that the second image patches extracted from the second sample image based on the reverse mask graph and the first image patches extracted from the first sample image based on the mask graph are complementary to each other. After the preset mask graph is known, the reverse mask graph is also known.

[0069] Among them, the image sizes of the first sample image and the second sample image can be the same or different, but the image patches extracted from the first sample image and the image patches extracted from the second sample image can be of the same size, so as to facilitate the splicing of the image patches extracted from the two sample images. Assume the first sample image is represented as X 1 , and the second sample image is represented as X 2 , the mask graph is represented as M, and the reverse mask graph is represented as 1 - M. The mixed image X m can be represented as X m = X 1 ⊙ M + X 2 ⊙ (1 - M), where ⊙ represents the Hadamard product.

[0070] In a possible implementation manner, splicing the first image patch and the second image patch to obtain a mixed image may include: filling the second image patch extracted from the second sample image into the local area covered by the mask graph in the first sample image to obtain a mixed image; or, alternatively, filling the first image patch extracted from the first sample image into the local area covered by the reverse mask graph in the second sample image to obtain a mixed image. The embodiments of the present disclosure do not limit this.

[0071] As described above, the encoder may include multiple sub-encoders. In a possible implementation manner, the encoder includes N sub-encoders, and each sub-encoder includes a multi-head attention mechanism layer, where N is a positive integer; among them, in step S12, encoding the mixed image through the encoder in the preset image reconstruction model to obtain the target feature map of the mixed image includes:

[0072] Step S121: Determine the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder, and determine the attention window adopted by the multi-head attention mechanism layer in each sub-encoder;

[0073] Step S122: Through N sub-encoders, encode the mixed image according to the attention masks and attention windows adopted in the N sub-encoders to obtain the target feature map of the mixed image.

[0074] As described above, the mixed image is an image formed by splicing image patches of two sample images, which means that the feature map encoded from the mixed image will contain the feature information of the two sample images. When the multi-head attention mechanism layer in the sub-encoder calculates the multi-head attention in the input feature map (or input vector), the feature information of one sample image is actually useless information for the other sample image. To calculate the effective multi-head attention in the input feature map (or input vector), the multi-head attention between the features belonging to the same sample image in the input feature map (or input vector) should be calculated.

[0075] Based on this, the multi-head attention mechanism layer in N sub-encoders can calculate the multi-head attention between the features belonging to the same sample image in the input feature map (or input vector) by combining the attention mask, which can enable the encoder to learn more effective "visual representations" in the mixed image, that is, to make the encoder output a target feature map containing more effective feature information. The attention mask is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image, that is, the attention mask can represent the features belonging to the two sample images in the input feature map (or input vector), so that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image.

[0076] Considering that if the global multi-head attention in the input feature map is calculated by each multi-head attention mechanism layer in each sub-encoder, it will generate a large amount of computation and the correlation between features that are far apart is not high. Therefore, an attention window can be set for the multi-head attention mechanism layer, and the attention window is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features within the same attention window. This can enable the multi-head attention mechanism layer to calculate the multi-head attention between the local features within the attention window in blocks, improving the processing efficiency and training efficiency of the encoder.

[0077] As described above, the specific model structure of the sub-encoder in the embodiments of the present disclosure is not limited. In addition to the multi-head attention mechanism layer in each sub-encoder, for example, it may also include an MLP layer, an LN layer, etc. Through N sub-encoders, according to the attention mask and attention window adopted in the N sub-encoders, the mixed image is encoded. It can be understood that after the N sub-encoders process the mixed image, and the multi-head attention mechanism layer of each sub-encoder is processed according to the attention mask and attention window adopted by the sub-encoder.

[0078] In the embodiments of the present disclosure, by determining the attention mask and attention window adopted by the multi-head attention mechanism in each sub-encoder, and then encoding the mixed image according to the attention mask and attention window, the encoder can learn more effective "visual representations" in the mixed image and improve the processing efficiency and training efficiency of the entire encoder.

[0079] As described above, a plurality of sub-encoders may be included in the encoder, and there is a downsampling layer between every two sub-encoders. The downsampling layer is used to multiply reduce the resolution of the input feature map (i.e., the feature map input to the sub-encoder) and multiply increase the number of channels of the input feature map. Then, the attention masks used in the multi-head attention mechanism layers in each sub-encoder should also gradually reduce the resolution correspondingly. In a possible implementation manner, in step S121, determining the attention masks used in the multi-head attention mechanism layers in each sub-encoder includes:

[0080] Determining the attention mask used by the first sub-encoder among the N sub-encoders according to the mask map used when splicing and mixing images; downsampling the attention mask used by the (n - 1)-th sub-encoder according to the scale of the feature map encoded by the n-th sub-encoder among the N sub-encoders to obtain the attention mask used by the n-th sub-encoder, where 2 ≤ n ≤ N. In this way, the attention masks used in the multi-head attention mechanism layers of each sub-encoder can be effectively obtained.

[0081] Among them, determining the attention mask used by the first sub-encoder among the N sub-encoders according to the mask map used when splicing and mixing images may include, for example: using the mask map used when extracting image patches from the first sample image as the attention mask used by the first sub-encoder; or using the inverse mask map of the mask map used when extracting image patches from the second sample image as the attention mask used by the first sub-encoder; or reconstructing the attention mask according to the mask map or the inverse mask map, etc. The embodiments of the present disclosure do not limit this.

[0082] Among them, the attention mask may adopt a binary mask or other types of masks, and the embodiments of the present disclosure do not limit this. It should be understood that the mask map used when splicing and mixing images may indicate the image patches extracted from the first sample image or may inversely indicate the image patches extracted from the second sample image. Therefore, the attention mask determined based on this mask map can represent the image patches in the mixed image that belong to different sample images respectively, and of course, can also represent the feature vectors in the input vector of the first sub-encoder that belong to different sample images respectively.

[0083] As described above, there is a downsampling layer between every two sub-encoders. The downsampling layer is used to multiply the resolution of the input feature map (i.e., the feature map input to the sub-encoder) and multiply the number of channels of the input feature map. The scale of the feature map can include the resolution of the input feature map (i.e., length and width, or size) and / or the number of channels (or depth). Then, the scale of the feature map encoded by the nth sub-encoder can be understood as the scale of the feature map output by a downsampling layer before the nth sub-encoder, that is, the scale of the feature map input to the nth sub-encoder.

[0084] Among them, since the downsampling layer between every two sub-encoders is used to multiply the resolution of the input feature map, the resolution of the attention mask obtained by downsampling the attention mask adopted by the (n - 1)th sub-encoder is also multiplied.

[0085] As described above, an attention window can be set for the multi-head attention mechanism layer, so that the multi-head attention mechanism layer can calculate the multi-head attention between local features within the attention window in blocks, improving the processing efficiency and training efficiency of the encoder. In a possible implementation manner, in step S121, determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder includes:

[0086] Determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder according to the window size preset for the multi-head attention mechanism layer in each sub-encoder; among them, the attention window includes at least one of the attention window for calculating global multi-head attention and the attention window for calculating local multi-head attention in blocks. By this method, compared with the related art that uses a complex cyclic shifted attention window for global attention modeling, using the attention window with the above preset window size can not only achieve global attention modeling but also reduce the complexity of the encoder, improving the processing efficiency and training efficiency of the encoder.

[0087] As described above, the encoder may include multiple sub-encoders, and each sub-encoder may include at least two multi-head attention mechanism layers. The window size of the attention window adopted by the multi-head attention mechanism window can be set for each multi-head attention mechanism layer in each sub-encoder.

[0088] Among them, when the window size of the attention window set in a certain multi-layer attention mechanism layer is the same as the size of the input feature map (i.e., resolution, length, and width), this attention window is the attention window used to calculate global multi-head attention, and the calculation range of global multi-head attention is the features in the entire input feature map (or input vector); when the window size of the attention window set in a certain multi-layer attention mechanism layer is different from the size of the input feature map, this attention window is the attention window used to calculate local multi-head attention in blocks, and the calculation range of calculating local multi-head attention in blocks is the features within the attention window.

[0089] It should be understood that those skilled in the art can set the window size adopted by each multi-head attention mechanism layer in the encoder according to actual needs. For example, assuming that there are 20 multi-head attention mechanism layers in the entire encoder, it can be set that the 6th, 12th, and 18th multi-head attention mechanism layers adopt global attention windows (i.e., attention windows used to calculate global multi-head attention), and the other 17 multi-head attention mechanism layers all adopt attention windows for calculating local multi-head attention in blocks. The window size of the attention window for calculating local multi-head attention in blocks can be set to 14×14 or 7×7, etc., and the embodiments of the present disclosure do not limit this.

[0090] It should be noted that the window sizes set for the above 20 multi-head attention mechanism layers are one implementation manner disclosed in the embodiments of the present disclosure. In fact, under the inspiration of the embodiments of the present disclosure, those skilled in the art can customize various window sizes, as long as the method of calculating multi-head attention by setting the attention window in the embodiments of the present disclosure is within the protection scope of the present disclosure.

[0091] As described above, the encoder may include N sub-encoders, and there is a downsampling layer between every two sub-encoders. The downsampling layer is used to reduce the resolution of the input feature map (i.e., the feature map input to the sub-encoder) by a factor and increase the number of channels of the input feature map by a factor. The multi-head attention mechanism layer of each sub-encoder processes the input feature map (or input vector) according to the attention mask and attention window adopted by the sub-encoder. In a possible implementation manner, in step S122, through N sub-encoders, according to the attention masks and attention windows adopted in the N sub-encoders, the mixed image is encoded to obtain the target feature map of the mixed image, including:

[0092] Convert the mixed image into an input vector of a specified dimension; through the first sub-encoder, encode the input vector according to the attention mask and attention window adopted in the first sub-encoder to obtain the first output feature map; downsample the (n-1)th output feature map to obtain the (n-1)th input feature map with reduced resolution and increased number of channels; through the nth sub-encoder, encode the (n-1)th input feature map according to the attention mask and attention window adopted in the nth sub-encoder to obtain the nth output feature map, where 2 ≤ n ≤ N; use the Nth output feature map obtained through the Nth sub-encoder as the target feature map. In this way, the encoded target feature map can be effectively obtained.

[0093] Among them, converting the mixed image into an input vector of a specified dimension can be understood as converting the mixed image into an input vector of a specified dimension that can be input into the first sub-encoder. In a possible implementation, converting the mixed image into an input vector of a specified dimension includes: performing channel unfolding and linear transformation on multiple image patches that are stitched into the mixed image to obtain a sequence vector; embedding the position encoding vector corresponding to the mixed image into the sequence vector to obtain the input vector, where the position encoding vector is used to indicate the position information of each of the multiple image patches in two sample images, and the position encoding vectors adopted by the image patches in different sample images are different. In this way, an input vector containing the position information of each of the multiple image patches that are stitched into the mixed image in two sample images can be obtained, which is convenient for the subsequent encoder and decoder to effectively utilize the position information in the input vector for encoding and decoding.

[0094] For example, in the above process of obtaining the sequence vector, assume that the image size of a three-channel RGB mixed image is H×W, and each image patch has 4×4 = 16 pixels. Since each pixel in the mixed image has three values of R, G, and B, then unfolding the multiple image patches that are stitched into the mixed image in the channel dimension (i.e., channel unfolding) can obtain an image, and then performing a linear transformation (or linear mapping, which is used to map the data into a vector space of a specified dimension) on each pixel in the image in the channel dimension to obtain a sequence vector with a specified dimension C, that is, obtaining a sequence vector.

[0095] It can be understood that since the position encoding vector can indicate the position information of each of multiple image patches in two sample images, and the position encoding vectors adopted by the image patches in different sample images are different, embedding the position encoding vector into the sequence vector can enable the encoder to know which sample image different features belong to during the process of encoding the mixed image, and can also enable the decoder to reconstruct the two sample images based on the position information indicated by the position encoding vector, that is, generate two reconstructed images corresponding to the two sample images. Among them, those skilled in the art can calculate the position encoding vector corresponding to the above-mentioned mixed image by using position encoding methods known in the art, such as relative position encoding or absolute position encoding, etc., and the embodiments of the present disclosure do not limit this.

[0096] As described above, the first image patch extracted from the first sample image can be spliced with the second image patch extracted from the second sample image to obtain a mixed image, which means that the multiple image patches spliced into the mixed image can include the first image patch and the second image patch.

[0097] It should be understood that the input vector obtained by embedding the position encoding vector into the sequence vector has the same scale as the sequence vector. For example, when embedding the position encoding vector into the sequence vector, an input vector can be obtained. Among them, those skilled in the art can adopt vector embedding methods known in the art to implement embedding the position encoding vector into the sequence vector to obtain the input vector, and the embodiments of the present disclosure do not limit this.

[0098] Continuing with the above input vector, for example, in the above process of obtaining the target feature map, assuming that there are 4 sub-encoders in the encoding, through the first sub-encoder, according to the attention mask and attention window adopted in the first sub-encoder, the input vector is encoded to obtain the first output feature map; downsampling the first output feature map can obtain the first input feature map with reduced resolution and increased number of channels, that is, obtain the first input feature map; through the second sub-encoder, according to the attention mask and attention window adopted in the second sub-encoder, the first input feature map is encoded to obtain the second output feature map, and then downsampling the second output feature map can obtain the second input feature map, and then through the third sub-encoder, according to the attention mask and attention window adopted in the third sub-encoder, the second input feature map is encoded to obtain The third output feature map, and so on, the fourth output feature map encoded by the fourth sub-encoder can be obtained; The fourth output feature map; the fourth output feature map obtained through the fourth sub-encoder can be used as the target feature map. The fourth output feature map of can be used as the target feature map.

[0099] As described above, the attention mask can represent the features belonging to the two sample images in the input feature map, and can also represent the features belonging to the two sample images in the output feature map. Then, the attention mask used when encoding the target feature map can represent the features belonging to the two sample images in the target feature map. To facilitate the decoder to efficiently generate the two reconstructed images corresponding to the two sample images respectively, in a possible implementation manner, in step S13, the decoder in the image reconstruction model decodes the target feature map to obtain the two decoded reconstructed images, including:

[0100] According to the attention mask used when encoding the target feature map, the target feature map is disassembled into two sub-feature maps; the decoder decodes the two sub-feature maps to obtain the two decoded reconstructed images. In this way, the decoder can efficiently generate the two reconstructed images corresponding to the two sample images respectively.

[0101] Among them, the attention mask used when encoding the target feature map is also the attention mask used by the Nth sub-encoder. Since the attention mask used when encoding the target feature map can represent the features belonging to the two sample images in the target feature map, then according to the attention mask used when encoding the target feature map, the two sub-feature maps disassembled from the target feature map respectively contain the features belonging to the two sample images. In this way, the decoder can decode the two sub-feature maps respectively to obtain the two reconstructed images corresponding to the two sample images respectively. Among them, the specific model structure and specific decoding process of the decoder in the embodiments of the present disclosure are not limited.

[0102] In a possible implementation manner, a linear transformation layer can also be added between the decoder and the encoder based on the model structure of the decoder to map the target feature map output by the encoder into the dimension required by the decoder, for example, convert it into a feature map with 512 channels.

[0103] As described above, the loss between the two reconstructed images and the two sample images can be determined based on the image differences between the two reconstructed images and the two sample images, and the image reconstruction model can be trained based on this loss. In a possible implementation, the two sample images include a first sample image and a second sample image, and the two reconstructed images include a first reconstructed image corresponding to the first sample image and a second reconstructed image corresponding to the second sample image. Among them, in step S14, the image reconstruction model is trained according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder, including:

[0104] Step S141: Determine a first image difference between the first reconstructed image and the first sample image, and a second image difference between the second reconstructed image and the second sample image.

[0105] As described above, the image difference between the entire image regions of each reconstructed image and the corresponding sample image can be determined, that is, the first image difference between the entire image regions of the first reconstructed image and the first sample image, and the second image difference between the entire image regions of the second reconstructed image and the second sample image can be determined. The first image difference and the second image difference can, for example, adopt error functions such as mean square error and absolute error, and can also adopt distance functions such as L1 distance and L2 distance. The embodiments of the present disclosure do not limit this.

[0106] As described above, the decoder generates two reconstructed images based on two sub-feature maps disassembled from the target feature map. That is to say, the decoder actually reconstructs the complete sample image based on the partial features of the sample image contained in each sub-feature map. Then, the image difference between each reconstructed image and the local image region filled with the other sample image in the corresponding sample image can also be determined. It should be understood that the smaller the image difference between the local image regions, the more effective visual representation information in the partial features belonging to the sample image in the mixed image can be extracted by the encoder. In this way, the loss determined based on the image difference between the local image regions can be used to train the encoder to extract more effective visual representations in the image, or to extract more effective feature information in the image, thereby improving the model performance of the trained target encoder.

[0107] In a possible implementation, determining a first image difference between a first reconstructed image and a first sample image, and a second image difference between a second reconstructed image and a second sample image includes: determining the first image difference between the first reconstructed image and the first sample image according to a reverse mask map used when extracting image patches from the second sample image; determining the second image difference between the second reconstructed image and the second sample image according to a mask map used when extracting image patches from the first sample image; wherein the mask map and the reverse mask map are opposite to each other. In this way, the image difference between each reconstructed image and the local image region filled by the other sample image in the corresponding sample image can be effectively obtained.

[0108] It should be understood that the reverse mask map used when extracting image patches from the second sample image can indicate the local image region in the first sample image that will be filled by the image patches of the second sample image; correspondingly, the mask map used when extracting image patches from the first sample image can indicate the local image region in the second sample image that will be filled by the image patches of the first sample image. Therefore, the first image difference between the first reconstructed image and the local image region in the first sample image can be determined based on the reverse mask map, and the second image difference between the second reconstructed image and the local image region in the second sample image can be determined based on the mask map.

[0109] For example, assume the first sample image is X 1 The corresponding first reconstructed image is denoted as Y 1 , the second sample image is X 2 The corresponding second reconstructed image is denoted as Y 2 , the first image difference can be expressed as (Y 1 - X 1 ) ⊙ (1 - M), and the second image difference can be expressed as (Y 2 - X 2 ) ⊙ M, where M represents the mask map and (1 - M) represents the reverse mask map.

[0110] Step S142: Determine a loss according to the first image difference and the second image difference, and train an encoder according to the loss to obtain a target encoder.

[0111] As described above, the sum between the first image difference and the second image difference can be used as the loss between the two reconstructed images and the two sample images. In a possible implementation, following the above example, the loss L between the two reconstructed images and the two sample images can be determined with reference to formula (1).

[0112]

[0113] Wherein, Denotes the square of the 2-norm, and ⊙ denotes the Hadamard product.

[0114] As described above, model training usually includes multiple iterative trainings. Then, multiple mixed images can be used to execute the above steps S11 to S14 (including steps S141 to S142) multiple times to iteratively train the encoder and the decoder multiple times until the training end index is reached, and the trained target encoder is obtained. The training end index can include, for example, loss convergence, loss set to 0, the number of iterations reaching the specified number of training times, etc. The embodiments of the present disclosure do not limit this.

[0115] In the embodiments of the present disclosure, by determining the image difference between each reconstructed image and the local image region filled with another sample image in the corresponding sample image, the loss determined based on this image difference can be used to train the encoder to extract more effective visual representations in the image, or more effective feature information in the image, so as to improve the model performance of the trained target encoder.

[0116] Figure 3 Shows a schematic framework diagram of an image reconstruction model according to an embodiment of the present disclosure, as Figure 3 shown, the image reconstruction model includes: an encoder and a decoder. The encoder includes four sub-encoders, and there is a downsampling layer between every two sub-encoders, and there is an image conversion layer before the first sub-encoder.

[0117] As Figure 3 shown, the mixed image is an image formed by splicing image patches in two sample images according to a mask map. This mask map can be used as the attention mask adopted by the first sub-encoder at the same time, and then the attention mask adopted by the first sub-encoder is downsampled step by step to obtain the attention mask adopted by each sub-encoder;

[0118] As Figure 3 shown, the image conversion layer can be used to convert the mixed image into an input vector of a specified dimension, that is, the multiple image patches spliced into the mixed image can be channel-expanded and linearly transformed to obtain a sequence vector, and the position encoding vector corresponding to the mixed image is embedded into the sequence vector to obtain an input vector; then the first sub-encoder is used to process the input vector to obtain a first output feature map, and then the downsampling layer is used to downsample the first output feature map to obtain a first output feature map, and then so on to obtain the target feature map output by the fourth sub-encoder.

[0119] As Figure 3As shown, according to the attention mask used when encoding the target feature map (i.e., the attention mask used by the fourth sub-encoder), the target feature map is disassembled into two sub-feature maps; then the decoder is used to decode the two sub-feature maps to obtain two decoded reconstructed images.

[0120] Training the model to obtain the target encoder according to the image reconstruction model of the present disclosure embodiment can reduce the negative impact of artificial input (i.e., filling the image blocks covered by the mask with meaningless special symbols) on model training, reduce redundant calculations, reduce waste of computing resources, obtain higher training efficiency and model performance; it can also be used for training homogeneous structure models and pyramid structure models, has high versatility, and the model complexity is low, which is beneficial to improving the processing efficiency of the model.

[0121] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0122] In addition, the present disclosure also provides a model training device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any model training method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated here.

[0123] Figure 4 The block diagram showing the model training device according to the embodiment of the present disclosure is as Figure 4 As shown, the device includes:

[0124] An acquisition module 101, configured to acquire a mixed image, where the mixed image is an image formed by splicing image blocks in two sample images;

[0125] An encoding module 102, configured to encode the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image;

[0126] A decoding module 103, configured to decode the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images;

[0127] A training module 104, configured to train the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder.

[0128] In a possible implementation, the encoder includes N sub-encoders, each sub-encoder includes a multi-head attention mechanism layer, where N is a positive integer; among them, the encoding module 102 includes: a determination sub-module, configured to determine the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder, and determine the attention window adopted by the multi-head attention mechanism layer in each sub-encoder; an encoding sub-module, configured to encode the mixed image through the N sub-encoders according to the attention mask and attention window adopted in the N sub-encoders, to obtain the target feature map of the mixed image; where the attention mask is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image, and the attention window is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features within the same attention window.

[0129] In a possible implementation, determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention mask adopted by the first sub-encoder among the N sub-encoders according to the mask map used when splicing the mixed image; downsampling the attention mask adopted by the (n - 1)-th sub-encoder according to the scale of the feature map encoded by the n-th sub-encoder among the N sub-encoders, to obtain the attention mask adopted by the n-th sub-encoder, where 2 ≤ n ≤ N.

[0130] In a possible implementation, determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder according to the window size preset for the multi-head attention mechanism layer in each sub-encoder; where the attention window includes at least one of an attention window for calculating global multi-head attention and an attention window for calculating local multi-head attention in a segmented manner.

[0131] In a possible implementation, the encoding of the mixed image by the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain the target feature map of the mixed image includes: converting the mixed image into an input vector of a specified dimension; encoding the input vector by the first sub-encoder according to the attention mask and attention window adopted in the first sub-encoder to obtain a first output feature map; downsampling the (n - 1)-th output feature map to obtain the (n - 1)-th input feature map with reduced resolution and increased number of channels; encoding the (n - 1)-th input feature map by the n-th sub-encoder according to the attention mask and attention window adopted in the n-th sub-encoder to obtain the n-th output feature map, where 2 ≤ n ≤ N; and using the N-th output feature map encoded by the N-th sub-encoder as the target feature map.

[0132] In a possible implementation, the converting of the mixed image into an input vector of a specified dimension includes: performing channel unfolding and linear transformation on the multiple image patches that are stitched into the mixed image to obtain a sequence vector; and embedding the position encoding vector corresponding to the mixed image into the sequence vector to obtain the input vector, where the position encoding vector is used to indicate the position information of each of the multiple image patches in the two sample images, and the position encoding vectors adopted for the image patches in different sample images are different.

[0133] In a possible implementation, the decoding module 103 includes: a disassembling sub-module, configured to disassemble the target feature map into two sub-feature maps according to the attention mask adopted during the encoding of the target feature map; and a decoding sub-module, configured to decode the two sub-feature maps by using the decoder to obtain two decoded reconstructed images.

[0134] In a possible implementation, the two sample images include a first sample image and a second sample image, and the obtaining module 101 includes: a first extraction sub-module, configured to determine a first image patch extracted from the first sample image according to the first sample image and a preset mask map; a second extraction sub-module, configured to determine a second image patch extracted from the second sample image according to the second sample image and the inverse mask map corresponding to the mask map; and a stitching sub-module, configured to stitch the first image patch and the second image patch to obtain the mixed image.

[0135] In a possible implementation, the two sample images include a first sample image and a second sample image, and the two reconstructed images include a first reconstructed image corresponding to the first sample image and a second reconstructed image corresponding to the second sample image. Among them, the training module 104 includes: a difference determination sub-module, configured to determine a first image difference between the first reconstructed image and the first sample image, and a second image difference between the second reconstructed image and the second sample image; a training sub-module, configured to determine the loss according to the first image difference and the second image difference, and train the image reconstruction model according to the loss to obtain a trained target encoder.

[0136] In a possible implementation, the determining the first image difference between the first reconstructed image and the first sample image, and the second image difference between the second reconstructed image and the second sample image includes: determining the first image difference between the first reconstructed image and the first sample image according to a reverse mask graph used when extracting image patches from the second sample image; determining the second image difference between the second reconstructed image and the second sample image according to a mask graph used when extracting image patches from the first sample image; where the mask graph and the reverse mask graph are opposite to each other.

[0137] In a possible implementation, the target encoder is applied to a network model for a downstream task, and the downstream task includes at least one of object detection, image completion, image segmentation, and image classification.

[0138] In the embodiments of the present disclosure, extracting the target feature map in the mixed image through the encoder is equivalent to learning the "visual representation" in the mixed image, and then using the decoder to respectively predict the reconstructed images based on the target feature map, which is equivalent to predicting the partial image patches in the two sample images that are covered by the other sample image. Among them, since the mixed image is obtained by splicing the image patches in the two sample images, that is, the mixed image input to the encoder comes from real sample images. Compared with using meaningless special symbols to fill the partial image patches covered by the mask, it can reduce the potential negative impact on model training caused by input inconsistency, while reducing the waste of computing resources, improving the overall model training efficiency and the performance of the trained model, and having high generality.

[0139] This method has a specific technical association with the internal structure of the computer system and can solve the technical problems of how to improve the hardware operation efficiency or execution effect (including reducing the data storage volume, reducing the data transmission volume, and increasing the hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in line with the natural law.

[0140] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure may be used to execute the methods described in the above method embodiments. The specific implementation may refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0141] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0142] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.

[0143] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.

[0144] The electronic device may be provided as a terminal, a server or other forms of devices.

[0145] Figure 5 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 5 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.

[0146] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.

[0147] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the above computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.

[0148] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0149] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0150] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0151] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0152] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0153] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0154] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0155] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] The computer program product can be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0157] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they will not be elaborated herein.

[0158] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0159] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the independent consent of the individual. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0160] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A model training method, characterized in that, the method includes: obtaining a mixed image, which is an image formed by splicing image patches in two sample images; encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image; decoding the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images; training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder; wherein, the encoder includes N sub-encoders, each sub-encoder includes a multi-head attention mechanism layer, and N is a positive integer; wherein, the encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image includes: determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder, and determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder; encoding the mixed image through the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain a target feature map of the mixed image; wherein, the attention mask is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image, and the attention window is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features within the same attention window.

2. The method according to claim 1, characterized in that, the determining the attention mask adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention mask adopted by the first sub-encoder among the N sub-encoders according to the mask map used when splicing the mixed image; downsampling the attention mask adopted by the (n - 1)-th sub-encoder according to the scale of the feature map encoded by the n-th sub-encoder among the N sub-encoders to obtain the attention mask adopted by the n-th sub-encoder, where 2 ≤ n ≤ N.

3. The method according to claim 1 or 2, characterized in that, the determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder includes: determining the attention window adopted by the multi-head attention mechanism layer in each sub-encoder according to the window size preset for the multi-head attention mechanism layer in each sub-encoder; wherein, the attention window includes at least one of an attention window for calculating global multi-head attention and an attention window for calculating local multi-head attention in blocks.

4. The method according to claim 1 or 2, characterized in that, the encoding the mixed image through the N sub-encoders according to the attention masks and attention windows adopted in the N sub-encoders to obtain a target feature map of the mixed image includes: converting the mixed image into an input vector of a specified dimension; Through the first sub-encoder, encode the input vector according to the attention mask and attention window adopted in the first sub-encoder to obtain the first output feature map; Downsample the (n-1)th output feature map to obtain the (n-1)th input feature map with reduced resolution and increased number of channels; Through the nth sub-encoder, encode the (n-1)th input feature map according to the attention mask and attention window adopted in the nth sub-encoder to obtain the nth output feature map, where 2 ≤ n ≤ N; Use the Nth output feature map encoded by the Nth sub-encoder as the target feature map.

5. The method according to claim 4, wherein, The converting the mixed image into an input vector of a specified dimension includes: Perform channel unfolding and linear transformation on multiple image patches that are stitched into the mixed image to obtain a sequence vector; Embed the position encoding vector corresponding to the mixed image into the sequence vector to obtain an input vector, where the position encoding vector is used to indicate the position information of each of the multiple image patches in the two sample images, and the position encoding vectors adopted by the image patches in different sample images are different.

6. The method according to claim 1 or 2, wherein, The decoding the target feature map through the decoder in the image reconstruction model to obtain two decoded reconstructed images includes: Decompose the target feature map into two sub-feature maps according to the attention mask adopted when encoding the target feature map; Use the decoder to decode the two sub-feature maps to obtain two decoded reconstructed images.

7. The method according to claim 1 or 2, wherein, The two sample images include a first sample image and a second sample image, and the obtaining the mixed image includes: Determine a first image patch extracted from the first sample image according to the first sample image and a preset mask map; Determine a second image patch extracted from the second sample image according to the second sample image and the inverse mask map corresponding to the mask map; Stitch the first image patch and the second image patch to obtain the mixed image.

8. The method according to claim 1 or 2, wherein, The two sample images include a first sample image and a second sample image, and the two reconstructed images include a first reconstructed image corresponding to the first sample image and a second reconstructed image corresponding to the second sample image, wherein, the training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder includes: Determine a first image difference between the first reconstructed image and the first sample image, and a second image difference between the second reconstructed image and the second sample image; Determine the loss according to the first image difference and the second image difference, and train the image reconstruction model according to the loss to obtain a trained target encoder.

9. The method according to claim 8, wherein, Determining the first image difference between the first reconstructed image and the first sample image, and the second image difference between the second reconstructed image and the second sample image includes: Determining the first image difference between the first reconstructed image and the first sample image according to the inverse mask map used when extracting image patches from the second sample image; Determining the second image difference between the second reconstructed image and the second sample image according to the mask map used when extracting image patches from the first sample image; wherein the mask map and the inverse mask map are opposite to each other.

10. The method according to claim 1 or 2, characterized in that the target encoder is applied to a network model for a downstream task, and the downstream task includes at least one of object detection, image inpainting, image segmentation, and image classification.

11. A model training device, characterized in that it includes: an acquisition module for acquiring a mixed image, which is an image formed by splicing image patches in two sample images; an encoding module for encoding the mixed image through an encoder in a preset image reconstruction model to obtain a target feature map of the mixed image; a decoding module for decoding the target feature map through a decoder in the image reconstruction model to obtain two decoded reconstructed images; a training module for training the image reconstruction model according to the loss between the two reconstructed images and the two sample images to obtain a trained target encoder; wherein the encoder includes N sub-encoders, each sub-encoder includes a multi-head attention mechanism layer, and N is a positive integer; wherein, the encoding module includes: a determination sub-module for determining the attention mask used by the multi-head attention mechanism layer in each sub-encoder and determining the attention window used by the multi-head attention mechanism layer in each sub-encoder; an encoding sub-module for encoding the mixed image through the N sub-encoders according to the attention mask and attention window used in the N sub-encoders to obtain a target feature map of the mixed image; wherein, the attention mask is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features of the same sample image, and the attention window is used to indicate that the multi-head attention mechanism layer calculates the multi-head attention between the features within the same attention window.

12. An electronic device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, on which computer program instructions are stored, characterized in that when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Model training method, image processing and registration method and related devices and equipment

    CN112348819A

  • Mask-based image deblurring model and method

    CN113538258A