Method of training vision transformer network, electronic device, and computer-program product

The method enhances the training of vision transformer networks by using a first encoder and multiple modules with random masks and model distillation, addressing the challenges of unsupervised learning and improving feature representation and robustness.

WO2025145324A1PCT designated stage expired Publication Date: 2025-07-10BOE TECHNOLOGY GROUP CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/070329
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing vision transformer networks face challenges in effectively training without labeled data and in enhancing feature representation and robustness for unsupervised learning tasks.

Method used

A method is introduced for training a vision transformer network using a first encoder to extract features, followed by multiple modules that include regression layers and encoders, employing random masks to filter features, and calculating losses based on predicted and masked regions, along with model distillation techniques to enhance feature representation and robustness.

Benefits of technology

The method improves the feature extraction and robustness of vision transformer networks, enabling effective unsupervised learning and adaptation to specific tasks without labeled data, enhancing the network's ability to capture and represent visual features comprehensively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024070329_10072025_PF_FP_ABST
    Figure CN2024070329_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of training a vision transformer network includes receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; and calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD OF TRAINING VISION TRANSFORMER NETWORK, ELECTRONIC DEVICE, AND COMPUTER-PROGRAM PRODUCTTECHNICAL FIELD

[0001] The present invention relates to display technology, more particularly, to a method of training a vision transformer network, an electronic device, and a computer-program product.BACKGROUND

[0002] Vision transformer network is a neural network model that uses the transformer architecture to encode image inputs into feature vectors. Typically, vision transformer network consists of two main components: the backbone and the head. The backbone is responsible for the encoding step of the network. The backbone takes the input images and outputs a vector of features. The head is responsible for making the predictions. The head maps the encoded feature vectors to the prediction scores.SUMMARY

[0003] In one aspect, the present disclosure provides a method of training a vision transformer network, comprising extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder; wherein the one or more modules comprises a first module; wherein the method further comprises receiving, by the first encoder, a first image, and a random mask; outputting, by the first encoder, extracted features to the first module; wherein the first module comprises one or more regression layers and a second encoder; wherein the method further comprises receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; and calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

[0004] Optionally, the method further comprises receiving a second image; wherein the first image is an unlabeled image, and the second image is an enhanced image derived from the first image.

[0005] Optionally, the first encoder comprises a block feature encoding layer and one or more multi-head self-attention modules; a respective multi-head self-attention module comprises a multi-head attention layer and a feed-forward layer; the multi-head attention layer comprises one or more fully connected layers that handle multi-head attention calculations; and  the feed-forward layer comprises one or more fully connected layers for additional processing of attention outputs.

[0006] Optionally, the second encoder is an encoder based on a structure of the first encoder, and obtained by dynamically updating weight parameters of the first encoder; and wherein the method further comprises receiving, by the second encoder, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0007] Optionally, the method further comprises receiving, by a first decoder of the first module, the predicted feature vectors of masked regions of the first image from the one or more regression layers; and determining, by the first decoder, predicted category score vectors based on the predicted feature vectors of masked regions of the first image.

[0008] Optionally, the method further comprises receiving, by a second decoder of the first module, a second image; and determining, by the second decoder, category score vectors based on the second image.

[0009] Optionally, the method further comprises calculating a second loss based on the predicted category score vectors and the category score vectors. In one example, the second loss is a cross-entropy loss.

[0010] Optionally, the one or more modules further comprises a second module; wherein the method further comprises receiving, by a third decoder of the second module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and performing, by the third decoder, image reconstruction to obtain a reconstructed image.

[0011] Optionally, the method further comprises calculating a third loss based on the features of unmasked regions of the first image and the reconstructed image.

[0012] Optionally, the one or more modules further comprises a third module; wherein the method further comprises extracting, by a third encoder of a third module, features of a first image to obtain features extracted by the third encoder.

[0013] Optionally, the method further comprises performing, by the third module, model distillation to impose constraints on the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0014] Optionally, the method further comprises using the features extracted by the third encoder as feature guidance for the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0015] Optionally, the method further comprises comparing, by the third module, the features extracted by the third encoder and the features of masked regions of the first image  obtained by filtering features extracted by the first encoder using the random mask, and minimizing, by the third module, a difference between the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask, thereby performing the feature guidance.

[0016] Optionally, the method further comprises calculating a fourth loss based on the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0017] Optionally, the method further comprises performing an iterative learning process based on a weighted sum of a first loss, a second loss, a third loss, and a fourth loss; wherein the first loss is calculated based on predicted feature vectors of masked regions of the first image and feature vectors of masked regions of the first image; the second loss is calculated based on predicted category score vectors and category score vectors; the third loss is calculated based on features of unmasked regions of the first image and a reconstructed image; and the fourth loss is calculated based on features extracted by the third encoder and features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0018] Optionally, the weighted sum is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*leva;

[0019] wherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.

[0020] Optionally, the method further comprises obtaining the first image by applying at least one of data augmentation algorithms, first-scale resizing, or normalization, to an input image; and obtaining a second image by replacing scaling and / or normalization in the first image with a second scale.

[0021] Optionally, the random mask is a binary matrix that is generated with a degree of randomness and is used to selectively hide or obscure certain parts of the first image; and the random mask comprises a masked part and an unmasked part; wherein the method further comprises receiving, by the first encoder, the first image and the unmasked part of the random mask as inputs.

[0022] In another aspect, the present disclosure provides an electronic device, comprising a memory; and one or more processors; wherein the memory and the one or more processors are connected with each other; and the memory stores computer-executable instructions for controlling the one or more processors to extract, by a first encoder, features of a first image,  and process, by one or more modules, extracted features; wherein the one or more modules comprises a first module; wherein the memory further stores computer-executable instructions for controlling the one or more processors to receive, by the first encoder, a first image, and a random mask; and output, by the first encoder, extracted features to the first module; wherein the first module comprises one or more regression layers and a second encoder; wherein the memory further stores computer-executable instructions for controlling the one or more processors to receive, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receive, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the second encoder, feature vectors of masked regions of the first image; and calculate a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

[0023] In another aspect, the present disclosure provides a computer-program product, comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder; wherein the one or more modules comprises a first module; wherein the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by the first encoder, a first image, and a random mask; outputting, by the first encoder, extracted features to the first module; wherein the first module comprises one or more regression layers and a second encoder; wherein the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; and calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

[0024] BRIEF DESCRIPTION OF THE FIGURES

[0025] The following drawings are merely examples for illustrative purposes according to various disclosed embodiments and are not intended to limit the scope of the present invention.

[0026] FIG. 1 is a schematic diagram illustrating the structure of a vision transformer network in some embodiments according to the present disclosure.

[0027] FIG. 2 is a schematic diagram illustrating the structure of a vision transformer network in some embodiments according to the present disclosure.

[0028] FIG. 3 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.

[0029] FIG. 4 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.

[0030] FIG. 5 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.

[0031] FIG. 6 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.

[0032] FIG. 7 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.

[0033] FIG. 8 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure.DETAILED DESCRIPTION

[0034] The disclosure will now be described more specifically with reference to the following embodiments. It is to be noted that the following descriptions of some embodiments are presented herein for purpose of illustration and description only. It is not intended to be exhaustive or to be limited to the precise form disclosed.

[0035] The present disclosure provides, inter alia, a method of training a vision transformer network, an electronic device, and a computer-program product that substantially obviate one or more of the problems due to limitations and disadvantages of the related art. In one aspect, the present disclosure provides a method of training a vision transformer network. In some embodiments, the method includes extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder. Optionally, the one or more modules comprises a first module. Optionally, the method further comprises receiving, by the first encoder, a first image, a second image, and a random mask; and outputting, by the first encoder, extracted features to a first module. Optionally, the first module comprises one or more regression layers and a second encoder. Optionally, the method further comprises receiving, by one or more regression layers of the first module, features of  unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; and calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

[0036] FIG. 1 is a schematic diagram illustrating the structure of a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 1, the vision transformer network includes a first encoder EC1 configured to extract features of a first image IM1, and one or more modules configured to process the features extracted by the first encoder EC1. Optionally, the one or more modules include a first module MD1, a second module MD2, and a third module MD3.

[0037] The first encoder EC1 is equivalent to a backbone in the context of a vision transformer model, and the one or more modules are equivalent to one or more heads in the context of the vision transformer model. Backbone and head are essential components of the architecture of the vision transformer model, each with distinct roles in processing and transforming visual data. The backbone, also known as the feature extractor or encoder, is the part of the vision transformer model responsible for processing the input image and extracting meaningful features. It typically consists of the initial layers and transformations that convert the raw image into a set of feature representations. In a vision transformer model, the backbone takes the input image and converts it into a sequence of fixed-size feature vectors. This process often includes breaking the image into non-overlapping patches, linearly embedding them, and adding positional encodings. The backbone can include multiple layers of self-attention and feed-forward networks to capture complex relationships between patches. The backbone's primary role is to capture essential visual information from the input data, enabling the model to understand the content and structure of the image. It transforms the image into a format that can be further processed by the model’s head.

[0038] The head, sometimes referred to as the classifier or decoder, is the part of the vision transformer model responsible for making predictions based on the features extracted by the backbone. It typically consists of one or more fully connected layers that map the extracted features to the desired output. The head takes the feature representations generated by the backbone and performs tasks specific to the computer vision problem at hand. For image classification, the head may include a softmax layer to predict class probabilities. In other tasks like object detection or semantic segmentation, the head may consist of additional layers  tailored to those tasks. The head's role is to transform the high-level features from the backbone into task-specific predictions or output. It tailors the vision transformer’s capabilities to the particular computer vision task, making it versatile and adaptable to a wide range of applications.

[0039] The present disclosure provides a multi-head vision transformer network configured for unsupervised learning. Each module of the one or more modules is configured to perform a specific task within the unsupervised learning process. In some embodiments, the first image IM1 is an unlabeled image from a scene. Multiple images input to the first encoder EC1 is used to train a vision transformer model without any ground truth labels. In one example, an unlabeled image is an image that does not have associated metadata describing its content. In the context of machine learning and computer vision, labels are annotations or tags that provide information about the image, such as the identification of objects within the image, the classification of the scene, or any other descriptive attributes.

[0040] In some embodiments, the vision transformer network according to the present disclosure is configured to perform an iterative learning process, in which the vision transformer network processes the first image IM1 and adjusts feature representations over one or more iterations. Once the first encoder EC1 has been trained in an unsupervised manner, learned features from the iterative learning process can be transferred and utilized for downstream tasks DT. Transfer learning allows the model to adapt its knowledge to specific tasks without starting from scratch.

[0041] In some embodiments, the first module MD1 is configured to use regression and classification guidance through a teacher classification decoder, as well as feature structure separation via encoding and decoding. This enhances the feature learning capability of the feature extraction network. In one example, the first module MD1 includes a context autoencoder ( “CAE” ) . A context autoencoder is a type of neural network architecture designed to capture and encode contextual information along with the input data during the encoding and decoding processes.

[0042] The encoder part of the convolutional autoencoder typically includes convolutional layers, pooling layers, and optionally some fully connected layers. Convolutional layers are used to capture local patterns and features in the input image. Pooling layers reduce the spatial dimensions of the feature maps, leading to a compressed representation. The final layer of the encoder typically outputs a lower-dimensional representation of the input data, often referred to as a bottleneck layer or latent space. The bottleneck layer in the encoder represents a highly compressed version of the input image, containing important features and patterns. The dimension of this latent space is a hyperparameter and can be adjusted depending on the desired level of compression.

[0043] The decoder part of the network is responsible for reconstructing the original input from the compressed representation. The decoder part typically includes convolutional layers and up-sampling layers (e.g., transposed convolutions or nearest-neighbor up-sampling) to increase the spatial dimensions. The final layer of the decoder produces the reconstructed image.

[0044] The loss function used in training the convolutional autoencoder measures the similarity between the input image and the reconstructed image. In one example, the loss function is mean squared error, which penalizes the difference between pixel values in the original and reconstructed images. Training a convolutional autoencoder involves minimizing this loss function, which encourages the network to learn a compact representation of the input data in the bottleneck layer while preserving important features for reconstruction.

[0045] In some embodiments, the second module MD2 is configured to reconstruct image features based on a teacher classification decoder. This not only helps the model learn relevant scene features for image reconstruction but also provides a basis for visual performance analysis. In one example, the second module MD2 includes a model that explores the limits of masked visual representation learning at scale (e.g., an “EVA” model) . Instead of predicting features, the EVA model utilizes a large-scale teacher encoder to guide and enhance the feature extraction process. It combines the concept of model distillation to impose constraints on the feature maps, aiming to achieve two critical goals, enhancing feature representation and improving feature robustness. The EVA model introduces a novel approach by leveraging a large-scale teacher encoder, which provides feature guidance to the smaller feature extraction networks. This guidance enriches the representation by adding a dimension to the feature maps, thus increasing their complexity and expressiveness. Consequently, this enhances the ability of the smaller feature extraction networks to represent visual features more effectively. This approach contributes to the model's capacity to capture and represent features in the input data comprehensively.

[0046] In addition to enhancing feature representation, the EVA model’s utilization of the teacher encoder contributes to the robustness of the feature extraction network. The teacher encoder’s guidance helps the smaller network maintain a high level of stability and resilience in the face of variations and challenges in the input data. This aspect is particularly valuable for real-world applications where data can exhibit considerable diversity and complexity.

[0047] To achieve these goals, the EVA model employs a unique loss function. Features of masked regions and features of unmasked regions of an image may be obtained by filtering features extracted by an encoder using a random mask. The EVA model employs a unique loss function that measures the cosine similarity between the feature vectors generated by the teacher encoder and the feature maps corresponding to the unmasked regions in the input data. This loss function incentivizes the smaller network to produce feature representations that are  more similar to the teacher encoder's features, aligning them closely with the comprehensive and high-quality features extracted by the large-scale model. Minimizing the cosine similarity loss is central to improving the feature representation and robustness.

[0048] In some embodiments, the third module MD3 is configured to perform image reconstruction. This not only helps the model learn relevant scene features for image reconstruction but also provides a basis for visual performance analysis. In one example, the third module MD3 includes a masked autoencoder (MAE) . A masked autoencoder is an encoder configured to learn representations of data with missing or masked information. The missing or masked information could be data with certain features or elements deliberately hidden or incomplete. The encoder part of the masked autoencoder takes the input data, which includes the masked or missing values, and transforms it into a lower-dimensional representation. This encoder should learn to capture the most important features from the available information while dealing with the missing parts of the data. The reduced-dimensional representation created by the encoder is often referred to as the latent space. It contains the compressed information about the input data and should ideally capture the essence of the data, even with missing values. The decoder part of the masked autoencoder takes the information from the latent space and attempts to reconstruct the complete input data, including the missing or masked values. One of the primary advantages of the masked autoencoder is their ability to learn from incomplete or partially missing data. The encoder learns to fill in the gaps or predict the missing values during the reconstruction process.

[0049] The training objective for the masked autoencoder is to minimize the difference between the reconstructed data and the ground truth data. The loss function used during training typically measures the similarity between the predicted (reconstructed) data and the original (ground truth) data. In one example, the loss function is mean squared error, which penalizes the difference between the reconstructed data and the ground truth data.

[0050] The vision transformer network is configured to calculate a weighted sum WS of the loss functions from the one or more modules (e.g., the first module MD1, the second module MD2, and the third module MD3) .

[0051] FIG. 2 is a schematic diagram illustrating the structure of a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 2, the vision transformer network includes a first encoder EC1 configured to extract features of a first image IM1, and one or more modules configured to process the features extracted by the first encoder EC1. Optionally, the one or more modules include a first module MD1, a second module MD2, and a third module MD3.

[0052] In some embodiments, the first encoder EC1 is configured to receive a first image IM1, a second image IM2, and a random mask RM. In some embodiments, the first image  IM1 is an unlabeled image, and the second image IM2 is an enhanced image derived from the first image IM1. In some embodiments, the first image IM1 is an image obtained by applying data augmentation algorithms, followed by the first-scale resizing and normalization, to an input image. In some embodiments, the second image IM2 is an image obtained by replacing both the scaling and normalization in the first image IM1 with a second scale.

[0053] In some embodiments, the random mask RM is a binary matrix that is generated with a degree of randomness and is used to selectively hide or obscure certain parts of an image or data. The mask is typically composed of elements with two possible values: 0 and 1. The value 0 indicates that the corresponding part of the image or data is not obscured and is considered "unmasked" or fully visible. The value 1 indicates that the corresponding part of the image or data is obscured, hidden, or "masked, " making it partially or completely invisible. In some embodiments, the random mask RM has a specified mask proportion, which refers to the proportion or fraction of the image that is masked during the data preprocessing. Optionally, the specified mask proportion is in a range of 40%to 80%, e.g., 40%to 50%, 50%to 60%, 60%to 70%, or 70%to 80%. In one example, the specified mask proportion is 60%.

[0054] Various appropriate algorithms may be used to derive the second image IM2 from the first image IM1. Examples of appropriate algorithms include various data augmentation techniques such as random flipping, random grayscale, random cropping, multi-scale augmentation, and CutPaste augmentation. Additional examples of appropriate algorithms include multi-scale resizing (e.g., dual-scale resizing) and multi-scale normalization (e.g., dual-scale normalization) .

[0055] Multi-scale augmentation is a data augmentation technique used in machine learning and computer vision. It involves randomly resizing training data, such as images, to different scales and then expanding or padding them to a uniform or consistent size. This process introduces variations in the scale of the data, which can help the model generalize better by learning from images of different sizes.

[0056] CutPaste augmentation is another data augmentation method used in machine learning, particularly for image processing tasks. It involves selecting a random region or patch from one part of an image, cutting or cropping it, and then pasting or copying it to a random location within the same image. This technique creates variations in the spatial arrangement of objects within the image, effectively generating new training samples with rearranged object positions.

[0057] In some embodiments, the first encoder EC1 includes a block feature encoding layer and one or more multi-head self-attention modules. A block feature encoding layer typically refers to a layer within a neural network, often used in computer vision tasks, that is responsible for encoding or extracting features from a specific region or block of the input data.  This layer processes a portion of the input data, typically through convolutional operations, to capture local patterns and details within that block. A multi-head self-attention module is a neural network component that computes multiple sets of attention weights simultaneously for different positions or elements within a sequence or input data. Each set of attention weights is calculated based on the relationships between elements and is used to capture contextual information and dependencies.

[0058] In some embodiments, the block feature encoding layer has 192 convolutional kernels. In one example, each convolutional kernel has a size of 16 x 16 pixels. In another example, the convolution operation of the block feature encoding layer uses a stride of 16.

[0059] In some embodiments, a respective multi-head self-attention module is equipped with multiple attention heads. In one example, the respective multi-head self-attention module is equipped with three attention heads. In some embodiments, the respective multi-head self-attention module includes a multi-head attention layer and a feed-forward layer. In one example, the multi-head attention layer includes one or more (e.g., 2) fully connected layers that handle multi-head attention calculations. In another example, the feed-forward layer includes one or more (e.g., 2) fully connected layers for additional processing of the attention outputs.

[0060] In some embodiments, the first encoder EC1 is configured to output extracted features to the first module MD1. In some embodiments, the first module MD1 is configured to introduce one or more regression layers RL between a first decoder DC1 (e.g., a feature classification decoder) and the first encoder EC1 (e.g., a feature extraction encoder) . It restricts the output of the one or more regression layers RL for unmasked region features to be consistent with the masked region features. This separation enhances the feature extraction capability of the vision transformer network, effectively decoupling the feature extraction and feature classification stages. Additionally, this algorithm imposes constraints on the first decoder DC1 (e.g., the feature classification decoder) to further enhance feature representation. The loss function includes two weighted components: the first component is the mean squared error loss between an output of the one or more regression layers RL for unmasked input and the input to a second encoder EC2 (e.g., a teacher encoder) for masked data. The second component is the cross-entropy loss between the classification results for unmasked input and the output of a second decoder DC2 (e.g., a teacher feature classifier) .

[0061] In some embodiments, the first encoder EC1 is configured to output extracted features to the second module MD2. In some embodiments, the second module MD2 is configured to introduce a certain proportion of random mask noise during the training input process. This allows the vision transformer network to perform feature reconstruction in the masked regions using the information from unmasked areas, thereby enhancing the robustness of the vision transformer network. In the present disclosure, adjustments and refinements are made to a  decoding phase of the second module MD2, enabling it to perform decoding and reconstruction on feature maps extracted using the vision transformer network. The loss function for the decoder involves calculating the mean squared error between the input image and the output reconstruction image.

[0062] In some embodiments, the first encoder EC1 is configured to output extracted features to the third module MD3. In some embodiments, the third module MD3 is configured to employs a large-sized teacher encoder to guide the features, rather than predict features directly. By combining the characteristics of model distillation, it imposes constraints on the feature maps. This approach has a dual purpose: it enhances the feature expression capacity of the small-sized feature extraction network while also increasing the feature robustness based on the teacher encoder. The loss function for the third module MD3 includes a cosine similarity loss between the feature vectors output by a third encoder EC3 (e.g., a teacher encoder) and the features in the unmasked regions.

[0063] As discussed above, the random mask RM is a binary matrix that is generated with a degree of randomness and is used to selectively hide or obscure certain parts of an image or data. The mask is typically composed of elements with two possible values: 0 and 1. The value 0 indicates that the corresponding part of the image or data is not obscured and is considered "unmasked" or fully visible. The value 1 indicates that the corresponding part of the image or data is obscured, hidden, or "masked, " making it partially or completely invisible.

[0064] In some embodiments, the random mask RM includes a masked part (corresponding to the value 1) and an unmasked part (corresponding to the value 0) . Referring to FIG. 2, the first encoder EC1 is configured to receive the first image IM1 and the unmasked part UMP of the random mask RM as inputs. In some embodiments, the first encoder EC1 is configured to transmit an output to the first module MD1.

[0065] In some embodiments, the first module MD1 includes one or more regression layers RL. In some embodiments, the one or more regression layers RL are configured to receive features of unmasked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, and is configured to determine predicted feature vectors of masked regions of the first image IM1 based on the features of unmasked regions of the first image IM1.

[0066] In some embodiments, the first module MD1 further includes a second encoder EC2. In some embodiments, the second encoder EC2 is configured to receive features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, and is configured to determine feature vectors of masked regions of the first image IM1.

[0067] In some embodiments, the second encoder EC2 is an encoder based on the structure of the first encoder EC1, and obtained by dynamically updating weight parameters of the first encoder EC1. In some embodiments, the second encoder EC2 is configured to proportionally update weight parameters while iteratively optimizing the first encoder EC1. Optionally, an input to the second encoder EC2 is the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM.

[0068] In some embodiments, the vision transformer network further includes a first loss function calculator LS1. In some embodiments, the first loss function calculator LS1 is configured to receive the predicted feature vectors of masked regions of the first image IM1 from the one or more regression layers RL, and the feature vectors of masked regions of the first image IM1 from the second encoder EC2, and is configured to calculate a first loss (e.g., a separation feature loss) based on the predicted feature vectors of masked regions of the first image IM1 and the feature vectors of masked regions of the first image IM1. In one example, the first loss is a mean square error loss.

[0069] In some embodiments, the first module MD1 further includes a first decoder DC1. In some embodiments, the first decoder DC1 is configured to receive the predicted feature vectors of masked regions of the first image IM1 from the one or more regression layers RL, and is configured to determine predicted category score vectors based on the predicted feature vectors of masked regions of the first image IM1.

[0070] In some embodiments, the first module MD1 further includes a second decoder DC2. In some embodiments, the second decoder DC2 is configured to receive the second image IM2, and is configured to determine category score vectors based on the second image IM2. The category score vectors serve as a reference for feature classification.

[0071] In some embodiments, the vision transformer network further includes a second loss function calculator LS2. In some embodiments, the second loss function calculator LS2 is configured to receive the predicted category score vectors from the first decoder DC1 and the category score vectors from the second decoder DC2, and is configured to calculate a second loss (e.g., a feature classification loss) based on the predicted category score vectors and the category score vectors. In one example, the second loss is a cross-entropy loss.

[0072] In some embodiments, the second module MD2 includes a third decoder DC3. In some embodiments, the third decoder DC3 is configured to receive features of unmasked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, and is configured to perform image reconstruction to obtain a reconstructed image.

[0073] In some embodiments, the vision transformer network further includes a third loss function calculator LS3. In some embodiments, the third loss function calculator LS3 is  configured to receive the features of unmasked regions of the first image IM1 and the reconstructed image, and configured to calculate a third loss based on the features of unmasked regions of the first image IM1 and the reconstructed image. In one example, the third loss is a mean square error loss.

[0074] In some embodiments, the third module MD3 includes a third encoder EC3. In some embodiments, the third encoder EC3 is configured to extract features of a first image IM1 to obtain features extracted by the third encoder EC3.

[0075] In some embodiments, the third module MD3 is configured to use the features extracted by the third encoder EC3 as feature guidance FG for the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM. The term feature guidance refers to the process of using information from one set of features (usually from a pre-trained model or teacher model, for example, the third encoder EC3 according to the present disclosure) to guide or enhance the feature learning in another model, often referred to as the student model (for example, the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM) . The teacher model provides supervision to the student model by helping it learn more robust and expressive features. The feature guidance FG can help the student model learn better representations, improve its performance, or adapt to specific tasks.

[0076] In some embodiments, the third module MD3 is configured to compare the features extracted by the third encoder EC3 and the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, and minimize a difference between the features extracted by the third encoder EC3 and the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, thereby performing the feature guidance FG.

[0077] In some embodiments, the vision transformer network further includes a fourth loss function calculator LS4. In some embodiments, the fourth loss function calculator LS4 is configured to receive the features extracted by the third encoder EC3 and the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM, and configured to calculate a fourth loss based on the features extracted by the third encoder EC3 and the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM. In one example, the fourth loss is a cosine similarity loss.

[0078] In some embodiments, the third module MD3 is further configured to perform model distillation to impose constraints on the features of masked regions of the first image IM1 obtained by filtering features extracted by the first encoder EC1 using the random mask RM. By combining the concept of the model distillation, the student model is trained to mimic the  output of the teacher model, which helps in transferring the knowledge from the teacher model to the student model.

[0079] In some embodiments, the third encoder EC3 is based on and loaded with publicly available feature extraction network parameters.

[0080] In some embodiments, the third encoder EC3 is configured to undergo an iterative optimization, e.g., a synchronous iterative optimization.

[0081] In alternative embodiments, the third encoder EC3 is configured not to undergo an iterative optimization.

[0082] In some embodiments, the vision transformer network is configured to perform an iterative learning process based on a weighted sum WS of the first loss, the second loss, the third loss, and the fourth loss.

[0083] In some embodiments, the weighted sum WS is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*leva

[0084] wherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.

[0085] In another aspect, the present disclosure provides a vision transformer model obtained by training the vision transformer network disclosed herein. In some embodiments, performance of the vision transformer model can be tested by various appropriate methods.

[0086] In some embodiments, during the model inference process, data augmentation (random flipping, random grayscale, random cropping, multi-scale augmentation, and CutPaste augmentation) , multi-scale resizing (e.g., dual-scale resizing) , or multi-scale normalization (e.g., dual-scale normalization) is not performed on the first image IM1. In some embodiments, only a reconstructed image is output from the third decoder of the second module. Details or features of the reconstructed image is evaluated through visual analysis or analyzed by a target scene feature evaluation algorithm, to analyze the similarity between the reconstructed image and the input image.

[0087] In some embodiments, during the model inference process, data augmentation (e.g., random resizing, random color adjustment, random contrast adjustment, and random noise) is performed on the first image IM1 to obtain multiple enhanced images based on the first image IM1. In some embodiments, only an output (e.g., the predicted category score vectors) from the first decoder DC1 is saved. The multiple enhanced images are averaged, and the average of  the multiple enhanced images is compared with an output (the category score vectors) from the second decoder DC2.

[0088] The vision transformer model may have the structure of the vision transformer network discussed in the present disclosure. In some embodiments, the vision transformer model includes at least the first encoder EC1 discussed in the present disclosure, which is obtained by training the vision transformer network disclosed herein. In some embodiments, the first encoder EC1 includes a block feature encoding layer and one or more multi-head self-attention modules. In some embodiments, the block feature encoding layer has 192 convolutional kernels. In one example, each convolutional kernel has a size of 16 x 16 pixels. In another example, the convolution operation of the block feature encoding layer uses a stride of 16. In some embodiments, a respective multi-head self-attention module is equipped with multiple attention heads. In one example, the respective multi-head self-attention module is equipped with three attention heads. In some embodiments, the respective multi-head self-attention module includes a multi-head attention layer and a feed-forward layer. In one example, the multi-head attention layer includes one or more (e.g., 2) fully connected layers that handle multi-head attention calculations. In another example, the feed-forward layer includes one or more (e.g., 2) fully connected layers for additional processing of the attention outputs.

[0089] In another aspect, the present disclosure provides a method of training a vision transformer network. FIG. 3 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 3, the method in some embodiments includes extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder. Optionally, the one or more modules include a first module, a second module, and a third module.

[0090] In some embodiments, the method further includes receiving, by the first encoder, a first image, a second image, and a random mask. In some embodiments, the first image is an unlabeled image, and the second image is an enhanced image derived from the first image. In some embodiments, the method further includes applying data augmentation algorithms, followed by the first-scale resizing and normalization, to an input image to obtain the first image. In some embodiments, the method further includes replacing both the scaling and normalization in the first image with a second scale to obtain the second image.

[0091] In some embodiments, the first encoder includes a block feature encoding layer and one or more multi-head self-attention modules. In some embodiments, the block feature encoding layer has 192 convolutional kernels. In one example, each convolutional kernel has a size of 16 x 16 pixels. In another example, the convolution operation of the block feature encoding layer uses a stride of 16. In some embodiments, a respective multi-head self- attention module is equipped with multiple attention heads. In one example, the respective multi-head self-attention module is equipped with three attention heads. In some embodiments, the respective multi-head self-attention module includes a multi-head attention layer and a feed-forward layer. In one example, the multi-head attention layer includes one or more (e.g., 2) fully connected layers that handle multi-head attention calculations. In another example, the feed-forward layer includes one or more (e.g., 2) fully connected layers for additional processing of the attention outputs.

[0092] In some embodiments, the method further includes generating the random mask, with a degree of randomness, to selectively hide or obscure certain parts of an image or data. The mask is typically composed of elements with two possible values: 0 and 1. The value 0 indicates that the corresponding part of the image or data is not obscured and is considered "unmasked" or fully visible. The value 1 indicates that the corresponding part of the image or data is obscured, hidden, or "masked, " making it partially or completely invisible. In some embodiments, the random mask has a specified mask proportion, which refers to the proportion or fraction of the image that is masked during the data preprocessing. Optionally, the specified mask proportion is in a range of 40%to 80%, e.g., 40%to 50%, 50%to 60%, 60%to 70%, or 70%to 80%. In one example, the specified mask proportion is 60%.

[0093] Various appropriate algorithms may be used to derive the second image from the first image. Examples of appropriate algorithms include various data augmentation techniques such as random flipping, random grayscale, random cropping, multi-scale augmentation, and CutPaste augmentation. Additional examples of appropriate algorithms include multi-scale resizing (e.g., dual-scale resizing) and multi-scale normalization (e.g., dual-scale normalization) .

[0094] In some embodiments, the method further includes outputting, by the first encoder, extracted features to a first module. Optionally, the method includes outputting, by the first encoder, extracted features to one or more regression layers between a first decoder (e.g., a feature classification decoder) and the first encoder (e.g., a feature extraction encoder) .

[0095] In some embodiments, the method further includes outputting, by the first encoder, extracted features to a second module. In some embodiments, the method further includes introducing, by the second module, a certain proportion of random mask noise during the training input process.

[0096] In some embodiments, the method further includes outputting, by the first encoder, extracted features to a third module. In some embodiments, the method further includes employing, by the third module, a large-sized teacher encoder to guide the features, rather than predict features directly.

[0097] FIG. 4 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 4, in some embodiments, the method further includes receiving, by the first encoder, the first image and an unmasked part of the random mask; and transmitting, by the first encoder, an output to the first module. Optionally, the method includes receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image.

[0098] FIG. 5 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 5, in some embodiments, the method further includes receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image.

[0099] In some embodiments, the second encoder EC2 is an encoder based on the structure of the first encoder EC1, and obtained by dynamically updating weight parameters of the first encoder EC1. In some embodiments, the method further includes proportionally updating weight parameters while iteratively optimizing the first encoder, thereby obtaining the second encoder.

[0100] In some embodiments, the method further includes calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image. In one example, the first loss is a mean square error loss.

[0101] FIG. 6 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 6, in some embodiments, the method further includes receiving, by a first decoder of the first module, the predicted feature vectors of masked regions of the first image from the one or more regression layers; and determining, by the first decoder, predicted category score vectors based on the predicted feature vectors of masked regions of the first image.

[0102] In some embodiments, the method further includes receiving, by a second decoder of the first module, the second image; and determining, by the second decoder, category score vectors based on the second image. The category score vectors serve as a reference for feature classification.

[0103] In some embodiments, the method further includes calculating a second loss based on the predicted category score vectors and the category score vectors. In one example, the second loss is a cross-entropy loss.

[0104] FIG. 7 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 7, in some embodiments, the method further includes receiving, by a third decoder of a second module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and performing, by the third decoder, image reconstruction to obtain a reconstructed image.

[0105] In some embodiments, the method further includes calculating a third loss based on the features of unmasked regions of the first image and the reconstructed image. In one example, the third loss is a mean square error loss.

[0106] FIG. 8 is a flow chart illustrating a method of training a vision transformer network in some embodiments according to the present disclosure. Referring to FIG. 8, in some embodiments, the method further includes extracting, by a third encoder of a third module, features of a first image to obtain features extracted by the third encoder.

[0107] In some embodiments, the method further includes using the features extracted by the third encoder as feature guidance for the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0108] In some embodiments, the method further includes calculating a fourth loss based on the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask. In one example, the fourth loss is a cosine similarity loss.

[0109] In some embodiments, the method further includes performing, by the third module, model distillation to impose constraints on the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0110] In some embodiments, the third encoder is based on and loaded with publicly available feature extraction network parameters.

[0111] In some embodiments, the third encoder is configured to undergo an iterative optimization, e.g., a synchronous iterative optimization.

[0112] In alternative embodiments, the third encoder is configured not to undergo an iterative optimization.

[0113] In some embodiments, the method further includes performing an iterative learning process based on a weighted sum of the first loss, the second loss, the third loss, and the fourth loss.

[0114] In some embodiments, the weighted sum WS is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*leva

[0115] wherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.

[0116] In another aspect, the present disclosure provides a computer-program product comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by a processor to cause the processor to perform extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder. Optionally, the one or more modules include a first module, a second module, and a third module.

[0117] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by the first encoder, a first image, a second image, and a random mask. In some embodiments, the first image is an unlabeled image, and the second image is an enhanced image derived from the first image. In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform applying data augmentation algorithms, followed by the first-scale resizing and normalization, to an input image to obtain the first image. In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform replacing both the scaling and normalization in the first image with a second scale to obtain the second image.

[0118] In some embodiments, the first encoder includes a block feature encoding layer and one or more multi-head self-attention modules. In some embodiments, the block feature encoding layer has 192 convolutional kernels. In one example, each convolutional kernel has a size of 16 x 16 pixels. In another example, the convolution operation of the block feature encoding layer uses a stride of 16. In some embodiments, a respective multi-head self-attention module is equipped with multiple attention heads. In one example, the respective multi-head self-attention module is equipped with three attention heads. In some embodiments, the respective multi-head self-attention module includes a multi-head attention layer and a feed-forward layer. In one example, the multi-head attention layer includes one or more (e.g., 2) fully connected layers that handle multi-head attention calculations. In another example, the feed-forward layer includes one or more (e.g., 2) fully connected layers for additional processing of the attention outputs.

[0119] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform generating the random mask, with a degree of randomness, to selectively hide or obscure certain parts of an image or data. The mask is typically composed of elements with two possible values: 0 and 1. The value 0 indicates that the corresponding part of the image or data is not obscured and is considered "unmasked" or fully visible. The value 1 indicates that the corresponding part of the image or data is obscured, hidden, or "masked, " making it partially or completely invisible. In some embodiments, the random mask has a specified mask proportion, which refers to the proportion or fraction of the image that is masked during the data preprocessing. Optionally, the specified mask proportion is in a range of 40%to 80%, e.g., 40%to 50%, 50%to 60%, 60%to 70%, or 70%to 80%. In one example, the specified mask proportion is 60%.

[0120] Various appropriate algorithms may be used to derive the second image from the first image. Examples of appropriate algorithms include various data augmentation techniques such as random flipping, random grayscale, random cropping, multi-scale augmentation, and CutPaste augmentation. Additional examples of appropriate algorithms include multi-scale resizing (e.g., dual-scale resizing) and multi-scale normalization (e.g., dual-scale normalization) .

[0121] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform outputting, by the first encoder, extracted features to a first module. Optionally, the computer-readable instructions are executable by a processor to further cause the processor to perform outputting, by the first encoder, extracted features to one or more regression layers between a first decoder (e.g., a feature classification decoder) and the first encoder (e.g., a feature extraction encoder) .

[0122] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform outputting, by the first encoder, extracted features to a second module. In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform introducing, by the second module, a certain proportion of random mask noise during the training input process.

[0123] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform outputting, by the first encoder, extracted features to a third module. In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform employing, by the third module, a large-sized teacher encoder to guide the features, rather than predict features directly.

[0124] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by the first encoder, the first image and an unmasked part of the random mask; and transmitting, by the first encoder, an  output to the first module. Optionally, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image.

[0125] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image.

[0126] In some embodiments, the second encoder EC2 is an encoder based on the structure of the first encoder EC1, and obtained by dynamically updating weight parameters of the first encoder EC1. In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform proportionally updating weight parameters while iteratively optimizing the first encoder, thereby obtaining the second encoder.

[0127] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform calculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image. In one example, the first loss is a mean square error loss.

[0128] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by a first decoder of the first module, the predicted feature vectors of masked regions of the first image from the one or more regression layers; and determining, by the first decoder, predicted category score vectors based on the predicted feature vectors of masked regions of the first image.

[0129] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by a second decoder of the first module, the second image; and determining, by the second decoder, category score vectors based on the second image. The category score vectors serve as a reference for feature classification.

[0130] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform calculating a second loss based on the predicted category score vectors and the category score vectors. In one example, the second loss is a cross-entropy loss.

[0131] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform receiving, by a third decoder of a second  module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and performing, by the third decoder, image reconstruction to obtain a reconstructed image.

[0132] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform calculating a third loss based on the features of unmasked regions of the first image and the reconstructed image. In one example, the third loss is a mean square error loss.

[0133] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform extracting, by a third encoder of a third module, features of a first image to obtain features extracted by the third encoder.

[0134] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform using the features extracted by the third encoder as feature guidance for the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0135] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform calculating a fourth loss based on the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask. In one example, the fourth loss is a cosine similarity loss.

[0136] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform performing, by the third module, model distillation to impose constraints on the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0137] In some embodiments, the third encoder is based on and loaded with publicly available feature extraction network parameters.

[0138] In some embodiments, the third encoder is configured to undergo an iterative optimization, e.g., a synchronous iterative optimization.

[0139] In alternative embodiments, the third encoder is configured not to undergo an iterative optimization.

[0140] In some embodiments, the computer-readable instructions are executable by a processor to further cause the processor to perform performing an iterative learning process based on a weighted sum of the first loss, the second loss, the third loss, and the fourth loss.

[0141] In some embodiments, the weighted sum WS is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*leva

[0142] wherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.

[0143] In another aspect, the present disclosure provides an electronic device, comprising a memory; and one or more processors; wherein the memory and the one or more processors are connected with each other. In some embodiments, the memory stores computer-executable instructions for controlling the one or more processors to extract, by a first encoder, features of a first image, and process, by one or more modules, extracted features. In some embodiments, the one or more modules comprises a first module; wherein the memory further stores computer-executable instructions for controlling the one or more processors to receive, by the first encoder, a first image, a second image, and a random mask; and output, by the first encoder, extracted features to a first module; wherein the first module comprises one or more regression layers and a second encoder; wherein the memory further stores computer-executable instructions for controlling the one or more processors to receive, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image; receive, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the second encoder, feature vectors of masked regions of the first image; and calculate a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

[0144] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by the first encoder, a first image, a second image, and a random mask. In some embodiments, the first image is an unlabeled image, and the second image is an enhanced image derived from the first image. In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to apply data augmentation algorithms, followed by the first-scale resizing and normalization, to an input image to obtain the first image. In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to replace both the scaling and normalization in the first image with a second scale to obtain the second image.

[0145] In some embodiments, the first encoder includes a block feature encoding layer and one or more multi-head self-attention modules. In some embodiments, the block feature encoding layer has 192 convolutional kernels. In one example, each convolutional kernel has a size of 16 x 16 pixels. In another example, the convolution operation of the block feature encoding layer uses a stride of 16. In some embodiments, a respective multi-head self-attention module is equipped with multiple attention heads. In one example, the respective multi-head self-attention module is equipped with three attention heads. In some embodiments, the respective multi-head self-attention module includes a multi-head attention layer and a feed-forward layer. In one example, the multi-head attention layer includes one or more (e.g., 2) fully connected layers that handle multi-head attention calculations. In another example, the feed-forward layer includes one or more (e.g., 2) fully connected layers for additional processing of the attention outputs.

[0146] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to generate the random mask, with a degree of randomness, to selectively hide or obscure certain parts of an image or data. The mask is typically composed of elements with two possible values: 0 and 1. The value 0 indicates that the corresponding part of the image or data is not obscured and is considered "unmasked" or fully visible. The value 1 indicates that the corresponding part of the image or data is obscured, hidden, or "masked, " making it partially or completely invisible. In some embodiments, the random mask has a specified mask proportion, which refers to the proportion or fraction of the image that is masked during the data preprocessing. Optionally, the specified mask proportion is in a range of 40%to 80%, e.g., 40%to 50%, 50%to 60%, 60%to 70%, or 70%to 80%. In one example, the specified mask proportion is 60%.

[0147] Various appropriate algorithms may be used to derive the second image from the first image. Examples of appropriate algorithms include various data augmentation techniques such as random flipping, random grayscale, random cropping, multi-scale augmentation, and CutPaste augmentation. Additional examples of appropriate algorithms include multi-scale resizing (e.g., dual-scale resizing) and multi-scale normalization (e.g., dual-scale normalization) .

[0148] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to output, by the first encoder, extracted features to a first module. Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to output, by the first encoder, extracted features to one or more regression layers between a first decoder (e.g., a feature classification decoder) and the first encoder (e.g., a feature extraction encoder) .

[0149] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to output, by the first encoder, extracted features to a  second module. In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to introduce, by the second module, a certain proportion of random mask noise during the training input process.

[0150] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to output, by the first encoder, extracted features to a third module. In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to employ, by the third module, a large-sized teacher encoder to guide the features, rather than predict features directly.

[0151] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by the first encoder, the first image and an unmasked part of the random mask; and transmit, by the first encoder, an output to the first module. Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image.

[0152] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the second encoder, feature vectors of masked regions of the first image.

[0153] In some embodiments, the second encoder EC2 is an encoder based on the structure of the first encoder EC1, and obtained by dynamically updating weight parameters of the first encoder EC1. In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to proportionally update weight parameters while iteratively optimizing the first encoder, thereby obtaining the second encoder.

[0154] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to calculate a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image. In one example, the first loss is a mean square error loss.

[0155] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by a first decoder of the first module, the predicted feature vectors of masked regions of the first image from the one or more regression  layers; and determine, by the first decoder, predicted category score vectors based on the predicted feature vectors of masked regions of the first image.

[0156] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by a second decoder of the first module, the second image; and determine, by the second decoder, category score vectors based on the second image. The category score vectors serve as a reference for feature classification.

[0157] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to calculate a second loss based on the predicted category score vectors and the category score vectors. In one example, the second loss is a cross-entropy loss.

[0158] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to receive, by a third decoder of a second module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and perform, by the third decoder, image reconstruction to obtain a reconstructed image.

[0159] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to calculate a third loss based on the features of unmasked regions of the first image and the reconstructed image. In one example, the third loss is a mean square error loss.

[0160] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to extract, by a third encoder of a third module, features of a first image to obtain features extracted by the third encoder.

[0161] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to use the features extracted by the third encoder as feature guidance for the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0162] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to calculate a fourth loss based on the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask. In one example, the fourth loss is a cosine similarity loss.

[0163] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to perform, by the third module, model distillation to impose constraints on the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.

[0164] In some embodiments, the third encoder is based on and loaded with publicly available feature extraction network parameters.

[0165] In some embodiments, the third encoder is configured to undergo an iterative optimization, e.g., a synchronous iterative optimization.

[0166] In alternative embodiments, the third encoder is configured not to undergo an iterative optimization.

[0167] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to perform an iterative learning process based on a weighted sum of the first loss, the second loss, the third loss, and the fourth loss.

[0168] In some embodiments, the weighted sum WS is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*leva

[0169] wherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.

[0170] Various illustrative neural networks, segments, units, channels, modules, and other operations described in connection with the configurations disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. Such neural networks, segments, units, channels, modules, and operations may be implemented or performed with a general purpose processor, a digital signal processor (DSP) , an ASIC or ASSP, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to produce the configuration as disclosed herein. For example, such a configuration may be implemented at least in part as a hard-wired circuit, as a circuit configuration fabricated into an application-specific integrated circuit, or as a firmware program loaded into non-volatile storage or a software program loaded from or into a data storage medium as machine-readable code, such code being instructions executable by an array of logic elements such as a general purpose processor or other digital signal processing unit. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. A software module may reside in a non-transitory storage medium such as RAM (random-access memory) , ROM (read-only memory) , nonvolatile RAM (NVRAM) such as flash RAM, erasable programmable ROM (EPROM) , electrically erasable programmable ROM (EEPROM) ,  registers, hard disk, a removable disk, or a CD-ROM; or in any other form of storage medium known in the art. An illustrative storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0171] The foregoing description of the embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form or to exemplary embodiments disclosed. Accordingly, the foregoing description should be regarded as illustrative rather than restrictive. Obviously, many modifications and variations will be apparent to practitioners skilled in this art. The embodiments are chosen and described in order to explain the principles of the invention and its best mode practical application, thereby to enable persons skilled in the art to understand the invention for various embodiments and with various modifications as are suited to the particular use or implementation contemplated. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents in which all terms are meant in their broadest reasonable sense unless otherwise indicated. Therefore, the term “the invention” , “the present invention” or the like does not necessarily limit the claim scope to a specific embodiment, and the reference to exemplary embodiments of the invention does not imply a limitation on the invention, and no such limitation is to be inferred. The invention is limited only by the spirit and scope of the appended claims. Moreover, these claims may refer to use “first” , “second” , etc. following with noun or element. Such terms should be understood as a nomenclature and should not be construed as giving the limitation on the number of the elements modified by such nomenclature unless specific number has been given. Any advantages and benefits described may not apply to all embodiments of the invention. It should be appreciated that variations may be made in the embodiments described by persons skilled in the art without departing from the scope of the present invention as defined by the following claims. Moreover, no element and component in the present disclosure is intended to be dedicated to the public regardless of whether the element or component is explicitly recited in the following claims.

Claims

1.A method of training a vision transformer network, comprising extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder;wherein the one or more modules comprises a first module;wherein the method further comprises:receiving, by the first encoder, a first image, and a random mask; andoutputting, by the first encoder, extracted features to the first module;wherein the first module comprises one or more regression layers and a second encoder;wherein the method further comprises:receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image;receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; andcalculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.2.The method of claim 1, further comprising receiving a second image;wherein the first image is an unlabeled image, and the second image is an enhanced image derived from the first image.3.The method of claim 1, wherein the first encoder comprises a block feature encoding layer and one or more multi-head self-attention modules;a respective multi-head self-attention module comprises a multi-head attention layer and a feed-forward layer;the multi-head attention layer comprises one or more fully connected layers that handle multi-head attention calculations; andthe feed-forward layer comprises one or more fully connected layers for additional processing of attention outputs.4.The method of claim 1, wherein the second encoder is an encoder based on a structure of the first encoder, and obtained by dynamically updating weight parameters of the first encoder; andwherein the method further comprises receiving, by the second encoder, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.5.The method of any one of claims 1 to 4, further comprising:receiving, by a first decoder of the first module, the predicted feature vectors of masked regions of the first image from the one or more regression layers; anddetermining, by the first decoder, predicted category score vectors based on the predicted feature vectors of masked regions of the first image.6.The method of claim 5, further comprising receiving, by a second decoder of the first module, a second image; and determining, by the second decoder, category score vectors based on the second image.7.The method of claim 6, further comprising calculating a second loss based on the predicted category score vectors and the category score vectors, wherein the second loss is a cross-entropy loss.8.The method of any one of claims 1 to 7, wherein the one or more modules further comprises a second module;wherein the method further comprises:receiving, by a third decoder of the second module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; andperforming, by the third decoder, image reconstruction to obtain a reconstructed image.9.The method of claim 8, further comprising calculating a third loss based on the features of unmasked regions of the first image and the reconstructed image.10.The method of any one of claims 1 to 9, wherein the one or more modules further comprises a third module;wherein the method further comprises extracting, by a third encoder of the third module, features of a first image to obtain features extracted by the third encoder.11.The method of claim 10, further comprising performing, by the third module, model distillation to impose constraints on the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.12.The method of claim 10, further comprising using the features extracted by the third encoder as feature guidance for the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.13.The method of claim 12, further comprising:comparing, by the third module, the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask, andminimizing, by the third module, a difference between the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask, thereby performing the feature guidance.14.The method of claim 10, further comprising calculating a fourth loss based on the features extracted by the third encoder and the features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.15.The method of any one of claims 1 to 14, further comprising performing an iterative learning process based on a weighted sum of a first loss, a second loss, a third loss, and a fourth loss;wherein the first loss is calculated based on predicted feature vectors of masked regions of the first image and feature vectors of masked regions of the first image;the second loss is calculated based on predicted category score vectors and category score vectors;the third loss is calculated based on features of unmasked regions of the first image and a reconstructed image; andthe fourth loss is calculated based on features extracted by the third encoder and features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask.16.The method of claim 15, wherein the weighted sum is calculated according to: WS=αmae*lmae+αcae* (lce+β*lmse) +αeva*levawherein WS stands for the weighted sum; lmse stands for the first loss; lce stands for the second loss; lmae stands for the third loss; leva stands for the fourth loss; αcae stands for a first weight; β stands for a second weight; αmae stands for a third weight; and αeva stands for a fourth weight.17.The method of any one of claims 1 to 16, further comprising obtaining the first image by applying at least one of data augmentation algorithms, first-scale resizing, or normalization, to an input image; andobtaining a second image by replacing scaling and / or normalization in the first image with a second scale.18.The method of any one of claims 1 to 17, wherein the random mask is a binary matrix that is generated with a degree of randomness and is used to selectively hide or obscure certain parts of the first image; andthe random mask comprises a masked part and an unmasked part;wherein the method further comprises receiving, by the first encoder, the first image and the unmasked part of the random mask as inputs.19.An electronic device, comprising:a memory; andone or more processors;wherein the memory and the one or more processors are connected with each other; andthe memory stores computer-executable instructions for controlling the one or more processors to extract, by a first encoder, features of a first image, and process, by one or more modules, extracted features;wherein the one or more modules comprises a first module;wherein the memory further stores computer-executable instructions for controlling the one or more processors to:receive, by the first encoder, a first image, and a random mask; andoutput, by the first encoder, extracted features to the first module;wherein the first module comprises one or more regression layers and a second encoder;wherein the memory further stores computer-executable instructions for controlling the one or more processors to:receive, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image;receive, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determine, by the second encoder, feature vectors of masked regions of the first image; andcalculate a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.20.A computer-program product, comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform extracting, by a first encoder, features of a first image; and processing, by one or more modules, the features extracted by the first encoder;wherein the one or more modules comprises a first module;wherein the computer-readable instructions are executable by a processor to further cause the processor to perform:receiving, by the first encoder, a first image, and a random mask;outputting, by the first encoder, extracted features to the first module;wherein the first module comprises one or more regression layers and a second encoder;wherein the computer-readable instructions are executable by a processor to further cause the processor to perform:receiving, by one or more regression layers of the first module, features of unmasked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the one or more regression layers of the first module, predicted feature vectors of masked regions of the first image based on the features of unmasked regions of the first image;receiving, by a second encoder of the first module, features of masked regions of the first image obtained by filtering features extracted by the first encoder using the random mask; and determining, by the second encoder, feature vectors of masked regions of the first image; andcalculating a first loss based on the predicted feature vectors of masked regions of the first image and the feature vectors of masked regions of the first image.

Citation Information

Patent Citations

  • Attention-based image generation neural networks

    CN109726794A

  • Video processing model training method, device and equipment

    CN116310643A

  • Visual target tracking method based on mask auto-encoder

    CN116434115A

  • Non-autoregressive machine translation system and method and electronic equipment

    CN116502654A

  • Living body detection model training method and device, medium and electronic equipment

    CN116721315A