Method and system for pre-training a vision transformer using knowledge distillation, and a pre-trained vision transformer

The knowledge distillation framework aligns image-text features to efficiently pre-train vision transformers, reducing data processing overhead and improving performance in vision tasks by aligning image-text representations.

JP2025540663APending Publication Date: 2025-12-16LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025528711
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-20
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing vision transformer pre-training methods require large datasets, leading to excessive data processing overhead and are not suitable for self-supervised learning on uncurated datasets, and existing token sparsification frameworks fail to address image-text misalignment.

Method used

A knowledge distillation framework is employed to pre-train a vision transformer by aligning image-text feature representations using a student encoder relative to a teacher encoder, incorporating a token sparsification layer to accelerate data learning and improve efficiency.

Benefits of technology

This approach reduces data processing overload, enables rapid pre-training on large-scale unscreened image-text pairs, and mitigates image-text misalignment, resulting in a lightweight transformer with superior performance in various vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025540663000001_ABST
    Figure 2025540663000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for pre-training Vision Transformers on large uncurated datasets in a self-supervised manner according to a knowledge distillation framework, reducing data processing overhead and rapidly learning simplified Vision Transformers.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for pre-training a vision transformer through knowledge distillation, and a vision transformer pre-trained thereby. [Background technology]

[0002] Recently, the emergence of Vision-Language Pretraining (VLP), which is pre-trained on large-scale general-domain data, has led to rapid development of artificial intelligence-based computer vision processing technology.

[0003] In particular, Vision Transformers trained on large-scale image-text datasets using techniques such as global self-attention and contrastive language-image pretraining (see Reference 2) have shown groundbreaking advances in downstream tasks, including diverse and challenging vision tasks.

[0004] However, to fully train global self-attention, which is primarily driven by vision transformers, a large dataset is required, which has the problem of excessive data processing overhead.

[0005] On the other hand, in a previous paper, a framework based on token sparsification was proposed as a method to accelerate the pre-training of vision transformers. However, in the previous paper, it is only applicable to supervision-learning pipelines (e.g., classification, detection, or dense prediction) with pre-defined labels, and is not suitable for pre-training with self-supervision of unexplored text-images. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Youwei Liang, Chongjian GE, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. EVit: Expediting vision transformers via token reorganizations. In International Conference on Learning Representations, 2022. [Non-patent document 2] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision, on 26 Feb 2021 Summary of the Invention [Problem to be solved by the invention]

[0007] The present invention proposes a method and system for pre-training vision transformers using a self-supervised learning method on large uncurated datasets to reduce data processing overhead and quickly learn simplified vision transformers.

[0008] In particular, according to one embodiment of the present invention, a method and system for pre-training a vision transformer that uses a self-supervised learning method to address the image-text misalignment problem that can occur when applying an existing token sparsification framework to contrastive language-image pre-training can be provided.

[0009] More specifically, according to one embodiment of the present invention, a method and system for pre-training a vision transformer can be provided, including a knowledge distillation framework for contrastive language-image pre-training that can solve the problem of data efficiency due to token sparsification.

[0010] Furthermore, according to an embodiment of the present invention, a variety of vision tasks and applications involving vision tasks can be provided using such a pre-trained vision transformer. [Means for solving the problem]

[0011] A vision transformer pre-training method and system according to one embodiment of the present invention employs a knowledge distillation framework in which a student encoder is pre-trained in a way that it learns an image-text alignment matrix relative to a teacher encoder.

[0012] In detail, the knowledge distillation framework according to an embodiment of the present invention is pre-trained using a knowledge distillation method so that alignment matrices for image-text feature representations for the text encoder 10 and the teacher encoder and alignment matrices for image-text feature representations for the text encoder 10 and the student encoder match each other. This enables efficient knowledge extraction using a student encoder that is simpler than the teacher encoder and mitigates image-text mismatches that naturally exist in large datasets.

[0013] In this case, the student encoder according to an embodiment of the present invention includes a token sparsification layer to accelerate data learning and improve image-text matching efficiency. [Effects of the Invention]

[0014] A method and system for pre-training a vision transformer according to a knowledge distillation framework according to an embodiment of the present invention reduces data processing overload and enables rapid pre-training for large-scale unscreened image-text pair datasets.

[0015] Furthermore, according to the method and system for pre-training a vision transformer in accordance with the knowledge distillation framework according to the embodiment of the present invention, a lightweight vision transformer can be pre-trained by knowledge distillation.

[0016] Furthermore, the method and system for pre-training a vision transformer according to the knowledge distillation framework according to an embodiment of the present invention solves the problem of image-text misalignment through token sparsification, allowing for faster data processing.

[0017] Furthermore, a vision transformer pre-trained according to the knowledge distillation framework according to an embodiment of the present invention can exhibit excellent performance in various vision tasks, particularly in image segmentation by efficiently identifying patch tokens with high attentional importance. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 illustrates an example block diagram of a computing system for pre-training a vision transformer (Expediting Contrastive Language-Image Pretraining (ECLIPS)) according to a knowledge distillation framework and executing the pre-trained vision transformer, in accordance with an embodiment of the present invention. [Figure 2] FIG. 1 illustrates an example block diagram of a computing device for pre-training a vision transformer and executing the pre-trained vision transformer according to a knowledge distillation framework in accordance with an embodiment of the present invention. [Figure 3] FIG. 10 illustrates an example block diagram of another aspect of a computing device for pre-training a vision transformer and executing the pre-trained vision transformer according to a knowledge distillation framework in accordance with an embodiment of the present invention. [Figure 4] FIG. 1 is a conceptual diagram of a method for pre-training a vision transformer according to a knowledge distillation framework in accordance with an embodiment of the present invention. [Figure 5] 1 illustrates a meta-architecture of a framework for pre-training vision transformers by knowledge distillation according to an embodiment of the present invention. [Figure 6] 10A-10C are graphs comparing existing pre-training methods and models to illustrate the effectiveness of a pre-trained method and a pre-trained Vision Transformer according to an embodiment of the present invention. [Figure 7]FIG. 10 is a diagram comparing attention tokens at the deep layers of a vision transformer pre-trained by an embodiment of the present invention and an existing vision transformer. DETAILED DESCRIPTION OF THE INVENTION

[0019] Because the present invention can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. The advantages and features of the present invention, as well as methods for achieving them, will become clearer with reference to the embodiments described in detail below in conjunction with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be realized in various forms. In the following embodiments, terms such as "first," "second," etc., are used without any limiting meaning but to distinguish one element from another. Furthermore, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "include" or "have" mean the presence of a feature or element described in the specification and do not preclude the possibility that one or more other features or elements may be added. Furthermore, the size of elements in the drawings may be exaggerated or reduced for ease of explanation. For example, the size and thickness of each element in the drawings are arbitrarily illustrated for ease of explanation, and the present invention is not necessarily limited to those shown in the drawings.

[0020] FIG. 1 illustrates an example block diagram of a computing system that pre-trains a vision transformer according to a knowledge distillation framework and executes an application that includes the pre-trained vision transformer, in accordance with an embodiment of the present invention.

[0021] Referring to FIG. 1, a computing system 1000 according to one embodiment of the present invention includes a user computing device 110, a training computing system 150, and a server computing system 130, each of which is communicatively connected via a network 170.

[0022] According to various embodiments of the present invention, 1) the user computing device 110 can pre-train the vision transformer 120 locally and execute an application including the learned vision transformer 120; 2) the server computing system 130, which communicates with the user computing device 110, can pre-train the vision transformers 120 and / or 140 and provide the application including the vision transformers 120 and / or 140 to the user computing device 110 directly or in the form of a web service; and 3) the user computing device 110 and the server computing system 130 can cooperate with each other to pre-train the vision transformers 120 and / or 140 or execute the pre-trained vision transformers 120 and / or 140 to provide various application services.

[0023] Additionally, according to various embodiments of the present invention, the user computing device 110 and / or the server computing system 130 may train the model 120 by interacting with a training computing system 150 communicatively connected via a network 170. In this case, the training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.

[0024] That is, the vision transformer pre-training method according to the embodiment can be implemented in the following ways: 1) the user computing device 110 can pre-train the vision transformer 120 directly locally; 2) the server computing system 130 and the user computing device 110 can interact with each other via a network and perform pre-training; and 3) a separate training computing system 150 can pre-train the vision transformer using various training and learning techniques.

[0025] The training computing system 150 may also be implemented in a manner in which the pre-trained vision transformers 120 and / or 140 are transmitted to the user computing device 110 and / or server computing system 130 via a network to provide and / or update the vision transformers.

[0026] In some embodiments, the training computing system 150 may be part of the server computing system 130 or part of the user computing device 110 .

[0027] Additionally, the present invention provides a vision transformer pre-training method and system that can be included in applications that perform additional work, such as fine-tuning the pre-trained vision transformer, to perform various downstream tasks.

[0028] The user computing devices 110 may include any other type of computing device, such as smartphones, mobile phones, digital broadcasting devices, personal digital assistants (PDAs), portable multimedia players (PMPs), desktops, wearable devices, embedded computing devices and / or tablet PCs.

[0029] Such a user computing device 110 includes at least one or more processors 111 and memory 112, where the processor 111 may comprise at least one or a plurality of electrically connected processors including central processing units (CPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.

[0030] The memory 112 may include one or more non-transitory / transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, or a combination thereof, and may also include web storage of a server on the internet that performs memory storage functions. Such memory 112 can store data 113 and instructions 114 required by the at least one processor 111 to perform operations such as pre-training a vision transformer 120 or executing an application that includes a pre-trained vision transformer 120.

[0031] In one embodiment, the user computing device 110 can store at least one or more machine learning models (e.g., vision transformer 120).

[0032] In particular, the Vision Transformer 120 in one embodiment may be a variety of machine learning models, such as multiple neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models, or may be comprised of combinations thereof.

[0033] The neural networks may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, and / or other forms of neural networks.

[0034] In one embodiment, the user computing device 110 can receive at least one or more vision transformers 120 from the server computing system 130 via the network 170, store them in memory, and then execute the stored vision transformers 120 using the processor 111 to operate applications having a variety of vision-based tasks.

[0035] In another embodiment, the server computing system 130 may include at least one machine learning model (e.g., the Vision Transformer 140) to perform operations according to the Vision Transformer 140 and communicate data related thereto with the user computing device 110 to provide downstream tasks to the user using the Vision Transformer 140. For example, the user computing device 110 may perform downstream tasks including the Vision Transformer 140 by having the server computing system 130 use the Vision Transformer 140 over the web to provide output in response to user input. Alternatively, the Vision Transformers 120 and / or 140 may be implemented such that at least a portion of the Vision Transformers 120 and / or 140 executes on the user computing device 110 and the remainder executes on the server computing system 130.

[0036] The user computing device 110 may also include at least one or more input components that sense user input. For example, the user input components may include a touch sensor (e.g., a touch screen and / or a touch pad) that senses touch of a user's input medium (e.g., a finger or a stylus), an image sensor that senses user motion input, a microphone that senses user voice input, a button, a mouse, and / or a keyboard, etc. The user input components may also include an interface and an external controller (e.g., a mouse, a keyboard, etc.) when receiving input to an external controller via an interface.

[0037] The server computing system 130 includes at least one or more processors 131 and memory 132. Here, the processor 131 may be comprised of at least one or a plurality of electrically connected processors including central processing units (CPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.

[0038] The memory 132 may include one or more non-transitory / transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., or a combination thereof. Such memory 132 may store data 133 and instructions 134 necessary for the processor 131 to pre-train the vision transformer 140 or to perform various vision tasks (e.g., image detection, classification, segmentation, etc.) using the vision transformer 140.

[0039] In one embodiment, the server computing system 130 may be implemented to include at least one computing device. For example, the server computing system 130 may be implemented to operate multiple computing devices using a serial computing architecture, a parallel computing architecture, or a combination thereof. The server computing system 130 may also include multiple computing devices connected via a network.

[0040] The server computing system 130 may also store at least one or more vision transformers 140. For example, the server computing system 130 may include neural networks or / and other multi-layer nonlinear models as the vision transformers 140. Exemplary neural networks may include feedforward neural networks, deep neural networks, recursive neural networks, and convolutional neural networks.

[0041] The training computing system 150 includes at least one or more processors 151 and memory 152. Here, the processor 151 may be comprised of at least one or a plurality of electrically connected processors including central processing units (CPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.

[0042] The memory 152 may include one or more non-transitory / transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. Such memory 152 can store data 153 and instructions 154 required by the processor 151 to train the vision transformer.

[0043] For example, the training computing system 150 may include a model trainer 160 that pre-trains the vision transformers 120 and / or 140 stored on the user computing device 110 and / or server computing system 130 using various training or learning techniques, such as backpropagation of errors (according to the framework shown in FIG. 5).

[0044] For example, the model trainer 160 can update one or more parameters of the vision transformers 120 or / and 140 in a backpropagation manner based on a defined loss function.

[0045] In some implementations, performing backpropagation of errors may include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight reduction, dropout, knowledge distillation, etc.) to improve the generalization ability of the trained vision transformers 120 and / or 140.

[0046] In particular, model trainer 160 can train vision transformers 120 and / or 140 based on a set of training data. The training data may include different multi-modal data, such as images, audio samples, text, etc. Examples of image types that can be used include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images, and / or various other forms of imagery.

[0047] Such trainer data and input data for downstream tasks can be provided by the user computing device 110 or / and the server computing system 130. When the training computing device trains the vision transformer 120 on the specific data of the user computing device 110, the vision transformer 120 can be characterized as a personalized model.

[0048] Model trainer 160 then includes computer logic utilized to provide the desired functionality. Model trainer 160 may be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in one embodiment, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In another implementation, model trainer 160 includes one or more sets of computer-executable instructions stored on a computer-readable storage medium such as RAM, a hard disk, or optical or magnetic media.

[0049] Network 170 may include, but is not limited to, a 3GPP (3rd Generation Partnership Project) network, an LTE (Long Term Evolution) network, a WIMAX (World Interoperability for Microwave Access) network, the Internet, a LAN (Local Area Network), a Wireless LAN (Wireless Local Area Network), a WAN (Wide Area Network), a PAN (Personal Area Network), a Bluetooth (registered trademark) network, a satellite broadcasting network, an analog broadcasting network, and / or a DMB (Digital Multimedia Broadcasting) network.

[0050] In general, communications over network 170 may occur using any type of wired and / or wireless connection, using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0051] FIG. 2 illustrates an example block diagram of a computing device for pre-training a vision transformer and executing the pre-trained vision transformer according to a knowledge distillation framework in accordance with an embodiment of the present invention.

[0052] 2, the computing devices 100 included in the user computing device 110, the server computing system 130, and the training computing system 150 include multiple applications (e.g., Application 1 through Application N). Each application may include a machine learning library and one or more vision transformers. For example, the applications may include an image processing (e.g., detection, classification, segmentation, etc.) application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, a chat-bot application, etc.

[0053] In an embodiment, the computing device 100 may include a model trainer 160 for pre-training vision transformers, and may store and operate the pre-trained vision transformers to perform various vision tasks using the vision transformers on input data.

[0054] Each application on computing device 100 may communicate with numerous other components on computing device 100, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In one embodiment, each application may communicate with each device component using an API (e.g., a public API). In one embodiment, the API used by each application may be specific to that application.

[0055] FIG. 3 shows another example block diagram of a computing device 200 for pre-training a vision transformer and executing the pre-trained vision transformer according to a knowledge distillation framework in accordance with an embodiment of the present invention.

[0056] 3, computing device 200 includes multiple applications (e.g., Application 1 through Application N). Each application can communicate with a central intelligence layer. For example, the applications may include an image processing application, a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In one embodiment, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).

[0057] The central intelligence layer may then include multiple vision transformers. For example, as shown in FIG. 3, at least a portion of each vision transformer may be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications may share a single vision transformer. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer may be included within or implemented separately from the operating system of computing device 200.

[0058] The central intelligence layer can communicate with a central device data layer, which may be a centralized data repository for computing device 200. As shown in Figure 3, the central device data layer can communicate with numerous other components of computing device 200, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a proprietary API).

[0059] The techniques described herein can refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from said systems. It will be appreciated that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, and divisions of work and functionality between and among components. For example, the processes described herein can be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications can be implemented in a single system or systems distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0060] Hereinafter, the process in which the computing system 1000 pre-trains a vision transformer using a knowledge distillation framework will be described in detail with reference to FIGS. 4 to 6.

[0061] The vision transformer described in this invention refers to a vision-language model (VLP) pre-trained on a large-scale dataset to provide a joint representation between two heterogeneous forms of data: a vision-language (image-text) pair.

[0062] The vision transformer according to this embodiment may include a single-stream model that transforms input data in which images and text are combined, and a dual (multi) stream model that processes image text using separate encoders.

[0063] In the following embodiments, for convenience of work, we will describe a vision transformer with a dual-stream architecture that is pre-trained with a control target on an image-text matching dataset.

[0064] A vision transformer pre-training method according to an embodiment can facilitate pre-training on a dataset of contrasting image-text pairs using a self-distilled encoder.

[0065] 4 and 5, a vision transformer pre-training architecture according to an embodiment includes a text encoder 10, a teacher image encoder 20, and a student image encoder 30. The student image encoder 30 and the teacher image encoder 20 may include a multi-head self-attention layer and a feedforward network. The student image encoder 30 may further include a token sparsification layer.

[0066] Here, the image-text dataset for pre-training is an uncurated dataset, meaning data that has not been subjected to labeling or captioning, for example.

[0067] In an embodiment, in order to clearly verify the efficiency of the vision transformer pre-training method according to the present embodiment, the dataset for pre-training may include at least one of large-scale open source datasets CC (Conceptual Captions) 3M, YFCC (Yahoo Flickr Creative Commons) 15M, and 88M.

[0068] Additionally, downstream datasets for validating the performance of the vision transformer pre-trained by this embodiment may include zero-shot image-text in Flickr30K or / and MS-COCO.

[0069] Thereafter, the computing system 1000 can classify the image-text pairs of the pre-training dataset according to the batch size, map image-text pairs that have already been matched within the batch size to positive pairs, and match positively matched text of other images to negative pairs.

[0070] The computing system 1000 can then input the text to the text encoder 10 in batches to output a text feature representation (T).

[0071] The computing system 1000 may also patch batches of images and input the image patches to the teacher image encoder 20 to output a first image feature representation (I'). An overbar will be replaced with '.

[0072] Next, the computing system 1000 can map the text feature representation (T) and the first image feature representation (I') according to the already matched positive pairs and negative pairs to generate a first alignment metric (A').

[0073] The computing system 1000 then trains the output text feature representation (T) and the first alignment metric (A') according to the positive / negative criteria to which the image feature representation has already been mapped using a similarity control alignment method (e.g., InfonCE loss (a method that trains to maximize the similarity of positive pairs and minimize the similarity of negative pairs)), and performs learning on data pairs with hard-aligned labels.

[0074] At this time, in the process of training the first alignment metric (A') composed of the first image feature representation (I') for similarity alignment, the teacher image encoder 20 may be a momentum teacher model with a stop gradient (Momentum Teacher with Stop Gradient), and therefore, backpropagation (sg) to the teacher image encoder 20 can be blocked during similarity alignment.

[0075] In particular, the similarity may refer to the dot product between the image feature representation and the text feature representation (T).

[0076] Thereafter, the computing system 1000 can be trained using a loss function to make the spatial distance between positive feature representations closer and the spatial distance between negative feature representations greater for similarity alignment.

[0077] In other words, we can perform contrastive learning by defining the loss function so that it becomes equal to 1.

[0078]

number

[0079] For example, as described above, the computing system 1000 can apply the loss function InfoNCE Loss to the similarity metric for training.

[0080] The computing system 1000 can then input the batched image patches to the student image encoder 30 to output a second image feature representation (I).

[0081] At this time, the student image encoder 30 can accelerate pre-training by including a token sparsification layer to reconstruct patch tokens.

[0082] In detail, the student image encoder 30 calculates the attention value (self-attention) between image patches, and can discard tokens that are below a pre-defined standard depending on the calculated attention value between each image patch.

[0083] For example, the student image encoder 30 can discard inattentive tokens according to a fixed ratio (1-κ) depending on the attention value between each patch in the 4th, 7th, and 10th transformer layers of the self-attention layer, where κ is the token retention ratio.

[0084] Then, the computing system 1000 can generate second alignment metrics by mapping the text feature representation (T) and the second image feature representation (I) that has undergone token sparsification according to the already matched positive pairs and negative pairs.

[0085] Next, the computing system 1000 can perform knowledge distillation such that the second alignment metric (A) predicts the output value of the first alignment metric (A') aligned by similarity mapping, unlike existing knowledge distillation methods.

[0086] That is, the computing system 1000 can perform knowledge distillation by training the student image encoder 30 to match the second alignment metric (A) through soft alignment according to the first alignment metric (A'). In this case, the text encoder 10 may be a Momentum Teacher with Stop Gradient model that blocks backpropagation (sg) to the text encoder 10 during knowledge distillation.

[0087] In particular, the computing system 1000 can perform knowledge distillation such that the second alignment metric (A) follows the parameters of the first alignment metric (A').

[0088] Specifically, the computing system 1000 may update the parameters of the first alignment metric (A') using the second alignment metric (A) as an exponential moving average (EMA).

[0089] The computing system 1000 can repeat the training of the teacher image encoder 20 and the training of the student image encoder 30 and the parameter update process n times (e.g., 1 to 3 times). During the repeated process, the text encoder 10 and the teacher image encoder 20 can prevent backpropagation (SG) by using a stopping gradient to prevent mutual collapsing.

[0090] The calculation process for the pre-training will be described in detail below using specific formulas.

[0091] Specifically, a function A for the momentum teacher image encoder 20 with stopping gradients and a function A representing the first alignment metric (A') for the momentum text encoder 10 with stopping gradients are  ̄ ij and a function A representing the second alignment metric (A) ij can be defined as the following equation 2.

[0092]

number

[0093] where sg is the stopping gradient and I  ̄ j=f  ̄ I (x j I ) and I j =f I (x j I ) are the image feature representations for the j-th image output by the teacher image encoder 20 and the student image encoder 30, respectively, and T i =f T (x i T ) is the text feature representation (T) for the i-th text.

[0094] Also, A∈R N×N ̄ is the alignment metric for the image feature representation and the text feature representation, N is the batch size of image-text pairs, and sim is a function for cosine similarity.

[0095] Then, as described above, the loss for the first alignment metric (A') can be obtained using the InfoNCE loss (Equation 3 below).

[0096]

number

[0097] L T =L N (A) is the InfoNCE loss, and in the following, we use A in Equation 2 according to Equation 3. ij =sim(T i ,I j ) is the InfoNCE loss for L CLIP (A  ̄ ) is defined as

[0098] Next, as described above, the computing system 1000 performs knowledge distillation to predict that the second alignment metric (A) will match the first alignment metric (A').

[0099] In detail, when we define the distillation loss as the KL divergence for each row and column between the first alignment metric (A') and the second alignment metric (A), and σ is the softmax function, the KL divergence between the first alignment metric (A') and the second alignment metric (A) is D KL (A  ̄ ||A) can be expressed as the following equation 4, and the distillation loss can be calculated using equation 5.

[0100]

number

[0101] Here, the overall distillation loss is L distill (A  ̄ ,A) is the average of the KL loss for the row vector and column vector of the first alignment metric parameters and the second alignment metric parameters, respectively, and can be defined as follows:

[0102]

number

[0103] Then, to accelerate the training of the student image encoder 30, we balance the knowledge distillation training with the teacher image encoder 20 training, and the final loss L of the student image encoder 30 is student is defined by the following equation (6), and therefore the final loss L is defined by the following equation (7).

[0104]

number

[0105]

number

[0106] Here, λ is a parameter that balances the KL divergence loss and the InfoNCE loss, and is calculated based on the exponential moving average (ema) in the embodiment.

[0107] As described above, the teacher image encoder 20 and the text encoder 10 can perform updates using stopping gradients to prevent backpropagation.

[0108] For details, θ fI and θ f ̄I are the parameters of the student encoder and the momentum teacher image encoder 20, respectively, and θ f ̄I (t) The update can be performed by the following equation 8.

[0109]

number

[0110] Experimental results showed that the most efficient training was achieved when m was 0.994.

[0111] That is, by the steps of Equations 2 to 8, the encoder of the vision transformer can be pre-trained by contrastive instruction and knowledge distillation.

[0112] Below, we will describe a comparison of the effect of pre-training a vision transformer using knowledge distillation according to an embodiment of the present invention with existing techniques.

[0113] For comparison, refer to FIG. 6, which is a graph comparing the existing pre-training method (EViT) in which token sparsification is directly applied to the existing contrastive language-image pre-training method, and the vision transformer pre-training method using knowledge distillation (ECLIPSE) according to the present invention.

[0114] The vision transformer pre-training method (ECLIPSE) according to an embodiment of the present invention trains a simplified vision transformer with processing speed approximately 101% faster than the existing pre-training method (EViT), and it has been confirmed that the performance of the pre-trained vision transformer is also relatively superior in terms of zero-shot image accuracy. Furthermore, it has been confirmed that it exhibits superior performance in pre-training speed, vision transformer capacity, and zero-shot accuracy compared to CLIP (Contrastive Language-Image Pretraining), a representative contrastive language-image pre-training model that does not apply token sparsification. Here, the backbone used to compare the performance of each model is ViT-B / 16.

[0115] [Table 1]

[0116] More specifically, Table 1 shows the zero-shot accuracy of the CLIP model and the ECLIPSE model of the present invention for ImageNet-1k Top-1. It can be seen that the accuracy of the ECLIPSE model is superior to that of the CLIP model.

[0117] In particular, Figure 7 shows the attention level for each patch token calculated for the CLIP model and the ECLIPSE model of the present invention, and it can be seen that the ECLIPSE model more accurately identifies patches related to objects (people) that need to be carefully observed in the image. In other words, it can be seen that the ECLIPSE model shows superior performance in image segmentation tasks due to token sparsification.

[0118] Therefore, the computing system 1000 can utilize the vision transformer of the present invention to perform a variety of vision tasks more efficiently and with more accurate performance than existing models.

[0119] For example, the vision transformer of the present invention can perform vision tasks such as image classification, segmentation, object detection, image generation, automatic caption generation, image retrieval, and image description.

[0120] Additionally, the computing system 1000 can execute various applications, including a vision transformer that has excellent performance for such vision tasks, to perform various artificial intelligence tasks.

[0121] Furthermore, this framework, which utilizes token sparsification and knowledge distillation for contrastive language-image pre-training, could be extended to pre-training for additional modalities such as audio at the level of a layperson.

[0122] The above-described embodiments of the present invention may be embodied in the form of program instructions executable by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include, alone or in combination, program instructions, data files, data structures, and the like. The program instructions recorded on the computer-readable recording medium may be specially designed and constructed for the present invention, or may be known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include not only machine language code, such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc. A hardware device may be replaced by one or more software modules to perform processes according to the present invention, and vice versa. [Industrial Applicability]

[0123] The present invention has industrial applicability because it is an invention that provides the basis for a method for pre-training a vision transformer using a computing system and for performing tasks such as vision tasks using the pre-trained vision transformer.

Claims

1. 1. A method for pre-training a vision transformer on a vision-language dataset, wherein a computing system including a memory and a processor comprises: obtaining a dataset consisting of a plurality of image-text pairs; inputting the n-th batch of text in the acquired dataset into a text encoder to output text feature representations; inputting an n-th batch of images in the acquired dataset into a teacher image encoder to output a first image feature representation; inputting the nth batch of images in the acquired dataset into a student image encoder to output a second image feature representation; generating a first alignment metric for the text feature representation and the first image feature representation; training the first alignment metric such that the text feature representation and the first image feature representation are aligned to have similarity according to a positive and negative mapping relationship of the n batches of image-text pairs; and performing knowledge distillation on the second alignment metric to predict an output of the learned first alignment metric. A method for pre-training vision transformers via knowledge distillation.

2. The step of inputting the image data to the teacher image encoder and outputting a first image feature representation includes: patching the nth batch of images; and outputting a first image feature representation by passing the patched image patches through a plurality of self-attention layers and a feedforward network layer. The method for pre-training a vision transformer by knowledge distillation according to claim 1.

3. The step of inputting the student image encoder and outputting a second image feature representation includes: and inputting the patched image patches into a token sparsification layer based on output values ​​from a plurality of self-attention layers to perform token sparsification. The method for pre-training a vision transformer by knowledge distillation according to claim 2.

4. learning the first alignment metric such that the text feature representation and the first image feature representation are aligned; determining a positive feature representation pair and a negative feature representation pair between the text feature representation and the first image feature representation according to a mapping relationship of the image-text pair; and training the encoder with a loss function that encourages the positive feature representation pairs to be closer and the negative feature representation pairs to be farther apart for similarity alignment. The method for pre-training a vision transformer by knowledge distillation according to claim 1.

5. training an encoder using the loss function, When learning for similarity alignment using the loss function, applying a momentum stopping gradient to the teacher image encoder to block backpropagation. The method for pre-training a vision transformer by knowledge distillation according to claim 4.

6. The step of performing knowledge distillation on the second alignment matrix includes: knowledge distillation to predict an output value of a first alignment metric according to the similarity alignment; The method for pre-training a vision transformer by knowledge distillation according to claim 5.

7. The step of performing knowledge distillation on the second alignment matrix includes: blocking backpropagation to the text encoder during the knowledge distillation. The method for pre-training a vision transformer by knowledge distillation according to claim 6.

8. The step of performing knowledge distillation on the second alignment matrix includes: knowledge distillation such that parameters of the second alignment metric follow parameters of the first alignment metric; The method for pre-training a vision transformer by knowledge distillation according to claim 7.

9. The step of knowledge distilling such that parameters of the second alignment metric follow parameters of the first alignment metric comprises: updating parameters of the second alignment metric to parameters of the first alignment metric with an exponential moving average (EMA); The method for pre-training a vision transformer by knowledge distillation according to claim 8.

10. The distillation loss between the function A' of the first alignment metric and the function A of the second alignment metric is the KL divergence D KL (A  ̄ ||A), and the KL divergence D KL (A  ̄ ||A) is defined by the following equation 4: The method for pre-training a vision transformer by knowledge distillation according to claim 1. [Equation 4] where σ is the softmax function.

11. The step of performing knowledge distillation on the second alignment matrix includes: Overall distillation loss L distill (A  ̄ , A) is the average of the KL loss for the row vector and the column vector, and is characterized by being defined as the following equation 5: The method for pre-training a vision transformer by knowledge distillation according to claim 10. [Equation 5]

12. The final loss L of the student image encoder distill (A  ̄ , A) are defined by the following equation 6: The final loss L of the encoder including the teacher image encoder and the student image encoder is defined by the following equation (7): The method for pre-training a vision transformer by knowledge distillation according to claim 11. [Equation 6] Here, λ is a parameter that balances the KL divergence loss and the InfoNCE loss, and is set based on the exponential moving average (ema) in the embodiment. [Equation 7]

13. and performing a stopping gradient update to prevent backpropagation between the teacher image encoder and the text encoder. The method for pre-training a vision transformer by knowledge distillation according to claim 12.

14. A vision task execution model comprising a vision transformer pre-trained according to claim 1.

15. a computing system including a memory and a processor, the computing system including a pre-trained vision transformer; a text encoder that inputs the nth batch of text and outputs text feature representations; a teacher image encoder that receives an nth batch of images and outputs a first image feature representation; a student image encoder that inputs the nth batch of images and outputs a second image feature representation; and a first alignment metric trained to similarity align the text feature representation and the first image feature representation; a second alignment metric for the text feature representation and the second image feature representation is knowledge distilled to predict an output value of the first alignment metric; Pre-trained vision transformers via knowledge distillation.