Vision transformer device and method using token merging
The vision transformer device accelerates inference by merging tokens based on similarity across multiple transformer blocks, addressing the delay in inference time and maintaining high accuracy without additional model learning.
Patent Information
- Application Number
- PCT/KR2024/018690
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-20
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-05
AI Technical Summary
Vision Transformers face delays in inference time due to their large model size, and existing token pruning techniques lead to information loss and require fine-tuning to maintain accuracy.
A vision transformer device and method that accelerates inference by comparing similarity between adjacent tokens using image data features and performing token merging at once across multiple transformer blocks without additional model learning.
This approach improves overall inference speed while maintaining high accuracy by reducing the number of tokens through serial merging, thereby minimizing information loss and the need for fine-tuning.
Smart Images

Figure KR2024018690_05062025_PF_FP_ABST
Abstract
Description
Vision transformer device and method using token merging
[0001] The present invention relates to a vision transformer device and method using token merging, and more particularly, to a technology for implementing acceleration of inference related to image recognition and processing based on a vision transformer by performing token merging at once without additional model learning by comparing similarities between adjacent tokens using features of image data.
[0002] This application is the result of research conducted with the support of the National IT Industry Promotion Agency (NIPA) (No. RS-2022-00155966, Artificial Intelligence Convergence Innovation Talent Development (Ewha Womans University), National Project Number: 1711179344, No. 2021-0-02068, Artificial Intelligence Innovation Hub Research and Development) with funding from the government (Ministry of Science and ICT) in 2024.
[0003] With the emergence of the transformer model, which demonstrates powerful performance in identifying and generating sentence structures and semantic relationships, transformer-based models are showing remarkable performance in various fields such as natural language processing, computer vision, and speech recognition.
[0004] In particular, the vision transformer (ViT) model shows superior performance compared to classical CNN models in image classification tasks.
[0005] However, despite the excellent accuracy of the Vision Transformer model, the Vision Transformer has the problem of delay in inference time due to the large model size.
[0006] To solve this problem, studies have been proposed to apply the token pruning technique to vision transformers, which reduces the amount of computation required by pruning unnecessary tokens within the image.
[0007] However, in conventional technology, information loss occurs as tokens are removed, resulting in reduced accuracy and requiring fine-tuning to improve model performance.
[0008] To overcome the problems of token pruning techniques, token merging was proposed.
[0009] In conventional technology, token merging gradually combines similar tokens to reduce the total number of tokens and prevent loss of information.
[0010] However, there is a limit to the acceleration of inference speed as incremental merging across multiple transformer blocks is required.
[0011] The present invention aims to provide a vision transformer device and method that implements acceleration of inference related to image recognition and processing based on a vision transformer by comparing similarities between adjacent tokens by utilizing the features of image data and performing token merging at once without additional model learning.
[0012] The present invention aims to provide a vision transformer device and method that maintains high accuracy while improving the overall inference speed by performing chain merging of tokens by comparing the similarity with adjacent tokens based on one of a plurality of tokens constituting a token matrix in one of a plurality of transformer blocks.
[0013] A vision transformer device according to one embodiment of the present invention may include an input processing unit that divides a recognition target image into a plurality of patches and processes any one of the plurality of patches into input data in the form of a token, a transformer encoding unit that performs transformer encoding on the input data using a plurality of transformer blocks, and performs token merging on selected tokens based on a similarity between a reference token selected from among a plurality of tokens constituting a token matrix based on the input data and an adjacent token and a reference similarity in only one of the plurality of transformer blocks, and an inference processing unit that performs object recognition inference on the recognition target image based on a result of token merging processing output through the plurality of transformer blocks.
[0014] The transformer encoding processing unit may select at least one similar token based on the similarity between the selected reference token and a token adjacent to at least one of the above, below, left, and right of the selected reference token and the similarity based on the similarity based on the similarity based on the similarity based on the above, below, left, and right of the selected at least one similar token, and may select at least one additional similar token based on the similarity between the above, below, left, and right of the selected at least one similar token and the similarity based on the similarity based on the similarity based on the above, below, left, and right of the selected at least one similar token, and may store information about the at least one additional similar token, and may merge the selected reference token, the above, at least one similar token, and the above, at least one additional similar token at a time to output the result of the token merging processing through the plurality of transformer blocks.
[0015] The above transformer encoding processing unit may perform an operation on an attention distribution for tokens constituting the token matrix, and, based on the calculated attention distribution and the spatial locality of the tokens, compare the similarity between the selected reference token and a token adjacent to at least one of the upper, lower, left, and right of the selected reference token with the reference similarity, and select a token having a similarity determined by the comparison to be greater than the reference similarity as the at least one similar token.
[0016] The above transformer encoding processing unit can calculate the attention distribution by applying a softmax function to the attention scores for the tokens constituting the token matrix.
[0017] The above transformer encoding processing unit can set the reference similarity to decrease so as to increase the token merging ratio for tokens constituting the token matrix, and can set the reference similarity to increase so as to decrease the token merging ratio.
[0018] The above transformer encoding processing unit can set the reference similarity to decrease by considering the background image ratio in the recognition target image, and can set the reference similarity to increase by considering the number of objects to be identified in the recognition target image.
[0019] The above-mentioned inference processing unit can perform the object recognition inference by classifying a class for an object in the recognition target image through head processing of a multilayer perceptron based on the reduced number of tokens constituting the token matrix as adjacent backgrounds with uniform colors are merged into one token in the token merging processing result.
[0020] According to one embodiment of the present invention, a vision transformer method may include a step of, in an input processing unit, dividing a recognition target image into a plurality of patches and processing any one of the plurality of patches into input data in the form of a token; a step of, in a transformer encoding unit, performing transformer encoding on the input data using a plurality of transformer blocks, wherein, in only one of the plurality of transformer blocks, token merging is performed on selected tokens based on a similarity between a selected reference token and an adjacent token and a reference similarity among a plurality of tokens constituting a token matrix based on the input data; and a step of, in an inference processing unit, performing object recognition inference on the recognition target image based on a result of token merging processing output through the plurality of transformer blocks.
[0021] The step of performing token merging on selected tokens based on the similarity between a selected reference token and adjacent tokens among the plurality of tokens constituting a token matrix based on the input data and the reference similarity in only one transformer block among the plurality of transformer blocks may include the steps of selecting at least one similar token based on the similarity between the selected reference token and a token adjacent to at least one of the above, below, left, and right of the selected reference token and the reference similarity, storing information about the selected similar token, selecting at least one additional similar token based on the similarity between the at least one selected similar token and a token adjacent to at least one of the above, below, left, and right and the reference similarity, storing information about the at least one additional similar token, and merging the selected reference token, the at least one selected similar token, and the at least one selected additional similar token at a time to output the token merging processing result through the plurality of transformer blocks.
[0022] The step of performing token merging on selected tokens based on the similarity between a selected reference token and adjacent tokens and the reference similarity among the plurality of tokens constituting the token matrix based on the input data in only one of the plurality of transformer blocks may include the step of performing an attention distribution operation on the tokens constituting the token matrix, comparing the similarity between the selected reference token and a token adjacent to at least one of the upper, lower, left, and right of the selected reference token and the reference similarity based on the calculated attention distribution and the spatial locality of the tokens, and selecting a token whose compared similarity is greater than the reference similarity as the at least one similar token.
[0023] The present invention can provide a vision transformer device and method that implements acceleration of inference related to image recognition and processing based on a vision transformer by comparing similarities between adjacent tokens using the characteristics of image data and performing token merging at once without additional model learning.
[0024] The present invention can provide a vision transformer device and method that improves the overall inference speed while maintaining high accuracy in performing chain merging of tokens by comparing the similarity with adjacent tokens based on one of a plurality of tokens constituting a token matrix in one of a plurality of transformer blocks.
[0025] FIG. 1 is a drawing illustrating a vision transformer device using token merging according to one embodiment of the present invention.
[0026] FIG. 2 is a drawing illustrating a token merging block of a transformer encoding unit in a vision transformer device according to an embodiment of the present invention.
[0027] FIG. 3 and FIG. 4 are drawings explaining a vision transformer method using token merging according to one embodiment of the present invention.
[0028] FIGS. 5A to 5E are drawings explaining a token merging procedure in a transformer encoding unit of a vision transformer device according to one embodiment of the present invention.
[0029] FIG. 6 is a drawing illustrating an additional token merging procedure in a transformer encoding unit of a vision transformer device according to one embodiment of the present invention.
[0030] FIG. 7 is a drawing explaining the result of token merging processing in a transformer encoding unit in a vision transformer device according to one embodiment of the present invention.
[0031] FIG. 8 and FIG. 9 are drawings explaining simulation results for inference performance in a vision transformer device according to one embodiment of the present invention.
[0032] Below, various embodiments of this document are described with reference to the attached drawings.
[0033] The examples and terms used herein are not intended to limit the technology described in this document to a particular embodiment, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments.
[0034] In the following description of various embodiments, if it is determined that a detailed description of a related known function or configuration may unnecessarily obscure the gist of the invention, the detailed description will be omitted.
[0035] The terms described below are defined based on their functions in various embodiments, and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification.
[0036] In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0037] A singular expression may include a plural expression unless the context clearly indicates otherwise.
[0038] In this document, expressions such as "A or B" or "at least one of A and / or B" may include all possible combinations of the items listed together.
[0039] Expressions such as "first," "second," "first," or "second," may modify the components without regard to order or importance, and are used only to distinguish one component from another, but do not limit the components.
[0040] When it is said that a component (e.g., a first component) is “(functionally or communicatively) connected” or “connected” to another component (e.g., a second component), the component may be directly connected to the other component, or may be connected via another component (e.g., a third component).
[0041] In this specification, “configured to” may be used interchangeably with “suitable for,” “capable of,” “modified to,” “made to,” “capable of,” or “designed to,” depending on the context, for example, in terms of hardware or software.
[0042] In some contexts, the expression "a device configured to" may mean that the device is "capable of" doing something in conjunction with other devices or components.
[0043] For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0044] Also, the term 'or' means 'inclusive or' rather than 'exclusive or'.
[0045] That is, unless otherwise stated or clear from context, the expression 'x utilizes a or b' means any one of the natural inclusive permutations.
[0046] The terms '..bu', '..gi', etc. used below mean a unit that processes at least one function or operation, and this can be implemented by hardware, software, or a combination of hardware and software.
[0047] FIG. 1 is a drawing illustrating a vision transformer device using token merging according to one embodiment of the present invention.
[0048] FIG. 1 illustrates components of a vision transformer device using token merging according to one embodiment of the present invention.
[0049] Referring to FIG. 1, a vision transformer device (100) using token merging according to one embodiment of the present invention includes an input processing unit (110), a transformer encoding unit (120), and an inference processing unit (130).
[0050] For example, the vision transformer device (100) is a deep learning model that divides an input image into small pieces, infers an object located in each piece, and recognizes the entire image based on the object inferred from each piece.
[0051] A vision transformer device (100) according to one embodiment of the present invention may be a device based on a deep learning model that performs object inference and image recognition without additional model learning by reducing the time consumed for token merging based on a framework that merges tokens at once.
[0052] The vision transformer device (100) according to one embodiment of the present invention performs token merging at once before deep learning model inference, thereby reducing the time consumed for token merging compared to the conventional technology that performs merging for each transformer block, thereby reducing the inference time.
[0053] An input processing unit (110) according to one embodiment of the present invention can divide a recognition target image into a plurality of patches and process any one of the plurality of patches into input data in the form of a token.
[0054] For example, the input processing unit (110) can divide an image into patches and perform data processing in the form of tokens as input data for application to a vision transformer model, which is a deep learning model that performs various image recognition and processing tasks.
[0055] For example, the input processing unit (110) performs input data processing for token merging to reduce the number of tokens and prevent information loss.
[0056] According to one embodiment of the present invention, a transformer encoding unit (120) performs transformer encoding on input data using a plurality of transformer blocks, and performs token merging on selected tokens based on the similarity between a selected reference token and an adjacent token and the reference similarity among a plurality of tokens constituting a token matrix based on the input data in only one of the plurality of transformer blocks.
[0057] The transformer encoding processing unit (120) can select at least one similar token based on the similarity between the selected reference token and a token adjacent to at least one of the upper, lower, left, and right sides of the selected reference token and the reference similarity.
[0058] The transformer encoding processing unit (120) stores information about the selected similar tokens, and can select at least one additional similar token based on the similarity between at least one selected similar token and at least one adjacent token from above, below, left, and right, and the above-mentioned standard similarity.
[0059] The transformer encoding processing unit (120) stores information about at least one additional similar token, and can merge the selected reference token, the selected at least one similar token, and the selected at least one additional similar token at once to output the token merging processing result through multiple transformer blocks.
[0060] The transformer encoding processing unit (120) performs an operation on an attention distribution for tokens constituting a token matrix, and compares the similarity between a selected reference token and a token adjacent to at least one of the upper, lower, left, and right sides of the selected reference token based on the calculated attention distribution and the spatial locality of the tokens, and the reference similarity, and selects a token whose compared similarity is greater than the reference similarity as at least one similar token.
[0061] The transformer encoding processing unit (120) can calculate the attention distribution by applying a softmax function to the attention scores for tokens constituting the token matrix.
[0062] The transformer encoding processing unit (120) can set a reduced reference similarity to increase the token merging ratio for tokens constituting the token matrix, and can set an increased reference similarity to decrease the token merging ratio.
[0063] The transformer encoding processing unit (120) can set a reduced reference similarity by considering the background image ratio in the recognition target image, and can set an increased reference similarity by considering the number of objects to be identified in the recognition target image.
[0064] Unlike the conventional technology that performs token merging for each transformer block, the transformer encoding processing unit (120) can reduce the overall time required for merging by performing token merging for only one of a plurality of transformer blocks.
[0065] The transformer encoding processing unit (120) can adjust the degree of token merging based on the standard similarity according to the settings, and the smaller this value, the more tokens are merged.
[0066] The fact that adjacent patches on an image are more likely to have similar backgrounds or features can be exploited.
[0067] Adjacent backgrounds with uniform colors can be merged at once, and based on this, only adjacent tokens in the vertical, horizontal, and right directions can be compared to significantly reduce the number of similarity comparisons between tokens.
[0068] One token merging according to one embodiment of the present invention repeatedly performs the process of visiting adjacent tokens for all tokens present in a given image and then merging similar tokens.
[0069] That is, a single token merge recursively repeats the process of comparing the similarity of a given token with the four adjacent tokens in the vertical, horizontal, and right directions.
[0070] The inference processing unit (130) can perform object recognition inference on a recognition target image based on the token merging processing results output through multiple transformer blocks.
[0071] The inference processing unit (130) can perform object recognition inference to classify a class for an object in a recognition target image through head processing of a multilayer perceptron based on the reduced number of tokens constituting the token matrix as adjacent backgrounds with uniform colors are merged into one token in the token merging processing result.
[0072] Accordingly, the present invention can provide a vision transformer device and method that implements acceleration of inference related to image recognition and processing based on a vision transformer by comparing similarities between adjacent tokens by utilizing the features of image data and performing token merging at once without additional model learning.
[0073] FIG. 2 is a drawing illustrating a token merging block of a transformer encoding unit in a vision transformer device according to an embodiment of the present invention.
[0074] FIG. 2 illustrates and explains in more detail the procedures performed in each token merging block of the transformer encoding unit in a vision transformer device according to one embodiment of the present invention.
[0075] Referring to FIG. 2, the transformer encoding unit (200) of the present invention includes a first transformer block (210), a second transformer block (220), and a third transformer block (230).
[0076] The number of the first transformer block (210), the second transformer block (220), and the third transformer block (230) may be increased or decreased as an example.
[0077] The transformer encoding unit (200) of the present invention performs token merging only in the first transformer block (210), thereby providing similar accuracy and improved inference speed as the prior art that performs token merging in all of the first transformer block (210), the second transformer block (220), and the third transformer block (230).
[0078] In step (S201), the transformer encoding unit (200) selects a reference token and compares the similarity of the selected reference token with four adjacent tokens based on the reference similarity.
[0079] In step (S202), the transformer encoding unit (200) selects a token among four adjacent tokens whose similarity is higher than the reference similarity.
[0080] In step (S203), the transformer encoding unit (200) additionally selects a reference token as an extension target among the selected adjacent tokens and then performs a similarity comparison as in step (S201).
[0081] In step (S204), the transformer encoding unit (200) selects a token among four adjacent tokens whose similarity is higher than the standard similarity, similar to step (S202).
[0082] That is, the transformer encoding unit (200) compares the similarity of one token with the four adjacent tokens above, below, left, and right.
[0083] In relation to spatial locality, similar characteristics are utilized between adjacent tokens on the image.
[0084] Similarity criteria values selected based on attention distribution are used for similarity comparison.
[0085] In a vision transformer device according to an embodiment of the present invention, each procedure performed in a token merging block of a transformer encoding unit improves the overall inference speed while maintaining high accuracy without additional model learning.
[0086] The present invention considers token similarity comparison and spatial locality with a small number of times.
[0087] The transformer encoding unit reduces the token merging time by performing token merging in only one of multiple transformer blocks.
[0088] Preferably, token merging is performed in the first transformer block.
[0089] FIG. 3 and FIG. 4 are drawings explaining a vision transformer method using token merging according to one embodiment of the present invention.
[0090] FIG. 3 illustrates a procedure for performing vision transformer acceleration using a vision transformer method using token merging according to one embodiment of the present invention.
[0091] Referring to FIG. 3, in step (S301), a vision transformer method using token merging according to an embodiment of the present invention processes input data.
[0092] That is, a vision transformer method using token merging according to one embodiment of the present invention can divide a recognition target image into a plurality of patches and process any one of the plurality of patches as input data in the form of a token.
[0093] In step (S302), a vision transformer method using token merging according to an embodiment of the present invention performs token merging in only one transformer block among a plurality of transformer blocks.
[0094] That is, a vision transformer method using token merging according to an embodiment of the present invention performs transformer encoding on input data using a plurality of transformer blocks, and performs token merging on selected tokens based on the similarity between a selected reference token and an adjacent token and the reference similarity among a plurality of tokens constituting a token matrix based on the input data in only one of the plurality of transformer blocks.
[0095] In step (S303), a vision transformer method using token merging according to an embodiment of the present invention performs object recognition inference based on the result of token merging processing.
[0096] That is, a vision transformer method using token merging according to one embodiment of the present invention can perform object recognition inference in a recognition target image based on the token merging processing results output through a plurality of transformer blocks.
[0097] FIG. 4 illustrates a token merging at once (ToMato) procedure related to vision transformer acceleration using a vision transformer method using token merging according to one embodiment of the present invention.
[0098] Referring to FIG. 4, in step (S401), a vision transformer method using token merging according to an embodiment of the present invention inputs a token matrix.
[0099] That is, the vision transformer method using token merging according to one embodiment of the present invention can input a token matrix that has been converted from an input image that is a recognition target.
[0100] In step (S402), a vision transformer method using token merging according to an embodiment of the present invention calculates an attention distribution.
[0101] That is, the vision transformer method using token merging according to one embodiment of the present invention can calculate the attention distribution by applying the attention score and softmax function.
[0102] In step (S403), a vision transformer method using token merging according to an embodiment of the present invention compares the similarity between a selected reference token and an adjacent token.
[0103] That is, the vision transformer method using token merging according to one embodiment of the present invention compares the similarity between the current reference token and four adjacent tokens.
[0104] In step (S404), the vision transformer method using token merging according to one embodiment of the present invention determines whether the similarity of the comparison token is greater or less than the reference similarity.
[0105] That is, the vision transformer method using token merging according to one embodiment of the present invention can determine whether there is a token among adjacent tokens having a similarity greater than a reference similarity, and if there is a comparison token having a similarity greater than the reference similarity, proceeds to step (S405), and if it is less than the reference similarity, proceeds to step (S406).
[0106] In step (S405), a vision transformer method using token merging according to an embodiment of the present invention stores comparison token information and selects the comparison token as a reference token.
[0107] That is, the vision transformer method using token merging according to one embodiment of the present invention stores token information when the similarity is greater than the reference similarity.
[0108] In step (S406), the vision transformer method using token merging according to one embodiment of the present invention merges all previously stored tokens.
[0109] That is, the vision transformer method using token merging according to one embodiment of the present invention can store all tokens corresponding to token information stored at once, as determined by the similarity being greater than the reference similarity.
[0110] In step (S407), the vision transformer method using token merging according to one embodiment of the present invention checks whether merging has been completed for all tokens.
[0111] That is, the vision transformer method using token merging according to one embodiment of the present invention sequentially determines whether merging is complete by performing a new merging process by moving on to the next token when merging for a specific token is completed based on the result of the standard similarity determination, and if merging is complete, proceeds to step (S408), and if merging is not complete, proceeds to step (S403).
[0112] In step (S408), the vision transformer method using token merging according to one embodiment of the present invention outputs the token merging processing result.
[0113] That is, the vision transformer method using token merging according to one embodiment of the present invention can output a token merging processing result in which merging of tokens in a token matrix is completed.
[0114] Accordingly, the present invention can provide a vision transformer device and method that improves the overall inference speed while maintaining high accuracy in performing chain token merging by comparing the similarity with adjacent tokens based on one of a plurality of tokens constituting a token matrix in one of a plurality of transformer blocks.
[0115] FIGS. 5A to 5E are drawings explaining a token merging procedure in a transformer encoding unit of a vision transformer device according to one embodiment of the present invention.
[0116] Referring to Fig. 5a, it shows a token matrix (500) and an attention distribution (501) for the token matrix (500).
[0117] With respect to the token matrix (500), after multi-head attention is completed, an attention distribution (501) is calculated as a distribution of the results of attention performance among tokens excluding class tokens and distillation tokens.
[0118] Referring to Fig. 5b, similarity comparison between adjacent tokens is performed in the attention distribution (511) starting from the i-th token, which is the reference token in the token matrix (510).
[0119] The attention distribution (511) shows that the solid line is similar and the dotted line shows that there is a difference.
[0120] Referring to Figure 5c, the result of selecting adjacent tokens excluding token (i+14) from among adjacent tokens starting from the i-th token, which is the reference token, in the token matrix (520) is shown.
[0121] The attention distribution (521) shows that the token corresponding to the dotted line is not selected from the token matrix (520).
[0122] That is, it can be confirmed that only tokens exceeding the sim (similarity threshold which affects merging rate) value corresponding to the similarity criterion in the token matrix (520) and attention distribution (521) are selected.
[0123] For example, in relation to similarity comparison, the value of the similarity criterion may be set to 0.1e-7.
[0124] Referring to FIG. 5d, the token matrix (530) and attention distribution (531) show recursive similarity comparisons based on newly merged tokens.
[0125] Referring to FIG. 5e, the token matrix (540) and attention distribution (541) show that token information is stored after recursively comparing similarity based on newly merged tokens.
[0126] The token merging procedure according to the images and tables in FIGS. 5a to 5e is described in an integrated manner.
[0127] In a token merging procedure according to one embodiment of the present invention, token merging begins immediately after multi-head attention is applied in the first decoder, and a procedure is performed to determine a token to be merged for four adjacent tokens in the upper, lower, left, and right directions.
[0128] The same process of merging four adjacent tokens based on the tokens to be merged is repeated.
[0129] After recursively finding similar tokens, when no more similar tokens can be found among adjacent tokens, all found similar tokens are merged into one.
[0130] The recursive token merging process starts from the first token and proceeds to the last token.
[0131] FIG. 6 is a drawing illustrating an additional token merging procedure in a transformer encoding unit of a vision transformer device according to one embodiment of the present invention.
[0132] FIG. 6 illustrates an image-based example of an additional token merging procedure in a transformer encoding unit of a vision transformer device according to an embodiment of the present invention.
[0133] Referring to FIG. 6, in the image (600), the portion (601) is expanded to an adjacent token, and in the image (610), the portion (611) is merged.
[0134] As we move from image (620) to the next token, a new merging process is performed with part (621).
[0135] If there is no more token among adjacent tokens that exceeds the similarity criterion, the stored tokens are merged into one token, and then the merging process is repeated by moving on to the next token. If the token has already been merged, the process proceeds to the next part (621).
[0136] The merging is determined to be complete from the image (620) to the part (621), and the merging process for the entire token is performed by merging from the image (630) to the part (631).
[0137] In relation to merging for the entire token, the result of the token merging process is provided as an image (640) of the actual final result.
[0138]
[0139] FIG. 7 is a drawing explaining the result of token merging processing in a transformer encoding unit in a vision transformer device according to one embodiment of the present invention.
[0140] FIG. 7 illustrates a token merging processing result in a transformer encoding unit of a vision transformer device according to an embodiment of the present invention.
[0141] Referring to FIG. 7, an input recognition target image (700) is shown, and a token merging processing result (710) output by applying the recognition target image (700) to the present invention is shown.
[0142] The token merging processing result (710) represents the result of effectively merging patches semantically.
[0143] FIG. 8 and FIG. 9 are drawings explaining simulation results for inference performance in a vision transformer device according to one embodiment of the present invention.
[0144] FIG. 8 compares the latency and accuracy of a vision transformer device according to an embodiment of the present invention with those of a prior art device in relation to simulation results for inference performance.
[0145] Referring to FIG. 8, the graph (800) shows a guide line (801) related to the present invention and a guide line (802) related to the prior art.
[0146] Comparing the indicator line (801) and the indicator line (802), it can be confirmed that the accuracy decrease due to the decrease in delay speed of the present invention is slight.
[0147] That is, it can be confirmed that the accuracy of the guide line (802) varies greatly depending on the inference time, but the accuracy of the guide line (801) does not vary greatly even if the inference time decreases.
[0148] FIG. 9 illustrates the values of the similarity reference diagram and the latency related to the inference speed in relation to the simulation results for the inference performance in a vision transformer device according to one embodiment of the present invention.
[0149] Referring to FIG. 9, a graph (900) shows the relationship between the value of the similarity reference diagram and the delay speed related to the inference speed with a guide line (901) related to the present invention.
[0150] That is, the graph (900) shows that as the similarity criterion value decreases, the inference speed increases.
[0151] The more uniform the background or features in the image, the higher the merging rate, as related to the similarity criterion value.
[0152] The more important or complex features an image has in relation to the similarity criterion value, the lower the merging ratio.
[0153] For example, the present invention uses a similarity criterion value of 10 -7 This reduces the total number of tokens from approximately 8.16% to 20.92%.
[0154] When the value representing the number of tokens merged in each transformer block is set to "2", approximately 15% of tokens are ultimately merged.
[0155] Accordingly, the present invention can flexibly set the delay speed and merging range according to the similarity criterion value setting.
[0156] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0157] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may be distributed on network-connected computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0158] Although the embodiments described above have been described with limited drawings, those skilled in the art will recognize that various modifications and variations can be made based on the above description. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0159] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. An input processing unit that divides a recognition target image into multiple patches and processes one of the multiple patches into input data in the form of a token; A transformer encoding unit that performs transformer encoding on the input data using a plurality of transformer blocks, and performs token merging on selected tokens based on the similarity between a reference token selected from among a plurality of tokens forming a token matrix based on the input data and an adjacent token and a reference similarity in only one of the plurality of transformer blocks; and It is characterized by including an inference processing unit that performs object recognition inference in the recognition target image based on the token merging processing result output through the plurality of transformer blocks. Vision transformer device.
2. In paragraph 1, The transformer encoding processing unit is characterized in that it selects at least one similar token based on the similarity between the selected reference token and a token adjacent to at least one of the above, below, left, and right of the selected reference token and the similarity between the above, below, left, and right of the selected reference token and the similarity between the above, below, left, and right of the selected at least one similar token and the similarity between the above, below, left, and right of the selected at least one similar token and the similarity between the above, below, left, and right of the selected at least one additional similar token and the similarity between the above, below, left, and right of the selected at least one similar token and the similarity between the above, below, left, and right of the selected at least one similar token, and the above, ... Vision transformer device.
3. In paragraph 2, The transformer encoding processing unit performs an operation on an attention distribution for tokens constituting the token matrix, and compares the similarity between the selected reference token and a token adjacent to at least one of the upper, lower, left, and right of the selected reference token based on the calculated attention distribution and the spatial locality of the tokens with the reference similarity, and selects a token having a compared similarity greater than the reference similarity as the at least one similar token. Vision transformer device.
4. In paragraph 3, The above transformer encoding processing unit is characterized in that it calculates the attention distribution by applying a softmax function to the attention scores for the tokens constituting the token matrix. Vision transformer device.
5. In paragraph 1, The transformer encoding processing unit is characterized in that it sets the reference similarity to decrease so as to increase the token merging ratio for tokens constituting the token matrix, and sets the reference similarity to increase so as to decrease the token merging ratio. Vision transformer device.
6. In paragraph 5, The transformer encoding processing unit is characterized in that it sets the reference similarity to decrease by considering the background image ratio in the recognition target image, and sets the reference similarity to increase by considering the number of target objects to be identified in the recognition target image. Vision transformer device.
7. In paragraph 1, The above inference processing unit is characterized in that the number of tokens constituting the token matrix is reduced as adjacent backgrounds with uniform colors are merged into one token in the result of the token merging processing, and the object recognition inference is performed to classify the class of the object in the recognition target image through head processing of a multilayer perceptron based on the reduced tokens. Vision transformer device.
8. In the input processing unit, a step of dividing the recognition target image into multiple patches and processing one of the multiple patches into input data in the form of a token; In a transformer encoding unit, a step of performing transformer encoding on the input data using a plurality of transformer blocks, performing token merging on selected tokens based on the similarity between a selected reference token and an adjacent token and the reference similarity among a plurality of tokens forming a token matrix based on the input data in only one transformer block among the plurality of transformer blocks; and In the inference processing unit, it is characterized by including a step of performing object recognition inference on the recognition target image based on the token merging processing result output through the plurality of transformer blocks. Vision Transformer Method.
9. In paragraph 8, A step of performing token merging on selected tokens based on the similarity between a selected reference token and adjacent tokens and the reference similarity among a plurality of tokens forming a token matrix based on the input data in only one of the plurality of transformer blocks is as follows: A method for generating a token merging process, comprising: selecting at least one similar token based on the similarity between the selected reference token and a token adjacent to at least one of the above, below, left, and right of the selected reference token and the reference similarity, storing information about the selected similar token, selecting at least one additional similar token based on the similarity between the selected at least one similar token and a token adjacent to at least one of the above, below, left, and right of the selected at least one similar token and the reference similarity, storing information about the at least one additional similar token, and merging the selected reference token, the selected at least one similar token, and the selected at least one additional similar token at once and outputting the token merging processing result through the plurality of transformer blocks. Vision Transformer Method.
10. In paragraph 9, A step of performing token merging on selected tokens based on the similarity between a selected reference token and adjacent tokens and the reference similarity among a plurality of tokens forming a token matrix based on the input data in only one of the plurality of transformer blocks is as follows: It is characterized by including a step of performing an attention distribution operation on tokens constituting the token matrix, comparing the similarity between the selected reference token and a token adjacent to at least one of the upper, lower, left, and right of the selected reference token based on the calculated attention distribution and the spatial locality of the tokens, and selecting a token whose compared similarity is greater than the reference similarity as the at least one similar token. Vision Transformer Method.
Citation Information
Patent Citations
Chromatography device and method of use
KR1020230148122A
Pressure vessel
KR1020250033502A
Pre-training of computer vision foundational models
US20230162481A1
Deep learning-based anomaly detection in images
US20230281959A1
KR20230159998A