Image processing system, image processing method, and program

By employing a sparse transformer unit that selectively performs matrix multiplication operations based on calculated differences between feature vectors, the image processing time is substantially reduced for high-resolution tasks in ViT-based systems.

JP2025077817APending Publication Date: 2025-05-19NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023190300
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

Image processing using Vision Transformer (ViT) experiences increased processing time for high-resolution tasks such as object detection and pose estimation due to the high amount of calculation required.

Method used

The implementation of a sparse transformer unit that uses a matrix formed by first and second feature vectors to calculate differences and extract a feature vector to be operated on, reducing the number of matrix multiplication operations by only processing the feature vector to be operated on and using previous results for other vectors.

Benefits of technology

This approach significantly reduces the processing time of each transformer processing unit and consequently the entire image processing apparatus, enhancing efficiency in high-resolution tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025077817000001_ABST
    Figure 2025077817000001_ABST
Patent Text Reader

Abstract

To provide an image processing system that shortens the processing time of image processing using a ViT.SOLUTION: An image processing system includes a plurality of sparse transformer units. Each sparse transformer unit includes an extraction unit that calculates, using a matrix formed by arranging a plurality of first feature vectors at a first time point as rows and a matrix formed by arranging a plurality of second feature vectors at a second time point as rows, a difference between each of the first feature vectors and the second feature vectors, and extracts, based on the difference, a feature vector to be calculated; and a plurality of matrix multiplication operators that perform matrix multiplication operations using the plurality of first feature vectors. Each matrix multiplication operator includes a transformer processing unit that performs a matrix multiplication operation on the feature vector to be calculated, does not perform the matrix multiplication operation on the feature vectors among the first feature vectors that are not targets of the calculation, and uses a result of the matrix multiplication operation at the second time point.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus, an image processing method, and a program for executing image processing at high speed.

Background Art

[0002] As a model for performing image recognition processing using Transformer (a deep learning model) used in the field of natural language processing, ViT (Vision Transformer) is known.

[0003] As a related technique, Patent Document 1 discloses an image processing apparatus that improves the performance of feature extraction by multi-task learning using a feature extractor using Transformer. The image processing apparatus of Patent Document 1 divides an image including an object to generate a plurality of partial images. Next, the image processing apparatus converts the divided partial images into tokens that are fixed-dimensional vectors, and adds class tokens having a fixed dimension corresponding to the tokens to the column of the converted tokens. Next, the image processing apparatus updates the column of tokens to which the class tokens are added based on the relevance between the tokens, and extracts a feature amount of the object from the encoded representation corresponding to the updated class tokens. Further, an attribute of the object is determined from the encoded representation corresponding to the updated class tokens.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the image processing using ViT described above, when processing high-resolution tasks such as practical object detection and pose estimation, the amount of calculation increases, so the processing time becomes long.

[0006] An example of the object of the present disclosure is to shorten the processing time of image processing using ViT.

Means for Solving the Problems

[0007] To achieve the above object, an image processing apparatus according to one aspect of the present disclosure has a plurality of sparse transformer units, wherein the sparse transformer unit uses a matrix formed by using a plurality of first feature vectors at a first time point as rows and a matrix formed by using a plurality of second feature vectors at a second time point before the first time point as rows, and for each of the first feature vectors, calculates a difference between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference, an extraction unit that extracts a feature vector to be operated on from among the first feature vectors; has a plurality of matrix multiplication units that execute matrix multiplication operations using the plurality of first feature vectors, and each of the matrix multiplication units executes a matrix multiplication operation on the feature vector to be operated on, does not execute a matrix multiplication operation on the feature vectors other than the feature vector to be operated on among the first feature vectors, and uses the result of the matrix multiplication operation at the second time point, and a transformer processing unit. It is characterized by this.

[0008] Also, to achieve the above object, an image processing method according to one aspect of the present disclosure is such that an image processing apparatus executes a plurality of sparse transformer processes, wherein the sparse transformer process Using a matrix composed of a plurality of first feature vectors at a first time point as rows and a matrix composed of a plurality of second feature vectors at a second time point before the first time point as rows, for each of the first feature vectors, calculate the difference between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference, an extraction process of extracting a feature vector to be operated on from among the first feature vectors, Having a plurality of matrix multiplication units that execute matrix multiplication operations using the plurality of first feature vectors, each of the matrix multiplication units executes a matrix multiplication operation on the feature vector to be operated on, does not execute a matrix multiplication operation on the feature vectors other than the feature vector to be operated on among the first feature vectors, and a transformer process that uses the result of the matrix multiplication operation at the second time point, It is characterized by executing.

[0009] Furthermore, in order to achieve the above object, a program in one aspect of the present disclosure Causes a computer to execute a plurality of sparse transformer processes, The sparse transformer process Using a matrix composed of a plurality of first feature vectors at a first time point as rows and a matrix composed of a plurality of second feature vectors at a second time point before the first time point as rows, for each of the first feature vectors, calculate the difference between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference, an extraction process of extracting a feature vector to be operated on from among the first feature vectors, Having a plurality of matrix multiplication units that execute matrix multiplication operations using the plurality of first feature vectors, each of the matrix multiplication units executes a matrix multiplication operation on the feature vector to be operated on, does not execute a matrix multiplication operation on the feature vectors other than the feature vector to be operated on among the first feature vectors, and a transformer process that uses the result of the matrix multiplication operation at the second time point, It is characterized by executing.

Advantages of the Invention

[0010] According to the present disclosure as described above, the processing time of image processing using ViT can be shortened.

Brief Description of Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0012] First, an overview will be described to facilitate the understanding of the embodiments described hereinafter. FIG. 1 is a diagram for explaining an example of a conventional image processing apparatus. The image processing apparatus 1 shown in FIG. 1 is an apparatus that uses ViT to perform, for example, object detection, pose estimation, and the like. For details of ViT, refer to Reference [1].

[0013] [1] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, [online], Published: 13 Jan 2021, Last Modified: 17 Sept 2023, International Conference on Learning Representations ICLR 2021, [searched on September 27, 2023], Internet <URL:https: / / openreview.net / forum?id=YicbFdNTTy>

[0014] The image processing apparatus 1 includes a feature vector generation unit 2, a transformer processing unit 3 (3 1 , 3 2 ···3 n : a plurality of layers), and a determination unit 4. Note that n is a positive integer of 3 or more. In reality, the transformer processing unit 3 is arranged in, for example, several hundreds or more. Also, the transformer processing unit 3 is a general Transformer used in ViT.

[0015] When the feature vector generation unit 2 first receives an input image, it divides the input image into m preset images (patch dimensions). Next, the feature vector generation unit 2 generates a feature vector with d preset feature amounts (feature quantity dimensions) for each of the m divided images (patches). That is, the feature vector generation unit 2 generates an input matrix X 1 (an m×d matrix). For details of the feature vector generation unit 2, refer to reference [1].

[0016] Next, the first transformer processing unit 3 1 When the input matrix X 1 generated by the feature vector generation unit 2 is input, it executes a transformer process and outputs an output matrix Y 1 (= the input matrix X 2 (an m×d matrix)). Next, the transformer processing unit 3 2 When the input matrix X 1 generated by the transformer processing unit 3 2 is input, it executes a transformer process and outputs an output matrix Y 2 (= the input matrix X 3 (an m×d matrix)).

[0017] In this way, the transformer process is sequentially executed. When the input matrix X n generated by the transformer processing unit 3 n-1 is input to the last transformer processing unit 3 n , the transformer process is executed, and an output matrix Y n (= the input matrix X n+1 (an m×d matrix)) is output.

[0018] Next, the determination unit 4 makes a determination based on the input matrix X n+1 (an m×d matrix) and outputs a determination result (for example, an inference result such as object detection or pose estimation). Note that the determination unit 4 is, for example, a class filter (MLP Head) using an MLP (Multi Layer Perceptron).

[0019] Next, the transformer processing unit 3 will be described. FIG. 2 is a diagram for explaining an example of a conventional transformer processing unit. The transformer processing unit 3 (Transformer) shown in FIG. 2 includes matmul (matrix multiplication calculators) 3a, 3b, 3c, an attention (matrix multiplication calculator) 3d, a softmax 3e, a matmul (matrix multiplication calculator) 3f, and an MLP 3g. Note that the above-described matmul and attention perform matrix multiplication operations.

[0020] In the transformer processing unit 3 shown in FIG. 2, first, the input matrix X (m×d matrix) is input to each of matmul 3a, 3b, and 3c. Matmul 3a performs a matrix multiplication operation (linear transformation) using the input input matrix X (m×d matrix) and the matrix Wa (d×k matrix) of learned weight parameters stored in a storage device (not shown), and outputs a matrix Da (key (K): m×k matrix). Here, k and d represent the number of feature amounts (dimensions) of each patch.

[0021] Also, matmul 3b performs a matrix multiplication operation (linear transformation) using the input matrix X (m×d matrix) and the matrix Wb (d×k matrix) of learned weight parameters stored in the above-described storage device, and outputs a matrix Db (query (Q): m×k matrix).

[0022] Furthermore, matmul 3c performs a matrix multiplication operation (linear transformation) using the input matrix X (m×d matrix) and the matrix Wc (d×d matrix) of learned weight parameters stored in the above-described storage device, and outputs a matrix Dc (value (V): m×d matrix).

[0023] Next, attention 3d (Attention mechanism) performs a matrix multiplication operation (inner product) using the matrix Da (m×k matrix) and the matrix Db (m×k matrix), and outputs a matrix Dd (similarity: m×m matrix).

[0024] Next, softmax3e applies the softmax function to matrix Dd (an m×m matrix) to calculate matrix De (importance: an m×m matrix). Note that each row of matrix De stores the importance of each patch with respect to other patches.

[0025] Next, matmul3f performs a matrix multiplication operation using matrix Dc (value (V): an m×d matrix) and matrix Dd (importance: an m×m matrix), and outputs matrix Df (an m×d matrix). Next, MLP3g outputs output matrix Y (an m×d matrix) using matrix Df.

[0026] However, in the image processing using the above-mentioned ViT, for example, when processing high-resolution tasks such as practical object detection and pose estimation, the amount of computation of the matrix multiplication operation executed by the transformer processing unit 3 becomes extremely large, so the processing time becomes long.

[0027] Through such a process, the inventor found the problem of reducing the amount of computation of the matrix multiplication operation executed by the transformer processing unit, and at the same time derived means for solving the related problem. As a result, the processing time of the image processing using ViT can be shortened.

[0028] Hereinafter, embodiments will be described with reference to the drawings. In the drawings described below, elements having the same function or corresponding functions are denoted by the same reference numerals, and repeated descriptions thereof may be omitted.

[0029] (Embodiment) Using FIG. 3, the configuration of a plurality of sparse transformer units included in the image processing apparatus in the embodiment will be described. FIG. 3 is a diagram for explaining an example of the sparse transformer unit.

[0030] [Device Configuration] The sparse transformer unit 10 shown in FIG. 3 is used to shorten the processing time of image processing using ViT. Also, as shown in FIG. 3, the sparse transformer unit 10 has an extraction unit 11 and a transformer processing unit 12.

[0031] The extraction unit 11 uses a matrix (input matrix X t (m×d matrix)) composed of a plurality of first feature vectors input at the first time point t as rows, and a matrix (input matrix X t-1 (m×d matrix)) composed of a plurality of second feature vectors input at the second time point t-1 before the first time point t as rows. For each first feature vector, the difference Sub (=|X t -X t-1 |) between the first feature vector and the second feature vector corresponding to the first feature vector is calculated, and based on the difference Sub, the feature vector to be operated on (X’ t (m’×k matrix)) is extracted from among the first feature vectors.

[0032] The transformer processing unit 12 has a plurality of matrix multiplication units that execute matrix multiplication operations using a plurality of first feature vectors (input matrix X t (m×d matrix)). Each matrix multiplication unit executes a matrix multiplication operation on the feature vector to be operated on, does not execute a matrix multiplication operation on the feature vectors other than the feature vector to be operated on among the first feature vectors, and uses the result of the matrix multiplication operation at the second time point.

[0033] In this way, in the embodiment, the matrix multiplication operation is executed only on the feature vector to be operated on, and for the feature vectors other than the feature vector to be operated on, the result of the matrix multiplication operation at the second time point t-1 is used. Therefore, the amount of matrix multiplication operations can be reduced compared to the conventional transformer processing unit 3. As a result, the processing time of each transformer processing unit can be shortened. Furthermore, since the processing time of the transformer processing unit can be shortened, the processing time of the entire image processing apparatus can be shortened.

[0034] [System Configuration] Next, with reference to FIG. 4, the configuration of the image processing apparatus in the embodiment will be described in more detail. FIG. 4 is a diagram showing an example of a system having the image processing apparatus.

[0035] ●Regarding the image processing apparatus The image processing apparatus 100 includes a feature vector generation unit 2 and sparse transformer units 10 (10 1 (extraction unit 11 1 , transformer processing unit 12 1 ), 10 2 (extraction unit 11 2 , transformer processing unit 12 2 ), ··· 10 n (extraction unit 11 n , transformer processing unit 12 n )) and a determination unit 4. Note that the determination unit 4 has already been described, so the description thereof will be omitted. Also, n is a positive integer of 3 or more.

[0036] Further, the image processing apparatus 100 is, for example, an information processing apparatus such as a CPU (Central Processing Unit), or a programmable device such as an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or a circuit, a server computer, a personal computer, a mobile terminal, etc. equipped with any one or more of them.

[0037] Also, during operation, the image processing apparatus 100 acquires information necessary for inference, such as the structure of the neural network and learned parameters, stored in a storage device (not shown). The storage device is, for example, a database, a server computer, a circuit having a memory, etc.

[0038] ●Regarding the feature vector generation unit The feature vector generation unit 2 sequentially acquires the input images. Next, for each acquired image, the feature vector generation unit 2 divides the acquired image into m preset images. Next, for each of the m (patch dimension number) divided images (patches), the feature vector generation unit 2 generates a feature vector having d (feature quantity dimension number) preset feature quantities. That is, the feature vector generation unit 2 generates an input matrix X 1 t (m×d matrix).

[0039] Note that at the first time point t, a plurality of feature vectors generated by the feature vector generation unit 2 are represented as the input matrix X 1 t (m×d matrix), and at the second time point t-1, a plurality of feature vectors generated by the feature vector generation unit 2 are represented as the input matrix X 1 t-1 (m×d matrix).

[0040] ●Regarding the sparse transformer unit FIG. 5 is a diagram for explaining an example of the configuration of the sparse transformer unit. Each of the sparse transformer units 10 has an extraction unit 11 and a transformer processing unit 12.

[0041] ●Regarding the extraction unit The extraction unit 11 reduces the amount of matrix multiplication operations of the matmul (matrix multiplication calculator: first calculation unit) 12a, 12b, 12c, and the sparse attention (matrix multiplication calculator) 12d included in the transformer processing unit 12 by using the following method (1) or (2).

[0042] ●Regarding method (1) The extraction unit 11 extracts a feature vector to be calculated (matrix X’ t (m’×d matrix)) from the input matrix X t (m×d matrix) by method (1). Also, the extraction unit 11 extracts the input matrix X tOutput the operation target identification information including the information representing the row number of the feature vector to be operated on in the (m×d matrix) to the transformer processing unit 12. Hereinafter, when the extraction unit 11 adopts method (1), it may be represented as the first extraction unit 11a.

[0043] ● Regarding the first extraction unit 11a Using FIG. 4, the extraction unit 11a 1 from 11a n will be described. In the example of FIG. 4, first, the input matrix X at the first time point t generated by the feature vector generation unit 2 1 t (m×d matrix) and the input matrix X at the second time point t-1 1 t-1 (m×d matrix) are input into the extraction unit 11 1 (the first extraction unit 11a 1 ). Then, the extraction unit 11 1 (the first extraction unit 11a 1 ) extracts the feature vector to be operated on (matrix X 1 t (m×d matrix)) from among a plurality of first feature vectors. 1 ’ t (m 1 ’×d matrix)).

[0044] Also, the extraction unit 11 1 (the first extraction unit 11a 1 ) obtains the operation target identification information representing the row number of the feature vector to be operated on in the input matrix X 1 t (m×d matrix) used in sparse attention 12d.

[0045] Next, in the example of FIG. 4, the input matrix X at the first time point t generated by the extraction unit 11 1 and the input matrix X at the second time point t-1 2 t (m×d matrix) are input into the extraction unit 11 2 t-1 (m×d matrix). Then, the extraction unit 11 2 inputs them. Then, the extraction unit 11 2 (the first extraction unit 11a2 ) extracts the feature vector to be operated on (matrix X' (m'×d matrix)) from among a plurality of first feature vectors (input matrix X (m×d matrix)). 2 t (m×d matrix)) from among a plurality of first feature vectors (input matrix X 2 ’ t (m 2 ’×d matrix)) is extracted.

[0046] Also, the extraction unit 11 2 (the first extraction unit 11a 2 ) obtains operation target identification information representing the row number of the feature vector to be operated on in the input matrix X (m×d matrix) used in the sparse attention 12d. 2 t (m×d matrix) used in the sparse attention 12d.

[0047] Next, in the example of FIG. 4, the input matrix X (m×d matrix) at the first time point t and the input matrix X (m×d matrix) at the second time point t−1 generated by the extraction unit 11 are input to the extraction unit 11. Then, the extraction unit 11 (the first extraction unit 11a n-1 generates the input matrix X n t (m×d matrix) at the first time point t and the input matrix X n t-1 (m×d matrix) at the second time point t−1 are input to the extraction unit 11 n . Then, the extraction unit 11 n (the first extraction unit 11a n ) extracts the feature vector to be operated on (matrix X' (m'×d matrix)) from among a plurality of first feature vectors (input matrix X n t (m×d matrix)). n ’ t (m n ’×d matrix)) is extracted.

[0048] Also, the extraction unit 11 n (the first extraction unit 11a n ) obtains operation target identification information representing the row number of the feature vector to be operated on in the input matrix X (m×d matrix) used in the sparse attention 12d. n t (m×d matrix) used in the sparse attention 12d.

[0049] The extraction unit 11 (the first extraction unit 11a) will be specifically described with reference to FIG. 5. The extraction unit 11 (the first extraction unit 11a) first uses a plurality of first feature vectors (input matrix X t (m×d matrix)) input at the first time point t and a plurality of second feature vectors (input matrix X t-1 (m×d matrix)) input at the second time point t - 1 to calculate, for each first feature vector, the difference Sub (= |X t - X t-1 |) between the first feature vector and the second feature vector corresponding to the first feature vector. Note that the difference Sub is a matrix.

[0050] Next, the extraction unit 11 (the first extraction unit 11a) calculates, for each first feature vector of the difference Sub, the sum Sum of the feature amounts (elements) included in the first feature vector.

[0051] Next, when each of the calculated sums Sum is equal to or greater than a preset threshold (the first threshold) Th1, the extraction unit 11 (the first extraction unit 11a) extracts the first feature vector corresponding to the sum Sum equal to or greater than the threshold Th1 as the feature vector (matrix X’ t (m’×d matrix)) to be operated on.

[0052] Using FIG. 6, the extraction unit 11 1 (the first extraction unit 11a 1 ) will be specifically described. FIG. 6 is a diagram for explaining an example of the extraction unit (method (1)). Table 61 at the first time point t in FIG. 6 shows the relationship between the first divided images (patches: p1, p2, p3, p4, p5) divided into five by the feature vector generation unit 2 at the first time point t and the first feature vectors (rows) of each of the first divided images (patches).

[0053] Also, table 62 at the second time point t - 1 in FIG. 6 shows the relationship between the second divided images (patches: p1, p2, p3, p4, p5) divided into five by the feature vector generation unit 2 at the second time point t - 1 and the second feature vectors (rows) of each of the second divided images (patches).

[0054] Also, in the example of FIG. 6, each of the five first feature vectors has eight feature amounts. That is, Table 61 at the first time point t in FIG. 6 is an input matrix X with a patch dimension m = 5 and a feature amount dimension d = 8. 1 t (5×8 matrix). Table 62 at the second time point t - 1 in FIG. 6 is an input matrix X with a patch dimension m = 5 and a feature amount dimension d = 8. 1 t-1 (5×8 matrix).

[0055] Table 63 in FIG. 6 represents the difference Sub 1 (=|X 1 t -X 1 t-1 |). The difference Sub 1 is a matrix obtained by calculating the difference between corresponding feature amounts (elements) of the matrix in Table 61 and the matrix in Table 62 and obtaining the absolute value of the calculated difference. For example, the difference value Sub11 (bold) in Table 63 of FIG. 6 can be calculated like the number 1.

[0056] (number 1) Sub11 = |C11a - C11b|

[0057] Next, the first extraction unit 11a 1 calculates the sum (sum matrix Sum) of the elements included in each row corresponding to the first divided image (patches: p1, p2, p3, p4, p5) of Table 63 (difference Sub 1 ).

[0058] For example, the sum Sum1 of patch p1 in Table 63 is the sum of the eight elements (difference values: Sub11, Sub12, Sub13 ··· Sub18) included in the row of patch p1. For example, the sum Sum1 of patch p1 can be calculated like the number 2.

[0059] (number 2) Sum1 = Sub11 + Sub12 + Sub13 + ··· + Sub18

[0060] In this way, the first extraction unit 11a1 Calculates the total sums Sum1, Sum2, Sum3, Sum4, and Sum5 for each of the first divided images (patches: p1, p2, p3, p4, p5). Note that the total sum matrix Sum is represented like Equation 3.

[0061] (Equation 3) Sum = {Sum1, Sum2, Sum3, ···, Sum5}

[0062] However, the total sum Sum is not limited to the total sum (reduce sum) shown in Equation 2. The total sum Sum may use, for example, a weighted sum.

[0063] Next, the first extraction unit 11a 1 When the total sum Sum for each row corresponding to the first divided image (patch) is equal to or greater than a preset threshold Th1, it can be determined that there has been a change in the divided image (patch) corresponding to the total sum Sum that is equal to or greater than the threshold Th1. Therefore, the first divided image (patch) with a total sum equal to or greater than the threshold Th1 is extracted. The threshold Th1 is determined, for example, through experiments, simulations, etc.

[0064] Next, the first extraction unit 11a 1 Extracts the first feature vector (row) corresponding to the extracted first divided image (patch) as the feature vector to be operated on. That is, only the feature vector corresponding to the first divided image (patch) to be operated on is extracted from the m first feature vectors (rows), enabling dimensionality reduction (m → m’ (m > m’)).

[0065] When there is one divided image (patch) with a total sum equal to or greater than the threshold Th1, for example, when the object to be operated on is p2, the first feature vector (1×8 matrix) corresponding to p2 at the first time point t in Table 61 is extracted. When there are multiple divided images (patches) with a total sum equal to or greater than the threshold Th1, for example, when the objects to be operated on are p2 and p3, the two first feature vectors (2×8 matrix) corresponding to p2 and p3 at the first time point t in Table 61 are extracted.

[0066] Furthermore, the extraction unit 11a 1 (the first extraction unit 11a 1 ) generates operation target identification information including information representing the row number of the feature vector to be operated on in the input matrix X 1 t (m×d matrix).

[0067] After that, the first extraction unit 11a 1 (the first extraction unit 11a 1 ) outputs the extracted feature vector to be operated on (matrix X 1 ’ t (m 1 ’×d matrix)) to the transformer processing unit 12 1 .

[0068] Next, the extraction unit 11 2 ···11 n (the first extraction unit 11a 2 ···11a n ) will be specifically described. In the example of FIG. 4, for each of the extraction units 11 2 ···11 n (the first extraction unit 11a 2 ···11a n ), for each of the sparse transformer units 10 2 ···10 n-1 , the output matrix Y 1 t (matrix X 2 t ) at the first time point t, Y 2 t (matrix X 3 t )···Y n-1 t (matrix X n t ) and the output matrix Y 1 t-1 (matrix X 2 t-1 ) at the second time point t - 1, Y 2 t-1 (matrix X 3 t-1 )···Y n-1 t-1 (matrix X n t-1Using as the input, the above-described first extraction unit 11a 1 executes the process as described above.

[0069] In this way, the extraction unit 11 1 ···11 n (the first extraction unit 11a 1 ···11a n ) each extracts a feature vector to be operated on (matrix X' t (m'×d matrix)) from among a plurality of first feature vectors (input matrix X t (m×d matrix)).

[0070] Furthermore, the extraction unit 11 (the first extraction unit 11a) generates operation target identification information including information representing the row numbers of the feature vectors to be operated on in the sparse attention 12d in the input matrix X t (m×d matrix).

[0071] After that, the extraction unit 11 (the first extraction unit 11a) outputs the feature vector to be operated on (matrix X' t (m'×d matrix)) to the matmul12a, 12b, 12c of the transformer processing unit 12. Also, the extraction unit 11 (the first extraction unit 11a) outputs the operation target identification information to the sparse attention 12d.

[0072] ● Regarding method (2) FIG. 7 is a diagram for explaining an example of the configuration of the sparse transformer unit. In method (2), the extraction unit 11 extracts the feature vector to be operated on (matrix X' t (m'×d matrix)) from the input matrix X t (m×d matrix) in the same manner as in method (1). That is, the extraction unit 11 extracts the feature vector to be operated on (matrix X' t (m'×d matrix)) in order to reduce the amount of calculation of the matrix multiplication operations of the matmul12a, 12b, 12c of the transformer processing unit 12.

[0073] In addition, in the extraction unit 11 of method (2), the operation target identification information used to reduce the amount of calculation of the matrix multiplication operation of the sparse attention 12d in the transformer processing unit 12 is extracted using the operation target extraction matrix.

[0074] Note that hereinafter, when method (2) is adopted in the extraction unit 11, it may be referred to as the second extraction unit 11b.

[0075] ● Regarding the second extraction unit 11b Using FIG. 4, the extraction unit 11b 1 from 11b n will be described. In the example of FIG. 4, first, the input matrix X 1 t (m×d matrix) at the first time point t generated by the feature vector generation unit 2 and the input matrix X 1 t-1 (m×d matrix) at the second time point t−1 are input into the extraction unit 11 1 (the second extraction unit 11b 1 ). Then, the extraction unit 11 1 (the second extraction unit 11b 1 ) extracts the feature vector (matrix X 1 t (m×d matrix)) of the operation target from among a plurality of first feature vectors. 1 ’ t (m 1 ’×d matrix)).

[0076] In addition, the extraction unit 11 1 (the first extraction unit 11b 1 ) generates operation target identification information including information representing the positions of the elements of the matrix of the operation target in the input matrix X 1 t (m×d matrix) used in the sparse attention 12d.

[0077] Next, in the example of FIG. 4, the input matrix X 1 at the first time point t generated by the extraction unit 11 2 t (m×d matrix) and the input matrix X 2t-1 (an m×d matrix) and are input to the extraction unit 11 2 . Then, the extraction unit 11 2 (the second extraction unit 11b 2 ) extracts a feature vector to be operated on (matrix X 2 t (an m×d matrix)) from among a plurality of first feature vectors (input matrix X 2 ’ t (an m 2 ’×d matrix)).

[0078] Also, the extraction unit 11 2 (the first extraction unit 11b 2 ) generates operation target identification information including information representing the positions of the elements of the matrix to be operated on in the input matrix X 2 t (an m×d matrix) used in sparse attention 12d.

[0079] Next, in the example of FIG. 4, the input matrix X n-1 generated by the extraction unit 11 at the first time point t n t (an m×d matrix) and the input matrix X n t-1 (an m×d matrix) at the second time point t-1 are input to the extraction unit 11 n . Then, the extraction unit 11 n (the second extraction unit 11b n ) extracts a feature vector to be operated on (matrix X n t (an m×d matrix)) from among a plurality of first feature vectors (input matrix X n ’ t (an m n ’×d matrix)).

[0080] Also, the extraction unit 11 n (the first extraction unit 11b n ) generates operation target identification information including information representing the positions of the elements of the matrix to be operated on in the input matrix X n t (an m×d matrix) used in sparse attention 12d.

[0081] Using FIG. 7, the extraction unit 11 (first extraction unit 11b) will be specifically described. Similar to method (1), the extraction unit 11 (second extraction unit 11b) first extracts the feature vector (matrix X' t (m'×d matrix)) to be operated on. Also, the extraction unit 11 (first extraction unit 11b) generates identification information for the operation target to be used for reducing the amount of matrix multiplication operation of the sparse attention 12d in the transformer processing unit 12 using the matrix to be extracted for the operation target.

[0082] Using FIG. 8, the extraction unit 11 (first extraction unit 11b) will be specifically described. FIG. 8 is a diagram for explaining an example of the extraction unit (method (2)). In method (2), the extraction unit 11 (first extraction unit 11b) first performs a matrix multiplication operation (Sum×Sum) using the sum Sum. Table 81 in FIG. 8 shows the result (matrix to be extracted for the operation target) of the matrix multiplication operation (Sum×Sum) using the sum Sum.

[0083] Next, the extraction unit 11 (first extraction unit 11b) compares each element of the matrix to be extracted for the operation target with a preset threshold value (second threshold value) Th2, and if it is equal to or greater than the threshold value Th2, it selects it as an element to be operated on. For example, in the case of the value Sc11 of the element in Table 81, if Sc11 is equal to or greater than the threshold value Th2, the position (p1, p1) of the element corresponding to Sc11 is extracted. Conversely, if Sc11 is less than the threshold value Th2, Sc11 is not extracted.

[0084] Next, the extraction unit 11 (first extraction unit 11b) generates identification information for the operation target including information representing the positions of the selected elements and outputs it to the sparse attention (matrix multiplier) 12d.

[0085] ● Regarding the transformer processing unit The transformer processing unit 12 includes matmul (matrix multiplication calculator: first calculation unit) 12a, 12b, 12c, sparse attention (matrix multiplication calculator) 12d, softmax 12e, matmul (matrix multiplication calculator) 12f, and MLP 12g. Note that the above-mentioned matmul and sparse attention perform matrix multiplication operations.

[0086] In the transformer processing unit 12 shown in FIGS. 5 and 7, first, the input matrix X' t (m'×d matrix) is input to each of matmul 12a, 12b, and 12c. Matmul 12a uses the input input matrix X' t (m'×d matrix) and the matrix Wa (d×k matrix) composed of the learned weight parameters stored in a storage device (not shown) to perform a matrix multiplication operation (linear transformation), and outputs the matrix Da' (key (K): m'×k matrix).

[0087] Also, matmul 12b uses the input matrix X' t (m'×d matrix) and the matrix Wb (d×k matrix) composed of the learned weight parameters stored in the above-mentioned storage device to perform a matrix multiplication operation (linear transformation), and outputs the matrix Db' (query (Q): m'×k matrix).

[0088] Furthermore, matmul 12c uses the input matrix X' t (m'×d matrix) and the matrix Wc (d×d matrix) composed of the learned weight parameters stored in the above-mentioned storage device to perform a matrix multiplication operation (linear transformation), and outputs the matrix Dc' (value (V): m'×d matrix).

[0089] As described above, each of matmul (matrix multiplication calculator: first calculation unit) 12a, 12b, and 12c performs a matrix multiplication operation using the input matrix X' t (m'×d matrix) with reduced dimensions, so the calculation amount of the matrix multiplication operation can be reduced.

[0090] The sparse attention12d (Attention mechanism: attention processing unit) first obtains a matrix Da' (m'×k matrix), a matrix Db' (m'×k matrix), and operation target identification information. Next, attention12d performs a matrix multiplication operation (inner product) using the matrix Da' (m'×k matrix) and the matrix Db' (m'×k matrix), and outputs a matrix Dd' (similarity: m×m matrix). However, sparse attention12d performs the matrix multiplication operation based on the operation target identification information. Here, since the matrix Dd' is the sum-of-products operation of sparse vectors of the multiplication vector and the multiplicand vector, the dimension of the output matrix is m×m instead of m'×m'.

[0091] When the operation target identification information described in method (1) is obtained, sparse attention12d uses the rows and columns containing the elements indicated by the row numbers of the feature vectors of the operation target in the input matrix X t (m×d matrix) as the elements of the operation target in the matrix multiplication operation, and only performs the matrix multiplication operation on the elements of the operation target. Also, sparse attention12d does not perform the matrix multiplication operation on the elements outside the operation target, and uses the result of the matrix multiplication operation at the second time point t - 1 (internal feature quantity: CA t-1 )

[0092] FIG. 9 is a diagram for explaining the sparse attention of method (1). For example, when the row number of the feature vector is p2, as shown in FIG. 9, only the elements corresponding to row p2 and column p2 (shaded) corresponding to the row number p2 are subjected to the matrix multiplication operation.

[0093] When the operation target identification information described in method (2) is obtained, sparse attention12d uses the elements corresponding to the positions of the extracted elements as the elements of the operation target in the matrix multiplication operation, and only performs the matrix multiplication operation on the elements of the operation target. Also, sparse attention12d does not perform the matrix multiplication operation on the elements outside the operation target, and uses the result of the matrix multiplication operation at the second time point t - 1 (internal feature quantity: CA t-1 )

[0094] FIG. 10 is a diagram for explaining the sparse attention of method (2). For example, when the positions of the extracted elements are (p1, p2), (p2, p1), and (p2, p2), as shown in FIG. 10, matrix multiplication operations are performed only on the elements corresponding to (p1, p2), (p2, p1), and (p2, p2) (shaded).

[0095] softmax12e applies the softmax function to matrix Dd’ (an m×m matrix) to calculate matrix De’ (importance: an m×m matrix). Note that the importance of each patch with respect to other patches is stored in each row of matrix De’.

[0096] matmul12f performs a matrix multiplication operation using matrix Dc’ (value (V): an m’×d matrix) and matrix Dd’ (importance: an m×m matrix), and outputs matrix Df’ (an m×d matrix). Next, MLP12g outputs output matrix Y (an m×d matrix) using matrix Df’. Here, since matrix Dd’ is the sum-of-products operation of sparse vectors of the multiplication vector and the multiplicand vector, the dimension of the output matrix is m×m instead of m’×m’.

[0097] [Device Operation] Next, the operation of the image processing apparatus in the embodiment will be described with reference to FIG. 11. FIG. 11 is a diagram for explaining the image processing apparatus. In the following description, reference will be made to the figures as appropriate. Also, in the embodiment, an image processing method is implemented by operating the image processing apparatus. Therefore, the description of the image processing method in the embodiment is replaced by the following description of the operation of the image processing apparatus.

[0098] As shown in FIG. 11, first, the feature vector generation unit 2 sequentially acquires the input image (step A1). Note that the image has been pre-processed (such as resizing, color conversion, image cropping, rotation, etc.).

[0099] Next, the feature vector generation unit 2 sequentially acquires images, divides the acquired images into a preset number of images, and generates a feature vector for each of the divided images (step A2).

[0100] Specifically, in step A2, the feature vector generation unit 2 divides the input image into a preset number m (patch dimension number) of images. Next, in step A2, the feature vector generation unit 2 generates a feature vector having a preset number d (feature amount dimension number) of feature amounts for each of the m divided images (patches). That is, the feature vector generation unit 2 generates an input matrix X 1 (an m×d matrix).

[0101] Next, the extraction unit 11 extracts a feature vector to be operated on (matrix X’t (m’×d matrix)) from the input matrix Xt (m×d matrix) (step A3).

[0102] Specifically, in step A3, the extraction unit 11 uses a plurality of first feature vectors (input matrix X t (an m×d matrix)) input at the first time point t and a plurality of second feature vectors (input matrix X t-1 (an m×d matrix)) input at the second time point t−1 before the first time point t, and for each of the first feature vectors, calculates the difference Sub (=|X t −X t-1 |) between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference Sub, extracts the feature vector to be operated on (X’ t (m’×k matrix)) from among the first feature vectors.

[0103] Next, in step A3, the extraction unit 11 (the first extraction unit 11a) calculates the sum Sum of the feature amounts (elements) included in the first feature vector for each of the first feature vectors of the difference Sub.

[0104] Next, in step A3, when each of the calculated sums Sum is equal to or greater than a preset threshold Th1, the extraction unit 11 (the first extraction unit 11a) extracts the first feature vector corresponding to the sum Sum equal to or greater than the threshold Th1 from the feature vectors (matrix X’ t (m’×d matrix)) to be calculated.

[0105] Next, the extraction unit 11 generates calculation target identification information (step A4). Specifically, in step A4, the extraction unit 11 generates calculation target identification information by the above-described method (1) or (2).

[0106] In the case of method (1), the extraction unit 11 (the first extraction unit 11a) generates calculation target identification information including information representing the row numbers of the feature vectors to be calculated used in sparse attention 12d in the input matrix X t (m×d matrix) as described above.

[0107] In the case of method (2), the extraction unit 11 (the first extraction unit 11b) generates calculation target identification information including information representing the positions of the elements of the matrix to be calculated used in sparse attention 12d in the input matrix X t (m×d matrix) as described above.

[0108] The transformer processing unit 12 executes transformer processing (layer) (step A5). Specifically, in step A5, in the transformer processing unit 12, first, the input matrix X’ t (m’×d matrix) is input to each of matmul12a, 12b, and 12c.

[0109] Next, matmul12a performs a matrix multiplication operation (linear transformation) using the input input matrix X’ t (m’×d matrix) and the matrix Wa (d×k matrix) composed of learned weight parameters stored in a storage device (not shown), and outputs the matrix Da’ (key (K): m’×k matrix).

[0110] Also, matmul12b performs a matrix multiplication operation (linear transformation) using the input matrix X' t (m'×d matrix) and the matrix Wb (d×k matrix) composed of the learned weight parameters stored in the above-described storage device, and outputs a matrix Db' (query (Q): m'×k matrix).

[0111] Also, matmul12c performs a matrix multiplication operation (linear transformation) using the input matrix X' t (m'×d matrix) and the matrix Wc (d×d matrix) composed of the learned weight parameters stored in the above-described storage device, and outputs a matrix Dc' (value (V): m'×d matrix).

[0112] Furthermore, sparse attention12d (Attention mechanism) first obtains the matrix Da' (m'×k matrix), the matrix Db' (m'×k matrix), and the operation target identification information. Next, attention12d performs a matrix multiplication operation (inner product) using the matrix Da' (m'×k matrix) and the matrix Db' (m'×k matrix), and outputs a matrix Dd' (similarity: m×m matrix). However, sparse attention12d performs the matrix multiplication operation based on the operation target identification information. Here, since the matrix Dd' is the sum-of-products operation of sparse vectors of the multiplication vector and the multiplicand vector, the dimension of the output matrix is m×m instead of m'×m'.

[0113] When the operation target identification information described in method (1) is obtained, sparse attention12d performs a matrix multiplication operation only on the elements of the row and column corresponding to the row number based on the row number of the feature vector of the operation target in the input matrix Xt (m×d matrix). Also, sparse attention12d does not perform a matrix multiplication operation on the elements outside the operation target, and uses the result of the matrix multiplication operation (internal feature amount: CA t-1 ) at the second time point t - 1. Note that the result of the matrix multiplication operation (internal feature amount: CA t ) at the first time point t is stored in the storage device.

[0114] When the operation target identification information described in method (2) is acquired, sparse attention 12d performs a matrix multiplication operation only on the elements corresponding to the positions of the extracted elements. Also, sparse attention 12d does not perform a matrix multiplication operation on elements outside the operation target, and uses the result of the matrix multiplication operation (internal feature amount: CA t-1 ) at the second time point t - 1. Note that the result of the matrix multiplication operation (internal feature amount: CA t ) at the first time point t is stored in the storage device.

[0115] Softmax 12e applies the softmax function to the matrix Dd’ (m×m matrix) to calculate the matrix De’ (importance: m×m matrix). Note that the importance of each patch with respect to other patches is stored in each row of the matrix De’.

[0116] Matmul 12f performs a matrix multiplication operation using the matrix Dc’ (value (V): m’×d matrix) and the matrix Dd’ (importance: m×m matrix), and outputs the matrix Df’ (m×d matrix). Next, MLP 12g outputs the output matrix Y (m×d matrix) using the matrix Df’.

[0117] Next, when all the transformer processing units 10 (layers) included in the image processing apparatus 100 are executed (step A6: Yes), the determination unit 4 makes a determination based on the input matrix X n+1 (m×d matrix) and outputs a determination result (for example, an inference result such as object detection or pose estimation).

[0118] Also, when not all the transformer processing units 10 (layers) have been executed (step A6: No), the process proceeds to step A3 and the processes from steps A3 to A5 are executed.

[0119] [Effects of the Embodiment] According to the embodiment as described above, matrix multiplication operations are only performed on the feature vectors to be calculated, and for the feature vectors not to be calculated, the results of the matrix multiplication operations at the second time point t-1 are used. Therefore, the amount of matrix multiplication operations can be reduced compared to the conventional transformer processing unit. As a result, the processing time of each transformer processing unit can be shortened. Furthermore, since the processing time of the transformer processing unit can be shortened, the processing time of the entire image processing apparatus can be shortened.

[0120] Note that in the embodiment, the transformer processing unit has been described. However, the technology of the embodiment can also be used for components other than the transformer processing unit. Also, softmax may be replaced with another function, and the location where attention is taken among K, Q, and V may also be changed.

[0121] [Program] The program in the embodiment may be a program that causes a computer to execute steps A1 to A7 shown in FIG. 11. By installing and executing this program on a computer, the image processing apparatus and the image processing method in the embodiment can be realized. In this case, the processor of the computer functions as the feature vector generation unit 2, the plurality of sparse transformer units 10 (extraction unit 11, transformer processing unit 12), and the determination unit 4, and performs processing.

[0122] Also, the program in the embodiment may be executed by a computer system constructed by a plurality of computers. In this case, for example, each computer may function as any one of the feature vector generation unit 2, the plurality of sparse transformer units 10 (extraction unit 11, transformer processing unit 12), and the determination unit 4.

[0123] [Physical Configuration] Here, a computer that realizes the image processing apparatus by executing the program in the embodiment will be described with reference to FIG. 12. FIG. 12 is a diagram for explaining an example of a computer that realizes the image processing apparatus in the embodiment.

[0124] As shown in FIG. 12, the computer 110 includes a CPU (Central Processing Unit) 111, a main memory 112, a storage device 113, an input interface 114, a display controller 115, a data reader / writer 116, and a communication interface 117. These components are connected to each other via a bus 121 so as to be capable of data communication with each other. Note that the computer 110 may include a GPU or an FPGA in addition to or instead of the CPU 111.

[0125] The CPU 111 expands the program in the embodiment composed of a code group stored in the storage device 113 into the main memory 112, and executes various operations by executing each code in a predetermined order. The main memory 112 is typically a volatile storage device such as a DRAM (Dynamic Random Access Memory).

[0126] Also, the program in the embodiment is provided in a state stored in a computer-readable recording medium 120. Note that the program in the embodiment may be distributed on the Internet connected via the communication interface 117.

[0127] Specific examples of the storage device 113 include a hard disk drive and a semiconductor storage device such as a flash memory. The input interface 114 mediates data transmission between the CPU 111 and input devices 118 such as a keyboard and a mouse. The display controller 115 is connected to a display device 119 and controls the display on the display device 119.

[0128] The data reader / writer 116 mediates data transmission between the CPU 111 and the recording medium 120, and executes reading of the program from the recording medium 120 and writing of the processing result in the computer 110 to the recording medium 120. The communication interface 117 mediates data transmission between the CPU 111 and other computers.

[0129] In addition, specific examples of the recording medium 120 include general-purpose semiconductor memory devices such as CF (Compact Flash (registered trademark)) and SD (Secure Digital), magnetic recording media such as Flexible Disk, or optical recording media such as CD-ROM (Compact Disk Read Only Memory).

[0130] Note that the image processing apparatus 100 in the embodiment can also be realized by using hardware corresponding to each part, for example, an electronic circuit, instead of a computer installed with a program. Furthermore, the image processing apparatus 100 may be partially realized by a program and the remaining part may be realized by hardware. In the embodiment, the computer is not limited to the computer shown in FIG. 12.

[0131] [Appendix] Regarding the above embodiments, the following appendix is further disclosed. Some or all of the above-described embodiments can be expressed by (Appendix 1) to (Appendix 24) described below, but are not limited to the following description.

[0132] (Appendix 1) The image processing apparatus has a plurality of sparse transformer units, The sparse transformer unit uses a matrix composed of a plurality of first feature vectors at a first time point as rows and a matrix composed of a plurality of second feature vectors at a second time point before the first time point as rows, and for each of the first feature vectors, calculates a difference between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference, extracts a feature vector to be operated on from among the first feature vectors, an extraction unit; It has a plurality of matrix multiplication units that perform matrix multiplication operations using the plurality of first feature vectors. Each of the matrix multiplication units performs a matrix multiplication operation on the feature vector to be operated on, and does not perform a matrix multiplication operation on the feature vectors other than the ones to be operated on among the first feature vectors, and has a transformer processing unit that uses the result of the matrix multiplication operation at the second time point. Image processing apparatus.

[0133] (Appendix 2) The image processing apparatus further has a feature vector generation unit that sequentially acquires images, divides the acquired images into a preset number of images, and generates feature vectors for each of the divided images. The image processing apparatus according to Appendix 1.

[0134] (Appendix 3) The extraction unit uses the first feature vector corresponding to the first divided image generated by dividing the first image acquired at the first time point and the second feature vector corresponding to the second divided image generated by dividing the second image acquired at the second time point, and for each of the first feature vectors, calculates the absolute value of the difference between the first feature vector and the second feature vector corresponding to the second divided image at the same position as the first divided image as a difference. For each of the first feature vectors, calculates the sum of the feature amounts included in the first feature vector, and when the calculated sum is equal to or greater than a preset first threshold value, extracts the first feature vector corresponding to the sum equal to or greater than the first threshold value as the feature vector to be operated on. The image processing apparatus according to Appendix 2.

[0135] (Appendix 4) The transformer processing unit has a first calculation unit. Each of the plurality of matrix multiplication units included in the first arithmetic unit executes a matrix multiplication operation using a matrix composed of the feature vectors to be processed and a matrix composed of weight parameters obtained in advance by learning. The image processing apparatus according to Supplementary Note 1.

[0136] (Supplementary Note 5) The extraction unit further generates operation target identification information including information representing the row numbers of the feature vectors to be processed. The image processing apparatus according to Supplementary Note 1.

[0137] (Supplementary Note 6) The transformer processing unit has an attention processing unit, when the matrix multiplication unit included in the attention processing unit executes a matrix multiplication operation, the elements included in the rows and columns indicated by the row numbers of the feature vectors to be processed are used as the elements to be processed in the matrix multiplication operation, and the matrix multiplication operation is executed only for the elements to be processed, and the matrix multiplication operation is not executed for the elements other than the elements to be processed, and the result of the matrix multiplication operation at the second point in time is used. The image processing apparatus according to Supplementary Note 5.

[0138] (Supplementary Note 7) The extraction unit further executes an integration operation using the feature vectors to be processed to generate an operation target extraction matrix, and when the elements of the operation target extraction matrix are equal to or greater than a second threshold value set in advance, the elements equal to or greater than the first threshold value are selected, and operation target identification information including information representing the positions of the selected elements in the operation target extraction matrix is generated. The image processing apparatus according to Supplementary Note 3.

[0139] (Supplementary Note 8) The transformer processing means has an attention processing unit, When the matrix multiplication unit included in the attention processing unit executes a matrix multiplication operation, an element corresponding to the position of the element is used as an element to be operated on in the matrix multiplication operation, and the matrix multiplication operation is executed only on the element to be operated on, and the matrix multiplication operation is not executed on the elements other than the element to be operated on, and the result of the matrix multiplication operation at the second point in time is used. The image processing apparatus according to Supplementary Note 7.

[0140] (Supplementary Note 9) The image processing apparatus executes a plurality of sparse transformer processes. The sparse transformer process Using a matrix composed of a plurality of first feature vectors at a first point in time as rows and a matrix composed of a plurality of second feature vectors at a second point in time before the first point in time as rows, for each of the first feature vectors, calculating a difference between the first feature vector and the second feature vector corresponding to the first feature vector, and based on the difference, an extraction process of extracting a feature vector to be operated on from among the first feature vectors. Having a plurality of matrix multiplication units that execute a matrix multiplication operation using the plurality of first feature vectors, and each of the matrix multiplication units executes a matrix multiplication operation on the feature vector to be operated on, and does not execute a matrix multiplication operation on the feature vectors other than the feature vector to be operated on among the first feature vectors, and uses the result of the matrix multiplication operation at the second point in time. A transformer process is executed. Image processing method.

[0141] (Supplementary Note 10) The image processing apparatus further Sequentially acquires images, divides the acquired images into a preset number of images, and generates feature vectors for each of the divided images. The image processing method according to Supplementary Note 9.

[0142] (Supplementary Note 11) The extraction process Using the first feature vector corresponding to the first divided image generated by dividing the first image acquired at the first time point and the second feature vector corresponding to the second divided image generated by dividing the second image acquired at the second time point, for each of the first feature vectors, the absolute value of the difference between the first feature vector and the second feature vector corresponding to the second divided image at the same position as the first divided image is calculated as a difference, For each of the first feature vectors, the sum of the feature amounts included in the first feature vector is calculated, and when the calculated sum is equal to or greater than a preset first threshold value, the first feature vector corresponding to the sum equal to or greater than the first threshold value is extracted as a feature vector to be an operation target. The image processing method according to Appendix 10.

[0143] (Appendix 12) The transformer process has a first operation process, Each of the plurality of matrix multiplication operations executed by the first operation process executes a matrix multiplication operation using a matrix composed of the feature vectors to be an operation target and a matrix composed of weight parameters obtained in advance by learning. The image processing method according to Appendix 9.

[0144] (Appendix 13) The extraction process further generates operation target identification information including information representing the row number of the feature vector to be an operation target. The image processing method according to Appendix 9.

[0145] (Appendix 14) The transformer process has an attention process, When the attention process executes a matrix multiplication operation, the elements included in the rows and columns indicated by the row numbers of the feature vectors to be an operation target are used as elements to be an operation target in the matrix multiplication operation, and the matrix multiplication operation is executed only for the elements to be an operation target, and the matrix multiplication operation is not executed for the elements outside the operation target, and the result of the matrix multiplication operation at the second time point is used. The image processing method described in Supplementary Note 13.

[0146] (Supplementary Note 15) The extraction process further performs an accumulation operation using the feature vectors of the operation target to generate an operation target extraction matrix, and when an element of the operation target extraction matrix is equal to or greater than a preset second threshold, selects the element that is equal to or greater than the first threshold, and generates operation target identification information including information representing the position of the selected element in the operation target extraction matrix. The image processing method described in Supplementary Note 11.

[0147] (Supplementary Note 16) The transformer process has an attention process, when the attention process performs a matrix multiplication operation, sets the element corresponding to the position of the element as the element of the operation target in the matrix multiplication operation, performs the matrix multiplication operation only on the element of the operation target, does not perform the matrix multiplication operation on the element outside the operation target, and uses the result of the matrix multiplication operation at the second time point. The image processing method described in Supplementary Note 15.

[0148] (Supplementary Note 17) Causes a computer to execute a plurality of sparse transformer processes, The sparse transformer process uses a matrix composed of a plurality of first feature vectors at a first time point as rows and a matrix composed of a plurality of second feature vectors at a second time point before the first time point as rows, calculates the difference between each of the first feature vectors and the second feature vector corresponding to the first feature vector, and based on the difference, extracts the feature vectors of the operation target from among the first feature vectors (extraction process), It has a plurality of matrix multiplication units that execute matrix multiplication operations using the plurality of first feature vectors. Each of the matrix multiplication units executes a matrix multiplication operation on the feature vector to be processed, and does not execute a matrix multiplication operation on the feature vectors other than the ones to be processed among the first feature vectors. It performs a transformer process that uses the result of the matrix multiplication operation at the second point in time, A program that executes

[0149] (Appendix 18) The computer is further caused to sequentially acquire images, divide the acquired images into a preset number of images, and generate feature vectors for each of the divided images, The program according to Appendix 17.

[0150] (Appendix 19) The extraction process is Using the first feature vector corresponding to the first divided image generated by dividing the first image acquired at the first point in time and the second feature vector corresponding to the second divided image generated by dividing the second image acquired at the second point in time, for each of the first feature vectors, calculate the absolute value of the difference between the first feature vector and the second feature vector corresponding to the second divided image at the same position as the first divided image as the difference, For each of the first feature vectors, calculate the sum of the feature amounts included in the first feature vector. When the calculated sum is equal to or greater than a preset first threshold value, extract the first feature vector corresponding to the sum equal to or greater than the first threshold value as the feature vector to be processed, The program according to Appendix 18.

[0151] (Appendix 20) The transformer process has a first arithmetic process, Each of the plurality of matrix multiplication operations executed by the first arithmetic processing uses a matrix composed of the feature vectors to be processed and a matrix composed of weight parameters obtained in advance by learning, and executes a matrix multiplication operation. The program described in Supplementary Note 17.

[0152] (Supplementary Note 21) The extraction process further generates operation target identification information including information representing the row numbers of the feature vectors to be processed. The program described in Supplementary Note 17.

[0153] (Supplementary Note 22) The transformer process has an attention process. When the attention process executes a matrix multiplication operation, the elements included in the rows and columns indicated by the row numbers of the feature vectors to be processed are used as the elements to be processed in the matrix multiplication operation, and the matrix multiplication operation is executed only for the elements to be processed, and the matrix multiplication operation is not executed for the elements outside the processing target, and the result of the matrix multiplication operation at the second point in time is used. The program described in Supplementary Note 21.

[0154] (Supplementary Note 23) The extraction process further executes an integration operation using the feature vectors to be processed to generate a matrix of rows to be processed. When the elements of the matrix of rows to be processed are equal to or greater than a preset second threshold, the elements equal to or greater than the first threshold are selected, and operation target identification information including information representing the positions of the selected elements in the matrix of rows to be processed is generated. The program described in Supplementary Note 19.

[0155] (Supplementary Note 24) The transformer process has an attention process. When the attention process executes a matrix multiplication operation, an element corresponding to the position of the element is used as an element to be operated on in the matrix multiplication operation, and the matrix multiplication operation is executed only on the element to be operated on, and the matrix multiplication operation is not executed on the elements other than the element to be operated on, and the result of the matrix multiplication operation at the second point in time is used. The program according to Supplementary Note 23.

[0156] As described above, the invention has been described with reference to the embodiments, but the invention is not limited to the above-described embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the invention within the scope of the invention.

Industrial Applicability

[0157] According to the above description, the processing time of image processing using ViT can be shortened. Also, it is useful in fields where processing using Transformer is required.

Explanation of Signs

[0158] 1 Image processing apparatus 2 Feature vector generation unit 3 Transformer processing unit 3a, 3b, 3c, 3f matmul 3d attention 3e softmax 3g MLP 4 Determination unit 10 Sparse Transformer unit 10 11 Determination unit 12 Transformer processing unit 100 Image processing apparatus 110 Computer 111 CPU 112 Main memory 113 Storage device 114 Input interface 115 Display controller 116 Data reader / writer 117 Communication interface 118 Input device 119 Display device 120 Recording medium 121 Bus

Claims

1. The image processing device includes a plurality of sparse transformer means, The sparse transformer means an extraction means for calculating, for each of the first feature vectors, a difference between the first feature vector and the second feature vector corresponding to the first feature vector, using a matrix formed of rows each representing a plurality of first feature vectors at a first time point and a matrix formed of rows each representing a plurality of second feature vectors at a second time point prior to the first time point, and extracting a feature vector to be calculated from among the first feature vectors based on the difference; a transformer processing means for performing a matrix multiplication operation using the plurality of first feature vectors, each of the matrix multiplication operators performing a matrix multiplication operation on the feature vectors to be multiplied and not performing a matrix multiplication operation on feature vectors not to be multiplied among the first feature vectors, and using a result of the matrix multiplication operation at the second time point; Image processing device.

2. The image processing device further comprises: a feature vector generating means for sequentially acquiring images, dividing the acquired images into a preset number of images, and generating a feature vector for each divided image; The image processing device according to claim 1 .

3. The extraction means includes: using the first feature vector corresponding to a first divided image generated by dividing a first image acquired at the first time point and the second feature vector corresponding to a second divided image generated by dividing a second image acquired at the second time point, calculating, for each of the first feature vectors, an absolute value of a difference between the first feature vector and the second feature vector corresponding to the second divided image at the same position as the first divided image, and setting the absolute value as a difference; calculating a sum of feature amounts included in each of the first feature vectors, and if the calculated sum is equal to or greater than a first threshold value set in advance, extracting the first feature vector corresponding to the sum equal to or greater than the first threshold value as a feature vector to be calculated; The image processing device according to claim 2 .

4. The transformer processing means includes a first calculation means, each of the plurality of matrix multiplication calculators included in the first calculation means performs a matrix multiplication operation using a matrix configured by the feature vector of the calculation target and a matrix configured by using weight parameters previously obtained by learning; The image processing device according to claim 1 .

5. The extraction means further comprises: generating operation target identification information including information representing a row number of the feature vector of the operation target; The image processing device according to claim 1 .

6. The transformer processing means has an attention processing means, When the matrix multiplication unit of the attention processing means executes a matrix multiplication operation, elements included in the row and column indicated by the row number of the feature vector to be calculated are treated as elements to be calculated in the matrix multiplication operation, the matrix multiplication operation is executed only for the elements to be calculated, the matrix multiplication operation is not executed for elements other than the elements to be calculated, and the result of the matrix multiplication operation at the second point in time is used. The image processing device according to claim 5 .

7. The extraction means further comprises: performing an integration operation using the feature vectors of the operation targets to generate an operation target extraction matrix, and when an element of the operation target extraction matrix is ​​equal to or greater than a second threshold value set in advance, selecting the element equal to or greater than the first threshold value, and generating operation target identification information including information indicating a position of the selected element in the operation target extraction matrix; The image processing device according to claim 3 .

8. The transformer processing means has an attention processing means, When the matrix multiplication unit of the attention processing means executes a matrix multiplication operation, an element corresponding to the position of the element is set as an element to be multiplied in the matrix multiplication operation, the matrix multiplication operation is executed only for the element to be multiplied, the matrix multiplication operation is not executed for elements other than the element to be multiplied, and a result of the matrix multiplication operation at the second point in time is used. The image processing device according to claim 7.

9. An image processing device executes a plurality of sparse transformer processes; The sparse transformer process includes: an extraction process of calculating, for each of the first feature vectors, a difference between the first feature vector and the second feature vector corresponding to the first feature vector, using a matrix formed of rows each representing a plurality of first feature vectors at a first time point and a matrix formed of rows each representing a plurality of second feature vectors at a second time point prior to the first time point, and extracting a feature vector to be calculated from among the first feature vectors based on the difference; a transformer process including a plurality of matrix multiplication calculators that execute a matrix multiplication operation using the plurality of first feature vectors, each of the matrix multiplication calculators executes a matrix multiplication operation on the feature vectors to be multiplied and does not execute a matrix multiplication operation on feature vectors among the first feature vectors that are not to be multiplied, and uses a result of the matrix multiplication operation at the second time point; An image processing method that performs

10. Having a computer execute multiple sparse transformer operations; The sparse transformer process includes: an extraction process of calculating, for each of the first feature vectors, a difference between the first feature vector and the second feature vector corresponding to the first feature vector, using a matrix formed of rows each representing a plurality of first feature vectors at a first time point and a matrix formed of rows each representing a plurality of second feature vectors at a second time point prior to the first time point, and extracting a feature vector to be calculated from among the first feature vectors based on the difference; a transformer process including a plurality of matrix multiplication calculators that execute a matrix multiplication operation using the plurality of first feature vectors, each of the matrix multiplication calculators executes a matrix multiplication operation on the feature vectors to be multiplied and does not execute a matrix multiplication operation on feature vectors among the first feature vectors that are not to be multiplied, and uses a result of the matrix multiplication operation at the second time point; A program that executes.

Citation Information

Patent Citations

  • Image processing apparatus and control method and program of image processing apparatus

    JP2023092206A