Method and device for extracting image features, electronic equipment and storage medium

By obtaining the target image window and determining its mask mode in Swin Transformer, calculating the target proportion and adjusting the general matrix multiplication operation, the problem of GEMM redundant calculation after sliding window operation is solved, and resource consumption savings and calculation efficiency improvements are achieved.

CN120047691APending Publication Date: 2025-05-27SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510117740.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

After the sliding window operation of Swin Transformer, there is redundant calculation in the prior art, which leads to a large resource consumption, especially when hardware implementation.

Method used

By obtaining the target image window and determining its corresponding mask mode, calculating the target proportion, and adjusting the calculation method of general matrix multiplication, avoiding unnecessary redundant calculations. The specific method includes determining the target proportion according to the mask pattern, and performing a general matrix multiplication operation on the first calculation matrix and the second calculation matrix based on the ratio to obtain the target matrix.

Benefits of technology

It effectively saves resource consumption, especially when hardware is implemented, it can significantly reduce the consumption of hardware resources and improve computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047691A_ABST
    Figure CN120047691A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for extracting image features, electronic equipment and a storage medium, and the method comprises the steps: obtaining a target image window, and determining a mask mode corresponding to the target image window; determining a first calculation matrix and a second calculation matrix corresponding to the target image window; determining a target proportion according to the mask mode; wherein the target proportion represents the proportion occupied by the operation result which does not need to be subjected to mask processing when multiplication operation is carried out on the first calculation matrix and the second calculation matrix; performing general matrix multiplication on the first calculation matrix and the second calculation matrix according to a target proportion to obtain a target matrix; the target matrix is an image feature required in the image processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method, device, electronic device and storage medium for extracting image features. Background Art

[0002] Transformer is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks. It has become the basis of NLP models and is also widely used in image and multimodal models. It performs well in different modal data sets and is the focus of research in the field of artificial intelligence (AI). Swin transformer applies it to the image field and has achieved excellent performance in downstream tasks such as image classification, object detection, and semantic segmentation. However, due to the emergence of various transformer-based models and the lack of a unified model, hardware acceleration design for specific models such as swin transformer will bring more resource overhead.

[0003] The current mainstream transformer acceleration is to accelerate the general parts of self-attention (SA), layer normalization (Layer Normalization, Layernorm), and general matrix multiplication (GEMM). Among them, the idea of ​​GEMM reconfiguration is: converting a large matrix into multiple small matrices for parallel calculation; GEMM optimization pays more attention to sparse compression operations, such as simplifying data close to 0 by focusing on the most significant bit (MSB), or arranging matrix elements and selecting the top k elements (top-k elements) to achieve the purpose of streamlining calculations, thereby reducing hardware resource consumption.

[0004] In the calculation of the swin transformer, for the sliding window shift, the reconfigurable GEMM still has redundant calculations after data scheduling, and the calculations increase as the feature map scale increases. Since the matrix data is not approximately 0 before entering the normalization function softmax, the sparse method cannot avoid redundant calculations in the softmax layer. Summary of the invention

[0005] The embodiments of the present disclosure provide a method, device, electronic device and storage medium for extracting image features.

[0006] First aspect, embodiments of the present disclosure provide a method for extracting image features, the method comprising:

[0007] Obtain a target image window and determine a mask pattern corresponding to the target image window;

[0008] Determine a first calculation matrix and a second calculation matrix corresponding to the target image window;

[0009] Determine a target ratio according to the mask pattern; wherein, the target ratio represents: the ratio of the operation results that do not need to be masked when performing a multiplication operation on the first calculation matrix and the second calculation matrix;

[0010] Perform a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix; the target matrix is the image feature required in the image processing process.

[0011] In some embodiments, the obtaining the target image window includes:

[0012] Obtain an initial image, the initial image includes a first number of initial image windows;

[0013] Perform a sliding window operation on the initial image to obtain a first image, the first image includes a second number of intermediate image windows; the second number is greater than the first number;

[0014] Perform an image block movement on the first image to obtain a target image, the target image includes the first number of the target image windows;

[0015] Wherein, the operation results that do not need to be masked are: the operation results between the image blocks from the same intermediate image window.

[0016] In some embodiments, determining the mask pattern corresponding to the target image window includes:

[0017] If the target image window does not contain the image blocks obtained by movement, determine that the mask pattern corresponding to the target image window is a first mask pattern;

[0018] If the target image window contains the image blocks obtained by moving from a first position, determine that the mask pattern corresponding to the target image window is a second mask pattern, the first position is in a first direction of the first image;

[0019] If the target image window contains the image blocks obtained by moving from a second position, determine that the mask pattern corresponding to the target image window is a third mask pattern, the second position is in a second direction of the first image;

[0020] If the target image window contains an image block obtained by moving from a third position, determine that the mask pattern corresponding to the target image window is a fourth mask pattern, where the third position is at the intersection of the first direction and the second direction of the first image.

[0021] In some embodiments, the sliding distance of the sliding window operation is half of the side length of the target image window; the determining the target ratio according to the mask pattern includes:

[0022] If the mask pattern is the first mask pattern, determine that the target ratio is 1;

[0023] If the mask pattern is the second mask pattern or the third mask pattern, determine that the target ratio is 1 / 2;

[0024] If the mask pattern is the fourth mask pattern, determine that the target ratio is 1 / 4.

[0025] In some embodiments, both the number of rows of the first calculation matrix and the number of columns of the second calculation matrix are a third quantity, where the third quantity is the number of image blocks included in the target image window, and the number of loops of the general matrix multiplication is the product of the target ratio and the third quantity.

[0026] In some embodiments, performing a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix includes:

[0027] Determine the reciprocal of the target ratio as the loop distribution constant;

[0028] Sequentially divide the third quantity of row vectors in the first calculation matrix into the loop distribution constant number of first intervals, and sequentially divide the third quantity of column vectors in the second calculation matrix into the loop distribution constant number of second intervals; the number of row vectors included in each first interval is equal to the number of column vectors included in each second interval and is equal to the number of loops;

[0029] In the i-th loop of the general matrix multiplication, multiply the i-th row vector in the a-th first interval of the first calculation matrix by each column vector in the a-th second interval of the second calculation matrix; where i is an integer greater than 0 and less than or equal to the number of loops, and a is an integer greater than 0 and less than or equal to the loop distribution constant;

[0030] Wherein, the result of multiplying each row vector and each column vector is an element in the target matrix.

[0031] In some embodiments, the first calculation matrix is a query matrix, and the second calculation matrix is a key matrix; the image processing is based on the swin transformer model.

[0032] Second, embodiments of the present disclosure provide an apparatus for extracting image features, including:

[0033] An acquisition unit, configured to acquire a target image window and determine a mask pattern corresponding to the target image window;

[0034] A first determination unit, configured to determine a first calculation matrix and a second calculation matrix corresponding to the target image window;

[0035] A second determination unit, configured to determine a target ratio according to the mask pattern; wherein, the target ratio represents: the proportion of the operation results that do not need to be masked during the multiplication operation of the first calculation matrix and the second calculation matrix;

[0036] A calculation unit, configured to perform a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix; the target matrix is the image feature required during the image processing.

[0037] Third, embodiments of the present disclosure provide an electronic device, including a memory and a processor, wherein:

[0038] The memory is configured to store a computer program that can run on the processor;

[0039] The processor is configured to execute the method according to any one of the first aspects when running the computer program.

[0040] Fourth, embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and when the computer program is executed by at least one processor, the method according to any one of the first aspects is implemented.

[0041] Embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a storage medium for extracting image features. When performing a general matrix multiplication operation, in combination with the proportion of the operation results that do not need to be masked, in the case where there are operation results that need to be masked, the operation method of the general matrix multiplication is adjusted, avoiding unnecessary redundant calculations, effectively saving resource consumption, especially when implemented in hardware, it can effectively reduce the consumption of hardware resources. Description of the Drawings

[0042] Figure 1 It is a schematic flowchart of the method for extracting image features provided by the embodiments of the present disclosure;

[0043] Figure 2 Schematic of window change for the sliding window operation provided by the embodiments of the present disclosure Figure 1 ;

[0044] Figure 3 Schematic of shifting and integrating an image window provided by the embodiments of the present disclosure Figure 1 ;

[0045] Figure 4 Schematic of window change for the sliding window operation provided by the embodiments of the present disclosure Figure 2 ;

[0046] Figure 5 Schematic of shifting and integrating an image window provided by the embodiments of the present disclosure Figure 2 ;

[0047] Figure 6 Schematic of the target image window corresponding to the mask mode provided by the embodiments of the present disclosure Figure 1 ;

[0048] Figure 7 Schematic of the target image window corresponding to the mask mode provided by the embodiments of the present disclosure Figure 2 ;

[0049] Figure 8 Schematic of matrix multiplication provided by the embodiments of the present disclosure Figure 1 ;

[0050] Figure 9 Schematic of matrix multiplication provided by the embodiments of the present disclosure Figure 2 ;

[0051] Figure 10 Schematic of the loop process of general matrix multiplication provided by the embodiments of the present disclosure Figure 1 ;

[0052] Figure 11 Schematic of the loop process of general matrix multiplication provided by the embodiments of the present disclosure Figure 2 ;

[0053] Figure 12 Schematic of the loop process of general matrix multiplication provided by the embodiments of the present disclosure Figure 3 ;

[0054] Figure 13 Schematic of logical clustering provided by the embodiments of the present disclosure;

[0055] Figure 14 Schematic of the device structure for extracting image features provided by the embodiments of the present disclosure;

[0056] Figure 15 Schematic of implementing general matrix multiplication using hardware provided by the embodiments of the present disclosureFigure 1 ;

[0057] Figure 16 Schematic diagram of implementing general matrix multiplication using hardware provided by embodiments of the present disclosure Figure 2 ;

[0058] Figure 17 Schematic structural diagram of an electronic device provided by embodiments of the present disclosure. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. It can be understood that the specific embodiments described herein are only used to explain the relevant disclosure, rather than limiting the disclosure. Additionally, it should be noted that, for the sake of description, only parts related to the relevant disclosure are shown in the drawings.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are only for the purpose of describing the embodiments of the present disclosure and are not intended to limit the present disclosure.

[0061] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0062] It should be noted that the terms "first / second / third" involved in the embodiments of the present disclosure are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0063] To solve the problems such as large resource consumption caused by computational redundancy after the sliding window operation of the Swin Transformer, an embodiment of the present disclosure provides a method for extracting image features. The method includes: obtaining a target image window and determining a mask pattern corresponding to the target image window; determining a first calculation matrix and a second calculation matrix corresponding to the target image window; determining a target ratio according to the mask pattern, where the target ratio represents the proportion of the calculation results that do not need to be masked during the multiplication operation of the first calculation matrix and the second calculation matrix; performing a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix, and the target matrix is the image feature required in the image processing process. In this way, in the embodiment of the present disclosure, when performing the general matrix multiplication operation, in combination with the proportion of the calculation results that do not need to be masked, in the case where there are calculation results that need to be masked, the operation mode of the general matrix multiplication is adjusted, avoiding unnecessary redundant calculations, effectively saving resource consumption, especially when implemented in hardware, it can effectively reduce the hardware resource consumption.

[0064] The following will describe each embodiment of the present disclosure in detail with reference to the accompanying drawings.

[0065] In an embodiment of the present disclosure, refer to Figure 1 , which shows a schematic flowchart of a method for extracting image features provided by an embodiment of the present disclosure. As Figure 1 shown, the method may include:

[0066] S101: Obtain a target image window and determine a mask pattern corresponding to the target image window.

[0067] It should be noted that the method for extracting image features provided by the embodiment of the present disclosure can be specifically applied to the image processing process based on the Swin Transformer model. The method can be implemented by a device for extracting image features, and the device for extracting image features can be integrated in an electronic device. Here, the electronic device can be a smart phone, a tablet computer, a notebook computer, a server, etc., and no specific limitation is made thereto.

[0068] For ease of understanding, some basic concepts that may be involved in the embodiment of the present disclosure are defined as follows:

[0069] Define the length and width of the picture as H and W respectively. For example, if both H and W are 224, the size of the picture is: 224×224 pixels;

[0070] Define the side length of the image block as pat. Taking pat = 4 as an example, that is, every 4×4 pixels are regarded as an image block (patch), then there are 56×56 image blocks in a 224×224 picture;

[0071] Define the side length of the window as M. Taking M = 4 as an example, that is, every 4×4 image patches are regarded as an image window, then there are 14×14 image windows in a 224×224 picture.

[0072] In the embodiments of the present disclosure, obtaining a target image window may include:

[0073] Obtain an initial image, which includes a first number of initial image windows;

[0074] Perform a sliding window operation on the initial image to obtain a first image, which includes a second number of intermediate image windows; the second number is greater than the first number;

[0075] Perform an image patch movement on the first image to obtain a target image, which includes a first number of target image windows.

[0076] It should be noted that, taking Figure 2 as an example, it shows a schematic diagram of the window change of a sliding window operation provided by the embodiments of the present disclosure. Figure 1 Among them, (a), (b), and (c) respectively represent: before sliding the window, after sliding the window, and after restoration. Figure 2 A small square in represents 1 image patch. Still assuming M = 4, Figure 2 The bold square frame in represents an image window.

[0077] In Figure 2 , the initial image is as shown in (a), and each image window therein is denoted as an initial image window, and the first number is 9. In Figure 2 , M = 4. Assuming that the row and column sliding sizes of the sliding window operation are both M / 2, then the first image obtained by sliding the window is as shown in (b), and each image window therein is denoted as an intermediate image window, and the second number is 16. Then, move the image patches in the first image in the manner shown in Figure 3 to obtain a target image, and each image window in the target image is denoted as a target image window.

[0078] That is to say, the basic sliding window operation of the swin transformer is: taking the output of the local self-attention layer as the input of the self-attention layer after sliding the window size in the column (col) and row (row) directions for secondary calculation, so as to obtain features that can perceive the graphics in adjacent image windows. Among them,

[0079] wherein, represents the floor of M / 2. In the embodiments of the present disclosure, for the convenience of description, M is taken as an even number as an example. In the example shown in Figure 2 , each window sliding operation slides 2 image patches in both the row and column directions. As shown in Figure 2As shown, with the image window 1 as a reference, after windowing, the position of the image window 1 changes, moving down and to the right by 2 image blocks respectively. The number of image windows after sliding is increased compared to the original image window (originally 9 image windows, and it becomes 16 image windows after window sliding), and their sizes are different (the image windows at the edges are incomplete). Therefore, for unified calculation, the image windows at the edges are shifted and integrated so that the image windows seen during calculation still maintain the same size and quantity. Specifically, as Figure 3 shown, the incomplete image window in the upper left corner is moved to the lower right corner, the incomplete image window above is moved to the lower part, and the incomplete image window on the left is moved to the right. The moved incomplete image window and the adjacent incomplete image window form a complete image window. In the embodiment of the present disclosure, the target image window is any one of the image windows obtained after window sliding operation and shifting and integration.

[0080] Figure 4 and Figure 5 also respectively show schematic diagrams of window sliding operation and shifting and integration of image windows in the case of M = 2, which is similar to the case of M = 4 and will not be elaborated here.

[0081] Performing self-attention calculation on the shifted edge image window without any operation will introduce confusion in position information. Therefore, it is necessary to use a mask to mask this operation. Different types of target image windows correspond to different mask patterns. As Figure 3 shown, there are four different types of target image windows, corresponding to four mask patterns. Specifically, determining the mask pattern corresponding to the target image window may include:

[0082] If the target image window does not contain the image blocks obtained by moving, then determine that the mask pattern corresponding to the target image window is the first mask pattern;

[0083] If the target image window contains the image blocks obtained by moving from the first position, then determine that the mask pattern corresponding to the target image window is the second mask pattern, and the first position is in the first direction of the first image;

[0084] If the target image window contains the image blocks obtained by moving from the second position, then determine that the mask pattern corresponding to the target image window is the third mask pattern, and the second position is in the second direction of the first image;

[0085] If the target image window contains the image blocks obtained by moving from the third position, then determine that the mask pattern corresponding to the target image window is the fourth mask pattern, and the third position is at the intersection of the first direction and the second direction of the first image.

[0086] It should be noted that, taking Figure 3For example, the first direction is the left side, and the second direction is the upper side. That is, the first position can be located on the left side of the first image, the second position is located above the first image, and the third position can be located at the upper left corner of the first image. The target image windows corresponding to the four mask patterns are respectively as Figure 6 shown:

[0087] The first mask pattern (MASK0) corresponds to: the complete image window obtained after the sliding window, without being moved and integrated;

[0088] The second mask pattern (MASK1) corresponds to: the target image window obtained by moving from the left side to the right side of the first image and then splicing and integrating.

[0089] The third mask pattern (MASK2) corresponds to: the target image window obtained by moving from the upper side to the lower side of the first image and then splicing and integrating.

[0090] The fourth mask pattern (MASK3) corresponds to: the target image window obtained by moving from the upper left corner to the lower right corner of the first image and then splicing and integrating.

[0091] Corresponding to Figure 4 , the target image windows corresponding to the four mask patterns are as Figure 7 shown, which will not be elaborated here. It can be understood that Figure 6 and Figure 7 shown are only examples. If the sliding window direction is different, the positions will change accordingly.

[0092] S102: Determine the first calculation matrix and the second calculation matrix corresponding to the target image window.

[0093] It should be noted that in the swin transformer, the first calculation matrix is the query matrix (Q matrix), and the second calculation matrix is the key matrix (K matrix). The Q matrix is obtained by multiplying the embedding matrix E of the input data by the Q weight matrix W_Q, and the K matrix is obtained by multiplying the embedding matrix E of the input data by the K weight matrix W_K. Among them, the embedding matrix E refers to the matrix composed of representing each image patch as a vector. The Q weight matrix W_Q and the K weight matrix W_K are trainable parameters. Initially, the values of these weight matrices are usually randomly initialized. During the training process, the model can update the values of the elements in these weight matrices through the backpropagation algorithm and gradient descent.

[0094] Exemplarily, as Figure 6 shown, for the case of M = 4, each image patch in the target image window is sequentially labeled as 0, 1, 2, 3, ……, 15 in the order from left to right and from top to bottom. On this basis, as Figure 8As shown, they are arranged from top to bottom in the order of the labels. Each row corresponds to a row vector of the Q matrix of an image block, and 16 row vectors form the Q matrix corresponding to the target image window. They are arranged from left to right in the order of the labels. Each column corresponds to a column vector of the K matrix of an image block, and 16 column vectors form the K matrix corresponding to the target image window.

[0095] For the case of M = 2, the relevant schematic diagram is as Figure 9 shown and will not be elaborated here.

[0096] S103: Determine the target ratio according to the mask pattern; where the target ratio represents the proportion of the operation results that do not need to be masked during the multiplication operation of the first calculation matrix and the second calculation matrix.

[0097] S104: Perform a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain the target matrix; the target matrix is the image feature required in the image processing process.

[0098] It should be noted that in the swin transformer, it is necessary to perform a multiplication operation on the Q matrix and the K matrix to obtain the final required target matrix. As Figure 8 shown, for 4 different types of target image windows, it can be understood that in Figure 8 the first three columns, different filling patterns represent the position information of the image blocks. White indicates that the position of the image block has not changed after the sliding window. Fine diagonal lines (dense diagonal lines) indicate that the image block has moved from the first position after the sliding window. Coarse diagonal lines (sparse diagonal lines) indicate that the image block has moved from the second position after the sliding window. The grid indicates that the image block has moved from the third position after the sliding window. The result after multiplying the Q matrix and the K matrix is as Figure 8 shown in the last column of. Among them, the white area represents the result of multiplying the image blocks at the same position, and the area filled with patterns represents the result of multiplying the image blocks at different positions. Here, the meaning of the same position is: before the sliding window operation and merging the windows, it is within the same image window, that is, within an intermediate image window in the first image.

[0099] For the first mask pattern MASK0, the results obtained by multiplying the Q matrix and the K matrix are all the results of multiplying the image blocks at the same position. For the second mask pattern MASK1 and the third mask pattern MASK2, half of the results obtained by multiplying the Q matrix and the K matrix are the results of multiplying the image blocks at the same position, and the other half are the results of multiplying the image blocks at different positions. For the fourth mask pattern MASK3, 1 / 4 of the results obtained by multiplying the Q matrix and the K matrix are the results of multiplying the image blocks at the same position, and the other 3 / 4 are the results of multiplying the image blocks at different positions.

[0100] Since the self-attention mechanism of the Swin Transformer calculates the image patches within the same image window after window sliding, in fact, only the results of multiplying the image patches at the same position are needed, and the calculation results for image patches at different positions are redundant. That is to say, if no operation is performed on the edge image window after shifting for self-attention calculation, it will introduce chaos in position information. Therefore, a mask is needed to mask this operation. For Figure 8 Regarding the calculation results shown in the last column, only the operation results corresponding to white are needed and do not need to be masked, while the operation results corresponding to the pattern filling are redundant and need to be masked. That is to say, among them, the operation results that do not need to be masked are: the operation results between the image patches from the same intermediate image window. The case of M = 2 is as Figure 9 shown and will not be elaborated here.

[0101] It can be seen that when the sliding distance of the window sliding operation is half of the side length of the target image window; determining the target ratio according to the mask mode includes:

[0102] If the mask mode is the first mask mode MASK0, then the target ratio is determined to be 1; that is, all operation results are needed and there is no part that needs to be masked;

[0103] If the mask mode is the second mask mode MASK1 or the third mask mode MASK2, then the target ratio is determined to be 1 / 2; that is, half of the operation results are redundant and need to be masked;

[0104] If the mask mode is the fourth mask mode MASK3, then the target ratio is determined to be 1 / 4; that is, 3 / 4 of the operation results are redundant and need to be masked.

[0105] As Figure 8 shown, the mask value MASK of the white area is 0, and the mask value MASK of the area filled with patterns is -100. This is because the mask operation is: x = x + mask value, where x represents the exponent value for subsequent exponential operation on the natural exponent. For the white part, the result of the exponential operation is: e x+0 = e x ; for the part filled with patterns, since the x value is usually small (much less than 100), therefore, e x-100 ≈ 0, which is equivalent to filtering the corresponding position value to 0 with -100. The calculation for the masked part is redundant, and the value of this part will only be set to 0 after the softmax layer. Therefore, the sparse strategy cannot accelerate this part.

[0106] Since there are 4 fixed mask patterns during masking, the embodiments of the present disclosure design GEMM (General Matrix Multiplication) based on the target ratio, so as to skip the part that needs to be masked during the operation, thereby simplifying the operation and saving resources. Here, the higher the target ratio, the more operation results are required, and the fewer redundant results need to be masked. Then, the more loop times GEMM requires, that is, there is a positive correlation between the loop times of GEMM and the target ratio.

[0107] It should also be noted that, as Figure 6 or Figure 7 shown, the target ratio can also be understood as: in the intermediate image window obtained after sliding window processing, the ratio of the number of image blocks to the total number of image blocks in the complete image window.

[0108] The improved GEMM based on the target ratio will be described in detail below. The solution is as follows:

[0109] Perform general matrix multiplication on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix, including:

[0110] Determine the reciprocal of the target ratio as the loop distribution constant;

[0111] Successively divide the third number of row vectors in the first calculation matrix into a loop distribution constant number of first intervals, and successively divide the third number of column vectors in the second calculation matrix into a loop distribution constant number of second intervals; the number of row vectors included in the first interval is equal to the number of column vectors included in the second interval and is equal to the loop times;

[0112] In the i-th loop of the general matrix multiplication, multiply the i-th row vector in the a-th first interval of the first calculation matrix by each column vector in the a-th second interval of the second calculation matrix; where i is an integer greater than 0 and less than or equal to the loop times, and a is an integer greater than 0 and less than or equal to the loop distribution constant;

[0113] Each loop of the general matrix multiplication obtains at least one row vector in the target matrix.

[0114] It should be noted that for each vector corresponding to an image block, it is necessary to perform an inner product (i.e., vector × matrix) with the vectors corresponding to 16 image blocks in the target image window, and each vector has to repeat this process, which ultimately manifests as matrix × matrix. The preliminary preparation and multiplication process of GEMM can include the following steps:

[0115] Step 1: For the four mask patterns, Figure 8 or Figure 9The matrix operation of the same filling pattern is required to obtain the result, and the multiplication of different filling patterns of the Q matrix and the K matrix is redundant. Determine the following parameters:

[0116] M: Window size. The size of a target image window is M×M image patches; the sliding length of each sliding window operation is

[0117] W×H: Image size. A picture has (W / M)×(H / M) image windows.

[0118] Take M = 2 as an example. As Figure 9 shown, each target image window includes 2×2 = 4 image patches.

[0119] For the first mask pattern MASK0, the target ratio = 1, then the cyclic distribution constant = 1 / 1 = 1;

[0120] For the second mask pattern MASK1 and the third mask pattern MASK2, the target ratio = 1 / 2, then the cyclic distribution constant = 1 / (1 / 2) = 2;

[0121] For the fourth mask pattern MASK3, the target ratio = 1 / 4, then the cyclic distribution constant = 1 / (1 / 4) = 4.

[0122] Step 2: Regard the matrix of the second mask pattern MASK1 as the data arrangement of the third mask pattern MASK2, and classify it as the same situation when entering GEMM, and perform logical rearrangement on the data under the fourth mask pattern MASK3.

[0123] It should be noted that for the first mask pattern MASK0, the overlapping width between the target image window after sliding the window and the original initial image window is 0. One image patch in the target image window needs to perform operations with all M×M image patches, and the actual number of image patches participating in the operation is M×M.

[0124] For the second mask pattern MASK1 and the third mask pattern MASK2, the overlapping width between the target image window after sliding the window and the initial image window before sliding is One image patch in the target image window needs to perform operations with (M×M) / 2 image patches in the same intermediate image window, that is, the actual number of image patches participating in the operation in the target image window is Take Figure 8 the image patches corresponding to MASK2 in as an example. The widths corresponding to both types of image patches are ×M = 8.

[0125] For the third mask pattern MASK3, the overlapping width between the target image window after the sliding window and the initial image window before the sliding window is One image block in the target image window needs to be operated on with (M / 2)×(M / 2) image blocks in the same intermediate image window, that is, the number of image blocks actually participating in the operation in the target image window is Taking Figure 8 the image blocks corresponding to MASK3 in as an example, the widths corresponding to the four types of image blocks are all

[0126] The second mask module MASK1 and the third mask pattern can be regarded as the same type in the operation. When entering GEMM, they are classified as the same situation. Then, the four mask patterns correspond to three working modes when entering GEMM:

[0127] The first mask pattern MASK0 corresponds to the first working mode: Since in the target image window corresponding to the first mask pattern MASK0, the target ratio is 100%, that is, all operation results are required and there is no redundancy, so the operation is performed according to the standard GEMM.

[0128] The second mask pattern MASK1 and the third mask pattern MASK2 correspond to the second working mode: Since in the target image windows corresponding to the second mask pattern MASK1 and the third mask pattern MASK2, the target ratio is 50%, that is, only half of the operation results are required and there is 50% redundancy, so 50% of the redundant workload can be reduced.

[0129] The fourth mask pattern MASK3 corresponds to the third working mode: Since in the target image window corresponding to the fourth mask pattern MASK3, the target ratio is 25%, that is, only 1 / 4 of the operation results are required and there is 75% redundancy, so 70% of the redundant workload can be reduced.

[0130] Step 3: Multiply-accumulate calculation.

[0131] It should be noted that the number of rows of the first calculation matrix and the number of columns of the second calculation matrix are both the third quantity, that is, the number of image blocks included in the target image window. The number of loops of the general matrix multiplication is the product of the target ratio and the third quantity. For example, in Figure 8 , M = 4, the Q matrix includes 16 row vectors, and the K matrix includes 16 column vectors; another example is in Figure 9 , M = 2, the Q matrix includes 4 row vectors, and the K matrix includes 4 column vectors; if M = 8, then the corresponding Q matrix includes 64 row vectors, and the corresponding K matrix includes 64 column vectors.

[0132] Taking Figure 9Taking the case of M = 2 as an example, for the first working mode, the standard GEMM is as follows Figure 10 As shown, here, taking the Q matrix and the K matrix both being 4×4 matrices as an example. Among them, the Q matrix includes: row vector 0 [a, b, c, d], row vector 1 [e, f, g, h], row vector 2 [i, j, k, l], row vector 3 [m, n, o, p]; the K matrix includes: column vector 0 [A, B, C, D], column vector 1 [E, F, G, H], column vector 2 [I, J, K, L], column vector 3 [M, N, O, P].

[0133] It can be understood that the multiplication of matrices is actually the multiplication of the elements of vectors and then addition. Here, multipliers can be used to implement element multiplication. When using multiple multipliers for GEMM, each multiplier includes two input terminals: the X terminal and the Y terminal. Assume that the X terminal is used to input the data of the Q matrix and the Y terminal is used to input the data of the K matrix; that is, calculate Q[M, M]×K[M, M], where Q[M, M] represents the Q matrix with both the number of rows and columns being M, and K[M, M] represents the K matrix with both the number of rows and columns being M, and fix K[M, M] at the Y terminal of the multiplier. For example Figure 10 As shown, for the 4×4 Q matrix and K matrix, 16 multipliers are required. For ease of description, assume that the 16 multipliers are arranged in a 4×4 array, and the Y terminals of the 16 multipliers are respectively fixed to the 4×4 elements in the K matrix: A, B, C,..., N, O, P. Among them, the X terminals of each multiplier can be connected to the first Distribution Network, and the output terminals of each multiplier can be connected to the second Reduction Network.

[0134] For MASK0, the loop distribution constant = 1, then the 4 row vectors in the first matrix form a first interval, and the 4 column vectors in the second matrix form a second interval; as Figure 10 shown, when performing the GEMM operation, a = 1, after fixing the K matrix

[0135] In the 1st cycle (or period), the 1st vector (row vector 0 [a, b, c, d]) in the 1st and only first interval in the Q matrix is multiplied by each column vector in the 1st and only second interval in the K matrix respectively; that is, the 1st row vector of the Q matrix is transmitted to the X terminal of each column multiplier in the GEMM through the first Distribution Network, and the calculation results of each column can be input into the Adder Tree or buffered through the second Distribution Network; the products of each column are added to obtain four sum values, and the four sum values form the 1st row vector of the target matrix.

[0136] In the second loop, the second vector (row vector 1 [e, f, g, h]) in the first and only first interval in the Q matrix is multiplied by each column vector in the first and only second interval in the K matrix respectively; that is, the second row vector of the Q matrix is transmitted to the X end of each column multiplier of the GEMM through the first distribution network, and the calculation results are cached in the same way as in the first loop to obtain the second row vector of the target matrix;

[0137] In the third loop, the third vector (row vector 2 [i, j, k, l]) in the first and only first interval in the Q matrix is multiplied by each column vector in the first and only second interval in the K matrix respectively; that is, the third row vector of the Q matrix is transmitted to the X end of each column multiplier of the GEMM through the first distribution network, and the calculation results are cached to obtain the third row vector of the target matrix;

[0138] In the fourth loop, the fourth vector (row vector 3 [m, n, o, p]) in the first and only first interval in the Q matrix is multiplied by each column vector in the first and only second interval in the K matrix respectively; that is, the fourth row vector of the Q matrix is transmitted to the X end of each column multiplier of the GEMM through the first distribution network, and the calculation results are cached to obtain the fourth row vector of the target matrix.

[0139] In this way, the conventional GEMM operation between 4×4 matrices is completed, that is, MASK0 corresponds to the conventional form of GEMM, no special operation is performed, and the number of loops is 4. If it is a larger matrix, repeat this loop step, continue to distribute each row vector to each column of the GEMM until a general matrix multiplication operation is completed.

[0140] The above Figure 10 and the related descriptions correspond to the complete GEMM operation without masks. For the working modes that require masking, this solution improves the distribution network on this basis, specifically as follows:

[0141] For the second working mode corresponding to the second mask mode MASK1 and the third mask mode MASK2, for the target image window corresponding to the second mask mode MASK1, it is necessary to perform logical clustering on it so that the format is consistent with the target image window corresponding to MASK2, and then obtain the corresponding Q matrix and K matrix. That is, the second working mode is actually the working mode corresponding to the third mask mode MASK2.

[0142] As Figure 11 shown, the assumptions of the Q matrix and the K matrix and the fixing method of the K matrix are all the same as Figure 10The same as above, which will not be elaborated here. If the loop distribution constant = 2, then both the Q matrix and the K matrix are divided into 2 intervals. Among them, the Q matrix is divided into the first interval 1 and the first interval 2. The first interval 1 includes row vector 0 and row vector 1, and the first interval 2 includes row vector 2 and row vector 3; the K matrix is divided into the second interval 1 and the second interval 2. The second interval 1 includes column vector 0 and column vector 1, and the second interval 2 includes column vector 2 and column vector 3. When a = 1, 2, after fixing the K matrix:

[0143] In the first loop, the first vector (row vector 0[a, b, c, d]) in the first first interval (i.e., the first interval 1) of the Q matrix is multiplied by each column vector in the first second interval (the second interval 1) of the K matrix respectively; at the same time, the first vector (row vector 2[i, j, k, l]) in the second first interval (i.e., the first interval 2) of the Q matrix is multiplied by each column vector in the second second interval (the second interval 2) of the K matrix respectively. That is, through the first distribution network, the first row vector in the first interval of the Q matrix is transmitted to the X terminals of the first and second column multipliers of the GEMM, and at the same time, the first row vector in the second interval of the Q matrix is transmitted to the X terminals of the third and fourth column multipliers of the GEMM; finally, through the second distribution network, the calculation results of each column are input into the Adder Tree or cached.

[0144] In the second loop, the second vector (row vector 1[e, f, g, h]) in the first first interval (i.e., the first interval 1) of the Q matrix is multiplied by each column vector in the first second interval (the second interval 1) of the K matrix respectively; at the same time, the second vector (row vector 3[m, n, o, p]) in the second first interval (i.e., the first interval 2) of the Q matrix is multiplied by each column vector in the second second interval (the second interval 2) of the K matrix respectively. That is, through the second distribution network, the second row vector in the first interval of the Q matrix is transmitted to the X terminals of the first and second column multipliers of the GEMM, and at the same time, the second row vector in the second interval of the Q matrix is transmitted to the X terminals of the third and fourth column multipliers of the GEMM; finally, through the second distribution network, the calculation results of each column are input into the Adder Tree or cached. In this way, the calculation results of each column are added to obtain the result of multiplying the corresponding row vector and column vector of this column, and this result is an element in the target matrix.

[0145] From Figure 11 It can be seen that each multiplication calculation only occurs between image blocks filled with the same pattern, and there is no redundant multiplication calculation of different patterns. The number of loops is (1 / 2)×4 = 2. In this way, only two loops are required to complete the GEMM operation of a 4×4 matrix.

[0146] Determine whether zero-padding decompression is required for the write-back storage method according to whether the softmax skips the operation.

[0147] The third working mode corresponding to the fourth mask mode MASK3:

[0148] Such as Figure 12 shown, the assumptions of the Q matrix and the K matrix and the fixing method of the K matrix are all the same as Figure 10 the same, which will not be elaborated here. The loop distribution constant = 4, then both the Q matrix and the K matrix are divided into 4 intervals. Among them, the Q matrix is divided into the first interval 1, the first interval 2, the first interval 3, and the first interval 4. The first interval 1 includes the row vector 0, the first interval 2 includes the row vector 1, the first interval 3 includes the row vector 2, and the first interval 4 includes the row vector 3; the K matrix is divided into the second interval 1, the second interval 2, the second interval 3, and the second interval 4. The second interval 1 includes the column vector 0, the second interval 2 includes the column vector 1, the second interval 3 includes the column vector 2, and the second interval 4 includes the column vector 3. a = 1, 2, 3, 4. After fixing the K matrix:

[0149] In the first loop, multiply the first vector (row vector 0[a, b, c, d]) in the first first interval (i.e., the first interval 1) of the Q matrix by the column vector in the first second interval (the second interval 1) of the K matrix; at the same time, multiply the first vector (row vector 1[e, f, g, h]) in the second first interval (i.e., the first interval 2) of the Q matrix by the column vector in the second second interval (the second interval 2) of the K matrix; at the same time, multiply the first vector (row vector 2[i, j, k, l]) in the third first interval (i.e., the first interval 3) of the Q matrix by the column vector in the third second interval (the second interval 3) of the K matrix; at the same time, multiply the first vector (row vector 3[m, n, o, p]) in the fourth first interval (i.e., the first interval 4) of the Q matrix by the column vector in the fourth second interval (the second interval 4) of the K matrix. That is, transmit the first row vector in the first interval of the Q matrix to the X end of the first column multiplier of the GEMM through the first distribution network, at the same time transmit the first row vector in the second interval of the Q matrix to the X end of the second column multiplier of the GEMM, at the same time transmit the first row vector in the third interval of the Q matrix to the X end of the third column multiplier of the GEMM, and at the same time transmit the first row vector in the fourth interval of the Q matrix to the X end of the fourth column multiplier of the GEMM. Finally, input the calculation result of each column into the AdderTree or cache it through the second distribution network. In this way, the calculation results of each column are added to obtain the result of multiplying the corresponding row vector and column vector of this column, and this result is an element in the target matrix.

[0150] From Figure 12It can be seen that each multiplication calculation only occurs between image blocks filled with the same pattern, and there is no redundant multiplication calculation for different patterns. The number of loops is (1 / 4) × 4 = 1. In this way, only one loop is required to complete the GEMM operation of a 4×4 matrix.

[0151] The above is an example with a 4×4 matrix. It can be understood that the number of rows of the Q matrix = the number of columns of the K matrix = the window size = M. It can be understood that assuming counting starts from 0, the M row vectors of the Q matrix are successively denoted as Q[0], Q[1], Q[2], ……, Q[M - 1], and the M column vectors in the K matrix are successively denoted as K[0], K[1], K[2], ……, K[M - 1]. The M column X ports of the multiplier are successively denoted as the 0-column X port, the 1-column X port, the 2-column X port, ……, the (M 2 - 1)-column X port. After fixing the K matrix at the end, for the second working mode, the loop is as follows:

[0152] Loop 1: The input Q[0] of the Q matrix is transferred to the column X port, and the input of the Q matrix is transferred to the port;

[0153] Loop 2: Based on the previous loop index, the data read address is incremented by 1, and it is repeated times;

[0154] Loop N: Determine whether zero-padding decompression is required for the write-back storage method according to whether the softmax skips the operation.

[0155] For the third working mode, the loop is as follows:

[0156] Loop 1: The input Q[0] of the Q matrix is transferred to the column X port, and the input of the Q matrix is transferred to the column X port; the input of the Q matrix is transferred to the port; the input of the Q matrix is transferred to the port;

[0157] Loop 1: Based on the previous loop index, the data read address is incremented by 1, and it is repeated times;

[0158] Loop N: Determine whether zero-padding decompression is required for the write-back storage method according to whether the softmax skips the operation.

[0159] It should be noted that for the second working mode and the third working mode, before performing GEMM, the image blocks need to be logically clustered to ensure that redundant operations are accurately avoided during the GEMM operation. The foregoing only shows a simple logical clustering method when M = 2. Here, taking M = 4 as an example, the logical clustering process of the foregoing second mask mode MASK1 and fourth mask mode MASK is described.

[0160] As Figure 13 shown, for the Q matrix of the second mask mode MASK1, there are two types of image blocks. All the white image blocks are clustered together (numbered 0, 1, 4, 5, 8, 9, 12, 13), and all the image blocks filled with thin diagonal lines are clustered together (numbered 2, 3, 6, 7, 10, 11, 14, 15). Then, in the order after clustering, all the row vectors are arranged in the order of first white, and after all the white ones are arranged, then all the image blocks filled with thin diagonal lines are arranged to obtain the Q matrix for the foregoing GEMM operation. The corresponding K matrix is clustered in the same way. Here, the vectors corresponding to white filling can be arranged in the front, or the vectors corresponding to white filling can be arranged in the back, as long as the Q matrix and the K matrix adopt the same clustering method. Figure 13 The logical clustering method of the Q matrix and K matrix for the fourth mask mode MASK3 is also shown, which will not be elaborated here.

[0161] That is to say, in the embodiments of the present disclosure, the first mask mode MASK0 corresponds to an ordinary matrix and does not require logical clustering; for the second mask mode MASK1 and the third mask mode MASK2, the situation of the second mask mode MASK1 needs to be converted into the third mask mode MASK2. Half of the image blocks after sliding the window within the target image window are logically clustered together, and the other half of the shifted image blocks are also logically clustered together, and the output matrix position is fixed and related to M. For the fourth mask mode MASK3, there are four types of image blocks, and the number of each type of image block is M / 4. Each M 2 / 4 image blocks are logically clustered together, and the output matrix position is fixed and related to M. Finally, after logical clustering, the GEMM operation is performed in the above manner, effectively reducing the number of operations and saving hardware resources. 2 / 4, and each M

[0162] It should also be noted that the foregoing examples are all based on a fixed K matrix. In practice, it is also possible to first fix the Q matrix and then correspondingly adjust the input method of each column vector in the K matrix.

[0163] According to the above operations, the hardware resources saved in a single masked GEMM matrix operation are:

[0164]

[0165] Among them, P bot is the proportion of the target image window of the second mask mode in all target image windows, and is 2 / 9 in Figure 3 ; Save bot is the performance improvement brought by the processing of this solution; P rig is the proportion of the target image window of the third mask mode in all target image windows, and is 2 / 9 in Figure 3 ; Save rig is the performance improvement brought by the processing of this solution; P cor is the proportion of the target image window of the fourth mask mode in all target image windows, and is 1 / 9 in Figure 3 ; Save cor is the performance improvement brought by the processing of this solution.

[0166] Regarding M as an even number, the above formula can be simplified as:

[0167]

[0168] When pat = 4, W = H = 224, and M = 7, the improvement is 11.56%;

[0169] When pat = 16, W = H = 224, and M = 7, the improvement is 32.01%.

[0170] For the overall estimation of the swin transformer, the overall network improvement of the swin transformer is in:

[0171]

[0172] Among them, and both represent the number of stages of operations performed in the swin transformer, and improve i represents the corresponding improvement. According to the Swin-T and Swin-S(B, L) parameters, the network improvements can be obtained as 12.5% (fewer high-sampling layers) and 14.08% (more high-sampling layers) respectively. Among them, Swin-T, Swin-S, Swin-B, and Swin-L respectively represent swin transformers with different parameters.

[0173] In another embodiment of the present disclosure, referring to Figure 14 which shows a schematic structural diagram of the composition of a device 20 for extracting image features provided by the embodiment of the present disclosure. As shown in Figure 14 , the device 20 may include:

[0174] An acquisition unit 201, configured to acquire a target image window and determine a mask mode corresponding to the target image window;

[0175] A first determination unit 202, configured to determine a first calculation matrix and a second calculation matrix corresponding to the target image window;

[0176] A second determination unit 203, configured to determine a target ratio according to the mask mode; wherein, the target ratio represents: the ratio of the operation results that do not need to be masked during the multiplication operation of the first calculation matrix and the second calculation matrix;

[0177] An operation unit 204, configured to perform a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix; the target matrix is an image feature required in the image processing process.

[0178] In some embodiments, the acquisition unit 201 is configured to acquire an initial image, where the initial image includes a first number of initial image windows; perform a sliding window operation on the initial image to obtain a first image, where the first image includes a second number of intermediate image windows; the second number is greater than the first number; perform an image block movement on the first image to obtain a target image, where the target image includes a first number of target image windows; wherein, the operation results that do not need to be masked are: the operation results between image blocks from the same intermediate image window.

[0179] In some embodiments, the first determination unit 202 is configured to, if the target image window does not include the image blocks obtained by movement, determine that the mask mode corresponding to the target image window is the first mask mode; if the target image window includes the image blocks obtained by moving from the first position, determine that the mask mode corresponding to the target image window is the second mask mode, and the first position is in the first direction of the first image; if the target image window includes the image blocks obtained by moving from the second position, determine that the mask mode corresponding to the target image window is the third mask mode, and the second position is in the second direction of the first image; if the target image window includes the image blocks obtained by moving from the third position, determine that the mask mode corresponding to the target image window is the fourth mask mode, and the third position is at the intersection of the first direction and the second direction of the first image.

[0180] In some embodiments, the sliding distance of the sliding window operation is half of the side length of the target image window, and the second determination unit 203 is configured to, if the mask mode is the first mask mode, determine that the target ratio is 1; if the mask mode is the second mask mode or the third mask mode, determine that the target ratio is 1 / 2; if the mask mode is the fourth mask mode, determine that the target ratio is 1 / 4.

[0181] In some embodiments, the number of rows of the first calculation matrix and the number of columns of the second calculation matrix are both a third quantity, where the third quantity is the number of image patches included in the target image window, and the number of loops of the general matrix multiplication is the product of the target ratio and the third quantity.

[0182] In some embodiments, the operation unit 204 is configured to determine the reciprocal of the target ratio as the loop distribution constant; sequentially divide the third quantity of row vectors in the first calculation matrix into a loop distribution constant number of first intervals, and sequentially divide the third quantity of column vectors in the second calculation matrix into a loop distribution constant number of second intervals; the number of row vectors included in the first interval is equal to the number of column vectors included in the second interval and is equal to the number of loops; in the i-th loop of the general matrix multiplication, multiply the i-th row vector in the a-th first interval of the first calculation matrix by each column vector in the a-th second interval of the second calculation matrix; where i is an integer greater than 0 and less than or equal to the number of loops, and a is an integer greater than 0 and less than or equal to the loop distribution constant; where the result of multiplying each of the row vectors and column vectors is an element in the target matrix.

[0183] In some embodiments, the first calculation matrix is a query matrix, and the second calculation matrix is a key matrix; the image processing is image processing based on the swin transformer model.

[0184] It should be noted that the operations on the Q matrix and the K matrix can be implemented by a multiplier, a selector, a first distribution network, a second distribution network, etc. For example Figure 15 As shown, for the second mask pattern MASK1 and the third mask pattern MASK2, the second mask pattern MASK1 is the same as the third mask pattern MASK2 after logical clustering, and there are two types of image patches. Then, in each loop, a 2-1 selector is used to select one of the vectors corresponding to the two types of image patches and input it to the X terminal of the multiplier. One end of the 2-1 selection receives the vector of one type of image patch, the other end receives the vector of the other type of image patch, and the output end is connected to the first distribution network, and the first distribution network is also connected to the multiplier to feed the output of the selector to the X terminal of the adder. It should also be noted that this solution can reuse the adder, or there is 1 adder corresponding to every two row vectors. In Figure 15 only the 2-1 adder connected to vector 0 and vector 8 is shown, and vector 1 and vector 9, vector 2 and vector 10, vector 3 and vector 11,..., vector 7 and vector 15 can all be correspondingly connected to a 2-1 adder.

[0185] Another example Figure 16As shown, for the fourth mask pattern MASK2, there are four image blocks. In each loop, a 4-1 selector is used to select one of the vectors corresponding to the four image blocks and input it to the X terminal of the multiplier. Among them, the four input terminals of the 4-1 selection respectively receive the vectors of the four image blocks, and the output terminal is connected to the first distribution network. The first distribution network is also connected to the multiplier to supply the output of the selector to the X terminal of the adder. It should also be noted that this solution can reuse the adder, or there is 1 adder corresponding to every four row vectors. In Figure 16 only the 4-1 adder connected to vector 0, vector 2, vector 8, and vector 10 is shown. Vector 1 + vector 3 + vector 9 + vector 11, vector 4 + vector 6 + vector 12 + vector 14, and vector 5 + vector 7 + vector 13 + vector 15 can all be correspondingly connected to a 4-1 adder.

[0186] In other embodiments, the device 20 can also be implemented by software, or by a combination of software and hardware, and no specific limitation is made thereto.

[0187] It should be noted that the image processing device 20 provided in the embodiments of the present disclosure is used to implement the method for extracting image features in the foregoing embodiments. For the details not disclosed in this embodiment, please refer to the description of the foregoing embodiments for understanding, and will not be elaborated here.

[0188] The embodiments of the present disclosure perform data logic rearrangement on GEMM according to the fixed mask pattern of the swin transformer; a two-way (or four-way) selector is added to the distribution network on the basis of the conventional GEMM, so as to realize the distribution selection of the input vector data; the mask rule of GEMM can be applied to the normalization operation, reducing the computing storage energy consumption on the attention data path.

[0189] The advantages of this solution are at least: effectively reducing the unnecessary data computing energy consumption and operation time brought by the sliding window mask in the swin transformer GEMM, with the gain increasing as the feature map scale increases, and the gain is 12.5% - 14.08% according to the number of attentions in the high sampling layer; the mask pattern can be extended to the normalization layer (softmax layer), and no corresponding operation is performed on the masked data, thereby reducing the computing energy consumption of the softmax layer; the GEMM hardware resources introduced by supporting this mask scheduling are less, and it does not affect the hardware versatility.

[0190] Understandably, in this embodiment, a "unit" may be a part of a circuit, a part of a processor, a part of a program or software, etc. Of course, it may also be a module or non-modular. Moreover, the components in this embodiment may be integrated in one processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software function module.

[0191] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0192] Therefore, this embodiment provides a computer storage medium that stores a computer program. When the computer program is executed by at least one processor, it implements the steps of the method described in any one of the foregoing embodiments.

[0193] This embodiment also provides a computer program product that includes a computer program. When the computer program is executed by at least one processor, it implements the steps of the method described in any one of the foregoing embodiments.

[0194] Based on the above computer storage medium and computer program product, refer to Figure 17 , which shows a schematic structural diagram of the composition of an electronic device 30 provided by an embodiment of the present disclosure. As Figure 17 shown, the electronic device 30 may include: a communication interface 701, a memory 702, and a processor 703; each component is coupled together through a bus system 704. It can be understood that the bus system 704 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 17Each of the various buses is labeled as the bus system 704. Among them, the communication interface 701 is used for receiving and sending signals during the process of receiving and sending information to and from other external network elements;

[0195] The memory 702 is used for storing computer programs that can run on the processor 703;

[0196] The processor 703 is used for executing the method described in any one of the foregoing when running the computer program.

[0197] It can be understood that the memory 702 in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 702 of the systems and methods described herein is intended to include, but not be limited to, these and any other suitable types of memory.

[0198] The processor 703 may be an integrated circuit chip with the ability to process signals. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 703 or instructions in the form of software. The above-mentioned processor 703 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 702, and the processor 703 reads the information in the memory 702 and combines its hardware to complete the steps of the above method.

[0199] It can be understood that the embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For a hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present disclosure, or a combination thereof.

[0200] For a software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0201] In another embodiment of the present disclosure, another electronic device is further provided, which includes the device 20 described in any one of the foregoing embodiments.

[0202] As described above, only the preferred embodiments of the present disclosure are given, and they are not intended to limit the protection scope of the present disclosure.

[0203] It should be noted that in the present disclosure, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0204] The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the superiority or inferiority of the embodiments.

[0205] The methods disclosed in several method embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments.

[0206] The features disclosed in several product embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new product embodiments.

[0207] The features disclosed in several method or device embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0208] As described above, only the specific implementation manners of the present disclosure are given, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for extracting image features, the method comprising: Acquire a target image window, and determine a mask mode corresponding to the target image window; Determine a first calculation matrix and a second calculation matrix corresponding to the target image window; Determining a target ratio according to the mask mode; wherein the target ratio represents: the ratio of operation results that do not need to be masked when performing a multiplication operation on the first calculation matrix and the second calculation matrix; A general matrix multiplication operation is performed on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix; the target matrix is ​​an image feature required in the image processing process.

2. The method according to claim 1, wherein: The step of acquiring the target image window comprises: Acquire an initial image, wherein the initial image includes a first number of initial image windows; Performing a sliding window operation on the initial image to obtain a first image, wherein the first image includes a second number of intermediate image windows; the second number is greater than the first number; Performing image block movement on the first image to obtain a target image, wherein the target image includes the first number of the target image windows; The operation results that do not need to be masked are: operation results between the image blocks from the same intermediate image window.

3. The method according to claim 2, wherein: Determining a mask mode corresponding to the target image window includes: If the target image window does not include the moved image block, determining that the mask mode corresponding to the target image window is the first mask mode; If the target image window includes an image block moved from a first position, determining that the mask mode corresponding to the target image window is a second mask mode, and the first position is located in a first direction of the first image; If the target image window includes an image block moved from a second position, determining that the mask mode corresponding to the target image window is a third mask mode, and the second position is located in a second direction of the first image; If the target image window includes an image block moved from a third position, it is determined that the mask mode corresponding to the target image window is a fourth mask mode, and the third position is located at the junction of the first direction and the second direction of the first image.

4. The method according to claim 3, wherein: The sliding distance of the sliding window operation is half of the side length of the target image window; and the determining the target ratio according to the mask mode includes: If the mask mode is the first mask mode, determining the target ratio to be 1; If the mask mode is the second mask mode or the third mask mode, determining the target ratio to be 1 / 2; If the mask mode is the fourth mask mode, the target ratio is determined to be 1 / 4.

5. The method according to claim 1, wherein: The number of rows of the first calculation matrix and the number of columns of the second calculation matrix are both a third number, the third number is the number of image blocks contained in the target image window, and the number of cycles of the general matrix multiplication is the product of the target ratio and the third number.

6. The method according to claim 5, wherein: Performing a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix includes: determining the inverse of the target ratio as a cyclic partition constant; Sequentially divide the third number of row vectors in the first calculation matrix into the cyclic allocation constant number of first intervals, and sequentially divide the third number of column vectors in the second calculation matrix into the cyclic allocation constant number of second intervals; the number of row vectors included in the first interval is equal to the number of column vectors included in the second interval, which is equal to the number of cycles; In the i-th loop of the general matrix multiplication, the i-th row vector in the a-th first interval of the first calculation matrix is ​​multiplied with each column vector in the a-th second interval of the second calculation matrix; wherein i is an integer greater than 0 and less than or equal to the number of loops, and a is an integer greater than 0 and less than or equal to the loop allocation constant; The result obtained by multiplying each row vector and the column vector is an element in the target matrix.

7. The method according to any one of claims 1 to 6, wherein: The first calculation matrix is ​​a query matrix, the second calculation matrix is ​​a key matrix; and the image processing is image processing based on a swin transformer model.

8. A device for extracting image features, comprising: An acquisition unit, used for acquiring a target image window and determining a mask mode corresponding to the target image window; A first determining unit, used to determine a first calculation matrix and a second calculation matrix corresponding to the target image window; a second determination unit, configured to determine a target ratio according to the mask mode; wherein the target ratio represents: a ratio of operation results that do not need to be masked when performing a multiplication operation on the first calculation matrix and the second calculation matrix; A calculation unit is used to perform a general matrix multiplication operation on the first calculation matrix and the second calculation matrix according to the target ratio to obtain a target matrix; the target matrix is ​​an image feature required in the image processing process.

9. An electronic device comprising a memory and a processor, wherein: The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the method according to any one of claims 1 to 7 when running the computer program.

10. A computer storage medium storing a computer program, wherein the computer program is executed by at least one processor to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Two-dimensional adaptive constant false alarm rate detection method of gridding complementary reference sliding window

    CN120742288A

  • A two-dimensional adaptive constant false alarm rate detection method of grid complementary reference sliding window

    CN120742288B