Method and system for processing images based on local attention
Patent Information
- Application Number
- KR1020240176234
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2026-09-23
- Estimated Expiration
- 2044-12-02
Smart Images

Figure 112024133048234-PAT00026_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a technology for processing images based on a local attention mechanism. Background Technology
[0002] General self-attention mechanisms can learn global context by effectively identifying long-range dependencies, but they have limitations in recognizing fine local features.
[0003] Furthermore, when applied to recognize objects corresponding to specific patterns through image search, it presents a problem requiring significant computation time due to linear time complexity regarding the embedding dimension and quadratic complexity regarding the input size.
[0004] To solve this problem, a technique has been proposed that performs a primary search using global features and a secondary search using local features. Existing local attention mechanisms have the advantage of improving local feature recognition by applying self-attention to a pre-configured window size rather than the entire global feature map.
[0005] However, due to the nature of these existing local attention mechanisms extracting keys and queries within a fixed window size, they fail to reflect surrounding information extending beyond the window size, resulting in a problem where the ability to identify long-range dependencies is degraded. Prior art literature
[0006] U.S. Patent Application No. 17347416 Republic of Korea Patent Application No. 10-2021-0164889 The problem to be solved
[0007] The objective of the present invention is to solve the above problem by providing a local attention mechanism that can reflect surrounding information more than existing local attention mechanisms by updating a key considering a range wider than the region of interest and performing an attention operation with the updated key.
[0008] The objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood from the description below. means of solving the problem
[0009] A local attention-based image processing method according to one aspect of the present invention for achieving the aforementioned purpose relates to a method performed by an electronic device, comprising: a step of generating a linearly transformed key, value, and query according to a region of interest consisting of channels belonging to a preset window size obtained in a sliding window method from a feature map of an input image; a step of generating a final key feature map by updating the key value by applying a convolution filter, which is a region wider than the region of interest, to the initial key feature map generated in the step of generating the linearly transformed key, value, and query; a step of generating a score by performing an attention operation on the final key feature map and the query; and a step of outputting the score by multiplying it by the value.
[0010] A local attention-based image processing system according to another aspect of the present invention includes a convolution network that receives an input image and performs a convolution operation to output a feature map for the input image, and a local attention network that receives the feature map for the input image output from the convolution network and performs a local attention operation to output the result. The local attention network includes a linear transformation unit that generates a linearly transformed key, value, and query according to a region of interest consisting of channels belonging to a preset window size obtained in a sliding window manner from the feature map for the input image; a multiscale context kernel unit that applies a convolution filter, which is a region wider than the region of interest, to the initial key feature map generated in the step of generating the linearly transformed key, value, and query to update the key value and produce a final key feature map; a first multiplier that calculates a score by performing an attention operation on the final key feature map and the query; and a second multiplier that multiplies the score with the value and outputs the result. Effects of the invention
[0011] According to the present invention, by reflecting surrounding information more than existing local attention mechanisms, there is an effect of improving the function of identifying long-range dependency.
[0012] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in the claims. Brief explanation of the drawing
[0013] FIG. 1 is a flowchart of a local attention-based image processing method according to one embodiment of the present invention. FIG. 2 is a block diagram of a local attention-based image processing system according to another embodiment of the present invention. FIG. 3 is a diagram specifically illustrating a local attention layer in a local attention-based image processing system according to another embodiment of the present invention. FIG. 4 is a diagram specifically illustrating the multiscale context kernel portion of a local attention layer in a local attention-based image processing system according to another embodiment of the present invention. FIG. 5 is a diagram visualizing an original image including landmarks and a map activated thereon by the existing technology and the present invention. Specific details for implementing the invention
[0014] The advantages and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the claims. Meanwhile, the terms used in this specification are for describing the embodiments and are not intended to limit the present invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text.
[0015] The present invention relates to an image processing technology based on local attention operations.
[0016] In particular, the present invention is characterized by the technical feature of overcoming the limitation of existing local attention mechanisms, which perform attention only on a fixed region of interest and fail to consider the relationships between all pixels.
[0017] These technical features can be achieved by a configuration that updates the value of each pixel of an initial key feature map, which is linearly transformed from a region of interest according to a preset expansion factor, updates it to a final key feature map that considers a wider range of context than the region of interest, performs query and attention operations on the updated final key feature map to calculate a score, and outputs the score by multiplying the value.
[0018] A local attention-based image processing method according to one embodiment of the present invention can be performed by a local attention-based image processing system according to another embodiment of the present invention.
[0019] For the convenience of the following explanation, drawing symbols for functionally consistent content will be consistent, and explanations will be avoided.
[0020] Referring to FIG. 2, a local attention-based image processing system (10) may be configured to include a convolution network (110), a local attention network (120), and a pooling layer (130).
[0021] The convolution network (110) may produce global features for the input image by receiving the input image and performing a convolution operation.
[0022] Multiple convolution networks (110) can be connected in series with each other.
[0023] The convolution network (110) may include a positional embedding layer that performs positional embedding by receiving the input image or the output of the immediately connected convolution network, a convolution layer that performs a convolution operation by receiving the output of the positional embedding layer, and a feedforward layer that receives the output of the convolution layer to produce global features for the input image.
[0024] The local attention network (120) may receive a feature map for an input image output from a convolution network (110), perform a local attention operation, and output it.
[0025] Multiple local attention networks (120) can be connected in series with each other.
[0026] The local attention network (120) may be configured to include a positional embedding layer (121), a local attention layer (123), and a 2 Layer FFN (125).
[0027] The positional embedding layer (121) may receive a feature map for an input image output from a convolution network (110) or a feature map output from a local attention network connected immediately prior to it, and perform positional embedding.
[0028] The local attention layer (123) may receive the feature map output from the positional embedding layer (121), perform a local attention operation, and output it.
[0029] Referring to FIG. 3, the local attention layer (123) may include a linear transformation unit (310), a multiscale context kernel unit (320), a first multiplier (330), and a second multiplier (340).
[0030] The linear transformation unit (310) may generate a key (1011), value (1012), and query (1013) that are linearly transformed according to a region of interest (1001) consisting of channels belonging to a preset window size, which is acquired in a sliding window manner from a feature map of an input image (S310).
[0031] The multiscale context kernel (320, Multiscale Context kernel, MCK) may apply a convolutional filter, which is composed of a region wider than the region of interest, to the initial key feature map generated by the linear transformation unit (310) to update the key value and produce a final key feature map (1030) (S320).
[0032] The multiscale context kernel (320) may produce a final key feature map (1030) in which the key value is updated by applying an expanded convolution filter according to a preset expansion coefficient to the initial key feature map generated by the linear transformation (310) to consider the context of a region wider than the region of interest (S320).
[0033] Specifically, referring to FIG. 4, the multiscale context kernel section (320) may be configured to include extended convolutional filters (321, 322, 323), a connector (324), and a point-by-point convolutional filter (325, PWConv).
[0034] The expansion convolution filters (321, 322, 323) may be composed of each expansion coefficient corresponding to preset expansion coefficients, receive an initial key feature map generated in a linear transformation unit, and produce a key feature map (1021, 1022, 1023) updated for each expansion coefficient.
[0035] Original feature map For , based on the i-th and its surrounding j-th (j=1, … ,k) neighbor pixels, each confirmed coefficient ( Key values updated through an extended convolutional filter based on =1,2,3) ( ) can be expressed as shown in the following mathematical formula.
[0036]
[0037] Here, represents the preset kernel radius of the extended convolutional filter, and represents the weight for each pixel in the kernel, and represents a pre-set positional bias.
[0038] The connector (324) may concatenate the initial key feature map (1011) generated in the linear transformation unit (310) and the updated key feature maps (1021, 1022, 1023) produced in the extended convolution filters (321, 322, 323) channel by channel.
[0039] The pointwise convolution filter (325) may perform pointwise convolution on the output of the connector (324) to reduce it by a preset input dimension, thereby producing a final key feature map (1030) with updated key values having the same size as the initial key feature map (1011) generated in the linear transformation unit (310).
[0040] Final Key Feature Map( ) can be expressed as shown in the following mathematical formula.
[0041]
[0042] Here, channel-wise concatenation means pointwise convolution.
[0043] According to the above configuration, information from a wide coverage area can be captured while maintaining a small kernel size.
[0044] The first multiplier (330) may calculate a score (1040) by performing an attention operation on the final key feature map (1030) calculated in the multiscale context kernel section (320) and the query (1012) generated in the linear transformation section (310).
[0045] Score ) can be expressed as shown in the following mathematical formula.
[0046]
[0047] Here, represents the i-th extracted (projected) query feature, and represents the dot product, and represents the position of the neighbor pixel located at the j-th position relative to the i-th pixel, and represents a pre-set positional bias for the neighbor pixel located at the j-th position relative to the j-th pixel.
[0048] The second multiplier (340) may output the result (1050) of multiplying the score (1040) calculated in the first multiplier (330) with the value (1013) generated in the linear transformation unit (310).
[0049] The second multiplier (340) may output the score (1040) and the value (1013) by the hadamard multiplication.
[0050] The output of the second multiplier (340) can be expressed as shown in the following mathematical formula.
[0051]
[0052] Here, is a pre-configured scaling parameter, and , , represents the value of the j-th neighbor pixel based on the i-th pixel.
[0053] Based on the above configuration, multi-scale contexts can be integrated and local attention optimized.
[0054] According to the present invention, multi-scale information can be effectively integrated through various expansion factors, and based thereon, spatially rich representations can be generated.
[0055] In particular, by including positional information, high accuracy can be achieved when applied to image search, etc.
[0056] The 2 Layer FFN (125) may be composed of two feed-forward network layers and receive the output (1050) of the second multiplier (340) and pass it through the two feed-forward network layers to produce an output.
[0057] The pooling layer (130) may receive the output of the local attention network (120), perform a pooling operation, and output the result.
[0058] According to the present invention, a key feature map used directly for attention operations in a conventional local attention mechanism is updated by considering surrounding channels using an extended convolution for each channel of the key feature map, and by performing an attention operation with the updated key feature map, it is possible to reflect finer local features better than in a conventional local attention mechanism while reflecting the context of a wide area.
[0059] Therefore, we can expect the advantage of improved finer discrimination ability compared to existing local attention mechanisms and the effective identification of long-range dependencies.
[0060] Figure 5(a) is the original image containing landmarks, (b) is a visualization of the map activated in the NA Transformer based on the existing Neighborhood Attention, and (c) is a visualization of the map activated based on the local attention mechanism according to an embodiment of the present invention.
[0061] Referring to Figure 5, it can be seen that in (c), the left pillar of the landmark is effectively activated compared to (b), and the area that is activated but is not a landmark in the upper left corner of the image in (b) is hardly activated in (c).
[0062] In other words, it can be confirmed that the local attention-based image processing technology according to the present invention demonstrates superior performance in activating landmark features compared to existing technologies.
[0063] The present invention can provide higher performance than existing methods in computer vision tasks, such as image search, by providing an improved local attention mechanism.
[0064] Accordingly, the local attention-based image processing technology provided in the present invention can contribute to improving the overall accuracy and efficiency of deep learning models.
[0065] Meanwhile, the blocks of the attached block diagram and the steps of the flowchart may be implemented as computer instructions that perform specified functions by being loaded into the processor or memory of an electronic device capable of data processing (e.g., a general-purpose computer, a specialized computer, a portable notebook computer, or a network computer). Since these computer program instructions can be stored in computer-readable memory, the functions described in the blocks of the block diagram or the steps of the flowchart may be produced as manufactured products containing means of instruction to perform them.
[0066] A person skilled in the art to which the present invention pertains will understand that the present invention may be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the claims and their equivalents should be interpreted as being included within the scope of the present invention. Explanation of the symbols
[0067] 10 : Image Processing System 110 : Convolutional Network 120: Local Attention Network 121 : Positional Embedding Layer 123 : Local Attention Layer 310 : Linear transformation part 320: Multiscale Context Kernel 330 : 1st multiplier 340 : Second multiplier 125:2 Layer FFN 130 : Pooling layer
Claims
Claim 1 A local attention-based image processing method comprising: a step of generating a linearly transformed key, value, and query according to a region of interest consisting of channels belonging to a preset window size, which is acquired in a sliding window manner from a feature map of an input image; a step of generating a final key feature map by updating key values by applying a convolutional filter, which is a region wider than the region of interest, to the initial key feature map generated in the step of generating the linearly transformed key, value, and query; a step of generating a score by performing an attention operation on the final key feature map and the query; and a step of multiplying the score by the value and outputting it; wherein the step of generating the final key feature map comprises generating a key feature map updated for each expansion coefficient by applying an expansion convolutional filter composed of each expansion coefficient according to preset expansion coefficients to the initial key feature map, and generating a final key feature map with updated key values by combining the initial key feature map and the key feature maps updated for the preset expansion coefficients by channel and then using a point-by-point convolutional filter to reduce the input dimension. Claim 2 delete Claim 3 delete Claim 4 A convolution network that receives an input image, performs a convolution operation, and outputs a feature map for the input image; and a local attention network that receives a feature map for an input image output from the convolution network and outputs a local attention operation; wherein the local attention network comprises: a linear transformation unit that generates a linearly transformed key, value, and query according to a region of interest composed of channels belonging to a preset window size obtained in a sliding window manner from a feature map for an input image; a multiscale context kernel unit that applies a convolution filter, which is composed of an area wider than the region of interest, to the initial key feature map generated in the step of generating the linearly transformed key, value, and query to update the key value and produce a final key feature map; a first multiplier that calculates a score by performing an attention operation on the final key feature map and the query; and a second multiplier that multiplies the score with the value and outputs the result; wherein the multiscale context kernel unit comprises expansion convolution filters that receive the initial key feature map generated by the linear transformation unit, which is composed of each expansion coefficient corresponding to preset expansion coefficients, and produce a key feature map updated for each expansion coefficient; and the initial key feature map generated by the linear transformation unit and A local attention-based image processing system comprising a connector that combines the updated key feature maps produced by the above-mentioned extended convolutional filters by channel, and a point-by-point convolutional filter that reduces the output of the connector by a preset input dimension to produce a final key feature map with updated key values. Claim 5 delete Claim 6 delete