Object Detection Method Based on the Interaction of Transformer Global and Local Attention
By introducing a global and local attention interaction mechanism into the Transformer model, the problem of high computational cost and insufficient interaction of the Transformer model is solved, and more efficient object detection effect is achieved.
Patent Information
- Application Number
- CN202210399175.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-15
AI Technical Summary
The Transformer model in the prior art has high computational cost, high complexity, and insufficient global and local interaction, resulting in low accuracy and accuracy of target detection results.
The object detection method based on Transformer's global and local attention interaction is adopted. By dividing the to be processed images into 4*4 image tokens, and performing multiple global local attention feature transformations, combining image token merging and cross-scale global local attention calculation, the global and local information interaction of feature extraction is enhanced.
It effectively alleviates the resource consumption problem of Transformer in intensive prediction tasks, improves the accuracy and accuracy of object detection, and provides an efficient feature extraction framework that is better than the previous SOTA effect.
Smart Images

Figure CN114743017B_ABST
Abstract
Description
Background Art
[0002] Object detection has always been a core task in the field of computer vision. Computers collect, store, and learn images of the real world, extract deep features, and finally accurately and efficiently capture the regions of interest in the images, and draw bounding boxes around the objects to obtain their category information and two-dimensional coordinate information. With the development of the intelligent and information era, object detection technology is increasingly penetrating into practical applications, such as autonomous driving, face recognition, public safety, etc., and has great practical research significance and value in the academic or business community.
[0003] Currently, the mainstream object detection methods are divided into deep learning methods based on traditional convolution and new model detection methods based on the Transformer self-attention mechanism. Traditional convolution methods are divided into two-stage and single-stage categories according to whether candidate boxes are generated. The two-stage method first learns to generate candidate boxes and then performs localization based on regression; the single-stage method does not generate candidate boxes but directly performs regression tasks based on the entire image. The Transformer model was first applied in the field of natural language understanding (NLP). It uses the encoder-decoder and self-attention mechanism to achieve parallel computing of information and breaks through the timing limitation of traditional convolution methods. The encoder is composed of several self-attention modules and feed-forward neural networks stacked together. Among them, the self-attention mechanism calculates the attention coefficients of the query vector Q and a series of key-value vectors K to represent the importance between data or features, and then acts on the value vector V, thereby screening a large amount of redundant information and focusing on its own information, reducing the dependence on external information. The overall structure of the decoder is similar to that of the encoder, except that it has a multi-head attention mechanism for interacting with the output of the encoder. Subsequently, the Transformer gradually expands to the visual field. Compared with the traditional convolution model, the detection model based on the Transformer self-attention mechanism, as the backbone network for information extraction, can not only more effectively judge the object category and position information by capturing high-level semantic features of the image, but also achieve parallel computing processing.
[0004] Generally speaking, the existing technologies still have the following problems: The network structures of the two-stage and single-stage methods based on deep learning are huge and complex, and the long-distance information dependence between pixels is lost, resulting in low detection accuracy; The Transformer model based on the self-attention mechanism makes up for the shortcomings of the vision limitation of the deep network learning model and has the ability to model long-distance features. However, the quadratic complexity of the global interaction of the self-attention mechanism hinders its application in dense prediction tasks. In addition, the extraction of global information is too concentrated, resulting in insufficient local and global interaction. Summary of the Invention
[0005] To solve the above problems in the prior art, namely, the high computational cost, high complexity, and insufficient global and local interactions of the Transformer model, resulting in low accuracy and precision of object detection results, the present invention provides an object detection method based on global and local attention interaction of the Transformer. The object detection method includes:
[0006] Divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and perform global-local attention feature transformation on the projected first initial feature map for a first set number of times to obtain a first feature map;
[0007] Merge the image tokens of the first feature map, and perform global-local attention feature transformation on the merged initial second feature map for a second set number of times to obtain a second feature map;
[0008] Merge the image tokens of the second feature map, and perform global-local attention feature transformation on the merged initial third feature map for a third set number of times to obtain a third feature map;
[0009] Merge the image tokens of the third feature map, and perform global-local attention feature transformation on the merged initial fourth feature map for a fourth set number of times to obtain a fourth feature map;
[0010] Input the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain object detection results.
[0011] In some preferred embodiments, the method of image token merging is as follows:
[0012] Merge every adjacent 2*2 image tokens of the first feature map / second feature map / third feature map into 1 image token, and finally achieve 2-fold downsampling of the resolution of the feature map and 2-fold upsampling of the feature dimension through a linear projection layer to obtain the initial second feature map / initial third feature map / initial fourth feature map.
[0013] In some preferred embodiments, the method of global-local attention feature transformation is as follows:
[0014] Perform layer normalization on the first initial feature map / second initial feature map / third initial feature map / fourth initial feature map;
[0015] Divide the layer-normalized feature map into non-overlapping multiple local windows with a set size, and perform multi-head self-attention calculation for each local window respectively to obtain local feature maps;
[0016] Respectively perform residual connections between the local feature maps and the corresponding initial feature maps, and respectively perform downsampling operations. Concatenate the downsampled image tokens into a global window, and perform multi-head self-attention calculation on the global window to obtain a global feature map;
[0017] Perform cross-scale global-local attention calculation on the local window and the global window to obtain a local feature map incorporating global information;
[0018] Perform window merging, layer normalization, and multi-layer perceptron operations on the local feature map incorporating global information to obtain the first feature map / second feature map / third feature map / fourth feature map.
[0019] In some preferred embodiments, the multi-head self-attention calculation is expressed as:
[0020]
[0021] where Q, K, and V represent the query matrix, key-value matrix, and value matrix obtained by splitting the feature map after expanding the feature dimension by 3 times through a linear layer. Each tensor of the matrix represents the pixel features of the window, B is the relative position offset matrix representing the relative positions between pixels, T represents matrix transpose, represents the relation matrix, represents the attention relation matrix, Softmax is a function that converts a set of attention coefficients into a probability distribution ranging from [0,1] and summing to 1, and d represents the number of channels;
[0022] The local multi-head self-attention calculation splits the number of channels of the query matrix Q, key-value matrix K, and value matrix into several groups, each group belonging to 1 head. Each head independently performs self-attention calculation, and the results of each head are horizontally concatenated, which is expressed as:
[0023] MultiHead(Q, K, V) = Concat(head 1 , …, head i , …, head h )
[0024] where h is the number of heads in the local multi-head self-attention calculation. In the stage of obtaining the first feature map, h = 3. Subsequently, in the stages of obtaining the second feature map, third feature map, and fourth feature map, h increases by a factor of 2. head i , i ∈ [1, h] is the result of performing self-attention calculation on the i-th group of query matrix Q, key-value matrix K, and value matrix V, and Concat is horizontal concatenation.
[0025] In some preferred embodiments, the cross-scale global-local attention calculation is expressed as:
[0026]
[0027] Among them, Q L is the local window query matrix, and each tensor of the matrix represents the pixel features of the local window. K G , V G are the global window key matrix and value matrix, and each tensor of the matrix represents the pixel features of the global window;
[0028] The global-local multi-head self-attention calculation splits the channel numbers of the query matrix Q L , the key matrix K G and the value matrix V G into several groups, each group belongs to 1 head, and each head independently performs self-attention calculation, and the results of each head are horizontally concatenated, which is expressed as:
[0029] GL-MultiHead(Q L , K G , V G ) = Concat(head GL-1 , …, head GL-i , …, head GL-h )
[0030] Among them, GL-h is the number of heads in the global-local multi-head self-attention calculation. In the stage of obtaining the first feature map, GL-h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, GL-h increases by a factor of 2. head GL-i , GL-i ∈ [GL-1, GL-h] is the result of the self-attention calculation of the i-th group of the query matrix Q L , the key matrix K G and the value matrix V G , and Concat is horizontal concatenation.
[0031] In some preferred embodiments, the first set number is 2, the second set number is 6, the third set number is 12, and the fourth set number is 1.
[0032] In some preferred embodiments, the target detection result includes the bounding box, target category, and position coordinates of the region of interest in the image to be processed.
[0033] On the other hand, the present invention proposes a target detection system based on the interaction between global and local attention of Transformer. The target detection system includes a preprocessing module, a stage one module, a stage two module, a stage three module, a stage four module, a feature fusion and target detection module;
[0034] The preprocessing module is configured to divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and obtain a first initial feature map;
[0035] The stage one module is configured to perform a first set number of global-local attention feature transformations on the first initial feature map to obtain a first feature map;
[0036] The stage two module is configured to perform image token merging on the first feature map, and perform a second set number of global-local attention feature transformations on the merged initial second feature map to obtain a second feature map;
[0037] The stage three module is configured to perform image token merging on the second feature map, and perform a third set number of global-local attention feature transformations on the merged initial third feature map to obtain a third feature map;
[0038] The stage four module is configured to perform image token merging on the third feature map, and perform a fourth set number of global-local attention feature transformations on the merged initial fourth feature map to obtain a fourth feature map;
[0039] The object detection module is configured to input the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain an object detection result.
[0040] In a third aspect of the present invention, an electronic device is proposed, including:
[0041] At least one processor; and
[0042] A memory communicatively connected to at least one of the processors; wherein,
[0043] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above object detection method based on the global and local attention interaction of Transformer.
[0044] In a fourth aspect of the present invention, a computer-readable storage medium is proposed, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above object detection method based on the global and local attention interaction of Transformer.
[0045] Advantages of the present invention:
[0046] (1) The object detection method based on the interaction of global and local attention of Transformer of the present invention uses a hierarchical vision Transformer backbone based on window attention to alleviate the quadratic complexity problem of Transformer in dense prediction tasks and the pixel size of images, greatly saving resource consumption and improving efficiency.
[0047] (2) The object detection method based on the interaction of global and local attention of Transformer of the present invention uses a global-local interaction mechanism to perform cross-scale and hierarchical interactions between each local window and the global window with rich high-level semantic information, making more full use of global information, solving the problem of insufficient interaction caused by the high concentration of global information, and further improving the accuracy and precision of subsequent object detection results.
[0048] (3) The object detection method based on the interaction of global and local attention of Transformer of the present invention can be used as a general feature extraction framework, combined with various detectors, with higher detection accuracy on the object detection public dataset COCO, and better performance than the Swin Transformer that achieved the SOTA effect before, providing an effective method for dense detection fields such as autonomous driving, face recognition, and vehicle detection. Brief Description of the Drawings
[0049] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present application will become more obvious:
[0050] Figure 1 is a schematic flowchart of the object detection method based on the interaction of global and local attention of Transformer of the present invention;
[0051] Figure 2 is a summary architecture diagram of the object detection method based on the interaction of global and local attention of Transformer of the present invention;
[0052] Figure 3 is a schematic diagram of local and global attention calculation of an embodiment of the object detection method based on the interaction of global and local attention of Transformer of the present invention;
[0053] Figure 4 is a schematic diagram of global and local attention interaction of an embodiment of the object detection method based on the interaction of global and local attention of Transformer of the present invention;
[0054] Figure 5 is a visualization diagram of the object detection result of an embodiment of the object detection method based on the interaction of global and local attention of Transformer of the present invention. Detailed implementation manners
[0055] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0056] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0057] A target detection method based on the interaction of global and local attention of Transformer according to the present invention, the target detection method includes:
[0058] Divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and perform global-local attention feature transformation on the projected first initial feature map for a first set number of times to obtain a first feature map;
[0059] Merge the image tokens of the first feature map, and perform global-local attention feature transformation on the merged initial second feature map for a second set number of times to obtain a second feature map;
[0060] Merge the image tokens of the second feature map, and perform global-local attention feature transformation on the merged initial third feature map for a third set number of times to obtain a third feature map;
[0061] Merge the image tokens of the third feature map, and perform global-local attention feature transformation on the merged initial fourth feature map for a fourth set number of times to obtain a fourth feature map;
[0062] Input the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain the target detection result.
[0063] To more clearly illustrate the target detection method based on the interaction of global and local attention of Transformer according to the present invention, the following combines Figure 1 Expand and describe each step in the embodiments of the present invention in detail.
[0064] The target detection method based on the interaction of global and local attention of Transformer in the first embodiment of the present invention includes steps S10 - step S50, and each step is described in detail as follows:
[0065] Step S10: Divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and perform the first set number of global-local attention feature transformations on the projected first initial feature map to obtain the first feature map.
[0066] As Figure 2 shown, it is the summary architecture diagram of the object detection method based on the global and local attention interaction of Transformer in the present invention. The feature processing process before object detection is divided into four stages. The process of obtaining the first initial feature map and the first feature map is called Stage 1, including image token embedding and 2 global-local interaction modules. The process of obtaining the second initial feature map and the second feature map is called Stage 2, including token downsampling and 6 global-local interaction modules. The process of obtaining the third initial feature map and the third feature map is called Stage 3, including token downsampling and 12 global-local interaction modules. The process of obtaining the fourth initial feature map and the fourth feature map is called Stage 4, including token downsampling and 1 global-local interaction module.
[0067] In an embodiment of the present invention, the size of the image to be processed is H*W*3, where H*W is the height and width of the image to be processed, and 3 is the original feature dimension of the image to be processed. The image to be processed is divided into (H / 4)*(W / 4) non-overlapping 4*4 image tokens through two-dimensional convolution, and the original feature dimension 3 of the image to be processed is converted to C through linear projection to obtain a feature map of (H / 4)*(W / 4)*C, that is, the first initial feature map.
[0068] As Figure 3 shown, it is the schematic diagram of local and global attention calculation of an embodiment of the object detection method based on the global and local attention interaction of Transformer in the present invention. The specific process is as follows:
[0069] Step S11: Perform layer normalization on the first initial feature map to accelerate training convergence and enhance the stability of the data feature distribution. Its calculation method is shown in Equation (1):
[0070]
[0071] where x and x′ are the pixel information of the image features of the feature map before and after normalization respectively, and μ and σ are the mean and variance of the pixels in the channels of the feature map before normalization.
[0072] Step S12: Divide the feature map after layer normalization into non-overlapping multiple local windows with a set size, and perform multi-head self-attention calculation on each local window respectively to obtain local feature maps.
[0073] In one embodiment of the present invention, for more effective modeling, taking image tokens as units, the 7×7 image tokens evenly divide the first initial feature map into multiple local windows in a non-overlapping manner.
[0074] Perform local multi-head self-attention calculation in units of local windows to enhance the correlation between pixels in the local windows. The calculation method is shown in Equation (2):
[0075]
[0076] Where Q, K, and V represent the query matrix, key-value matrix, and value matrix obtained by splitting the feature map after expanding the feature dimension by 3 times through a linear layer. Each tensor of the matrix represents the pixel features of the window. B is the relative position offset matrix representing the relative positions between pixels, T represents matrix transpose, represents the relationship matrix, represents the attention relationship matrix. Softmax is a function that converts a set of attention coefficients into a probability distribution with a range in [0, 1] and a sum of 1. d represents the number of channels.
[0077] The local multi-head self-attention calculation splits the channels of the query matrix Q, key-value matrix K, and value matrix into several groups. Each group belongs to 1 head, and each head independently performs self-attention calculation, and the results of each head are horizontally concatenated. Its representation is shown in Equation (3):
[0078] MultiHead(Q, K, V) = Concat(head 1 , …, head i , …, head h ) (3)
[0079] Where h is the number of heads in the local multi-head self-attention calculation. In the stage of obtaining the first feature map, h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, h increases by a factor of 2. head i , i ∈ [1, h] is the result of the self-attention calculation of the i-th group of the query matrix Q, key-value matrix K, and value matrix V. Concat is horizontal concatenation.
[0080] The specific process of performing local multi-head self-attention calculation in units of local windows to enhance the correlation between pixels in the local windows is described as follows:
[0081] Step S121, expand the feature dimension of the first feature map by 3 times through a linear layer and split it into the query matrix Q, key-value matrix K, and value matrix V.
[0082] Step S122: Calculate the feature inner product between each pixel in the query matrix Q and the key matrix K. To prevent the inner product from being too large, divide it by the square root of d to obtain the relationship matrix.
[0083] Step S123: Since the relative positions of the pixels in the local window are within the range of [-7 + 1, 7 - 1] in both the height and width dimensions, with a total of 13 values, two-dimensional relative position encoding is adopted. Set a learnable variable with a shape of 13 * 13. Obtain the relative position encoding from the relative encoding position index and add it to the relationship matrix to obtain the attention relationship matrix.
[0084] Step S124: Perform a softmax calculation on the attention relationship matrix in the last dimension to obtain the local attention relationship map. The calculation method is shown in Equation (4):
[0085]
[0086] where, z i represents the i-th inner product value in the attention relationship matrix, C is the number of tensors in the attention relationship matrix, and through the Softmax function, the output values of multi-classification can be converted into a probability distribution ranging from [0, 1] and summing to 1.
[0087] Step S125: Multiply the local relationship map by the value matrix V to obtain the local feature map after local window self-attention calculation.
[0088] Step S13: Perform residual connections between the local feature map and the first initial feature map respectively to update the local feature map, which is used to solve the problem of multi-layer network training and can make the model pay more attention to the current different parts. Then, perform downsampling operations on each updated local feature map through convolution respectively. To ensure the effectiveness of the global local attention calculation process, the number of image tokens downsampled for each local window is related to the stage. In the first stage, each local window is downsampled to 1 image token. As the number of stages increases, the number of image tokens generated by downsampling each local window increases by a factor of 4. Concatenate the downsampled image tokens into a global window and perform multi-head self-attention calculation on the global window to obtain the global feature map.
[0089] The global multi-head self-attention calculation process of the global feature map is the same as the local multi-head self-attention calculation process, that is, the process of Step S121 - Step S125.
[0090] Step S14, as Figure 4As shown in the figure, the global and local attention interaction schematic diagram of an embodiment of the object detection method based on the global and local attention interaction of Transformer of the present invention is used to perform cross-scale global-local attention calculation on the local window and the global window, breaking through the limitation that the Q, K, and V matrices come from the same feature space. The global information is supplemented to the local window through global-local interaction to obtain a local feature map integrated with global information. The calculation method is shown in Equation (5):
[0091]
[0092] Among them, Q L is the local window query matrix, and each tensor of the matrix represents the pixel features of the local window. K G , V G are the global window key matrix and value matrix, and each tensor of the matrix represents the pixel features of the global window.
[0093] The global-local multi-head self-attention calculation splits the channel numbers of the query matrix Q L , the key matrix K G and the value matrix V G into several groups. Each group belongs to 1 head, and each head independently performs self-attention calculation, and the results of each head are horizontally concatenated. Its representation is shown in Equation (6):
[0094] GL-MultiHead(Q L , K G , V G ) = Concat(head GL-1 , …, head GL-i , …, head GL-h ) (6)
[0095] Among them, GL-h is the number of heads in the global-local multi-head self-attention calculation. In the stage of obtaining the first feature map, GL-h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, GL-h increases by a factor of 2. head GL-i , GL-i ∈ [GL-1, GL-h] is the result of the self-attention calculation of the i-th group of the query matrix Q L , the key matrix K G and the value matrix V G . Concat is horizontal concatenation.
[0096] The specific process of performing cross-scale global-local attention calculation on the local window and the global window and supplementing the global information to the local window through global-local interaction includes:
[0097] Step S141, calculate each Q in the local window query matrix with a size of 7*7 Land the inner product of each K in the global window of size M*N. To prevent the inner product from being too large, divide it by the square root of d to obtain the relationship matrix. G
[0098] Step S142, similar to steps S123 and S124, add the relative position encoding to the relationship matrix in step S141 and perform Softmax calculation in the last dimension to obtain the global-local attention calculation graph of size (m*n)*(M*N).
[0099] Step S143, multiply the global-local attention calculation graph by V G to finally obtain the global-local interaction window with the same size as the original local window.
[0100] Step S15, perform window merging, layer normalization, and multi-layer perceptron operations on the local feature map integrating global information to obtain the first feature map.
[0101] Reshape the image features of multiple global-local interaction windows into an overall image feature, perform layer normalization on the overall image feature, and the method of layer normalization is the same as that in step S11. Pass the overall image feature after layer normalization through a multi-layer perceptron to obtain the first feature map.
[0102] Step S20, perform image token merging on the first feature map, and perform global-local attention feature transformation on the merged initial second feature map for a second set number of times to obtain the second feature map.
[0103] To make full use of image features and detect targets at different scales, the specific process of network stratification is as follows:
[0104] Step S21, select elements at intervals of 2 in the row and column directions of the first feature map output at the end of stage one (i.e., merge every adjacent 2*2 image tokens in the first feature map into 1 image token), and splice the selected elements into a tensor.
[0105] Step S22, project the tensor linearly to downsample the resolution of the first feature map in stage one by 2 times and increase the feature dimension by 2 times.
[0106] Step S23, project the tensor linearly to downsample the resolution of the "overall image feature" in step S15 by 2 times and increase the feature dimension by 2 times, denoted as the second initial feature map in stage two.
[0107] Step S24: Feed the second initial feature in Stage 2 into the "Global-Local Attention Module" in Stage 1 and repeat it 6 times (i.e., the second set number is 6) to process the second initial feature to obtain the second feature map in Stage 2. The processing process of the Global-Local Attention Module in Stage 2 for the feature map is the same as that in Stage 1, except that the input feature maps are different.
[0108] Step S30: Perform image token merging on the second feature map, and perform global-local attention feature transformation on the merged initial third feature map for a third set number of times to obtain the third feature map.
[0109] In Stage 3, 12 Global-Local Attention Modules (i.e., the third set number is 12) are used to perform the third initial feature processing to obtain the third feature map in Stage 3.
[0110] Step S40: Perform image token merging on the third feature map, and perform global-local attention feature transformation on the merged initial fourth feature map for a fourth set number of times to obtain the fourth feature map.
[0111] In Stage 4, 1 Global-Local Attention Module (i.e., the fourth set number is 1) is used to perform the fourth initial feature processing to obtain the fourth feature map in Stage 4.
[0112] Step S50: Input the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain the target detection result.
[0113] The target detection result includes the bounding box, target category, and position coordinates of the region of interest in the image to be processed.
[0114] As Figure 5 shown, it is the visualization diagram of the target detection result of an embodiment of the target detection method based on the interaction between the global and local attention of Transformer in the present invention. The picture is from the COCO dataset, and the target detection module is the CascadeMask R-CNN network.
[0115] Although the above steps are described in the above sequential order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0116] The target detection system based on the interaction between the global and local attention of Transformer in the second embodiment of the present invention. This target detection system includes a preprocessing module, a Stage 1 module, a Stage 2 module, a Stage 3 module, a Stage 4 module, a feature fusion and target detection module;
[0117] The preprocessing module is configured to divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and obtain the first initial feature map;
[0118] The stage one module is configured to perform a first set number of global-local attention feature transformations on the first initial feature map to obtain the first feature map;
[0119] The stage two module is configured to perform image token merging on the first feature map, and perform a second set number of global-local attention feature transformations on the merged initial second feature map to obtain the second feature map;
[0120] The stage three module is configured to perform image token merging on the second feature map, and perform a third set number of global-local attention feature transformations on the merged initial third feature map to obtain the third feature map;
[0121] The stage four module is configured to perform image token merging on the third feature map, and perform a fourth set number of global-local attention feature transformations on the merged initial fourth feature map to obtain the fourth feature map;
[0122] The object detection module is configured to input the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain the object detection result.
[0123] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0124] It should be noted that the object detection system based on the interaction of global and local attention of Transformer provided in the above embodiments is only illustrated by dividing the above functional modules. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present invention, they are only used to distinguish each module or step, and are not regarded as an improper limitation of the present invention.
[0125] An electronic device according to the third embodiment of the present invention includes:
[0126] At least one processor; and
[0127] A memory communicatively connected to at least one of the processors; wherein,
[0128] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned object detection method based on the interaction between global and local attention of Transformer.
[0129] A computer-readable storage medium according to a fourth embodiment of the present invention, the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned object detection method based on the interaction between global and local attention of Transformer.
[0130] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-mentioned storage device and processing device can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0131] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0132] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.
[0133] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, so that a process, method, article, or device / equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to these processes, methods, articles, or devices / equipment.
[0134] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, those skilled in the art can easily understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A target detection method based on the interaction of global and local attention of Transformer, characterized in that, the target detection method includes: Dividing the image to be processed into 4*4 image tokens, linearly projecting them into high-dimensional vectors, and performing the first set number of global-local attention feature transformations on the projected first initial feature map to obtain the first feature map; Merging the image tokens of the first feature map, and performing the second set number of global-local attention feature transformations on the merged initial second feature map to obtain the second feature map; Merging the image tokens of the second feature map, and performing the third set number of global-local attention feature transformations on the merged initial third feature map to obtain the third feature map; Merging the image tokens of the third feature map, and performing the fourth set number of global-local attention feature transformations on the merged initial fourth feature map to obtain the fourth feature map; Inputting the feature information of the second feature map, the third feature map, and the fourth feature map into the detection head respectively to obtain the target detection result; The method of the global-local attention feature transformation is: Performing layer normalization on the first initial feature map / second initial feature map / third initial feature map / fourth initial feature map; Dividing the layer-normalized feature map into non-overlapping multiple local windows with a set size, and performing multi-head self-attention calculations for each local window respectively to obtain local feature maps; Performing residual connections between the local feature maps and the corresponding initial feature maps respectively, and performing downsampling operations respectively. Concatenating the downsampled image tokens into a global window, and performing multi-head self-attention calculation on the global window to obtain a global feature map; Performing cross-scale global-local attention calculation on the local window and the global window to obtain a local feature map integrated with global information; Performing window merging, layer normalization, and multi-layer perceptron operations on the local feature map integrated with global information to obtain the first feature map / second feature map / third feature map / fourth feature map; The multi-head self-attention calculation is expressed as: Among them, Q, K, and V represent the query matrix, key-value matrix, and value matrix obtained by splitting the feature map after expanding the feature dimension by 3 times through a linear layer. Each tensor of the matrix represents the pixel features of the window. B is the relative position offset matrix representing the relative positions between pixels, T represents matrix transpose, represents the relationship matrix, represents the attention relationship matrix. Softmax is a function that converts a set of attention coefficients into a probability distribution with a range in [0, 1] and a sum of 1. d represents the number of channels; The local multi-head self-attention calculation splits the channel numbers of the query matrix Q, the key-value matrix K, and the value matrix V into several groups, each group belongs to 1 head, and each head independently performs self-attention calculation, and horizontally concatenates the results of each head, which is expressed as: MultiHead(Q, K, V) = Concat(head 1 , …, head i , …, head h ) Among them, h is the number of heads in the local multi-head self-attention calculation. In the stage of obtaining the first feature map, h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, h increases by a factor of 2, head i , i ∈ [1, h] is the result of the self-attention calculation of the i-th group of query matrix Q, key-value matrix K, and value matrix V, and Concat is horizontal splicing; The cross-scale global-local attention calculation is expressed as: Among them, Q L is the local window query matrix, and each tensor of the matrix represents the pixel features of the local window. K G , V G are the global window key matrix and value matrix, and each tensor of the matrix represents the pixel features of the global window; The global-local multi-head self-attention calculation splits the number of channels of the query matrix Q L , the key-value matrix K G and the value matrix V G into several groups. Each group belongs to one head, and each head independently performs self-attention calculation, and the results of each head are horizontally concatenated, which is expressed as: GL-MultiHead(Q L ,K G ,V G ) = Concat(head GL-1 , …, head GL-i , …, head GL-h ) Among them, GL-h is the number of heads in the global-local multi-head self-attention calculation. In the stage of obtaining the first feature map, GL-h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, GL-h increases by a factor of 2, head GL-i , GL-i ∈ [GL-1, GL-h] is the query matrix Q of the i-th group L , key-value matrix K G and value matrix V G of the result of self-attention calculation, and Concat is horizontal splicing.
2. The target detection method based on the interaction of global and local attention of Transformer according to claim 1, characterized in that, the method of the image token merging is: Merging every adjacent 2*2 image tokens of the first feature map / second feature map / third feature map into 1 image token, and finally realizing 2-fold downsampling of the resolution of the feature map and 2-fold upsampling of the feature dimension through a linear projection layer to obtain the initial second feature map / initial third feature map / initial fourth feature map.
3. The target detection method based on the interaction of global and local attention of Transformer according to claim 1, characterized in that, The first set number is 2, the second set number is 6, the third set number is 12, and the fourth set number is 1.
4. The object detection method based on the interaction of global and local attention of Transformer according to claim 1, characterized in that, the object detection result includes the bounding box, object category and position coordinates of the region of interest of the image to be processed.
5. An object detection system based on the interaction of global and local attention of Transformer, characterized in that, the object detection system includes a preprocessing module, a stage one module, a stage two module, a stage three module, a stage four module, a feature fusion and object detection module; the preprocessing module is configured to divide the image to be processed into 4*4 image tokens, linearly project them into high-dimensional vectors, and obtain a first initial feature map; the stage one module is configured to perform the global-local attention feature transformation for the first set number of times on the first initial feature map to obtain a first feature map; the stage two module is configured to perform image token merging on the first feature map, and perform the global-local attention feature transformation for the second set number of times on the merged initial second feature map to obtain a second feature map; the stage three module is configured to perform image token merging on the second feature map, and perform the global-local attention feature transformation for the third set number of times on the merged initial third feature map to obtain a third feature map; the stage four module is configured to perform image token merging on the third feature map, and perform the global-local attention feature transformation for the fourth set number of times on the merged initial fourth feature map to obtain a fourth feature map; the feature fusion and object detection module is configured to input the feature information of the second feature map, the third feature map and the fourth feature map into the detection head respectively to obtain the object detection result; the method of the global-local attention feature transformation is: perform layer normalization processing on the first initial feature map / second initial feature map / third initial feature map / fourth initial feature map; divide the layer-normalized feature map into a plurality of non-overlapping local windows with a set size, and perform multi-head self-attention calculation for each local window respectively to obtain local feature maps; perform residual connection on the local feature maps and the corresponding initial feature maps respectively, and perform downsampling operations respectively, splice the downsampled image tokens into a global window, and perform multi-head self-attention calculation on the global window to obtain global feature maps; perform cross-scale global-local attention calculation on the local window and the global window to obtain local feature maps incorporating global information; perform window merging, layer normalization and multi-layer perceptron operations on the local feature maps incorporating global information to obtain the first feature map / second feature map / third feature map / fourth feature map; the multi-head self-attention calculation is expressed as: Among them, Q, K, and V represent the query matrix, key-value matrix, and value matrix obtained by splitting the feature map after expanding the feature dimension by 3 times through a linear layer. Each tensor of the matrix represents the pixel features of the window. B is the relative position offset matrix representing the relative positions between pixels, T represents matrix transpose, represents the relationship matrix, represents the attention relationship matrix. Softmax is a function that converts a set of attention coefficients into a probability distribution with a range in [0, 1] and a sum of 1. d represents the number of channels; The local multi-head self-attention calculation splits the number of channels of the query matrix Q, the key-value matrix K and the value matrix V into several groups, each group belongs to 1 head, each head independently performs self-attention calculation, and splices the results of each head horizontally, which is expressed as: MultiHead(Q, K, V) = Concat(head 1 , …, head i , …, head h ) Among them, h is the number of heads in the local multi-head self-attention calculation. In the stage of obtaining the first feature map, h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, h increases by a factor of 2, head i , i ∈ [1, h] is the result of the self-attention calculation of the i-th group of query matrix Q, key-value matrix K, and value matrix V, and Concat is horizontal splicing; The cross-scale global-local attention calculation is expressed as: Among them, Q L is the local window query matrix, and each tensor of the matrix represents the pixel features of the local window. K G , V G are the global window key matrix and value matrix, and each tensor of the matrix represents the pixel features of the global window; The global-local multi-head self-attention calculation splits the number of channels of the query matrix Q L , the key-value matrix K G and the value matrix V G into several groups. Each group belongs to one head, and each head independently performs self-attention calculation, and the results of each head are horizontally concatenated, which is expressed as: GL-MultiHead(Q L ,K G ,V G ) = Concat(head GL-1 , …, head GL-i , …, head GL-h ) Among them, GL-h is the number of heads in the global-local multi-head self-attention calculation. In the stage of obtaining the first feature map, GL-h = 3. Subsequently, in the stages of obtaining the second, third, and fourth feature maps, GL-h increases by a factor of 2, head GL-i , GL-i ∈ [GL-1, GL-h] is the result of self-attention calculation for the i-th group of query matrix Q L , key-value matrix K G and value matrix V G . Concat is for horizontal splicing.
6. An electronic device, characterized in that it includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the object detection method based on the interaction between the global and local attention of the Transformer according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the object detection method based on the interaction between the global and local attention of the Transformer according to any one of claims 1-4.
Citation Information
Patent Citations
Rotation equivariant space local attention remote sensing image target detection method
CN113850129A
System for and method of deep learning diagnosis of plaque erosion through optical coherence tomography
WO2022051211A1