Lightweight transformer for high resolution images
Patent Information
- Application Number
- CN202280033330.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-08
- Filing Date
- 2022-05-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-05-10
AI Technical Summary
然而,与变换器相关的计算复杂性随着像素数的增加而平方增加;因此,已知变换器的实现在计算上是昂贵的
Smart Images

Figure CN117296060B_ABST
Abstract
Description
Background Technology
[0001] Neural Architecture Search (NAS) is a technique for automatically designing artificial neural networks (ANNs), a commonly used model in machine learning. NAS has been used to design networks with architectures that outperform hand-designed architectures. NAS methods can be categorized based on the search space, search strategy, and performance estimation strategy. The search space defines the types of ANNs(s) that can be designed and optimized, the search strategy defines the process used to explore the search space, and the performance estimation strategy evaluates the performance of the ANN based on its design.
[0002] In image and computer visualization tasks, high-resolution representations (HR) are crucial for dense prediction tasks such as segmentation, detection, and pose estimation. In previous NAS methods focused on image classification, learning HR representations was often neglected. While NAS methods have achieved success in automatically designing efficient image classification models and improving model efficiency for dense prediction tasks such as semantic segmentation and pose estimation, existing NAS methods either directly extend the search space designed for image classification or only search for feature aggregation heads. This lack of consideration for the specificity of dense prediction hinders the performance advancement of NAS methods relative to the best handcrafted models.
[0003] In principle, dense prediction tasks require the integrity of both global context and high-resolution representation. The former is crucial for clarifying blurred local features on each pixel, while the latter is useful for accurately predicting fine details such as semantic boundaries and keypoint locations. However, the integrity of global context and high-resolution representation is not the focus of major NAS classification algorithms. Typically, multi-scale features are combined at the network ends, and recent methods have improved performance by placing multi-scale feature processing within the network backbone. Furthermore, multi-scale convolutional representations do not provide the global foreground of the image because dense prediction tasks typically have high input resolution, while the network usually covers a fixed receptive field. Therefore, global attention strategies such as Squeeze-and-Excitation Network (SENet) or nonlocal networks have been proposed to enrich the convolutional features of images. Transformers are widely used in natural language processing and have achieved good results when combined with convolutional neural networks for image classification and object detection. However, the computational complexity associated with transformers increases quadratically with the number of pixels; therefore, the implementation of known transformers is computationally expensive.
[0004] Embodiments have been described in relation to these and other general considerations. While relatively specific problems have been discussed, it should be understood that the examples described herein are not intended to limit the solution of the specific problems identified in the background section above. Summary of the Invention
[0005] Based on examples in this disclosure, systems and methods for High-Resolution Neural Architecture Search (HR-NAS) are described. By efficiently encoding multi-scale contextual information while preserving high-resolution representations, the HR-NAS implementation described herein can find efficient and accurate networks for various tasks. To better encode multi-scale image context in the HR-NAS search space, a lightweight transformer with computational complexity that can dynamically change with respect to different objective functions and computational budgets is used. To maintain the high-resolution representation of the learned network, HR-NAS employs a multi-branch architecture that provides convolutional encoding across multiple feature resolutions. Therefore, HR-NAS can be trained using an efficient, fine-grained search strategy that effectively explores the search space and determines the optimal architecture given various tasks and computational resources.
[0006] According to examples of this disclosure, a method for obtaining attention features is described. The method may include: receiving multiple labels associated with image features in a first-dimensional space at a projector of a transformer; generating projected features at the projector of the transformer by concatenating the multiple labels with a position map, the projected features having a second-dimensional space smaller than the first-dimensional space; receiving the projected features at an encoder of the transformer and generating an encoded representation of the projected features using self-attention; decoding the encoded representation at a decoder of the transformer and obtaining a decoded output; and projecting the decoded output onto the first-dimensional space and adding image features from the first-dimensional space to obtain the attention features associated with the image features.
[0007] According to an example of this disclosure, a system is described. The system may include one or more storage devices storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to implement a neural network for generating image attention features by processing image features combined with a two-dimensional location map. The neural network may include: a projector of a transformer configured to receive a plurality of tags associated with image features in a first-dimensional space and generate projected features by concatenating the plurality of tags with the two-dimensional location map, the projected features having a second-dimensional space smaller than the first-dimensional space; an encoder of the transformer configured to receive the projected features and generate an encoded representation of the projected features using self-attention; and a decoder configured to decode the encoded representation and obtain a decoded output, wherein the decoded output is projected onto the first-dimensional space and combined with image features in the first-dimensional space to obtain attention features.
[0008] According to examples of this disclosure, a non-transient computer-readable storage medium is described, comprising instructions executable by one or more processors to perform a method. The method may include: receiving, at a projector of a transformer, a plurality of tags associated with image features in a first-dimensional space; generating, at the projector of the transformer, projection features by concatenating the plurality of tags with a position map, the projection features having a second-dimensional space smaller than the first-dimensional space; receiving the projection features at an encoder of the transformer and generating an encoded representation of the projection features using self-attention; decoding the encoded representation at a decoder of the transformer and obtaining a decoded output; and projecting the decoded output onto the first-dimensional space and adding image features from the first-dimensional space to obtain attention features associated with the image features.
[0009] This synopsis is provided to introduce a set of concepts in a simplified form, which will be further described in the detailed description below. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] Refer to the accompanying diagram below for examples of non-restrictive and non-exhaustive enumeration.
[0011] Figure 1 Details of an example transformer according to the examples in this disclosure are depicted.
[0012] Figure 2 Details of a multi-branch search space for dense prediction, which includes both multi-scale features and global context, are depicted according to an example of this disclosure.
[0013] Figure 3 Additional details are depicted for the multi-branch search space for dense prediction according to the examples of this disclosure.
[0014] Figure 4 Additional details of the search block according to the example of this disclosure are described.
[0015] Figure 5 Another example of a multi-branch search space for dense prediction is depicted according to the examples in this disclosure.
[0016] Figure 6 Details of a method for generating attention maps using a transformer, as illustrated in the examples of this disclosure, are described.
[0017] Figure 7 Details of a method for performing network architecture search according to an example of this disclosure are described.
[0018] Figure 8 This is a block diagram illustrating the physical components (e.g., hardware) of a computing system that can be used to implement aspects of this disclosure.
[0019] Figures 9A to 9B A mobile computing device is shown that can be used to practice various aspects of this disclosure.
[0020] Figure 10 An aspect of the architecture of a system for processing data, according to an example of this disclosure, is shown. Detailed Implementation
[0021] In the following detailed description, reference is made to the accompanying drawings, which form a part of the description, and specific embodiments or examples are illustrated therein by way of illustration. These aspects may be combined, other aspects may be utilized, and structural changes may be made. Embodiments may be implemented as methods, systems, or devices without departing from this disclosure. Therefore, embodiments may take the form of hardware implementations, entirely software implementations, or implementations combining software and hardware aspects. Accordingly, the following detailed description should not be considered limiting, and the scope of this disclosure is defined by the appended claims and their equivalents.
[0022] NAS methods have achieved significant success in automatically designing efficient image classification models. NAS has also been applied to improve model efficiency for dense prediction tasks, such as semantic segmentation and pose estimation. However, existing NAS methods for dense prediction either directly extend the search space designed for image classification or only search for feature aggregation heads. This lack of consideration for the specificity of dense prediction hinders the performance improvement of NAS methods relative to the best handcrafted models.
[0023] In principle, dense prediction tasks require the integrity of both global background and high-resolution representation. The former is crucial for clarifying blurred local features on each pixel, while the latter is useful for accurately predicting fine details such as semantic boundaries and keypoint locations. However, these principles, particularly HR representation, are not the focus of major NAS classification algorithms. Typically, multi-scale features are combined at the network's end, and recent methods have shown that performance can be improved by placing multi-scale feature processing within the network backbone. Furthermore, multi-scale convolutional representations cannot provide the global foreground of an image because dense prediction tasks typically have high input resolution, while networks often cover a fixed receptive field. Therefore, global attention strategies such as SENet or nonlocal networks have been proposed to enrich the convolutional features of images. Transformers are widely used in natural language processing and have achieved good results when combined with convolutional neural networks for image classification and object detection. However, the computational complexity associated with transformers increases quadratically with the number of pixels; therefore, implementing transformers is computationally expensive. According to examples in this disclosure, in-network multi-scale features and transformers are combined with NAS methods to obtain NAS that enables dynamic task objectives and resource constraints.
[0024] In the example, a dynamic downprojection strategy is utilized to overcome the computationally expensive costs associated with implementing transformers with image pixels. Therefore, a lightweight, plug-and-play transformer architecture that can be combined with convolutional neural structures is described. Furthermore, to search the fusion space of multi-scale convolutions and transformers, appropriate feature normalization, selection and balancing of fusion strategies are required. Thus, various model selections for multiple tasks can be generalized and optimized using the number of transformer-based queries.
[0025] According to examples in this disclosure, a super network, also known as "SuperNet," is first defined, where each layer of SuperNet includes a multi-branch parallel module followed by a fusion module. The parallel module includes search blocks with multiple resolutions, and the fusion module includes search blocks for feature fusion that determine how features from different resolutions are fused. Based on computational budget and task objectives, a fine-grained progressive shrinkage search strategy can be used to efficiently prune redundant blocks in the network and channels of the convolution and transformer queries, resulting in an efficient model. According to examples in this disclosure, an efficient transformer that can be easily combined with convolutional networks for image and computer vision tasks is described. According to examples in this disclosure, a multi-resolution search space including both convolutions and transformers is described to model multi-scale information and global context within the network for dense prediction tasks. Therefore, a transformer integrated into a resource-constrained NAS search space for image and computer vision tasks is described. According to examples in this disclosure, a resource-aware search method for determining efficient architectures for different tasks is described.
[0026] Figure 1 A neural network system, also referred to as transformer 102, is depicted according to an example of this disclosure. Transformer 102 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components, and techniques described below can be implemented. The transformer includes a projector 110, an encoder 104, and a decoder 106. Generally, both encoder 104 and decoder 106 are attention-based, i.e., both apply an attention mechanism (e.g., a multi-head self-attention configuration) to their respective received inputs while transforming the input sequence. In some cases, neither the encoder nor the decoder includes any convolutional layers or any recurrent layers. Projector 110 uses pointwise convolution (with batch normalization) to transform the channel dimension of the feature map from c + d p (in c This represents the channel number of the input feature X, and d p Location map P The channel number is reduced to a smaller dimension n, where nThis indicates the number of queries. Projector 110 can use bilinear interpolation to adjust the spatial dimensions of the feature map. s × s In other words, to reduce computational costs, the input features are projected using the projection function P(⦁). Projected to n × s × s The reduced size, of which n Indicates the number of queries, and s × s This is the reduced size of the space. Therefore, the projection process can be represented as... Where Concat represents the cascade operator, and the input sequence is... It is projection and planar embedding. It is a positional encoding that compensates for the loss of spatial information during self-attention. When d p When =2, It can be a two-dimensional location map that compensates for the loss of spatial information during self-attention. Compared to sinusoidal location encoding and learned embeddings, it contains two channels (i.e., d...). p =2) Two-dimensional position diagram P It is more efficient in terms of computational requirements for lightweight vision models. A two-dimensional position map can be obtained using the following formula:
[0027]
[0028]
[0029] 1×1 convolution and bilinear interpolation can be performed to obtain the projection P(⦁) and inverse projection in transformer 102. Original image features X 112 can be divided into n 108 labels are used to achieve a low-dimensional space. Each label 108 can be concatenated with a 2D location map P 114 at position 116 to reach the projected feature P 118. That is, the input image features X 112 is transformed into n tags A set, and n tags Each tag in the set includes embedded location information. s 2 Dimensional semantics. Then, the projected features can be... As a query, key and value Provided to encoder 104.
[0030] Encoder 104 includes a multi-head self-attention A(⦁) configuration 122, which allows encoder 104 to jointly engage with information at different locations. More specifically, the multi-head self-attention configuration A(⦁) 122 can be defined as: A , where head i = Attention ,in h It's the number of heads. d It is the hidden dimension of the subspace it participates in, and It is the learned embedding (weight).
[0031] Using residual connections, in addition and normalization operation 124, the output of multi-head self-attention A configuration 122 is combined with the input of multi-head self-attention A 122. The output of addition and normalization operation 124 is the encoder self-attention residual output, which is provided to position-wise feedforward network 126. Position-wise feedforward network F(⦁) 126 may include two linear transformations with ReLU activation between them; position-wise feedforward network F(⦁) 126 is applied to the participating features. Where the expansion ratio F(⦁) is set to 4, for example, , b 1 and b 2 These represent the weights and biases of the linear layer, respectively.
[0032] Therefore, encoder 104 can be made by This means that the label-by-label attention is first calculated. Then, a linear transformation is applied across spatial locations to obtain global attention features. F A residual connection is employed, from the addition &norm operation 124 around the feedforward network 126 to the addition &norm 128. The output of the encoder 104 is provided to the decoder 106.
[0033] Decoder 106 follows a similar process to encoder 104; the output from encoder 104 is provided to multi-head self-attention configuration A130, which also receives semantic queries. S 132. In other words, Q , K and V The multi-head self-attention configuration A130 is provided. The multi-head self-attention configuration A130 uses the output of encoder 104 as keys and values, and employs learnable semantic embeddings. (For example, n learnable s) 2The set of semantic embeddings is used as the query. Using residual connections, at append and normalization operations 138, the output of multi-head self-attention A configuration 130 is combined with the input of multi-head self-attention A130 to generate decoder self-attention residual output. The decoder self-attention residual output is fed to position-wise feedforward network F(⦁) configuration 136. Residual connections are used from addition & norm operations 134 around feedforward network 136 to addition & norm operations 138. Then, through the inverse projection function Project the output of decoder 106 back to the original feature size. Then add it to the image features. X 112. Since image modeling is not a prediction task and there is no temporal relationship between queries in semantic embedding, the first multi-head attention configuration in the standard transformer decoder (i.e., the first multi-head attention configuration that provides input to multi-head attention configuration 130) can be omitted from decoder 106.
[0034] The time complexities of multi-head self-attention and feedforward networks are respectively and ,in s 2 , d and n Located in the lower-dimensional space of the projection. Because s 2 Because it is a small space size for projection, the total time complexity of transformer 102 is (FlOP). and n 2 d Approximately linear. Therefore, in some examples, the transformer 102 can be utilized in a fine-grained search strategy to reduce and select appropriate... n To further improve the efficiency of converter 102.
[0035] Unrestricted differences between transformer 102 and the standard transformer include the use of a projection function P(⦁) to learn self-attention in a low-dimensional space; and the use of a two-dimensional position map. P Instead of sinusoidal position encoding; the first multi-head attention and spatial encoding in the standard transformer decoder are omitted; and the output of encoder 104 is directly used as the key and value of decoder 106 with residual connections (e.g., residual connections around multi-head self-attention A configuration 130).
[0036] Based on the examples in this disclosure, Figure 2A multi-branch search space 202 for dense prediction is depicted, which includes multi-scale features and global context while maintaining high-resolution representation throughout the neural network. SuperNet 204 is a multi-branch network comprising multiple search blocks 210, each search block including at least one convolutional layer 214; for example, search block 210 may also include a transformer 212. Transformer 212 may be the same as or similar to transformer 102 previously described in this disclosure. Unlike previous search methods for specific tasks, the network search network can be customized for a variety of dense prediction tasks. The multi-branch search space may include a parallel module 208 and a fusion module 206. In the example, the parallel module 208 and the fusion module 206 are configured alternately. For example, the fusion module may be used after the parallel modules to exchange information between multiple branches. In the example, the parallel module 208 and the fusion module 206 utilize search block 210.
[0037] Figure 3 Additional details are depicted regarding the multi-branch search space for dense prediction, as illustrated in the examples of this disclosure. Figure 3 As shown, after one or more convolutional layers 304 reduce the feature solution to, for example, one-quarter of the image size, low-resolution convolutional branches are gradually added to high-resolution convolutional branches using feature fusion modules 306, 314, etc. The multi-resolution branches are connected in parallel using parallel modules (e.g., parallel modules 308, 312, 316, etc.). At 318, the multi-branch features are concatenated together and connected to the final classification / regression layer.
[0038] Parallel modules 320, which may be the same as or similar to parallel modules 308, 312, 316, etc., typically achieve a larger receptive field and multi-scale features by stacking search blocks in each branch. For example, search block 334A may reside between feature maps 322 and 324; search block 334B may reside between feature maps 324 and 326. Search blocks 334A and 334B may be the same or different. Feature maps 322, 324, and 326 are illustrative examples of higher-resolution feature maps. Similarly, search block 334C may reside between feature maps 328 and 330; search block 334D may reside between feature maps 330 and 332. Search blocks 334C and 334D may be the same or different. Search blocks 334A, 334B, 334C, and 334D may be the same or different. Feature maps 328, 330, and 332 are illustrative examples of feature maps with lower resolution than feature maps 322, 324, and 326. In these examples, parallel module 320 includes... Branches, containing nc 1 ,…nc m Convolutional layers, where each branch has nw1 ,…nw m Channel. In other words, a parallel module can be represented as [ m, [ nc 1 ,…, nc m ],[ nw 1 ,…, nwm ]).
[0039] Fusion module 336, which can be the same as or similar to fusion modules 306 and 314, can have... m in and m out The two parallel modules of the branch are used to perform feature interactions between multiple branches using element-wise addition. For each output branch, a search block is used to fuse adjacent input branches to unify the feature map size. For example, an 8× output branch contains information from 4×, 8×, and 16× input branches. Feature transformation from high resolution to low resolution is achieved through search blocks and upsampling. For example, search blocks, represented as arrows in fusion module 336, can reside between feature maps 338 and 334, 338 and 340, 342 and 340, 342 and 344, 342 and 348, 346 and 344, 346 and 348, and 346 and 350. As in the parallel modules, search blocks can be the same as each other or can be different from each other.
[0040] Figure 4 Additional details of a search block 406 according to an example of this disclosure are depicted. Search block 406 may be identical to search block 404 in a parallel module and / or search block 410 in a fusion module. In the example, the search block includes a convolutional layer 412 and at least one transformer 430, wherein the number of convolutional channels and the number of queries / tags in the at least one transformer are searchable parameters. In the example, the convolutional layer 412 in search block 406 is organized according to the efficient structure of an inverted residual block and includes at least one transformer 430 to enhance the global context. In some examples, the convolutional layer 412 may differ from... Figure 4 The configuration shown may otherwise include configurations different from those shown. Figure 4 The configuration shown is illustrated. Similarly, in some examples, search block 406 may include different... Figure 4 The modified converter of at least one converter 430 shown, or at least one converter 430 may be omitted entirely.
[0041] if c Let represent the channel number of the input feature X, and for simplicity, the spatial dimension h×w is omitted. Then, the first layer 414 can be defined as a 1×1 pointwise convolution.C 0 The first layer is defined as a 1×1 pointwise layer. To expand the input features to have 3 using convolutions 416, 418, and 420. r The expansion ratio is high-dimensional. Different kernel sizes of 3×3, 5×5, and 7×7 were added to the three parts of the expanded feature, respectively. Three deep convolutional layers. Then, the outputs of layers 424, 422, and 426 are concatenated, followed by pointwise convolutional layers. To reduce the number of channels to c’ (in the parallel module) c ' = c At the same time, it will have n A transformer T for each query is applied to the input features. X This is done to obtain global self-attention, which is then added to the final output. In this way, the transformer T is considered a residual path that enhances the global context within each search block. The information flow in the search block can be written as: ,in Indicates the first convolutional layer The i-th part of the output, such as Figure 4 As shown. In the example, convolution is used in the transformer. In the step size 2 and half-size inverse projection To reduce the number of search blocks. In this way, the entire SuperNet (e.g., Figure 3 (302) is constructed by shrinking the search block as described in this paper, by shrinking the depthwise convolutional channels of the transformer T while preserving multi-scale and global information. The query / tag functionality makes such a model easily adaptable to limited computational budgets.
[0042] SuperNet (e.g., Figure 3 The 302 network is a multi-branch network comprising search blocks, where each search block can include a mixture of convolutional layers and transforms. Unlike previous search methods tailored to specific tasks, networks for various dense prediction tasks can be customized to obtain optimal feature combinations for different tasks. For example, a resource-aware channel / query-based fine-grained search strategy can be used to explore optimal feature combinations for different tasks.
[0043] In the example, a lightweight model is generated using a progressively shrinking neural architecture search paradigm by discarding some convolutional channels and transformer queries during training. Within the search block (e.g., 406), 1×1 convolutional layers are utilized. This ensures that each unit has a fixed input and output size. Conversely, depthwise convolution... The interactions between channels can be minimized, making it easy to remove unimportant channels during the search process. For example, if If the channels in the convolution are unimportant and are removed, then the convolution... They can be adjusted to be and (where c and c) Representing convolution (Number of channels). Similarly, using projection P(⦁) and inverse projection... The transformer T can be designed to include a variable number of queries and tags. If the queries are discarded, the projection P(⦁) and Can be processed in low-dimensional space Size characteristics. Therefore, the tags and features of both the encoder's and decoder's transformers are automatically scaled. As an example, a search block (e.g., 406) could contain... There are learnable sublayers, among which c This is the number of channels in search block 406. r It is the expansion ratio, and n It is the number of markers.
[0044] In the example, a factor α > 0 can be learned along with the network weights to scale the output in each learnable sublayer (e.g., 406) of the search block. Less important channels and queries can be progressively discarded while maintaining the overall performance of the search block. In some examples, a resource-aware penalty to α might push other important factors close to zero. For example, the computational cost γ > 0 for each sublayer (e.g., 406) of the search block is used as a weighted penalty to fit a limited computational budget.
[0045] in As mentioned above; i It is an index of the sublevel. It is the number of residual queries (tags), and It is the first i The computational cost of each sub-layer. Therefore, in three-depth convolutions... In this context, γ can be a fixed value, while in the transformer T, it is a dynamic value set based on the number of residual queries. After adding a resource-aware penalty term, the overall training loss is:
[0046] in, Let λ represent the standard classification / regression loss with a weight decay term specific to the task, and λ represent the coefficient of the L1 penalty term. Weight decay can help limit the values of network weights to prevent them from becoming too large and making important factors α difficult to learn. Within a few epochs, as a time interval, sublayers with important factors less than a threshold ε can be removed, and the statistics of the batch normalization (BN) layers can be recalibrated. If all labels / queries of the transformer are removed, the transformer will degenerate into a residual path. When the search ends, the structure of the residuals can be used directly without fine-tuning.
[0047] Resource-aware L1 regularization can find a trade-off between accuracy and efficiency for different amounts of resource budget. Considering that FLOP is the most widely used and readily available metric, and is approximated as a lower bound on latency, FLOP can be used as a penalty weight. Similar approaches can be applied to other metrics. Furthermore, during the search process, multi-branch SuperNets can be customized for different tasks. For different tasks, different convolutional channels and transformer labels are preserved for different branches, thereby identifying the optimal low-level / high-level and local / global feature combinations for a specific task.
[0048] Figure 5 Additional details of a multi-branch search space for dense prediction according to an example of this disclosure are depicted. In the example, the multi-branch search space includes a high-resolution convolutional stream received at a first level, and high-to-low resolution streams are added sequentially to form new levels, and the multi-resolution streams are connected in parallel. Thus, the resolution of the parallel streams in a later level includes the resolution from the previous level, plus the added lower resolution. According to an example of this disclosure, a first fusion module 503 may receive a high-resolution convolutional stream 502 as input, wherein the high-resolution convolutional stream may be a first resolution 510. The first fusion module 503 may be the same as or similar to fusion module 306. The first fusion module 503 may add high-to-low resolution streams corresponding to a second step or resolution 512. For example, as indicated by the arrow, it may be combined with search block 406 ( Figure 4 The same or similar search block 524 can initiate a convolutional stream of second resolution 512.
[0049] Can be with Figure 3 Parallel modules 504, which are identical or similar to parallel modules 308 and / or 320, can stack search blocks indicated by arrows in each branch, where the first branch may correspond to a first resolution 510 and the second branch may correspond to a second resolution 512. The search blocks in parallel module 504 can be... Figure 4 The search block 406 is the same as or similar. It can be compared with... Figure 3Another fusion module 505, identical or similar to fusion module 336, can exchange information across multi-resolution representations (e.g., features of a first resolution 510 and features of a second resolution 512). Therefore, fusion module 505 can upsample feature information from the second resolution 512 and fuse that information with feature information from the first resolution 510. Similarly, fusion module 505 can downsample feature information from the first resolution 510 and fuse that information with feature information from the second resolution 512. Similar to fusion module 503, fusion module 505 can add a high-to-low resolution stream corresponding to the third step or resolution 514.
[0050] Parallel module 506 may be located between fusion module 505 and fusion module 507. Fusion module 507 may upsample feature information from the second resolution 512 and fuse this information with feature information from the first resolution 510. Similarly, fusion module 507 may downsample feature information from the first resolution 510 and fuse this information with feature information from the second resolution 512 and feature information upsampled from the third resolution 514. Fusion module 507 may downsample feature information from the second resolution 512 and fuse this information with feature information from the third resolution 514. Similar to fusion modules 503 and 505, fusion module 507 may add a high-to-low resolution stream corresponding to the fourth step or resolution 516. In the example, fusion module 507 and... Figure 3 The fusion module 314 is the same as or similar to it.
[0051] Parallel module 508 can reside between fusion module 507 and fusion module 509. Fusion module 509 can operate in a similar manner to fusion module 507, fusing feature information from various resolutions and adding a high-to-low resolution stream corresponding to the fifth step or resolution 518. In the example, the number of parallel modules and fusion modules can differ. Figure 3 , Figure 4 and / or Figure 5 The number described in the example. In the example, the number of fusion modules and feature modules may be more or less than the number of fusion modules and feature modules shown.
[0052] In the example, the search block indicated by the arrow can be search block 532A and / or 532B, where search block 532A can be related to search block 406 ( Figure 4Search block 406 may include a convolutional layer 412 and a transformer 430, similar to or the same as search block 406. In some examples, search block 532A may perform a feature transformation from low resolution to high resolution; in some examples, the resolution of the feature transformation may remain unchanged. In some examples, the search block implementing the high-to-low resolution feature transformation may implement search block 532B, wherein search block 532A may be similar to search block 406. Figure 4 Similarly, search block 406 may include convolutional layer 412 and transformer 430. Search block 532B may be referred to as a shrinking search block.
[0053] Figure 6 The details of a method 600 for generating attention maps using a transformer according to examples of this disclosure are described. Figure 6 The general sequence of steps in method 600 is shown. Generally, method 600 begins at 602 and ends at 618. Method 600 may include more or fewer steps, or may differ from the steps shown. Figure 6 The steps are arranged in the order shown. Method 600 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. In the example, aspects of method 600 are executed by one or more processing devices such as a computer or server. Furthermore, method 600 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), neural processing unit, or other hardware device. References will be incorporated herein by reference. Figures 1 to 5 Method 600 is explained by describing the system, components, modules, software, data structures, user interface, etc.
[0054] The method begins at 602, where the process can proceed to 604. At 604, one or more input feature maps can be received. To reduce computational cost, at 606, the input features are projected using the projection function P(⦁). X Projected onto a reduced size. Compared to sinusoidal position encoding and learned embeddings, a two-dimensional position map containing two channels... P It is more efficient in terms of computational requirements for lightweight visual models.
[0055] The encoder of the converter may include a multi-head self-attention configuration A(⦁), which allows the encoder to collectively engage with information from different locations. Furthermore, using a residual connection layer, the output of the multi-head self-attention configuration is combined with the input of the multi-head self-attention configuration A to generate an encoder self-attention residual output. The encoder self-attention residual output is provided to a feedforward network. At 608, the output from the encoder is provided to the decoder's multi-head self-attention configuration A, where the decoder's multi-head self-attention configuration A also receives semantic queries at 610. That is, the key from the encoder portion of the converter...K Sum V The multi-head self-attention configuration A is provided to the decoder; query Q is a learnable semantic embedding. S (For example, n A learnable s 2 (A set of semantic embeddings). Then, at 612, the decoder can be based on... Q , K and V Obtain the output. That is, the multi-head self-attention configuration A uses the encoder. F The output is used as keys and values, and learnable semantic embeddings are used as queries. Using residual connection layers, the output of the decoder's multi-head self-attention A configuration is combined with the input of multi-head self-attention A to produce the decoder's self-attention residual output. The output is fed to the location feedforward network F(⦁). The residual connections feed the inputs of the location feedforward network surrounding the feedforward network to the addition and normalization operations. Then, at 614, the decoder's output is passed through the inverse projection function. Projected back to the original feature size This is done to obtain attention features. These features can then be added to the image features. X In the example, the output of the transformer can be added to the convolutional layer within the search block (e.g., 406), as previously described. Method 600 can end at 618.
[0056] Figure 7 Details of a method 700 for performing network architecture search according to an example of this disclosure are described. Figure 7 The general sequence of steps in method 700 is shown. Generally, method 700 begins at 702 and ends at 716. Method 700 may include more or fewer steps, or may differ from the steps shown. Figure 7 The steps are arranged in the order shown. Method 700 can be executed as a set of computer-executable instructions that are executed by a computer system and encoded or stored on a computer-readable medium. In the example, aspects of method 700 are executed by one or more processing devices such as a computer or server. Furthermore, method 700 can be executed by gates or circuits associated with a processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), system-on-a-chip (SoC), neural processing unit, or other hardware device. References will be incorporated herein by reference. Figures 1 to 6 Method 700 is explained by describing the system, components, modules, software, data structures, user interface, etc.
[0057] The method begins at 702, where the process can proceed to 704. At 704, SuperNet is built or otherwise configured. SuperNet can be used with SuperNet 302 ( Figure 3 The search blocks are the same as or similar to the search blocks described above, and typically include one or more parallel modules and one or more fusion modules, wherein each parallel module and each fusion module may include a search block as described above (e.g., Figure 4 Each search block may include convolutional layers and transformers as previously described in the examples according to this disclosure. In the example, the convolutional layers of SuperNet can reduce the spatial dimension of image features. For example, the spatial dimension of image features can be reduced by a factor of four. Starting from the high-resolution branch of SuperNet, at 706, image features of a first resolution can be generated, for example, using a first plurality of stacked search blocks in a first parallel module. At 708, the first parallel module can generate image features of a second resolution. For example, the first parallel module may include a plurality of stacked search blocks at the first resolution level and a plurality of stacked search blocks at the second resolution level. Thus, image features of the first resolution can be generated by a plurality of stacked search blocks, and image features of the second resolution can be generated by a second plurality of stacked search blocks. At 710, the fusion module can generate multi-scale image features of the first resolution and multi-scale image features of the second resolution by fusing the image features of the first resolution and the image features of the second resolution. In the example, the search blocks in the fusion module can adjust the spatial dimension or resolution of image features by upsampling or downsampling depending on which branch the fusion module is located on. For example, high-to-low resolution image feature transformation can be achieved by shrinking the search blocks, while low-to-high resolution feature transformation can be achieved by using different search blocks. Therefore, the output branch of the fusion module can include information from multiple branches of SuperNet. In some examples, SuperNet can be pruned at 712. That is, some convolutional channels and transformer query portions of the search block can be discarded as previously described. Method 700 can end at 714.
[0058] Figure 8 This is a block diagram illustrating the physical components (e.g., hardware) of a computing system 800, which can be used to implement various aspects of this disclosure. The computing system components described below can be adapted to the computing and / or processing devices described above. In a basic configuration, the computing system 800 may include at least one processing unit 802 and a system memory 804. Depending on the configuration and type of the computing device, the system memory 804 may include, but is not limited to, volatile memory (e.g., random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM)), flash memory, or any combination of these memories.
[0059] System memory 804 may include an operating system 805 and one or more program modules 806 suitable for running software application 820, such as one or more components supported by the system described herein. As an example, system memory 804 may include one or more of the following: a transformer 821, a projector 822, an encoder 823, a decoder 824, SuperNet 825, a parallel module 826, a fusion module 827, a search block 828, and / or a convolution configuration 829. Transformer 821 may be the same as or similar to the previously described transformer 102. Projector 822 may be the same as or similar to the previously described projector 110. Encoder 823 may be the same as or similar to the previously described transformer 102. Decoder 824 may be the same as or similar to the previously described decoder 106. SuperNet 825 may be the same as or similar to the previously described SuperNet 302. Parallel module 826 may be the same as or similar to the previously described parallel module 320. Fusion module 827 may be the same as or similar to the previously described fusion module 336. Search block 828 may be the same as or similar to search block 406 previously described. Convolution configuration 829 may be the same as or similar to convolutional layer 412 as previously described. One or more components described in system memory 804 may include one or more of the other components described in system memory 804. For example, converter 821 may include encoder 823 and decoder 824. For example, operating system 805 may be adapted to control the operation of computing system 800.
[0060] Furthermore, the examples disclosed herein can be implemented in conjunction with graphics libraries, other operating systems, or any other application, and are not limited to any particular application or system. This basic configuration is... Figure 8 The components within the dashed line 808 are shown. The computing system 800 may have additional features or functions. For example, the computing system 800 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. This additional storage... Figure 8 The image shows a removable storage device 809 and a non-removable storage device 810.
[0061] As described above, multiple program modules and data files can be stored in system memory 804. When executed on processing unit 802, program module 806 (e.g., software application 820) can perform processing including but not limited to the aspects described herein. Other program modules that can be used according to various aspects of this disclosure may include email and contact applications, word processing applications, spreadsheet applications, database applications, PowerPoint presentation applications, drawing or computer-aided programs, etc.
[0062] Furthermore, embodiments of this disclosure can be implemented in circuits, discrete electronic components, packages containing logic gates or integrated electronic chips, circuits utilizing microprocessors, or on a single chip containing electronic components or a microprocessor. For example, embodiments of the invention can be implemented using a system-on-a-chip (SoC), wherein... Figure 8 Each or many of the components shown can be integrated onto a single integrated circuit. Such a SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “programmed”) onto a chip substrate as a single integrated circuit. When operating via the SOC, the functions described herein with respect to the client switching protocol can be operated via dedicated logic integrated onto a single integrated circuit (chip) along with other components of the computing system 800. Embodiments of the invention can also be implemented using other techniques capable of performing logical operations, such as AND, OR, and NOT, including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, embodiments of this disclosure can be implemented within a general-purpose computer or in any other circuit or system.
[0063] The computing system 800 may also have one or more input devices 812, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. One or more input devices 812 may include an image sensor. It may also include multiple output devices 814 such as a display, speaker, printer, etc. The above devices are examples, and other devices may also be used. The computing system 800 may include one or more communication connections 816 that allow communication with other computing devices / systems 850. Examples of suitable communication connections 816 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuitry; universal serial bus (USB), parallel and / or serial ports.
[0064] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, or program modules. System memory 804, removable storage device 809, and non-removable storage device 810 are examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic tape, disk storage or other magnetic storage devices, or any other manufactured product that can be used to store information and is accessible by computing system 800. Any such computer storage medium may be part of computing system 800. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0065] Communication media can be implemented by computer-readable instructions, data structures, program modules, or other data in modulated data signals, such as carrier waves or other transmission mechanisms, and include any information delivery medium. The term "modulated data signal" can describe a signal having one or more characteristics set or altered in a manner that encodes information in the signal. By way of example and not limitation, communication media can include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0066] Figures 9A to 9B A mobile computing device 900 is illustrated, such as a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., which can be used to practice the examples of this disclosure. In some examples, the mobile computing device 900 can utilize a trained search space and / or a trained model to perform one or more tasks, such as an image classification task. In other examples, the mobile computing device 900 can provide information to and receive information from a system such as computing system 800. In some examples, the mobile computing device 900 can be the same as or similar to computing system 800. In some aspects, the client can be a mobile computing device. Reference Figure 9A This illustrates one aspect of a mobile computing device 900 used to implement these aspects. In a basic configuration, the mobile computing device 900 is a handheld computer with both input and output elements. The mobile computing device 900 typically includes a display 905 and one or more input buttons 910 that allow the user to input information into the mobile computing device 900. The display 905 of the mobile computing device 900 can also be used as an input device (e.g., a touchscreen display).
[0067] If included, the optional side input element 915 allows for further user input. The side input element 915 can be a rotary switch, a button, or any other type of manual input element. Alternatively, the mobile computing device 900 may incorporate more or fewer input elements. For example, in some embodiments, the display 905 may not be a touchscreen.
[0068] In another alternative embodiment, the mobile computing device 900 is a portable telephone system, such as a cellular phone. The mobile computing device 900 may also include an optional keypad 935. The optional keypad 935 may be a physical keypad or a "soft" keypad generated on a touchscreen display.
[0069] In various embodiments, output elements include a display 905 for displaying a graphical user interface (GUI), a visual indicator 920 (e.g., a light-emitting diode), and / or an audio transducer 925 (e.g., a speaker). In some aspects, the mobile computing device 900 incorporates a vibration transducer for providing haptic feedback to the user. In yet another aspect, the mobile computing device 900 incorporates input and / or output ports, such as audio inputs (e.g., a microphone jack), audio outputs (e.g., a headphone jack), and video outputs (e.g., an HDMI port) for sending or receiving signals from external devices.
[0070] Figure 9B This is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, the mobile computing device 900 can be combined with a system (e.g., architecture) 902 to implement some aspects. In one embodiment, the system 902 is implemented as a "smartphone" capable of running one or more applications (e.g., browser, email, calendar, contact manager, messaging client, games, media client / player, and other applications). In some aspects, the system 902 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and wireless phone.
[0071] One or more applications 966 may be loaded into memory 962 and run on or associated with operating system 964. Examples of applications include telephone dialer programs, email programs, imaging programs, multimedia programs, video programs, word processing programs, spreadsheet programs, internet browser programs, messaging programs, map programs, etc. System 902 also includes a non-volatile storage area 968 within memory 962. Non-volatile storage area 968 may be used to store persistent information that should not be lost when system 902 is powered off. Applications 966 may use and store information in non-volatile storage area 968, such as emails or other messages used by email applications. A synchronization application (not shown) also resides on system 902 and is programmed to interact with a corresponding synchronization application residing on the host computer to keep the information stored in non-volatile storage area 968 synchronized with the corresponding information stored on the host computer. As should be understood, other applications may be loaded into memory 962 and run on the mobile computing device 900 described herein.
[0072] System 902 has a power supply 970 that can be implemented as one or more batteries. The power supply 970 may also include an external power source, such as an AC adapter or a live docking station for replenishing or charging the batteries.
[0073] System 902 may also include a radio interface layer 972 that performs functions of transmitting and receiving radio frequency communications. Radio interface layer 972 facilitates wireless connectivity between system 902 and the "external world" through a communications operator or service provider. Transmissions to and from radio interface layer 972 are conducted under the control of operating system 964. In other words, communications received by radio interface layer 972 can be propagated to application program 966 via operating system 964, and vice versa.
[0074] A visual indicator 920 can be used to provide visual notifications, and / or an audio interface 974 can be used to generate audible notifications via an audio transducer 925. In the illustrated embodiment, the visual indicator 920 is a light-emitting diode (LED), and the audio transducer 925 is a speaker. These devices can be directly coupled to a power supply 970 such that when activated, they remain on for the duration specified by the notification mechanism, even if the processor 960 and other components may be turned off to conserve battery power. The LED can be programmed to remain lit indefinitely until the user takes action to indicate the device's power-on status. The audio interface 974 is used to provide and receive audible signals to and from the user. For example, in addition to being coupled to the audio transducer 925, the audio interface 974 can also be coupled to a microphone to receive sound input, such as for telephone conversations. According to embodiments of this disclosure, the microphone can also be used as an audio sensor to control notifications, as described below. System 902 may also include a video interface 976, which enables the vehicle camera 930 to operate to record still images, video streams, etc.
[0075] The mobile computing device 900 implementing system 902 may have additional features or functions. For example, the mobile computing device 900 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. This additional storage... Figure 9B The non-volatile storage region 968 is shown in the middle.
[0076] As described above, data / information generated or captured by mobile computing device 900 and stored via system 902 can be locally stored on mobile computing device 900, or the data can be stored on any number of storage media accessible by the devices via radio interface layer 972 or via wired connections between mobile computing device 900 and individual computing devices associated with mobile computing device 900, such as server computers in a distributed computing network, like the Internet. It should be understood that such data / information can be accessed via mobile computing device 900 via radio interface layer 972 or via a distributed computing network. Similarly, according to known data / information transmission and storage devices, including email and collaborative data / information sharing systems, such data / information can be easily transferred between computing devices for storage and use.
[0077] Figure 10 This illustration shows one aspect of the architecture of a system for processing data received at a computing system from a remote source, such as a personal computer 1004, a tablet computing device 1006, or a mobile computing device 1008, as described above. The personal computer 1004, tablet computing device 1006, or mobile computing device 1008 may include one or more applications. The content at the server device 1002 may be stored in different communication channels or other storage types.
[0078] As described above, server device 1002 and / or personal computer 1004, tablet computing device 1006, or mobile computing device 1008 may use one or more of the previously described program modules or software applications 804. Figure 8 For example, server device 1002 may include transformer 1021 and / or SuperNet 1025; SuperNet 1025 may include a network model trained for a specific task (e.g., image classification) in an untrained state and / or after training.
[0079] Server device 1002 can provide data to / from client computing devices such as personal computer 1004, tablet computing device 1006, and / or mobile computing device 1008 (e.g., smartphone) via network 1015. As an example, the aforementioned computer system can be implemented in personal computer 1004, tablet computing device 1006, and / or mobile computing device 1008 (e.g., smartphone). In addition to receiving graphics data that can be preprocessed at the graphics originating system or post-processed at the receiving computing system, any of these embodiments of the computing device can obtain content from storage 1016.
[0080] Furthermore, the aspects and functions described herein can operate on distributed systems (e.g., cloud-based computing systems), where application functions, memory, data storage and retrieval, and various processing functions can operate remotely to each other via distributed computing networks such as the Internet or intranets. Various types of user interfaces and information can be displayed via in-vehicle computing device displays or via remote display units associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with on a wall projected thereon. Interaction with numerous computing systems on which embodiments of the present invention can be practiced includes keystroke input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures used to control the functions of the computing device, and so on.
[0081] For example, aspects of this disclosure have been described above with reference to block diagrams and / or operating instructions of methods, systems, and computer program products according to various aspects of this disclosure. Functions / actions indicated in the boxes may not occur in the order shown in any flowchart. For example, depending on the functions / actions involved, two consecutively displayed blocks may actually be executed substantially simultaneously, or sometimes these blocks may be executed in reverse order.
[0082] This disclosure relates to systems and methods for obtaining attention features, based at least on the examples provided in the following sections:
[0083] (A1) In one aspect, some examples include a method for obtaining attention features. The method may include: receiving a plurality of tags associated with image features in a first-dimensional space at a projector of a transformer; generating projection features at the projector of the transformer by concatenating the plurality of tags with a position map, the projection features having a second-dimensional space smaller than the first-dimensional space; receiving the projection features at an encoder of the transformer and generating an encoded representation of the projection features using self-attention; decoding the encoded representation at a decoder of the transformer and obtaining a decoded output; and projecting the decoded output onto the first-dimensional space and adding image features from the first-dimensional space to obtain attention features associated with the image features.
[0084] (A2) In some examples of A1, the method further includes: at the encoder of the transformer, applying self-attention to the projected features using a multi-head self-attention configuration that receives the projected features as keys, values and queries from the projector.
[0085] (A3) In some examples of A1-A2, the method further includes: applying self-attention to the result of the projected features and combining the key, value and query from the projector to generate encoder self-attention residual output; and processing the encoder self-attention residual output to generate an encoded representation.
[0086] (A4) In some examples of A1-A3, the method further includes: at the decoder of the transformer, applying self-attention to the encoded representation using a multi-head self-attention configuration that receives keys and values as inputs from the encoder and one or more semantic embeddings as queries.
[0087] (A5) In some examples of A1-A4, the method further includes: applying self-attention to the result of the encoded representation and combining it with keys and values from the encoder and one or more semantic embeddings to generate a decoder self-attention residual output; and processing the decoder self-attention residual output to generate a decoded output, wherein the decoded output is in a second-dimensional space.
[0088] (A6) In some examples of A1-A5, the projection features are obtained using bilinear interpolation.
[0089] (A7) In some examples of A1-A6, the location map includes a two-dimensional location map.
[0090] In another aspect, some examples include computing systems that include one or more processors and memory coupled to the one or more processors, the memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described herein (e.g., A1-A7 described above).
[0091] In yet another aspect, some examples include a non-transient computer-readable storage medium storing one or more programs for execution by one or more processors of a storage device, the one or more programs including instructions for performing any of the methods described herein (e.g., A1-A7 described above).
[0092] The descriptions and illustrations of one or more aspects provided in this application are not intended to limit or restrict the scope of the disclosure claimed in any way. The aspects, examples, and details provided in this application are considered sufficient to convey the best mode of possession and enable others to make and use the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. Various features (both structural and methodological) are intended to be selectively included or omitted, whether shown and described in combination or separately, to produce embodiments with a particular set of features. Having been provided with the descriptions and illustrations of this application, those skilled in the art can contemplate variations, modifications, and substitutions within the spirit of the broader aspects of the general inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.
Claims
1. A method for obtaining attention features, the method comprising: At the projector of the transformer, a plurality of tags associated with image features in a first-dimensional space are received, wherein the image features are divided into the plurality of tags; At the projector of the converter, a projection feature is generated by concatenating the plurality of markers with a position map, the projection feature having a second dimension space smaller than the first dimension space; The projected features are received at the encoder of the converter, and self-attention is used to generate an encoded representation of the projected features; The encoded representation is decoded at the decoder of the converter, and a decoded output is obtained; as well as The decoded output is projected onto the first dimensional space, and the image features in the first dimensional space are added to obtain attention features associated with the image features.
2. The method according to claim 1, further comprising: At the encoder of the transformer, self-attention is applied to the projected features using a multi-head self-attention configuration that receives the projected features as keys, values, and queries from the projector.
3. The method according to claim 2, further comprising: The result of applying the self-attention to the projected features is combined with the key, value, and query from the projector to generate the encoder self-attention residual output; as well as The encoder self-attention residual output is processed to generate the encoded representation.
4. The method according to claim 2, further comprising: At the decoder of the transformer, self-attention is applied to the encoded representation using a multi-head self-attention configuration that receives keys and values as inputs from the encoder and one or more semantic embeddings as queries.
5. The method according to claim 4, further comprising: The result of applying the self-attention to the encoded representation is combined with the key and the value from the encoder, as well as one or more semantic embeddings, to generate the decoder self-attention residual output; as well as The decoder self-attention residual output is processed to generate the decoded output. The decoding output is located in the second-dimensional space.
6. The method of claim 1, wherein the projection features are obtained using bilinear interpolation.
7. The method according to claim 1, wherein the location map includes a two-dimensional location map.
8. A system comprising: One or more storage devices store instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to implement a neural network for generating image attention features by processing image features combined with a two-dimensional location map, the neural network comprising: The projector of the converter is configured to receive a plurality of markers associated with image features in a first-dimensional space, and to generate projection features by concatenating the plurality of markers with the two-dimensional location map, the projection features having a second-dimensional space smaller than the first-dimensional space; The encoder of the transformer is configured to receive projected features and use self-attention to generate an encoded representation of the projected features; and A decoder is configured to decode the encoded representation and obtain a decoded output. The decoded output is projected onto the first dimensional space and combined with the image features in the first dimensional space to obtain the attention features.
9. The system of claim 8, wherein the encoder is configured to apply self-attention to the projected features at the encoder of the converter using a multi-head self-attention configuration, the multi-head self-attention configuration receiving the projected features as keys, values, and queries from the projector.
10. The system of claim 9, wherein the encoder is configured to: The result of applying the self-attention to the projected features is combined with the key, value, and query from the projector to generate the encoder self-attention residual output; and The encoder self-attention residual output is processed to generate the encoded representation.
11. The system of claim 9, wherein the decoder of the transformer is configured to apply self-attention to the encoded representation using a multi-head self-attention configuration, the multi-head self-attention configuration receiving keys and values as inputs and one or more semantic embeddings as queries from the encoder.
12. The system of claim 11, wherein the decoder is configured to: The result of applying the self-attention to the encoded representation is combined with the key and value from the encoder, along with one or more semantic embeddings, to generate a decoder self-attention residual output; and The decoder self-attention residual output is processed to generate the decoded output, wherein the decoded output is in the second-dimensional space.
13. The system of claim 8, wherein the projection features are obtained using bilinear interpolation.
14. A non-transient computer-readable storage medium comprising instructions executable by one or more processors to perform a method, the method comprising: At the projector of the transformer, a plurality of tags associated with image features in a first-dimensional space are received, wherein the image features are divided into the plurality of tags; At the projector of the converter, a projection feature is generated by concatenating the plurality of markers with a position map, the projection feature having a second dimension space smaller than the first dimension space; The projected features are received at the encoder of the converter, and self-attention is used to generate an encoded representation of the projected features; The encoded representation is decoded at the decoder of the converter, and a decoded output is obtained; as well as The decoded output is projected onto the first dimensional space, and the image features in the first dimensional space are added to obtain attention features associated with the image features.
15. The computer-readable storage medium of claim 14, wherein the method further comprises: At the encoder of the transformer, self-attention is applied to the projected features using a multi-head self-attention configuration that receives the projected features as keys, values, and queries from the projector.
16. The computer-readable storage medium of claim 15, wherein the method further comprises: The result of applying the self-attention to the projected features is combined with the key, value, and query from the projector to generate the encoder self-attention residual output; as well as The encoder self-attention residual output is processed to generate the encoded representation.
17. The computer-readable storage medium of claim 15, wherein the method further comprises: At the decoder of the transformer, self-attention is applied to the encoded representation using a multi-head self-attention configuration that receives keys and values as inputs from the encoder and one or more semantic embeddings as queries.
18. The computer-readable storage medium of claim 17, wherein the method further comprises: The result of applying the self-attention to the encoded representation is combined with the key and value from the encoder and one or more semantic embeddings to generate the decoder self-attention residual output; as well as The decoder self-attention residual output is processed to generate the decoded output, wherein the decoded output is in the second-dimensional space.
19. The computer-readable storage medium of claim 14, wherein the projection feature is obtained using bilinear interpolation.
20. The computer-readable storage medium of claim 14, wherein the location map comprises a two-dimensional location map.
Citation Information
Patent Citations
Driver action recognition method based on self-attention mechanism
CN112016459A
One-dimensional convolution position coding method of visual depth adaptive neural network
CN112801280A