A method and system for visual task processing based on agent attention
Patent Information
- Application Number
- CN202311676395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-07
AI Technical Summary
[0003]但是,Softmax注意力计算所有查询-键对之间的相似性,存在与令牌数量呈二次复杂度的计算复杂性
[0055]本申请实施例提供了一种基于代理注意力的视觉任务处理方法,所述方法包括:获取待处理图像;将所述待处理图像输入预先训练的视觉任务处理模型,得到视觉任务处理结果,所述视觉任务处理结果为以下任意一项:图像分类结果、目标检测结果、语义分割结果和图像生成结果;其中,所述视觉任务处理模型为将transformer模型的原生的注意力模块替换为代理注意力模块而得到的,所述代理注意力模块包括第一softmax注意力模块和第二softmax注意力模块,所述第一softmax注意力模块使用第一三元组(A,K,V)进行注意力计算,得到代理特征VA,所述第二softmax注意力模块使用第二三元组(Q,A,VA)进行注意力计算,A是对Q进行处理得到的,Q,K,V为所述原生的注意力模块进行注意力计算所使用的三元组,Q表示查询令牌、A表示查询令牌Q的代理令牌、K表示键令牌、V表示值令牌。本申请将执行视觉处理任务时用到的视觉任务处理模型的原生的注意力模块替换为包括第一softmax注意力模块和第二softmax注意力模块的代理注意力模块,通过使用替换了代理注意力模块的视觉任务处理模型执行视觉处理任务,不仅具有较高的特征表现力,还具有低计算复杂度的优点。
Smart Images

Figure CN117671371B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual processing technology, and in particular to a visual task processing method and system based on agent attention. Background Technology
[0002] Currently, natural language processing (NLP) technology is rapidly gaining prominence in the field of computer vision, particularly in image classification, object detection, semantic segmentation, and multimodal tasks, where it has achieved significant success. Among these, applying softmax attention or linear attention to visual tasks is currently a widely used method for attention computation.
[0003] However, Softmax attention calculates the similarity between all query-key pairs, resulting in computational complexity that is quadratic with the number of tokens. Linear attention applies mapping functions to the query token Q and key token K respectively to change the computation order, which reduces computational complexity but suffers from insufficient expressive power.
[0004] Therefore, there is an urgent need for a new visual task processing method based on agent attention. Summary of the Invention
[0005] In view of the above problems, embodiments of this application provide a visual task processing method and system based on proxy attention, so as to overcome the above problems or at least partially solve the above problems.
[0006] In a first aspect, this application provides a visual task processing method based on proxy attention, the method comprising:
[0007] Obtain the image to be processed;
[0008] The image to be processed is input into a pre-trained visual task processing model to obtain a visual task processing result, which is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result;
[0009] The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) AAttention calculation is performed, where A is obtained by processing Q, and Q, K, V are the triples used by the original attention module for attention calculation. Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token.
[0010] Optionally, the agent attention module uses the following formula to calculate attention:
[0011] O A =σ(QA) T +B2)σ(AK T +B1)V
[0012] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters.
[0013] Optionally, the agent attention module uses the following formula to calculate attention:
[0014] O=σ(QA T +B2)σ(AK T +B1)V+DWC(V)
[0015] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters; and DWC represents depthwise convolution.
[0016] Optionally, the proxy token A is set as a learnable parameter, which is learned during the training of the proxy attention module, wherein the number n of the proxy token A is set as a hyperparameter, and n is an integer greater than or equal to 1.
[0017] Optionally, the proxy token A is obtained by performing the following operations on the query token Q:
[0018] The query token Q is obtained by performing a pooling operation, a transformation point association operation, or a token merging operation on the query token Q.
[0019] Optionally, the number of proxy tokens A is less than the number of query tokens Q.
[0020] Optionally, the method further includes:
[0021] The attention module in the object detection model to be trained is replaced with the proxy attention module, and the replaced object detection model is trained using the first training sample to obtain the target object detection model.
[0022] The object detection model to be trained includes: RetinaNet model, Mask R-CNN model, or CascadeMask R-CNN model.
[0023] Optionally, the method further includes:
[0024] The attention module in the segmentation model to be trained is replaced with the proxy attention module, and the replaced segmentation model to be trained is trained using the second training sample to obtain the target segmentation model.
[0025] The segmentation model to be trained includes either the SemanticFPN model or the UpperNet model.
[0026] Optionally, the method further includes:
[0027] The target stable diffusion model is obtained by replacing the attention module in the original stable diffusion model with the proxy attention module.
[0028] The native stable diffusion models include: the Stable Diffusion model or the ToMeSD model.
[0029] A second aspect of this application provides a visual task processing system based on proxy attention, the system comprising:
[0030] The acquisition module is used to acquire the image to be processed;
[0031] The input module is used to input the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result, wherein the visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result;
[0032] The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) A Attention calculation is performed, where A is obtained by processing Q, and Q, K, V are the triples used by the original attention module for attention calculation. Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token.
[0033] Optionally, the system includes: a first computing submodule;
[0034] The first calculation submodule uses the following formula to calculate attention:
[0035] O A =σ(QA) T +B2)σ(AK T +B1)V
[0036] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters.
[0037] Optionally, the system includes: a second computing submodule;
[0038] The second calculation submodule uses the following formula to calculate attention:
[0039] O=σ(QA T +B2)σ(AK T +B1)V+DWC(V)
[0040] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters; and DWC represents depthwise convolution.
[0041] Optionally, the system includes: a setting submodule;
[0042] The setting submodule is used to set the agent token A as a learnable parameter so that it can be learned during the training of the agent attention module. The number n of the agent token A is set as a hyperparameter, where n is an integer greater than or equal to 1.
[0043] Optionally, the proxy token A is obtained by performing the following operation on the query token Q, wherein the setting submodule includes:
[0044] A subunit is set up to obtain the proxy token A by performing a pooling operation on the query token Q, a transformation point association operation on the query token Q, or a token merging operation on the query token Q.
[0045] Optionally, the system further includes:
[0046] The first replacement submodule is used to replace the attention module in the object detection model to be trained with the proxy attention module, and to train the replaced object detection model to be trained with the first training sample to obtain the target object detection model.
[0047] The object detection model to be trained includes: RetinaNet model, Mask R-CNN model, or CascadeMask R-CNN model.
[0048] Optionally, the system further includes:
[0049] The second replacement submodule is used to replace the attention module in the segmentation model to be trained with the proxy attention module, and to train the replaced segmentation model to be trained using the second training sample to obtain the target segmentation model.
[0050] The segmentation model to be trained includes either the SemanticFPN model or the UpperNet model.
[0051] Optionally, the system further includes:
[0052] The third replacement submodule is used to replace the attention module in the original stable diffusion model with the proxy attention module to obtain the target stable diffusion model;
[0053] The native stable diffusion models include: the Stable Diffusion model or the ToMeSD model.
[0054] The beneficial effects of this application are:
[0055] This application provides a visual task processing method based on proxy attention. The method includes: acquiring an image to be processed; inputting the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result, wherein the visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result; wherein the visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module, the proxy attention module including a first softmax attention module and a second softmax attention module, wherein the first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain a proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) AAttention calculation is performed, where A is obtained by processing Q. Q, K, and V are the triples used by the original attention module for attention calculation, where Q represents the query token, A represents the surrogate token of the query token Q, K represents the key token, and V represents the value token. This application replaces the original attention module of the visual task processing model used in performing visual processing tasks with a surrogate attention module including a first softmax attention module and a second softmax attention module. By using the visual task processing model with the surrogate attention module replaced, visual processing tasks are performed, resulting in not only higher feature representation but also lower computational complexity. Attached Figure Description
[0056] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart illustrating the steps of a visual task processing method based on proxy attention provided in an embodiment of this application.
[0058] Figure 2 This is a schematic diagram of the model structure of a proxy attention module provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram of the information processing flow of a proxy attention module provided in an embodiment of this application;
[0060] Figure 4 This is a schematic diagram of a visual task processing system based on proxy attention provided in an embodiment of this application. Detailed Implementation
[0061] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0062] In a first aspect, this application provides a visual task processing method based on proxy attention, the method as follows: Figure 1 As shown, it specifically includes:
[0063] Step S101: Obtain the image to be processed;
[0064] Specifically, in this step, the system performs image acquisition, import, or other related operations to obtain image data to be processed. This can be done by acquiring images from a camera, storage device, network source, etc.
[0065] Step S102: Input the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result. The visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result.
[0066] The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) A) Attention calculation is performed. A is obtained by processing Q. Q, K, and V are the triples used by the original attention module for attention calculation. Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token.
[0067] In this step, the acquired image to be processed is input into a pre-trained visual task processing model to obtain the result of the visual task processing.
[0068] Specifically, a pre-trained visual task processing model is used. This model has been trained on a large-scale dataset and has learned the features and patterns of the image processing task. This model design includes a proxy attention module, which replaces the native attention module of the Transformer model. The proxy attention module in the visual task processing model consists of a first softmax attention module and a second softmax attention module. These two modules work together to achieve an efficient and flexible attention mechanism for the input image. The first softmax attention module uses a first triplet (A, K, V) for attention computation. The proxy feature V obtained through the first softmax attention module... A Global information was captured. The second softmax attention module uses the second triplet (Q, A, V). AAttention is calculated. Here, A is obtained by processing Q, where A is the surrogate token of the query token Q, K is the key token, and V is the value token. Finally, the result is processed by a visual task processing model containing the surrogate attention module to obtain the result of the visual task. This result can be an image classification result, an object detection result, a semantic segmentation result, or an image generation result, depending on the model design and task settings.
[0069] In a preferred embodiment of this application, a method is provided as follows: Figure 2 The diagram shows the model structure of the proxy attention module. This proxy attention module consists of two traditional Softmax attention modules: a first Softmax attention module and a second Softmax attention module. The first Softmax attention module is applied to the triple (A, K, V), where the proxy token A acts as a query token, used to aggregate information from the value token V, and attention is calculated between the proxy token A and the key token K. The second Softmax attention module is applied to the triple (Q, A, V). A Attention computation is performed on V, where V A This is the output of the first Softmax attention module, which is then processed by the second Softmax attention module to focus on the triples (Q, A, V). A Attention calculations are performed to obtain the final output of the proxy attention module. Intuitively, the proxy token A introduced in this application can be regarded as a proxy for the query token Q, because the proxy token A directly aggregates information from the key token K and value token V, and then passes the aggregated information to the query token Q. The query token Q no longer needs to communicate directly with the native key token K and value token V.
[0070] In some embodiments, the proxy attention module described in this application can be represented by the following formula (1):
[0071] in, This indicates the first Softmax attention module;
[0072] This indicates the second Softmax attention module.
[0073] As can be seen from formula (1), the proxy attention module proposed in this application consists of a first Softmax attention module and a second Softmax attention module. The first Softmax attention module can be regarded as proxy aggregation, and the second Softmax attention module can be regarded as proxy broadcasting. That is, in the proxy aggregation stage, in the first Softmax attention module, proxy token A is regarded as the proxy for query token Q, attention calculation is performed, and the information of all value tokens V is aggregated to obtain proxy feature V.A Then, during the proxy broadcast phase, A is treated as a key token, and V... A Treating it as a value token, a second attention calculation is performed with the query token Q, broadcasting the agent features from globally to each query token, resulting in the final output O.
[0074] In this application, the proxy token A acts as a proxy for the query token Q, aggregating global information from the key token K and the value token V and passing it back to the query token Q. This design aims to avoid calculating pairwise similarity between the query token Q and the key token K, while maintaining information exchange between each query-key pair through the proxy token A.
[0075] In this application, formula (1) can be further simplified to formula (2):
[0076] O A =σ(QA) T )σ(AK T )V (2)
[0077] Where σ(·) is the Softmax function.
[0078] In a preferred embodiment of this application, in order to better utilize location information, a proxy bias is introduced for the proxy attention module. Specifically, the proxy attention module with added proxy bias is shown in the following formula (3):
[0079] O A =σ(QA) T +B2)σ(AK T +B1)V (3)
[0080] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters. In practical applications, each surrogate bias can be constructed using three bias components. Specifically, three initial bias components need to be obtained first, and then the final surrogate bias is obtained by transforming and matrix adding the three initial bias components.
[0081] In this embodiment, the purpose of introducing surrogate bias in the surrogate attention module is to improve the model's performance when processing inputs at different locations or regions. Surrogate bias, by considering positional information, makes the model more likely to focus on specific parts of the input, thereby improving the model's expressive power and performance. The introduction of surrogate bias helps the model handle different locations or regions in the input sequence more flexibly, because in some tasks, such as vision tasks, different locations may have different importance. This is particularly helpful for tasks with long-range dependencies or requiring global information. In the surrogate attention module mentioned in this paper, surrogate bias introduces spatial information into attention calculation, enabling the model to better capture the structural and semantic information of the input by increasing the attention region of different surrogate tokens. This improvement can enhance the model's performance in various computer vision tasks (such as image classification, object detection, semantic segmentation, image generation, etc.).
[0082] In summary, adding surrogate bias to the surrogate attention module enhances the model's modeling of spatial structure: surrogate bias helps the model better understand the spatial structure of the input and improves its ability to focus on different locations. It also improves the model's expressive power: considering location information allows the model to handle various inputs more flexibly, thus enhancing its expressive capabilities.
[0083] In a preferred embodiment of this application, in order to improve the performance of the proxy attention module in handling feature diversity, this embodiment introduces depthwise convolution. Specifically, the proxy attention module with added depthwise convolution is shown in the following formula (4):
[0084] O=σ(QA T +B2)σ(AK T +B1)V+DWC(V) (4)
[0085] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters; and DWC represents depthwise convolution.
[0086] In this embodiment, adding deep convolution to the surrogate attention module enhances feature diversity. As a form of generalized linear attention, the surrogate attention module, due to its linear structure, may sometimes be insufficient to capture complex feature variations. The introduction of deep convolution helps maintain feature diversity by performing more complex transformations on features through convolution operations, thereby improving the model's ability to represent different features and enabling it to better adapt to different data patterns and structures. This is highly beneficial for handling visual tasks with diverse features, such as image classification, object detection, semantic segmentation, and image generation. Therefore, the introduction of deep convolution can further improve the performance of the surrogate attention module when processing visual tasks.
[0087] In a preferred embodiment, the proxy token A is set as a learnable parameter, which is learned during the training of the proxy attention module, wherein the number n of proxy tokens A is set as a hyperparameter, and n is an integer greater than or equal to 1.
[0088] Specifically, the proxy token A is configured as a learnable parameter. This means that during the training of the proxy attention module, the model automatically adjusts the value of the proxy token A through learning. The number n of proxy tokens A is set as a hyperparameter, representing a value that needs to be manually selected before model training. Typically, n is an integer greater than or equal to 1, which controls the number of proxy tokens, reducing computational complexity while maintaining global context modeling capabilities.
[0089] This setup allows the model to dynamically adjust the parameters of proxy token A during the learning process to adapt to different tasks and data. This also gives the model a degree of flexibility, allowing the number of proxy tokens and the learning process to be adjusted according to the requirements of a specific problem.
[0090] In some embodiments, the number of proxy tokens A can be set to be less than the number of query tokens Q.
[0091] Specifically, the number of proxy tokens A can be set to be less than the number of query tokens Q. In practical applications, the number of proxy tokens A can be set to be significantly less than the number of query tokens Q. Since the number of proxy tokens is small, the computational cost of calculating attention is relatively low, which improves the efficiency of the model.
[0092] In a preferred embodiment, the proxy token A is obtained by performing the following operations on the query token Q:
[0093] The query token Q is obtained by performing a pooling operation, a transformation point association operation, or a token merging operation on the query token Q.
[0094] In this embodiment, the proxy token A is obtained by performing the following operation on the query token Q.
[0095] Specifically, there are several ways to obtain proxy token A:
[0096] Through pooling: Proxy token A is obtained by pooling query token Q. This means that the model performs some form of aggregation on the query tokens to generate proxy token A.
[0097] Specifically, in the model, one or more query tokens Q are selected. These tokens can be specific positions in the input sequence or other task-related information. For each selected query token Q, its value is obtained. This typically refers to representational information extracted from the input features, such as embedding vectors or other feature representations. A pooling operation is applied to the values of the query tokens Q. In practice, common pooling operations include average pooling and max pooling. In average pooling, the values of the query tokens Q are averaged; in max pooling, the maximum value among the query tokens is selected. The result of the pooling operation is the surrogate token A. The surrogate token A contains aggregated information of the selected query tokens Q, representing the features of this set of query tokens Q. This method of obtaining surrogate tokens through pooling operations allows the model to integrate partial or global information from the input sequence to generate a more representative token, which can improve model efficiency, reduce computational complexity, and, in some cases, provide a more representative surrogate token.
[0098] By associating with transformation points: Proxy token A is obtained by associating with query token Q using transformation points. This may involve associating the query token with certain transformation points to generate proxy token A.
[0099] Specifically, the model selects one or more query tokens Q, which can represent specific locations in the input sequence or other task-related information. Deformation points are introduced, which can be predefined or learned by the model. These deformation points serve as reference points for the model to associate with the query tokens; they can be spatial locations, feature values, etc. The association between the query token Q and the introduced deformation points is calculated. This can be achieved by measuring their similarity or applying some association function. The purpose of the association is to capture the relationship between the query token Q and the deformation points. Finally, the result of the deformation point association is used to generate a surrogate token A. This approach allows the model to generate surrogate tokens in a more dynamic and flexible way by associating with deformation points. By introducing deformation points, the model can better capture structural or contextual information in the input sequence, making the generated surrogate tokens more expressive.
[0100] The proxy token A is obtained by performing a token merging operation on the query token Q. This may involve merging the query token with other tokens to form the proxy token A.
[0101] Specifically, one or more query tokens Q are selected in the model. These tokens can represent specific positions in the input sequence or other task-related information. Other tokens are selected for merging; these can be tokens from other positions related to query token Q or other task-related information. A token merging operation is performed, merging query token Q with the other selected tokens. The specific form of the merging operation may depend on the model design and can be simple vector concatenation, weighted averaging, etc., which are not limited here. The result of the merging operation forms a surrogate token A. Through the token merging operation, the model can integrate information from different tokens into the surrogate token, making the surrogate token more expressive. This method can be used to capture the correlation between different positions or other information, improving the model's ability to model the input sequence.
[0102] In this application, the proxy token A can be obtained by learning during the training of the proxy attention module, or by performing operations such as pooling, deformation point association, or token merging on the query token Q. Regardless of which method is used to obtain the proxy token A, this application sets the number of proxy tokens A n as a hyperparameter.
[0103] In a preferred embodiment, the attention module in the object detection model to be trained is replaced with the proxy attention module, and the replaced object detection model is trained with the first training sample to obtain the target object detection model.
[0104] The object detection model to be trained includes: RetinaNet model, Mask R-CNN model, or CascadeMask R-CNN model.
[0105] In this embodiment, the specific operation of applying the proxy attention module mentioned in this application to the object detection model is as follows: the proxy attention module replaces the original attention module in the object detection model to be trained. Then, the replaced object detection model is trained using the first training sample. During the training process, the model will adjust the weight of the proxy attention module according to the features and labels of the first training sample, thereby gradually improving the performance on the target object detection task. In practical applications, the object detection model to be trained can be a RetinaNet model, a Mask R-CNN model, or a Cascade Mask R-CNN model.
[0106] In this application, by applying the proxy attention module to the aforementioned object detection model, a series of experiments were conducted using 1x and 3x timelines with different detection heads. The results showed that the model exhibited consistent improvements across all configurations. Agent-PVT improved the average accuracy by +3.9 to +4.7 compared to the PVT model, while Agent-Swin improved the average accuracy by +1.5 compared to the Swin model. These significant improvements are attributed to the large receptive field introduced by the proxy attention module, thus demonstrating its effectiveness in high-resolution scenes.
[0107] In a preferred embodiment, the attention module in the segmentation model to be trained is replaced with the proxy attention module, and the replaced segmentation model to be trained is trained using a second training sample to obtain the target segmentation model.
[0108] The segmentation model to be trained includes either the SemanticFPN model or the UpperNet model.
[0109] In this embodiment, the specific operation of applying the surrogate attention module mentioned in this application to the segmentation model is as follows: the attention module in the segmentation model to be trained is replaced with the surrogate attention module, and the replaced segmentation model is trained using a second training sample to finally obtain the target segmentation model. In practical applications, the segmentation model to be trained can be a SemanticFPN model or an UpperNet model.
[0110] In this application, by applying a surrogate attention module to the aforementioned segmentation model, computational complexity is reduced, making the segmentation model more efficient in performing tasks such as inference and training. Furthermore, while maintaining computational efficiency, the surrogate attention module can better capture global contextual information in the image. This is particularly important for segmentation tasks, as it helps the model better understand the overall structure and environment of the target. Since the surrogate attention module performs well in high-resolution scenes, applying it to the segmentation model can improve the model's performance when processing complex and detailed images. The surrogate attention module, with its large receptive field, helps the model better model long-distance relationships between pixels.
[0111] In a preferred embodiment of this application, the attention module in the original stable diffusion model is replaced with the proxy attention module to obtain the target stable diffusion model;
[0112] The native stable diffusion models include: the Stable Diffusion model or the ToMeSD model.
[0113] In this embodiment, the specific operation of applying the proxy attention module mentioned in this application to the native stable diffusion model is as follows: the attention module in the native stable diffusion model is replaced with the proxy attention module, and then the target stable diffusion model is obtained. In practical applications, the native stable diffusion model can be a Stable Diffusion model or a ToMeSD model.
[0114] In this application, by applying the surrogate attention module to the aforementioned stable diffusion model, the generation process of the stable diffusion model has a faster computation speed and less image generation time due to the lower computational complexity of the surrogate attention module. The more effective integration of global information by the surrogate attention module can improve the performance of the stable diffusion model in the generation task, thereby improving the quality of the generated images or sequences. The introduction of the surrogate attention module can reduce computational overhead, especially for long sequence or high-resolution image generation tasks.
[0115] In a preferred embodiment, such as Figure 3 The diagram illustrates the information processing flow of a proxy attention module according to an embodiment of this application: For an original image, the original image is divided into a 4*4 window. For this 4*4 image, a linear transformation and grouping operation are performed to obtain its corresponding query token Q, key token K, and value token V. Query token Q, key token K, and value token V are all 4*4 matrices. For query token Q, a pooling operation is performed to obtain proxy token A. After pooling, the original query token Q is transformed from a 4*4 matrix into a 2*2 proxy token A. Next, attention is calculated using proxy token A, key token K, and value token V to aggregate key token K and value token V, thus obtaining proxy features (V). A Agent Features, where the agent features are also a 2x2 matrix, introduce agent bias when aggregating agent token A with key token K and value token V. Then, the original query token Q, agent token A, and agent feature V are... A A second attention calculation is performed, and the result of the second attention calculation is added to the depthwise convolution to obtain the final query features, which are 4*4 matrices.
[0116] This application provides a visual task processing method based on proxy attention. The method includes: acquiring an image to be processed; inputting the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result, wherein the visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result; wherein the visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module, the proxy attention module including a first softmax attention module and a second softmax attention module, wherein the first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain a proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) A Attention calculation is performed, where A is obtained by processing Q. Q, K, and V are the triples used by the original attention module for attention calculation, where Q represents the query token, A represents the surrogate token of the query token Q, K represents the key token, and V represents the value token. This application replaces the original attention module of the visual task processing model used in performing visual processing tasks with a surrogate attention module including a first softmax attention module and a second softmax attention module. By using the visual task processing model with the surrogate attention module replaced, visual processing tasks are performed, resulting in not only higher feature representation but also lower computational complexity.
[0117] Based on the same inventive concept, a second aspect of the embodiments of this application provides a visual task processing system based on agent attention, the system as follows: Figure 4 As shown, it includes:
[0118] The acquisition module 201 is used to acquire the image to be processed;
[0119] Input module 202 is used to input the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result, wherein the visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result;
[0120] The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V)A Attention calculation is performed, where A is obtained by processing Q, and Q, K, V are the triples used by the original attention module for attention calculation. Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token.
[0121] Optionally, the system includes: a first computing submodule;
[0122] The first calculation submodule uses the following formula to calculate attention:
[0123] O A =σ(QA) T +B2)σ(AK T +B1)V
[0124] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters.
[0125] Optionally, the system includes: a second computing submodule;
[0126] The second calculation submodule uses the following formula to calculate attention:
[0127] O=σ(QA T +B2)σ(AK T +B1)V+DWC(V)
[0128] Where σ(·) represents the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters; and DWC represents depthwise convolution.
[0129] Optionally, the system includes: a setting submodule;
[0130] The setting submodule is used to set the agent token A as a learnable parameter so that it can be learned during the training of the agent attention module. The number n of the agent token A is set as a hyperparameter, where n is an integer greater than or equal to 1.
[0131] Optionally, the proxy token A is obtained by performing the following operation on the query token Q, wherein the setting submodule includes:
[0132] A subunit is set up to obtain the proxy token A by performing a pooling operation on the query token Q, a transformation point association operation on the query token Q, or a token merging operation on the query token Q.
[0133] Optionally, the system further includes:
[0134] The first replacement submodule is used to replace the attention module in the object detection model to be trained with the proxy attention module, and to train the replaced object detection model to be trained with the first training sample to obtain the target object detection model.
[0135] The object detection model to be trained includes: RetinaNet model, Mask R-CNN model, or CascadeMask R-CNN model.
[0136] Optionally, the system further includes:
[0137] The second replacement submodule is used to replace the attention module in the segmentation model to be trained with the proxy attention module, and to train the replaced segmentation model to be trained using the second training sample to obtain the target segmentation model.
[0138] The segmentation model to be trained includes either the SemanticFPN model or the UpperNet model.
[0139] Optionally, the system further includes:
[0140] The third replacement submodule is used to replace the attention module in the original stable diffusion model with the proxy attention module to obtain the target stable diffusion model;
[0141] The stable diffusion model to be trained includes either the Stable Diffusion model or the ToMeSD model.
[0142] Each embodiment in this specification focuses on the differences from other embodiments. For the same or similar parts between the embodiments, please refer to each other.
[0143] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0144] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0147] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0148] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0149] The above provides a detailed description of a visual task processing method and system based on proxy attention. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A visual task processing method based on agent attention, characterized in that, The method includes: Obtain the image to be processed; The image to be processed is input into a pre-trained visual task processing model to obtain a visual task processing result, which is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result; The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) A Attention calculation is performed, where A is obtained by processing Q, and Q, K, V are the triples used by the original attention module for attention calculation, where Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token. The agent attention module uses at least the following formula for attention calculation: in, B1 and B2 represent the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters.
2. The visual task processing method based on proxy attention according to claim 1, characterized in that, The agent attention module uses the following formula to calculate attention: in, B1 and B2 represent the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters; DWC represents depthwise convolution.
3. The visual task processing method based on proxy attention according to claim 1, characterized in that, The proxy token A is set as a learnable parameter, which is learned during the training of the proxy attention module. The number n of the proxy token A is set as a hyperparameter, where n is an integer greater than or equal to 1.
4. The visual task processing method based on proxy attention according to claim 1, characterized in that, The proxy token A is obtained by performing the following operations on the query token Q: The query token Q is obtained by performing a pooling operation, a transformation point association operation, or a token merging operation on the query token Q.
5. The visual task processing method based on proxy attention according to claim 1, characterized in that, The number of proxy tokens A is less than the number of query tokens Q.
6. The visual task processing method based on proxy attention according to claim 1, characterized in that, The method further includes: The attention module in the object detection model to be trained is replaced with the proxy attention module, and the replaced object detection model is trained using the first training sample to obtain the target object detection model. The object detection model to be trained includes: RetinaNet model, Mask R-CNN model, or Cascade Mask R-CNN model.
7. The visual task processing method based on proxy attention according to claim 1, characterized in that, The method further includes: The attention module in the segmentation model to be trained is replaced with the proxy attention module, and the replaced segmentation model to be trained is trained using the second training sample to obtain the target segmentation model. The segmentation model to be trained includes either the SemanticFPN model or the UpperNet model.
8. The visual task processing method based on proxy attention according to claim 1, characterized in that, The method further includes: The target stable diffusion model is obtained by replacing the attention module in the original stable diffusion model with the proxy attention module. The native stable diffusion models include: the Stable Diffusion model or the ToMeSD model.
9. A visual task processing system based on agent attention, characterized in that, The system includes: The acquisition module is used to acquire the image to be processed; The input module is used to input the image to be processed into a pre-trained visual task processing model to obtain a visual task processing result, wherein the visual task processing result is any one of the following: image classification result, object detection result, semantic segmentation result, and image generation result; The visual task processing model is obtained by replacing the native attention module of the transformer model with a proxy attention module. The proxy attention module includes a first softmax attention module and a second softmax attention module. The first softmax attention module uses a first triple (A,K,V) to perform attention calculation to obtain the proxy feature V. A The second softmax attention module uses the second triplet (Q, A, V) A Attention calculation is performed, where A is obtained by processing Q, and Q, K, V are the triples used by the original attention module for attention calculation, where Q represents the query token, A represents the proxy token of the query token Q, K represents the key token, and V represents the value token. The agent attention module uses at least the following formula for attention calculation: in, B1 and B2 represent the softmax function; B1 and B2 represent surrogate biases, where B1 and B2 are learnable parameters.