Image processing method, image processing device, computer equipment and storage medium

By using multi-scale multi-head self-attention transformation processing technology in image processing, multiple feature recognition of the feature map of the target image is solved, and the problem of low accuracy of image feature recognition in the prior art is achieved, and more efficient image processing is achieved.

CN115131651BActive Publication Date: 2025-05-06CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210550511.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-05-06
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

In the prior art, when recognizing image features, it is easy to cause excessive calculation of high-resolution images, and some features of low-resolution images cannot be correctly recognized, which affects the smooth progress of image tasks.

Method used

By pre-trained feature recognition model groups, multiple multi-scale multi-head self-attention transformation processing is performed on the target feature map of the target image, at least one granularity feature included in the target feature map is identified, and then the target image processing task is performed on the target image based on the feature resolution corresponding to the identified granularity feature.

Benefits of technology

Improve the accuracy of identifying granularity features, ensuring that the target image processing task can be carried out smoothly and meet processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131651B_ABST
    Figure CN115131651B_ABST
Patent Text Reader

Abstract

The present invention provides an image processing method, an image processing device, a computer device and a storage medium. The image processing method comprises: obtaining a target image and extracting a target feature map of the target image. The target feature map is input into a feature recognition model group, and multiple multi-scale multi-head self-attention transformation processes are performed on the target feature map to identify at least one granularity feature included in the target feature map, wherein multi-head self-attention transformation processes of different scales are used to identify the feature map of the target feature map at different feature resolutions. Based on the feature resolution corresponding to at least one granularity feature, a target image processing task is performed on the target image. Through the present invention, by performing multiple multi-scale multi-head self-attention transformation processes on the target feature map of the target image, the accuracy of identifying the granularity features included in the target feature map can be improved, thereby ensuring that the target image processing task can be carried out smoothly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an image processing method, an image processing device, a computer device and a storage medium. Background Art

[0002] Image processing is a technology that uses computer programs to analyze images to achieve the desired results. With the rapid development of computer technology, image processing technology has also developed rapidly. As deep learning models based on self-attention mechanisms have made great achievements in the field of natural language, deep learning models based on self-attention mechanisms have also been gradually applied in the field of image processing technology. A variety of image tasks can be processed through deep neural network models based on the self-attention mechanism architecture.

[0003] In related technologies, when performing image tasks, a deep neural network model based on a deep self-attention network (Transformer) architecture can be used to perform target image processing tasks on the feature map of the target feature map, and then the multi-head self-attention module in the Transformer architecture can be used to perform global feature recognition on the target feature map at a fixed feature resolution to complete the image task. Among them, Transformer is a deep learning model based on the self-attention mechanism.

[0004] However, when this method is used to perform image tasks, since a single-scale multi-head self-attention module is used to perform global feature recognition on the feature map of the target feature map, it is easy to cause excessive computational effort required for high-resolution images, and some features of low-resolution images cannot be correctly recognized, thus affecting the smooth progress of the image task. Summary of the invention

[0005] Therefore, the technical problem to be solved by the present invention is to overcome the defect of low accuracy of recognition results when recognizing image features in the prior art, thereby providing an image processing method, an image processing device, a computer device and a storage medium.

[0006] According to a first aspect, the present invention provides an image processing method, the method comprising:

[0007] Acquire a target image, and extract a target feature map of the target image;

[0008] Inputting the target feature map into a feature recognition model group, performing multiple multi-scale multi-head self-attention transformation processes on the target feature map to identify at least one granularity feature included in the target feature map, wherein the multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions;

[0009] Based on the feature resolution corresponding to the at least one granularity feature, a target image processing task is performed on the target image.

[0010] In this way, a pre-trained feature recognition model group can be used to perform multiple multi-scale multi-head self-attention transformation processing on the target feature map of the target image, thereby improving the accuracy of identifying granular features, and then performing the target image processing task on the target image based on the feature resolution corresponding to the identified granular features, so that the target image processing task can proceed smoothly and meet the processing requirements of the target image processing task.

[0011] In combination with the first aspect, the feature recognition model group includes multiple connected feature recognition models, different feature recognition models are used to identify the granular features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map;

[0012] The step of inputting the target feature map into a feature recognition model group and performing multiple multi-scale multi-head self-attention transformation processes on the target feature map includes:

[0013] Inputting the target feature map into a feature recognition model group, and sequentially performing multiple multi-scale multi-head self-attention transformation processes on the target feature map through the multiple connected feature recognition models;

[0014] Among them, each feature recognition model performs at least one multi-scale multi-head self-attention transformation processing on the target feature map, different scales correspond to different feature resolutions, and the feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution for identifying the target feature map.

[0015] In combination with the first aspect or the first embodiment of the first aspect, in a second embodiment of the first aspect, performing multiple multi-scale multi-head self-attention transformation processes on the target feature map to identify at least one granularity feature included in the target feature map includes:

[0016] Determining a target number of times each feature recognition model in the feature recognition model group performs the multi-scale multi-head self-attention transformation processing on the target feature map according to a correspondence between a preset feature resolution threshold and the number of times the multi-scale multi-head self-attention transformation processing is performed;

[0017] According to the target number, each feature recognition model is controlled to perform a multi-scale multi-head self-attention transformation process on the target feature map each time to identify at least one granularity feature included in the target feature map.

[0018] In combination with the second embodiment of the first aspect, in a third embodiment of the first aspect, controlling each feature recognition model to perform a multi-scale multi-head self-attention transformation process on the target feature map each time according to the target number of times, and identifying at least one granularity feature included in the target feature map, includes:

[0019] Based on the correspondence between the preset feature resolution and the number of attention groups, determine the target number of attention groups corresponding to each target feature resolution in the current feature extraction model;

[0020] Extracting the first feature map of the target feature map at each target scale respectively;

[0021] Group each first feature map according to the corresponding number of target attention groups to obtain a feature map group corresponding to each first feature map;

[0022] Performing at least one multi-head attention transformation on each feature map group to identify granular features included in each first feature map;

[0023] According to the target number of the current feature extraction model, the granularity features included in all the first feature maps are accumulated to obtain the granularity features recognized by the current feature extraction model at the target feature resolution threshold.

[0024] In this way, in each feature recognition model, at least one multi-scale multi-head self-attention transformation process is performed on the target feature map according to the target number of times to identify the granularity features of the target feature map at different feature resolutions. This can effectively avoid or partially avoid the occurrence of fine-grained features being missed, so that both high-granularity features and fine-grained features in the target image can be effectively recognized, thereby helping to improve the accuracy of identifying granular features.

[0025] In combination with the first embodiment of the first aspect, in a fourth embodiment of the first aspect, extracting a first feature map of the target feature map at each target scale includes:

[0026] According to the corresponding relationship between the feature resolution and the recognition granularity, determining the target recognition granularity corresponding to the target feature map under the current target size;

[0027] Extract the target first feature map of the target feature map at the target recognition granularity.

[0028] In combination with the fourth embodiment of the first aspect, in a fifth embodiment of the first aspect, after obtaining the granularity feature identified by the current feature extraction model at the target feature resolution threshold, the method further includes:

[0029] Using a pre-trained feedforward neural network model, a feedforward neural network transformation is performed on the granularity features identified under the target feature resolution threshold to obtain a transformation result;

[0030] The transformation result is accumulated to the target feature map, the target feature map is updated, and the updated target feature map is input into the next feature extraction model, or the next multi-scale multi-head self-attention transformation processing is performed on the updated target feature map.

[0031] In combination with the first aspect, in a sixth embodiment of the first aspect, respectively extracting a first feature map of the target feature map at each target scale includes:

[0032] Obtaining a target output dimension of the current feature extraction model, the current feature extraction model being the first feature extraction model to process the target feature map;

[0033] Decomposing the high dimension and the wide dimension in the target output dimension according to a specified number to obtain a plurality of sub-wide dimensions and a plurality of sub-high dimensions of the same number;

[0034] According to the order in which the sub-width dimensions and the sub-height dimensions are arranged alternately in sequence, the feature arrangement order of the target feature map is reorganized to extract the first feature map at each target scale.

[0035] According to a second aspect, the present invention further provides an image processing device, the device comprising:

[0036] An extraction unit, used to acquire a target image and extract a target feature map of the target image;

[0037] an identification unit, configured to input the target feature map into a feature recognition model group, perform a plurality of multi-scale multi-head self-attention transformation processes on the target feature map, and identify at least one granularity feature included in the target feature map, wherein the multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions;

[0038] A processing unit is used to perform a target image processing task on the target image based on the feature resolution corresponding to the at least one granularity feature.

[0039] In combination with the second aspect, in a first embodiment of the second aspect, the feature recognition model group includes multiple connected feature recognition models, different feature recognition models are used to identify granular features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map;

[0040] The identification unit comprises:

[0041] The recognition subunit is used to input the target feature map into the feature recognition model group, and perform multiple multi-scale multi-head self-attention transformation processes on the target feature map in sequence through the multiple connected feature recognition models;

[0042] Among them, each feature recognition model performs at least one multi-scale multi-head self-attention transformation processing on the target feature map, different scales correspond to different feature resolutions, and the feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution for identifying the target feature map.

[0043] In combination with the second aspect or the first embodiment of the second aspect, in a second embodiment of the second aspect, the identification unit includes:

[0044] A first determining unit is used to determine a target number of times each feature recognition model in the feature recognition model group performs a multi-scale multi-head self-attention transformation process on the target feature map according to a correspondence between a preset feature resolution threshold and a number of times the multi-scale multi-head self-attention transformation process is performed;

[0045] A processing unit is used to control each feature recognition model to perform a multi-scale multi-head self-attention transformation process on the target feature map according to the target number of times, so as to identify at least one granularity feature included in the target feature map.

[0046] In combination with the second embodiment of the second aspect, in a third embodiment of the second aspect, the processing unit includes:

[0047] A second determination unit is used to determine the target number of attention groups corresponding to each target feature resolution in the current feature extraction model based on a preset correspondence between feature resolution and attention group number;

[0048] An extraction unit, used to extract a first feature map of the target feature map at each target scale respectively;

[0049] A division unit, used for grouping each first feature map according to the corresponding number of target attention groups, to obtain a feature map group corresponding to each first feature map;

[0050] A feature recognition unit, configured to perform at least one multi-head attention transformation on each feature map group to recognize a granular feature included in each first feature map;

[0051] The accumulation unit is used to accumulate the granularity features included in all the first feature maps according to the target number of times of the current feature extraction model to obtain the granularity features recognized by the current feature extraction model under the target feature resolution threshold.

[0052] In combination with the third embodiment of the second aspect, in a fourth embodiment of the second aspect, the extraction unit includes:

[0053] A third determining unit is used to determine the target recognition granularity corresponding to the target feature map under the current target size according to the corresponding relationship between the feature resolution and the recognition granularity;

[0054] The recognition unit is used to extract the target first feature map of the target feature map at the target recognition granularity.

[0055] In combination with the third embodiment of the second aspect, in a fifth embodiment of the second aspect, the device further includes:

[0056] A transformation unit, used to perform a feedforward neural network transformation on the granularity feature identified under the target feature resolution threshold through a pre-trained feedforward neural network model to obtain a transformation result;

[0057] An updating unit is used to accumulate the transformation result to the target feature map, update the target feature map, and input the updated target feature map into the next feature extraction model, or perform the next multi-scale multi-head self-attention transformation processing on the updated target feature map.

[0058] In combination with the third embodiment of the second aspect, in a sixth embodiment of the second aspect, the extraction unit includes:

[0059] An acquisition unit, configured to acquire a target output dimension of the current feature extraction model, wherein the current feature extraction model is a first feature extraction model that processes the target feature map;

[0060] A decomposition unit, used to decompose the high dimension and the wide dimension in the target output dimension according to a specified number to obtain a plurality of sub-wide dimensions and a plurality of sub-high dimensions of the same number;

[0061] The extraction subunit is used to reorganize the feature arrangement order of the target feature map according to the order in which the sub-width dimension and the sub-height dimension are arranged alternately in sequence, so as to extract the first feature map at each target scale.

[0062] According to the third aspect, an embodiment of the present invention further provides a computer device, comprising a memory and a processor, wherein the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the image processing method of the first aspect and any one of its optional embodiments by executing the computer instructions.

[0063] According to a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the image processing method of the first aspect and any one of its optional embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0065] Figure 1 is a flowchart of an image processing method proposed according to an exemplary embodiment.

[0066] Figure 2 It is a schematic diagram of a model structure proposed according to an exemplary embodiment.

[0067] Figure 3 The present invention is a flowchart of a method for identifying image features according to an exemplary embodiment.

[0068] Figure 4 It is a schematic diagram of another model structure proposed according to an exemplary embodiment.

[0069] Figure 5 It is a schematic diagram of a process of recombining features according to an exemplary embodiment.

[0070] Figure 6 It is a structural block diagram of an image processing device proposed according to an exemplary embodiment.

[0071] Figure 7 It is a schematic diagram of the hardware structure of a computer device proposed according to an exemplary embodiment. DETAILED DESCRIPTION

[0072] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0073] In the related technology, when a deep neural network model based on the Transformer architecture performs a target image processing task on a target feature map, the feature map of the target feature map is transformed through a fully connected layer to obtain three feature maps with the same output dimension, and then the feature map of the target feature map is recognized through the multi-head self-attention module in the Transformer architecture.

[0074] However, when this method is used for recognition, a single-scale multi-head self-attention module is used to perform global feature recognition on the feature map of the target feature map. If the feature size corresponding to some features of the target feature map corresponding to the image is lower than the feature size corresponding to the single-scale, it is easy for these features to fail to be correctly recognized, thereby affecting the accuracy of feature recognition. If the resolution of the target feature map is too high, the amount of calculation will increase at a square rate, thereby affecting the smooth progress of the image task.

[0075] To solve the above problems, an image processing method is provided in an embodiment of the present invention, which is used in a computer device. It should be noted that its execution subject can be an image processing device, which can be implemented as part or all of the computer device through software, hardware, or a combination of software and hardware. The computer device can be a terminal, a client, or a server. The server can be a single server or a server cluster composed of multiple servers. The terminal in the embodiment of the present application can be a smart phone, a personal computer, a tablet computer, a wearable device, and other intelligent hardware devices such as an intelligent robot. In the following method embodiments, the execution subject is taken as an example for explanation.

[0076] The computer device in this embodiment includes a pre-trained feature recognition model group for identifying image features to be processed. The image processing method provided by the present invention can perform multiple multi-scale multi-head self-attention transformation processes on the target feature map of the target image through the pre-trained feature recognition model group to identify at least one granularity feature included in the target feature map to improve the accuracy of identifying granularity features, and then when performing the target image processing task on the target image based on the feature resolution corresponding to the granularity feature, the target image processing task can be ensured to proceed smoothly.

[0077] Figure 1 FIG. 1 is a flowchart of an image processing method according to an exemplary embodiment. Figure 1 As shown, the image processing method includes the following steps S101 to S103.

[0078] In step S101, a target image is acquired, and a target feature map of the target image is extracted.

[0079] In an embodiment of the present invention, the target image is an image for which a target image processing task is to be performed. The target image processing task is an image processing task that a user needs to perform, such as high-definition monitoring, power transmission inspection defect detection, remote sensing image classification, etc., which are not limited in the present invention. In one example, the target image can be obtained through the cloud or a local database.

[0080] By performing feature transformation processing on the target image, a feature map of the target image is obtained, wherein the feature points in the feature map can be referred to as image blocks. In one example, the target image can be subjected to feature transformation processing by a pre-trained feature transformation processing model, thereby obtaining a target feature map of the target image. The feature vector corresponding to each image block is accumulated with a vector of dimension C that can represent its position, thereby obtaining a target feature map of the target image. In one example, the vector representing the position can be multiplied by a set of pre-set values ​​through the index of the position, and then the obtained products are concatenated after sine (sin) and cosine (cos) transformations are performed respectively.

[0081] In one implementation scenario, a target image with a dimension of 3*H*W is input into a feature transformation processing model, and the target image is subjected to feature transformation through the convolution layer and transposition layer of the feature transformation processing model, thereby obtaining a feature map with a dimension of (H / p)*(W / p)*C, and then through the position embedding model, the feature vector corresponding to each image block is accumulated with a vector with a dimension of C that can represent its position, thereby obtaining a target feature map of the target image. Among them, H represents the high dimension, W represents the wide dimension, p represents the step size of the convolution layer, and C represents the number of feature channels of the output dimension of the feature transformation processing model. In one example, the target feature map of the target image can be extracted through the feature transformation processing model (patch embedding layer) in the ViT-Base model (a deep neural network model based on the Transformer architecture). Among them, the feature transformation processing model is composed of a convolution layer and a transposition layer. The size of the convolution kernel in the convolution layer is p*p, and the corresponding values ​​of H, W, C and p can be: H=W=224, C=768, p=16.

[0082] In step S102, the target feature map is input into the feature recognition model group, and multiple multi-scale multi-head self-attention transformation processes are performed on the target feature map to identify at least one granularity feature included in the target feature map.

[0083] In an embodiment of the present invention, the feature recognition model group includes a plurality of pre-trained feature recognition models for performing multiple feature recognitions on a target feature map. The target feature map is input into the feature recognition model group, and multiple feature recognition models in the feature recognition model group perform multiple multi-scale multi-head self-attention transformation processes on the target feature map to identify at least one granular feature included in the target feature map. The multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions.

[0084] In one embodiment, the plurality of feature recognition models in the feature recognition model group are in a connected positional relationship. That is, the output end of the previous feature recognition model is connected to the input end of the next feature recognition model, and the output of the previous feature recognition model is the input of the next feature recognition model. In addition, different feature recognition models are used to identify the granularity features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map, and different scales correspond to different feature resolutions. The feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution for identifying the target feature map. The target feature map is input into the feature recognition model group, and through multiple connected feature recognition models, the target feature map is sequentially subjected to multiple multi-scale multi-head self-attention transformation processes to identify the granularity features of the target feature map at different feature resolutions, so as to identify at least one granularity feature included in the target feature map, thereby improving the recognition accuracy.

[0085] In step S103, a target image processing task is performed on the target image based on the feature resolution corresponding to at least one granularity feature.

[0086] In an embodiment of the present invention, based on the feature resolution corresponding to each identified granular feature, an image processing task suitable for execution on a target image is determined; when the image processing task suitable for execution on a target image includes a target image task, the target image processing task is executed on the target image.

[0087] Through the above embodiments, it is possible to perform multiple multi-scale multi-head self-attention transformation processing on the target feature map of the target image through a pre-trained feature recognition model group, thereby improving the accuracy of identifying granular features, and then performing the target image processing task on the target image based on the feature resolution corresponding to the identified granular features, so that the target image processing task can proceed smoothly and meet the processing requirements of the target image processing task.

[0088] In one implementation scenario, the arrangement of multiple feature recognition models in the feature recognition model group can be as follows: Figure 2 As shown, for the convenience of description, an example is given in which the feature recognition model group includes three feature recognition models. Figure 2It is a schematic diagram of a model structure proposed according to an exemplary embodiment. The target feature map is input into the feature recognition model 1, and at least one multi-scale multi-head self-attention transformation process is performed on the target feature map, and at least one granularity feature under multiple resolution thresholds corresponding to the feature recognition model 1 is identified in the target feature map, and the identified granularity feature is superimposed on the target feature map to form a new target feature map, and the new target feature map is input into the feature recognition model 2, and at least one multi-scale multi-head self-attention transformation process is performed on the new target feature map through the feature recognition model 2, and at least one granularity feature under multiple resolution thresholds corresponding to the feature recognition model 2 is identified in the new target feature map. And so on, until the feature recognition model 3 completes the feature recognition of the target feature map, and obtains at least one granularity feature included in the target feature map.

[0089] In another implementation scenario, the number of feature recognition modules in the feature recognition model group can be pre-specified, and the arrangement order of each feature recognition module in the feature recognition model group depends on its corresponding output dimension. In order to better identify the granular features included in the target feature map and avoid missed recognition or misidentification, when sorting multiple feature recognition models, the feature recognition models with larger output dimensions are arranged in the front position, and the feature recognition models with smaller output dimensions are arranged in the back position, which helps to improve the accuracy of identifying granular features. That is, in the feature recognition model group, the feature recognition model with a larger output dimension is arranged in the front position in the feature recognition model group, and vice versa, the feature recognition model with a smaller output dimension is arranged in the back position in the feature recognition model group.

[0090] In another implementation scenario, the output dimension corresponding to each feature recognition module can be determined using the following formula: k ×(H / p k )×(W / p k ), p k =p / 2 N-k , where k is the order of the current feature recognition module, C is the number of feature channels, H and W are the width and height of the target feature map, p is the feature resolution, and N is the total number of feature recognition modules in the feature recognition module group. For example: Figure 2 The feature recognition model group structure shown in the figure takes the total number of feature recognition modules as 3 and p as 16 as an example. The output dimensions of each feature recognition module are C1×(H)×(W), C1×(H / 4)×(W / 4), C2×(H / 8)×(W / 8), and C3×(H / 16)×(W / 16). k 、H / p k and W / p k The order of precedence is not limited in the present invention.

[0091] In one embodiment, the number of times each feature recognition model performs multi-scale multi-head self-attention transformation processing on the target feature map depends on its corresponding feature resolution threshold. According to the correspondence between the preset feature resolution threshold and the number of times the multi-scale multi-head self-attention transformation processing is performed, the target number of times each feature recognition model in the feature recognition model group performs multi-scale multi-head self-attention transformation processing on the target feature map is determined, and then according to the target number of times, each feature recognition model is controlled to perform each multi-scale multi-head self-attention transformation processing on the target feature map to identify at least one granularity feature included in the target feature map. In one example, the higher the feature resolution threshold corresponding to the feature recognition model, the more corresponding target times, which helps to avoid the occurrence of missed recognition of granularity features, and thus makes the obtained recognition result more accurate.

[0092] The following will specifically describe the process in which each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map to identify at least one granular feature included in the target feature map. Since each feature recognition model in the feature recognition model group performs the same principle of multi-scale multi-head self-attention transformation process on the target feature map according to the target number of times, the following embodiment will take the process in which one of the feature recognition models performs a multi-scale multi-head self-attention transformation process on the target feature map as an example for description.

[0093] Figure 3 FIG. 1 is a flow chart of a method for identifying granularity features according to an exemplary embodiment. Figure 3 As shown, the method for identifying particle size features includes the following steps.

[0094] Figure 3 FIG. 1 is a flow chart of a method for identifying granularity features according to an exemplary embodiment. Figure 3 As shown, the method for identifying particle size features includes the following steps.

[0095] In step S301, based on the corresponding relationship between the preset feature resolution and the number of attention groups, the target number of attention groups corresponding to each target feature resolution in the current feature extraction model is determined.

[0096] In the embodiment of the present invention, the number of attention groups can be understood as the number of times self-attention transformation is required for the multi-head self-attention transformation process at each scale, wherein different scales correspond to different feature resolutions.

[0097] In order to improve the accuracy of identifying target feature maps at different feature resolutions, the correspondence between feature resolution and the number of attention groups is pre-configured, and then the number of target attention groups corresponding to each target feature resolution in the current feature extraction model is determined respectively, so as to determine the number of multi-head self-attention transformation processing at each target feature resolution.

[0098] In step S302, the first feature map of the target feature map at each target scale is extracted respectively.

[0099] In the embodiment of the present invention, in order to improve the accuracy of identifying the granularity features of the target feature map at each target size, the first feature map of the target feature map at each target scale is extracted respectively.

[0100] In one embodiment, the target recognition granularity corresponding to the target recognition feature map at the current target size can be determined based on the corresponding relationship between the feature resolution and the recognition granularity, and then the first target feature map of the target feature map at the target recognition granularity can be extracted based on the target recognition granularity. In one implementation scenario, the corresponding relationship between the feature resolution and the recognition granularity can be an inverse relationship, that is, when the target recognition granularity is L, the corresponding target feature resolution is 1 / L. According to the target recognition granularity, when extracting the first feature map of the target feature map at the target recognition granularity, the target feature map can be divided into L×L square areas, and then multiple feature points are extracted in different areas to obtain the first feature map of the target feature map at the current target scale.

[0101] In one example, any of the following feature extraction methods can be used to extract feature points in each square area: The first method: input the first feature map into C k The first method is to use a fully connected layer with a dimension and a 1-dimensional output to generate the weight of each feature point in the current square area, normalize the weight through a normalization function (softmax), and multiply it with the input feature, and then sum it in the current square area to obtain the feature point of the current square area, so that the obtained feature point can fully express the characteristics of the current square area. The second method: all point features in the current square area are processed by average pooling, and the processed result is used as the feature point of the current square area. The third method: using the identity transformation method, the only feature point contained in the current square area is used as the feature point of the current square area. The specific method for obtaining the feature points in the square area can be determined according to the actual processing requirements, and is not limited in the present invention.

[0102] In step S303, each first feature map is grouped according to the corresponding number of target attention groups to obtain a feature map group corresponding to each first feature map.

[0103] In the embodiment of the present invention, each first feature map is grouped according to the number of target attention groups corresponding to it, thereby obtaining a feature map group corresponding to the first feature map. A ×W A ) is divided into feature maps of the target attention groups, and then a feature map group corresponding to each first feature map is obtained.

[0104] In step S304, at least one multi-head attention transformation is performed on each feature map group to identify the granular features included in each first feature map.

[0105] In an embodiment of the present invention, at least one multi-head attention transformation is performed for each feature map group to identify the association relationship between the feature points included in each feature map group, and then identify the granular features included in each first feature map. In one example, the process of identifying the association relationship between the feature points included in each feature map group can be performed using any recognition method in the prior art, which will not be described in detail here.

[0106] In step S305, according to the target number of the current feature extraction model, the granularity features included in all the first feature maps are accumulated to obtain the granularity features recognized by the current feature extraction model under the target feature resolution threshold.

[0107] In an embodiment of the present invention, the number of times the current feature recognition model performs multi-scale multi-head self-attention transformation processing on the target feature map is determined according to the target number, and the granularity features obtained each time the first feature map is recognized are accumulated to enhance the recognition of the granularity features, thereby obtaining the granularity features recognized by the current feature extraction model under the target feature resolution threshold.

[0108] Through the above embodiments, in each feature recognition model, a multi-scale multi-head self-attention transformation process is performed on the target feature map at least once according to the target number of times to identify the granularity features of the target feature map at different feature resolutions. This can effectively avoid or partially avoid the occurrence of fine-grained features being missed, so that both high-granularity features and fine-grained features in the target image can be effectively recognized, thereby helping to improve the accuracy of identifying granular features.

[0109] In one implementation scenario, the corresponding relationship between feature resolution and the number of attention groups can be determined using the following formula: Among them, L is the recognition granularity which is inversely proportional to the feature resolution, H A and W A Respectively represent the height and width of the attention window, H A =W A , G is the number of attention groups, pk =p / 2 N-k , p is the feature resolution, N is the total number of feature recognition modules in the feature recognition module group, and k is the arrangement position of the current feature recognition module in the feature recognition module group. If L = 1, H = W = 1024, H A =W A , G=3, then the target scales of the multi-head self-attention transformation processing of the target feature map by the current feature recognition module are: (8,8,8,1), (4,8,8,4), (1,16,16,16). If L=1, H=3072, W=4096, p=16, G=3, then the target scales of the multi-head self-attention transformation processing of the target feature map by the current feature recognition module are: (16,12,16,1), (4,16,16,12), (1,16,16,192). If L=1, H=W=256, p=16, G=2, then the target scales of the multi-head self-attention transformation processing of the target feature map by the current feature recognition module are: (2,8,8,1), (1,8,8,4).

[0110] In one embodiment, after obtaining the granular features identified by the current feature extraction model under the target feature resolution threshold, in order to facilitate the execution of the next multi-scale multi-head self-attention transformation process or the multi-scale multi-head self-attention transformation process by the next feature extraction model, the granular features identified under the target feature resolution threshold are subjected to a feedforward neural network (FFN) transformation through a pre-trained feedforward neural network model to obtain a transformation result. The transformation result is accumulated to the target feature map in the form of a residual, and the target feature map is updated. In one implementation scenario, the feedforward neural network model includes: a fully connected layer, an activation function, and a fully connected layer. Among them, the activation function can be GELU (a high-performance neural network activation function).

[0111] In another embodiment, before performing the multi-head attention transformation on the feature map group each time, the input feature map group is normalized by layer norm (LayerNorm) or directional normalization (BatchNorm). The transformation result is regularized (drop_path) and then added to the target feature map output by the previous feature extraction model as the output of the current feature extraction model to obtain an updated target feature map.

[0112] In one implementation scenario, the process of performing a dual-scale multi-head self-attention transformation on the first feature map can be as follows: Figure 4 shown. Figure 4It is another schematic diagram of a model structure proposed according to an exemplary embodiment. After the target feature map is input into the current feature transformation submodule, it is normalized by the layered standard (LayerNorm or LN) to obtain the normalized target feature map. The first feature map of the target feature map at the first scale is extracted. The first feature map at the first scale is grouped according to the number of target attention groups corresponding to the first scale to obtain a first feature map group. If the number of target attention groups corresponding to the first scale is greater than 1, multiple groups of multi-head attention transformation processing are performed on the first feature map group. In order to improve the recognition accuracy, enhance the robustness of recognition, and avoid the occurrence of feature missing, in the process of performing at least one multi-head attention transformation processing on the first feature map group, random depth processing is performed on the first feature map group so that each feature map in the first feature map group can effectively perform feature recognition. The features in the first feature map group are extracted by weighted pooling, and then the granular features included in the first feature map at the first scale are identified according to the extracted features, and the granular features at the first scale are accumulated to the target feature map in the form of residuals to obtain an updated target feature map. The updated target feature map is normalized again by the layering standard. After the normalization, the first feature map at the second scale is extracted. The first feature map at the second scale is grouped according to the number of target attention groups corresponding to the second scale to obtain a second feature map group. If the number of target attention groups corresponding to the second scale is 1, a single-group multi-head attention transformation is performed. The second feature map group is subjected to random depth processing to identify the granular features of the target feature map at the second scale. After the new target feature map is normalized by the layering standard, a feedforward neural network transformation is performed through a feedforward neural network to obtain a transformation result, and the transformation result is accumulated to the updated target feature map in the form of a residual to obtain a new target feature map for the next multi-scale multi-head self-attention transformation, or input into the next feature extraction model for granular feature recognition.

[0113] Among them, the random depth processing of the feature map group includes: for all output feature points in the feature map group, all values ​​are set to 0 with a probability of p, and the values ​​of all feature points are set to the original 1 / (1-p) with a probability of (1-p), thereby expanding the number of feature points to be processed for feature recognition.

[0114] In one embodiment, if the current feature extraction model is the first feature extraction model to process the target feature map, then before respectively extracting the first feature map of the target feature map at each target scale, the target feature map is feature reorganized based on the target output dimension of the current feature recognition model, so that when subsequently extracting feature points in different regions or performing multi-scale multi-head self-attention transformation processing, the operation of reorganizing feature data can be reduced, thereby helping to speed up the recognition process and improve recognition efficiency.

[0115] In the process of reorganization, the high dimension and width dimension in the target output dimension are decomposed according to the specified number to obtain multiple sub-width dimensions and multiple sub-height dimensions of the same number. For example: decompose the W dimension and H dimension into n sub-width dimensions and n sub-height dimensions: w0×w1×w2×…×w n , h0×h1×h2×…×h n In one example, since the target feature map may be a rectangular feature map with different side lengths, in order to ensure that the number of sub-width dimensions and sub-height dimensions is the same, the values ​​corresponding to w0 and h0 are determined according to the difference between the height dimension and the width dimension, and w1=h1=w2=h2…=w n =h n For example, when the difference between the W dimension and the H dimension is 1, the sub-width dimension and sub-height dimension after decomposition are: w0 = 1, h0 = 2, w1 = h1 = w2 = h2 ... = w n =h n = 2. The target feature map is decomposed, and the feature arrangement order of the target feature map is reorganized in the order of alternating sub-width dimensions and sub-height dimensions to obtain a reorganized target feature map, and then the first feature map at each target scale is extracted through the reorganized target feature map.

[0116] In one implementation scenario, the process of reorganizing the target feature map can be as follows: Figure 5 shown. Figure 5 It is a schematic diagram of a process of reorganizing features proposed according to an exemplary embodiment. The left figure is a target feature map before reorganization, and the multiple sub-width dimensions and the multiple sub-height dimensions are: w0=h0=w1=h1=w2=h2…=w8=h8=2. After the target feature map is decomposed according to the multiple sub-width dimensions and the multiple sub-height dimensions, it is reorganized in a row-first manner, and then the target feature map after reorganization in the right figure is obtained.

[0117] In another implementation scenario, a computer device for performing image processing on a target image may include the following models: a feature transformation processing model, a position embedding model, a feature recognition model group, an output unit, and an execution unit.

[0118] Among them, the feature transformation processing model is used to perform feature transformation on the target image, and then obtain a feature map corresponding to the target image.

[0119] The position embedding model is used to accumulate the feature points in the feature map according to their corresponding feature vectors and the vector of dimension C that can represent their positions to obtain the target feature map of the target image.

[0120] The feature recognition model group is used to recognize at least one granularity feature included in the target feature map.

[0121] The output unit is used to output a target feature map superimposed with at least one particle size feature.

[0122] The execution unit is used to execute a target image processing task on the target image based on the feature resolution corresponding to the identified granularity feature.

[0123] Through the above-mentioned embodiment, the target feature map is subjected to multiple multi-scale multi-head self-attention transformation processing through a pre-trained feature recognition model group, and the granular features in the target feature map are identified, which can reduce the amount of calculation and video memory occupied by executing a single-scale multi-head self-attention transformation processing, thereby helping to improve the accuracy of granular feature recognition and reduce the occurrence of missed recognition or misrecognition. In addition, by performing multiple multi-scale multi-head self-attention transformation processing on the target feature map, the problem that a single feature cannot obtain feature information from outside the group during the single-scale grouping autonomous attention transformation process can be solved, thereby helping to reduce the accuracy loss caused by the autonomous attention transformation grouping.

[0124] In another implementation scenario, based on the ImageNet dataset and the OpenImageV4 dataset, the accuracy of performing multi-scale multi-head self-attention transformation on the feature map to identify granular features can be verified, and the following experimental results can be obtained (not limited to the following experiments):

[0125] If the first feature map with H=W=256 and p=16 is used for verification, the output of the dual-scale multi-head self-attention transformation process ((2,8,8,1), (1,8,8,4)) is connected to the output module of the classification task, trained on 200 categories of data in the ImageNet dataset, and verified using the validation set, the classification accuracy reaches 87.2%, which exceeds 86.8% of the ViT model under the same setting under similar computational complexity. It can be seen that the image processing method provided by the present invention can achieve an accuracy not lower than that of the ViT model under low feature resolution.

[0126] If the first feature map of H=W=512 and p=16 is used for verification. The output of the three-scale multi-head self-attention transformation processing ((4,8,8,1), (2,8,8,4), (1,8,8,16)) is connected to the output module of the classification task, and trained by the MAE method on the training set of the OpenImageV4 data set. When the number of pictures per graphics card per iteration is 64, the video memory occupancy is 20.4GB, which is better than the ViT model with the same configuration (it cannot be trained on the ViT model because it exceeds the maximum video memory). It can be seen that the image processing method provided by the present invention can effectively save video memory during training under high feature resolution. Moreover, after 80 rounds of training, the mean square error of the picture pixel prediction on the verification set is 0.713, indicating that the feature extraction network model can effectively encode and predict picture information with higher feature resolution.

[0127] Based on the same inventive concept, the present invention also provides an image processing device.

[0128] Figure 6 FIG. 1 is a structural block diagram of an image processing device according to an exemplary embodiment. Figure 6 As shown, the image processing device includes an extraction unit 601 , a recognition unit 602 and a processing unit 603 .

[0129] The extraction unit 601 is used to obtain a target image and extract a target feature map of the target image;

[0130] The recognition unit 602 is used to input the target feature map into the feature recognition model group, perform multiple multi-scale multi-head self-attention transformation processes on the target feature map, and identify at least one granularity feature included in the target feature map, wherein the multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions;

[0131] The processing unit 603 is used to perform a target image processing task on the target image based on a feature resolution corresponding to at least one granularity feature.

[0132] In one embodiment, the feature recognition model group includes a plurality of connected feature recognition models, and different feature recognition models are used to recognize the granular features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map. The recognition unit 602 includes: a recognition subunit, which is used to input the target feature map into the feature recognition model group, and sequentially perform multiple multi-scale multi-head self-attention transformation processes on the target feature map through multiple connected feature recognition models. Among them, each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map, and different scales correspond to different feature resolutions. The feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution of the target feature map.

[0133] In another embodiment, the identification unit 602 includes: a first determination unit, configured to determine the target number of times each feature recognition model in the feature recognition model group performs multi-scale multi-head self-attention transformation processing on the target feature map according to the correspondence between the preset feature resolution threshold and the number of times the multi-scale multi-head self-attention transformation processing is performed. A processing unit, configured to control each feature recognition model to perform each multi-scale multi-head self-attention transformation processing on the target feature map according to the target number of times, and identify at least one granularity feature included in the target feature map.

[0134] In another embodiment, the processing unit includes: a second determination unit, which is used to determine the target number of attention groups corresponding to each target feature resolution in the current feature extraction model based on the corresponding relationship between the preset feature resolution and the number of attention groups. An extraction unit, which is used to extract the first feature map of the target feature map at each target scale. A division unit, which is used to group each first feature map according to the corresponding target number of attention groups to obtain a feature map group corresponding to each first feature map. A feature identification unit, which is used to perform at least one multi-head attention transformation on each feature map group to identify the granularity features included in each first feature map. An accumulation unit, which is used to accumulate the granularity features included in all the first feature maps according to the target number of the current feature extraction model to obtain the granularity features identified by the current feature extraction model under the target feature resolution threshold.

[0135] In another embodiment, the extraction unit includes: a third determination unit, configured to determine the target recognition granularity corresponding to the target feature map at the current target size according to the corresponding relationship between the feature resolution and the recognition granularity. A recognition unit, configured to extract the target first feature map of the target feature map at the target recognition granularity.

[0136] In another embodiment, the device further includes: a transformation unit, which is used to perform a feedforward neural network transformation on the granularity features identified under the target feature resolution threshold through a pre-trained feedforward neural network model to obtain a transformation result. An updating unit, which is used to accumulate the transformation result to the target feature map, update the target feature map, and input the updated target feature map to the next feature extraction model, or perform the next multi-scale multi-head self-attention transformation processing on the updated target feature map.

[0137] In another embodiment, the extraction unit includes: an acquisition unit for acquiring a target output dimension of a current feature extraction model, the current feature extraction model being the first feature extraction model for processing a target feature map. A decomposition unit for decomposing a high dimension and a wide dimension in the target output dimension according to a specified number to obtain a plurality of sub-wide dimensions and a plurality of sub-high dimensions of the same number. An extraction subunit for reorganizing the feature arrangement order of the target feature map in the order in which the sub-wide dimensions and the sub-high dimensions are arranged alternately in sequence, so as to extract the first feature map at each target scale.

[0138] The specific limitations and beneficial effects of the above-mentioned image processing device can be found in the above-mentioned limitations on the image processing method, which will not be repeated here. The above-mentioned modules can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0139] Figure 7 FIG. 1 is a schematic diagram of a hardware structure of a computer device according to an exemplary embodiment. Figure 7 As shown, the device includes one or more processors 710 and a memory 720, and the memory 720 includes a persistent memory, a volatile memory, and a hard disk. Figure 7 A processor 710 is taken as an example. The device may also include: an input device 730 and an output device 740.

[0140] The processor 710, the memory 720, the input device 730 and the output device 740 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0141] The processor 710 may be a central processing unit (CPU). The processor 710 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips. A general-purpose processor may be a microprocessor or the processor may be any conventional processor.

[0142] The memory 720 is a non-transient computer-readable storage medium, including a persistent memory, a volatile memory, and a hard disk, and can be used to store non-transient software programs, non-transient computer executable programs, and modules, such as program instructions / modules corresponding to the business management method in the embodiment of the present application. The processor 710 executes various functional applications and data processing of the server by running the non-transient software programs, instructions, and modules stored in the memory 720, that is, implementing any of the above-mentioned image processing methods.

[0143] The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data required for use, etc. In addition, the memory 720 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 720 may optionally include a memory remotely arranged relative to the processor 710, and these remote memories may be connected to the data processing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0144] The input device 730 can receive input digital or character information and generate key signal input related to user settings and function control. The output device 740 can include display devices such as a display screen.

[0145] One or more modules are stored in the memory 720, and when executed by one or more processors 710, the execution is as follows: Figure 1-Figure 5 The method shown.

[0146] The above-mentioned product can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not described in detail in this embodiment, please refer to Figure 1-Figure 5 Related description of the illustrated embodiment.

[0147] The embodiment of the present invention also provides a non-transitory computer storage medium, which stores computer executable instructions, and the computer executable instructions can execute the authentication method in any of the above method embodiments. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memory.

[0148] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. An image processing method, characterized in that: The method comprises: Acquire a target image, and extract a target feature map of the target image; Inputting the target feature map into a feature recognition model group, performing multiple multi-scale multi-head self-attention transformation processes on the target feature map to identify at least one granularity feature included in the target feature map, wherein multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions; performing multiple multi-scale multi-head self-attention transformation processes on the target feature map to identify at least one granularity feature included in the target feature map includes: Determining a target number of times each feature recognition model in the feature recognition model group performs the multi-scale multi-head self-attention transformation processing on the target feature map according to a correspondence between a preset feature resolution threshold and the number of times the multi-scale multi-head self-attention transformation processing is performed; According to the target number, controlling each feature recognition model to perform a multi-scale multi-head self-attention transformation process on the target feature map each time to identify at least one granularity feature included in the target feature map; The method of controlling each feature recognition model to perform a multi-scale multi-head self-attention transformation process on the target feature map according to the target number of times, and identifying at least one granularity feature included in the target feature map, includes: Based on the correspondence between the preset feature resolution and the number of attention groups, determine the target number of attention groups corresponding to each target feature resolution in the current feature extraction model; Extracting the first feature map of the target feature map at each target scale respectively; Group each first feature map according to the corresponding number of target attention groups to obtain a feature map group corresponding to each first feature map; Performing at least one multi-head attention transformation on each feature map group to identify granular features included in each first feature map; According to the target number of the current feature extraction model, the granularity features included in all the first feature maps are accumulated to obtain the granularity features recognized by the current feature extraction model at the target feature resolution threshold; Based on the feature resolution corresponding to the at least one granularity feature, a target image processing task is performed on the target image.

2. The method according to claim 1, characterized in that The feature recognition model group includes a plurality of connected feature recognition models, and different feature recognition models are used to recognize the granular features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map; The step of inputting the target feature map into a feature recognition model group and performing multiple multi-scale multi-head self-attention transformation processes on the target feature map includes: Inputting the target feature map into a feature recognition model group, and sequentially performing multiple multi-scale multi-head self-attention transformation processes on the target feature map through the multiple connected feature recognition models; Among them, each feature recognition model performs at least one multi-scale multi-head self-attention transformation processing on the target feature map, different scales correspond to different feature resolutions, and the feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution for identifying the target feature map.

3. The method according to claim 1, characterized in that Extracting a first feature map of the target feature map at each target scale includes: According to the corresponding relationship between the feature resolution and the recognition granularity, determining the target recognition granularity corresponding to the target feature map under the current target size; A first feature map of the target feature map at the target recognition granularity is extracted.

4. The method according to claim 1, characterized in that After obtaining the granularity feature identified by the current feature extraction model at the target feature resolution threshold, the method further includes: Using a pre-trained feedforward neural network model, a feedforward neural network transformation is performed on the granularity features identified under the target feature resolution threshold to obtain a transformation result; The transformation result is accumulated to the target feature map, the target feature map is updated, and the updated target feature map is input into the next feature extraction model, or the next multi-scale multi-head self-attention transformation processing is performed on the updated target feature map.

5. The method according to claim 1, characterized in that The step of respectively extracting a first feature map of the target feature map at each target scale includes: Obtaining a target output dimension of the current feature extraction model, the current feature extraction model being the first feature extraction model to process the target feature map; Decomposing the high dimension and the wide dimension in the target output dimension according to a specified number to obtain a plurality of sub-wide dimensions and a plurality of sub-high dimensions of the same number; According to the order in which the sub-width dimensions and the sub-height dimensions are arranged alternately in sequence, the feature arrangement order of the target feature map is reorganized to extract the first feature map at each target scale.

6. An image processing device, characterized in that: The device comprises: An extraction unit, used to acquire a target image and extract a target feature map of the target image; an identification unit, configured to input the target feature map into a feature recognition model group, perform a plurality of multi-scale multi-head self-attention transformation processes on the target feature map, and identify at least one granularity feature included in the target feature map, wherein the multi-head self-attention transformation processes of different scales are used to identify feature maps of the target feature map at different feature resolutions; The identification unit comprises: A first determining unit is used to determine a target number of times each feature recognition model in the feature recognition model group performs a multi-scale multi-head self-attention transformation process on the target feature map according to a correspondence between a preset feature resolution threshold and a number of times the multi-scale multi-head self-attention transformation process is performed; A processing unit, configured to control each feature recognition model to perform a multi-scale multi-head self-attention transformation process on the target feature map according to the target number of times, and identify at least one granularity feature included in the target feature map; The processing unit comprises: A second determination unit is used to determine the target number of attention groups corresponding to each target feature resolution in the current feature extraction model based on a preset correspondence between feature resolution and attention group number; An extraction unit, used to extract a first feature map of the target feature map at each target scale respectively; A division unit, used for grouping each first feature map according to the corresponding number of target attention groups, to obtain a feature map group corresponding to each first feature map; A feature recognition unit, configured to perform at least one multi-head attention transformation on each feature map group to recognize a granular feature included in each first feature map; An accumulation unit, configured to accumulate the granularity features included in all the first feature maps according to the target number of times of the current feature extraction model, to obtain the granularity features recognized by the current feature extraction model at the target feature resolution threshold; A processing unit is used to perform a target image processing task on the target image based on the feature resolution corresponding to the at least one granularity feature.

7. The device according to claim 6, characterized in that The feature recognition model group includes a plurality of connected feature recognition models, and different feature recognition models are used to recognize the granular features of the target feature map at different feature resolution thresholds, and each feature recognition model performs at least one multi-scale multi-head self-attention transformation process on the target feature map; The identification unit comprises: The recognition subunit is used to input the target feature map into the feature recognition model group, and perform multiple multi-scale multi-head self-attention transformation processes on the target feature map in sequence through the multiple connected feature recognition models; Among them, each feature recognition model performs at least one multi-scale multi-head self-attention transformation processing on the target feature map, different scales correspond to different feature resolutions, and the feature resolution threshold corresponding to each feature recognition model is the maximum feature resolution for identifying the target feature map.

8. The device according to claim 6, characterized in that The extraction unit comprises: A third determining unit is used to determine the target recognition granularity corresponding to the target feature map under the current target size according to the corresponding relationship between the feature resolution and the recognition granularity; The recognition unit is used to extract a first feature map of the target feature map at the target recognition granularity.

9. The device according to claim 6, characterized in that The device also includes: A transformation unit, used to perform a feedforward neural network transformation on the granularity feature identified under the target feature resolution threshold through a pre-trained feedforward neural network model to obtain a transformation result; An updating unit is used to accumulate the transformation result to the target feature map, update the target feature map, and input the updated target feature map into the next feature extraction model, or perform the next multi-scale multi-head self-attention transformation processing on the updated target feature map.

10. The device according to claim 6, characterized in that The extraction unit comprises: An acquisition unit, configured to acquire a target output dimension of the current feature extraction model, wherein the current feature extraction model is a first feature extraction model that processes the target feature map; A decomposition unit, used to decompose the high dimension and the wide dimension in the target output dimension according to a specified number to obtain a plurality of sub-wide dimensions and a plurality of sub-high dimensions of the same number; The extraction subunit is used to reorganize the feature arrangement order of the target feature map according to the order in which the sub-width dimension and the sub-height dimension are arranged alternately in sequence, so as to extract the first feature map at each target scale.

11. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image processing method according to any one of claims 1 to 5 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the image processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image deblurring method and device based on artificial intelligence, equipment and storage medium

    CN112561826A