A method, device, computer equipment and storage medium for generating attention image

By dividing and sorting the feature areas of the original image feature map and generating an attention image transfer sequence, the problem of large computational cost in generating attention images is solved and the efficiency and accuracy are improved.

CN114022488BActive Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111045879.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-07
Publication Date
2025-10-03
Estimated Expiration
2041-09-07

AI Technical Summary

Technical Problem

The process of generating attention images is computationally intensive, time-consuming, and inefficient.

Method used

The feature map of the original image is divided into feature areas to obtain N feature matrices, which are sorted from small to large in matrix scale to generate an attention image transfer sequence. The attention image corresponding to each feature matrix is ​​generated through the preset attention image generation strategy.

Benefits of technology

The amount of matrix calculation is reduced, the efficiency of generating attention images is improved, and the accuracy of feature extraction and the quality of image generation are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022488B_ABST
    Figure CN114022488B_ABST
Patent Text Reader

Abstract

The present application relates to the fields of artificial intelligence and digital medicine, and provides a method for generating an attention image. The method divides a feature map obtained by extracting features from an original map using a neural network model into feature regions, and obtains a feature matrix each time the division is performed. After N divisions, N feature matrices are obtained; wherein N is an integer greater than 1; the N feature matrices are sorted in order of matrix scale from small to large to obtain an attention image transfer sequence; according to a preset attention image generation strategy, an attention image corresponding to each feature matrix is ​​generated based on the attention image transfer sequence, and finally, the attention image corresponding to the Nth feature matrix is ​​obtained through the attention image generation strategy. This method predicts the correlation of the corresponding area of ​​the fine-grained feature map based on the area with smaller autocorrelation of the coarse-grained feature map, thereby reducing the amount of matrix calculation and improving the efficiency of generating attention images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, apparatus, computer equipment, and storage medium for generating an attention image. Background Art

[0002] With the continuous development of artificial intelligence, a large number of semantic segmentation networks have emerged. Traditional semantic segmentation networks mainly simulate long-range dependencies by stacking multiple convolutions, but the dependencies obtained are weak. To obtain dense pixel-level correlations, non-local networks use the self-attention mechanism to generate attention images, allowing a single feature at any position to perceive the features at all other positions, which can produce more powerful pixel-level representation capabilities. The self-attention mechanism can capture the spatial dependency between any two positions in the feature map and obtain dependency information between long-range pixels. However, in the process of generating attention images, calculating the correlation between each pixel often requires matrix multiplication of large matrices, which often requires a large amount of computation and is time-consuming, resulting in low efficiency. Summary of the Invention

[0003] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for generating attention images in response to the above technical problems, so as to solve the problems of large amount of calculation, long time consumption and low efficiency in the process of generating attention images.

[0004] A first aspect of an embodiment of the present application provides a method for generating an attention image, comprising:

[0005] Perform feature area division on the feature map obtained by extracting features from the original map using the neural network model, obtaining a feature matrix each time the map is divided, and N feature matrices are obtained after N divisions, where N is an integer greater than 1;

[0006] Sort the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence;

[0007] According to the preset attention image generation strategy, the attention image corresponding to each feature matrix is ​​generated based on the attention image transfer sequence, and through the attention image generation strategy, the attention image corresponding to the Nth feature matrix is ​​finally obtained.

[0008] In the above scheme, the feature map obtained by extracting features from the original map using the neural network model is divided into feature regions, and each division obtains a feature matrix. After N divisions, the N feature matrices obtained include:

[0009] The feature map obtained by extracting features from the original map using a neural network model is divided into feature regions, and a feature matrix is ​​obtained each time the feature map is divided. After N divisions, N feature matrices are obtained. Each time the feature map is divided, the feature map is divided into regions of equal size.

[0010] In the above scheme, the feature map obtained by extracting features from the original map using the neural network model is divided into feature regions, and each division obtains a feature matrix. After N divisions, the N feature matrices obtained include:

[0011] The feature map obtained by extracting features from the original map using a neural network model is divided into feature regions. A feature matrix is ​​obtained each time the feature map is divided. After N divisions, N feature matrices are obtained, where each element value in the first feature matrix of the N feature matrices is the mean of the eigenvalues ​​of the feature region.

[0012] In the above scheme, the preset attention image generation strategy includes:

[0013] According to the i-th characteristic matrix, obtain the i+1-th characteristic matrix; where i is an integer greater than 1 and less than N;

[0014] According to the first threshold, obtaining the area with larger autocorrelation in the attention image of the i-th feature matrix;

[0015] According to the (i+1)th feature matrix, obtaining the self-attention value of the area with larger autocorrelation in the attention image of the (i+1)th feature matrix;

[0016] In the above solution, the preset attention image generation strategy also includes:

[0017] According to the first threshold, obtaining an area with smaller autocorrelation in the attention image of the i-th feature matrix;

[0018] According to the area with smaller autocorrelation in the attention image of the i-th feature matrix, the self-attention value of the corresponding area in the attention image of the i+1-th feature matrix is ​​predicted.

[0019] In the above solution, obtaining the self-attention value of the area with larger autocorrelation in the attention image of the i+1th feature matrix according to the i+1th feature matrix includes:

[0020] The self-attention value of the area with larger autocorrelation in the attention image of the i+1th feature matrix is ​​obtained by the dot product of the i+1th feature matrix.

[0021] In the above solution, the method further includes:

[0022] Normalizing the Nth attention image corresponding to the N feature matrices to obtain a normalized image;

[0023] Obtaining an h feature map of the feature map by performing a linear transformation on the feature map;

[0024] A self-attention feature map is obtained according to the normalized image and the h feature map.

[0025] A second aspect of an embodiment of the present application provides a device for generating an attention image, including:

[0026] Division unit: divides the feature map obtained by extracting features from the original map using the neural network model into feature regions to obtain N feature matrices; where N is an integer greater than 1;

[0027] Sorting unit: sorting the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence;

[0028] Processing unit: generates an attention image corresponding to each of the feature matrices based on the attention image transfer sequence according to a preset attention image generation strategy.

[0029] A third aspect of an embodiment of the present application provides a computer device comprising: a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, the computer instructions being used to enable the computer to execute the steps of a method for generating an attention image.

[0030] Implementing the method for generating an attention image provided by the embodiment of the present application has the following beneficial effects:

[0031] An embodiment of the present application provides a method for generating an attention image. The method involves dividing the feature map of an original image into feature regions to obtain N feature matrices, where N is an integer greater than 1. The N feature matrices are sorted in ascending order of matrix scale to obtain an attention image transfer sequence. An attention image corresponding to each feature matrix is ​​generated based on the attention image transfer sequence according to a preset attention image generation strategy. This method predicts the correlation of corresponding regions in a fine-grained feature map based on regions with low autocorrelation in the coarse-grained feature map, thereby reducing the amount of matrix computation required and improving the efficiency of generating the attention image. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0033] Figure 1 This is a flowchart of a method for generating an attention image in one embodiment of the present application;

[0034] Figure 2 is a flowchart of a method for generating an attention image in another embodiment of the present application;

[0035] Figure 3 This is a structural block diagram of a method and apparatus for generating an attention image provided by an embodiment of the present application;

[0036] Figure 4 This is a structural block diagram of a server-side device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0038] The method for generating an attention image provided in this embodiment can be executed by the server.

[0039] The method for generating an attention image involved in this application is applied to the field of artificial intelligence, so that when generating feature attention images, the amount of calculation can be reduced and the efficiency can be improved.

[0040] See also Figure 1 , Figure 1 A flowchart of an implementation method for generating an attention image provided in an embodiment of the present application is shown.

[0041] like Figure 1 As shown, a method for generating an attention image includes:

[0042] S11: performing feature area division on the feature map obtained by extracting features from the original map using the neural network model, obtaining a feature matrix for each division, and performing N divisions to obtain N feature matrices; wherein N is an integer greater than 1.

[0043] In step S11, the neural network model is used to extract the features of the original image to obtain a feature map, which is then divided into regions. Each region is represented by an eigenvalue, and each region division can obtain a feature matrix composed of eigenvalues. The feature map is divided into regions N times to obtain N feature matrices, wherein each region division is a region reduction division based on the previous region division, so the feature matrix is ​​a matrix that increases successively.

[0044] In this embodiment, using a tumor image as an example, a convolutional neural network is used to extract features from the original image. Feature values ​​of the original image pixels are obtained through training. Initial region division is performed based on the feature regions in the original image, resulting in an initial matrix for the feature map. The size of the initial matrix can be customized based on the location of the features in the original image, ensuring that the extracted features are located in at least one region. In this embodiment, the initial region division is of 2×1 dimensions, resulting in a 2×1 feature matrix. Region division is then performed sequentially. In this embodiment, the feature map is divided four times, ultimately dividing it into 8×8 dimensions, resulting in an 8×8 feature matrix. During region division, the size of each region varies depending on the dimensions of the resulting feature map. If the feature map is 2480×2480 pixels, dividing it into 2×1 dimensions results in a region of 1240×2480 pixels per dimension. Dividing it into 2×2 dimensions results in a region of 1240×1240 pixels per dimension. As the number of divisions increases, the size of each region decreases.

[0045] Here, the feature map of the original image is obtained by a convolutional neural network. The convolutional neural network is obtained through training. The training network in this embodiment includes an input layer and an output layer, a convolutional layer, a fully connected layer, a pooling layer and an activation function. The convolutional neural network usually uses a convolutional layer and a pooling layer as feature extraction units to extract target features in the original image and obtain a feature map.

[0046] As an embodiment of the present application, step S11 specifically includes:

[0047] The feature area division of the feature map obtained by extracting features from the original image using the neural network model, obtaining a feature matrix each time the division is performed, and N feature matrices are obtained by dividing the feature map N times, including: the feature area division of the feature map obtained by extracting features from the original image using the neural network model, obtaining a feature matrix each time the division is performed, and N feature matrices are obtained by dividing the feature map N times, wherein each time the area is divided, the feature map is divided into areas of equal size. The feature area division of the feature map obtained by extracting features from the original image using the neural network model, obtaining a feature matrix each time the division is performed, and N feature matrices are obtained by dividing the feature map N times, including: the feature area division of the feature map obtained by extracting features from the original image using the neural network model, obtaining a feature matrix each time the division is performed, and N feature matrices are obtained, wherein each element value in the first feature matrix of the N feature matrices is the mean of the eigenvalues ​​of the feature area in which it is located.

[0048] In this embodiment, taking a tumor image in a medical image as an example, based on the feature map obtained from the original image, the feature map is initially divided into 2×1 regions. The feature matrix obtained after the region division is a 2×1 feature matrix. During the region division, each region is of equal size. If the region division cannot guarantee the equal size of each region, the feature map needs to be supplemented by pixels with a eigenvalue of 0 to make each region of equal size. For example, when the feature map is 2480×2480 pixels in size, the feature map is divided into 2×1 feature regions, and the size of each region is 1240×1240 pixels.

[0049] Here, when obtaining the feature matrix, each element value in the first feature matrix is ​​obtained through the pixel eigenvalue, and the element value in the feature matrix is ​​the mean of the eigenvalues ​​of the corresponding area.

[0050] S12: Sort the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence;

[0051] In step S12, an attention image corresponding to the feature matrix is ​​obtained according to the feature matrix of the feature map, wherein the attention image obtains N feature matrices according to the elements in the feature matrix, and generates N attention images, and the sequence of the attention images corresponds to the sequence of the feature matrices.

[0052] In this embodiment, taking the tumor image in the medical image as an example, according to the features in the original image, the feature map is divided into 2×1 areas, and the initial feature matrix is ​​a 2×1 matrix. The feature map is finally divided into 8×8 areas, and finally four feature matrices and four attention images are obtained. They are sorted from small to large according to the scale of the feature matrix, which are 2×1, 2×2, 4×4, and 8×8 matrices, respectively. Among them, the area corresponding to each element in the small-scale feature matrix in the feature map and the area corresponding to each element in the large-scale feature matrix in the feature map have a corresponding relationship with each other.

[0053] Here, the matrix feature map is divided into regions to obtain feature matrices. The size of the feature matrix starts from 2×1, and the corresponding attention image N1 is obtained. The second feature matrix is ​​a 2×2 matrix, and the corresponding attention image N2 is obtained. The third feature matrix is ​​a 4×4 matrix, and the corresponding attention image N3 is obtained. The fourth feature matrix is ​​an 8×8 matrix, and the corresponding attention image N4 is obtained. The attention image transfer sequence is N1 to N2, N2 to N3, and N3 to N4.

[0054] S13: According to the preset attention image generation strategy, the attention image corresponding to each feature matrix is ​​generated based on the attention image transfer sequence, and finally the attention image corresponding to the Nth feature matrix is ​​obtained through the attention image generation strategy.

[0055] In step S13, an attention image is generated based on the feature matrix. The attention image corresponding to the small-scale feature matrix is ​​then used to generate the attention image corresponding to the large-scale feature matrix. The attention image reflects the correlation between the regions corresponding to each element in the feature matrix. Generating an attention image that reflects the correlation between each image region and the target region can improve feature extraction accuracy. Directly calculating the relationship between any two regions in the feature map allows for the acquisition of the image's global geometric features in a single step.

[0056] In this embodiment, taking a tumor image from a medical image as an example, the attention image corresponding to each feature matrix is ​​generated according to the attention image transfer sequence. The attention image corresponding to the final feature matrix is ​​calculated sequentially based on the small-scale feature matrix. In this embodiment, the corresponding attention image N1 is calculated based on the feature matrix being a 2×1 matrix. Based on the obtained attention image N1 and the 2×2 feature matrix, the attention image N2 corresponding to the 2×2 feature matrix is ​​obtained again. Finally, the attention image N4 corresponding to the fourth 8×8 feature matrix is ​​obtained in sequence.

[0057] As an embodiment of the present application, step S13 specifically includes:

[0058] According to the i-th feature matrix, obtain the i+1-th feature matrix; wherein i is an integer greater than 1 and less than N; according to the first threshold, obtain the region with larger autocorrelation in the attention image of the i-th feature matrix; according to the i+1-th feature matrix, obtain the self-attention value of the region with larger autocorrelation in the attention image of the i+1-th feature matrix. According to the first threshold, obtain the region with smaller autocorrelation in the attention image of the i-th feature matrix; according to the region with smaller autocorrelation in the attention image of the i-th feature matrix, predict the self-attention value of the corresponding region in the attention image of the i+1-th feature matrix; according to the i+1-th feature matrix, obtain the self-attention value of the region with larger autocorrelation in the attention image of the i+1-th feature matrix, including: obtaining the self-attention value of the region with larger autocorrelation in the attention image of the i+1-th feature matrix by performing a dot product of the i+1-th feature matrix.

[0059] In this embodiment, the second feature matrix is ​​obtained according to the first feature matrix, the fourth feature matrix is ​​obtained in sequence, and then the corresponding attention image is obtained according to the fourth feature matrix. For example, a 2×2 feature matrix is ​​obtained according to a 2×1 feature matrix, wherein the process of obtaining the 2×2 feature matrix is ​​as follows: the feature matrix is ​​a 2×1 size matrix, and an attention image of size 2×2 is obtained by multiplying the feature matrix by points, wherein the elements in the 2×1 feature matrix are the mean values ​​of the feature values ​​of the region; then the element values ​​in the 2×1 feature matrix are updated iteratively, and the number of iterations is 3. When the feature matrix is ​​a 2×1 size matrix, the values ​​of the elements in the feature matrix are Then the autocorrelation between regions can be obtained as Then the attention values ​​in the attention image are According to the formula a=a*aa+b*ab, b=a*ba+b*bb, update the values ​​of the elements in the 2×1 feature matrix, and then obtain the autocorrelation between the regions in the feature map based on the updated element values, so as to obtain the attention image corresponding to the 2×1 feature matrix; according to the attention value in the obtained attention image, obtain the value of the elements in the 2×2 feature matrix, so the 2×2 feature matrix is According to this method, the third characteristic matrix and the fourth characteristic matrix are obtained in sequence.

[0060] Furthermore, in this embodiment, a first threshold is set, and the attention value in the attention image is compared with the first threshold. When the attention value is greater than the first threshold, the corresponding feature map area in the attention image is an area with large autocorrelation. In this area, when obtaining the attention image corresponding to the next feature matrix, it is necessary to calculate the attention value of the attention image based on the element value of the feature map corresponding area in the next feature matrix.

[0061] For example, in this embodiment, the first threshold is set to 0.0001, and the attention image obtained according to the 2×1 feature matrix is The area with larger autocorrelation is the feature map area corresponding to the attention value of 0.59 and the attention value of 0.49. The 2×2 feature matrix is ​​obtained based on the 2×1 feature matrix. Then, according to the attention image corresponding to the 2×1 feature matrix, we find the area with larger autocorrelation in the 2×2 feature matrix. The area with larger autocorrelation in the 2×2 feature matrix is ​​the feature map area corresponding to element ab and element ba. When calculating the attention image according to the 2×2 feature matrix, we only need to calculate the attention value of the area with larger autocorrelation. The attention image corresponding to the 2×2 feature matrix is ​​a 4×4 matrix, so the obtained attention image is

[0062] Furthermore, based on the comparison between the attention value in the attention image and the first threshold, when the attention value is less than the first threshold, the corresponding feature map area in the attention image is an area with smaller autocorrelation. In this area, when the attention image corresponding to the next feature matrix is ​​obtained, the self-attention value of the corresponding area in the attention image of the next feature matrix is ​​predicted.

[0063] For example, in this embodiment, the first threshold is set to 0.0001, and the attention image obtained according to the 2×1 feature matrix is The area with smaller autocorrelation is the feature map area corresponding to the attention value of 0.00001 and the attention value of 0.00002. The 2×2 feature matrix is ​​obtained based on the 2×1 feature matrix. Then, according to the attention image corresponding to the 2×1 feature matrix, we find the area with smaller autocorrelation in the 2×2 feature matrix. The area with smaller autocorrelation in the 2×2 feature matrix is ​​the feature map area corresponding to element aa and element bb. When calculating the attention image according to the 2×2 feature matrix, we predict the attention value of the corresponding area based on the attention value of 0.00001 and the attention value of 0.00002. The attention image corresponding to the 2×2 feature matrix is ​​a 4×4 matrix, so the obtained attention image is

[0064] In the above scheme, when obtaining the attention image corresponding to the large-scale matrix, the attention value of the corresponding area in the attention image corresponding to the large-scale matrix is ​​predicted based on the area with small autocorrelation in the attention image corresponding to the small-scale matrix, which reduces the amount of calculation between matrices, saves time, and improves the efficiency of obtaining the attention image.

[0065] See also Figure 2 , Figure 2This is a flowchart of generating an attention image provided by another embodiment of the present application, relative to Figure 1 In a corresponding embodiment, the method for generating an attention image provided in this embodiment further includes step S14 after steps S11 to S13, which is described in detail as follows:

[0066] S14: Generate a self-attention feature map based on the obtained attention image and feature map.

[0067] In step S14, the attention image is the attention image corresponding to the Nth feature matrix. The attention image contains the correlation between the regions corresponding to the Nth feature matrix and has richer target information. Each region in the attention image corresponding to the Nth feature matrix has the correlation between the expected region and the region, which can represent the importance of each region relative to the target region. Therefore, the attention value corresponding to each region in the attention image can be used as a weight and added to the feature map to generate a self-attention feature map.

[0068] In this embodiment, taking the tumor image in the medical image as an example, N is four, and the fourth feature matrix is ​​an 8×8 matrix. Then the feature image contains 64 feature areas, and the attention value in the obtained attention image is the correlation between any two areas. Then the attention image contains 64×64 correlation information. The greater the correlation, the greater the possibility of the area containing the target information. The 64×64 correlation information is used as the weight of whether the target area is included. The weight is added to the feature map, which can expand the distinction between the target features and the non-target features, so that a more accurate self-attention feature map can be obtained.

[0069] As an embodiment of the present application, step S14 specifically includes:

[0070] Normalizing the Nth attention image corresponding to the N feature matrices to obtain a normalized image; obtaining an h feature map of the feature map by performing a linear transformation on the feature map; and obtaining a self-attention feature map based on the normalized image and the h feature map.

[0071] In this embodiment, after obtaining the attention image according to the feature matrix, the attention image is about the correlation between the various regions in the feature map. The greater the correlation, the greater the possibility of the target feature in the region. By normalizing the attention image, the sum of the attention values ​​in each row is made to 1, and a normalized image is obtained. The attention value in each row is the correlation between any region in the feature matrix and the remaining regions. After normalization, each attention value represents the position coefficient of the target feature. In the subsequent processing process, using the normalized attention image can improve the stability of the model.

[0072] In this embodiment, the feature map is linearly transformed according to a convolutional neural network to obtain an h feature map of the feature map, wherein the convolutional neural network uses a 1×1 convolutional neural network. After obtaining the h feature map, the h feature map is divided into regions, and the division size is equal to the N feature matrices. In this embodiment, the h feature map is divided into 8×8 regions, and then the h feature map is expanded into a (8×8)×1 feature map according to the divided regions. Then, matrix multiplication is performed based on the normalized image and the expanded h image, so that the position coefficient of the target feature is applied to the feature map, and then the self-attention feature map is obtained through the reshape operation.

[0073] Here, by normalizing the attention image, the information of all feature positions can be effectively captured, and the generated image distribution can better approximate the real image distribution, which can improve the stability of the model and indirectly improve the quality of the generated image. Finally, the normalized image is weighted to the feature map obtained by the convolution operation, so that the features in the obtained self-attention image can retain the interdependence between long-distance pixels.

[0074] An embodiment of the present application provides a method for generating an attention image, which divides the feature map of the original image into feature regions to obtain N feature matrices; wherein N is an integer greater than 1; the N feature matrices are sorted in a matrix scale from small to large to obtain an attention image transfer sequence; and according to a preset attention image generation strategy, an attention image corresponding to each feature matrix is ​​generated based on the attention image transfer sequence. In the process of feature extraction, the attention image generates an attention image that reflects the correlation between each region and the target region, and also reflects the importance of each region relative to the target region, which is the correlation between regions. There is no need to calculate the correlation of each pixel in the feature map relative to the rest of the pixels, which greatly reduces the scale of the feature matrix. This method predicts the correlation of the corresponding area of ​​the fine-grained feature map based on the area with smaller autocorrelation of the coarse-grained feature map, reduces the amount of matrix calculation, and improves the efficiency of generating the attention image.

[0075] See also Figure 3 , Figure 3 This is a device interface for a method of generating an attention image provided by an embodiment of the present application. In this embodiment, the server side includes three units for executing Figures 1 to 2 For details of the steps in the corresponding embodiment, please refer to Figures 1 to 2 as well as Figures 1 to 2 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 3 The apparatus 30 for generating an attention image comprises: a dividing unit 31, a sorting unit 32, and a processing unit 33, wherein:

[0076] A division unit 31 performs feature region division on the feature map obtained by extracting features from the original map using a neural network model, obtaining a feature matrix each time the division is performed, and N feature matrices are obtained after N divisions are performed, where N is an integer greater than 1;

[0077] Sorting unit 32: sorting the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence;

[0078] Processing unit 33: generates an attention image corresponding to each feature matrix based on the attention image transfer sequence according to a preset attention image generation strategy, and finally obtains the attention image corresponding to the Nth feature matrix through the attention image generation strategy.

[0079] As an embodiment of the present application, the apparatus 30 for generating an attention image further includes: a generating unit 34 .

[0080] Generating unit 34: Generates a self-attention feature map based on the obtained attention image and feature map.

[0081] As an embodiment of the present application, the apparatus 30 for generating an attention image further includes:

[0082] The first execution unit 35 calculates the eigenvalue mean of the feature region where each element value in the first feature matrix of the N feature matrices is located.

[0083] The second execution unit 36: based on the i-th feature matrix, obtain the i+1-th feature matrix; wherein i is an integer greater than 1 and less than N; based on the first threshold, obtain the area with larger autocorrelation in the attention image of the i-th feature matrix; based on the i+1-th feature matrix, obtain the self-attention value of the area with larger autocorrelation in the attention image of the i+1-th feature matrix.

[0084] The third execution unit 37: according to the first threshold, obtain the area with smaller autocorrelation in the attention image of the i-th feature matrix; according to the area with smaller autocorrelation in the attention image of the i-th feature matrix, predict the self-attention value of the corresponding area in the attention image of the i+1-th feature matrix.

[0085] As an embodiment of the present application, the division unit 31 is specifically used to extract the features of the original image using a neural network model to obtain a feature map, and divide the feature map into regions. Each region is represented by an eigenvalue, and each region division can obtain a feature matrix composed of eigenvalues. The feature map is divided into regions N times to obtain N feature matrices.

[0086] As an embodiment of the present application, the processing unit 33 is specifically configured to obtain an i+1th feature matrix based on the i-th feature matrix, where i is an integer greater than 1 and less than N; obtain, based on the first threshold, a region with a large autocorrelation in the attention image of the i-th feature matrix; obtain, based on the i+1th feature matrix, a self-attention value for the region with a large autocorrelation in the attention image of the i+1th feature matrix; obtain, based on the first threshold, a region with a small autocorrelation in the attention image of the i-th feature matrix; and predict, based on the region with a small autocorrelation in the attention image of the i-th feature matrix, a self-attention value for the corresponding region in the attention image of the i+1th feature matrix.

[0087] As an embodiment of the present application, the apparatus 30 for generating an attention image further includes:

[0088] The fourth execution unit 38: normalizes the Nth attention image corresponding to the N feature matrices to obtain a normalized image; and obtains an h feature map of the feature map by performing a linear transformation on the feature map.

[0089] It should be understood that Figure 3 In the structural block diagram of the device for generating the attention image shown, each unit is used to perform Figures 1 to 2 The steps in the corresponding embodiment, and for Figures 1 to 2 Each step in the corresponding embodiment has been explained in detail in the above embodiment. Figures 1 to 2 as well as Figures 1 to 2 The relevant descriptions in the corresponding embodiments will not be repeated here.

[0090] In one embodiment, a computer device is provided. The computer device is a server, and its internal structure diagram can be as follows: Figure 4 As shown. The computer device 40 includes a processor 41, an internal memory 43, and a network interface 44 connected via a system bus 42. Among them, the processor 41 of the computer device is used to provide computing and control capabilities. The memory of the computer device 40 includes a readable storage medium 45 and an internal memory 43. The readable storage medium 45 stores an operating system 46, computer-readable instructions 47, and a database 48. The internal memory 43 provides an environment for the operation of the operating system 46 and the computer-readable instructions 47 in the readable storage medium 45. The database 48 of the computer device 40 is used to store data involved in the method for generating an attention image. The network interface 44 of the computer device 40 is used to communicate with an external terminal via a network connection. When the computer-readable instructions 47 are executed by the processor 41, a method for generating an attention image is implemented. The readable storage medium 45 provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0091] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0092] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0093] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for generating an attention image, characterized in that include: Performing feature region division on a feature map obtained by extracting features from the original image using a neural network model, obtaining a feature matrix each time the division is performed, and performing N divisions to obtain N feature matrices, where N is an integer greater than 1; sorting the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence; generating an attention image corresponding to each feature matrix based on the attention image transfer sequence according to a preset attention image generation strategy, and finally obtaining an attention image corresponding to the Nth feature matrix through the attention image generation strategy; The preset attention image generation strategy includes: According to the i-th characteristic matrix, obtain the i+1-th characteristic matrix; where i is an integer greater than 1 and less than N; According to the first threshold, obtain the area with large autocorrelation in the attention image of the i-th feature matrix; According to the i+1th feature matrix, the self-attention value of the area with large autocorrelation in the attention image of the i+1th feature matrix is ​​obtained.

2. The method for generating an attention image according to claim 1, wherein: The feature map obtained by extracting features from the original map using the neural network model is divided into feature regions, and each division obtains a feature matrix. After N divisions, the obtained N feature matrices include: The feature map obtained by extracting features from the original map using a neural network model is divided into feature regions, and a feature matrix is ​​obtained each time the feature map is divided. After N divisions, N feature matrices are obtained. Each time the feature map is divided, the feature map is divided into regions of equal size.

3. The method for generating an attention image according to claim 1, wherein: The feature map obtained by extracting features from the original map using the neural network model is divided into feature regions, and each division obtains a feature matrix. After N divisions, the obtained N feature matrices include: The feature map obtained by extracting features from the original map using a neural network model is divided into feature regions. A feature matrix is ​​obtained each time the feature map is divided. After N divisions, N feature matrices are obtained, where each element value in the first feature matrix of the N feature matrices is the mean of the eigenvalues ​​of the feature region.

4. The method for generating an attention image according to claim 1, wherein: The preset attention image generation strategy also includes: According to the first threshold, obtaining an area with small autocorrelation in the attention image of the i-th feature matrix; According to the area with small autocorrelation in the attention image of the i-th feature matrix, the self-attention value of the corresponding area in the attention image of the i+1-th feature matrix is ​​predicted.

5. The method for generating an attention image according to claim 1, wherein: The step of obtaining the self-attention value of the region with large autocorrelation in the attention image of the i+1th feature matrix according to the i+1th feature matrix includes: The self-attention value of the area with large autocorrelation in the attention image of the i+1th feature matrix is ​​obtained by the dot product of the i+1th feature matrix.

6. The method for generating an attention image according to claim 1, wherein: The method further comprises: Normalizing the Nth attention image corresponding to the N feature matrices to obtain a normalized image; Obtaining an h feature map of the feature map by performing a linear transformation on the feature map; A self-attention feature map is obtained according to the normalized image and the h feature map.

7. A device for generating an attention image, characterized in that include: Division unit: divides the feature map obtained by extracting features from the original map using the neural network model into feature regions, obtains a feature matrix each time the map is divided, and obtains N feature matrices after N divisions, where N is an integer greater than 1; Sorting unit: sorting the N feature matrices in ascending order of matrix scale to obtain an attention image transfer sequence; Processing unit: generating an attention image corresponding to each feature matrix based on the attention image transfer sequence according to a preset attention image generation strategy, and finally obtaining an attention image corresponding to the Nth feature matrix through the attention image generation strategy; The preset attention image generation strategy includes: According to the i-th characteristic matrix, obtain the i+1-th characteristic matrix; where i is an integer greater than 1 and less than N; According to the first threshold, obtain the area with large autocorrelation in the attention image of the i-th feature matrix; According to the i+1th feature matrix, the self-attention value of the area with large autocorrelation in the attention image of the i+1th feature matrix is ​​obtained.

8. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein: When the processor executes the computer-readable instructions, the processor implements the steps of the method for generating an attention image as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating an attention image as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Facial attribute recognition method and apparatus, and electronic device and storage medium

    WO2021063056A1