Image Feature Extraction Method, Device and Equipment Based on Self-Attention Mechanism

The self-attention mechanism in the neural network encoder-decoder model with multiple sliding windows enhances image feature extraction, addressing limitations of conventional methods by providing a more precise and comprehensive representation of image features.

CN114581682BActive Publication Date: 2025-07-15HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210163230.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-07-15
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

The existing neural network-based codec models have problems such as large computational cost and low computational efficiency in image feature extraction, and it is difficult to describe image features from multiple dimensions.

Method used

The Transformer model based on the self-attention mechanism is adopted to encode the image self-attention through multiple sliding windows, combine the cascaded convolutional layer and self-attention calculation layer to extract the image feature vectors, and use multiple sliding windows to extract image features from different dimensions.

Benefits of technology

It improves the accuracy and efficiency of image feature extraction, can describe images from multiple dimensions, enhances the expression ability of image features, and reduces the calculation amount of self-attention calculation layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581682B_ABST
    Figure CN114581682B_ABST
Patent Text Reader

Abstract

One or more embodiments of this specification provide a method, apparatus, and device for extracting image features based on a self-attention mechanism, including: obtaining a target image; inputting the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on multiple supported sliding windows to obtain a feature vector of the target image; wherein, each sliding window corresponds to a different image feature extraction dimension; obtaining the feature vector of the target image output by the neural network-based encoding and decoding model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of computer vision technology, and in particular, to a method, apparatus, and device for extracting image features based on a self-attention mechanism. Background Art

[0002] In related technologies, the neural network-based encoding and decoding model is a basic model based on the self-attention mechanism and has powerful representation capabilities, and is usually used in natural language processing. For example, if the natural language processing task is English to Chinese translation, the English sentence is input into the neural network-based encoding and decoding model for processing, and the model can output the corresponding Chinese sentence of the English sentence.

[0003] Inspired by the powerful representation capabilities of the neural network-based encoding and decoding model, some scholars have extended the neural network-based encoding and decoding model to the field of machine vision, such as extracting image features through the neural network-based encoding and decoding model. Compared with the commonly used convolutional neural network for image feature extraction, the neural network-based encoding and decoding model has better performance in image feature extraction. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide a method, apparatus, and device for extracting image features based on a self-attention mechanism.

[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:

[0006] According to the first aspect of one or more embodiments of this specification, a method for extracting image features based on a self-attention mechanism is proposed, including:

[0007] Obtain a target image;

[0008] Input the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image; wherein, various sliding windows correspond to different image feature extraction dimensions;

[0009] Obtain the feature vector of the target image output by the neural network-based encoding and decoding model.

[0010] Optionally, the neural network-based encoding and decoding model includes a first module; the first module includes: a first convolutional layer and a self-attention calculation layer;

[0011] The performing self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image includes:

[0012] The first convolutional layer in the first module performs convolution on the target image to obtain an initial feature matrix of the target image, and uses the initial feature matrix as an input matrix to input to the self-attention calculation layer of this module;

[0013] The self-attention calculation layer of the first module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, and determines the feature vector of the target image according to the output matrix of the first module.

[0014] Optionally, the encoding and decoding model based on a neural network further includes: at least one cascaded second module; the second module includes a second convolutional layer and a self-attention calculation layer;

[0015] The determining the feature vector of the target image according to the output matrix of the first module includes:

[0016] The self-attention calculation layer of the first module inputs the output matrix of the first module to the second module at the first position;

[0017] The second convolutional layer of each second module performs convolution on the output matrix output by the previous module, and uses the convolution result as an input matrix to input to the self-attention calculation layer of this second module;

[0018] The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module;

[0019] The self-attention calculation layer of the last second module uses the obtained output matrix as the feature vector of the target image.

[0020] Optionally, the performing self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module includes:

[0021] For each type of sliding window, perform self-attention encoding on the input matrix based on this type of sliding window to obtain an encoded matrix;

[0022] Perform convolution on the encoded matrix; wherein, the convolved encoded matrix represents the position features of each pixel point of the target image; the position features include local position features, and / or global position features;

[0023] Concatenate the convolved encoded matrix with the encoded matrix to obtain a concatenated matrix;

[0024] Fuse the concatenated matrices corresponding to each type of sliding window to obtain the output matrix of this module.

[0025] Optionally, the attention encoding of the input matrix based on the sliding window to obtain an encoded matrix includes:

[0026] Sliding the sliding window on the input matrix according to a preset step size;

[0027] For each sliding operation, determine the target area circled by the sliding window on the input matrix. For each element in the target area, calculate the correlation degree scores between the element and itself and other elements in the target area respectively, and obtain the self-attention encoding value corresponding to the element according to the correlation degree scores, and update the value of the element in the input matrix to the self-attention encoding value to obtain the encoded matrix.

[0028] Optionally, the first convolutional layer in the first module performs convolution on the target image to obtain an initial feature matrix of the target image, including:

[0029] The first convolutional layer performs at least one continuous convolution process on the target image to obtain an initial feature matrix of the target image; wherein, the size of the convolutional kernel corresponding to each convolution process is less than or equal to a preset threshold.

[0030] Optionally, the multiple sliding windows include: a first sliding window for extracting image features from the global dimension of the image and a second sliding window for extracting image features from the local dimension of the image.

[0031] Optionally, each sliding window includes: at least one effective calculation area; when there are multiple effective calculation areas, the multiple effective calculation areas are arranged at intervals in the sliding window;

[0032] The determination of the target area circled by the sliding window on the input matrix includes:

[0033] Taking the area circled by the effective area of the sliding window on the input matrix as the target area.

[0034] Optionally, the neural network-based encoding and decoding model is a Transformer model supporting the self-attention mechanism.

[0035] According to the second aspect of one or more embodiments of the present specification, an image feature extraction device based on the self-attention mechanism is proposed, including:

[0036] A first acquisition module for acquiring a target image;

[0037] An input module for inputting the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image; wherein, each sliding window corresponds to a different image feature extraction dimension;

[0038] A second acquisition module for acquiring the feature vector of the target image output by the neural network-based encoding and decoding model.

[0039] Optionally, the neural network-based encoding and decoding model includes a first module; the first module includes: a first convolutional layer and a self-attention calculation layer;

[0040] When performing self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image, the first convolutional layer in the first module is used to perform convolution on the target image to obtain an initial feature matrix of the target image, and the initial feature matrix is used as an input matrix and input into the self-attention calculation layer of this module;

[0041] The self-attention calculation layer of the first module is used to perform self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain an output matrix of this module, and determine the feature vector of the target image according to the output matrix of the first module.

[0042] Optionally, the neural network-based encoding and decoding model further includes: at least one cascaded second module; the second module includes a second convolutional layer and a self-attention calculation layer;

[0043] When determining the feature vector of the target image according to the output matrix of the first module, the self-attention calculation layer of the first module is used to input the output matrix of the first module into the first second module in the front row;

[0044] The second convolutional layer of each second module is used to perform convolution on the output matrix output by the previous module, and the convolution result is used as an input matrix and input into the self-attention calculation layer of this second module;

[0045] The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain an output matrix of this module;

[0046] The self-attention calculation layer of the last second module is used to use the obtained output matrix as the feature vector of the target image.

[0047] Optionally, when the self-attention calculation layer performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, for each type of sliding window, it performs self-attention encoding on the input matrix based on this type of sliding window to obtain an encoded matrix; performs convolution on the encoded matrix; wherein, the encoded matrix after convolution represents the position features of each pixel point of the target image; the position features include local position features and / or global position features; splices the encoded matrix after convolution with the encoded matrix to obtain a spliced matrix; fuses the spliced matrices corresponding to each type of sliding window to obtain the output matrix of this module.

[0048] Optionally, when the self-attention calculation layer performs attention encoding on the input matrix based on this sliding window to obtain an encoded matrix, it is used to slide this sliding window on the input matrix according to a preset step size; for each sliding operation, determine the target area circled by this sliding window on this input matrix, for each element in this target area, calculate the correlation degree score between this element and itself and other elements in this target area respectively, and obtain the self-attention encoding value corresponding to this element based on the correlation degree score, and update the value of this element in the input matrix to this self-attention encoding value to obtain the encoded matrix.

[0049] Optionally, the first convolutional layer in the first module, when performing convolution on the target image to obtain the initial feature matrix of this target image, is used to perform at least one continuous convolution process on the target image to obtain the initial feature matrix of this target image; wherein, the size of the convolutional kernel corresponding to each convolution process is less than or equal to a preset threshold.

[0050] Optionally, the multiple sliding windows include: a first sliding window for extracting image features from the global dimension of the image, and a second sliding window for extracting image features from the local dimension of the image.

[0051] Optionally, each type of sliding window includes: at least one effective calculation area; when there are multiple effective calculation areas, the multiple effective calculation areas are arranged at intervals in the sliding window;

[0052] When the self-attention calculation layer determines the target area circled by this sliding window on this input matrix, it is used to circle the area circled by the effective area of this sliding window on this input matrix as the target area.

[0053] Optionally, the neural network-based encoding and decoding model is a Transformer model supporting the self-attention mechanism.

[0054] According to the third aspect of one or more embodiments of this specification, an electronic device is proposed, including:

[0055] A processor;

[0056] A memory for storing instructions executable by the processor;

[0057] Wherein, the processor realizes the above-mentioned image feature extraction method based on the self-attention mechanism by running the executable instructions.

[0058] According to the fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is proposed, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method for extracting image features based on the self-attention mechanism are realized.

[0059] As can be seen from the above description, an electronic device can obtain a target image; input the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image; wherein, various sliding windows correspond to different image feature extraction dimensions; obtain the feature vector of the target image output by the neural network-based encoding and decoding model.

[0060] Since a variety of sliding windows correspond to a variety of different image feature extraction dimensions, the neural network-based encoding and decoding model in this specification can extract image features from multiple dimensions, so that the extracted image features can describe the image from multiple dimensions, making the image features more accurate in describing the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic structural diagram of an image feature extraction system shown in an exemplary embodiment of this specification

[0062] Figure 2 is a flowchart of a method for extracting image features based on the self-attention mechanism shown in an exemplary embodiment of this specification;

[0063] Figure 3a is a schematic diagram of a first module in a neural network-based encoding and decoding model shown in an exemplary embodiment of this specification;

[0064] Figure 3b is a schematic diagram of a convolution process shown in an exemplary embodiment of this specification;

[0065] Figure 4 is a schematic structural diagram of a neural network-based encoding and decoding model shown in an exemplary embodiment of this specification;

[0066] Figure 5a and Figure 5b is a schematic diagram of a sliding window shown in an exemplary embodiment of this specification;

[0067] Figure 6a It is a schematic diagram of a sliding window shown in an exemplary embodiment of this specification;

[0068] Figure 6b It is a schematic diagram of the relationship between a sliding window and a matrix shown in an exemplary embodiment of this specification;

[0069] Figure 7a It is a schematic diagram of an existing sliding window shown in an exemplary embodiment of this specification;

[0070] Figure 7b and Figure 7c and Figure 7d It is a schematic diagram of an existing sliding window sliding on a matrix shown in an exemplary embodiment of this specification;

[0071] Figure 8a It is a schematic diagram of a sliding window shown in an exemplary embodiment of this specification;

[0072] Figure 8b and Figure 8c and Figure 8d It is a schematic diagram of an existing sliding window sliding on a matrix shown in an exemplary embodiment of this specification;

[0073] Figure 9a It is a schematic diagram of a sliding window shown in an exemplary embodiment of this specification;

[0074] Figure 9b It is a schematic diagram of a sliding window circling a target area on a matrix shown in an exemplary embodiment of this specification;

[0075] Figure 10 It is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this specification.

[0076] Figure 11 It is a block diagram of an extraction device for image features based on a self-attention mechanism provided in an exemplary embodiment of this specification. Detailed implementation manners

[0077] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with one or more embodiments of this specification. On the contrary, they are only examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0078] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0079] Figure 1 It is a schematic structural diagram of an image feature extraction system shown in an exemplary embodiment of this specification.

[0080] As Figure 1 shown, the system may include: an electronic device and an encoding and decoding model based on a neural network.

[0081] Among them, the electronic device refers to a device with storage and computing capabilities, which may be a server, a server cluster, a virtual server, a cloud server, a data center, a computer, etc. Here, only an exemplary description of the electronic device is given, and no specific limitation is imposed on it.

[0082] The encoding and decoding model based on a neural network is an encoding and decoding model built on a neural network. The encoding and decoding model based on a neural network may be a model that supports the self-attention mechanism, such as a Transformer model that supports the self-attention mechanism, or other models. Here, only an exemplary description of the encoding and decoding model based on a neural network is given, and no specific limitation is imposed on it.

[0083] In this specification, the electronic device may obtain a target image and input the target image into a trained encoding and decoding model based on a neural network, so that the encoding and decoding model based on a neural network performs self-attention encoding on the target image based on multiple supported sliding windows to obtain a feature vector of the target image; among them, various sliding windows correspond to different image feature extraction dimensions; obtain the feature vector of the target image output by the encoding and decoding model based on a neural network.

[0084] Since multiple sliding windows correspond to multiple different image feature extraction dimensions, the encoding and decoding model based on a neural network in this specification can extract image features from multiple dimensions, so that the extracted image features can describe the image from multiple dimensions, making the description of the image by the image features more accurate.

[0085] See Figure 2 , Figure 2 is a flowchart of a method for extracting image features based on the self-attention mechanism shown in an exemplary embodiment of this specification. This method can be applied to an electronic device and may include the following steps.

[0086] Step 202: Obtain a target image;

[0087] In implementation, the electronic device may use the image input by an external device as the target image. Here, the external device may refer to a user terminal or other electronic devices, such as other servers and other external devices. This is only an exemplary description of the external device and is not specifically limited.

[0088] Optionally, the electronic device locally stores images. When the electronic device receives a task instruction for image feature extraction, it obtains the image stored locally that is the same as the target image indicated by the task instruction. For example, if the image identifier carried in the task instruction is 1, the electronic device will obtain the image with the identifier 1 from the images stored locally as the target image.

[0089] Optionally, the electronic device may periodically obtain the images stored locally that have not been subjected to feature extraction as the target images.

[0090] This is only an exemplary description of obtaining the target image and is not specifically limited.

[0091] Step 204: Input the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on multiple supported sliding windows, and obtains a feature vector of the target image; where various sliding windows correspond to different image feature extraction dimensions.

[0092] In implementing Step 204, a trained neural network-based encoding and decoding model (hereinafter simply referred to as the encoding and decoding model) is set on the electronic device or other devices. The electronic device may input the target image into the trained encoding and decoding model, and the encoding and decoding model may perform self-attention encoding on the target image based on multiple supported sliding windows to obtain the feature vector of the target image.

[0093] The following details "the neural network-based encoding and decoding model performs self-attention encoding on the target image based on multiple supported sliding windows and obtains the feature vector of the target image" through Steps A1 to A2.

[0094] As Figure 3a shown, in an optional implementation manner, the neural network-based encoding and decoding model includes a first module. The first module includes a first convolutional layer and a self-attention calculation layer. The first convolutional layer is cascaded with the self-attention calculation layer, and the output result of the first convolutional layer is the input of the self-attention calculation layer.

[0095] Step A1: The first convolutional layer performs convolution on the target image to obtain an initial feature matrix of the target image, and takes the initial feature matrix as an input matrix and inputs it into the self-attention calculation layer of this module.

[0096] In an alternative way, the first convolutional layer corresponds to multiple convolutional kernels, and the size of each convolutional kernel is less than or equal to a preset threshold. For example, if the preset threshold is 3*3, then the size of each convolutional kernel is less than or equal to 3*3.

[0097] In the embodiments of this specification, the first convolutional layer performs at least one continuous convolution process on the target image to obtain an initial feature matrix of the target image; wherein, each convolution process uses one convolutional kernel.

[0098] For example, as Figure 3b shown, assume that the first convolutional layer has two convolutional kernels. These two convolutional kernels are convolutional kernel 1 and convolutional kernel 2 respectively. The size of each convolutional kernel is 3*3, and assume that the size of the initially input image is 224*224. After the first convolutional layer performs convolution processing on the target image using convolutional kernel 1, the size of the obtained matrix is 112*112. Then, the first convolutional layer uses convolutional kernel 2 to perform convolution processing on the matrix with a size of 112*112 to obtain an initial feature matrix, and the size of this initial feature matrix is 56*56.

[0099] The advantage of performing continuous convolution using multiple convolutional kernels with a size smaller than the preset threshold is that:

[0100] The existing convolution method usually obtains an initial feature matrix by performing one convolution on the target image using a convolutional kernel with a larger size. However, the method adopted in this specification is to perform continuous convolution on the target image using cascaded convolutional kernels with smaller sizes. Since the size of the convolutional kernel used in each convolution is smaller, and subsequent multiple convolutions are also performed on the basis of one convolution, the convolution granularity is finer, so that this convolution method can enhance the model's ability to extract image features.

[0101] Of course, in practical applications, the first convolutional layer can also perform convolution on the target image using a large convolutional kernel with a size larger than the size threshold to obtain an initial feature matrix. Here, it is only an exemplary description of the way for the first convolutional layer to obtain the initial feature matrix, and no specific limitation is imposed on it.

[0102] In the embodiments of this specification, after obtaining the initial feature matrix, the first convolutional layer can take the initial feature matrix as an input matrix and input it into the self-attention calculation layer of the first module.

[0103] Step A2: The self-attention calculation layer of the first module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, and determines the feature vector of the target image according to the output matrix of the first module.

[0104] In an alternative method, the neural network-based encoding and decoding model further includes: at least one cascaded second module; the second module includes a second convolutional layer and a self-attention calculation layer;

[0105] The self-attention calculation layer of the first module inputs the output matrix of the first module into the second module at the first position;

[0106] The second convolutional layer of each second module performs convolution on the output matrix output by the previous module to obtain the input matrix of this second module, and inputs the input matrix into the self-attention calculation layer of this second module;

[0107] The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module;

[0108] The self-attention calculation layer of the last second module outputs the obtained output matrix as the feature vector of the target image.

[0109] As Figure 4 shown, Figure 4 the neural network-based encoding and decoding model in includes 4 modules, namely Module 1, Module 2, Module 3, and Module 4. Module 1 is connected to Module 2, Module 2 is connected to Module 3, and Module 3 is connected to Module 4. Among them, the first module is Module 1, and the second modules are Module 2, Module 3, and Module 4.

[0110] The first convolutional layer of Module 1 performs convolution on the original image to obtain the initial feature matrix, and inputs the initial feature matrix as the input matrix into the self-attention calculation layer of Module 1. The self-attention calculation layer performs self-attention encoding on the input matrix based on multiple sliding windows to obtain the output matrix, and then inputs the output matrix into the second convolutional layer of Module 2.

[0111] The second convolutional layer of Module 2 performs convolution on the output matrix output by Module 1, and inputs the convolution result as the input matrix into the self-attention calculation layer of Module 2. The self-attention calculation layer performs self-attention encoding on the input matrix based on multiple sliding windows to obtain the output matrix, and then inputs the output matrix into the second convolutional layer of Module 3.

[0112] The second convolutional layer of module 3 convolves the output matrix output by module 2, and takes the convolution result as the input matrix and inputs it into the self-attention calculation layer of module 3. The self-attention calculation layer performs self-attention encoding on the input matrix based on multiple sliding windows to obtain an output matrix, and then inputs the output matrix into the second convolutional layer of module 4.

[0113] The first convolutional layer of module 4 convolves the output matrix output by module 3, and takes the convolution result as the input matrix and inputs it into the self-attention calculation layer of module 4. The self-attention calculation layer performs self-attention encoding on the input matrix based on multiple sliding windows to obtain an output matrix, and outputs the obtained output matrix as the feature vector of the original image.

[0114] The advantages of the encoding and decoding model designed in this way are as follows:

[0115] In practical applications, the extracted image features will be used for various image tasks, such as image detection tasks, image segmentation tasks, image classification tasks, etc. When the image task is image detection or image segmentation, it is often necessary to input an image with a larger resolution (i.e., a larger size) for feature extraction so that the extracted image features can meet such image tasks. However, inputting an image with a larger resolution for feature extraction will increase the computational load of the self-attention calculation layer.

[0116] To reduce the computational load of self-attention encoding, at least one second module is cascaded after the first module. The second convolutional layer of each second module will first convolve the matrix output by the previous module to reduce the matrix size and increase the number of channels, and then perform self-attention calculation on the convolved matrix. That is to say, through the way of module cascading, the matrix size input to the self-attention calculation layer of each module is successively reduced to achieve the purpose of increasing the receptive field and reducing the computational load of the self-attention calculation layer. At the same time, the loss caused by the reduction in size is compensated by increasing the number of channels.

[0117] Still taking Figure 4 as an example, assume that the size of the original image is H*W, where H is the length, W is the width, and the number of channels is C.

[0118] After being convolved by the first convolutional layer of module 1, the matrix size input to the self-attention calculation layer of this module 1 is H / 4*W / 4, and the number of channels remains C.

[0119] After being convolved by the second convolutional layer of module 2, the matrix size input to the self-attention calculation layer of this module 2 is H / 8*W / 8, and the number of channels is 2C.

[0120] After being convolved by the second convolutional layer of module 3, the matrix size input to the self-attention calculation layer of this module 3 is H / 16*W / 16, and the number of channels is 4C.

[0121] After being processed by the second convolutional layer of Module 4, the matrix size of the input to the self-attention calculation layer of Module 4 is H / 32 * W / 32, and the number of channels is 8C.

[0122] It can be seen that every time passing through the convolutional layer of a module, the size of the matrix output by the previous module is halved, and the number of channels is doubled. Halving the matrix size can increase the receptive field and reduce the computational amount of the self-attention calculation layer at the same time. Doubling the number of channels can make up for the loss of halving the size and ensure the model accuracy.

[0123] Of course, in practical applications, the encoding and decoding model of the neural network can also only include the first module. For example, when the resolution of the original image pixels is relatively low, the encoding and decoding model based on the neural network can also only include the first module. The first convolutional layer performs convolution on the target image to obtain the initial feature matrix of the target image, and takes the initial feature matrix as the input matrix and inputs it into the self-attention calculation layer of this module. The self-attention calculation layer of the first module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, and extracts the feature vector of the image from the output matrix output by this layer. Here, only an exemplary description of the model structure is given, and no specific limitation is imposed on it.

[0124] Since the self-attention encoding of each module in this specification involves multiple sliding windows, multiple sliding windows will be introduced first, and then the method of performing self-attention encoding on the input matrix input to this layer by the self-attention layer with multiple sliding windows will be introduced.

[0125] 1) Multiple sliding windows

[0126] This specification proposes multiple sliding windows, and different sliding windows represent different feature extraction methods. For example, the multiple sliding windows proposed in this specification include: the first sliding window for extracting image features from the global dimension of the image, and the second sliding window for extracting image features from the local dimension of the image. Among them, the first sliding window can include one or more windows, and the second sliding window can also include one or more windows.

[0127] For example, the sliding windows in this specification are composed of DSSA (Dilation Stripe Self-Attention) windows and CWSA (Cross Window Self-Attention) windows.

[0128] The DSSA window is the first sliding window for extracting image features from the global dimension of the image, while the CWSA window is the second sliding window for extracting image features from the local dimension of the image.

[0129] For example, by Figure 5aIt can be seen that Figure 5a in Figure 5a , the DSSA window is composed of a horizontal axis window and a vertical axis window, and the horizontal axis window and the vertical axis window are the first sliding windows for extracting image features from the global dimension.

[0130] From Figure 5b It can be seen that Figure 5b the CWSA window in Figure 5b is composed of an interval window and a fixed window, and the interval window and the fixed window are the second sliding windows for extracting image features from the local dimension.

[0131] Optionally, in the embodiments of this specification, an effective calculation area is set in the sliding window. The effective calculation area means that in each slide, when the sliding window encloses a matrix area, the elements located in the effective calculation area in this matrix area are subjected to self-attention encoding.

[0132] For example, as Figure 6a shown, assuming Figure 6a the shaded part in the vertical axis window in Figure 6a represents the effective calculation area, and the white horizontal bar represents the invalid calculation area.

[0133] As Figure 6b shown, Figure 6b represents a matrix, Figure 6b and each grid in Figure 6b represents an element in the matrix. Assuming that when the sliding window makes the first slide, the target area enclosed by the sliding window in the matrix is area A (i.e., Figure 6b the solid line frame area in Figure 6b ), self-attention calculation is performed between the elements located in the shaded part in area A, that is, Figure 6b self-attention calculation is performed between the elements located in the shaded part in the second column, the fourth column, and the sixth column of the solid line rectangle.

[0134] Optionally, in this specification, when there are multiple effective calculation areas in the sliding window, the multiple effective calculation areas are set at intervals.

[0135] For example, Figure 5a the shaded horizontal bars in the vertical axis window in Figure 5a represent the effective calculation areas, and the white horizontal bars represent the invalid calculation areas. This vertical axis window includes 3 effective calculation areas, and the 3 effective areas are set at intervals. For example, in the order from left to right, an invalid calculation area is set between the first effective calculation area and the second effective calculation area, and an invalid calculation area is set between the second effective calculation area and the third effective calculation area. Through this setting method, the 3 effective calculation areas are set at intervals.

[0136] For another example, as shown in the interval window in 5b, each black rectangle in the interval window represents an effective calculation area. An invalid calculation area is set between two adjacent effective calculation areas.

[0137] It should be noted that the advantage of setting an effective calculation area at intervals in the sliding window is as follows:

[0138] In the existing self-attention mechanism, all areas in the sliding window are effective areas. And after sliding according to the preset step size, the areas circled by the sliding window each time it slides on the matrix do not overlap, which makes it impossible to establish the correlation relationship between different areas.

[0139] However, in the sliding window adopted in this specification, the effective areas are set at intervals. Therefore, it is beneficial for the multiple areas circled by the sliding window sliding on the matrix multiple times to overlap, so the correlation relationship between areas can be established.

[0140] For example, the existing sliding window setting is as Figure 7a shown.

[0141] Figures 7b - 7d represents Figure 7a the process of the sliding window shown sliding in the matrix. Assuming the sliding step size is 1, the areas circled each time it slides do not overlap.

[0142] Assume that the sliding window slides three times. For the first slide, as Figure 7b described, the sliding window circles the first column. For the second slide, as Figure 7c shown, the sliding window circles the second column. For the third slide, as Figure 7d shown, the sliding window circles the third column. It can be seen from this that the areas circled in the three slides do not overlap, so the correlation relationship between areas cannot be established.

[0143] However, the sliding window in this specification is assumed to be as Figure 8a shown.

[0144] Figures 8b - 8d represents Figure 8a the sliding window shown sliding in the matrix. Assume the sliding step size is 1.

[0145] Assume that the sliding window slides three times. For the first slide, as Figure 8b described, the sliding window circles the shaded parts of the 2nd, 4th, and 6th columns. For the second slide, as Figure 8c shown, the sliding window circles the shaded parts of the 3rd, 5th, and 7th columns. For the third slide, as Figure 8d shown, the sliding window circles the shaded parts of the 4th, 6th, and 8th columns. It can be seen from this that there is an overlap between the first slide and the third slide (i.e., the shaded part of the 4th column), so the establishment of the correlation relationship between the elements of different areas can be realized.

[0146] 2) Taking the self-attention layer performing self-attention encoding on the input matrix input to this layer with multiple sliding windows as an example for illustration

[0147] In the embodiments of this specification, the self-attention encoding methods of the self-attention calculation layers in the first module and the second module are the same, except for the parameters used in the encoding and the encoding objects.

[0148] For example, the encoding object of the self-attention layer in the first module is the matrix output by the first convolutional layer of the first module. And the encoding object of the self-attention calculation layer in the second module is the matrix output by the second convolutional layer of the second module.

[0149] In implementation, whether it is the self-attention calculation layer in the first module or the self-attention calculation layer in the second module, the following operations are performed.

[0150] The self-attention calculation layer performs self-attention encoding on the input matrix based on each type of sliding window to obtain an encoded matrix. Then, the self-attention calculation layer convolves the encoded matrix and concatenates the convolved encoded matrix with the encoded matrix to obtain a concatenated matrix; where the convolved encoded matrix represents the position features of each pixel point of the target image; the position features include local position features or global position features. Finally, the self-attention calculation layer can fuse the concatenated matrices corresponding to each type of sliding window to obtain the output matrix of this module.

[0151] The following takes the example of a multi-sliding window self-attention encoding of the input matrix of an arbitrary self-attention layer through steps B1 to B4 for illustration.

[0152] Step B1: Perform attention encoding on the input matrix based on each type of sliding window to obtain an encoded matrix.

[0153] In implementation, the self-attention calculation layer slides the sliding window on the input matrix according to a preset step size;

[0154] For each sliding operation, the self-attention calculation layer determines the target area circled by the sliding window on the input matrix. For each element in the target area, calculate the correlation degree scores between this element and itself and other elements in the target area respectively, and obtain the self-attention encoding value corresponding to this element based on the correlation degree scores, and update this element to this self-attention encoding value in the input matrix to obtain the encoded matrix.

[0155] For example, assume the sliding window is as Figure 9a shown,

[0156] such as Figure 9b shown, assume Figure 9b the dashed rectangle represents the input matrix, and each grid in the dashed rectangle represents an element. Assume the area circled by the sliding window after the first slide Figure 9bas shown by the solid rectangle in

[0157] It can be seen that the sliding window encloses a 3*3 area after the first slide. Since the first and third columns of the sliding window are the effective calculation areas, the first and third columns in this 3*3 area are the target areas.

[0158] For the convenience of narration, Figure 9b the element in the first row and first column in

[0159] is denoted as a11, the element in the first row and second column is denoted as a12, …, and the element in the i-th row and j-th column is denoted as aij.

[0160] Taking a11 in the target area as an example for illustration. The self-attention layer can calculate the correlation degree scores of a11 with a11, a21, a31, a13, a23, and a33 respectively. Then, the self-attention calculation layer can calculate the encoded value of a11 based on these correlation degree scores.

[0161] By analogy, after the sliding window completes all slides, an encoded matrix can be obtained.

[0162] In addition, for the specific implementation method of "for each element in the target area, calculate the correlation degree scores of this element with itself and other elements in the target area respectively, and obtain the self-attention encoded value corresponding to this element based on the correlation degree scores", it can be obtained according to the existing calculation method of the self-attention mechanism. Here, the method of obtaining the self-attention calculation method is briefly introduced.

[0163] In implementation, for each element in the target area, the self-attention calculation layer can calculate the q vector, k vector, and v vector of this element.

[0164] Then, for each element, the self-attention calculation layer can multiply the q vector of this element with its own k vector respectively to obtain the correlation degree score of this element with itself. The self-attention layer multiplies the q vector of the value of this feature with the k vectors of other elements in the target area respectively to obtain the correlation degree scores between this element and other elements. Then, the self-attention calculation layer divides the correlation degree scores of this element with itself and other elements in the target area by a preset scale scalar, and then performs normalization and softmax processing, and multiplies the processing result with the v vector of this element to obtain the self-attention encoded value of this element.

[0165] In practical applications, self-attention encoding values of each element in the target region are obtained through matrix operations. Specifically, the self-attention calculation layer determines the Q matrix, K matrix, and V matrix corresponding to the target region. Then, the Q matrix is multiplied pointwise with the K matrix, and the result of the pointwise multiplication is divided by a preset scale scalar, followed by normalization processing and softmax processing. The processed result is multiplied with the V matrix to obtain the self-attention encoding values of each point in the target region.

[0166] Step B2: The self-attention calculation layer performs convolution on the encoding matrix.

[0167] Among them, the convolved encoding matrix represents the position features of each pixel point in the target image; the position features include local position features or global position features.

[0168] It should be noted that performing convolution on the encoding matrix is to obtain the features representing the local positions and / or global positions of each pixel point in the target image, so as to more fully express the position features of each pixel point.

[0169] Step B3: The self-attention calculation layer concatenates the convolved encoding matrix with the encoding matrix to obtain a concatenated matrix.

[0170] Step B4: The self-attention calculation layer fuses the concatenated matrices corresponding to each sliding window to obtain the output matrix of this module.

[0171] As can be seen from the above description, each sliding window corresponds to a concatenated matrix, and the self-attention calculation layer can fuse the concatenated matrices corresponding to each sliding window to obtain the output matrix of this module.

[0172] It should be noted that the existing self-attention calculation encodes a matrix to obtain an encoding matrix, and then uses the encoding matrix as the output matrix.

[0173] In this specification, convolution is performed on the encoding matrix obtained by each sliding window to obtain the convolution result representing the global positions and / or local positions of each pixel point in the image. Then, the convolution results corresponding to each sliding window and the encoding matrix are concatenated to obtain the concatenated matrices of each sliding window, and the concatenated matrices of each sliding window are fused to obtain the output matrix. Through the self-attention calculation of multiple sliding windows, various correlation relationships of image pixels are represented, and through the convolution of the encoding matrix, the local positions and / or global positions of image pixels are represented. Therefore, this output matrix represents richer image information, which is beneficial to enhancing the accuracy of the entire model's feature extraction.

[0174] Step 106: Obtain the features of the target image output by the neural network-based encoding and decoding model.

[0175] In implementation, in an optional implementation manner, when the neural network-based encoding and decoding model includes a first module and at least one second module, the self-attention calculation layer of the last second module determines the image features of the target image from the output matrix and outputs them. The electronic device can obtain the image features of the target image output by the neural network-based encoding and decoding model.

[0176] Certainly, in another optional implementation manner, when the neural network-based encoding and decoding model only includes the first module, the self-attention calculation layer of the first module determines the feature vector of the target image from the output matrix and outputs it. The electronic device can obtain the feature vector of the target image output by the neural network-based encoding and decoding model.

[0177] In addition, in the embodiments of this specification, after extracting the image features of the target image, the electronic device can perform various image processing tasks based on the image features of the target image. For example, the electronic device can perform image classification based on the image features of the target image. For example, if there is a bag in the image, then determine the color type, style type, etc. of the bag.

[0178] Certainly, the electronic device can perform image segmentation, etc. based on the image features of the target image. Here, only an exemplary description of the image processing tasks is given, and no specific limitation is imposed on them.

[0179] As can be seen from the above description, first, the neural network-based encoding and decoding model in this specification uses multiple sliding windows for self-attention encoding. Since multiple sliding windows correspond to multiple different image feature extraction dimensions, the neural network-based encoding and decoding model in this specification can extract image features from multiple dimensions, enabling the image features to describe the image from multiple dimensions and making the description of the image by the image features more accurate.

[0180] Second, after the self-attention calculation layer in this specification calculates the encoding matrix, it performs convolution on the encoding matrix and splices the convolved encoding matrix with the encoding matrix to obtain the output matrix of this module. On the one hand, the convolved encoding matrix can represent the position features of various dimensions of each pixel point in the image, making the expression of the position features more abundant. On the other hand, through the self-attention calculation of multiple sliding windows, various correlation relationships of the image pixels are represented, and through the convolution of the encoding matrix, the local position and / or global position of the image pixels are represented. Therefore, this output matrix represents more abundant image information, which is beneficial to enhancing the accuracy of the entire model's feature extraction.

[0181] Third, the first convolutional layer of the first module performs consecutive convolutions on the target image using cascaded convolutional kernels with smaller sizes. Since the size of the convolutional kernel used in each convolution is small and subsequent multiple convolutions are performed on the basis of one convolution, the convolution granularity is finer, so that this convolution method can enhance the ability of the image to extract features.

[0182] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Please refer to Figure 10 , at the hardware level, the device includes a processor 1002, an internal bus 1004, a network interface 1006, a memory 1008, and a non-volatile memory 1010. Of course, there may also be other hardware required for other services. One or more embodiments of this specification can be implemented based on software. For example, the processor 1002 reads the corresponding computer program from the non-volatile memory 1010 into the memory 1008 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.

[0183] Please refer to Figure 11 , the image feature extraction device based on the self-attention mechanism can be applied to a device such as Figure 10 shown to implement the technical solution of this specification. Among them, the image feature extraction device based on the self-attention mechanism may include:

[0184] A first acquisition unit 1101, configured to acquire a target image;

[0185] An input unit 1102, configured to input the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image; wherein, various sliding windows correspond to different image feature extraction dimensions;

[0186] A second acquisition unit 1103, configured to acquire the feature vector of the target image output by the neural network-based encoding and decoding model.

[0187] Optionally, the neural network-based encoding and decoding model includes a first module; the first module includes: a first convolutional layer and a self-attention calculation layer;

[0188] When performing self-attention encoding on the target image based on multiple supported sliding windows to obtain the feature vector of the target image, the first convolutional layer in the first module is used to perform convolution on the target image to obtain the initial feature matrix of the target image, and the initial feature matrix is used as the input matrix and input into the self-attention calculation layer of this module;

[0189] The self-attention calculation layer of the first module is used to perform self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, and determine the feature vector of the target image according to the output matrix of the first module.

[0190] Optionally, the neural network-based encoding and decoding model further includes: at least one cascaded second module; the second module includes a second convolutional layer and a self-attention calculation layer;

[0191] When determining the feature vector of the target image according to the output matrix of the first module, the self-attention calculation layer of the first module is used to input the output matrix of this first module into the second module at the first position;

[0192] The second convolutional layer of each second module is used to perform convolution on the output matrix output by the previous module, and use the convolution result as the input matrix and input into the self-attention calculation layer of this second module;

[0193] The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module;

[0194] The self-attention calculation layer of the last second module is used to use the obtained output matrix as the feature vector of the target image.

[0195] Optionally, when the self-attention calculation layer performs self-attention encoding on the input matrix based on multiple supported sliding windows to obtain the output matrix of this module, for each type of sliding window, it is used to perform self-attention encoding on the input matrix based on this type of sliding window to obtain an encoded matrix; perform convolution on the encoded matrix; wherein, the convolved encoded matrix represents the position features of each pixel point of the target image; the position features include local position features and / or global position features; splice the convolved encoded matrix with the encoded matrix to obtain a spliced matrix; fuse the spliced matrices corresponding to each type of sliding window to obtain the output matrix of this module.

[0196] Optionally, when the self-attention calculation layer performs attention encoding on the input matrix based on the sliding window to obtain an encoded matrix, it is configured to slide the sliding window on the input matrix according to a preset step size; for each sliding operation, determine the target area circled by the sliding window on the input matrix, calculate the correlation degree scores between each element in the target area and this element and other elements in the target area respectively, obtain the self-attention encoded value corresponding to this element based on the correlation degree scores, and update the value of this element in the input matrix to the self-attention encoded value to obtain the encoded matrix.

[0197] Optionally, the first convolutional layer in the first module, when performing convolution on the target image to obtain the initial feature matrix of the target image, is configured to perform at least one continuous convolution process on the target image to obtain the initial feature matrix of the target image; wherein, the size of the convolution kernel corresponding to each convolution process is less than or equal to a preset threshold.

[0198] Optionally, the multiple sliding windows include: a first sliding window for extracting image features from the global dimension of the image and a second sliding window for extracting image features from the local dimension of the image.

[0199] Optionally, each sliding window includes: at least one effective calculation area; when there are multiple effective calculation areas, the multiple effective calculation areas are arranged at intervals in the sliding window;

[0200] When the self-attention calculation layer determines the target area circled by the sliding window on the input matrix, it is configured to use the area circled by the effective area of the sliding window on the input matrix as the target area.

[0201] Optionally, the neural network-based encoding and decoding model is a Transformer model that supports the self-attention mechanism.

[0202] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0203] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0204] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0205] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0206] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0207] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0208] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0209] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0210] The above description is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. An image feature extraction method based on self-attention mechanism, comprising: Obtaining a target image; Inputting the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image; wherein, various sliding windows correspond to different image feature extraction dimensions; Obtaining the feature vector of the target image output by the neural network-based encoding and decoding model.

2. The method according to claim 1, wherein the neural network-based encoding and decoding model comprises a first module; the first module comprises: The first convolutional layer and the self-attention calculation layer; Performing self-attention encoding on the target image based on a variety of supported sliding windows to obtain a feature vector of the target image, comprising: The first convolutional layer in the first module performs convolution on the target image to obtain an initial feature matrix of the target image, and takes the initial feature matrix as an input matrix and inputs it into the self-attention calculation layer of this module; The self-attention calculation layer of the first module performs self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain an output matrix of this module, and determines the feature vector of the target image according to the output matrix of the first module.

3. The method according to claim 2, wherein the neural network-based encoding and decoding model further comprises: At least one cascaded second module; The second module includes a second convolutional layer and a self-attention calculation layer; Determining the feature vector of the target image according to the output matrix of the first module, comprising: The self-attention calculation layer of the first module inputs the output matrix of the first module into the first second module in the front; The second convolutional layer of each second module performs convolution on the output matrix output by the previous module, and takes the convolution result as an input matrix and inputs it into the self-attention calculation layer of this second module; The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain an output matrix of this module; The self-attention calculation layer of the last second module takes the obtained output matrix as the feature vector of the target image.

4. According to the method described in claim 2 or 3, performing self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain an output matrix of this module, comprising: For each kind of sliding window, performing self-attention encoding on the input matrix based on this kind of sliding window to obtain an encoded matrix; Performing convolution on the encoded matrix; wherein, the convolved encoded matrix represents the position features of each pixel point of the target image; the position features include local position features and / or global position features; Concatenating the convolved encoded matrix with the encoded matrix to obtain a concatenated matrix; Fusing the concatenated matrices corresponding to each kind of sliding window to obtain an output matrix of this module.

5. According to the method described in claim 4, performing attention encoding on the input matrix based on this kind of sliding window to obtain an encoded matrix, comprising: Sliding this sliding window on the input matrix according to a preset step size; For each sliding operation, determine the target area circled by the sliding window on the input matrix. For each element in the target area, calculate the correlation degree scores between the element and itself, as well as other elements in the target area, and obtain the self-attention encoding value corresponding to the element based on the correlation degree scores. Then, update the value of the element in the input matrix to the self-attention encoding value to obtain the encoded matrix.

6. The method according to claim 2, wherein the first convolutional layer in the first module performs convolution on the target image to obtain an initial feature matrix of the target image, including: The first convolutional layer performs at least one continuous convolution process on the target image to obtain an initial feature matrix of the target image; wherein, the size of the convolutional kernel corresponding to each convolution process is less than or equal to a preset threshold.

7. The method according to claim 1, wherein the multiple sliding windows include: A first sliding window for extracting image features from the global dimension of the image, and a second sliding window for extracting image features from the local dimension of the image.

8. The method according to claim 5, wherein each sliding window comprises: At least one effective calculation area; when there are multiple effective calculation areas, the multiple effective calculation areas are arranged at intervals in the sliding window; The determining the target area circled by the sliding window on the input matrix includes: Taking the area circled by the effective area of the sliding window on the input matrix as the target area.

9. The method according to claim 1, wherein the neural network-based encoding and decoding model is a Transformer model supporting the self-attention mechanism.

10. An image feature extraction device based on the self-attention mechanism, comprising: A first acquisition unit for acquiring a target image; An input unit for inputting the target image into a trained neural network-based encoding and decoding model, so that the neural network-based encoding and decoding model performs self-attention encoding on the target image based on multiple supported sliding windows to obtain a feature vector of the target image; wherein, each sliding window corresponds to a different image feature extraction dimension; A second acquisition unit for acquiring the feature vector of the target image output by the neural network-based encoding and decoding model.

11. The apparatus according to claim 10, wherein the neural network-based encoding and decoding model comprises a first module; the first module comprises: A first convolutional layer and a self-attention calculation layer; When performing self-attention encoding on the target image based on multiple supported sliding windows to obtain a feature vector of the target image, the first convolutional layer in the first module performs convolution on the target image to obtain an initial feature matrix of the target image, and inputs the initial feature matrix as an input matrix into the self-attention calculation layer of this module; The self-attention calculation layer of the first module is used to perform self-attention encoding on the input matrix based on multiple supported sliding windows to obtain an output matrix of this module, and determine the feature vector of the target image according to the output matrix of the first module.

12. The device according to claim 11, wherein the neural network-based encoding and decoding model further comprises: At least one cascaded second module; The second module includes a second convolutional layer and a self-attention calculation layer; When determining the feature vector of the target image according to the output matrix of the first module, the self-attention calculation layer of the first module is used to input the output matrix of the first module into the first second module in the queue; The second convolutional layer of each second module is used to perform convolution on the output matrix output by the previous module, and use the convolution result as the input matrix to the self-attention calculation layer of this second module; The self-attention calculation layer of the second module performs self-attention encoding on the input matrix based on a variety of supported sliding windows to obtain the output matrix of this module; The self-attention calculation layer of the last second module is used to use the obtained output matrix as the feature vector of the target image.

13. An electronic device, comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the method according to any one of claims 1-9 by running the executable instructions.

14. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are realized.

Citation Information

Patent Citations

  • Remote sensing image target extraction method based on deep neural network

    CN112712500A

  • Image tag identification method and device, electronic equipment and readable storage medium

    CN113627466A