Image feature extraction method, device, electronic device and storage medium
By performing matrix transformation and window division on the image, and using the self-attention module for feature extraction, the problem of large calculation volume and difficult to guarantee accuracy of the Transformer model is solved, and efficient and accurate image feature extraction is achieved.
Patent Information
- Application Number
- CN202111604195.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The Transformer model has a large amount of calculation and low calculation efficiency in the process of image feature extraction, which makes it difficult to guarantee the calculation accuracy and cannot meet the needs of image feature extraction.
By performing matrix transformation processing on the image to be processed, a preset number of feature maps is obtained and divided into multiple windows according to the preset window size. Then, these windows are input to the feature extraction layer, and feature extraction is performed using the self-attention module, and finally the feature extraction result of the image is obtained.
Through window division and self-attention module, this method can extract image features more efficiently, improve the extraction accuracy, and use less training data during the training process to achieve a model that meets the accuracy requirements.
Smart Images

Figure CN114419325B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to an image feature extraction method, device, electronic device and storage medium. Background Art
[0002] Image feature extraction is the most basic problem in the field of computer vision. It refers to the process of extracting the feature information of the input image through analysis and learning of the input image, and then classifying the input image according to the feature information. In the process of image feature extraction, feature extraction of the input image is very important. An efficient image feature extraction network can effectively improve the efficiency of downstream image feature extraction.
[0003] In the prior art, the Transformer model can be used to extract feature information of the input image. The Transformer model is a deep learning model that uses an Encoder-Decoder framework, wherein the Encoder includes a two-layer network structure of Self-Attention and Feed Forward Neural Network, and the Decoder includes a three-layer network structure of Self-Attention, Encoder-Decoder Attention and Feed Forward Neural Network.
[0004] However, the training of the Transformer model requires a large amount of training data and continuous iterative adjustment of the model parameters to achieve high-precision feature extraction. This will lead to a large amount of computation during the training of the Transformer model and low computational efficiency. In other words, when the amount of training data is considerable, the computational accuracy of the Transformer model is difficult to guarantee. Summary of the invention
[0005] The present disclosure provides an image feature extraction method, device, electronic device and storage medium to at least solve the problem that the image feature extraction process in the related art has a large amount of calculation, low calculation efficiency, and difficulty in ensuring calculation accuracy, and it is increasingly difficult to meet the needs of image feature extraction. The technical solution of the present disclosure is as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an image feature extraction method, comprising:
[0007] Get the image to be processed;
[0008] Performing matrix transformation on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels;
[0009] Divide each feature map into multiple windows according to the preset window size;
[0010] Input the multiple windows into a first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on the pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows;
[0011] The obtained feature block is input to the next feature extraction layer for processing to obtain the feature block corresponding to the next feature extraction layer, and the step of inputting the obtained feature block to the next feature extraction layer for processing is returned, until the obtained feature block is input to the last feature extraction layer for processing to obtain the feature extraction result of the image to be processed.
[0012] Optionally, performing matrix transformation processing on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed includes:
[0013] Adjusting the size of the image to be processed to obtain an intermediate image;
[0014] The intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
[0015] Optionally, the feature block corresponding to the first feature extraction layer includes a first calculation result, and the first self-attention module adopts the following steps to perform a first self-attention operation on the pixels in each window respectively to obtain a first calculation result corresponding to each window:
[0016] For each window, the pixels in the window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs as the first calculation result corresponding to each window.
[0017] Optionally, the feature block corresponding to the first feature extraction layer includes a second calculation result, and the second self-attention module performs a second self-attention operation on the multiple windows to obtain a second calculation result by:
[0018] Upsample all pixels in each window to obtain the sampling points corresponding to each window;
[0019] The sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs within the multiple windows as the second calculation result.
[0020] Optionally, the step of inputting the obtained feature block to a next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer includes:
[0021] Reorganize the obtained feature blocks to obtain a reorganized feature block, wherein the size of the feature block is N times that of the reorganized feature block, the number of channels of the reorganized feature block is N times that of the feature block, and N is an integer greater than 1;
[0022] Inputting the recombined feature block into a multi-layer perceptron for mapping processing to obtain a mapping feature block;
[0023] The mapped feature block is input to the next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer.
[0024] Optionally, each feature extraction layer includes a plurality of first self-attention modules and a plurality of second self-attention modules, and the plurality of windows are input into the first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, including:
[0025] Inputting the multiple windows into multiple first self-attention modules in the first feature extraction layer in sequence, performing a first self-attention operation on the pixels in each window, and obtaining a first processing result corresponding to each first self-attention module;
[0026] Inputting the multiple windows sequentially into multiple second self-attention modules in the first feature extraction layer, respectively performing a second self-attention operation on the multiple windows, and obtaining a second processing result corresponding to each second self-attention module;
[0027] The obtained first processing result and the obtained second processing result are used as feature blocks corresponding to the first feature extraction layer.
[0028] According to a second aspect of an embodiment of the present disclosure, there is provided an image feature extraction device, comprising:
[0029] An acquisition unit, configured to acquire an image to be processed;
[0030] A transformation unit is configured to perform matrix transformation processing on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels;
[0031] A division unit is configured to divide each feature map into a plurality of windows according to a preset window size;
[0032] An extraction unit is configured to input the multiple windows into a first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows;
[0033] The processing unit is configured to execute the step of inputting the obtained feature block to the next feature extraction layer for processing, obtaining the feature block corresponding to the next feature extraction layer, returning to the step of inputting the obtained feature block to the next feature extraction layer for processing, until the obtained feature block is input to the last feature extraction layer for processing, thereby obtaining the feature extraction result of the image to be processed.
[0034] Optionally, the transformation unit is configured to perform:
[0035] Adjusting the size of the image to be processed to obtain an intermediate image;
[0036] The intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
[0037] Optionally, the feature block corresponding to the first feature extraction layer includes a first calculation result, and the extraction unit is configured to execute:
[0038] For each window, the pixels in the window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs as the first calculation result corresponding to each window.
[0039] Optionally, the feature block corresponding to the first feature extraction layer includes a second calculation result, and the extraction unit is configured to execute:
[0040] Upsample all pixels in each window to obtain the sampling points corresponding to each window;
[0041] The sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs within the multiple windows as the second calculation result.
[0042] Optionally, the processing unit is configured to execute:
[0043] Reorganize the obtained feature blocks to obtain a reorganized feature block, wherein the size of the feature block is N times that of the reorganized feature block, the number of channels of the reorganized feature block is N times that of the feature block, and N is an integer greater than 1;
[0044] Inputting the recombined feature block into a multi-layer perceptron for mapping processing to obtain a mapping feature block;
[0045] The mapped feature block is input to the next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer.
[0046] Optionally, each feature extraction layer includes a plurality of first self-attention modules and a plurality of second self-attention modules, and the extraction unit is configured to perform:
[0047] Inputting the multiple windows into multiple first self-attention modules in the first feature extraction layer in sequence, performing a first self-attention operation on the pixels in each window, and obtaining a first processing result corresponding to each first self-attention module;
[0048] Inputting the multiple windows sequentially into multiple second self-attention modules in the first feature extraction layer, respectively performing a second self-attention operation on the multiple windows, and obtaining a second processing result corresponding to each second self-attention module;
[0049] The obtained first processing result and the obtained second processing result are used as feature blocks corresponding to the first feature extraction layer.
[0050] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device for extracting image features, including:
[0051] processor;
[0052] a memory for storing instructions executable by the processor;
[0053] Wherein, the processor is configured to execute the instructions to implement the image feature extraction method described in the first item above.
[0054] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an image feature extraction electronic device, the image feature extraction electronic device is enabled to perform the image feature extraction method described in the first item above.
[0055] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the image feature extraction method described in the first item above.
[0056] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0057] Acquire an image to be processed; perform matrix transformation on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels; divide each feature map into multiple windows according to a preset window size; input the multiple windows into a first feature extraction layer for processing to obtain a feature block corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on the pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows; input the obtained feature block into a next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer, and return to the step of inputting the obtained feature block into the next feature extraction layer for processing, until the obtained feature block is input into the last feature extraction layer for processing to obtain a feature extraction result of the image to be processed.
[0058] In this way, the feature map of the image to be processed is divided into windows. On the one hand, for each window, each pixel in the window can be processed by the first self-attention module, and the first calculation result obtained can reflect the local attention characteristics between the pixels in each window. On the other hand, each window is processed by the second self-attention module, and the second calculation result obtained can reflect the global attention characteristics between the windows. Therefore, this scheme has higher accuracy in extracting image features. Correspondingly, less training data can be used during training to obtain a model that meets the accuracy requirements.
[0059] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0061] Figure 1 The figure is a flow chart of a method for extracting image features according to an exemplary embodiment.
[0062] Figure 2 This is the logical diagram of the first self-attention module.
[0063] Figure 3 This is the logical diagram of the second self-attention module.
[0064] Figure 4It is a logical schematic diagram of an image feature extraction method according to an exemplary embodiment.
[0065] Figure 5 It is a block diagram of an image feature extraction device according to an exemplary embodiment.
[0066] Figure 6 A block diagram of an electronic device for image feature extraction is shown according to an exemplary embodiment.
[0067] Figure 7 The invention is a block diagram of a device for extracting image features according to an exemplary embodiment. DETAILED DESCRIPTION
[0068] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0069] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0070] Figure 1 is a flow chart of an image feature extraction method according to an exemplary embodiment. Figure 1 As shown, the image feature extraction method includes the following steps.
[0071] In step S11, an image to be processed is obtained.
[0072] The classification, recognition, transformation and other processing of images usually need to be based on the features of the image, so the feature extraction of the image is particularly important. In the present disclosure, the image to be processed is the image that needs to be feature extracted, which can be any image format, such as RGB image, etc. The size of the image to be processed is not limited. Generally, the larger the size of the image to be processed, the more feature information it contains.
[0073] In step S12, matrix transformation is performed on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels.
[0074] In this step, the image to be processed is subjected to matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed. Specifically, first, the size of the image to be processed is adjusted to obtain an intermediate image, and then the intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
[0075] Among them, adjusting the size of the image to be processed can be achieved by performing a Resize operation on the image to be processed. For example, when the image to be processed is smaller than the required size, the size of the image to be processed can be increased by interpolation, or when the image to be processed is larger than the required size, the size of the image to be processed can be reduced by downsampling, or the image to be processed can be directly cropped, etc., without specific limitation.
[0076] The size of the intermediate image can be determined according to the requirements of the convolutional mapping layer for the input image, so that the intermediate image obtained better meets the requirements of the convolutional mapping layer, and the subsequent processing results are more accurate. For example, after adjusting the size of the image to be processed, a 224*224*3 intermediate image can be obtained. Then, the intermediate image is input to the convolutional mapping layer, and the intermediate image can be mapped into a 56*56*98 feature map, which is equivalent to extracting 98 56*56 feature maps.
[0077] It can be understood that each feature map includes local feature information of the image to be processed at the corresponding position. Therefore, by processing the feature map, the feature information of the image to be processed can be extracted.
[0078] In step S13, each feature map is divided into multiple windows according to a preset window size.
[0079] In this step, each feature map is traversed according to a preset window size, and each feature map can be divided into multiple windows, and these windows do not overlap.
[0080] For example, continuing with the above example, if the preset window size is 7*7, the window size represents the number of pixels, then a 56*56 feature map will be divided into 64 windows, each window including 7*7=49 pixels (tokens).
[0081] In step S14, multiple windows are input into the first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on the pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between multiple windows.
[0082] In this step, the first feature extraction layer includes two parts, namely the first self-attention module and the second self-attention module, that is, inputting multiple windows into the first feature extraction layer for processing is also divided into two parts. Specifically, multiple windows can be sequentially input into multiple first self-attention modules in the first feature extraction layer, and the first self-attention operation is performed on the pixels in each window to obtain the first processing result corresponding to each first self-attention module; then, multiple windows are sequentially input into multiple second self-attention modules in the first feature extraction layer to perform the second self-attention operation between the multiple windows, respectively, to obtain the second processing result corresponding to each second self-attention module; and then, the obtained first processing result and the obtained second processing result are used as the feature blocks corresponding to the first feature extraction layer.
[0083] In this way, using the first self-attention module and the second self-attention module, self-attention operations can be performed on the pixels in each window and the pixels between windows respectively, and the local attention features between the pixels in each window and the global attention features between the windows are analyzed, so that the final feature extraction result can take into account both the local and the global, and the obtained image feature extraction result is more accurate.
[0084] Among them, the first self-attention module is used to extract local attention features between pixels in each window. Specifically, for each window, the pixels in the window can be input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs, as the first calculation result corresponding to each window.
[0085] like Figure 2 As shown, it is a logical schematic diagram of the first self-attention module. The 4*4 square matrix on the far left represents the feature map, where each square represents a pixel in the feature map. If the preset window size is 2*2, then each feature map can be divided into 4 windows, each window includes 4 pixels. In this way, the pixels in each window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs, and then the first calculation result of the feature map can be obtained. Since each pixel is input into the first self-attention module for correlation analysis, the size of the first calculation result of the feature map is 4*4.
[0086] The second self-attention module is used to extract the global attention features between each window. Specifically, all the pixels in each window can be upsampled first to obtain the sampling points corresponding to each window. Then, the sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs in multiple windows as the second calculation result.
[0087] like Figure 3 As shown, it is a logical schematic diagram of the second self-attention module. The 4*4 square matrix on the far left represents the feature map, where each square represents a pixel in the feature map. If the preset window size is 2*2, then each feature map can be divided into 4 windows, each window includes 4 pixels. First, the pixels in each window are upsampled to obtain the sampling points corresponding to the 4 windows, and then the second self-attention operation is performed between each window to obtain the second calculation result. Furthermore, the size of the second calculation result can be adjusted by downsampling and other methods, so that the size of the adjusted second calculation result is the same as the size of the first calculation result. In this way, it is convenient to stack the second calculation result and the first calculation result to obtain the feature block of the image to be processed.
[0088] In the present disclosure, the number of layers of the first self-attention module and the second self-attention module in each feature extraction layer is equal, but the number of layers of the first self-attention module and the second self-attention module between each feature extraction layer may be different. For example, the number of layers of the self-attention module in the first feature extraction layer may be 2, indicating that the first feature extraction layer has two layers of first self-attention modules and two layers of second self-attention modules, and the number of layers of the self-attention module in the second feature extraction layer may be 6, indicating that the second feature extraction layer has six layers of first self-attention modules and six layers of second self-attention modules.
[0089] In step S15, the obtained feature block is input to the next feature extraction layer for processing to obtain the feature block corresponding to the next feature extraction layer, and the step of inputting the obtained feature block to the next feature extraction layer for processing is returned, until the obtained feature block is input to the last feature extraction layer for processing to obtain the feature extraction result of the image to be processed.
[0090] In one implementation, before the obtained feature block is input to the next feature extraction layer for processing, a space-to-depth transformation operation may be performed on the feature block, which may specifically include the following steps:
[0091] First, the obtained feature blocks are reorganized to obtain reorganized feature blocks. The size of the feature blocks is N times that of the reorganized feature blocks, and the number of channels of the reorganized feature blocks is N times that of the feature blocks, where N is an integer greater than 1. Then, the reorganized feature blocks are input into a multilayer perceptron for mapping processing to obtain mapped feature blocks. Furthermore, the mapped feature blocks are input into the next feature extraction layer for processing to obtain new feature blocks.
[0092] For example, the value of N can be 4, that is, the size of the feature block is first reduced to one-fourth of the original size, the number of channels is quadrupled, and then the number of channels is halved through a multi-layer perceptron to achieve feature extraction of the feature block. In this way, the impact of background and many meaningless information on the feature extraction results can be further reduced, and the key information of the image to be processed is extracted step by step to obtain the final image feature extraction result.
[0093] In one implementation, before the obtained feature block is input to the next feature extraction layer for processing, the feature block can be first input to the pooling layer connected to the first feature extraction layer for pooling operation, and then the feature block after the pooling operation is input to the next feature extraction layer for processing, and so on, that is, a pooling layer can be set between each feature extraction layer for pooling operation, and the pooling operation is to downsample the feature blocks corresponding to each feature extraction layer, and maximum pooling, average pooling, overlapping pooling and pyramid pooling can be adopted, and the present disclosure does not limit this. Through the pooling operation, the dimension of the feature block can be reduced, thereby effectively reducing the model parameters and preventing overfitting.
[0094] like Figure 4 , which is a logical schematic diagram of an image feature extraction method according to an exemplary embodiment, wherein the dimension of the image to be processed is H×W×3, and the image feature extraction model specifically adopts a Transformer model, which includes 4 feature extraction layers, each of which includes a first self-attention module and a second self-attention module, and the dimensions of the output feature blocks are respectively and The number of layers of the first self-attention module and the second self-attention module in each layer are 2, 2, 6, and 2 respectively.
[0095] Specifically, if the dimension of the image to be classified is 224×224×3, the image to be processed is input to the convolutional mapping layer, and the mapping dimension of the convolutional mapping layer is set to C=98, then 98 56*56 feature maps can be extracted, and then the windows are divided into 7*7 sizes, and the obtained windows are processed by the first feature extraction layer, the second feature extraction layer, the third feature extraction layer and the fourth feature extraction layer in turn, and finally the feature extraction result of the image to be processed is obtained. Among them, a pooling layer can be set between each layer of feature extraction layer for pooling operation.
[0096] In the present disclosure, the image feature extraction results can be applied to different fields such as target detection, image semantic segmentation and video classification. For example, the output feature extraction results can be input into MLP (Multilayer Perceptron) for processing, and the obtained transformation matrix size is 784*1000, which is the predicted score of the image to be processed in 1000 categories. Finally, according to the predicted score of the image to be processed in 1000 categories, the classification result of the image to be processed can be obtained.
[0097] From the above, it can be seen that the technical solution provided by the embodiment of the present disclosure divides the feature map of the image to be processed in the form of windows. On the one hand, for each window, the first self-attention module can be used to process each pixel in the window, and the first calculation result obtained can reflect the local attention characteristics between the pixels in each window. On the other hand, the second self-attention module is used to process each window, and the second calculation result obtained can reflect the global attention characteristics between each window. Therefore, this solution has higher accuracy in extracting image features. Correspondingly, during training, less training data can be used to obtain a model that meets the accuracy requirements.
[0098] Moreover, this solution is an end-to-end architecture with low complexity and low parameters. It is easy to migrate to various medium and low performance devices, so it can also be easily embedded in various business scenarios for practical application.
[0099] Figure 5 is a block diagram of an image feature extraction device according to an exemplary embodiment, the device comprising:
[0100] An acquisition unit 201 is configured to acquire an image to be processed;
[0101] The transformation unit 202 is configured to perform matrix transformation processing on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels;
[0102] A division unit 203 is configured to divide each feature map into a plurality of windows according to a preset window size;
[0103] The extraction unit 204 is configured to input the multiple windows into a first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows;
[0104] The processing unit 205 is configured to execute the step of inputting the obtained feature block to the next feature extraction layer for processing, obtaining the feature block corresponding to the next feature extraction layer, returning to the step of inputting the obtained feature block to the next feature extraction layer for processing, until the obtained feature block is input to the last feature extraction layer for processing, thereby obtaining the feature extraction result of the image to be processed.
[0105] In one implementation, the transform unit 202 is configured to perform:
[0106] Adjusting the size of the image to be processed to obtain an intermediate image;
[0107] The intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
[0108] In one implementation, the feature block corresponding to the first feature extraction layer includes a first calculation result, and the extraction unit 204 is configured to execute:
[0109] For each window, the pixels in the window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs as the first calculation result corresponding to each window.
[0110] In one implementation, the feature block corresponding to the first feature extraction layer includes the second calculation result, and the extraction unit 204 is configured to execute:
[0111] Upsample all pixels in each window to obtain the sampling points corresponding to each window;
[0112] The sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs within the multiple windows as the second calculation result.
[0113] In one implementation, the processing unit 205 is configured to execute:
[0114] Reorganize the obtained feature blocks to obtain a reorganized feature block, wherein the size of the feature block is N times that of the reorganized feature block, the number of channels of the reorganized feature block is N times that of the feature block, and N is an integer greater than 1;
[0115] Inputting the recombined feature block into a multi-layer perceptron for mapping processing to obtain a mapping feature block;
[0116] The mapped feature block is input to the next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer.
[0117] In one implementation, each feature extraction layer includes a plurality of first self-attention modules and a plurality of second self-attention modules, and the extraction unit 204 is configured to perform:
[0118] Inputting the multiple windows into multiple first self-attention modules in the first feature extraction layer in sequence, performing a first self-attention operation on the pixels in each window, and obtaining a first processing result corresponding to each first self-attention module;
[0119] Inputting the multiple windows sequentially into multiple second self-attention modules in the first feature extraction layer, respectively performing a second self-attention operation on the multiple windows, and obtaining a second processing result corresponding to each second self-attention module;
[0120] The obtained first processing result and the obtained second processing result are used as feature blocks corresponding to the first feature extraction layer.
[0121] From the above, it can be seen that the technical solution provided by the embodiment of the present disclosure divides the feature map of the image to be processed in the form of windows. On the one hand, for each window, the first self-attention module can be used to process each pixel in the window, and the first calculation result obtained can reflect the local attention characteristics between the pixels in each window. On the other hand, the second self-attention module is used to process each window, and the second calculation result obtained can reflect the global attention characteristics between each window. Therefore, this solution has higher accuracy in extracting image features. Correspondingly, during training, less training data can be used to obtain a model that meets the accuracy requirements.
[0122] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0123] Figure 6 The invention is a block diagram of an electronic device for image feature extraction according to an exemplary embodiment.
[0124] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of an electronic device to perform the above method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0125] In an exemplary embodiment, a computer program product is also provided, which, when executed on a computer, enables the computer to implement the above-mentioned method for extracting image features.
[0126] From the above, it can be seen that the technical solution provided by the embodiment of the present disclosure divides the feature map of the image to be processed in the form of windows. On the one hand, for each window, the first self-attention module can be used to process each pixel in the window, and the first calculation result obtained can reflect the local attention characteristics between the pixels in each window. On the other hand, the second self-attention module is used to process each window, and the second calculation result obtained can reflect the global attention characteristics between each window. Therefore, this solution has higher accuracy in extracting image features. Correspondingly, during training, less training data can be used to obtain a model that meets the accuracy requirements.
[0127] Figure 7 It is a block diagram of a device 800 for extracting image features according to an exemplary embodiment.
[0128] For example, apparatus 800 may be a mobile phone, a computer, a digital broadcast electronic device, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0129] Reference Figure 6 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0130] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0131] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0132] The power supply component 807 provides power to the various components of the device 800. The power supply component 807 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0133] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0134] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0135] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0136] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0137] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0138] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the methods described in the first and second aspects.
[0139] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the device 800 to complete the above method. Optionally, for example, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0140] In an exemplary embodiment, a computer program product containing instructions is also provided. When the computer program product is run on a computer, the computer is enabled to execute the image feature extraction method described in the first embodiment above.
[0141] From the above, it can be seen that the technical solution provided by the embodiment of the present disclosure divides the feature map of the image to be processed in the form of windows. On the one hand, for each window, the first self-attention module can be used to process each pixel in the window, and the first calculation result obtained can reflect the local attention characteristics between the pixels in each window. On the other hand, the second self-attention module is used to process each window, and the second calculation result obtained can reflect the global attention characteristics between each window. Therefore, this solution has higher accuracy in extracting image features. Correspondingly, during training, less training data can be used to obtain a model that meets the accuracy requirements.
[0142] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0143] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for extracting image features, characterized in that: include: Get the image to be processed; Performing matrix transformation on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels; Divide each feature map into multiple windows according to the preset window size; Input the multiple windows into a first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on the pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows; Input the obtained feature block to the next feature extraction layer for processing, obtain the feature block corresponding to the next feature extraction layer, return to the step of inputting the obtained feature block to the next feature extraction layer for processing, until the obtained feature block is input to the last feature extraction layer for processing, and obtain the feature extraction result of the image to be processed; The step of inputting the obtained feature block to the next feature extraction layer for processing to obtain the feature block corresponding to the next feature extraction layer includes: Reorganize the obtained feature blocks to obtain a reorganized feature block, wherein the size of the feature block is N times that of the reorganized feature block, the number of channels of the reorganized feature block is N times that of the feature block, and N is an integer greater than 1; Inputting the recombined feature block into a multi-layer perceptron for mapping processing to obtain a mapping feature block; Inputting the mapped feature block to the next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer; Or, the step of inputting the obtained feature block to a next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer includes: A pooling layer is set between each feature extraction layer, and the feature block is input into the pooling layer for pooling operation before being input into the next feature extraction layer for processing; The feature blocks after the pooling operation are implanted into the next feature extraction layer for processing to obtain the feature blocks corresponding to the next feature extraction layer.
2. The image feature extraction method according to claim 1, characterized in that: The step of performing matrix transformation on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed includes: Adjusting the size of the image to be processed to obtain an intermediate image; The intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
3. The image feature extraction method according to claim 1, characterized in that: The feature block corresponding to the first feature extraction layer includes a first calculation result, and the first self-attention module adopts the following steps to perform a first self-attention operation on the pixels in each window respectively: For each window, the pixels in the window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs as the first calculation result corresponding to each window.
4. The image feature extraction method according to claim 1, characterized in that: The feature block corresponding to the first feature extraction layer includes a second calculation result, and the second self-attention module performs a second self-attention operation on the multiple windows by adopting the following steps: Upsample all pixels in each window to obtain the sampling points corresponding to each window; The sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs within the multiple windows as the second calculation result.
5. The image feature extraction method according to claim 1, characterized in that: Each feature extraction layer includes a plurality of first self-attention modules and a plurality of second self-attention modules. The plurality of windows are input into the first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, including: Inputting the multiple windows into multiple first self-attention modules in the first feature extraction layer in sequence, performing a first self-attention operation on the pixels in each window, and obtaining a first processing result corresponding to each first self-attention module; Inputting the multiple windows sequentially into multiple second self-attention modules in the first feature extraction layer, respectively performing a second self-attention operation on the multiple windows, and obtaining a second processing result corresponding to each second self-attention module; The obtained first processing result and the obtained second processing result are used as feature blocks corresponding to the first feature extraction layer.
6. An image feature extraction device, characterized in that: include: An acquisition unit, configured to acquire an image to be processed; A transformation unit is configured to perform matrix transformation processing on the image to be processed to obtain a preset number of feature maps corresponding to the image to be processed, each feature map having the same number of pixels; A division unit is configured to divide each feature map into a plurality of windows according to a preset window size; An extraction unit is configured to input the multiple windows into a first feature extraction layer for processing to obtain feature blocks corresponding to the first feature extraction layer, wherein the first feature extraction layer includes a first self-attention module and a second self-attention module, the first self-attention module is used to perform a first self-attention operation on pixels in each window respectively, and the second self-attention module is used to perform a second self-attention operation between the multiple windows; a processing unit configured to execute the step of inputting the obtained feature block to a next feature extraction layer for processing, obtaining a feature block corresponding to the next feature extraction layer, returning to the step of inputting the obtained feature block to the next feature extraction layer for processing, until the obtained feature block is input to the last feature extraction layer for processing, thereby obtaining a feature extraction result of the image to be processed; Wherein, the processing unit is configured to execute: Reorganize the obtained feature blocks to obtain a reorganized feature block, wherein the size of the feature block is N times that of the reorganized feature block, the number of channels of the reorganized feature block is N times that of the feature block, and N is an integer greater than 1; Inputting the recombined feature block into a multi-layer perceptron for mapping processing to obtain a mapping feature block; Inputting the mapped feature block to the next feature extraction layer for processing to obtain a feature block corresponding to the next feature extraction layer; Or, the processing unit is configured to execute: A pooling layer is set between each feature extraction layer, and the feature block is input into the pooling layer for pooling operation before being input into the next feature extraction layer for processing; The feature blocks after the pooling operation are implanted into the next feature extraction layer for processing to obtain the feature blocks corresponding to the next feature extraction layer.
7. The image feature extraction device according to claim 6, characterized in that: The transformation unit is configured to perform: Adjusting the size of the image to be processed to obtain an intermediate image; The intermediate image is input into the convolution mapping layer for matrix transformation processing to obtain a preset number of feature maps corresponding to the image to be processed.
8. The image feature extraction device according to claim 6, characterized in that: The feature block corresponding to the first feature extraction layer includes a first calculation result, and the extraction unit is configured to execute: For each window, the pixels in the window are input into the first self-attention module for correlation analysis to obtain the weight of each pixel in the window to which it belongs as the first calculation result corresponding to each window.
9. The image feature extraction device according to claim 6, characterized in that: The feature block corresponding to the first feature extraction layer includes a second calculation result, and the extraction unit is configured to execute: Upsample all pixels in each window to obtain the sampling points corresponding to each window; The sampling points corresponding to each window are input into the second self-attention module for correlation analysis to obtain the weight of the window to which each sampling point belongs within the multiple windows as the second calculation result.
10. The image feature extraction device according to claim 6, characterized in that: Each feature extraction layer includes a plurality of first self-attention modules and a plurality of second self-attention modules, and the extraction unit is configured to perform: Inputting the multiple windows into multiple first self-attention modules in the first feature extraction layer in sequence, performing a first self-attention operation on the pixels in each window, and obtaining a first processing result corresponding to each first self-attention module; Inputting the multiple windows sequentially into multiple second self-attention modules in the first feature extraction layer, respectively performing a second self-attention operation on the multiple windows, and obtaining a second processing result corresponding to each second self-attention module; The obtained first processing result and the obtained second processing result are used as feature blocks corresponding to the first feature extraction layer.
11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image feature extraction method as described in any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device for image feature extraction, the electronic device for image feature extraction is enabled to perform the image feature extraction method as claimed in any one of claims 1 to 5.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the image feature extraction method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Medical image segmentation method and device and storage medium
CN109872306A
Target detection method based on scene level and region suggestion self-attention module
CN110516670A