Feature Extraction Method, Device, Equipment and Storage Medium for Images
Through the feature extraction method of convolutional neural network combined with channel and spatial attention mechanism, the problem of semantic information and position information separation in image feature extraction is solved, and the accuracy of target detection and the experience of autonomous driving are improved.
Patent Information
- Application Number
- CN202211219646.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Although the existing image feature extraction methods can provide rich semantic information, the position information is blurred, resulting in a decrease in the accuracy of target detection and the sense of autonomous driving experience.
The convolutional neural network model is used to combine channel attention mechanism and spatial attention mechanism to extract the image feature and fuse layer by layer to obtain the target feature map to ensure that the feature map contains rich semantic information and accurate position information.
It improves the accuracy of target detection and the experience of autonomous driving, and enhances the efficiency and accuracy of feature extraction, especially in complex driving scenarios.
Smart Images

Figure CN115620017B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular, to a method, apparatus, device, and storage medium for feature extraction of images. Background Art
[0002] In driving scenarios and some special scenarios such as industrial and mining areas with complex environments, there are quite high requirements for the real-time performance and accuracy of target detection. Feature extraction of images is the basis for various intensive computational downstream tasks such as target detection or image segmentation. The accuracy of the image feature extraction results directly affects the accuracy of subsequent downstream tasks such as target detection and image segmentation.
[0003] Currently, most methods for feature extraction of images obtain feature maps of images through convolutional neural networks. Although the feature maps obtained by this feature extraction method have rich semantic information, the position information is relatively vague, which affects the accuracy of target detection and thus affects the experience of autonomous driving. Summary of the Invention
[0004] This application provides a method, apparatus, device, and storage medium for feature extraction of images, which can ensure that the obtained feature maps include both rich semantic information and accurate position information, improve the accuracy of target detection, and enhance the experience of autonomous driving.
[0005] To achieve the above object, this application adopts the following technical solutions:
[0006] In the first aspect of the embodiments of this application, a method for feature extraction of images is provided. The method includes:
[0007] Obtain an image to be processed;
[0008] Perform convolution processing on the image using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4;
[0009] Perform feature extraction on each of the original feature maps from the second to the nth original feature maps using a preset channel attention mechanism model to obtain channel attention feature maps corresponding to each of the original feature maps from the second to the nth original feature maps;
[0010] Perform feature extraction on the nth original feature map using a preset spatial attention mechanism model to obtain a spatial attention feature map corresponding to the nth original feature map;
[0011] Perform layer-by-layer feature fusion based on the channel attention feature maps corresponding to the second to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the first original feature map to obtain a target feature map of the image;
[0012] Perform object detection in the autonomous driving scenario based on the target feature map.
[0013] In one embodiment, perform layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map to obtain the target feature map of the image, including:
[0014] Determine the fusion map corresponding to the nth original feature map based on the spatial attention feature map and the channel attention feature map corresponding to the nth original feature map;
[0015] Perform layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps, and the 1st original feature map to obtain the target feature map of the image.
[0016] In one embodiment, perform layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps, and the 1st original feature map to obtain the target feature map of the image, including:
[0017] Perform layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps to obtain the fusion map corresponding to the 2nd original feature map;
[0018] Obtain the target feature map of the image based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map.
[0019] In one embodiment, perform layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps to obtain the fusion map corresponding to the 2nd original feature map, including:
[0020] Start from i = n - 1 and perform at least one fusion processing process until i = 2 to obtain the fusion map corresponding to the 2nd original feature map, where i is an integer from (n - 1) to 2;
[0021] Wherein, the mth fusion processing process includes: obtaining the fusion map corresponding to the ith original feature map according to the channel attention feature map corresponding to the ith original feature map, the spatial attention feature map corresponding to the nth original feature map, and the fusion map corresponding to the (i + 1)th original feature map, and i is an integer from (n - 1) to 2 in sequence.
[0022] In one embodiment, obtaining the fused feature map corresponding to the i-th original feature map based on the channel attention feature map corresponding to the i-th original feature map, the spatial attention feature map corresponding to the n-th original feature map, and the fused feature map corresponding to the (i + 1)-th original feature map includes:
[0023] Performing deconvolution processing on the fused feature map corresponding to the (i + 1)-th original feature map to obtain a reference map corresponding to the (i + 1)-th original feature map;
[0024] Performing upsampling processing on the spatial attention feature map corresponding to the n-th original feature map to obtain an intermediate map;
[0025] Performing information integration processing on the reference map corresponding to the (i + 1)-th original feature map, the intermediate map, and the channel attention feature map corresponding to the i-th original feature map to obtain the fused feature map corresponding to the i-th original feature map.
[0026] In one embodiment, obtaining the target feature map of the image based on the fused feature map corresponding to the second original feature map and the first original feature map includes:
[0027] Performing upsampling on the fused feature map corresponding to the second original feature map to obtain a reference map corresponding to the second original feature map;
[0028] Performing information integration processing on the reference map corresponding to the second original feature map and the first original feature map to obtain the target feature map.
[0029] In one embodiment, determining the fused feature map corresponding to the n-th original feature map based on the spatial attention feature map and the channel attention feature map corresponding to the n-th original feature map includes:
[0030] Performing information integration processing on the spatial attention feature map corresponding to the n-th original feature map and the channel attention feature map corresponding to the n-th original feature map to obtain the fused feature map corresponding to the n-th original feature map.
[0031] An embodiment of the present application provides a feature extraction device for an image, and the device includes:
[0032] An acquisition module, configured to acquire an image to be processed;
[0033] A convolution module, configured to perform convolution processing on the image by using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4;
[0034] A first processing module, configured to perform feature extraction on each of the second to n-th original feature maps by using a preset channel attention mechanism model to obtain channel attention feature maps corresponding to each of the second to n-th original feature maps;
[0035] A second processing module, configured to extract features from the nth original feature map by using a preset spatial attention mechanism model, so as to obtain a spatial attention feature map corresponding to the nth original feature map;
[0036] A determination module, configured to perform layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map, so as to obtain a target feature map of the image;
[0037] A detection module, configured to perform target detection in an autonomous driving scenario based on the target feature map.
[0038] In a third aspect of the embodiments of the present application, a computer device is provided, where the device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the method for feature extraction of an image in the first aspect of the embodiments of the present application is implemented.
[0039] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method for feature extraction of an image in the first aspect of the embodiments of the present application is implemented.
[0040] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0041] For the method for feature extraction of an image provided in the embodiments of the present application, by obtaining an image to be processed, and performing convolution processing on the image by using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4, then, extracting features from each of the 2nd to nth original feature maps by using a preset channel attention mechanism model to obtain channel attention feature maps corresponding to each of the 2nd to nth original feature maps, and extracting features from the nth original feature map by using a preset spatial attention mechanism model to obtain a spatial attention feature map corresponding to the nth original feature map. Finally, layer-by-layer feature fusion is performed based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map to obtain a target feature map of the image. Since the deep feature maps have rich semantic information and the shallow feature maps have accurate position information, the target feature map obtained after layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map includes both rich semantic information and accurate position information. Further, such a target feature map is used for target detection, which can improve the accuracy of target detection and the experience of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Schematic internal structure diagram of a computer device provided by an embodiment of the present application;
[0043] Figure 2 Flowchart of a method for feature extraction of an image provided by an embodiment of the present application Figure 1 ;
[0044] Figure 3 Flowchart of a method for feature extraction of an image provided by an embodiment of the present application Figure 2 ;
[0045] Figure 4 Schematic diagram of a feature extraction process of an image provided by an embodiment of the present application;
[0046] Figure 5 Structural diagram of an image feature extraction device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality" is two or more.
[0049] In addition, the use of "based on" or "according to" means open and inclusive, because a process, step, calculation, or other action "based on" or "according to" one or more conditions or values may in practice be based on additional conditions or values beyond those stated.
[0050] In special scenarios such as driving scenarios and some industrial and mining areas with complex environments, there are quite high requirements for the real-time performance and accuracy of target detection. Feature extraction of images is the basis for various intensive computational downstream tasks such as target detection or image segmentation. The accuracy of the image feature extraction results directly affects the accuracy of subsequent downstream tasks such as target detection and image segmentation.
[0051] Currently, most of the image feature extraction methods obtain the feature map of the image through a convolutional neural network. However, although the feature map obtained by this feature extraction method has rich semantic information, the position information is relatively vague, which affects the accuracy of object detection and thus the experience of autonomous driving.
[0052] To solve the above problems, the embodiments of the present application provide a method for extracting features of an image. By obtaining the image to be processed and using n convolutional structures in a preset convolutional neural network model to perform convolutional processing on the image, n original feature maps are obtained, where n is an integer greater than 4. Then, using a preset channel attention mechanism model to extract features from each of the second to the nth original feature maps, channel attention feature maps corresponding to each of the second to the nth original feature maps are obtained, and using a preset spatial attention mechanism model to extract features from the nth original feature map, a spatial attention feature map corresponding to the nth original feature map is obtained. Finally, based on the channel attention feature maps corresponding to the second to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the first original feature map, layer-by-layer feature fusion is performed to obtain the target feature map of the image. Since the deep feature maps have rich semantic information and the shallow feature maps have accurate position information, the target feature map obtained by layer-by-layer feature fusion based on the channel attention feature maps corresponding to the second to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the first original feature map includes both rich semantic information and accurate position information. Further, such a target feature map is used for object detection, which can improve the accuracy of object detection and the experience of autonomous driving.
[0053] In addition, the target feature map obtained by the image feature extraction method provided by the embodiments of the present application has the same effect when used for detecting drivable areas, lane line detection, etc. in autonomous driving scenarios.
[0054] The execution subject of the image feature extraction method provided by the embodiments of the present application can be a computer device, a terminal device, or a server. Among them, the terminal device can be various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, etc., and the present application does not make specific limitations in this regard.
[0055] Figure 1 It is a schematic internal structure diagram of a computer device provided by the embodiments of the present application. As Figure 1As shown in the figure, the computer device includes a processor and a memory connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The computer program can be executed by the processor to implement the steps of a method for feature extraction of an image provided in each of the above embodiments. The internal memory provides a cache operating environment for the operating system and the computer program in the non-volatile storage medium.
[0056] Those skilled in the art can understand that Figure 1 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0057] Based on the above execution subject, an embodiment of the present application provides a method for feature extraction of an image. As Figure 2 shown in the figure, the method includes the following steps:
[0058] Step 201, obtain the image to be processed.
[0059] In the autonomous driving scenario, the image to be processed is obtained by an in-vehicle camera. The resolution of the image is usually 1280*720, and the image may contain information such as vehicles, pedestrians, road surfaces, and lane lines. Usually, object detection needs to be performed on the vehicle, pedestrian, or other obstacle information included in the image.
[0060] Step 202, perform convolution processing on the image using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4.
[0061] The 1st to nth original feature maps are obtained in sequence through a preset convolutional neural network model. Usually, the length and width scales of the next-layer feature map are 1 / 2 of the previous layer, and the number of channels is 2 times that of the previous layer, which specifically depends on the backbone network. The present application does not make specific limitations on this.
[0062] It should be noted that the following entire process is described by taking n = 5 as an example, but the present application is not limited to n = 5.
[0063] Among them, if n = 5, there are 5 convolution structures and 5 convolution feature maps. The 1st convolution structure outputs the 1st original feature map, the 2nd convolution structure outputs the 2nd original feature map, the 3rd convolution structure outputs the 3rd original feature map, the 4th convolution structure outputs the 4th original feature map, and the 5th convolution structure outputs the 5th original feature map.
[0064] The above process is as follows: Input the image into the first convolutional structure to obtain the first original feature map, then input the first convolutional feature map into the second convolutional structure to obtain the second original feature map, then input the second original feature map into the third convolutional structure to obtain the third original feature map, then input the third original feature map into the fourth convolutional structure to obtain the fourth original feature map, and finally input the fourth original feature map into the fifth convolutional structure to obtain the fifth original feature map.
[0065] Step 203: Use the preset channel attention mechanism module to extract features from each of the second to the nth original feature maps, and obtain the channel attention feature maps corresponding to each of the second to the nth original feature maps.
[0066] That is to say, input the second original feature map, the third original feature map, the fourth original feature map, and the fifth original feature map into the model based on the channel attention mechanism for feature extraction respectively, and obtain the second channel attention feature map corresponding to the second original feature map, the third channel attention feature map corresponding to the third original feature map, the fourth channel attention feature map corresponding to the fourth original feature map, and the fifth channel attention feature map corresponding to the fifth original feature map. Since the extracted features have more selective attention to the channels compared to the original features. Therefore, select the second to the nth original feature maps, so as to ensure that semantic information at multiple scales can be obtained.
[0067] Specifically, for the model based on the channel attention mechanism, first use global average pooling to convert the feature map with size (H, W, C) into a feature map with size (1, 1, C), and the pooling is implemented according to the following formula:
[0068]
[0069] where H, W, and C are the height, width, and number of channels of the feature map respectively, and x c (i,j) is the value of the c-th channel at the i-th row and j-th column. Use two fully connected layers to obtain the attention weights between channels, use a sigmoid function to obtain the normalized weights, and finally multiply with the original feature map in the channel dimension to obtain the channel attention feature map.
[0070] s c = Sigmoid(fc2(Relu(fc1(p c ,w1))),w2)
[0071] Step 204: Use the preset spatial attention mechanism module to extract features from the nth original feature map, and obtain the spatial attention feature map corresponding to the nth original feature map.
[0072] The 5th original feature map is input into a preset model based on the spatial attention mechanism for feature extraction, and the 5th spatial attention feature map corresponding to the 5th original feature map is obtained. The extracted features have more selective attention to spatial positions compared to the original features. n is selected because the size of the nth original feature map is small, and the computational cost of adding an attention module to it is relatively small. In this way, while ensuring that the obtained feature map includes both rich semantic information and accurate position information, the feature extraction efficiency can be improved.
[0073] Among them, the attention module of the model based on the spatial attention mechanism uses RCCA (Recurrent Criss-Cross Attention), that is, a recurrent cross-shaped attention module. Since the CCA module only focuses on the cross-shaped spatial region of the same row and the same column as the current element. Although the CCA pays attention to the context information in the horizontal and vertical directions, it is still very sparse. And RCCA obtains a wider range of context information through continuous cycling of the CCA module in a criss-cross and iterative manner. The cyclic RCCA defaults to executing CCA twice. After the first CCA, the attention weights in the horizontal and vertical directions are obtained. In the second CCA, the current horizontal and vertical directions already contain the context information obtained after the previous operation. In this way, only two CCA iterations are needed to indirectly obtain global semantic information. In this way, not only can rich global semantic information be obtained, but also the feature extraction efficiency can be improved. At the same time, RCCA has advantages in terms of computational speed and memory consumption.
[0074] Step 205: Based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map, perform layer-by-layer feature fusion to obtain the target feature map of the image.
[0075] Use the 2nd channel attention feature map, the 3rd channel attention feature map, the 4th channel attention feature map, the 5th channel attention feature map, the 5th spatial attention feature map, and the 1st original feature map to perform layer-by-layer feature fusion to obtain the target feature map of the image. Based on the basic feature pyramid network, add an attention mechanism to enhance the feature extraction effect. At the same time, the shallow network has strong position information, and the deep network has strong semantic information. Among them, the original input image is an image, and the feature map is a series of matrices used to describe the input image.
[0076] Step 206: Perform object detection in the autonomous driving scenario based on the target feature map.
[0077] Through the feature extraction method provided by the embodiments of the present application, the target feature map of the image can be obtained. Based on this target feature map, tasks such as object detection and semantic segmentation can be performed.
[0078] An embodiment of the present application provides a method for extracting features of an image. By obtaining an image to be processed and performing convolution processing on the image using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4. Then, using a preset channel attention mechanism model to extract features from each of the original feature maps from the 2nd to the nth original feature maps to obtain channel attention feature maps corresponding to each of the original feature maps from the 2nd to the nth original feature maps, and using a preset spatial attention mechanism model to extract features from the nth original feature map to obtain a spatial attention feature map corresponding to the nth original feature map. Finally, based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map, layer-by-layer feature fusion is performed to obtain the target feature map of the image. Since the deep feature maps have rich semantic information and the shallow feature maps have accurate position information, the target feature map obtained after layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map includes both rich semantic information and accurate position information. Further, such a target feature map is used for tasks such as target detection or image segmentation, which can improve the accuracy of task processing. In addition, due to the use of convolution, the size of the receptive field is limited, and there is still a lack of global spatial information in the feature maps after multiple convolutions. The use of the attention mechanism model can solve the problem of limited receptive field of the convolutional network.
[0079] Optionally, as Figure 3 shown, the process of step 205, based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map, performing layer-by-layer feature fusion to obtain the target feature map of the image can be as follows:
[0080] Step 301, based on the spatial attention feature map and the channel attention feature map corresponding to the nth original feature map, determine the fusion map corresponding to the nth original feature map.
[0081] Step 302, based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps, and the 1st original feature map, perform layer-by-layer feature fusion to obtain the target feature map of the image.
[0082] Optionally, the process of step 302 above may be: based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps, perform layer-by-layer feature fusion to obtain the fusion map corresponding to the 2nd original feature map, and based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map, obtain the target feature map of the image.
[0083] Specifically, the process of performing layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps to obtain the fusion map corresponding to the 2nd original feature map may be:
[0084] Start from i = n - 1 and perform at least one fusion process until i = 2 to obtain the fusion map corresponding to the 2nd original feature map, where i is an integer from (n - 1) to 2;
[0085] Wherein, the mth fusion process includes: according to the channel attention feature map corresponding to the ith original feature map, the spatial attention feature map corresponding to the nth original feature map, and the fusion map corresponding to the (i + 1)th original feature map, obtain the fusion map corresponding to the ith original feature map, and i is successively an integer from (n - 1) to 2.
[0086] Taking n = 5 as an example, the implementation processes of step 301 and step 302 above may be: first, according to the 5th channel attention feature map and the 5th spatial attention feature map, obtain the 5th fusion map. Then, according to the 5th fusion map, the 4th channel attention feature map, and the 5th spatial attention feature map, obtain the 4th fusion map. Then, according to the 4th fusion map, the 3rd channel attention feature map, and the 5th spatial attention feature map, obtain the 3rd fusion map. Then, according to the 3rd fusion map, the 2nd channel attention feature map, and the 5th spatial attention feature map, obtain the 2nd fusion map. Finally, according to the 2nd fusion map and the 1st original feature map, obtain the target feature map of the image. The target feature map obtained by performing layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map includes both rich semantic information and accurate position information.
[0087] Optionally, the process of step 301 determining the fusion map corresponding to the nth original feature map based on the spatial attention feature map and the channel attention feature map corresponding to the nth original feature map may be:
[0088] Perform information integration processing on the spatial attention feature map corresponding to the nth original feature map and the channel attention feature map corresponding to the nth original feature map to obtain the fusion map corresponding to the nth original feature map.
[0089] Among them, the information integration process can be the add process. For example, the 5th spatial attention feature map and the 5th channel attention feature map are subjected to the add process to obtain the 5th fusion map.
[0090] Optionally, the process of obtaining the target feature map of the image based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map can be as follows:
[0091] Upsample the fusion map corresponding to the 2nd original feature map to obtain the reference map corresponding to the 2nd original feature map, and then perform information integration processing on the reference map corresponding to the 2nd original feature map and the 1st original feature map to obtain the target feature map.
[0092] For example, upsample the 2nd fusion map to obtain the 2nd reference map. Then perform the add process on the 2nd reference map and the 1st original feature map to obtain the target feature map.
[0093] Optionally, the process of obtaining the fusion map corresponding to the i-th original feature map based on the channel attention feature map corresponding to the i-th original feature map, the spatial attention feature map corresponding to the n-th original feature map, and the fusion map corresponding to the (i + 1)-th original feature map can be as follows:
[0094] Perform deconvolution processing on the fusion map corresponding to the (i + 1)-th original feature map to obtain the reference map corresponding to the (i + 1)-th original feature map; perform upsampling processing on the spatial attention feature map corresponding to the n-th original feature map to obtain the intermediate map; perform information integration processing on the reference map corresponding to the (i + 1)-th original feature map, the intermediate map, and the channel attention feature map corresponding to the i-th original feature map to obtain the fusion map corresponding to the i-th original feature map.
[0095] Taking n = 5 as an example, that is, the process of obtaining the 4th - 2nd fusion maps above can be as follows:
[0096] Perform deconvolution processing on the 5th fusion map to obtain the 5th reference map, then perform upsampling processing on the 5th spatial attention feature map to obtain the intermediate map, and finally perform the add process on the 5th reference map, the intermediate map, and the 4th channel attention feature map to obtain the 4th fusion map. Perform deconvolution processing on the 4th fusion map to obtain the 4th reference map, then perform upsampling processing on the 5th spatial attention feature map to obtain the intermediate map, and finally perform the add process on the 4th reference map, the intermediate map, and the 3rd channel attention feature map to obtain the 3rd fusion map. Perform deconvolution processing on the 3rd fusion map to obtain the 3rd reference map, then perform upsampling processing on the 5th spatial attention feature map to obtain the intermediate map, and finally perform the add process on the 3rd reference map, the intermediate map, and the 2nd channel attention feature map to obtain the 2nd fusion map. Finally, the 4th - 2nd fusion maps are obtained.
[0097] The entire process of performing layer-by-layer feature fusion using the second-channel attention feature map, the third-channel attention feature map, the fourth-channel attention feature map, the fifth-channel attention feature map, the fifth spatial attention feature map, and the first original feature map to obtain the target feature map of the image can be referred to Figure 4 , where Figure 4 The data in the columns of 20x12 and 40x24 in represent the resolution of the image, and the data in the columns of 1 / 2 and 1 / 4 represent the size of the image compared to the original image. Figure 4 c1, c2, c3, c4, and c5 in correspond to the first original feature map, the second original feature map, the third original feature map, the fourth original feature map, and the fifth original feature map, respectively. Figure 4 p2, p3, p4, and p5 in correspond to the second fusion map, the third fusion map, the fourth fusion map, and the fifth fusion map, respectively. Figure 4 p1 in is the final target feature map. Among them, Figure 4 Head1, Head2, and Head3 in are the tasks to which the target feature map is to be input, and this task can be a target detection task and a semantic segmentation task. Figure 4 SAM in is a spatial attention mechanism model, and CAM is a channel attention mechanism model.
[0098] The target feature map obtained by the image feature extraction method provided by the embodiments of the present application has improved processing effects in both target detection tasks and semantic segmentation tasks compared to the feature maps obtained by traditional feature pyramid networks, and the data is shown in Table 1.
[0099] Table 1 Comparison chart of experimental results
[0100] Model Object detection mAp Semantic segmentation mIoU Feature map obtained by traditional feature pyramid network 0.401 0.371 Feature map obtained by the method of this application 0.490 0.383
[0101] As Figure 5 shown, the embodiments of the present application also provide an image feature extraction device, and the device includes:
[0102] An acquisition module 11, configured to acquire an image to be processed;
[0103] A convolution module 12, configured to perform convolution processing on the image using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4;
[0104] A first processing module 13, configured to use a preset channel attention mechanism model to perform feature extraction on each of the second to nth original feature maps to obtain a channel attention feature map corresponding to each of the second to nth original feature maps;
[0105] The second processing module 14 is configured to extract features from the nth original feature map by using a preset spatial attention mechanism model, so as to obtain a spatial attention feature map corresponding to the nth original feature map;
[0106] The determination module 15 is configured to perform layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map, so as to obtain a target feature map of the image;
[0107] The detection module 16 is configured to perform target detection in an autonomous driving scenario based on the target feature map.
[0108] In one embodiment, the determination module 15 is specifically configured to:
[0109] Based on the spatial attention feature map and the channel attention feature map corresponding to the nth original feature map, determine a fusion map corresponding to the nth original feature map;
[0110] Based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps, and the 1st original feature map, perform layer-by-layer feature fusion to obtain a target feature map of the image.
[0111] In one embodiment, the determination module 15 is specifically configured to:
[0112] Based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps, perform layer-by-layer feature fusion to obtain a fusion map corresponding to the 2nd original feature map;
[0113] Based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map, obtain a target feature map of the image.
[0114] In one embodiment, the determination module 15 is specifically configured to:
[0115] Based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to 2nd original feature maps, perform layer-by-layer feature fusion to obtain a fusion map corresponding to the 2nd original feature map;
[0116] Based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map, obtain a target feature map of the image.
[0117] In one embodiment, the determination module 15 is specifically configured to:
[0118] Start performing at least one fusion processing procedure starting from \(i = n - 1\) until \(i = 2\) to obtain the fusion graph corresponding to the second original feature map, where \(i\) is an integer from \((n - 1)\) to 2;
[0119] Among them, the \(m\)-th fusion processing procedure includes: obtaining the fusion graph corresponding to the \(i\)-th original feature map according to the channel attention feature map corresponding to the \(i\)-th original feature map, the spatial attention feature map corresponding to the \(n\)-th original feature map, and the fusion graph corresponding to the \((i + 1)\)-th original feature map, where \(i\) is an integer from \((n - 1)\) to 2 in sequence.
[0120] In one embodiment, the determination module 15 is specifically configured to:
[0121] Perform deconvolution processing on the fusion graph corresponding to the \((i + 1)\)-th original feature map to obtain a reference graph corresponding to the \((i + 1)\)-th original feature map;
[0122] Perform upsampling processing on the spatial attention feature map corresponding to the \(n\)-th original feature map to obtain an intermediate graph;
[0123] Perform information integration processing on the reference graph corresponding to the \((i + 1)\)-th original feature map, the intermediate graph, and the channel attention feature map corresponding to the \(i\)-th original feature map to obtain the fusion graph corresponding to the \(i\)-th original feature map.
[0124] In one embodiment, the determination module 15 is specifically configured to:
[0125] Perform upsampling on the fusion graph corresponding to the second original feature map to obtain a reference graph corresponding to the second original feature map;
[0126] Perform information integration processing on the reference graph corresponding to the second original feature map and the first original feature map to obtain the target feature map.
[0127] In one embodiment, the determination module 15 is specifically configured to:
[0128] Perform information integration processing on the spatial attention feature map corresponding to the \(n\)-th original feature map and the channel attention feature map corresponding to the \(n\)-th original feature map to obtain the fusion graph corresponding to the \(n\)-th original feature map.
[0129] The feature extraction device for images provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here.
[0130] For the specific limitations of the image feature extraction device, reference may be made to the limitations of the image feature extraction method in the foregoing text, which will not be elaborated here. Each module in the above image feature extraction device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the server in hardware form or independent of it, or stored in the memory in the server in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0131] In another embodiment of the present application, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the image feature extraction method as in the embodiments of the present application are implemented.
[0132] In another embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by the processor, the steps of the image feature extraction method as in the embodiments of the present application are implemented.
[0133] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer instructions. When the computer instructions run on the image feature extraction device, the image feature extraction device is enabled to execute each step in the method flow shown in the above method embodiment for the image feature extraction method.
[0134] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0135] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0136] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for feature extraction of an image, characterized in that, The method includes: Obtaining an image to be processed; Performing convolution processing on the image by using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4; Performing feature extraction on each of the original feature maps from the 2nd to the nth original feature maps by using a preset channel attention mechanism model to obtain the channel attention feature maps corresponding to each of the original feature maps from the 2nd to the nth original feature maps; Performing feature extraction on the nth original feature map by using a preset spatial attention mechanism model to obtain the spatial attention feature map corresponding to the nth original feature map; Performing layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map to obtain the target feature map of the image; Performing object detection in an autonomous driving scenario based on the target feature map.
2. The method according to claim 1, wherein The performing layer-by-layer feature fusion based on the channel attention feature maps corresponding to the 2nd to the nth original feature maps, the spatial attention feature map corresponding to the nth original feature map, and the 1st original feature map to obtain the target feature map of the image includes: Determining a fusion map corresponding to the nth original feature map based on the spatial attention feature map and the channel attention feature map corresponding to the nth original feature map; Performing layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps, and the 1st original feature map to obtain the target feature map of the image.
3. The method according to claim 2, wherein The performing layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps, and the 1st original feature map to obtain the target feature map of the image includes: Performing layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps to obtain the fusion map corresponding to the 2nd original feature map; Obtaining the target feature map of the image based on the fusion map corresponding to the 2nd original feature map and the 1st original feature map.
4. The method according to claim 3, wherein The performing layer-by-layer feature fusion based on the fusion map corresponding to the nth original feature map, the spatial attention feature map, and the channel attention feature maps corresponding to the (n - 1)th to the 2nd original feature maps to obtain the fusion map corresponding to the 2nd original feature map includes: Starting from i = n - 1 and performing at least one fusion processing process until i = 2 to obtain the fusion map corresponding to the 2nd original feature map, where i is an integer from (n - 1) to 2; Wherein, the mth fusion processing process includes: obtaining the fusion map corresponding to the ith original feature map according to the channel attention feature map corresponding to the ith original feature map, the spatial attention feature map corresponding to the nth original feature map, and the fusion map corresponding to the (i + 1)th original feature map, and i is successively an integer from (n - 1) to 2.
5. The method according to claim 4, characterized in that Obtaining the fused graph corresponding to the i-th original feature map based on the channel attention feature map corresponding to the i-th original feature map, the spatial attention feature map corresponding to the n-th original feature map, and the fused graph corresponding to the (i + 1)-th original feature map includes: Performing deconvolution processing on the fused graph corresponding to the (i + 1)-th original feature map to obtain a reference graph corresponding to the (i + 1)-th original feature map; Performing upsampling processing on the spatial attention feature map corresponding to the n-th original feature map to obtain an intermediate graph; Performing information integration processing on the reference graph corresponding to the (i + 1)-th original feature map, the intermediate graph, and the channel attention feature map corresponding to the i-th original feature map to obtain the fused graph corresponding to the i-th original feature map.
6. The method according to claim 3, characterized in that Obtaining the target feature map of the image based on the fused graph corresponding to the second original feature map and the first original feature map includes: Performing upsampling on the fused graph corresponding to the second original feature map to obtain a reference graph corresponding to the second original feature map; Performing information integration processing on the reference graph corresponding to the second original feature map and the first original feature map to obtain the target feature map.
7. The method according to claim 2, wherein Determining the fused graph corresponding to the n-th original feature map based on the spatial attention feature map and the channel attention feature map corresponding to the n-th original feature map includes: Performing information integration processing on the spatial attention feature map corresponding to the n-th original feature map and the channel attention feature map corresponding to the n-th original feature map to obtain the fused graph corresponding to the n-th original feature map.
8. An image feature extraction device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be processed; A convolution module, configured to perform convolution processing on the image by using n convolution structures in a preset convolutional neural network model to obtain n original feature maps, where n is an integer greater than 4; A first processing module, configured to perform feature extraction on each of the second to n-th original feature maps by using a preset channel attention mechanism model to obtain the channel attention feature map corresponding to each of the second to n-th original feature maps; A second processing module, configured to perform feature extraction on the n-th original feature map by using a preset spatial attention mechanism model to obtain the spatial attention feature map corresponding to the n-th original feature map; A determination module, configured to perform layer-by-layer feature fusion based on the channel attention feature maps corresponding to the second to n-th original feature maps, the spatial attention feature map corresponding to the n-th original feature map, and the first original feature map to obtain the target feature map of the image; A detection module, configured to perform target detection in an autonomous driving scenario based on the target feature map.
9. A computer device, characterized in that, It includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, it implements the image feature extraction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the image feature extraction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image saliency detection method based on feature selection and feature fusion
CN111275076A
Feature extraction method and device for super-resolution image reconstruction and storage medium
CN114511446A