Video target feature extraction method, device, computer equipment and storage medium
By extracting the features of video frame images layer by layer, using convolution and attention calculation to generate feature maps of different resolutions, and performing feature fusion and downsampling processing, the problem of low feature extraction efficiency in the prior art is solved, and more efficient video target feature extraction is achieved.
Patent Information
- Application Number
- CN202210590089.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The existing video target feature extraction method uses ResNet50 as the backbone network, resulting in low feature extraction efficiency.
The features of video frame images are extracted layer by layer, and feature maps with different resolutions are generated through convolution processing and attention calculation, and feature fusion and downsampling are performed to improve feature extraction efficiency.
The speed and efficiency of feature extraction of video targets is improved, and the accuracy and management efficiency of feature maps are enhanced.
Smart Images

Figure CN114973088B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video segmentation technology, and in particular to a method, device, computer equipment, and storage medium for extracting features of a video target. Background Art
[0002] Semi-supervised video segmentation involves using the segmentation results of a given first frame as a reference for propagation to subsequent frames. Feature extraction of video objects is a key step in semi-supervised video segmentation. The goal of video object segmentation is to segment objects of specific categories within the input video frame and generate their segmentation masks.
[0003] Existing methods for extracting features from video targets use an encoder network to extract features layer by layer, using the extracted features as target features. However, this method uses a ResNet50 network as its backbone. Due to its large capacity, the image extraction efficiency is low, resulting in low feature extraction efficiency for video targets. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method, apparatus, computer equipment and storage medium for extracting features of a video target, so as to improve the efficiency of extracting features of a video target.
[0005] In order to solve the above technical problems, the present invention provides a method for extracting features of a video target, including:
[0006] Obtain a current frame image in the video, and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence;
[0007] Performing convolution processing on the second resolution feature map to obtain a second convolution feature map, and performing attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weighted feature map;
[0008] Performing convolution processing on the first resolution feature map to obtain a first convolution feature map, multiplying the first convolution feature map with the second weight feature map to obtain a first weight feature map, and performing feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map;
[0009] Obtaining a second output feature map by concatenating the first output feature map and the third resolution feature map, and using the fourth resolution feature map as the third output feature map;
[0010] Based on the current frame image, downsampling and concatenating the first output feature map, the second output feature map, and the third output feature map to obtain a first basic feature map, a second basic feature map, and a third basic feature map;
[0011] A first target feature map, a second target feature map, and a third target feature map are obtained by performing convolution processing on the first basic feature map, the second basic feature map, and the third basic feature map respectively.
[0012] In order to solve the above technical problems, the present invention provides a feature extraction device for a video target, comprising:
[0013] A feature map extraction module is used to obtain a current frame image in the video and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence;
[0014] A second weighted feature map generation module is configured to perform convolution processing on the second resolution feature map to obtain a second convolution feature map, and perform attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weighted feature map;
[0015] a first output feature map generating module, configured to perform convolution processing on the first resolution feature map to obtain a first convolution feature map, multiply the first convolution feature map by the second weight feature map to obtain a first weight feature map, and perform feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map;
[0016] a second output feature map generating module, configured to obtain a second output feature map by cascading the first output feature map and the third resolution feature map, and use the fourth resolution feature map as the third output feature map;
[0017] a basic feature map generation module, configured to perform downsampling and cascading processing on the first output feature map, the second output feature map, and the third output feature map based on the current frame image to obtain a first basic feature map, a second basic feature map, and a third basic feature map;
[0018] The target feature map generation module is used to obtain a first target feature map, a second target feature map and a third target feature map by performing convolution processing on the first basic feature map, the second basic feature map and the third basic feature map respectively.
[0019] To solve the above technical problems, a technical solution adopted by the present invention is: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the feature extraction method of the video target described in any one of the above.
[0020] The embodiment of the present invention provides a method, device, computer equipment and storage medium for extracting features of a video target. The method includes: obtaining a current frame image in a video, and extracting features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map and a fourth resolution feature map in sequence; performing convolution processing on the second resolution feature map to obtain a second convolution feature map, and performing attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weight feature map; performing convolution processing on the first resolution feature map to obtain a first convolution feature map, and multiplying the first convolution feature map with the second weight feature map to obtain a first weight feature map, and multiplying the first weight feature map with the second weight feature map to obtain a first weight feature map. The convolution feature map is subjected to feature fusion to obtain a first output feature map; the first output feature map is cascaded with the third resolution feature map to obtain a second output feature map, and the fourth resolution feature map is used as the third output feature map; based on the current frame image, the first output feature map, the second output feature map and the third output feature map are downsampled and cascaded to obtain a first basic feature map, a second basic feature map and a third basic feature map; the first basic feature map, the second basic feature map and the third basic feature map are convoluted respectively to obtain a first target feature map, a second target feature map and a third target feature map. The embodiment of the present invention extracts the features of the current frame image layer by layer to obtain feature maps of different resolutions, and performs attention calculation based on the feature maps of different resolutions to obtain weight information of the feature maps. At the same time, the different feature maps are cascaded and convolved to improve the speed of matting, and the feature maps are downsampled to obtain the target feature map, which is beneficial to improving the feature extraction efficiency of the video target. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 This is a flowchart of an implementation of the method for extracting features from a video target provided by an embodiment of the present application;
[0023] Figure 2This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0024] Figure 3 This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0025] Figure 4 This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0026] Figure 5 This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0027] Figure 6 This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0028] Figure 7 This is another implementation flowchart of a sub-process in the feature extraction method of a video target provided in an embodiment of the present application;
[0029] Figure 8 Schematic diagram of a feature extraction device for a video target provided in an embodiment of the present application;
[0030] Figure 9 It is a schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0032] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0034] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0035] It should be noted that the feature extraction method of the video target provided in the embodiment of the present application is generally executed by a server, and accordingly, the feature extraction device of the video target is generally configured in the server.
[0036] See also Figure 1 , Figure 1 A specific implementation of a feature extraction method for a video target is shown.
[0037] It should be noted that the method of the present invention is not limited to the method of Figure 1 The process sequence shown is limited to the following steps:
[0038] S1: Obtain a current frame image in the video, and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence.
[0039] Specifically, this application is an improved version based on the STCN network architecture. STCN consists of three modules: the encoding (Encoder) network, the decoding (decoder) network and the associative memory module. The encoding network is divided into three sub-modules, corresponding to: KeyEncoder, ValueEncoderSO and ValueEncoder; the decoding network is a simple CNN network superposition based on the output of three different resolutions for feature fusion; the associative memory module is used to solve the approximate feature map of the fused features. The key point of this application is to improve the three sub-modules of the encoding network.
[0040] This application is based on the STCN network architecture, which is a semi-supervised video segmentation algorithm for extracting video target features. In an embodiment of the present application, the current frame image in the video is first obtained, and the current frame image is an image of a reference mask that has been marked, and then based on the reference mask image, the predicted mask image of the subsequent frame is obtained. In an embodiment of the present application, assuming that the current frame image in the input video is an RGB image with a width and height of 320x320, the input dimension is: 1x3x320x320, and the reference mask image is a single-channel binary image, and the input dimension is: 1x1x320x320. This application uses the EfficientNet-b3 network structure as the backbone network to extract the features of the network structure layer by layer for the current frame image, which are four feature maps of different resolutions, namely 1x384x10x10, 1x136x20x20, 1x48x40x40 and 1x32x80x80, that is, the first resolution feature map, the second resolution feature map, the third resolution feature map and the fourth resolution feature map.
[0041] S2: Convolve the second resolution feature map to obtain a second convolution feature map, and perform attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weight feature map.
[0042] See also Figure 2 , Figure 2 A specific implementation of step S2 is shown, which is described in detail as follows:
[0043] S21: Perform convolution processing on the second resolution feature map to obtain a second convolution feature map.
[0044] S22: Enlarging the first resolution feature map by a preset multiple to obtain a first resolution feature enlarged map.
[0045] S23: Cascading the second convolution feature map and the first resolution feature magnification map to obtain a cascade processing result, and convolving the cascade processing result to obtain a second cascade feature map.
[0046] S24: Obtain a second weighted feature map by performing attention calculation on the second cascade feature map.
[0047] Specifically, the second resolution feature map (1x136x20x20) is first passed through a convolution layer (input is 136, output is 128, kernel size is 3, step size is 1) and convolution is performed to obtain the second convolution feature map. Figure 1x128x20x20; then the first resolution feature map is magnified by a preset multiple to obtain a first resolution feature magnification map. The preset multiple of the embodiment of the present application is 2 times. The first resolution feature map is magnified by a preset multiple of 2 times to obtain a first resolution feature magnification map of 1x384x20x20. The second convolution feature map and the first resolution feature magnification map are cascaded to obtain a cascaded processing result. The cascaded processing result is a feature image of 1x512x20x20. The feature image is then input into a convolution layer (input is 512, output is 128, kernel size is 1, step size is 1) for convolution processing to obtain a second cascaded feature map of 1x128x20x20. Finally, attention calculation is performed on the second cascaded feature map to obtain a second weighted feature map, which is a feature map of 1x128x1x1. Among them, cascade refers to the mapping relationship between multiple objects in computer science. Establishing a cascade relationship between data improves management efficiency.
[0048] Specifically, by convolving the second resolution feature map to obtain a second convolution feature map, and performing attention calculation on the second convolution feature map to obtain a second weight feature map, the feature information of the feature map is effectively obtained, and its attention weight is obtained, providing a basis for subsequent feature extraction.
[0049] See also Figure 3 , Figure 3 A specific implementation of step S24 is shown, which is described in detail as follows:
[0050] S241: Performing average pooling on the second cascade feature map to obtain a pooled feature map, and performing convolution on the pooled feature map to obtain an initial weight feature map.
[0051] S242: Normalize and activate the initial weight feature map to obtain a second weight feature map.
[0052] Specifically, the second cascade feature map is input into the average pooling layer (pooling width is 20, pooling height is 20) for average pooling processing to obtain a pooled feature map, which includes 1x128x1x1 weight information. The pooled feature map is then passed through a convolution layer (input is 128, output is 128, kernel size is 1, step size is 1) for convolution processing to obtain an initial weight feature map, which includes 1x128x1x1 weight information. Finally, the initial weight feature map is passed through the normalization layer and the Hard Sigmoid layer for normalization and activation processing to obtain the second weight feature map, where the second weight feature map includes weight information 1x128x1x1.
[0053] S3: Convolve the first resolution feature map to obtain a first convolution feature map, multiply the first convolution feature map with the second weight feature map to obtain a first weight feature map, and perform feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map.
[0054] See also Figure 4 , Figure 4 A specific implementation of step S3 is shown, which is described in detail as follows:
[0055] S31: Perform convolution processing on the first resolution feature map to obtain a first convolution feature map.
[0056] S32: Multiply the first convolution feature map and the second weight feature map to obtain a first weight feature map.
[0057] S33: Amplify the first weight feature map by a preset multiple to obtain a first weight feature amplification map. S34: Based on the weight information of the second weight feature map, perform feature fusion on the first weight feature amplification map and the second convolution feature map to obtain a first output feature map.
[0058] Specifically, the first resolution feature map (1x384x10x10) is convolved through a convolution layer (input is 384, output is 128, kernel size is 3, step size is 1) to obtain the first convolution feature map of 1x128x10x10; then the attention weight of each layer of the second weight feature map is multiplied by the first convolution feature map to obtain the first weight feature with attention weight information. Figure 1 x128x10x10. By magnifying the first weight feature map by a preset multiple, which is 2 times, the first weight feature magnification is obtained. Figure 1 x128x20x20, where the first weighted feature amplification map includes weight information. Based on the weight information of the second weighted feature map, the first weighted feature amplification map is fused with the second convolution feature map to obtain the first output feature map.
[0059] See also Figure 5 , Figure 5 A specific implementation of step S34 is shown, which is described in detail as follows:
[0060] S341: Determine the weight information of the second convolution feature map based on the weight information of the second weight feature map.
[0061] S342: Multiply the first weighted feature amplification map and the second convolution feature map to obtain a multiplication result feature map, and add the multiplication result feature map and the second convolution feature map to obtain a fusion feature map.
[0062] S343: Perform convolution processing on the fused feature map to obtain a first output feature map.
[0063] Specifically, the weight information of the second weight feature map specifically includes the weight value W1 of each layer of the second weight feature map. According to the formula W2=1-W1, the weight information W2 of the second convolution feature map is obtained. The first weight feature amplification map is multiplied by the second convolution feature map to obtain a 1x128x20x20 multiplication result feature map, and the multiplication result feature map is added to the second convolution feature map to obtain a 1x128x20x20 fusion feature map. Finally, the fusion feature map is input into a convolution layer (input is 128, output is 128, kernel size is 1, step size is 1) for convolution processing to obtain a smoothed 1x128x20x20 first output feature map, which is the first low-resolution feature map output by the KeyEncoder module. Figure 1 x128x20x2.
[0064] S4: The first output feature map is cascaded with the third resolution feature map to obtain a second output feature map, and the fourth resolution feature map is used as the third output feature map.
[0065] See also Figure 6 , Figure 6 A specific implementation of step S4 is shown, which is described in detail as follows:
[0066] S41: Enlarging the first output feature map by a preset multiple to obtain a first output feature enlarged map.
[0067] S42: Cascade the first output feature magnification image and the third resolution feature image to obtain a cascade result.
[0068] S43: Obtain a second output feature map by performing convolution processing on the cascade result, and use the fourth resolution feature map as the third output feature map.
[0069] Specifically, the first output feature map is magnified by a preset multiple to obtain a first output feature magnification map, wherein the preset multiple is 2 times. After magnification by 2 times, a first output feature magnification map of 1x128x40x40 is obtained; the first output feature magnification map of 1x128x40x40 is cascaded with the third resolution feature map of 1x48x40x40 to obtain a cascade result, which is a feature map of 1x176x40x40. The cascade result is input into the convolution layer (input is 176, output is 128, kernel size is 3, and step size is 1) for convolution processing to obtain a second output feature map of 1x128x40x40, which is the fused feature map of the second output of the KeyEncoder module. At the same time, the fourth resolution feature map (1x32x80x80) is used as the third output feature map, which is the fused feature map of the third output of the KeyEncoder module.
[0070] S5: Based on the current frame image, the first output feature map, the second output feature map, and the third output feature map are downsampled and cascaded to obtain a first basic feature map, a second basic feature map, and a third basic feature map.
[0071] See also Figure 7 , Figure 7 A specific implementation of step S5 is shown, which is described in detail as follows:
[0072] S51: Based on the current frame image, downsample the first output feature map, the second output feature map, and the third output feature map to obtain a first sampling result, a second sampling result, and a third sampling result.
[0073] S52: Perform cascade processing on the first sampling result and the first output feature map to obtain a first basic feature map.
[0074] S53: Cascade the second sampling result and the second output feature map to obtain a second basic feature map.
[0075] S54: Cascade the third sampling result and the third output feature map to obtain a third basic feature map.
[0076] Specifically, based on the current frame image of 1x3x320x320 as input, the three output feature maps are downsampled accordingly to obtain a first sampling result of 1x3x20x20, a second sampling result of 1x3x40x40, and a third sampling result of 1x3x80x80. The first sampling result is then cascaded with the first output feature map to obtain a first basic feature map of 1x131x20x20; the second sampling result is cascaded with the second output feature map to obtain a second basic feature map of 1x131x40x40; and the third sampling result is cascaded with the third output feature map to obtain a third basic feature map of 1x35x80x80. There are two main purposes for reducing the image (also known as subsampling or downsampling): 1. To make the image fit the size of the display area; 2. To generate a thumbnail of the corresponding image.
[0077] In the example of the present application, the first output feature map, the second output feature map and the third output feature map are downsampled and cascaded based on the current frame image to obtain the first basic feature map, the second basic feature map and the third basic feature map, so as to reduce the image to the target size. At the same time, each sampling result is cascaded with the corresponding feature map so that it is associated with more image features, which is conducive to improving the accuracy of image feature extraction.
[0078] S6: Obtain a first target feature map, a second target feature map, and a third target feature map by performing convolution processing on the first basic feature map, the second basic feature map, and the third basic feature map respectively.
[0079] Specifically, the first basic feature map and the second basic feature map are input into the convolution layer (input is 131, output is 128, kernel size is 3, and step size is 1) for convolution processing to obtain a first target feature map of 1x128x20x20 and a second target feature map of 1x128x40x40; the third basic feature map is input into the convolution layer (input is 35, output is 32, kernel size is 3, and step size is 1) for convolution processing to obtain a third target feature map of 1x32x80x80.
[0080] The embodiment of the present invention provides a method, device, computer equipment and storage medium for extracting features of a video target. The method includes: obtaining a current frame image in a video, and extracting features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map and a fourth resolution feature map in sequence; performing convolution processing on the second resolution feature map to obtain a second convolution feature map, and performing attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weight feature map; performing convolution processing on the first resolution feature map to obtain a first convolution feature map, and multiplying the first convolution feature map with the second weight feature map to obtain a first weight feature map, and multiplying the first weight feature map with the second weight feature map to obtain a first weight feature map. The convolution feature map is subjected to feature fusion to obtain a first output feature map; the first output feature map is cascaded with the third resolution feature map to obtain a second output feature map, and the fourth resolution feature map is used as the third output feature map; based on the current frame image, the first output feature map, the second output feature map and the third output feature map are downsampled and cascaded to obtain a first basic feature map, a second basic feature map and a third basic feature map; the first basic feature map, the second basic feature map and the third basic feature map are convoluted respectively to obtain a first target feature map, a second target feature map and a third target feature map. The embodiment of the present invention extracts the features of the current frame image layer by layer to obtain feature maps of different resolutions, and performs attention calculation based on the feature maps of different resolutions to obtain weight information of the feature maps, and cascades and convolves the different feature maps, thereby improving the speed of matting, and downsampling the feature maps to obtain target feature maps, thereby helping to improve the feature extraction efficiency of video targets.
[0081] Please refer to Figure 8 , as a response to the above Figure 1 The present application provides an embodiment of a device for extracting features of a video target. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0082] like Figure 8 As shown, the feature extraction device of the video target of this embodiment includes: a feature map extraction module 71, a second weight feature map generation module 72, a first output feature map generation module 73, a second output feature map generation module 74, a basic feature map generation module 75 and a target feature map generation module 76, wherein:
[0083] A feature map extraction module 71 is used to obtain a current frame image in the video and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence;
[0084] A second weighted feature map generating module 72 is configured to perform convolution processing on the second resolution feature map to obtain a second convolution feature map, and perform attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weighted feature map;
[0085] A first output feature map generating module 73 is configured to perform convolution processing on the first resolution feature map to obtain a first convolution feature map, multiply the first convolution feature map with the second weight feature map to obtain a first weight feature map, and perform feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map;
[0086] a second output feature map generating module 74, configured to obtain a second output feature map by concatenating the first output feature map and the third resolution feature map, and use the fourth resolution feature map as the third output feature map;
[0087] A basic feature map generation module 75 is configured to perform downsampling and cascade processing on the first output feature map, the second output feature map, and the third output feature map based on the current frame image to obtain a first basic feature map, a second basic feature map, and a third basic feature map;
[0088] The target feature map generation module 76 is used to obtain a first target feature map, a second target feature map and a third target feature map by performing convolution processing on the first basic feature map, the second basic feature map and the third basic feature map respectively.
[0089] Furthermore, the second weight feature map generating module 72 includes:
[0090] a second convolution feature map generating unit, configured to perform convolution processing on the second resolution feature map to obtain a second convolution feature map;
[0091] A first resolution feature map magnification unit, configured to magnify the first resolution feature map by a preset multiple to obtain a first resolution feature magnification map;
[0092] A second cascade feature map generating unit is configured to obtain a cascade processing result by cascading the second convolution feature map and the first resolution feature magnification map, and performing convolution processing on the cascade processing result to obtain a second cascade feature map;
[0093] An attention calculation unit is used to obtain a second weighted feature map by performing attention calculation on the second cascade feature map.
[0094] Furthermore, the attention calculation unit includes:
[0095] An initial weight feature map generating subunit is used to obtain a pooled feature map by performing average pooling processing on the second cascade feature map, and to obtain an initial weight feature map by performing convolution processing on the pooled feature map;
[0096] The normalization processing subunit is used to perform normalization processing and activation processing on the initial weight feature map to obtain a second weight feature map.
[0097] Furthermore, the first output feature map generation module 73 includes:
[0098] a first convolution feature map generating unit, configured to perform convolution processing on the first resolution feature map to obtain a first convolution feature map;
[0099] A first weight feature map generating unit is configured to multiply the first convolution feature map by the second weight feature map to obtain a first weight feature map;
[0100] A first weight feature magnification map generating unit, configured to obtain a first weight feature magnification map by magnifying the first weight feature map by a preset multiple;
[0101] The weighted summation result generating unit is used to perform feature fusion on the first weight feature amplification map and the second convolution feature map based on the weight information of the second weight feature map to obtain a first output feature map.
[0102] Furthermore, the weighted summation result generating unit includes:
[0103] A basic weight information generating subunit, configured to determine the weight information of the second convolution feature map based on the weight information of the second weight feature map;
[0104] A feature map multiplication subunit is used to multiply the first weight feature amplification map with the second convolution feature map to obtain a multiplication result feature map, and add the multiplication result feature map with the second convolution feature map to obtain a fusion feature map;
[0105] The weighted sum result convolution subunit is used to perform convolution processing on the fused feature map to obtain a first output feature map.
[0106] Furthermore, the second output feature map generating module 74 includes:
[0107] The first output feature magnified image generating unit is configured to magnify the first output feature image by a preset multiple to obtain a first output feature magnified image.
[0108] a cascade result generating unit, configured to perform cascade processing on the first output feature magnification map and the third resolution feature map to obtain a cascade result;
[0109] The cascade result convolution unit is used to obtain a second output feature map by performing convolution processing on the cascade result, and use the fourth resolution feature map as the third output feature map.
[0110] Furthermore, the basic feature map generation module 75 includes:
[0111] a downsampling processing unit, configured to perform downsampling processing on the first output feature map, the second output feature map, and the third output feature map based on the current frame image to obtain a first sampling result, a second sampling result, and a third sampling result;
[0112] A first basic feature map generating unit, configured to perform cascade processing on the first sampling result and the first output feature map to obtain a first basic feature map;
[0113] A second basic feature map generating unit is configured to perform cascade processing on the second sampling result and the second output feature map to obtain a second basic feature map;
[0114] The third basic feature map generating unit is used to perform cascade processing on the third sampling result and the third output feature map to obtain a third basic feature map.
[0115] To solve the above technical problems, the present application also provides a computer device. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.
[0116] The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 8 with three components: a memory 81, a processor 82, and a network interface 83. However, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0117] Computer devices can be desktop computers, laptops, PDAs, cloud servers, etc. Computer devices can interact with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0118] The memory 81 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device 8. Of course, the memory 81 may also include both the internal storage unit of the computer device 8 and its external storage devices. In this embodiment, the memory 81 is generally used to store the operating system and various application software installed on the computer device 8, such as the program code of the feature extraction method of the video object. In addition, the memory 81 may also be used to temporarily store various types of data that have been output or are about to be output.
[0119] In some embodiments, the processor 82 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 82 is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to execute program code stored in the memory 81 or process data, such as executing the program code of the above-mentioned video object feature extraction method to implement various embodiments of the video object feature extraction method.
[0120] The network interface 83 may include a wireless network interface or a wired network interface. The network interface 83 is generally used to establish a communication connection between the computer device 8 and other electronic devices.
[0121] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program. The computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the feature extraction method of a video target as described above.
[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.
[0123] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A feature extraction method for a video target, characterized in that: include: Obtain a current frame image in the video, and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence; Performing convolution processing on the second resolution feature map to obtain a second convolution feature map, and performing attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weighted feature map; Performing convolution processing on the first resolution feature map to obtain a first convolution feature map, multiplying the first convolution feature map with the second weight feature map to obtain a first weight feature map, and performing feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map; Obtaining a second output feature map by concatenating the first output feature map and the third resolution feature map, and using the fourth resolution feature map as the third output feature map; Based on the current frame image, downsampling and concatenating the first output feature map, the second output feature map, and the third output feature map to obtain a first basic feature map, a second basic feature map, and a third basic feature map; A first target feature map, a second target feature map, and a third target feature map are obtained by performing convolution processing on the first basic feature map, the second basic feature map, and the third basic feature map respectively; The method of obtaining a second output feature map by cascading the first output feature map and the third resolution feature map, and using the fourth resolution feature map as the third output feature map, includes: Amplify the first output feature map by a preset multiple to obtain a first output feature amplified map, Performing a cascade process on the first output feature magnification image and the third resolution feature image to obtain a cascade result; The second output feature map is obtained by performing convolution processing on the cascade result, and the fourth resolution feature map is used as the third output feature map.
2. The feature extraction method of a video target according to claim 1, characterized in that: The convolution processing is performed on the second resolution feature map to obtain a second convolution feature map, and the attention calculation is performed on the second convolution feature map based on the first resolution feature map to obtain a second weight feature map, including: Performing convolution processing on the second resolution feature map to obtain the second convolution feature map; Enlarging the first resolution feature map by a preset multiple to obtain a first resolution feature enlarged map; Cascading the second convolution feature map with the first resolution feature magnification map to obtain a cascade processing result, and convolving the cascade processing result to obtain a second cascade feature map; The second weighted feature map is obtained by performing attention calculation on the second cascade feature map.
3. The feature extraction method of a video target according to claim 2, characterized in that: The step of performing attention calculation on the second cascade feature map to obtain the second weight feature map includes: Performing average pooling on the second cascade feature map to obtain a pooled feature map, and performing convolution on the pooled feature map to obtain an initial weighted feature map; The initial weight feature map is normalized and activated to obtain the second weight feature map.
4. The feature extraction method of a video target according to claim 1, wherein: The method includes performing convolution processing on the first resolution feature map to obtain a first convolution feature map, multiplying the first convolution feature map with the second weight feature map to obtain a first weight feature map, and performing feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map, including: Performing convolution processing on the first resolution feature map to obtain the first convolution feature map; Multiplying the first convolution feature map and the second weight feature map to obtain a first weight feature map; A first weight feature magnification graph is obtained by magnifying the first weight feature graph by a preset multiple; Based on the weight information of the second weight feature map, the first weight feature amplification map and the second convolution feature map are feature fused to obtain a first output feature map.
5. The feature extraction method of a video target according to claim 4, characterized in that: The step of performing feature fusion on the first weight feature amplification map and the second convolution feature map based on the weight information of the second weight feature map to obtain a first output feature map includes: Determining weight information of the second convolution feature map based on the weight information of the second weight feature map; Multiplying the first weighted feature amplification map and the second convolution feature map to obtain a multiplication result feature map, and adding the multiplication result feature map and the second convolution feature map to obtain a fusion feature map; Perform convolution processing on the fused feature map to obtain the first output feature map.
6. The feature extraction method of a video target according to claim 1, characterized in that: The method of downsampling and concatenating the first output feature map, the second output feature map, and the third output feature map based on the current frame image to obtain a first basic feature map, a second basic feature map, and a third basic feature map includes: Based on the current frame image, down-sampling the first output feature map, the second output feature map, and the third output feature map to obtain a first sampling result, a second sampling result, and a third sampling result; Performing cascade processing on the first sampling result and the first output feature map to obtain the first basic feature map; Cascading the second sampling result and the second output feature map to obtain the second basic feature map; The third sampling result and the third output feature map are cascaded to obtain the third basic feature map.
7. A feature extraction device for a video target, characterized in that: include: A feature map extraction module is used to obtain a current frame image in the video and extract features of the current frame image layer by layer to obtain a first resolution feature map, a second resolution feature map, a third resolution feature map, and a fourth resolution feature map in sequence; A second weighted feature map generation module is configured to perform convolution processing on the second resolution feature map to obtain a second convolution feature map, and perform attention calculation on the second convolution feature map based on the first resolution feature map to obtain a second weighted feature map; a first output feature map generating module, configured to perform convolution processing on the first resolution feature map to obtain a first convolution feature map, multiply the first convolution feature map by the second weight feature map to obtain a first weight feature map, and perform feature fusion on the first weight feature map and the second convolution feature map to obtain a first output feature map; a second output feature map generating module, configured to obtain a second output feature map by cascading the first output feature map and the third resolution feature map, and use the fourth resolution feature map as the third output feature map; a basic feature map generation module, configured to perform downsampling and cascading processing on the first output feature map, the second output feature map, and the third output feature map based on the current frame image to obtain a first basic feature map, a second basic feature map, and a third basic feature map; a target feature map generation module, configured to obtain a first target feature map, a second target feature map, and a third target feature map by performing convolution processing on the first basic feature map, the second basic feature map, and the third basic feature map, respectively; The second output feature map generation module includes: A first output feature magnified image generating unit is configured to magnify the first output feature image by a preset multiple to obtain a first output feature magnified image. a cascade result generating unit, configured to perform cascade processing on the first output feature magnification map and the third resolution feature map to obtain a cascade result; The cascade result convolution unit is used to obtain the second output feature map by performing convolution processing on the cascade result, and use the fourth resolution feature map as the third output feature map.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method for extracting features of a video target according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the feature extraction method for a video target according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and device, computer equipment and storage medium
CN111047516A
Image processing method and device, computer equipment and storage medium
CN114418909A