Video processing method, device, electronic device and medium
By obtaining the multi-scale similarity information and mask features between the reference frame and the current frame, the matching accuracy problem caused by the scale change of the target object in the video is solved, and the accuracy of video target segmentation is improved.
Patent Information
- Application Number
- CN202310356217.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-04-04
AI Technical Summary
In the prior art, the scale variation of target objects in different frames in a video leads to low accuracy of matching results.
By obtaining the multi-scale similarity information and mask features between the reference frame and the current frame, and using visual semantic features and mask representation features, the mask features of the current frame at different scales are obtained respectively, and combined with the content information features to improve the matching accuracy.
The accuracy of video target segmentation is improved, and the accuracy of obtaining mask information of target objects in the video is improved through multi-scale matching.
Smart Images

Figure CN116403142B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video technology, and in particular to a video processing method, device, electronic device, and medium. Background Art
[0002] With the development of video technology and various video platforms, users have higher and higher requirements for video editing or playback, and more and more users hope to be able to provide the function of tracking target objects in the video.
[0003] In the existing technology, the target object information in the subsequent frames is often obtained by performing similarity matching on the target object information in the given initial frame. However, in actual applications, the sizes of the target objects in different frames are often different. Under such scale changes, the accuracy of the matching results is low. Summary of the Invention
[0004] The present disclosure provides a video processing method, device, electronic device, and medium to at least address the problem of low accuracy in related technologies. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, including:
[0006] Obtaining, based on first similarity information between a reference frame of a video to be processed and a current frame and a mask representation feature of the reference frame, a first mask feature corresponding to the current frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, the first similarity information being used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame being used to represent mask information of an object to be segmented in the reference frame;
[0007] Acquire at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and a current frame having a second scale;
[0008] Based on the second similarity information and the mask representation feature, obtaining a second mask feature corresponding to the current frame;
[0009] Based on the first mask feature, the second mask feature, and the content information feature of the current frame, mask information of the object to be segmented in the current frame is acquired.
[0010] Optionally, the acquiring at least one second similarity information between the reference frame and the current frame includes:
[0011] For any scale coefficient among the at least one scale coefficient, performing a first sampling operation on information corresponding to the current frame in the first similarity information according to the any scale coefficient to obtain second similarity information between the reference frame and the current frame; the sampling coefficient of the first sampling operation is the any scale coefficient, and the second scale is a scale corresponding to the any scale coefficient;
[0012] Alternatively, for any of the scale coefficients, a first sampling operation is performed on the visual semantic features of the current frame according to the any of the scale coefficients to obtain current visual semantic features of the current frame corresponding to the second scale;
[0013] Based on the current visual semantic feature and the visual semantic feature of the reference frame, second similarity information between the reference frame and the current frame is obtained.
[0014] Optionally, the acquiring, based on the second similarity information and the mask representation feature, a second mask feature corresponding to the current frame includes:
[0015] Based on the second similarity information and the mask representation feature, obtaining an intermediate mask feature corresponding to the current frame;
[0016] performing a second sampling operation on the intermediate mask feature; wherein the sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; if the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; if the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation;
[0017] The second mask feature is acquired based on the intermediate mask feature after the second sampling operation.
[0018] Optionally, acquiring the second mask feature based on the intermediate mask feature after the second sampling operation includes:
[0019] Obtaining residual information of a target sampling operation in the first sampling operation and the second sampling operation; the target sampling operation is an upsampling operation;
[0020] The residual information is fused with the intermediate mask feature after the second sampling operation to obtain the second mask feature.
[0021] Optionally, the acquiring, based on the second similarity information and the mask representation feature, an intermediate mask feature corresponding to the current frame includes:
[0022] Converting the value of each element included in the second similarity information into a preset range to obtain normalized second similarity information;
[0023] The intermediate mask feature is determined based on the normalized second similarity information and the mask characterization feature.
[0024] Optionally, the method further includes:
[0025] Acquiring visual semantic features of the reference frame and visual semantic features of the current frame;
[0026] Acquire first similarity information between the reference frame and the current frame based on the visual semantic features of the reference frame and the visual semantic features of the current frame;
[0027] The obtaining, based on first similarity information between a reference frame of the video to be processed and the current frame and a mask representation feature of the reference frame, a first mask feature corresponding to the current frame includes:
[0028] Converting the value of each element included in the first similarity information into a preset range to obtain normalized first similarity information;
[0029] The first mask feature is obtained based on the normalized first similarity information and the mask representation feature.
[0030] Optionally, acquiring mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature, and content information features of the current frame includes:
[0031] Acquire a fusion feature based on the first mask feature, the second mask feature, and the content information feature;
[0032] The fused features are decoded to obtain decoded features, and mask information of the object to be segmented in the current frame is acquired based on the decoded features.
[0033] According to a second aspect of an embodiment of the present disclosure, there is provided a video processing apparatus, including:
[0034] A first feature acquisition module is configured to acquire a first mask feature corresponding to the current frame based on first similarity information between a reference frame of a video to be processed and a current frame and a mask representation feature of the reference frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent mask information of the object to be segmented in the reference frame;
[0035] A first similarity acquisition module is configured to acquire at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and a current frame at a second scale;
[0036] A second feature acquisition module is configured to acquire a second mask feature corresponding to the current frame based on the second similarity information and the mask representation feature;
[0037] The mask acquisition module is configured to acquire mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature and the content information feature of the current frame.
[0038] Optionally, the first similarity acquisition module is specifically configured to: perform a first sampling operation on information corresponding to the current frame in the first similarity information according to any scale coefficient of at least one scale coefficient, so as to obtain second similarity information between the reference frame and the current frame; the sampling coefficient of the first sampling operation is the any scale coefficient, and the second scale is a scale corresponding to the any scale coefficient;
[0039] Alternatively, the first similarity acquisition module includes:
[0040] A first sampling submodule is configured to perform a first sampling operation on the visual semantic features of the current frame according to any of the scale coefficients to obtain a current visual semantic feature of the current frame corresponding to a second scale;
[0041] The first similarity acquisition module is specifically configured to execute: acquiring second similarity information between the reference frame and the current frame based on the current visual semantic feature and the visual semantic feature of the reference frame.
[0042] Optionally, the second feature acquisition module includes:
[0043] an intermediate feature acquisition submodule, configured to acquire an intermediate mask feature corresponding to the current frame based on the second similarity information and the mask representation feature;
[0044] a second sampling submodule configured to perform a second sampling operation on the intermediate mask feature; a sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; if the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; if the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation;
[0045] The feature acquisition submodule is configured to acquire the second mask feature based on the intermediate mask feature after the second sampling operation.
[0046] Optionally, the feature acquisition submodule includes:
[0047] a residual information acquiring unit configured to acquire residual information of a target sampling operation in the first sampling operation and the second sampling operation; the target sampling operation is an upsampling operation;
[0048] The first fusion unit is configured to fuse the residual information with the intermediate mask feature after the second sampling operation to obtain the second mask feature.
[0049] Optionally, the intermediate feature acquisition submodule includes:
[0050] a first normalization unit configured to convert the value of each element included in the second similarity information into a preset range to obtain normalized second similarity information;
[0051] The intermediate feature determining unit is configured to determine the intermediate mask feature based on the normalized second similarity information and the mask characterization feature.
[0052] Optionally, the device further includes:
[0053] A semantic feature acquisition module is configured to acquire the visual semantic features of the reference frame and the visual semantic features of the current frame;
[0054] A second similarity acquisition module is configured to acquire first similarity information between the reference frame and the current frame based on the visual semantic features of the reference frame and the visual semantic features of the current frame;
[0055] The first feature acquisition module includes:
[0056] a second normalization submodule, configured to convert the numerical value of each element included in the first similarity information into a preset range to obtain normalized first similarity information;
[0057] The first feature acquisition module is specifically configured to acquire the first mask feature based on the normalized first similarity information and the mask representation feature.
[0058] Optionally, the mask acquisition module includes:
[0059] A second fusion submodule is configured to acquire a fusion feature based on the first mask feature, the second mask feature and the content information feature;
[0060] The mask acquisition module is specifically configured to perform: decoding the fusion feature to obtain a decoded feature, and obtaining mask information of the object to be segmented in the current frame based on the decoded feature.
[0061] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0062] processor;
[0063] a memory for storing instructions executable by the processor;
[0064] The processor is configured to execute the instructions to implement the method as described in any one of the first aspects.
[0065] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device executes the method as described in any one of the first aspects.
[0066] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes readable program instructions. When the readable program instructions are executed by a processor of an electronic device, the electronic device executes the method as described in any one of the first aspects.
[0067] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: In the embodiments of the present disclosure, a first mask feature corresponding to the current frame is obtained based on first similarity information between a reference frame and a current frame of a video to be processed and a mask representation feature of the reference frame; the first similarity information is obtained based on the visual semantic features of the reference frame and the visual semantic features of the current frame, and the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent the mask information of the object to be segmented in the reference frame; at least one second similarity information between the reference frame and the current frame is obtained; one second similarity information is used to represent the similarity between the reference frame and a current frame at a second scale; based on the second similarity information and the mask representation feature, a second mask feature corresponding to the current frame is obtained; based on the first mask feature, the second mask feature and the content information feature of the current frame, the mask information of the object to be segmented in the current frame is obtained. In this way, the matching results of the current frame and the reference frame at the first scale and the second scale can be obtained respectively through the first similarity information and the second similarity information, thereby improving the matching accuracy to a certain extent. At the same time, by obtaining the first and second mask features through the similarity information at different scales, detailed information of the object to be segmented in the current frame can be obtained at different scales, thereby improving the accuracy of obtaining the mask information in the current frame and further improving the accuracy of video target segmentation.
[0068] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0070] Figure 1 is a flowchart of a video processing method according to an exemplary embodiment of the present disclosure;
[0071] Figure 2 is a schematic diagram of a spatiotemporal memory reading module according to an exemplary embodiment of the present disclosure;
[0072] Figure 3 1 is a schematic diagram of a processing process of a spatiotemporal memory reading module according to an exemplary embodiment of the present disclosure;
[0073] Figure 4 is a block diagram of a video processing device according to an exemplary embodiment;
[0074] Figure 5 is a block diagram of a device for video processing according to an exemplary embodiment;
[0075] Figure 6 is a block diagram showing another apparatus for video processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0076] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0077] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0078] Figure 1 is a flow chart of a video processing method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the method may include the following steps:
[0079] Step 101: Obtain a first mask feature corresponding to the current frame based on first similarity information between a reference frame of a video to be processed and a current frame and a mask representation feature of the reference frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, and the first similarity information is used to represent the similarity between the reference frame and the current frame having a first scale; the mask representation feature of the reference frame is used to represent mask information of the object to be segmented in the reference frame.
[0080] The video to be processed is any video that has a target tracking requirement, and a video that receives a tracking instruction containing a tracking object can be determined as the video to be processed. The object to be segmented is the target that needs to be tracked in the video to be processed.
[0081] It should be noted that video target tracking can also be called video target segmentation (VOS), which refers to separating the target object from other content in the video by taking the target object as the foreground. Therefore, the tracked target is also the object to be segmented. Furthermore, in the process of video target segmentation, a mask is often generated for the object to be segmented in the initial frame of the video by receiving the target segmentation requirement. The mask information of the object to be segmented in subsequent frames is predicted one by one based on the mask of the initial frame, thereby achieving video target tracking or video target segmentation.
[0082] Furthermore, the above-mentioned current frame refers to a video frame that needs to predict the mask information of the object to be segmented, which can be called a query frame (query). The above-mentioned reference frame refers to a video frame that is before the current frame and has obtained the mask information of the object to be segmented, which can also be called a memory frame (memory). It can be one or more, and the embodiment of the present disclosure does not limit this. It can be understood that the object to be segmented is often moving in the video. The more reference frames there are, the more dynamic information of the object to be segmented that the reference frames can provide, and the timing information of the object to be segmented can be formed, thereby improving the accuracy of the mask prediction to a certain extent. Furthermore, in the embodiment of the present disclosure, after obtaining the mask information of the current frame, the current frame can be added to the reference frame to provide more information for subsequent video frames.
[0083] The first scale refers to the original scale of the current frame, and the first similarity information refers to the similarity information of each pixel between the current frame and the reference frame obtained by performing similarity matching between the original scaled frames. The visual semantic features refer to the depth features obtained from the image content of the video frame. The visual semantic features and mask representation features of the reference frame are obtained based on the image information and mask information of the reference frame, so that the mask representation features of the reference frame contain information about the object to be segmented, and the visual semantic features of the current frame are obtained based on the image information of the current frame.
[0084] Step 102: Obtain at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and a current frame with a second scale.
[0085] Step 103: Obtain a second mask feature corresponding to the current frame based on the second similarity information and the mask representation feature.
[0086] Among them, the above-mentioned second scale refers to a scale different from the first scale. For example, several second scales can be set to be 1 / 4, 1 / 2, 2 times, 4 times, etc. of the original scale. The number of second scales can be pre-set according to actual needs, and the embodiment of the present disclosure does not limit this. Specifically, the above-mentioned second similarity information can be obtained by enlarging or reducing the current frame to a certain multiple and then performing similarity matching with the reference frame of the original scale. Accordingly, the number of second similarity information is consistent with the number of the above-mentioned second scales. It can be understood that when the proportion of the object to be segmented in the reference frame and the current frame is different, directly using the current frame of the original scale to match the reference frame will result in poor matching effect and failure to obtain more accurate similarity information. Therefore, in the embodiment of the present disclosure, the current frame of the second scale can be matched with the reference frame, which can make the proportion of the object to be segmented in the current frame and the reference frame closer to each other to a certain extent. In this case, more accurate similarity information can be obtained.
[0087] Furthermore, the first mask feature and the second mask feature include information features required for generating mask information of the current frame. Specifically, the first similarity information and the second similarity information can be understood as the matching scores between the current frame and the reference frame. Thus, the corresponding pixel information can be read from the mask representation features of the reference frame according to the matching scores obtained for the current frame at different scales, and the information that can be used to determine the mask can be determined to obtain the first and second mask features.
[0088] Step 104: Obtain mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature, and the content information feature of the current frame.
[0089] Among them, the content information features of the above-mentioned current frame are obtained by encoding the current frame, which may include detailed content information in the current frame. Therefore, in the embodiment of the present disclosure, through the above-mentioned first mask features, the second mask features and the content information features of the current frame, the information belonging to the mask can be filtered out from the current frame to obtain the mask information in the current frame.
[0090] In addition, it should be noted that the visual semantic features of the above-mentioned different frames, the mask representation features of the reference frame, and the content information features of the current frame can be obtained through the memory encoding network (Memory Encoder) and the query encoding network (Query Encoder). Specifically, the reference frame and the mask information of the reference frame can be input into the memory encoding network to obtain the visual semantic features of the reference frame (Memory key, k M ) and mask characterization features (Memory value, v M), accordingly, the current frame can be input into the query encoding network to obtain the visual semantic features of the current frame (Query key, k Q ) and content information features (Query value, v Q ). Specifically, any deep learning basic network (backbone) can be used as the above-mentioned memory encoding network and query encoding network to extract deep features and generate corresponding Key and Value. For example, a convolutional neural network (ResNet50) can be used, which contains four convolution blocks and two separate convolution layers, so that key (H*W*C / 8) and value (H*W*C / 2) can be obtained respectively, where H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image. Furthermore, when there are multiple reference frames, the Key and Value of the reference frame also include the T dimension to represent different reference frames, that is, k M (T*H*W*C / 8), v M (T*H*W*C / 2). Of course, a lightweight neural network (Mobilenetv2 / v3) can also be used to obtain the key and value of the reference frame and the current frame, which is not limited in the present embodiment.
[0091] Specifically, the operation of obtaining the mask information in the current frame can be obtained by fusing the first mask feature, the second mask feature and the content information feature of the current frame in the C dimension, and then predicting it through a mask prediction network. The embodiment of the present disclosure does not limit this.
[0092] In summary, the video processing method provided by the embodiment of the present disclosure obtains a first mask feature corresponding to the current frame based on first similarity information between a reference frame and a current frame of a video to be processed and a mask representation feature of the reference frame; the first similarity information is obtained based on the visual semantic features of the reference frame and the visual semantic features of the current frame, and the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent the mask information of the object to be segmented in the reference frame; at least one second similarity information is obtained between the reference frame and the current frame; one second similarity information is used to represent the similarity between the reference frame and a current frame at a second scale; based on the second similarity information and the mask representation feature, a second mask feature corresponding to the current frame is obtained; based on the first mask feature, the second mask feature and the content information feature of the current frame, the mask information of the object to be segmented in the current frame is obtained. In this way, the matching results of the current frame and the reference frame at the first scale and the second scale can be obtained respectively through the first similarity information and the second similarity information, thereby improving the matching accuracy to a certain extent. At the same time, by obtaining the first and second mask features through the similarity information at different scales, detailed information of the object to be segmented in the current frame can be obtained at different scales, thereby improving the accuracy of obtaining the mask information in the current frame and further improving the accuracy of video target segmentation.
[0093] In an optional embodiment, the operation of obtaining at least one second similarity information between the reference frame and the current frame may include the following steps:
[0094] Step 201: For any scale coefficient among at least one scale coefficient, perform a first sampling operation on information corresponding to the current frame in the first similarity information according to the any scale coefficient to obtain second similarity information between the reference frame and the current frame; the sampling coefficient of the first sampling operation is the any scale coefficient, and the second scale is a scale corresponding to the any scale coefficient.
[0095] The scale factor can be pre-set and can be one or more. The first sampling operation can be a downsampling operation or an upsampling operation. It can be understood that the downsampling operation is used to reduce the dimension. Therefore, by reducing the dimensions in the H and W dimensions through the downsampling operation, the effect of reducing the scale can be achieved. Correspondingly, the upsampling operation is used to increase the dimension. Therefore, by increasing the dimensions in the H and W dimensions through the upsampling operation, the effect of enlarging the scale can be achieved.
[0096] Specifically, the above-mentioned scale coefficient can be set to an integer S greater than 1. Through upsampling or downsampling, you can choose to enlarge or reduce it by S times. The above-mentioned sampling coefficient can be selected from pre-set scale coefficients. Specifically, the sampling coefficient refers to a parameter for performing a sampling operation. By performing sampling operations according to different sampling coefficients, data of different scales can be obtained. It can be understood that one scale coefficient corresponds to one second scale.
[0097] Specifically, the first similarity information can be obtained by performing an inner product operation on the visual semantic features of the reference frame and the visual semantic features of the current frame, with the visual semantic features of the two being k and M (T*H*W*C / 8) and k Q (H*W*C / 8), the first sampling operation is downsampling, taking the scale factor as 2 as an example, the similarity information can be obtained by the inner product operation as A∈R HW *THW Before the first sampling operation, in order to facilitate subsequent calculations, A can be transformed (reshaped) to A∈R THW *1*H*W , so that the THW dimension can be fixed, and the first sampling operation is performed on the H and W dimensions of the current frame, that is, the first sampling operation is performed on the information corresponding to the current frame in the first similarity information to obtain the second similarity information A'∈R THW *1*H / 2*W / 2 , further, A' can be further reshaped to A'∈R THW*HW / 4 .
[0098] It can be seen that in this step, by performing the first sampling operation on the information corresponding to the current frame in the first similarity information after obtaining the first similarity information, the second similarity information can be directly obtained without performing the similarity operation again.
[0099] Alternatively, in step 202 , for any of the scale coefficients, a first sampling operation is performed on the visual semantic features of the current frame according to the scale coefficient to obtain current visual semantic features of the current frame corresponding to the second scale.
[0100] Step 203: Acquire second similarity information between the reference frame and the current frame based on the current visual semantic feature and the visual semantic feature of the reference frame.
[0101] Alternatively, before obtaining the first similarity information, the visual semantic features of the current frame (k Q ) performs the first sampling operation, i.e., k QAfter zooming in or out by S times in the H and W dimensions, a similarity operation is performed between the current visual semantic features at the second scale and the visual semantic features of the reference frame at the original scale to obtain second similarity information.
[0102] Specifically, the visual semantic features of the two are k M (T*H*W*C / 8) and k Q (H*W*C / 8), the first sampling operation is downsampling, taking the scale factor as 2 as an example, we can first perform k Q (H*W*C / 8) performs downsampling operation on H and W dimensions to obtain the current visual semantic feature k of the second scale Q’ (H / 2*W / 2*C / 8), then k Q’ With k M Perform inner product operation to obtain the second similarity information A'∈R THW*1*H / 2*W / 2 .
[0103] In an embodiment of the present disclosure, a first sampling operation is performed on the information corresponding to the current frame in the first similarity information according to any scale coefficient of at least one scale coefficient, so as to obtain the second similarity information between the reference frame and the current frame; the sampling coefficient of the first sampling operation is any scale coefficient, and the second scale is the scale corresponding to any scale coefficient; or, for any of the scale coefficients, a first sampling operation is performed on the visual semantic features of the current frame according to any scale coefficient to obtain the current visual semantic features corresponding to the current frame and the second scale; based on the current visual semantic features and the visual semantic features of the reference frame, the second similarity information between the reference frame and the current frame is obtained. In this way, by performing the first sampling operation on the information corresponding to the current frame in the first similarity information, the second similarity information can be obtained without an additional similarity calculation operation, thereby improving the calculation efficiency. Alternatively, the similarity information can be calculated after performing the first sampling operation on the visual semantic features of the current frame, so that the second similarity information can be obtained more intuitively.
[0104] In an optional embodiment, the operation of obtaining the second mask feature corresponding to the current frame based on the second similarity information and the mask representation feature may include the following steps:
[0105] Step 301: Based on the second similarity information and the mask representation feature, obtain the intermediate mask feature corresponding to the current frame.
[0106] Among them, the above mask characterization features refer to the above v M , is obtained through the reference frame and the mask information of the reference frame, which contains the detailed information of the mask of the reference frame.
[0107] Specifically, after obtaining the second similarity information, the similarity between the current frame and the reference frame at the second scale can be obtained. The similarity can be used to address the value feature in the reference frame, that is, according to the similarity information from v M The information that can determine the mask is selected to obtain the above-mentioned intermediate mask features.
[0108] Specifically, the method for obtaining the above-mentioned intermediate mask features can be to make the similarity information and v M Multiply them together to get .
[0109] Step 302: Perform a second sampling operation on the intermediate mask feature; the sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; when the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; when the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation.
[0110] Step 303: Acquire the second mask feature based on the intermediate mask feature after the second sampling operation.
[0111] Furthermore, since the second similarity information is enlarged or reduced in the H and W dimensions according to the scale factor S, the intermediate mask features obtained based on the second similarity information are also enlarged or reduced by S times in the H and W dimensions, while the first mask features in step 104 and the content information features of the current frame maintain the original scale in the H and W dimensions. If the mask information of the current frame is directly predicted using the intermediate mask features, the first mask features, and the content information features of the current frame, the different dimensions of the H and W dimensions of the three may lead to prediction failure or prediction error. Therefore, in order to maintain the consistency of the dimensions of the H and W dimensions of the three, the embodiment of the present disclosure can perform a second sampling operation on the intermediate mask features, where the second sampling operation refers to a sampling operation opposite to the first sampling operation. Through the second sampling operation, an intermediate mask feature that is consistent with the first mask feature and the content information features of the current frame in the H and W dimensions can be obtained.
[0112] It is understood that the first sampling operation and the second sampling operation correspond to the same scaling factor. That is, any scaling factor corresponds to a pair of the first sampling operation and the second sampling operation. Specifically, when the first sampling operation is downsampling, the second sampling operation is an upsampling operation; when the first sampling operation is upsampling, the second sampling operation is a downsampling operation.
[0113] Furthermore, the intermediate mask features after the second sampling operation may be directly used as the second mask features, or the second mask features may be obtained by further processing the intermediate mask features after the second sampling operation.
[0114] In the disclosed embodiment, an intermediate mask feature corresponding to the current frame is obtained based on the second similarity information and the mask representation feature; a second sampling operation is performed on the intermediate mask feature; the sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; if the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; if the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation; and the second mask feature is obtained based on the intermediate mask feature after the second sampling operation. In this way, by performing a second sampling operation on the intermediate mask feature, which is the opposite of the first sampling operation, the scale of the resulting second mask feature can be restored to the original scale, thereby achieving multi-scale matching while avoiding prediction errors caused by different scales.
[0115] In an optional embodiment, the operation of acquiring the second mask features based on the intermediate mask features after the second sampling operation may specifically include the following steps in the embodiment of the present disclosure:
[0116] Step 401: Obtain residual information of a target sampling operation in the first sampling operation and the second sampling operation; the target sampling operation is an upsampling operation.
[0117] Step 402: Fusing the residual information with the intermediate mask features after the second sampling operation to obtain the second mask features.
[0118] Regarding the above steps 401-402, specifically, since the upsampling operation often causes data loss or data inconsistency, in order to avoid this situation, the residual information of the upsampling operation can be obtained to ensure the consistency of feature information transmission.
[0119] Specifically, the above-mentioned residual information can be realized by setting up a residual network, such as a ResBlock module, wherein the network structure of the ResBlock module can be self-assembled and constructed according to actual needs, and the embodiments of the present disclosure do not limit this. Specifically, since the intermediate mask feature is obtained after the first and second sampling operations, the above-mentioned residual network can be directly connected to the second sampling module that performs the second sampling operation. The residual network can include several convolution modules, and the intermediate mask feature after two sampling operations is input into the residual network to obtain a convolution output through convolution processing, and the input and output of several convolution modules are added as the final output to obtain the above-mentioned second mask feature. It can be understood that the input of the convolution module refers to the above-mentioned intermediate mask feature, and the output of the convolution module refers to the above-mentioned residual information. Accordingly, the above-mentioned fusion operation refers to adding the residual information to the intermediate mask feature, that is, adding the elements of the corresponding positions of the two.
[0120] In the disclosed embodiment, residual information is obtained between the first sampling operation and the target sampling operation in the second sampling operation; the target sampling operation is an upsampling operation; and the residual information is fused with the intermediate mask features after the second sampling operation to obtain the second mask features. In this way, by obtaining the residual information and fusing it with the intermediate mask features to obtain the second mask features, problems such as data loss or inconsistency during the upsampling operation can be avoided, data consistency is maintained, and the accuracy of the second mask features can be further guaranteed.
[0121] In an optional embodiment, the operation of obtaining the intermediate mask features corresponding to the current frame based on the second similarity information and the mask representation features may include the following steps:
[0122] Step 501: Convert the value of each element included in the second similarity information into a preset range to obtain normalized second similarity information.
[0123] It should be noted that the above-mentioned first and second similarity information is actually a similarity matrix, in which any element represents the similarity value between pixels at corresponding positions in the reference frame and the current frame. Accordingly, the above-mentioned conversion of the numerical value of each element to the preset range refers to normalized mapping, so that each converted numerical value belongs to the above-mentioned preset range, making it convenient for subsequent calculations. Among them, in order to facilitate subsequent calculations, the above-mentioned preset range can be set to (0, 1), and of course it can also be set according to actual needs. The embodiment of the present invention does not limit this. For example, taking the preset range of (0, 1) as an example, the above-mentioned normalized mapping operation can be performed by the softmax module, and by normalizing the similarity information and mapping it to the (0, 1) interval, it is equivalent to assigning corresponding weights to different pixel points according to the similarity between the current frame and the reference frame.
[0124] Step 502: Determine the intermediate mask feature based on the normalized second similarity information and the mask characterization feature.
[0125] Among them, this step can be obtained by multiplying the normalized second similarity information with the above-mentioned mask representation feature. Specifically, the normalized second similarity information is equivalent to an attention mechanism, that is, focusing attention on the more important area (mask), thereby obtaining the normalized similarity information and the mask representation feature (v M ) can be multiplied from v M Useful information is filtered out, so that the mask of the current frame can be predicted according to the mask features.
[0126] In the disclosed embodiment, normalized second similarity information is obtained by converting the numerical values of each element included in the second similarity information to within a preset range; and the intermediate mask features are determined based on the normalized second similarity information and the mask characterization features. Thus, by converting the similarity information to within the preset range, the normalized similarity information can more intuitively represent the degree of similarity between each pixel in the current frame and the reference frame. Furthermore, by combining the normalized second similarity information with the mask characterization features, higher-weighted feature information in the reference frame can be obtained, thereby obtaining detailed features of the object to be segmented in the reference frame, further improving the accuracy of subsequent acquisition of mask information.
[0127] In an optional embodiment, the embodiment of the present disclosure may further include the following steps:
[0128] Step 601: Obtain visual semantic features of the reference frame and visual semantic features of the current frame.
[0129] Step 602: Obtain first similarity information between the reference frame and the current frame based on the visual semantic features of the reference frame and the visual semantic features of the current frame.
[0130] The visual semantic features of the reference frame and the current frame are referred to as the key. The visual semantic features of the different frames can be obtained through the memory encoding network (Memory Encoder) and the query encoding network (Query Encoder). Specifically, the reference frame and the mask information of the reference frame can be input into the memory encoding network to obtain the visual semantic features of the reference frame (Memory key, k M ) and mask characterization features (Memory value, v M ), accordingly, the current frame can be input into the query encoding network to obtain the visual semantic features of the current frame (Query key, k Q ) and content information features (Query value, v Q ). Specifically, the visual semantic feature k of the current frame can be used Q and the visual semantic features k of the reference frame M An inner product operation is performed to obtain the first similarity information.
[0131] The above operation of obtaining the first mask feature corresponding to the current frame based on the first similarity information between the reference frame and the current frame of the video to be processed and the mask representation feature of the reference frame may specifically include the following steps in the embodiment of the present disclosure:
[0132] Step 603: Convert the value of each element included in the first similarity information into a preset range to obtain normalized first similarity information.
[0133] Step 604: Obtain the first mask feature based on the normalized first similarity information and the mask representation feature.
[0134] Converting the values of each element to within a preset range refers to performing normalized mapping so that each converted value falls within the preset range, facilitating subsequent calculations. Specifically, the normalized mapping operation can be performed using a softmax module, which normalizes the similarity information to the interval (0, 1), effectively assigning weights to different pixels based on the similarity between the current frame and the reference frame.
[0135] The above step 604 can be obtained by multiplying the normalized first similarity information with the above mask representation feature. Specifically, the normalized first similarity information is equivalent to an attention mechanism, that is, focusing attention on the more important area (mask), thereby obtaining the normalized first similarity information and the mask representation feature (v M ) can be multiplied from v M Useful information is filtered out, so that the mask of the current frame can be predicted according to the mask features.
[0136] In the disclosed embodiment, the visual semantic features of the reference frame and the visual semantic features of the current frame are obtained; based on the visual semantic features of the reference frame and the visual semantic features of the current frame, the first similarity information between the reference frame and the current frame is obtained; the numerical values of each element contained in the first similarity information are converted to a preset range to obtain the normalized first similarity information; based on the normalized first similarity information and the mask characterization features, the first mask features are obtained. In this way, by converting the similarity information to a preset range, the normalized similarity information can more intuitively represent the degree of similarity between each pixel of the current frame and the reference frame. At the same time, by comparing the normalized first similarity information with the mask characterization features, feature information with higher weight in the reference frame can be obtained, and detailed features about the object to be segmented in the reference frame at the first scale can be obtained, further improving the accuracy of subsequent acquisition of the mask information.
[0137] In an optional embodiment, the operation of obtaining the mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature, and the content information feature of the current frame may further include the following steps:
[0138] Step 701: Acquire a fusion feature based on the first mask feature, the second mask feature, and the content information feature.
[0139] Among them, the first mask feature, the second mask feature and the content information feature obtained by the above steps are all features of the same scale, and all three contain three dimensions of H, W, and C, which are equivalent to three feature maps of the same size containing information from different angles. Therefore, this step can obtain the fusion feature of the three through a fusion operation, wherein the above fusion operation can be a fusion operation (concat) of the three in the channel, C dimension. Specifically, the fusion operation refers to splicing the features involved in the operation, and the dimension of the fusion feature in the C dimension is equivalent to the sum of the C dimensions of the first mask feature, the second mask feature and the content information feature. It can be understood that the fusion feature contains the mask information that needs to be paid attention to and the appearance information of the current frame obtained at different scales.
[0140] Step 702: Decode the fused features to obtain decoded features, and obtain mask information of the object to be segmented in the current frame based on the decoded features.
[0141] Among them, the above-mentioned decoding features can be obtained through a preset mask prediction network, and the above-mentioned prediction network can construct the mask information of the current frame through the above-mentioned fusion features. Specifically, the above-mentioned prediction network can be a pre-built decoding module (Decoder), wherein the Decoder can include a convolution layer, a normalization layer, a sampling layer, etc. Through the fusion features that include the mask information and the appearance information of the current frame, the mask information of the current frame can be gradually obtained.
[0142] In the disclosed embodiment, a fused feature is obtained based on the first mask feature, the second mask feature, and the content information feature; the fused feature is decoded to obtain a decoded feature, and mask information of the object to be segmented in the current frame is obtained based on the decoded feature. In this way, the first mask feature, the second mask feature, and the content information feature can be used to obtain a fused feature that includes the mask information of interest obtained at different scales and the appearance information of the current frame. Using this fused feature for mask prediction can obtain more accurate mask information and improve the accuracy of video object segmentation.
[0143] It should be noted that in the field of video target segmentation, the previous image frames in the video stream and their corresponding target masks are often stored in an external memory bank. When predicting the target mask for the current frame, several frames (referred to as memory frames) and their masks are first selected from the memory bank and input into the memory encoding network (Memory Encoder) to obtain the corresponding key and value deep features. The key and value form a key-value pair. The key is used for addressing, while the value stores some more detailed information used to generate the mask. The current frame is then input into the query encoding network (Query Decoder) to obtain the key and value deep features of the current frame image. The inner product operation is then performed on the key of the Memory Encoder and the key obtained by the Query Decoder to calculate the similarity, which is equivalent to a spatiotemporal attention mechanism that assigns weights to values at different times and regions. Multiply this similarity graph by the value in the Memory Encoder, which is the result of the space-time memory read. This result is concatenated with the value of the Query Decoder and sent to the final Decoder network for mask restoration prediction and calculation of the loss function during training.
[0144] In the existing space-time memory read module, spatiotemporal modeling is achieved by calculating similarity graphs between pixel pairs in the memory frame and the query frame and passing them to the network feature information, enabling mask segmentation and extraction of target objects in subsequent frames. However, this similarity graph in spatiotemporal modeling is based on single features of the memory and query frames. In actual production applications, the size of objects in the memory and query frames often varies significantly. For example, the target object in the memory frame occupies most of the image space, while the object in the query frame only occupies a small part of the image. Under such drastic scale changes, spatiotemporal modeling matching fails, resulting in inaccurate segmentation.
[0145] Figure 2 is a schematic diagram of a spatiotemporal memory reading module according to an exemplary embodiment of the present disclosure, such as Figure 2As shown, the spatiotemporal memory reading module provided by the embodiment of the present disclosure includes a fusion layer and matching layers corresponding to different scales, which can be used to calculate the similarity information and mask features obtained through the current frames of different scales. The multi-scale matching mechanism of memory and query features is implemented in this module, thereby effectively alleviating the problem of drastic scale changes of targets in video object tracking videos and improving segmentation accuracy.
[0146] Figure 3 FIG. 1 is a schematic diagram of a processing process of a spatiotemporal memory reading module according to an exemplary embodiment of the present disclosure. Figure 3 As shown, the k of Memory Encoder can be M (T*H*W*C / 8) and k obtained by Query Decoder Q (H*W*C / 8) performs inner product operation to calculate similarity, establish spatiotemporal attention, assign weights to values in different time and regions, and obtain similarity (correspondance) information A∈R HW*THW , and then two branches are generated:
[0147] Step 311: Perform a softmax operation on A to generate a similarity graph at the first scale, and then compare this similarity graph with the v in the Memory Encoder. M Multiply (T*H*W*C / 2) to get the recombined feature F1∈R H*W*C / 2 , that is, the first mask feature mentioned above.
[0148] Step 312: reshape A to A∈R THW*1*H*W , and then perform a downsampling (averagepooling) operation with stride = 2 to obtain the correspondence information A'∈R at the second scale THW*1*H / 2*W / 2 , reshape it to A'∈R THW*HW / 4 , and then compare this similarity graph with the v in Memory Encoder M Multiply (T*H*W*C / 2) to get the recombined feature F2∈R H / 2*W / 2*C / 2 , then it is upsampled by 2 times (Upsample), and then information fusion is performed through a ResBlock module to obtain F3∈R H*W*C / 2 , that is, the second mask feature mentioned above.
[0149] Finally, by combining F1, F3 and v Q The features are concat- ed in the channel dimension to obtain feature y, which is then fed into the subsequent decoder network to generate the predicted mask information for the current frame.
[0150] The spatiotemporal memory reading module provided in the embodiments of the present disclosure effectively alleviates the problem of drastic scale changes of targets in video object tracking videos by implementing a multi-scale matching mechanism of memory and query features, thereby improving segmentation accuracy.
[0151] Figure 4 is a block diagram of a video processing device according to an exemplary embodiment. Figure 4 As shown, the device 80 may include:
[0152] A first feature acquisition module 801 is configured to acquire a first mask feature corresponding to the current frame based on first similarity information between a reference frame and a current frame of a video to be processed and a mask representation feature of the reference frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent mask information of the object to be segmented in the reference frame;
[0153] A first similarity acquisition module 802 is configured to acquire at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and the current frame at a second scale;
[0154] A second feature acquisition module 803 is configured to acquire a second mask feature corresponding to the current frame based on the second similarity information and the mask representation feature;
[0155] The mask acquisition module 804 is configured to acquire mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature, and content information features of the current frame.
[0156] In an optional embodiment, the first similarity obtaining module 802 is specifically configured to execute: for any scale coefficient among at least one scale coefficient, perform a first sampling operation on information corresponding to the current frame in the first similarity information according to the any scale coefficient to obtain second similarity information between the reference frame and the current frame; the sampling coefficient of the first sampling operation is the any scale coefficient, and the second scale is a scale corresponding to the any scale coefficient;
[0157] Alternatively, the first similarity obtaining module 802 includes:
[0158] A first sampling submodule is configured to perform a first sampling operation on the visual semantic features of the current frame according to any of the scale coefficients to obtain a current visual semantic feature of the current frame corresponding to a second scale;
[0159] The first similarity acquisition module 802 is specifically configured to execute: based on the current visual semantic features and the visual semantic features of the reference frame, acquire second similarity information between the reference frame and the current frame.
[0160] In an optional embodiment, the second feature acquisition module 803 includes:
[0161] an intermediate feature acquisition submodule, configured to acquire an intermediate mask feature corresponding to the current frame based on the second similarity information and the mask representation feature;
[0162] a second sampling submodule configured to perform a second sampling operation on the intermediate mask feature; a sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; if the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; if the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation;
[0163] The feature acquisition submodule is configured to acquire the second mask feature based on the intermediate mask feature after the second sampling operation.
[0164] In an optional embodiment, the feature acquisition submodule includes:
[0165] a residual information acquiring unit configured to acquire residual information of a target sampling operation in the first sampling operation and the second sampling operation; the target sampling operation is an upsampling operation;
[0166] The first fusion unit is configured to fuse the residual information with the intermediate mask feature after the second sampling operation to obtain the second mask feature.
[0167] In an optional embodiment, the intermediate feature acquisition submodule includes:
[0168] a first normalization unit configured to convert the value of each element included in the second similarity information into a preset range to obtain normalized second similarity information;
[0169] The intermediate feature determining unit is configured to determine the intermediate mask feature based on the normalized second similarity information and the mask characterization feature.
[0170] In an optional embodiment, the device 80 further includes:
[0171] A semantic feature acquisition module is configured to acquire the visual semantic features of the reference frame and the visual semantic features of the current frame;
[0172] A second similarity acquisition module is configured to acquire first similarity information between the reference frame and the current frame based on the visual semantic features of the reference frame and the visual semantic features of the current frame;
[0173] The first feature acquisition module 801 includes:
[0174] a second normalization submodule, configured to convert the numerical value of each element included in the first similarity information into a preset range to obtain normalized first similarity information;
[0175] The first feature acquisition module is specifically configured to acquire the first mask feature based on the normalized first similarity information and the mask representation feature.
[0176] In an optional embodiment, the mask acquisition module 804 includes:
[0177] A second fusion submodule is configured to acquire a fusion feature based on the first mask feature, the second mask feature and the content information feature;
[0178] The mask acquisition module 804 is specifically configured to perform: decoding the fused features to obtain decoded features, and obtaining mask information of the object to be segmented in the current frame based on the decoded features.
[0179] In summary, the video processing device provided by the embodiment of the present disclosure obtains a first mask feature corresponding to the current frame based on first similarity information between a reference frame and a current frame of a video to be processed and a mask representation feature of the reference frame; the first similarity information is obtained based on the visual semantic features of the reference frame and the visual semantic features of the current frame, and the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent the mask information of the object to be segmented in the reference frame; at least one second similarity information is obtained between the reference frame and the current frame; one second similarity information is used to represent the similarity between the reference frame and a current frame at a second scale; based on the second similarity information and the mask representation feature, a second mask feature corresponding to the current frame is obtained; based on the first mask feature, the second mask feature and the content information feature of the current frame, the mask information of the object to be segmented in the current frame is obtained. In this way, the matching results of the current frame and the reference frame at the first scale and the second scale can be obtained respectively through the first similarity information and the second similarity information, thereby improving the matching accuracy to a certain extent. At the same time, by obtaining the first and second mask features through the similarity information at different scales, detailed information of the object to be segmented in the current frame can be obtained at different scales, thereby improving the accuracy of obtaining the mask information in the current frame and further improving the accuracy of video target segmentation.
[0180] According to one embodiment of the present disclosure, an electronic device is provided, comprising: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to implement the steps in the video processing method in any of the above embodiments when executing.
[0181] According to an embodiment of the present disclosure, a storage medium is further provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the steps in the video processing method in any of the above embodiments.
[0182] According to one embodiment of the present disclosure, a computer program product is further provided, which includes readable program instructions. When the readable program instructions are executed by a processor of an electronic device, the electronic device can perform the steps in the video processing method in any of the above embodiments.
[0183] Figure 59 is a block diagram of a device for video processing according to an exemplary embodiment. The device 900 may include a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output interface 912, a sensor component 914, a communication component 916, and a processor 920. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above-mentioned video processing method. In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 904 including instructions, and the above-mentioned instructions can be executed by the processor 920 of the device 900 to complete the above-mentioned method. Optionally, the storage medium may be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0184] Figure 6 is a block diagram showing another apparatus for video processing according to an exemplary embodiment.
[0185] The apparatus 1000 may include a processing component 1022, a memory 1032, an input / output interface 1058, a network interface 1050, and a power supply component 1026. The apparatus 1000 may be provided as a server. The application program stored in the memory 1032 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1022 is configured to execute the instructions to perform the above-described video processing method.
[0186] The user information (including but not limited to the user's device information, user personal information, etc.) and related data involved in this disclosure are all information authorized by the user or authorized by all parties.
[0187] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0188] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that: The method comprises: Obtaining, based on first similarity information between a reference frame of a video to be processed and a current frame and a mask representation feature of the reference frame, a first mask feature corresponding to the current frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, the first similarity information being used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame being used to represent mask information of an object to be segmented in the reference frame; the first scale being the original scale of the current frame; Obtaining at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and the current frame at a second scale; wherein the second scale is different from the first scale, and the second similarity information is obtained by enlarging or reducing the current frame by a certain factor and then performing similarity matching with the reference frame at the original scale; Based on the second similarity information and the mask representation feature, obtaining a second mask feature corresponding to the current frame; Based on the first mask feature, the second mask feature, and the content information feature of the current frame, mask information of the object to be segmented in the current frame is acquired.
2. The method according to claim 1, characterized in that The acquiring at least one second similarity information between the reference frame and the current frame includes: performing a first sampling operation on the visual semantic features of the current frame according to any scale coefficient to obtain current visual semantic features of the current frame corresponding to a second scale; the sampling coefficient of the first sampling operation is the any scale coefficient, and the second scale is a scale corresponding to the any scale coefficient; Based on the current visual semantic feature and the visual semantic feature of the reference frame, second similarity information between the reference frame and the current frame is obtained.
3. The method according to claim 2, characterized in that The acquiring, based on the second similarity information and the mask representation feature, a second mask feature corresponding to the current frame includes: Based on the second similarity information and the mask representation feature, obtaining an intermediate mask feature corresponding to the current frame; performing a second sampling operation on the intermediate mask feature; wherein the sampling coefficient of the second sampling operation is the same scale coefficient as the sampling coefficient of the first sampling operation; if the first sampling operation is an upsampling operation, the second sampling operation is a downsampling operation; if the first sampling operation is a downsampling operation, the second sampling operation is an upsampling operation; The second mask feature is acquired based on the intermediate mask feature after the second sampling operation.
4. The method according to claim 3, characterized in that The acquiring the second mask feature based on the intermediate mask feature after the second sampling operation includes: Obtaining residual information of a target sampling operation in the first sampling operation and the second sampling operation; the target sampling operation is an upsampling operation; The residual information is fused with the intermediate mask feature after the second sampling operation to obtain the second mask feature.
5. The method according to claim 3, characterized in that The acquiring, based on the second similarity information and the mask representation feature, the intermediate mask feature corresponding to the current frame includes: Converting the value of each element included in the second similarity information into a preset range to obtain normalized second similarity information; The intermediate mask feature is determined based on the normalized second similarity information and the mask characterization feature.
6. The method according to claim 1, characterized in that The method further comprises: Acquiring visual semantic features of the reference frame and visual semantic features of the current frame; Acquire first similarity information between the reference frame and the current frame based on the visual semantic features of the reference frame and the visual semantic features of the current frame; The obtaining, based on first similarity information between a reference frame of the video to be processed and the current frame and a mask representation feature of the reference frame, a first mask feature corresponding to the current frame includes: Converting the value of each element included in the first similarity information into a preset range to obtain normalized first similarity information; The first mask feature is obtained based on the normalized first similarity information and the mask representation feature.
7. The method according to any one of claims 1 to 6, characterized in that: The acquiring, based on the first mask feature, the second mask feature, and the content information feature of the current frame, mask information of the object to be segmented in the current frame includes: Acquire a fusion feature based on the first mask feature, the second mask feature, and the content information feature; The fused features are decoded to obtain decoded features, and mask information of the object to be segmented in the current frame is acquired based on the decoded features.
8. A video processing device, characterized in that: The device comprises: A first feature acquisition module is configured to acquire a first mask feature corresponding to the current frame based on first similarity information between a reference frame of a video to be processed and a current frame and a mask representation feature of the reference frame; the first similarity information is obtained based on visual semantic features of the reference frame and the current frame, the first similarity information is used to represent the similarity between the reference frame and the current frame at a first scale; the mask representation feature of the reference frame is used to represent mask information of the object to be segmented in the reference frame; the first scale is the original scale of the current frame; a first similarity acquisition module configured to acquire at least one second similarity information between the reference frame and the current frame; the second similarity information is used to represent the similarity between the reference frame and the current frame at a second scale; wherein the second scale is different from the first scale, and the second similarity information is obtained by enlarging or reducing the current frame by a certain factor and then performing similarity matching with the reference frame at the original scale; A second feature acquisition module is configured to acquire a second mask feature corresponding to the current frame based on the second similarity information and the mask representation feature; The mask acquisition module is configured to acquire mask information of the object to be segmented in the current frame based on the first mask feature, the second mask feature and the content information feature of the current frame.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device executes the method according to any one of claims 1 to 7.