Video Object Segmentation Method, Apparatus, and Electronic Device

By extracting and fusion of redundant feature information in the video target segmentation model and adjusting the coding network resolution using hollow convolution, the robustness and generalization of the model in various scenarios is solved, and the segmentation effect is improved.

CN115424184BActive Publication Date: 2025-07-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211241410.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-08
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

The existing video target segmentation model is poorly robust and generalized in a variety of scenarios, which hinders the further improvement of model performance.

Method used

By obtaining the current image frame of the video and its previous image frames and target masks, the encoding process is performed and inputting the spatiotemporal memory network is input to the spatiotemporal memory network, redundant feature information is extracted, and then fused with the third feature information is input to the decoding network for target segmentation. The resolution of the memory and query the encoding network is adjusted by using hollow convolution to obtain deeper semantic information.

Benefits of technology

It enhances the robustness and generalization of the video target segmentation model, improves the ability to distinguish target objects and backgrounds, and improves the representation ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424184B_ABST
    Figure CN115424184B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video object segmentation method, an apparatus, and an electronic device. The method includes: obtaining a current image frame of a video to be processed, at least one image frame in front of the current image frame, and object masks of the at least one image frame; performing encoding processing on the at least one image frame, the object masks of the at least one image frame, and the current image frame to obtain first encoded features of the at least one image frame and second encoded features of the current image frame; inputting the first encoded features and the second encoded features into a spatio-temporal memory network to obtain third feature information for predicting an object mask of the current image frame; extracting redundant feature information from the third feature information, where the redundant feature information is feature information with a similarity degree between different channels of the third feature information exceeding a preset value; and inputting the fused third feature information and redundant feature information into a decoding network to obtain an object mask of the current image frame, where the object mask is used for object segmentation of the current image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing, and in particular, to a video object segmentation method, an apparatus, and an electronic device. Background Art

[0002] Video Object Segmentation (VOS) is a fundamental ability for video scene understanding and video editing, and this technology has broad application prospects in fields such as intelligent short video editing, special effect production, and short video creation. The VOS technology refers to given a target object mask in the initial image frame of a certain video sequence, predicting the pixel-level segmentation mask result of the target object in subsequent image frames. With the development of deep learning, deep neural networks are applied to VOS. The high-level semantic features extracted based on the deep network can more accurately distinguish the target object and the background from complex scenes, thus greatly improving the effect of object segmentation. Therefore, the VOS segmentation technology based on deep learning has become one of the mainstream technologies.

[0003] In current video object segmentation models, generally, the output of the space-time memory network included in the model is directly fed into the subsequent decoder network in the model for final mask prediction; in this way, it often leads to poor robustness and generalization of the model in various scenarios, thus hindering the further improvement of the model performance. Summary of the Invention

[0004] The present disclosure provides a video object segmentation method, an apparatus, and an electronic device to at least solve the problem of poor robustness and generalization of the video object segmentation model in the related art.

[0005] According to a first aspect of an embodiment of the present disclosure, a video object segmentation method is provided, including: obtaining a current image frame of a video to be processed, at least one image frame corresponding to the current image frame, and the target mask of the at least one image frame, where the position of the at least one image frame in the video to be processed is before the current image frame; performing encoding processing on the at least one image frame, the target mask of the at least one image frame, and the current image frame to obtain first encoding features of the at least one image frame and second encoding features of the current image frame; inputting the first encoding features and the second encoding features into a space-time memory network to obtain third feature information for predicting the target mask of the current image frame; extracting redundant feature information from the third feature information, where the redundant feature information is feature information with a similarity degree exceeding a preset value between different channels of the third feature information; and inputting the fused third feature information and redundant feature information into a decoding network to obtain the target mask of the current image frame, where the target mask is used for object segmentation of the current image frame.

[0006] Optionally, extracting redundant feature information from the third feature information includes: inputting the third feature information into a redundant feature acquisition network to obtain first redundant feature information; performing normalization processing on the first redundant feature information to obtain second redundant feature information; and processing the second redundant feature information through an activation function to obtain redundant feature information.

[0007] Optionally, before inputting the third feature information into the redundant feature acquisition network to obtain first redundant feature information, it further includes: performing dimensionality reduction processing on the third feature information.

[0008] Optionally, encoding at least one image frame, the target mask of at least one image frame, and the current image frame to obtain the first encoded features of at least one image frame and the second encoded features of the current image frame includes: inputting at least one image frame and the target mask of at least one image frame into a memory encoding network to obtain a first key-value pair as the first encoded features, where the first key-value pair includes first feature information and first key information, the first feature information includes the encoded information of at least one image frame and the encoded information of the target mask, and the first key information includes addressing information for querying the first feature information; inputting the current image frame into a query encoding network to obtain a second key-value pair as the second encoded features, where the second key-value pair includes second feature information and second key information, the second feature information includes the encoded information of the current image frame, and the second key information includes addressing information for querying the second feature information.

[0009] Optionally, the memory encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting at least one image frame and the target mask of at least one image frame into the memory encoding network to obtain a first key-value pair, it further includes: adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network to obtain an adjusted memory encoding network, and the resolution of the output result of the Nth stage module of the adjusted memory encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; inputting at least one image frame and the target mask of at least one image frame into the memory encoding network to obtain a first key-value pair, including: inputting at least one image frame and the target mask of at least one image frame into the adjusted memory encoding network to obtain a first key-value pair with the first resolution.

[0010] Optionally, the query encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting the current image frame into the query encoding network to obtain the second key-value pair, it further includes: adjusting the parameters of the dilated convolution in the Nth stage module of the query encoding network to obtain an adjusted query encoding network, and the resolution of the output result of the Nth stage module of the adjusted query encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; inputting the current image frame into the query encoding network, and the obtained second key-value pair includes: inputting the current image frame into the adjusted query encoding network to obtain the second key-value pair with the first resolution.

[0011] Optionally, adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network includes: when the size parameter of the convolutional kernel in the Nth stage module remains unchanged, adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network so that the resolution of the output result of the adjusted Nth stage module is the same as the first resolution.

[0012] According to the second aspect of the embodiments of the present disclosure, there is provided a video object segmentation device, including: an image frame acquisition unit configured to acquire the current image frame of the video to be processed, at least one image frame corresponding to the current image frame, and the target mask of the at least one image frame, where the position of the at least one image frame in the video to be processed is before the current image frame; an encoding unit configured to perform encoding processing on the at least one image frame, the target mask of the at least one image frame, and the current image frame to obtain the first encoded feature of the at least one image frame and the second encoded feature of the current image frame; a third feature information acquisition unit configured to input the first encoded feature and the second encoded feature into a spatio-temporal memory network to obtain third feature information for predicting the target mask of the current image frame; a redundant feature information acquisition unit configured to extract redundant feature information from the third feature information, where the redundant feature information is the feature information whose similarity degree between different channels of the third feature information exceeds a preset value; a target mask acquisition unit configured to input the fused third feature information and redundant feature information into a decoding network to obtain the target mask of the current image frame, where the target mask is used for object segmentation of the current image frame.

[0013] Optionally, the redundant feature information acquisition unit is further configured to input the third feature information into a redundant feature acquisition network to obtain first redundant feature information; perform normalization processing on the first redundant feature information to obtain second redundant feature information; and process the second redundant feature information through an activation function to obtain redundant feature information.

[0014] Optionally, the redundant feature information acquisition unit is further configured to perform dimensionality reduction processing on the third feature information before inputting the third feature information into the redundant feature acquisition network to obtain the first redundant feature information.

[0015] Optionally, the encoding unit is further configured to input at least one image frame and the target mask of at least one image frame into the memory encoding network to obtain a first key-value pair as the first encoded feature, where the first key-value pair includes first feature information and first key information, the first feature information includes the encoded information of at least one image frame and the encoded information of the target mask, and the first key information includes addressing information for querying the first feature information; input the current image frame into the query encoding network to obtain a second key-value pair as the second encoded feature, where the second key-value pair includes second feature information and second key information, the second feature information includes the encoded information of the current image frame, and the second key information includes addressing information for querying the second feature information.

[0016] Optionally, the memory encoding network includes N stage modules, and one convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. The encoding unit is further configured to adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network to obtain an adjusted memory encoding network, and the resolution of the output result of the Nth stage module of the adjusted memory encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; input at least one image frame and the target mask of at least one image frame into the adjusted memory encoding network to obtain a first key-value pair with the first resolution.

[0017] Optionally, the query encoding network includes N stage modules, and one convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. The encoding unit is further configured to adjust the parameters of the dilated convolution in the Nth stage module of the query encoding network to obtain an adjusted query encoding network, and the resolution of the output result of the Nth stage module of the adjusted query encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; input the current image frame into the adjusted query encoding network to obtain a second key-value pair with the first resolution.

[0018] Optionally, the encoding unit is configured to adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network while keeping the size parameter of the convolutional kernel in the Nth stage module unchanged, so that the resolution of the output result of the adjusted Nth stage module is the same as the first resolution.

[0019] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the video object segmentation method according to the present disclosure.

[0020] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium. When instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the video object segmentation method according to the present disclosure as described above.

[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer instructions which, when executed by a processor, implement the video object segmentation method according to the present disclosure.

[0022] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0023] According to the video object segmentation method, device, and electronic device of the present disclosure, redundant feature information in the third feature information output by the spatio-temporal memory network is acquired, where the redundant feature information is feature information with a similarity degree between different channels of the third feature information exceeding a preset value. Based on the redundant feature information and the third feature information output by the spatio-temporal memory network, the target mask of the current image frame is jointly predicted. That is, the proportion of similar features in the third feature information output by the spatio-temporal memory network is amplified, so that the internal information of the third feature information can be better fused, the representation ability of the video object segmentation model is enhanced, and the robustness and generalization ability of the video object segmentation model are improved. Therefore, the present disclosure solves the problem of poor robustness and generalization ability of the video object segmentation model in the related art.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0026] Figure 1 is a schematic structural diagram of a video object segmentation model showing an exemplary embodiment according to the present disclosure;

[0027] Figure 2 is a structural diagram of an encoding network shown according to an exemplary embodiment;

[0028] Figure 3 is a structural diagram of a spatio-temporal memory network shown according to an exemplary embodiment;

[0029] Figure 4 is a schematic diagram of an implementation scenario of the video object segmentation method showing an exemplary embodiment according to the present disclosure;

[0030] Figure 5It is a flowchart of a video object segmentation method shown according to an exemplary embodiment;

[0031] Figure 6 It is a structural diagram of an adjusted memory encoding network shown according to an exemplary embodiment;

[0032] Figure 7 It is a structural diagram of an adjusted query encoding network shown according to an exemplary embodiment;

[0033] Figure 8 It is a schematic diagram of obtaining redundant feature information shown according to an exemplary embodiment;

[0034] Figure 9 It is a schematic diagram of obtaining redundant feature information including dimensionality reduction processing shown according to an exemplary embodiment;

[0035] Figure 10 It is a structural diagram of an improved spatio-temporal memory network shown according to an exemplary embodiment;

[0036] Figure 11 It is a block diagram of a video object segmentation device shown according to an exemplary embodiment;

[0037] Figure 12 It is a block diagram of an electronic device 1200 according to an embodiment of the present disclosure. Detailed implementation manners

[0038] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0040] It should be noted here that "at least one of several items" in this disclosure all represents three parallel situations, namely, "any one of the several items", "any combination of multiple items among the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including both A and B. Another example is "performing at least one of step one and step two", which means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing both step one and step two.

[0041] The general process of the current video object segmentation model for video object segmentation is as Figure 1 shown. The image frames in front of the current image frame in the video stream and their corresponding object masks (masks) are saved in an external memory bank. When predicting the object mask of the current image frame, first, several image frames (denoted as memory image frames) and their masks are selected from the above external memory bank and input into the Memory Encoder to obtain the corresponding key and value (the key and value form a key-value pair, where the role of the key is to address, and the value stores some more detailed information for generating the mask); and the current image frame is input into the Query Encoder to obtain the key and value of the current image frame. Secondly, several keys output by the Memory Encoder and the keys output by the Query Encoder are input into the Space-time Memory (sometimes also called Space-time Memory Read) to obtain the feature information for predicting the object mask of the current image frame, and this feature information is sent into the final Decoder network for object mask prediction. It should be noted that the training process of the model is also like this. The loss function is calculated based on the predicted mask and the actual mask, and the parameters of the model are adjusted through this loss function.

[0042] Among them, the processing processes of the Memory Encoder and the Query Encoder are as Figure 2 shown. The Memory encoder and the Query encoder first extract deep features through any deep learning backbone network (taking ResNet50 as an example in the figure), and then separate into two parallel branches, each generating its corresponding Key and Value through a 3x3 convolutional layer.

[0043] Among them, the processing process of the Space-time Memory (sometimes also called Space-time Memory Read) is as follows Figure 3 shown. After obtaining the keys and values output by the Memory Encoder and the keys and values output by the Query Encoder, these information are processed by the Space-time Memory Read module to assist in accurate pixel-wise mask prediction. Specifically, first, the key of the Query Encoder and each key of the Memory Encoder are respectively multiplied matrix-wise to obtain a similarity map, and further constrained to the range of 0-1 through a normalization (softmax) operation; then, the normalized similarity map is multiplied matrix-wise with the value output by the Memory Encoder to obtain an intermediate result, which is equivalent to assigning a time-space weight matrix to each value (i.e., values at different times and regions); finally, the value of the Query Encoder and the above intermediate result are concatenated in the channel dimension as the result y read from the Memory Encoder and sent to the subsequent Decoder network for the final mask prediction.

[0044] However, as Figure 3 shown, in the related art, the Space-time Memory Network directly sends the result y into the subsequent Decoder network for the final mask prediction. In this way, y does not perform good information fusion, resulting in the model not having good robustness and generalization ability for a variety of scenarios, thus hindering the further improvement of the model performance. Moreover, as Figure 2 shown, when the Memory Encoding Network and the Query Encoding Network generate their respective Keys and Values, only 3 stage modules of the basic network (such as ResNet50) are used, namely res2, res3, and res4, and the resolutions of the features output by each module are 1 / 4, 1 / 8, and 1 / 16 of the input image respectively. However, generally speaking, the basic network of the Space-time Memory Network includes 4 stage modules. For example, ResNet50 also includes the res5 module, and the resolution of the features it outputs is 1 / 32 of the input image. Since the deeper the network, the lower the resolution and the richer the deep semantic information. Related technologies such as Figure 2The features output by res4 are selected to generate key-value pairs (key and value) because the features output by res4 retain a relatively high resolution (1 / 16) and also have certain semantic information, achieving a compromise between resolution and deep semantic information. However, this method discards the last stage module, resulting in the inability to fully utilize the rich deep semantic information of the base model (such as ResNet50), and thus making it difficult to further improve the overall performance of the video object segmentation model.

[0045] To address the above problems, the present disclosure provides a video object segmentation method that can improve the robustness and generalization of the video object segmentation model. The following will be described by taking the scenario of video object segmentation as an example.

[0046] Figure 4 It is a schematic diagram of an implementation scenario showing a video object segmentation method according to an exemplary embodiment of the present disclosure. As Figure 4 described, this implementation scenario includes a server 100, a user terminal 110, and a user terminal 120. Among them, the number of user terminals is not limited to 2, and includes but is not limited to devices such as mobile phones and personal computers. The user terminal can be equipped with a camera for acquiring videos. The server can be a single server, a server cluster composed of several servers, or a cloud computing platform or a virtualization center.

[0047] The user terminal 110 or the user terminal 120 obtains a video through a camera. When the user wants to view the puppy in the video, at this time, the user uploads the video as a video to be processed to the server 100 through the user terminal 110 or the user terminal 120. After receiving the video to be processed, the server 100 processes it frame by frame. Taking the processing of the second image frame of the video as an example: The server 100 receives the annotation information of the first image frame in the video. The annotation information can be manually annotated, that is, the puppy in the first image frame is annotated. The annotation information can include the target mask of the image frame. At this time, the server 100 takes the first image frame as the memory image frame, stores the first image and its corresponding target mask in the memory, and then inputs the first image frame and its corresponding target mask into the memory encoding network to obtain the first key-value pair of the first image frame, where the first key-value pair includes the first feature information and the first key information. The first feature information includes the encoding information of the first image frame and the encoding information of the target mask. The first key information includes the addressing information for querying the first feature information, and inputs the current image frame into the query encoding network to obtain the second key-value pair of the current image frame, where the second key-value pair includes the second feature information and the second key information. The second feature information includes the encoding information of the current image frame. The second key information includes the addressing information for querying the second feature information. The server 100 then inputs the first key-value pair and the second key-value pair into the spatio-temporal memory network to obtain the third feature information for predicting the target mask of the current image frame; obtains the redundant feature information in the third feature information, where the redundant feature information is the feature information whose similarity between different channels of the third feature information exceeds a preset value; inputs the fused third feature information and redundant feature information into the decoding network to obtain the target mask of the current image frame; based on the target mask of the current image frame, performs target segmentation on the current image frame of the video to be processed. At the same time, when the segmentation effect corresponding to the target mask of the second image frame is relatively good, the second image frame and its corresponding target mask can also be stored in the memory as the memory image frame.

[0048] It should be noted that the user terminal 110 and the user terminal 120 can also complete this work independently without the server 100, and the present disclosure does not limit this.

[0049] Next, a video target segmentation method, device, and electronic device according to an exemplary embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.

[0050] Figure 5 is a flowchart of a video target segmentation method shown according to an exemplary embodiment. As Figure 5 shown, the video target segmentation method includes the following steps:

[0051] In step S501, the current image frame of the video to be processed, at least one image frame corresponding to the current image frame, and the target masks of the at least one image frame are obtained, where the positions of the at least one image frame in the video to be processed are before the current image frame. The video to be processed can be any type of video, and the present disclosure does not limit this. The target masks of the above at least one image frame can be obtained by manually annotating the target objects in the corresponding image frames, or can be obtained by the method of the present disclosure, and the present disclosure does not limit this. Generally, the target mask of the first image frame is obtained by manually annotating the target objects in the corresponding image frame.

[0052] In step S502, the at least one image frame, the target masks of the at least one image frame, and the current image frame are encoded to obtain the first encoded features of the at least one image frame and the second encoded features of the current image frame. In this step, generally, the at least one image frame and the current image frame are encoded separately, but the present disclosure does not limit this.

[0053] According to an exemplary embodiment of the present disclosure, encoding the at least one image frame, the target masks of the at least one image frame, and the current image frame to obtain the first encoded features of the at least one image frame and the second encoded features of the current image frame includes: inputting the at least one image frame and the target masks of the at least one image frame into a memory encoding network to obtain a first key-value pair as the first encoded features, where the first key-value pair includes first feature information and first key information, the first feature information includes the encoded information of the at least one image frame and the encoded information of the target masks, and the first key information includes addressing information for querying the first feature information; inputting the current image frame into a query encoding network to obtain a second key-value pair as the second encoded features, where the second key-value pair includes second feature information and second key information, the second feature information includes the encoded information of the current image frame, and the second key information includes addressing information for querying the second feature information. According to this embodiment, the memory encoding network and the query encoding network are respectively introduced to obtain the corresponding first key-value pair and second key-value pair, so that relatively accurate third feature information can be obtained subsequently.

[0054] Specifically, the above memory encoding network can be as Figure 2 shown, the first feature information can be the Value output by the memory encoding network, and the first key information can be the Key output by the memory encoding network. The present disclosure does not limit this. Specifically, the at least one image frame and the target masks of the at least one image frame in the memory encoding network (Memory encoder), first pass through any deep learning backbone network ( Figure 2Taking ResNet50 as an example, deep features are extracted, and then two parallel branches are separated, each passing through a 3x3 convolutional layer to generate their respective corresponding Key and Value.

[0055] According to an exemplary embodiment of the present disclosure, the memory encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting at least one image frame and the target mask of at least one image frame into the memory encoding network to obtain the first key-value pair, it further includes: adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network to obtain an adjusted memory encoding network, and the resolution of the output result of the Nth stage module of the adjusted memory encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; inputting at least one image frame and the target mask of at least one image frame into the memory encoding network to obtain the first key-value pair, including: inputting at least one image frame and the target mask of at least one image frame into the adjusted memory encoding network to obtain the first key-value pair with the first resolution. According to this embodiment, the resolution of the output of the Nth stage module of the memory encoding network is adjusted to be the same as the resolution of the output of the (N - 1)th stage module. Therefore, on the basis of maintaining the resolution of the generated key-value pair unchanged, that is, on the basis of ensuring that the output resolution of the memory encoding network is relatively high, deeper semantic information can be obtained, the representation ability of the network can be enhanced, the ability of the model to distinguish target objects and the background can be improved, and further the robustness and generalization of the model can be increased.

[0056] For example, taking N equal to 4 as an example, Figure 6 is a structural diagram of an adjusted memory encoding network shown according to an exemplary embodiment, as Figure 6 shown, the adjusted memory encoding network includes 4 stage modules, and the resolution of the output of the adjusted memory encoding network is the same as the resolution of the output of the 3rd stage module, both being 1 / 16. Specifically, at least one image frame and the target mask of at least one image frame in the memory encoding network first pass through any deep learning backbone network ( Figure 6 still taking ResNet50 as an example in

[0057] According to an exemplary embodiment of the present disclosure, adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network includes: when the size parameter of the convolutional kernel in the Nth stage module remains unchanged, adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network so that the resolution of the output result of the adjusted Nth stage module is consistent with the first resolution. According to this embodiment, when adjusting the Nth stage module, while ensuring that the size of the convolutional kernel remains unchanged, the adjusted memory encoding network can process the input image according to the original framework logic, thereby continuing the advantages of the original framework logic.

[0058] For example, still taking Figure 6 the memory encoding network structure shown as an example. Generally speaking, in the last stage (i.e., the above-mentioned Nth stage module) of the deep learning basic network, there is usually a convolutional layer conv (convolution), and this convolutional layer conv generally includes but is not limited to the following parameters: stride = s1, dilation = d1, padding = p1, kernel = k1, where the kernel parameter represents the size of the convolutional kernel, and the input can be of int type, such as k1 = 3 representing that the height = width = 3 of the convolutional kernel, or of tuple type, such as k1 = (3, 5) representing that the height = 3 and width = 5 of the convolutional kernel; the stride parameter represents the stride of the convolutional kernel, with a default value of 1, and the input can be of int type or tuple type. It should be noted that if it is of tuple type, the first int is for the height dimension and the second int is for the width dimension; the padding parameter represents the situation of padding zeros around the input feature matrix, with a default value of 0, and the input can also be of int type, such as p1 = 1 representing adding one row of 0 elements in the up and down directions and one column of 0 elements in the left and right directions (i.e., adding a circle of 0), and the input can also be of tuple type, such as p1 = (2, 1) representing adding two rows above and two rows below, one column on the left, and one column on the right; the dilation parameter is somewhat similar to the stride, and its actual meaning is: a filter with gaps between each point.

[0059] Assume that the input size of this convolutional layer conv is denoted as w in , and the output size is w out , then:

[0060]

[0061] Generally speaking, the above parameter combinations are in the following two forms:

[0062] (1) When s1 = 2, k1 = 3, d1 = 1, p1 = 1, the following relationship is obtained:

[0063]

[0064] (2) When s1 = 2, k1 = 5, d1 = 1, and p1 = 2, the following relationship is obtained:

[0065]

[0066] Therefore, the output size is approximately reduced to 1 / 2 of the input size. In order not to reduce the resolution, the principle of dilated convolution can be utilized, that is, adjust the parameters other than the kernel parameter size, such as adjusting the size of the dilation parameter. Without reducing the resolution while maintaining the same receptive field, the modified parameters can be:

[0067] (1) When s1 = 1, k1 = 3, d1 = 2, and p1 = 2, the following relationship is obtained

[0068]

[0069] (2) When s1 = 1, k1 = 5, d1 = 2, and p1 = 4, the following relationship is obtained

[0070]

[0071] In this way, since the size of the convolution kernel k1 is not changed by this adjustment method, the memory encoding network after adjusting the parameters can still normally read the pre-trained model weight information of the basic network, that is, the trained parameters of the model without adjustment can be used normally.

[0072] It should be noted that the present disclosure is not limited to the above two combinations of adjusted parameters, and can also be based on w in and w out equality relationship. First, let s1 = 1, and then adjust the two parameters d1 and p1 to achieve the same purpose.

[0073] More specifically, the above query encoding network can be as Figure 2 shown. The second feature information can be the Value output by the query encoding network, and the second key information can be the Key output by the query encoding network. In this regard, the present disclosure does not make any limitations. Specifically, in the memory encoding network (Memory encoder), the current image frame first passes through any deep learning backbone network ( Figure 2 taking ResNet50 as an example in

[0074] According to an exemplary embodiment of the present disclosure, the query encoding network includes N stage modules, and one convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting the current image frame into the query encoding network to obtain the second key-value pair, it further includes: adjusting the parameters of the dilated convolution in the Nth stage module of the query encoding network to obtain an adjusted query encoding network, and the resolution of the output result of the Nth stage module of the adjusted query encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; inputting the current image frame into the query encoding network, the obtained second key-value pair includes: inputting the current image frame into the adjusted query encoding network to obtain the second key-value pair with the first resolution. According to this embodiment, the resolution of the output of the Nth stage module of the query encoding network is adjusted to be the same as the resolution of the output of the (N - 1)th stage module. Therefore, on the basis that the resolution of the generated key-value pair remains unchanged, that is, on the basis of ensuring that the output resolution of the query encoding network is relatively high, deeper semantic information can be obtained, the representation ability of the network can be enhanced, the ability of the model to distinguish target objects and backgrounds can be improved, and further the robustness and generalization ability of the model can be increased.

[0075] For example, taking N equal to 4 as an example, Figure 7 is a structural diagram of an adjusted query encoding network shown according to an exemplary embodiment, as Figure 7 shown, the adjusted query encoding network includes 4 stage modules, and the resolution of the output of the adjusted query encoding network is the same as the resolution of the output of the 3rd stage module, both being 1 / 16. Specifically, in the query encoding network, the current image frame first passes through an arbitrary deep learning backbone network ( Figure 6 still taking ResNet50 as an example) to extract deep features, and then is separated into two parallel branches, each generating its corresponding Key and Value through a 3x3 convolutional layer.

[0076] According to an exemplary embodiment of the present disclosure, adjusting the parameters of the dilated convolution in the Nth stage module of the query encoding network includes: when the size parameter of the convolutional kernel in the Nth stage module remains unchanged, adjusting the parameters of the dilated convolution in the Nth stage module of the query encoding network so that the resolution of the output result of the adjusted Nth stage module is the same as the first resolution. According to this embodiment, when adjusting the Nth stage module, while ensuring that the size of the convolutional kernel remains unchanged, the adjusted query encoding network can process the input image according to the original framework logic, thereby continuing the advantages of the original framework logic.

[0077] Still taking Figure 7Take the query encoding network structure shown as an example. Generally speaking, in the last stage of the deep learning basic network (i.e., the above-mentioned Nth stage module), there is usually a convolutional layer conv (convolution), and this convolutional layer conv generally includes but is not limited to the following parameters: stride = s1, dilation = d1, padding = p1, kernel = k1. Among them, the kernel parameter represents the size of the convolutional kernel. The input can be of int type. For example, k1 = 3 means the height = width = 3 of the convolutional kernel, or it can be of tuple type. For example, k1 = (3, 5) means the height = 3 and width = 5 of the convolutional kernel; the stride parameter represents the stride of the convolutional kernel, with a default value of 1. The input can be of int type or tuple type. It should be noted that if it is of tuple type, the first int is used for the height dimension and the second int is used for the width dimension; the padding parameter represents the situation of padding zeros around the input feature matrix, with a default value of 0. Similarly, the input can be of int type. For example, p1 = 1 means adding one row of 0 elements in the up and down directions and one column of 0 pixels in the left and right directions (i.e., adding a circle of 0), and the input can also be of tuple type. For example, p1 = (2, 1) means adding two rows above and two rows below, adding one column on the left, and adding one column on the right; the dilation parameter is somewhat similar to the stride, and its actual meaning is: a filter with gaps between each point.

[0078] Assume that the input size of the convolutional layer conv is denoted as w in , and the output size is w out , then:

[0079]

[0080] Generally speaking, the above parameter combinations are in the following two forms:

[0081] (1) When s1 = 2, k1 = 3, d1 = 1, p1 = 1, the following relationship is obtained:

[0082]

[0083] (2) When s1 = 2, k1 = 5, d1 = 1, p1 = 2, the following relationship is obtained:

[0084]

[0085] Therefore, the output size is approximately reduced to 1 / 2 of the input size. In order not to reduce the resolution, the principle of dilated convolution can be used, that is, adjust the size of the parameters other than the kernel parameter. For example, adjust the size of the dilation parameter. Without reducing the resolution while maintaining the same receptive field, the modified parameters can be:

[0086] (1) When s1 = 1, k1 = 3, d1 = 2, and p1 = 2, the following relationship is obtained

[0087]

[0088] (2) When s1 = 1, k1 = 5, d1 = 2, and p1 = 4, the following relationship is obtained

[0089]

[0090] In this way, since this adjustment method does not change the size of the convolutional kernel k1, the memory encoding network after adjusting the parameters can still normally read the pre-trained model weight information of the basic network, that is, it can normally use the trained parameters of the model without adjustment.

[0091] It should be noted that the present disclosure is not limited to the above two combinations of adjusted parameters, and can also be based on w in and w out equality relationship, first set s1 = 1, and then adjust the two parameters d1 and p1 to achieve the same purpose.

[0092] Return Figure 5 , in step S503, the first encoded feature and the second encoded feature are input into the spatio-temporal memory network to obtain third feature information for predicting the target mask of the current image frame.

[0093] For example, the spatio-temporal memory network can be as Figure 3 shown. First, the key of the Query Encoder and each key of the Memory Encoder are respectively multiplied by a matrix to obtain a similarity map, and then it is further constrained to the range of 0-1 through a normalization (softmax) operation; then, the normalized similarity map is multiplied by the value output by the Memory Encoder through a matrix multiplication to obtain an intermediate result, which is equivalent to assigning a spatio-temporal weight matrix to each value (that is, different time and region values); finally, the value of the Query Encoder and the above intermediate result are concatenated in the channel dimension as the result y read from the Memory Encoder, that is, the above third feature information.

[0094] In step S504, redundant feature information in the third feature information is extracted, where the redundant feature information is feature information whose similarity degree between different channels of the third feature information exceeds a preset value. For example, if the third feature information contains 6 feature maps, which are respectively feature Figure 1 、featureFigure 2 , Feature Figure 3 , Feature Figure 4 , Feature Figure 5 and Feature Figure 6 For example, among them, Feature Figure 1 and Feature Figure 3 are two feature maps whose similarity exceeds a preset value, and Feature Figure 2 and Feature Figure 4 are two feature maps whose similarity exceeds a preset value. At this time, Feature Figure 1 and Feature 3 can form a feature set, and Feature Figure 2 and Feature Figure 4 can form a feature set, and the redundant feature information includes at least one feature map in each feature set. For example, it can include Feature Figure 1 and Feature Figure 2 , it can include Feature Figure 1 and Feature Figure 4 , and it can also include Feature Figure 1 , Feature Figure 3 and Feature Figure 2 , and the present disclosure does not limit this.

[0095] According to an exemplary embodiment of the present disclosure, extracting redundant feature information from the third feature information includes: inputting the third feature information into a redundant feature acquisition network to obtain first redundant feature information; performing normalization processing on the first redundant feature information to obtain second redundant feature information; and processing the second redundant feature information through an activation function to obtain redundant feature information. According to this embodiment, through the redundant feature acquisition network, normalization processing, and activation function, redundant feature information can be obtained conveniently and quickly.

[0096] For example, Figure 8 is a schematic diagram of obtaining redundant feature information shown according to an exemplary embodiment. As Figure 8 shown, the above-mentioned redundant feature acquisition network can be a 3x3 depth convolution, or other convolutions, such as grouped convolution. The role of convolution here is to generate the first redundant feature information, that is, to obtain similar feature maps. Therefore, as long as it can obtain redundant feature information in a low-cost manner. The above-mentioned normalization processing can use batch normalization, or other methods, and the present disclosure does not limit this. The above-mentioned activation function can use the ReLU function, or other activation functions, and the present disclosure also does not limit this.

[0097] Specifically, after the third feature information output by the spatio-temporal memory network as shown in Figure 3 can be input into Figure 8The module including 3x3 Depthwise convolution, Batch normalization, and ReLU function as shown obtains redundant feature information.

[0098] According to an exemplary embodiment of the present disclosure, before inputting the third feature information into the redundant feature acquisition network to obtain the first redundant feature information, it further includes: performing dimensionality reduction processing on the third feature information. According to this embodiment, before obtaining the redundant feature information, dimensionality reduction processing can be performed on the third feature information first, which can reduce the computational amount in the process of obtaining the redundant feature information and reduce costs.

[0099] For example, Figure 9 is a schematic diagram showing the acquisition of redundant feature information including dimensionality reduction processing according to an exemplary embodiment, as Figure 9 shown. The above-mentioned dimensionality reduction processing can be implemented through a 1x1 convolutional layer, Batch normalization, and ReLU function. Of course, the present disclosure does not limit this. Specifically, after the third feature information output by the spatio-temporal memory network as shown in Figure 3 is obtained, it can be input into the module including a 1x1 convolutional layer, Batch normalization, and ReLU function as shown in Figure 9 to obtain the dimensionality-reduced third feature information. Then, the dimensionality-reduced third feature information is input into the module including 3x3 Depthwise convolution, Batch normalization, and ReLU function as shown in Figure 8 to obtain redundant feature information. It should be noted that the channel dimension of the dimensionality-reduced third feature information is not limited to C / 2, and can be any C / n, but it is required that n is a positive integer, C / n is divisible, and C / n > 1.

[0100] Return Figure 5 , in step S505, after fusing the third feature information and the redundant feature information, it is input into the decoding network to obtain the target mask of the current image frame, where the target mask is used for target segmentation of the current image frame. In this step, the third feature information and the redundant feature information can be first concatenated, and then the concatenated result after concatenation is input into the decoding network to obtain the target mask of the current image frame. Then, based on the target mask of the current image frame, target segmentation is performed on the current image frame of the video to be processed. When target segmentation is completed for each image frame in the video to be processed, the video to be processed is completed with target segmentation.

[0101] It should be noted that the above structure for obtaining redundant feature information can also be included in the spatio-temporal memory network. In this way, it is equivalent to improving the original spatio-temporal memory network. That is, the improved spatio-temporal memory network is asFigure 10 As shown, the output of the improved spatio-temporal memory network is the concatenation result of the third feature information and the redundant feature information. At this time, it is the new y. The present disclosure does not limit the direct input of this concatenation result into the decoding network. That is to say, the present disclosure is equivalent to proposing an improved spatio-temporal memory reading module, which can better fuse the information inherent in y, enhance the representation ability of the entire model, and improve the robustness and generalization of the model. Generally, in a complex convolutional neural network, there are many similar channels, that is, there is similar feature information, namely the above-mentioned redundant feature information. The spatio-temporal memory reading module of the present disclosure is based on the following logic: the powerful feature extraction ability of the convolutional neural network is positively correlated with these similar feature information (feature similarity). Therefore, the present disclosure adds a part to amplify the feature similarity to further fuse the information inherent in y, thereby further improving the algorithm performance.

[0102] Moreover, the present disclosure also uses the dilated convolution feature to equivalently replace the downsampling module of the fourth stage module in the basic network architecture of the spatio-temporal memory network, such as the res5 module of ResNet50. Thus, under the condition of maintaining a similar receptive field, the same resolution (1 / 16) can still be maintained. Therefore, before and after the transformation, the network has a similar receptive field area, and the feature resolution used to generate the key-value pair remains unchanged, that is, it is still (1 / 16). At the same time, the rich deep encoding information of the pre-trained basic model (such as ResNet50) can be utilized, the representation ability of the network is enhanced, the ability of the model network to distinguish the target object and the background is improved, and the robustness and generalization of the algorithm are increased.

[0103] Figure 11 is a block diagram of a video object segmentation device shown according to an exemplary embodiment. Refer to Figure 11 , the device includes:

[0104] An image frame acquisition unit 110, configured to acquire a current image frame of a video to be processed, at least one image frame corresponding to the current image frame, and a target mask of the at least one image frame, wherein the position of the at least one image frame in the video to be processed is before the current image frame; An encoding unit 112, configured to perform encoding processing on the at least one image frame, the target mask of the at least one image frame, and the current image frame to obtain first encoding features of the at least one image frame and second encoding features of the current image frame; A third feature information acquisition unit 114, configured to input the first encoding features and the second encoding features into a spatio-temporal memory network to obtain third feature information for predicting the target mask of the current image frame; A redundant feature information acquisition unit 116, configured to extract redundant feature information from the third feature information, wherein the redundant feature information is feature information whose similarity between different channels of the third feature information exceeds a preset value; A target mask acquisition unit 118, configured to input the fused third feature information and redundant feature information into a decoding network to obtain the target mask of the current image frame, wherein the target mask is used for target segmentation of the current image frame.

[0105] According to an embodiment of the present disclosure, the redundant feature information acquisition unit 116 is further configured to input the third feature information into a redundant feature acquisition network to obtain first redundant feature information; perform normalization processing on the first redundant feature information to obtain second redundant feature information; process the second redundant feature information through an activation function to obtain redundant feature information.

[0106] According to an embodiment of the present disclosure, the redundant feature information acquisition unit 116 is further configured to perform dimensionality reduction processing on the third feature information before inputting the third feature information into the redundant feature acquisition network to obtain first redundant feature information.

[0107] According to an embodiment of the present disclosure, the encoding unit 112 is further configured to input the at least one image frame and the target mask of the at least one image frame into a memory encoding network to obtain a first key-value pair as the first encoding feature, wherein the first key-value pair includes first feature information and first key information, the first feature information includes encoding information of the at least one image frame and encoding information of the target mask, and the first key information includes addressing information for querying the first feature information; input the current image frame into a query encoding network to obtain a second key-value pair as the second encoding feature, wherein the second key-value pair includes second feature information and second key information, the second feature information includes encoding information of the current image frame, and the second key information includes addressing information for querying the second feature information.

[0108] According to an embodiment of the present disclosure, the memory encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. The encoding unit 112 is further configured to adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network to obtain an adjusted memory encoding network. The resolution of the output result of the Nth stage module of the adjusted memory encoding network is consistent with the first resolution of the output result of the (N - 1)th stage module. Input at least one image frame and the target mask of at least one image frame into the adjusted memory encoding network to obtain a first key-value pair with the first resolution.

[0109] According to an embodiment of the present disclosure, the query encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. The encoding unit 112 is further configured to adjust the parameters of the dilated convolution in the Nth stage module of the query encoding network to obtain an adjusted query encoding network. The resolution of the output result of the Nth stage module of the adjusted query encoding network is consistent with the first resolution of the output result of the (N - 1)th stage module. Input the current image frame into the adjusted query encoding network to obtain a second key-value pair with the first resolution.

[0110] According to an embodiment of the present disclosure, the encoding unit 112 is configured to adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network when the size parameter of the convolutional kernel in the Nth stage module remains unchanged, so that the resolution of the output result of the adjusted Nth stage module is consistent with the first resolution.

[0111] According to an embodiment of the present disclosure, an electronic device can be provided. Figure 12 It is a block diagram of an electronic device 1200 according to an embodiment of the present disclosure. The electronic device includes at least one memory 1201 and at least one processor 1202. A set of computer-executable instructions is stored in the at least one memory. When the set of computer-executable instructions is executed by the at least one processor, a video target segmentation method according to an embodiment of the present disclosure is executed.

[0112] As an example, the electronic device 1200 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 1000 does not have to be a single electronic device, and can also be an aggregate of any devices or circuits capable of executing the above instructions (or instruction sets) alone or jointly. The electronic device 1200 can also be a part of an integrated control system or a system manager, or can be configured to be interconnected with a local or remote (e.g., via wireless transmission) interface as a portable electronic device.

[0113] In the electronic device 1200, the processor 1202 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 1202 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0114] The processor 1202 may execute instructions or code stored in the memory, where the memory 1201 may also store data. The instructions and data may also be sent and received over a network via a network interface device, where the network interface device may employ any known transmission protocol.

[0115] The memory 1201 may be integrated with the processor 1202, for example, by arranging RAM or flash memory within an integrated circuit microprocessor and the like. Additionally, the memory 1201 may include a separate device, such as an external disk drive, a storage array, or other storage devices usable by any database system. The memory 1201 and the processor 1202 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 1202 can read files stored in the memory 1201.

[0116] Furthermore, the electronic device 1200 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device may be connected to each other via a bus and / or a network.

[0117] According to an embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the video object segmentation method of the embodiment of the present disclosure. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-RLTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0118] According to an embodiment of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the video object segmentation method of the embodiment of the present disclosure.

[0119] Those skilled in the art will readily conceive of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0120] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video object segmentation method, characterized in that, The video object segmentation method includes: Obtaining the current image frame of the video to be processed, at least one image frame corresponding to the current image frame, and the object mask of the at least one image frame, where the position of the at least one image frame in the video to be processed is before the current image frame; Performing encoding processing on the at least one image frame, the object mask of the at least one image frame, and the current image frame to obtain the first encoded feature of the at least one image frame and the second encoded feature of the current image frame, where the first encoded feature includes first feature information and first key information, the first feature information includes the encoded information of the at least one image frame and the encoded information of the object mask, the first key information includes the addressing information for querying the first feature information, the second encoded feature includes second feature information and second key information, the second feature information includes the encoded information of the current image frame, and the second key information includes the addressing information for querying the second feature information; Inputting the first encoded feature and the second encoded feature into a spatio-temporal memory network to obtain third feature information for predicting the object mask of the current image frame; Extracting redundant feature information from the third feature information, where the redundant feature information is the feature information whose similarity degree between different channels of the third feature information exceeds a preset value; Fusing the third feature information and the redundant feature information and inputting the result into a decoding network to obtain the object mask of the current image frame, where the object mask is used for object segmentation of the current image frame.

2. The video object segmentation method according to claim 1, wherein The extracting the redundant feature information from the third feature information includes: Inputting the third feature information into a redundant feature acquisition network to obtain first redundant feature information; Performing normalization processing on the first redundant feature information to obtain second redundant feature information; Processing the second redundant feature information through an activation function to obtain the redundant feature information.

3. The video object segmentation method according to claim 2, wherein Before inputting the third feature information into the redundant feature acquisition network to obtain first redundant feature information, it further includes: Performing dimensionality reduction processing on the third feature information.

4. The video object segmentation method according to any one of claims 1 to 3, characterized in that, The performing encoding processing on the at least one image frame, the object mask of the at least one image frame, and the current image frame to obtain the first encoded feature of the at least one image frame and the second encoded feature of the current image frame includes: Inputting the at least one image frame and the object mask of the at least one image frame into a memory encoding network to obtain a first key-value pair as the first encoded feature; Inputting the current image frame into a query encoding network to obtain a second key-value pair as the second encoded feature.

5. The video object segmentation method according to claim 4, characterized in that, The memory encoding network includes N stage modules, and a convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting the at least one image frame and the object mask of the at least one image frame into the memory encoding network to obtain a first key-value pair, it further includes: Adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network to obtain an adjusted memory encoding network, where the resolution of the output result of the Nth stage module of the adjusted memory encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; The step of inputting the at least one image frame and the target mask of the at least one image frame into the memory encoding network to obtain a first key-value pair includes: inputting the at least one image frame and the target mask of the at least one image frame into the adjusted memory encoding network to obtain the first key-value pair with the first resolution.

6. The video object segmentation method according to claim 4, wherein The query encoding network includes N stage modules, and one convolutional layer in the Nth stage module is a dilated convolution, where N is a positive integer greater than 2. Before inputting the current image frame into the query encoding network to obtain a second key-value pair, it further includes: Adjust the parameters of the dilated convolution in the Nth stage module of the query encoding network to obtain an adjusted query encoding network, where the resolution of the output result of the Nth stage module of the adjusted query encoding network is the same as the first resolution of the output result of the (N - 1)th stage module; The step of inputting the current image frame into the query encoding network to obtain a second key-value pair includes: inputting the current image frame into the adjusted query encoding network to obtain the second key-value pair with the first resolution.

7. The video object segmentation method according to claim 5 or 6, characterized in that The step of adjusting the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network includes: Under the condition that the size parameter of the convolution kernel in the Nth stage module remains unchanged, adjust the parameters of the dilated convolution in the Nth stage module of the memory encoding network and / or the query encoding network so that the resolution of the output result of the adjusted Nth stage module is the same as the first resolution.

8. A video object segmentation device, characterized in that, The video object segmentation device includes: An image frame acquisition unit configured to acquire the current image frame of the video to be processed, at least one image frame corresponding to the current image frame, and the target mask of the at least one image frame, where the position of the at least one image frame in the video to be processed is before the current image frame; An encoding unit configured to perform encoding processing on the at least one image frame, the target mask of the at least one image frame, and the current image frame to obtain a first encoded feature of the at least one image frame and a second encoded feature of the current image frame, where the first encoded feature includes first feature information and first key information, the first feature information includes the encoded information of the at least one image frame and the encoded information of the target mask, the first key information includes addressing information for querying the first feature information, the second encoded feature includes second feature information and second key information, the second feature information includes the encoded information of the current image frame, and the second key information includes addressing information for querying the second feature information; A third feature information acquisition unit configured to input the first encoded feature and the second encoded feature into a spatio-temporal memory network to obtain third feature information for predicting the target mask of the current image frame. A redundant feature information acquisition unit, configured to extract redundant feature information from the third feature information, where the redundant feature information is feature information with a similarity degree between different channels of the third feature information exceeding a preset value; A target mask acquisition unit, configured to input the fused third feature information and redundant feature information into a decoding network to obtain a target mask of the current image frame, where the target mask is used for target segmentation of the current image frame.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the video target segmentation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the video target segmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device

    CN110324617A

  • Video processing method and device

    CN114125462A