Small sample video segmentation model, method and system based on local reference features

By using local proxy encoder, support-query alignment encoder and timing alignment decoder in the small sample video segmentation model, combined with the optimal transmission algorithm and deformable self-attention layer, the semantic alignment of support-query and timing alignment of query videos is achieved, solving the problem of target pose changes and lack of context information in the prior art, and significantly improving the performance of video segmentation.

CN120032133AInactive Publication Date: 2025-05-23DEEP SPACE EXPLORATION LABORATORY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510497360.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing small sample video segmentation methods are difficult to maintain consistency and accuracy when dealing with target pose changes and lack of context information, resulting in erroneous correlation and classification.

Method used

Using a video segmentation model based on local reference features, the local proxy encoder, support-query alignment encoder and timing alignment decoder are combined with the optimal transmission algorithm and deformable self-attention layer to achieve semantic alignment of support-query and timing alignment of query video.

Benefits of technology

It significantly improves the video segmentation performance under the premise of small samples, can more accurately deal with the problem of target pose changes and lack of context information, and improves the consistency and accuracy of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032133A_ABST
    Figure CN120032133A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample video segmentation model, method and system based on local reference features, and belongs to the technical field of computer vision. Comprising a local proxy encoder, a support-query alignment encoder and a time sequence alignment decoder; after a non-target part of the support feature is filtered out, the support feature is input into a local agent encoder, similarity with an initial local agent is calculated, a correlation graph is obtained through projection, the most similar agent is distributed to each pixel in the correlation graph, an optimal distribution matrix is obtained, weighted pooling is applied, and an updated local agent is obtained; the query features are input into a support-query alignment encoder, and the query features and the updated local agent are projected and aggregated to obtain enhanced features; and the time sequence alignment decoder retains query features of a video preorder frame and a corresponding target mask, extracts foreground and background time sequence local agents and background time sequence local agents, activates the query features of the current frame to obtain a time sequence activation graph, and splices the time sequence activation graph with the enhanced features to obtain a prediction target mask.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a small sample video segmentation model, method and system based on local reference features. Background Art

[0002] Video segmentation is a basic visual task, and deep learning-based methods have achieved great success in this field. However, current video segmentation methods rely heavily on a large amount of time-consuming and labor-intensive dense annotation. In order to reduce the need for manual annotation, small sample video segmentation has attracted more and more attention. Its definition is to predict the target mask with unseen categories in an unlabeled video sequence (called the query set) through a limited number of labeled images (called the support set).

[0003] The most effective small sample video object segmentation methods can be roughly divided into two categories: prototype learning-based methods and affinity learning-based methods. On the one hand, prototype learning methods use masked average pooling on support feature maps to generate a compact vector (called prototype) to encode object semantics as a reference for query segmentation. However, objects in query videos often change postures constantly, and object prototypes are difficult to maintain consistency in various situations. On the other hand, affinity learning methods try to directly use the pixel similarity between support features and query features for segmentation. These methods use detailed pixel features as diverse object representations. However, due to the lack of contextual information, direct pixel correlation is often plagued by noisy pixels, background clutter, and intra-class differences, and erroneous correlations may lead to misclassification of query pixels. Overall, the above analysis shows that neither object prototypes nor pixel features are suitable as reference features. Summary of the invention

[0004] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a small sample video segmentation model, method and system based on local reference features, which solves the problems in the prior art.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A video segmentation model based on local reference features, including: a local proxy encoder, a support-query alignment encoder, and a temporal alignment decoder; Support features After filtering out the non-target part, the local proxy encoder is input and first compared with the initial local proxy Calculate the similarity and filter out the support features and the initial local agent Projection is performed to obtain the correlation map A; the most similar agent is assigned to each pixel of the correlation map A to obtain the optimal assignment matrix ; Apply weighted pooling, combined with the allocation matrix and the feedforward network to obtain the updated local agent ; Query Features The deformable self-attention layer in the input support-query alignment encoder is used to fuse the context information, and then the query features are transformed using linear projection. and update local proxy Project and aggregate to obtain enhanced features ; The time-aligned decoder retains the query features of the previous frames of the video And the corresponding target mask , from which the foreground temporal local proxy is extracted and background temporal local agent ; respectively use the local agent to extract the foreground time series and background temporal local agent Activate the query feature of the current frame to obtain the temporal activation map; enhance the feature It is spliced ​​with the temporal activation map and the predicted target mask is finally obtained through the convolutional neural network. .

[0006] The video segmentation method based on local reference features uses the above-mentioned video segmentation model based on local reference features, and includes the following steps: Acquisition of support images and query video frames ; Use the backbone network to extract the support images and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; The query feature and filtered support features Input the local reference feature-based video segmentation model and output the predicted target mask .

[0007] Further, the support features are filtered out The expression of the non-target part is: in, is the support feature after filtering. It is Hadamard. Indicates that the target mask is marked The size of the support feature is adjusted to Consistent.

[0008] Further, the updated local proxy is obtained by the local proxy encoder The steps are: 1) Using the initial local proxy With the support features after filtering Calculate similarity; 2) Support features after filtering Projected with local proxy features, we get the query ,key Sum : in, All are linear projection matrices; 3) Calculate the correlation graph A: in, is a scaling factor; 4) Assign each pixel to the most similar agent to obtain a better allocation matrix : in, is the entropy function, is a parameter that controls the smoothness of the mapping, The allocation matrix In coordinates The value at is the index of the local proxy, is the index of the feature map; Indicates the transmission plan The feasible range of represents an all-one vector of appropriate dimension; represents the number of local agents, h and w represent the height and width of the feature map respectively; 5) Apply weighted pooling, combined with the allocation matrix And a feed-forward network to update the local agent: in, represents a feed-forward network.

[0009] Furthermore, the enhanced features are obtained through the support-query alignment encoder The steps are: 1) Set the query feature Input to the deformable self-attention layer to fuse contextual information from other pixels; 2) Use linear projection to transform query features Projection as query , while the local agent Projection as key Sum ; 3) Aggregate support information from local agents to obtain enhanced features : .

[0010] Furthermore, the time-aligned decoder is used to obtain the predicted target mask The steps are: Step 1, through the previous frame Query features and predicted target mask , to calculate the uncertainty : Step 2, from the previous frame Query features Extract the lowest uncertainty Foreground feature points are used as candidate points to form a candidate set ; Randomly select an initial position , construct the initial position set and proxy collection ; Step 3, calculate the cosine similarity between each candidate point and the selected agent, and Among the candidate points, select the point that is most orthogonal to the proxy in the candidate set. ; Step 4: Select the location Update the location set and agent set; Step 5, repeat steps 3-4 until you select pixel features to obtain the temporal local proxy of the foreground : Step 5: Construct a candidate set by selecting background feature points to obtain a temporal local proxy for the background ; Step 6: Use the temporal local proxy of the foreground Activate the foreground of the current frame: in, Represents the index of height, width, channel and prototype respectively; is the foreground temporal activation map; Step 7: Use a temporal local proxy for the background To suppress the background area in the current frame: in, is the background temporal activation diagram; Step 8: Fusion Enhancement Features Get the predicted target mask : in, represents the concatenation operation of feature channels, Represents the segmentation header, consisting of a The convolution, Activate and a Convolutions are composed in sequence.

[0011] The video segmentation system based on local reference features includes: Image acquisition module: Acquisition of supporting images and query video frames ; Feature extraction and processing module: Use the backbone network to extract features from the support image and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; And, prediction module: query features and filtered support features Input the local reference feature-based video segmentation model and output the predicted target mask .

[0012] A computer storage medium stores a readable program, which can execute the above-mentioned video segmentation method based on local reference features when the program is running.

[0013] An electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned video segmentation method based on local reference features.

[0014] A computer program product includes computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned video segmentation method based on local reference features.

[0015] Beneficial effects of the present invention: The present invention makes full use of the advantages of local features as target reference features, and realizes small sample video segmentation by jointly exploring the semantic alignment of support-query and the temporal alignment of query video; the designed local proxy encoder introduces the optimal transmission algorithm to adaptively divide the target foreground area into semantically consistent local areas; the proposed support-query alignment decoder can pass the support information contained in the local proxy to the query; finally, the temporal alignment encoder effectively utilizes the temporal consistency and learns the temporal local proxy to assist the segmentation of the current frame; this method can significantly improve the video segmentation performance under the premise of small samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a framework diagram of a small sample video segmentation model based on local reference features of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] Example 1 like Figure 1 As shown, a small sample video segmentation model based on local reference features includes a local proxy encoder, a support-query alignment encoder and a temporal alignment decoder; Support features After filtering out the non-target part, the local proxy encoder is input and first compared with the initial local proxy Calculate the similarity and filter out the support features and the initial local agent Projection is performed to obtain the correlation map A; the most similar agent is assigned to each pixel of the correlation map A to obtain the optimal assignment matrix ; Apply weighted pooling, combined with the allocation matrix and the feedforward network to obtain the updated local agent ; Query Features The deformable self-attention layer in the input support-query alignment encoder is used to fuse the context information, and then the query features are transformed using linear projection. and update local proxy Project and aggregate to obtain enhanced features ; The time-aligned decoder retains the query features of the previous frames of the video And the corresponding target mask , from which the foreground temporal local proxy is extracted and background temporal local agent ; respectively use the local agent to extract the foreground time series and background temporal local agent Activate the query feature of the current frame to obtain the temporal activation map; enhance the feature It is spliced ​​with the temporal activation map and the predicted target mask is finally obtained through the convolutional neural network. .

[0020] Example 2 In this embodiment, a method for performing video segmentation using the small sample video segmentation model based on local reference features is proposed, comprising the following steps: S1, acquisition of supporting images and query video frames ; The query video is a video provided by the user or to be processed in batches. The supporting image is a video with the same category as the query video selected from the existing video library, and one of its frames is randomly selected.

[0021] S2, using the backbone network to respectively and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; The backbone network is a pre-trained residual connection network ResNet-50, which is composed of a series of convolutional neural networks. The support images and query video frames are input into the backbone network in sequence, and the output results of the convolutional neural network in the penultimate layer are used as support features and query features.

[0022] Filter out support features The expression of the non-target part is: in, is the support feature after filtering. is the Hadamard product (element-wise multiplication), Indicates that the target mask is marked The size of the support feature is adjusted to Consistent.

[0023] S3, query features and filtered support features Input such as Figure 1 In the small sample video segmentation model based on local reference features shown in Figure 1, the predicted target mask is output .

[0024] and all , where h, w, and c represent the height, width, and channel dimensions of the feature map, respectively; The specific contents of S3 include: S31, the filtered support features Input the local proxy encoder and get the updated local proxy ; 1) Initialize a set of local agents , and use them together with the filtered support features Calculate the similarity, where represents the number of local agents; 2) Project the support features and local proxy features to obtain the query ,key Sum , the formula is as follows: in, is the linear projection matrix; 3) The correlation graph can be obtained by the following formula: in, is a scaling factor; 4) The pixel-agent assignment problem is expressed as a discrete form of the optimal transmission problem, that is, each pixel is assigned to the most similar agent as much as possible to obtain a better assignment matrix , represents the number of local agents, h and w represent the height and width of the feature map respectively. The specific formula is as follows: in, is the entropy function, is a parameter that controls the smoothness of the mapping, The allocation matrix In coordinates The value at is the index of the local proxy, is the index of the feature map; Indicates the transmission plan The feasible range is defined as follows: in, represents an all-one vector of appropriate dimension; 5) Different agents can be responsible for different complementary local areas; apply weighted pooling, combined with the allocation matrix And a feed-forward network to update the local agent: in, represents a feed-forward network.

[0025] S32, query features and update local proxy Input support-query alignment encoder to obtain enhanced features ; The support-query alignment decoder aims to pass support information to the query and uses local agents as a bridge to align features on both sides; specifically: 1) Set the query feature Input to the deformable self-attention layer to fuse contextual information from other pixels; 2) Use linear projection to transform query features Projection as query , while the local agent Projection as key Sum ; 3) Aggregate the support information from the local agent through the following formula to obtain the enhanced feature : .

[0026] S33, will enhance the characteristics Input the time-aligned decoder to get the predicted target mask ; The role of the time alignment decoder is to effectively utilize the temporal consistency between query video frames to mine additional guidance information. To this end, a cache mechanism involving the first few frames is designed in this embodiment; the specific process is as follows: 1) In time frame , keep the previous frame Query features And predict the target mask ; Use the previous frame The predicted target mask Entropy is used to measure uncertainty , the calculation formula is: 2) From the previous frame Query features Extract the foreground feature points (candidate points) with the lowest uncertainty, the number is , to form a candidate set ; By randomly selecting an initial position , construct the initial position set and proxy collection ; and calculate each candidate point With selected agent The cosine similarity between is as follows: exist Among the candidate points, select the position of the point that is most orthogonal to the agent in the candidate set. The formula is as follows: 3) In this way, each selected pixel feature is far away from each other in the candidate set; then according to the selected position Update location collection and proxy collection , the format is as follows: 4) Repeat the above steps until you select pixel features; in this way, a temporal local proxy for the foreground can be obtained: 5) Using the same steps, we can construct a candidate set by selecting background feature points to obtain a temporal local proxy for the background ; 6) Then combine these proxies with the query features Interact to activate the target-related area in the current frame; considering that the foreground and background have different temporal properties, different alignment schemes are adopted; specifically: On the one hand, using the temporal local proxy of the foreground Activate the foreground of the current frame. The formula is as follows: in, Represents the index of height, width, channel and prototype respectively; is the foreground temporal activation map; On the other hand, using a temporally local proxy for the background To suppress the background area in the current frame, the formula is as follows: in, is the background temporal activation diagram; Finally, the foreground temporal activation map and background temporal activation map With enhanced features Fusion to get the predicted target mask : in, represents the concatenation operation of feature channels, Represents the segmentation header, consisting of a The convolution, Activate and a Convolutions are composed in sequence.

[0027] Based on similar inventive concepts, an embodiment of the present invention further provides a computer storage medium storing a readable program, which can execute the above-mentioned small sample video segmentation method based on local reference features when the program is run.

[0028] Based on similar inventive concepts, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned small sample video segmentation method based on local reference features.

[0029] Based on similar inventive concepts, an embodiment of the present invention further provides a computer program product, including computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the above-mentioned small sample video segmentation method based on local reference features.

[0030] Example 3 In this embodiment, a small sample video segmentation system based on local reference features is proposed, including: Image acquisition module: Acquisition of supporting images and query video frames ; Feature extraction and processing module: Use the backbone network to extract features from the support image and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; And, prediction module: query features and filtered support features Input into the small sample video segmentation model based on local reference features, and output the predicted target mask .

[0031] The video segmentation method of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0032] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. Video segmentation model based on local reference features, characterized by: include: Local proxy encoder, support-query alignment encoder, and temporal alignment decoder; Support features After filtering out the non-target part, the local proxy encoder is input and first compared with the initial local proxy Calculate the similarity and filter out the support features and the initial local agent Projection is performed to obtain the correlation map A; the most similar agent is assigned to each pixel of the correlation map A to obtain the optimal assignment matrix ; Apply weighted pooling, combined with the allocation matrix and the feedforward network to obtain the updated local agent ; Query Features The deformable self-attention layer in the input support-query alignment encoder is used to fuse the context information, and then the query features are transformed using linear projection. and update local proxy Project and aggregate to obtain enhanced features ; The time-aligned decoder retains the query features of the previous frames of the video And the corresponding target mask , from which the foreground temporal local proxy is extracted and background temporal local agent ; respectively use the local agent to extract the foreground time series and background temporal local agent Activate the query feature of the current frame to obtain the temporal activation map; enhance the feature It is spliced ​​with the temporal activation map and the predicted target mask is finally obtained through the convolutional neural network. .

2. A video segmentation method based on local reference features, using the video segmentation model based on local reference features according to claim 1, characterized in that: The following steps are involved: Acquisition of support images and query video frames ; Use the backbone network to extract the support images and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; The query feature and filtered support features Input the local reference feature-based video segmentation model and output the predicted target mask .

3. The video segmentation method based on local reference features according to claim 2, characterized in that: Filter out the support features The expression of the non-target part is: in, is the support feature after filtering. It is Hadamard. Indicates that the target mask is marked The size of the support feature is adjusted to Consistent.

4. The video segmentation method based on local reference features according to claim 2, characterized in that: The updated local proxy is obtained by the local proxy encoder The steps are: 1) Using the initial local proxy With the support features after filtering Calculate similarity; 2) Support features after filtering Projected with local proxy features, we get the query ,key Sum : in, All are linear projection matrices; 3) Calculate the correlation graph A: in, is a scaling factor; 4) Assign each pixel to the most similar agent to obtain a better allocation matrix : in, is the entropy function, is a parameter that controls the smoothness of the mapping, The allocation matrix In coordinates The value at is the index of the local proxy, is the index of the feature map; Indicates the transmission plan The feasible range of represents an all-one vector of appropriate dimension; represents the number of local agents, h and w represent the height and width of the feature map respectively; 5) Apply weighted pooling, combined with the allocation matrix And a feed-forward network to update the local agent: in, represents a feed-forward network.

5. The video segmentation method based on local reference features according to claim 4, characterized in that: Enhanced features obtained through support-query alignment encoder The steps are: 1) Set the query feature Input to the deformable self-attention layer to fuse contextual information from other pixels; 2) Use linear projection to transform query features Projection as query , while the local agent Projection as key Sum ; 3) Aggregate support information from local agents to obtain enhanced features : 。 6. The video segmentation method based on local reference features according to claim 5, characterized in that: Using the time-aligned decoder to obtain the predicted target mask The steps are: Step 1, through the previous frame Query features and predicted target mask , to calculate the uncertainty : Step 2, from the previous frame Query features Extract the lowest uncertainty Foreground feature points are used as candidate points to form a candidate set ; Randomly select an initial position , construct the initial position set and proxy collection ; Step 3, calculate the cosine similarity between each candidate point and the selected agent, and Among the candidate points, select the point that is most orthogonal to the proxy in the candidate set. ; Step 4: Select the location Update the location set and agent set; Step 5, repeat steps 3-4 until you select pixel features to obtain the temporal local proxy of the foreground : Step 5: Construct a candidate set by selecting background feature points to obtain a temporal local proxy for the background ; Step 6: Use the temporal local proxy of the foreground Activate the foreground of the current frame: in, Represents the index of height, width, channel and prototype respectively; is the foreground temporal activation map; Step 7: Use a temporal local proxy for the background To suppress the background area in the current frame: in, is the background temporal activation diagram; Step 8: Fusion Enhancement Features Get the predicted target mask : in, represents the concatenation operation of feature channels, Represents the segmentation header, consisting of a The convolution, Activate and a Convolutions are composed in sequence.

7. A video segmentation system based on local reference features, characterized in that: include: Image acquisition module: Acquisition of supporting images and query video frames ; Feature extraction and processing module: Use the backbone network to extract features from the support image and query video frames Extracting support features and query features , and annotated using the target mask Support features Filter out the non-target parts; And, prediction module: query features and filtered support features Input the local reference feature-based video segmentation model described in claim 1, and output the predicted target mask .

8. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the video segmentation method based on local reference features described in any one of claims 2 to 6.

9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the video segmentation method based on local reference features as described in any one of claims 2 to 6.

10. A computer program product comprising computer instructions, characterized in that The computer instructions instruct the computing device to perform operations corresponding to the video segmentation method based on local reference features as described in any one of claims 2-6.

Citation Information

Patent Citations

  • Small sample three-dimensional medical image segmentation system

    CN118470322A