Video instance segmentation method, system, medium and device based on hypergraph representation
By using hypergraph representation and weighted slice Wasserstein distance in video instance segmentation technology, the problem of instance segmentation and tracking in progressive occlusion environment is solved, and a more efficient and accurate video instance segmentation effect is achieved.
Patent Information
- Application Number
- CN202510376491.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing video instance segmentation technology is difficult to accurately segment and track instances in progressive occlusion environments, and the application of hypergraph technology in this field has not yet given a practical solution.
Using the video instance segmentation method based on hypergraph representation, the spatial and temporal consistency between frames is established by constructing hypergraphs, using cell cavity layer-overweighted hypergraph convolution and weighted slice Wasserstein distance, and the accurate tracking and segmentation of instances are achieved.
It significantly improves the video instance segmentation performance of the model in dynamic occlusion scenarios, and can more accurately capture complex relationships and changes between instances, ensuring the continuity and accuracy of instance tracking.
Smart Images

Figure CN119904783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video instance segmentation, and in particular to a method, system, medium and device for video instance segmentation based on hypergraph representation. Background Art
[0002] In the field of video instance segmentation, there are currently many challenges, and the problem of progressive occlusion is particularly prominent. Progressive occlusion can be caused by instance occlusion and camera occlusion, which makes part of the instance invisible, greatly increasing the difficulty of the segmentation task. In this case, maintaining spatiotemporal consistency between frames is crucial for accurate instance tracking. Existing studies, such as decoupling strategies such as Mask2Former-VIS, DVIS, and DVIS-DAQ, have constructed a target motion modeling framework based on explicit queries, but they rely too much on visible features and semantic similarity between instances. In a progressive occlusion environment, it is difficult to distinguish mutually occluded instances of the same type by relying solely on visible feature information within the field of view, which leads to recognition ambiguity and seriously affects segmentation accuracy.
[0003] Most existing video instance segmentation methods are based on video images, and few involve hypergraph-based technical solutions. Hypergraphs can express complex high-order relationships and have unique advantages in analyzing complex interactions between video instances. They can effectively model the rich structural properties between instances in the video and provide a more comprehensive perspective for solving occlusion problems.
[0004] However, conducting research on video instance segmentation on hypergraphs also faces many challenges. For example, the complexity of the hypergraph structure makes feature extraction and relationship modeling more difficult, and how to efficiently capture and utilize the complex relationships between instances in the hypergraph becomes a problem. At the same time, due to the characteristics of the hypergraph, there are also technical barriers in establishing spatiotemporal consistency between frames. At present, the existing technology has not yet provided a practical solution to the problem of how to effectively solve the problem of video instance segmentation on hypergraphs. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes a video instance segmentation method, system, medium and device based on hypergraph representation, which captures the rich structural attributes inherent in the video and extracts the interactive relationship between instances to enhance the spatiotemporal correspondence between local and global instance features between frames, establishes reliable spatiotemporal consistency between frames, and realizes accurate instance tracking and segmentation in complex dynamic scenes.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a video instance segmentation method based on hypergraph representation, comprising:
[0008] Constructing a hypergraph based on instance queries of video frames;
[0009] Allocating vector space for nodes and hyperedges of the hypergraph based on cell stacking, and defining feature associations between nodes and hyperedges through restriction mapping; aggregating the restriction mapping using a Laplacian operator to obtain enhanced structural features;
[0010] Generate projection directions from the vector space, and map the enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances, and weightedly obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value;
[0011] The WSW values of all cross-frame instance pairs in adjacent frames are calculated, and the pair with the smallest WSW value is selected as the best match; based on the best match, an inter-frame instance correspondence is established to achieve segmentation and continuous tracking of target instances in the video.
[0012] In a second aspect, the present invention provides a video instance segmentation system based on hypergraph representation, comprising:
[0013] The hypergraph construction module is configured to construct a hypergraph based on the instance query of the video frame;
[0014] The feature extraction module is configured to allocate vector space to nodes and hyperedges of the hypergraph based on cell stacking, and define feature associations between nodes and hyperedges through restriction mapping; and aggregate the restriction mapping using a Laplace operator to obtain enhanced structural features;
[0015] The distance calculation module is configured to generate a projection direction from the vector space, map the enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances, and weightedly obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as a WSW value;
[0016] The instance segmentation module is configured to calculate the WSW values of all cross-frame instance pairs in adjacent frames, select the pair with the smallest WSW value as the best match; establish the inter-frame instance correspondence based on the best match to achieve segmentation and continuous tracking of the target instance in the video.
[0017] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a video instance segmentation method based on hypergraph representation described in the first aspect.
[0018] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for video instance segmentation based on hypergraph representation described in the first aspect are implemented.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] (1) To solve the problem of progressive occlusion, the present invention proposes weighted temporal consistency for occlusion, adopts an importance weighting strategy, and emphasizes the contribution of key structural information in feature representation. This strategy consists of two key parts: on the one hand, weighted hypergraph convolution is used to extract enhanced structural features with the help of hyperedges. This method can effectively highlight the important interaction information between instances, so that the model has a clearer understanding of the complex relationship between instances. On the other hand, the weighted slice Wasserstein distance is used to measure the spatiotemporal consistency between adjacent frames. In this way, the changes in instances between different frames can be captured more accurately. This dual weighting mechanism significantly enhances the model's ability to cope with complex occlusions, thereby effectively improving the model's video instance segmentation performance in dynamic occlusion scenarios.
[0021] (2) When performing complex dynamic modeling, the present invention uses a complex dynamic modeling method based on hypergraph convolution and introduces weighted hypergraph convolution based on cell stacking. The core purpose of this operation is to capture local high-order subtle structure information. Cell stacking provides richer hierarchical structure information for the nodes and hyperedges of the hypergraph. With this rich information, the model can more accurately model the complex relationship between instances. Furthermore, the subtle structure information obtained by cell stacking is integrated into the hypergraph Laplacian operator. Through this integration, during the convolution process, the model can effectively capture hidden high-order structures, thereby achieving efficient feature propagation and aggregation. This series of operations allows the model to perform better when processing complex dynamic scenes and can better cope with various situations in complex dynamic environments.
[0022] (3) To achieve spatiotemporal consistency between frames, the present invention maintains spatiotemporal consistency through dynamic reasoning and uses a dynamic reasoning mechanism based on weighted slice Wasserstein distance to compare the structural features of adjacent frames. This not only maintains the structural characteristics of the hypergraph, but also accurately captures the correlation differences between instances. By maintaining these structural invariants, the correspondence between instances between frames can be accurately established even in complex situations with occlusion, thereby ensuring reliable instance tracking and guaranteeing the accuracy and consistency of video instance segmentation tasks between different frames.
[0023] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0025] Figure 1 A main flow chart of a video instance segmentation method based on hypergraph representation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0027] Embodiment 1
[0028] like Figure 1 As shown, this embodiment discloses a video instance segmentation method based on hypergraph representation, comprising the following steps:
[0029] S1: Constructing a hypergraph based on instance query of video frames;
[0030] S2: allocating vector space for nodes and hyperedges of the hypergraph based on cell stacking, and defining feature associations between nodes and hyperedges through restriction mapping; aggregating the restriction mapping using a Laplacian operator to obtain enhanced structural features;
[0031] S3: Generate projection directions from the vector space, and map the enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances, and weightedly obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value;
[0032] S4: Calculate the WSW values of all cross-frame instance pairs in adjacent frames, and select the pair with the smallest WSW value as the best match; establish an inter-frame instance correspondence relationship based on the best match to achieve segmentation and continuous tracking of the target instance in the video.
[0033] In order to solve the recognition ambiguity problem of existing decoupling strategies when dealing with progressive occlusion, this embodiment proposes a weighted dynamic structure inference (WDSI) framework, which constructs a hypergraph based on instance query, extracts the high-order structural information of the hypergraph based on the weighted hypergraph convolution (WSHC) of cellular sheaves, and then uses the dynamic reasoning mechanism based on weighted sliced Wasserstein distance (WSW) to compare the high-order structural features between adjacent frames, thereby establishing instance association and realizing instance tracking based on the association, thereby alleviating the problem of progressive occlusion.
[0034] Next, combine Figure 1 , a video instance segmentation method based on hypergraph representation disclosed in this embodiment is described in detail.
[0035] In S1, obtain the video frame and extract the instance-level query ,in represents the time step. Each query Indicates a frame queries, and The embedding vector generated by Mask2Former output contains appearance and spatial features. This embodiment will use these queries to build a hypergraph, perform hypergraph convolution to enhance structural features, and impose spatiotemporal consistency constraints.
[0036] First, the hypergraph is defined as ,in represents a node set, represents a hyperedge set, is the number of nodes, is the number of hyperedges, Defines the weight of a hyperedge. Any number of nodes can be connected. The number of nodes contained in each hyperedge writing , is called the degree of the hyperedge.
[0037] To query based Constructing a Hypergraph In this embodiment, the instance query is used as a hypergraph node and the node set is defined , where each node Corresponding to the query .
[0038] Hyperedges are used to model spatial dependencies by connecting multiple nodes to capture complex spatial relationships beyond simple pairwise connections.
[0039] Specifically, by calculating the eigenvector and The pairwise distance between them is used to construct the hyperedge. For each pair of instances and ,node and The distance between is defined as:
[0040]
[0041] in, represents the Euclidean norm. For each node , this embodiment is based on the minimum distance selection in the feature space Neighbors, forming a neighborhood set , used to capture local spatial context.
[0042] Then, define the hyperedge Contains nodes and its neighbors :
[0043]
[0044] The strength of the relationship within each hyperedge is determined by the weight It is calculated as follows:
[0045]
[0046] in, is the number of nodes in the neighborhood set. This weight effectively encodes spatial information and preserves the key arrangement properties of instances.
[0047] In S2, based on the constructed hypergraph, cell stacking is introduced to capture the high-order interaction information between instances.
[0048] Specifically, honeycomb stacking It is defined in the hypergraph The structure on which each node and each hyperedge are allocated vector spaces, called stalks, to represent their respective Dimensional features:
[0049]
[0050] Wherein, R represents the set of real numbers.
[0051] At the same time, for each hyperedge Node , define a linear mapping between a node and a hyperedge, called a restriction map, which is used to connect the feature space of the node and the hyperedge:
[0052]
[0053] in, Represents a multi-layer perceptron.
[0054] This layered structure constrains local consistency and models high-order interactions in the hypergraph, providing a solid mathematical framework for understanding and analyzing complex video data.
[0055] Furthermore, to enhance the representation of high-order interactive features, the SheafLaplacian is used for hypergraph feature propagation. The SheafLaplacian is used to aggregate information from neighboring nodes while maintaining high-order relationships and imposing local consistency constraints.
[0056] When transferring and updating node features, this embodiment fully considers the strength of the hyperedge, and when updating node features, not only the size of the hyperedge is considered, but also the difference between nodes is weighed by weight. This makes the stacked hypergraph neural network (SHNNs) have a good effect in modeling complex data interactions.
[0057] Specifically, for the hypergraph , whose linear stacked hypergraph Laplace ,go through After normalization, the definition of (hyperedge degree) is as follows:
[0058]
[0059] in, , Represents any node.
[0060] The linear stacked Laplacian operator is obtained by Acting on Features When , it is calculated as follows:
[0061]
[0062] in, , Respectively represent nodes , Features; , Respectively represent nodes With super edge Restriction mapping between express The transpose of .
[0063] After obtaining the aggregated restricted map, a weighted stacked hypergraph convolution (WSHC) layer is used to perform the following operations:
[0064]
[0065] in, is the activation function, and The sizes are and The identity matrix of is the number of nodes, is the stacking dimension. is the corresponding block diagonal matrix, each block Represents the degree matrix of each node. Represents the current layer index. Represents the normalized restriction map. , Respectively represent Layer, +1 layer of structural features.
[0066] No. The output of the first layer is used as the input of the next layer. After layer-by-layer weighted stacked hypergraph convolution processing and enhancement, a structural feature is constructed .
[0067] The stacked hypergraph neural network consists of multiple layers of weighted stacked hypergraph convolutional layers, which propagate and optimize node features by combining high-order relationships captured by the hypergraph. Stacked Laplacian Operator Through each hyperedge Defines constraints that enforce local consistency of node features. The expression The stacked Laplacian operators are fused when performing feature diffusion while keeping the constraints of the stacked structure.
[0068] Learnable weight matrix and Dynamically adjust feature transformation and aggregation. Uniformly scale the feature vector, while the ReLU activation function Introducing nonlinearity to model complex patterns.
[0069] In S3, the one-dimensional Wasserstein distance is equivalent to the one-dimensional probability measure and ( ), -Wasserstein distance is defined as follows:
[0070]
[0071] in, and They are and The cumulative distribution functions (CDFs) of . This formula provides a closed form solution for calculating the Wasserstein distance in one dimension, making it suitable for projective measures.
[0072] In order to generalize the Wasserstein distance to high-dimensional measurements, the sliced Wasserstein distance (SWD) measures and ) is projected into a one-dimensional subspace and the one-dimensional Wasserstein distances of these projections are averaged. SWD is defined as follows:
[0073]
[0074] in, and Respectively represent the direction (Right now The projection function is the measure of the unit sphere in the Will The points on are mapped to , so that the Wasserstein distance can be calculated in one-dimensional space.
[0075] Since the expected calculation cost in equation (11) is large, the Monte Carlo method is usually used for approximation, that is, Independent sampling Directions , and calculate the average:
[0076]
[0077] Among them, each and The measurement and Along direction Projection representation of . Number of projections Controls the accuracy of the Monte Carlo approximation.
[0078] Furthermore, after optimizing the relationship based on the high-order hypergraph, it is necessary to and Establish a corresponding relationship between instances.
[0079] Since objects may experience occlusion, motion, and appearance changes, accurate temporal instance tracking is essential. To this end, this embodiment improves the traditional slice Wasserstein distance and adopts the weighted slice Wasserstein (WSW) metric to determine the instance correspondence between consecutive frames.
[0080] For frame Query examples in and frame Examples in , define a random path:
[0081]
[0082] This path captures the directional difference between the two augmented queries, where and is the output of weighted stacked hypergraph convolution (WSHC).
[0083] The path is then normalized to obtain unit direction: ,if If ∈R is close to zero (i.e., the characteristics of the two instances are almost identical), a preset small constant is added to maintain stability.
[0084] Finally, a set of projection direction distributions is constructed that can highlight the spatial and appearance variations between different frames.
[0085] Specifically, the Wasserstein distance is calculated based on the projection: for each projection direction , this embodiment projects along this direction and , get a one-dimensional representation, and calculate the Wasserstein distance: , which measures the distance along The degree of alignment of instances in a direction. A larger Wasserstein distance indicates that the instances have greater dissimilarity in that direction, thereby revealing more significant changes. In order to strengthen the directions with higher information content, this embodiment weights the Wasserstein distance, highlights projections with greater differences with higher weights, and enhances the distinguishing ability of the WSW metric.
[0086] Finally, the frame and frame The WSW distance between them is expressed as:
[0087]
[0088] in, Represents the weight of each projection direction, which is used to adjust the contribution of the projection distance to the overall WSW distance. and They are and Along the unit direction The projection representation of . Represents the one-dimensional Wasserstein distance.
[0089] Furthermore, the weight Usually based on the Wasserstein distance Calculate, this distance measures the frame and Search and The specific calculation is as follows:
[0090]
[0091] in, is a monotonic function that gives higher weights to larger differences to highlight instances that change more significantly over time. This normalization ensures that the sum of all weights is 1, ensuring that all projection directions contribute evenly and preventing a certain direction from having an excessive impact on the final WSW result.
[0092] In S4, in order to establish the corresponding relationship between the instances, this embodiment calculates the WSW values of all instance pairs between adjacent frames: It measures the temporal consistency of the instances in the projection direction. A smaller WSW value indicates a stronger temporal consistency. Therefore, this embodiment selects the instance pair with the smallest WSW value as the best match:
[0093]
[0094] The selected instance pairs represent the instance matches with the strongest temporal consistency, thus ensuring robust and accurate object tracking.
[0095] Furthermore, in formula (14), when and is a discrete measure (with at most The computational complexity is as follows:
[0096] Random path sampling (sampling projection directions): (time & memory);
[0097] Projection sampling (sampling from von Mises-Fisher (vMF) or Power Spherical (PS) distribution): ;
[0098] One-dimensional Wasserstein distance calculation: Calculate all projections Need to increase time complexity .
[0099] Therefore, the overall time complexity is , the space complexity is It can be seen that these complexities are controllable, ensuring that the method has good scalability on large-scale datasets.
[0100] Further, the objective function, finally, the enhanced instance query representation As input, it is used for the class prediction head (class head) and the mask prediction head (mask head) to generate the class prediction result and mask coefficient output The total loss function of weighted dynamic structure inference (WDSI) is defined as follows:
[0101]
[0102] in, To predict the label, represents the mask loss associated with class and mask prediction.
[0103] This specific embodiment adopts weighted hypergraph convolution based on cell cavity stacking, which can capture local high-order subtle structural information, accurately model complex relationships between instances, achieve efficient feature propagation and aggregation, and enhance the model's ability to understand complex scenes. Secondly, by using a dynamic reasoning mechanism based on weighted slice Wasserstein distance, it not only maintains the structural characteristics of the hypergraph, but also accurately captures the differences in instance associations, accurately establishes inter-frame instance correspondences in the case of occlusion, and ensures reliable instance tracking. Furthermore, the dual weighting strategy highlights key structural information and enhances the model's ability to cope with complex occlusions. Overall, these innovations enable this solution to achieve outstanding results in video instance segmentation and video panorama segmentation tasks, and has good scalability, which can effectively meet the needs of practical applications such as autonomous driving, video editing and monitoring.
[0104] Embodiment 2
[0105] This embodiment provides a video instance segmentation system based on hypergraph representation, including:
[0106] The hypergraph construction module is configured to construct a hypergraph based on the instance query of the video frame;
[0107] The feature extraction module is configured to allocate vector space to nodes and hyperedges of the hypergraph based on cell stacking, and define feature associations between nodes and hyperedges through restriction mapping; and aggregate the restriction mapping using a Laplace operator to obtain enhanced structural features;
[0108] The distance calculation module is configured to generate a projection direction from the vector space, map the enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances, and weightedly obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as a WSW value;
[0109] The instance segmentation module is configured to calculate the WSW values of all cross-frame instance pairs in adjacent frames, select the pair with the smallest WSW value as the best match; establish the inter-frame instance correspondence based on the best match to achieve segmentation and continuous tracking of the target instance in the video.
[0110] Embodiment 3
[0111] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the method for video instance segmentation based on hypergraph representation as described in the first embodiment above are implemented.
[0112] Embodiment 4
[0113] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for video instance segmentation based on hypergraph representation as described in the first embodiment above are implemented.
[0114] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A video instance segmentation method based on hypergraph representation, characterized in that: include: Constructing a hypergraph based on instance queries of video frames; Allocating vector space to nodes and hyperedges of the hypergraph based on cell stacking, and defining feature associations between nodes and hyperedges through restricted mapping; A Laplacian operator is used to aggregate the restriction mapping to obtain enhanced structural features; generating projection directions from the vector space, and mapping enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; Calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances to obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value; The WSW values of all cross-frame instance pairs in adjacent frames are calculated, and the pair with the smallest WSW value is selected as the best match; based on the best match, an inter-frame instance correspondence is established to achieve segmentation and continuous tracking of target instances in the video.
2. The method for video instance segmentation based on hypergraph representation according to claim 1, characterized in that: The hypergraph is constructed based on the example query of the video frame, specifically: Treat instance queries as hypergraph nodes, where each node corresponds to the embedded features of an instance; Calculate the pairwise Euclidean distance between nodes, select the k nearest neighbors of each node based on the minimum distance, and form a hyperedge; The hyperedge weight is determined by the reciprocal mean of the distances between neighboring nodes.
3. The method for video instance segmentation based on hypergraph representation according to claim 1, characterized in that: The method of allocating vector space to the nodes and hyperedges of the hypergraph based on cell stacking, and defining feature associations between nodes and hyperedges by restricting mapping, specifically includes: Allocating vector space to nodes and hyperedges of the hypergraph for representing dimensional features; For each node belonging to a hyperedge, a linear mapping between the node and the hyperedge is defined by a multi-layer perceptron, called a restricted mapping, which is used to associate node and hyperedge features.
4. The method for video instance segmentation based on hypergraph representation according to claim 1, characterized in that: The Laplacian operator is used to aggregate the restriction mapping to obtain enhanced structural features, which specifically includes: The restriction map is normalized using the Laplace operator: ; in, , represents any node, represents a hyperedge; represents the hyperedge degree; , Respectively represent nodes With super edge Restriction mapping between , Respectively represent nodes , Features; express The transpose of The normalized restricted map is aggregated using weighted hypergraph convolution: ; ; in, is the activation function, represents the stacked Laplacian operator, and The sizes are and The identity matrix of is the number of nodes, is the stacking dimension; is the corresponding block diagonal matrix, each block represents the degree matrix of each node; Represents the current layer index; , Respectively represent Layer, +1 layer structural features; , represents the learnable weight matrix; represents the normalized restriction map; The first The output of the layer is used as the input of the next layer. After layer-weighted stacked hypergraph convolution processing and enhancement, enhanced structural features are obtained.
5. The method for video instance segmentation based on hypergraph representation according to claim 1, characterized in that: generating projection directions from the vector space, and mapping enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; Calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances to obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value, which specifically includes: Define the difference vector between adjacent frame instances and normalize it to get the unit projection direction; Along each unit projection direction, the structural features of adjacent frames are mapped to one-dimensional space, and the one-dimensional Wasserstein distance is calculated; Traverse the one-dimensional Wasserstein distances in all directions, configure the corresponding weights based on the one-dimensional Wasserstein distances, fuse all one-dimensional Wasserstein distances and weights, and obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value.
6. The method for video instance segmentation based on hypergraph representation according to claim 5, characterized in that: The WSW value is specifically: ; in, Represents one-dimensional Wasserstein distance; Represents the weight of each projection direction, which is used to adjust the contribution of the projection distance to the overall WSW distance; is the enhanced structural feature of adjacent frames t and t+1, each and They are and Along the unit direction The projection representation of .
7. The method for video instance segmentation based on hypergraph representation according to claim 1, characterized in that: The training process of establishing the inter-frame instance correspondence based on the best match to achieve segmentation and continuous tracking of the target instance in the video includes: Select the cross-frame instance pair with the smallest WSW value in adjacent frames as the best match; Input the enhanced instance query into the category prediction head and the mask prediction head to generate category prediction results and mask coefficients; The prediction process is optimized through the total loss function, and the segmentation and continuous tracking of the target instance are achieved based on the optimized category prediction results and mask coefficients.
8. A video instance segmentation system based on hypergraph representation, characterized in that: include: The hypergraph construction module is configured to construct a hypergraph based on the instance query of the video frame; A feature extraction module is configured to allocate vector space to nodes and hyperedges of the hypergraph based on cell stacking, and define feature associations between nodes and hyperedges through restriction mapping; A Laplacian operator is used to aggregate the restriction mapping to obtain enhanced structural features; A distance calculation module is configured to generate a projection direction from the vector space, and map the enhanced structural features of adjacent frames to a one-dimensional space along each projection direction; Calculate the Wasserstein distance of adjacent frames on the one-dimensional projection, and dynamically assign weights according to the distances to obtain the weighted slice Wasserstein distance of the hypergraph structural features of adjacent frames as the WSW value; The instance segmentation module is configured to calculate the WSW values of all cross-frame instance pairs in adjacent frames, select the pair with the smallest WSW value as the best match; establish the inter-frame instance correspondence based on the best match to achieve segmentation and continuous tracking of the target instance in the video.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a video instance segmentation method based on hypergraph representation as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the video instance segmentation method based on hypergraph representation as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
DCNN-based experimental operation key node scoring method and system
CN117726977A
Priori distance guided similarity memory matching video instance segmentation method
CN118429863A