Video sequence prediction method and system based on object segmentation guidance

The video prediction method guided by object segmentation explicitly extracts and encodes the structural representation information of objects, and uses a conditional diffusion model to generate future video frames. This solves the problem of insufficient consistency and realism of objects in long-term prediction in existing technologies, and achieves higher quality video prediction.

CN120932161BActive Publication Date: 2025-12-05JILIN HUAQIAO FOREIGN LANGUAGES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511464161.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-12-05
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing video prediction technologies lack explicit modeling of independent semantic objects in complex scenes, resulting in insufficient consistency and physical realism of objects in long-term predictions, making it difficult to guarantee the consistency of object identity and the continuity of motion.

Method used

A method based on object segmentation is adopted, which extracts and encodes the structural representation information of each object through a video object segmentation model. The conditional diffusion model is used to introduce object-level structured feature sequences during the iterative denoising process to generate future video frames, ensuring the identity, shape and motion constraints of objects.

Benefits of technology

It significantly improves the long-term consistency and physical realism of objects in predicted videos, reduces object distortion and violations of physical laws, enhances the credibility and detail sharpness of prediction results, and strengthens the interpretability and user control of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932161B_ABST
    Figure CN120932161B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video prediction, and more particularly to a video sequence prediction method and system based on object segmentation guidance, the prediction method comprising receiving a historical video frame sequence, and performing video object segmentation and tracking processing on the historical video frame sequence to generate structural representation information of each object in each frame and assign a unique tracking ID to each object; encoding the structural representation information and the tracking ID into an object-level structured feature sequence; inputting the object-level structured feature sequence as a guidance condition into a conditional diffusion model to generate latent space intermediate features representing future video frames through an iterative denoising process; and decoding the latent space intermediate features into a pixel-level future video frame sequence. The present application solves the problems of existing video prediction techniques in terms of object consistency, physical reality and error accumulation, and also expands the application potential in complex scenarios, providing support for video prediction in the fields of autonomous driving, robot perception, content creation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video prediction technology, and in particular to a video sequence prediction method and system based on object segmentation guidance. Background Technology

[0002] Video prediction is an important research area at the intersection of computer vision and artificial intelligence. Its goal is to generate future consecutive frames based on existing video frame sequences. This technology has wide-ranging value in many practical applications, such as behavior prediction in autonomous driving, robot environmental interaction simulation, video content completion and generation, and film and television special effects production.

[0003] Early video prediction methods relied primarily on traditional computer vision techniques, such as optical flow, which infers the content of the next frame by estimating pixel motion. However, these methods struggle to effectively handle issues such as occlusion, deformation, and lighting variations in complex scenes, often resulting in poor performance in terms of detail preservation and temporal consistency.

[0004] With the rapid development of deep learning, encoder-decoder structures based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have gradually become the mainstream method for video prediction. For example, models such as PredRNN improve the ability to predict long sequences to some extent by introducing a spatiotemporal memory mechanism. Nevertheless, these methods still generally suffer from problems such as blurred generated frames and severe loss of details, especially when generating long sequences, where the visual quality deteriorates significantly.

[0005] In recent years, generative models such as Generative Adversarial Networks (GANs) and Diffusion Models have been introduced into video prediction tasks, significantly improving the visual fidelity of generated frames. Diffusion Models, in particular, can synthesize high-quality, detailed video frames through a progressive denoising generation method. Representative works such as Video Diffusion Models (VDM) model video distribution in a latent space and utilize a 3D U-Net structure for spatiotemporal feature fusion, achieving remarkable generative results.

[0006] However, existing video prediction methods based on diffusion models still have a fundamental limitation: they lack the ability to explicitly model independent semantic objects (such as people, vehicles, and animals) in videos. Models perform overall spatiotemporal modeling at the pixel or feature block level, making it difficult to guarantee the consistency of identity, motion coherence, and structural integrity of various objects over long-term predictions. Therefore, in complex dynamic scenes, phenomena that violate physical laws, such as object distortion, disappearance, and unreasonable trajectories, are prone to occur.

[0007] To address this, researchers began exploring object-centric video prediction methods, attempting to decompose a scene into multiple independent moving entities and predict their appearance and motion separately. However, these methods typically rely on unsupervised object discovery mechanisms, and the extracted objects lack semantic clarity and stability, limiting their practical application effectiveness.

[0008] Meanwhile, video object segmentation and tracking technologies have made significant progress in recent years, such as Segment Anything Model (SAM) and its video extension model SAM2, which can efficiently and accurately extract the masks of any object in a video while maintaining identity consistency. Although these technologies have been widely used in video editing and analysis, there are currently no publicly available solutions for their deep integration with diffusion models to improve object-level consistency and physical plausibility in video prediction.

[0009] In summary, while existing video prediction technologies have continuously improved in terms of generation quality, they remain limited in long-term prediction of complex scenes due to the core problem of lacking explicit object-level modeling, resulting in insufficient consistency and physical realism in prediction results. Integrating high-precision video object segmentation and tracking information into prediction models has become a key direction for overcoming current technological bottlenecks. Summary of the Invention

[0010] In view of this, the present invention aims to provide a video sequence prediction method and system based on object segmentation guidance, so as to solve the technical problem that traditional video prediction methods are difficult to guarantee the long-term motion consistency and physical authenticity of objects.

[0011] To achieve the above objectives, the technical solution created by this invention is implemented as follows:

[0012] A video sequence prediction method based on object segmentation guidance includes the following steps:

[0013] S1: Receive the input historical video frame sequence, and use the video object segmentation model to perform video object segmentation and tracking processing on the historical video frame sequence, generate structural representation information of each object in each frame and assign a continuously unique tracking ID to each object, and obtain object-level spatiotemporal information composed of structural representation information and tracking ID.

[0014] S2: Encode object-level spatiotemporal information into object-level structured feature sequences, which include the appearance features, motion features, and dynamic spatial relationships between objects;

[0015] S3: The object-level structured feature sequence is used as a guiding condition input to the conditional diffusion model, and the potential spatial intermediate features representing future video frames are generated through an iterative denoising process;

[0016] S4: Decode the latent spatial intermediate features representing future video frames into a pixel-level sequence of future video frames.

[0017] Furthermore, in step S1, the video object segmentation model is either the SAM2 model or the Xmem model.

[0018] Furthermore, in step S1, the structural representation information of the object is any one of the object's segmentation mask, bounding box, key point skeleton, and three-dimensional pose parameters.

[0019] Furthermore, in step S2, when the structural representation information is a segmentation mask, the motion trajectory or optical flow information of each object is calculated based on the mask sequence of the object to obtain motion features.

[0020] Furthermore, in step S2, object-level spatiotemporal information is encoded using a GNN neural network or a Transformer architecture. During encoding, each tracked object is treated as a node in the graph, and the node attributes include the object's segmentation mask, appearance features, and motion features. The dynamic spatial relationships between objects are learned through the graph convolutional layers of the GNN neural network or the self-attention mechanism of the Transformer architecture.

[0021] Furthermore, in step S3, the conditional diffusion model is a denoising network with a U-Net architecture. The object-level structured feature sequence is injected into the denoising network through cross-attention mechanism, conditional normalization method or feature map channel dimension splicing method.

[0022] When injected through the cross-attention mechanism, a cross-attention module is introduced into the denoising network to inject the object-level structured feature sequence into the denoising network. At this time, the object-level structured feature sequence is used as the key and value, and the feature map of the denoising network is used as the query.

[0023] When injecting via conditional normalization, the AdaIN layer is used to inject the object-level structured feature sequence into the denoising network;

[0024] When injected via feature map channel-dimensional concatenation, the object-level structured feature sequence is concatenated with the feature map of the denoising network along the channel dimension, serving as the input to the convolutional layer of the denoising network.

[0025] Furthermore, when training the conditional diffusion model, the loss function used includes pixel-level reconstruction loss and object structure consistency loss; the object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frame and the segmentation results of the real frame.

[0026] Furthermore, in step S4, a convolutional network is used as a decoder to decode the intermediate features in the latent space.

[0027] A video sequence prediction system guided by object segmentation includes:

[0028] The input processing module is used to receive the input historical video frame sequence, perform video object segmentation and tracking processing on the historical video frame sequence, generate structural representation information of each object in each frame and assign a continuously unique tracking ID to each object, and obtain object-level spatiotemporal information composed of structural representation information and tracking ID.

[0029] The object spatiotemporal information encoding module is used to encode object-level spatiotemporal information into object-level structured feature sequences. The object-level structured feature sequences include the appearance features, motion features, and dynamic spatial relationships between objects.

[0030] The conditional diffusion prediction module is used to input the object-level structured feature sequence as a guiding condition into the conditional diffusion model, and generate potential spatial intermediate features representing future video frames through an iterative denoising process.

[0031] The output decoding module is used to decode the latent spatial intermediate features representing future video frames into a pixel-level sequence of future video frames.

[0032] Furthermore, the input processing module includes:

[0033] A data receiving unit, used to receive input historical video frame sequences;

[0034] The video object segmentation and tracking unit is used to perform video object segmentation and tracking processing on historical video frame sequences. The video object segmentation and tracking unit is implemented using a video object segmentation model, which is either the SAM2 model or the Xmem model.

[0035] Furthermore, the structural representation information of the object can be any one of the object's segmentation mask, bounding box, key point skeleton, or 3D pose parameters;

[0036] The object spatiotemporal information encoding module includes:

[0037] The motion feature extraction unit is used to calculate the motion trajectory or optical flow information of each object based on the object's mask sequence when the structural representation information is a segmentation mask, so as to obtain motion features.

[0038] The structural relationship encoding unit uses a GNN neural network or Transformer architecture to encode object-level spatiotemporal information. During encoding, each tracked object is treated as a node in the graph. The node attributes include the object's segmentation mask, appearance features, and motion features. The dynamic spatial relationships between objects are learned through the graph convolutional layers of the GNN neural network or the self-attention mechanism of the Transformer architecture.

[0039] Furthermore, conditional diffusion models include:

[0040] A denoising network based on the U-Net architecture;

[0041] The guided injection unit uses a cross-attention mechanism, conditional normalization, or feature map channel dimension splicing to inject object-level structured feature sequences into the denoising network.

[0042] When the injection unit uses the cross-attention mechanism for injection, a cross-attention module is introduced into the denoising network to inject the object-level structured feature sequence into the denoising network. At this time, the object-level structured feature sequence is used as the key and value, and the feature map of the denoising network is used as the query.

[0043] When the guided injection unit uses conditional normalization for injection, the AdaIN layer is used to inject the object-level structured feature sequence into the denoising network.

[0044] When the injection unit injects using the feature map channel dimension concatenation method, the object-level structured feature sequence is concatenated with the feature map of the denoising network in the channel dimension, and used as the input of the convolutional layer of the denoising network.

[0045] Furthermore, the video sequence prediction system based on object segmentation also includes a model training module for training the conditional diffusion model. The loss function used in the model training module includes pixel-level reconstruction loss and object structure consistency loss. The object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frame and the segmentation results of the real frame.

[0046] Furthermore, the output decoding module uses a convolutional network as a decoder to decode the intermediate features in the latent space.

[0047] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0048] (1) Significantly improves the long-term consistency of objects in predicted videos

[0049] This invention provides object-level strong guidance conditions for the conditional diffusion model by explicitly extracting and encoding the mask, tracking ID, and motion trajectory of each independent object. This enables the conditional diffusion model to strictly follow the object's identity, shape, and motion constraints when generating future frames, thereby effectively avoiding problems such as object distortion, disappearance, or identity confusion, and ensuring that objects maintain a stable shape and coherent motion trajectory in multi-frame prediction.

[0050] (2) Enhance the physical authenticity of prediction results

[0051] Because the conditional diffusion model is strongly constrained by the motion laws and spatial relationships of objects during the generation process, it forces the conditional model to follow the dynamic spatial relationships between objects (such as avoidance and following), making the generated object motion more in line with the physical laws of the real world. This greatly reduces the phenomena that violate physical common sense, such as illogical deformation of objects and penetration of other objects, which are common in existing technologies, thereby improving the credibility of the prediction results and making the prediction results more realistic.

[0052] (3) Effectively suppresses error accumulation in long-sequence prediction and improves long-term prediction stability.

[0053] In traditional frame-by-frame prediction methods, pixel-level errors accumulate and amplify with increasing frame count, leading to a rapid decline in prediction quality. This invention, through object-level structured guidance, limits the spread of individual pixel errors, enabling the conditional diffusion model to maintain clear object structure and scene layout when generating long sequences, thus mitigating quality degradation.

[0054] (4) Improve the robustness of prediction for complex dynamic scenarios

[0055] In complex scenarios with multiple objects and multiple movements, this invention can independently model each object and encode the dynamic spatial relationships between objects through GNN neural networks or Transformer architectures. This enables the conditional diffusion model to infer the motion state of each object separately, thereby allowing the conditional diffusion model to handle complex situations such as object interaction and occlusion more accurately, and the prediction results to be more reasonable and reliable.

[0056] (5) Improve the detail sharpness of predicted video frames

[0057] Existing diffusion models require significant modeling resources to maintain the overall scene structure, resulting in insufficient generation of object details. In this invention, the structure and motion constraints of objects are predetermined, eliminating the need for additional computing power to maintain the basic structure of objects. This allows more modeling resources to be focused on generating object textures and local details, resulting in richer details, clearer edges, and significantly improved visual fidelity in the final output video frames.

[0058] (6) Enhance the interpretability and user controllability of the model

[0059] This invention enables a clear object-level representation within the conditional diffusion model through explicit object segmentation, tracking, and encoding processes. Compared to existing end-to-end pixel regression models, the generation logic of the prediction results is easier to trace, improving interpretability. Furthermore, based on this object-level representation, it can be extended to achieve interactive video prediction (such as adjusting the prediction results by editing object masks or trajectories), providing support for personalized needs in fields such as autonomous driving and film and television production. Attached Figure Description

[0060] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0061] Figure 1 This is a schematic flowchart of the background noise suppression method for high-sensitivity images of space targets described in an embodiment of the present invention. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not constitute a limitation thereof.

[0063] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0064] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0065] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0066] The following will refer to Figure 1 The invention will be described in detail with reference to the embodiments.

[0067] like Figure 1 As shown, this embodiment of the invention provides a video sequence prediction method based on object segmentation guidance, comprising the following steps:

[0068] S1: Receives the input historical video frame sequence, and uses a video object segmentation model to perform video object segmentation and tracking processing on the historical video frame sequence, generates structural representation information of each object in each frame, and assigns a continuously unique tracking ID to each object, thereby obtaining object-level spatiotemporal information composed of structural representation information and tracking ID.

[0069] Video object segmentation models can employ the Segment Anything Model 2 (SAM2) model or the Xmem model. In addition to general segmentation models such as SAM2, for specific domains (such as autonomous driving), specialized semantic segmentation or instance segmentation models can also be used to extract specific categories of objects (such as vehicles and pedestrians) as guidance information.

[0070] Structural representation information is used to describe the structure and state of an object, and can be any of the following: object segmentation mask, bounding box, keypoint skeleton, or 3D pose parameters.

[0071] The segmentation mask is prioritized as structural representation information. For each input frame, the video object segmentation model automatically or semi-automatically identifies all objects in the scene and generates a binary segmentation mask for each object. Simultaneously, the video object segmentation model tracks the same object across consecutive frames, assigning each object a tracking ID that remains constant over time. The same logic applies to bounding boxes, keypoint skeletons, and 3D pose parameters.

[0072] After processing by the video object segmentation model, the historical video frame sequence is decomposed into two parallel data streams:

[0073] Raw frame sequence: contains visual information such as color and texture of the video.

[0074] Object Masking and Tracking ID Sequences: A structured representation that precisely describes the location, shape, and identity of each individual object at any given moment.

[0075] S2: Encode object-level spatiotemporal information into object-level structured feature sequences, which include the appearance features, motion features, and dynamic spatial relationships between objects.

[0076] In order for object-level spatiotemporal information to be understood by subsequent deep learning models, it is necessary to encode the object-level spatiotemporal information and convert it into features that can be understood by deep learning models.

[0077] The appearance features of an object refer to the static structural features that reflect the object, including the object's position (e.g., the pixel area corresponding to the mask) and the object's shape (e.g., the outline of a pedestrian or the shape of a vehicle).

[0078] The motion characteristics of an object refer to the features that reflect the dynamic motion state of the object, including the motion trajectory (e.g., the displacement of the mask centroid from t1 to t2) and the motion law (e.g., a pedestrian walking at a constant speed, a vehicle turning direction).

[0079] Dynamic spatial relationships between objects refer to the characteristics that reflect the dynamic interaction between objects, such as vehicle A following vehicle B, or pedestrians and bicycles not colliding.

[0080] For motion feature extraction from segmentation masks:

[0081] Based on the mask sequence of objects, the motion characteristics of each object can be calculated, such as the centroid trajectory calculated using the mask, or the optical flow calculated within the mask region.

[0082] Encoding of the segmentation mask:

[0083] To learn the interactions between objects (such as avoidance and following), graph neural networks (GNNs) or Transformer architectures can be used to encode object-level spatiotemporal information. During encoding, each tracked object is treated as a node in the graph. Node attributes include the object's segmentation mask, appearance features, and motion features. The dynamic spatial relationships between objects are learned through the graph convolutional layers of the GNN or the self-attention mechanism of the Transformer architecture. The resulting object-level structured feature sequence is a high-dimensional feature vector or feature map sequence, encoding the appearance features, motion features, and dynamic spatial relationships between all tracked objects.

[0084] In complex scenarios with multiple objects and multiple movements, this invention can independently model each object and encode the dynamic spatial relationships between objects through GNN neural networks or Transformer architectures. This enables the conditional diffusion model to infer the motion state of each object separately, thereby allowing the conditional diffusion model to handle complex situations such as object interaction and occlusion more accurately, and the prediction results to be more reasonable and reliable.

[0085] Motion feature extraction for bounding boxes:

[0086] Centroid trajectory: Calculate the displacement sequence of the bounding box center point (x+w / 2, y+h / 2) in consecutive frames.

[0087] Scale variation: Records the changes in the width and height (w, h) of the bounding box over time, used to characterize object scaling or distance changes.

[0088] Velocity and acceleration: The velocity vector is calculated by the displacement of the center point between frames, and the acceleration is calculated by the rate of change of velocity, reflecting the dynamic trend of the object.

[0089] Encoding of bounding boxes:

[0090] The coordinate parameters (x, y, w, h) of the bounding box are normalized to the interval [0, 1] to form a numerical vector input.

[0091] To enhance representation capabilities, appearance features within the bounding box region (such as region feature maps extracted by CNNs) can be concatenated into the coordinate vector to form a composite encoding.

[0092] In the temporal dimension, LSTM neural networks or Transformer architectures can be used to model the bounding box trajectory to obtain an encoded representation of the motion trend.

[0093] Motion feature extraction for keypoint skeletons:

[0094] Joint velocity: Two-dimensional coordinates (x, y) of each keypoint i ,y i Calculate the inter-frame difference to obtain the joint velocity vector.

[0095] Skeletal angle changes: The included angle is calculated by using the skeletal edges formed by key points to extract the joint rotation amplitude and rate of change.

[0096] Overall motion pattern: The motion of key points is represented as a spatiotemporal graph structure, and global motion features such as walking and waving patterns are extracted using a spatiotemporal graph convolutional network (ST-GCN).

[0097] Encoding of keypoint skeletons:

[0098] The two-dimensional coordinates (x, y) of key points of a human body or object i ,y i After being uniformly normalized, they are concatenated into a skeleton vector.

[0099] To model structural relationships, a graph structure can be constructed based on joint topology, and a graph convolutional network (GCN) can be used to extract the structural features of the skeleton.

[0100] In terms of temporal sequence, spatiotemporal graph convolutional networks (ST-GCN) can be used to model continuous frame skeleton sequences to obtain motion dynamic features.

[0101] Motion feature extraction for 3D pose parameters:

[0102] 3D joint displacement: the 3D coordinates (x, y) of each joint. i ,y i ,z i Calculate the displacement and velocity between consecutive frames.

[0103] Rotation angle variation: If a parametric model (such as SMPL) is used, the inter-frame differences can be calculated by representing the Euler angles or quaternions of each joint, and the rotation dynamic features can be extracted.

[0104] Motion trajectory and motion energy: By using joint velocity and angular velocity to calculate the overall motion energy, the dynamic amplitude and rhythm of the object can be described.

[0105] Encoding of 3D pose parameters:

[0106] The coordinates of the 3D key points (x) i ,y i ,z i The joint rotation angles of the parameterized human body model are used as input and normalized into a standard numerical vector.

[0107] The spatial dependencies are encoded using a 3D graph convolutional network (3D-GCN) or a Transformer architecture.

[0108] In terms of temporal dimension, RNN neural networks or temporal Transformer architectures can be combined to model the 3D pose changes of consecutive frames and capture the evolution of motion.

[0109] S3: The object-level structured feature sequence is used as a guiding condition input to the conditional diffusion model, and the potential spatial intermediate features representing future video frames are generated through an iterative denoising process.

[0110] Based on the appearance information of the original frame sequence, the object-level structured feature sequence is used as a guiding condition input to the conditional diffusion model. Under the strong guidance of the object-level structured feature sequence, potential spatial intermediate features representing future video frames are gradually generated from the noise.

[0111] The essence of intermediate features in latent space is intermediate features that are in latent space but have not yet been decoded into pixels. These intermediate features are eventually restored to pixel-level future video frames by the decoder.

[0112] Similar to standard video diffusion models, conditional diffusion models are also based on a progressive denoising process. At its core is a denoising network with a U-Net architecture. The input consists of latent spatial intermediate features of the noisy future video frames and the current time step; the output is the predicted noise.

[0113] When training the conditional diffusion model, the loss function used includes pixel-level reconstruction loss and object structure consistency loss. The object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frame and the segmentation results of the real frame, thereby enhancing the ability of the conditional diffusion model to generate structurally correct objects at the training level.

[0114] The key to this invention lies in how to inject object-level structured feature sequences into a denoising network. Specific implementation methods include the following:

[0115] The first type: Cross-Attention mechanism

[0116] In each or key layer of U-Net, a cross-attention module is introduced. The object-level structured feature sequence is used as the key and value, while the U-Net's own feature map is used as the query. In this way, the denoising network can focus on the location and region where the object should be when generating pixels, thus guiding denoising accordingly.

[0117] The second method: Conditional Normalization

[0118] For example, the AdaIN layer is used to incorporate object-level structured feature sequences into U-Net, which transforms object features into parameters (scale and bias) of affine transformation to modulate feature maps in U-Net.

[0119] The third method: Feature map channel dimension splicing

[0120] The object-level structured feature sequence is directly concatenated with the feature map of the denoising network in the channel dimension and used as the input to the convolutional layer of the denoising network.

[0121] The conditional diffusion model starts with pure Gaussian noise, and through repeated application of a denoising network and subtraction of the predicted noise, it generates a clear sequence of object-level structured features that represent future video frames after multiple time steps of iteration.

[0122] This invention provides object-level strong guidance conditions for the conditional diffusion model by explicitly extracting and encoding the mask, tracking ID, and motion trajectory of each independent object. This enables the conditional diffusion model to strictly follow the object's identity, shape, and motion constraints when generating future frames, thereby effectively avoiding problems such as object distortion, disappearance, or identity confusion, and ensuring that objects maintain a stable shape and coherent motion trajectory in multi-frame prediction.

[0123] Because the conditional diffusion model is strongly constrained by the motion laws and spatial relationships of objects during the generation process, it forces the conditional model to follow the dynamic spatial relationships between objects (such as avoidance and following), making the generated object motion more in line with the physical laws of the real world. This greatly reduces the phenomena that violate physical common sense, such as illogical deformation of objects and penetration of other objects, which are common in existing technologies, thereby improving the credibility of the prediction results and making the prediction results more realistic.

[0124] In traditional frame-by-frame prediction methods, pixel-level errors accumulate and amplify with increasing frame count, leading to a rapid decline in prediction quality. This invention, through object-level structured guidance, limits the spread of individual pixel errors, enabling the conditional diffusion model to maintain clear object structure and scene layout when generating long sequences, thus mitigating quality degradation.

[0125] S4: Decode the latent spatial intermediate features representing future video frames into a pixel-level sequence of future video frames.

[0126] A convolutional network is used as the decoder, similar to the decoder of a variational autoencoder (VAE), to decode the intermediate features in the latent space and restore them to a pixel-level sequence of future video frames, thereby obtaining the final prediction result.

[0127] Through the above process, this invention transforms the video prediction task from an informationless, end-to-end pixel regression problem into a more deterministic generation task guided by clear object structure information, thereby fundamentally improving the quality and long-term consistency of predictions.

[0128] This invention also provides a video sequence prediction system based on object segmentation guidance, comprising: an input processing module, an object spatiotemporal information encoding module, a conditional diffusion prediction module, and an output decoding module; wherein, the input processing module is used to receive the input historical video frame sequence, and perform video object segmentation and tracking processing on the historical video frame sequence, generating structural representation information of each object in each frame and assigning a persistently unique tracking ID to each object, thereby obtaining object-level spatiotemporal information composed of structural representation information and tracking ID; the object spatiotemporal information encoding module is used to encode the object-level spatiotemporal information into an object-level structured feature sequence, wherein the pixel-level structured feature sequence includes the appearance features, motion features, and dynamic spatial relationships between objects; the conditional diffusion prediction module is used to input the object-level structured feature sequence as a guiding condition to the conditional diffusion model, and generate potential spatial intermediate features representing future video frames through an iterative denoising process; the output decoding module is used to decode the potential spatial intermediate features representing future video frames into a pixel-level future video frame sequence.

[0129] The input processing module includes a data receiving unit and a video object segmentation and tracking unit. The data receiving unit is used to receive the input historical video frame sequence. The video object segmentation and tracking unit is used to perform video object segmentation and tracking processing on the historical video frame sequence. The video object segmentation and tracking unit is implemented using a video object segmentation model, which is either the SAM2 model or the Xmem model.

[0130] The structural representation information of an object can be any one of the following: object segmentation mask, bounding box, key point skeleton, or 3D pose parameters.

[0131] The object spatiotemporal information encoding module includes a motion feature extraction unit and a structural relationship encoding unit. The motion feature extraction unit is used to calculate the motion trajectory or optical flow information of each object based on the object's mask sequence when the structural representation information is a segmentation mask, in order to obtain motion features. The structural relationship encoding unit uses a GNN neural network or a Transformer architecture to encode the object-level spatiotemporal information. During encoding, each tracked object is treated as a node in the graph, and the node attributes include the object's segmentation mask, appearance features, and motion features. The dynamic spatial relationships between objects are learned through the graph convolutional layer of the GNN neural network or the self-attention mechanism of the Transformer architecture.

[0132] The conditional diffusion model includes a denoising network based on the U-Net architecture and a guided injection unit. The guided injection unit uses a cross-attention mechanism, a conditional normalization method, or a feature map channel dimension concatenation method to inject object-level structured feature sequences into the denoising network.

[0133] When the injection unit uses the cross-attention mechanism, a cross-attention module is introduced into the denoising network to inject the object-level structured feature sequence into the denoising network. At this time, the object-level structured feature sequence is used as the key and value, and the feature map of the denoising network is used as the query.

[0134] When the injection unit uses conditional normalization, the AdaIN layer is used to inject the object-level structured feature sequence into the denoising network.

[0135] When the injection unit injects using the feature map channel dimension concatenation method, the object-level structured feature sequence is concatenated with the feature map of the denoising network in the channel dimension, and used as the input of the convolutional layer of the denoising network.

[0136] The video sequence prediction system guided by object segmentation also includes a model training module for training the conditional diffusion model. The loss function used by the model training module includes pixel-level reconstruction loss and object structure consistency loss. The object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frame and the segmentation results of the real frame.

[0137] The output decoding module uses a convolutional network as a decoder. The convolutional network is a decoder of a variational autoencoder (VAE) to decode the intermediate features in the latent space and restore them to the pixel-level future video frame sequence, thereby obtaining the final prediction result.

[0138] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0139] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A video sequence prediction method based on object segmentation guidance, characterized in that, Includes the following steps: S1: Receive the input historical video frame sequence, and use the video object segmentation model to perform video object segmentation and tracking processing on the historical video frame sequence, generate structural representation information of each object in each frame and assign a continuously unique tracking ID to each object, and obtain object-level spatiotemporal information composed of structural representation information and tracking ID. In step S1, the structural representation information of the object is the keypoint skeleton; for the motion feature extraction of the keypoint skeleton: Joint velocity: Two-dimensional coordinates (x, y) of each keypoint i ,y i Calculate the inter-frame difference to obtain the joint velocity vector; Skeletal angle changes: The included angle is calculated by the skeletal edges formed by key points, and the joint rotation amplitude and rate of change are extracted; Overall motion pattern: The motion of key points is represented as a spatiotemporal graph structure, and global motion features are extracted using a spatiotemporal graph convolutional network; S2: Encode the object-level spatiotemporal information into an object-level structured feature sequence, which includes the object's appearance features, motion features, and dynamic spatial relationships between objects. In step S2, a GNN neural network or a Transformer architecture is used to encode the object-level spatiotemporal information. During encoding, each tracked object is treated as a node in the graph, and the node attributes include the object's segmentation mask, appearance features, and motion features. The dynamic spatial relationships between objects are learned through the graph convolutional layer of the GNN neural network or the self-attention mechanism of the Transformer architecture. S3: The object-level structured feature sequence is used as a guiding condition input to the conditional diffusion model, and the potential spatial intermediate features representing future video frames are generated through an iterative denoising process. In step S3, the conditional diffusion model is a denoising network with a U-Net architecture. The object-level structured feature sequence is injected into the denoising network through a cross-attention mechanism. When injected through the cross-attention mechanism, a cross-attention module is introduced into the denoising network to inject the object-level structured feature sequence into the denoising network. At this time, the object-level structured feature sequence is used as the key and value, and the feature map of the denoising network is used as the query. S4: Decode the latent spatial intermediate features representing future video frames into a pixel-level sequence of future video frames. When training the conditional diffusion model, the loss function used includes pixel-level reconstruction loss and object structure consistency loss. The object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frames and the segmentation results of the real frames.

2. A video sequence prediction system based on object segmentation guidance, characterized in that, include: The input processing module receives the input sequence of historical video frames and performs video object segmentation and tracking on the sequence. It generates structural representation information for each object in each frame and assigns a unique tracking ID to each object, obtaining object-level spatiotemporal information composed of the structural representation information and the tracking ID. The structural representation information of the object is a keypoint skeleton. Motion feature extraction is performed on the keypoint skeleton. Joint velocity: Two-dimensional coordinates (x, y) of each keypoint i ,y i Calculate the inter-frame difference to obtain the joint velocity vector; Skeletal angle changes: The included angle is calculated by the skeletal edges formed by key points, and the joint rotation amplitude and rate of change are extracted; Overall motion pattern: The motion of key points is represented as a spatiotemporal graph structure, and global motion features are extracted using a spatiotemporal graph convolutional network; The object spatiotemporal information encoding module encodes object-level spatiotemporal information into object-level structured feature sequences. These sequences include the object's appearance features, motion features, and dynamic spatial relationships between objects. The module employs a GNN neural network or a Transformer architecture to encode the object-level spatiotemporal information. During encoding, each tracked object is treated as a node in the graph, with node attributes including the object's segmentation mask, appearance features, and motion features. The module learns the dynamic spatial relationships between objects through the graph convolutional layers of the GNN neural network or the self-attention mechanism of the Transformer architecture. The conditional diffusion prediction module takes the object-level structured feature sequence as a guiding condition input to the conditional diffusion model and generates latent spatial intermediate features representing future video frames through an iterative denoising process. The conditional diffusion model is a denoising network with a U-Net architecture, which injects the object-level structured feature sequence into the denoising network through a cross-attention mechanism. When injecting through the cross-attention mechanism, a cross-attention module is introduced into the denoising network to inject the object-level structured feature sequence into the denoising network. At this time, the object-level structured feature sequence is used as the key and value, and the feature map of the denoising network is used as the query. The output decoding module is used to decode the latent spatial intermediate features representing future video frames into a pixel-level sequence of future video frames. When training the conditional diffusion model, the loss function used includes pixel-level reconstruction loss and object structure consistency loss. The object structure consistency loss is achieved by comparing the cross-union ratio loss between the segmentation results of the predicted frames and the segmentation results of the real frames.

Citation Information

Patent Citations

  • Video filling method and device based on diffusion model, equipment and storage medium

    CN117314733A

  • Weak supervision video saliency target detection method and system based on memory-edge guidance

    CN120564108A