Track control video generation method and device based on depth information and time-frequency optimization

Through multi-entity segmentation, depth estimation and time-frequency decomposition, combined with user instructions, optimize the 3D trajectory, use a multi-scale fusion network to generate control signals, and generate potential video representation sequences in the improved Stable Video Diffusion model, solving the problems of insufficient dynamic entity motion control accuracy and poor cross-frame consistency in the existing video generation method, and achieving smoother and more stable video generation.

CN120238709AActive Publication Date: 2025-07-01湖南马栏山视频先进技术研究院有限公司

Patent Information

Application Number
CN202510387765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing video generation methods are insufficient in dynamic entity motion control, poor cross-frame consistency, and it is difficult to effectively integrate depth information and time-frequency optimization, resulting in unsmooth motion control, insufficient spatial authenticity and time-frequency stability.

Method used

Through multi-entity segmentation, depth estimation and time-frequency decomposition, combined with user instructions, optimize the 3D trajectory, use a multi-scale fusion network to generate control signals, and generate video potential representation sequences in the improved Stable Video Diffusion model, and introduce annealing sampling and cross-modal attention mechanisms to improve motion smoothness and time-frequency stability.

Benefits of technology

It significantly improves the motion smoothness, spatial authenticity and timely frequency stability of the generated video, and solves the problems of insufficient motion control accuracy of dynamic entities and poor cross-frame consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238709A_ABST
    Figure CN120238709A_ABST
Patent Text Reader

Abstract

The invention provides a trajectory control video generation method and device based on depth information and time-frequency optimization, and relates to the technical field of image processing, and the method comprises the steps: optimizing a 3D trajectory through multi-entity segmentation, depth estimation and time-frequency decomposition in combination with a user instruction, and generating a control signal through a multi-scale fusion network; finally, the signals and original images are input into an improved Stable Video Diffusion model to generate a video potential representation sequence, the problems that an existing video generation method is insufficient in dynamic entity motion control precision and poor in cross-frame consistency are solved, and through 3D trajectory modeling guided by depth information and a time-frequency joint optimization mechanism, the video potential representation sequence is generated. And the motion smoothness, the space authenticity and the time-frequency stability of the generated video are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and video generation, and in particular to a trajectory control video generation method and device based on depth information and time-frequency optimization. Background Art

[0002] Existing video generation methods mainly rely on 2D image features or rough motion estimation, which makes it difficult to accurately control the 3D motion of dynamic entities, and there is a general problem of insufficient consistency across frames. Traditional methods (such as motion modeling based on keyframe interpolation) can achieve basic motion control by manually defining trajectory points or relying on sparse optical flow fields, but they cannot effectively capture the depth information and complex motion patterns of multiple entities, resulting in the accumulation of 3D trajectory reconstruction errors, which are specifically manifested as depth jump noise (such as unreasonable Z-axis jitter when the entity moves) and motion paths that deviate from physical laws (such as the lack of continuity in the deformation of non-rigid objects). In recent years, although deep learning-based video generation models (such as diffusion models and GANs) have improved in visual effects, their motion control still has significant limitations: the problem of insufficient control accuracy of dynamic entities mainly stems from the neglect of depth information and the coarseness of control point sampling. For example, existing methods mostly use feature alignment in 2D pixel space (such as the attention mechanism guided by optical flow) without introducing a depth estimation network, which leads to projection conflicts between foreground and background entities in 3D space and motion occlusion artifacts (such as an object passing through the surface of another object). At the same time, the uniform grid sampling strategy is difficult to cover high motion-sensitive areas such as joints and edges, resulting in the loss of local motion details.

[0003] In terms of cross-frame consistency, existing methods often rely on temporal convolution or recurrent networks to model inter-frame associations, but their ability to constrain high-frequency motion components (such as fast jitter) is weak. LSTM-based models are prone to motion amplitude attenuation (such as the entity gradually stagnating) in long video generation, while the Transformer architecture is difficult to achieve joint optimization in the time and frequency domains due to the computational complexity of the self-attention mechanism. In addition, the frequency domain energy distribution (such as the amplitude spectrum) of the generated video shows cross-frame mutations, causing high-frequency flicker noise (such as inconsistent lighting or texture jitter). Although existing frameworks (such as Stable Video Diffusion) directly splice control signals (such as segmentation masks) to the input, they do not design a multi-scale feature fusion mechanism, resulting in the gradual attenuation of control signals during the diffusion process of the latent space, and the annealing sampling stage is prone to over-smoothing effects (such as edge blur) or motion drift (such as the entity deviating from the predetermined path). Recent studies have attempted to enhance control accuracy through monocular depth estimation, but have not combined time-frequency analysis to suppress high-frequency noise; in addition, K-means clustering is used to extract control points, but the lack of spatial Gaussian weight guidance causes the control points to deviate from the kinematic key area.

[0004] Therefore, the existing video generation methods have problems of insufficient precision in controlling the movement of dynamic entities and poor cross-frame consistency. The above technical defects highlight the necessity of integrating depth information, time-frequency optimization, and multi-scale control, and also provide an improvement direction for this application. Summary of the Invention

[0005] In view of the above technical problems in the related art, the present invention proposes a trajectory control video generation method and device based on depth information and time-frequency optimization.

[0006] In a first aspect, the present invention provides a trajectory control video generation method based on depth information and time-frequency optimization, including the following steps:

[0007] S1. Perform multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determine the number of control points according to the entity region area, and extract a multi-scale control point set C covering the key motion regions through weighted clustering; the multi-scale control point set C is a 2D control point set;

[0008] S2. Extract the image depth map D through a depth estimation network based on the multi-scale control point set C, and map the multi-scale control point set C combined with the depth values to a global 3D trajectory set T;

[0009] S3. Perform time-frequency decomposition on the global 3D trajectory set T using discrete wavelet transform to obtain a low-frequency approximation component and a high-frequency detail component, and combine the user direction instruction U and the adaptive direction gain Adjust the high-frequency detail component to optimize the global 3D trajectory set T to obtain an optimized 3D trajectory set T';

[0010] S4. Input the multi-entity instance mask set M, the global 3D trajectory set T, and the optimized 3D trajectory set T' into a multi-scale fusion network to generate an entity-level optical flow field O i and multi-scale features and fuse the entity-level optical flow field O i and multi-scale features through a gated cross-scale attention mechanism to generate a multi-scale control signal S;

[0011] S5. Input the multi-scale control signal S and the original image I into an improved Stable Video Diffusion model, and generate a video latent representation sequence through annealing sampling and cross-modal attention during the latent space diffusion process of the improved Stable Video Diffusion model

[0012] Specifically, step S1 specifically includes the following steps:

[0013] S11. The original image Input to the pre-trained Mask R-CNN instance segmentation network to generate a set of multi-entity instance masks where H is the image height, W is the image width; 3 represents the three RGB channels; the entity instance mask element represents the pixel-level segmentation result of the i-th entity in the j-th frame, where i ∈ [1, N] represents the entity index and N is the total number of entities;

[0014] S12. For each entity instance mask element calculate its effective pixel area and dynamically determine the number of control points based on the area ratio and when the entity mask area change rate force to set where α is an empirical parameter for balancing the control point density; represents taking the maximum element value in represents taking the minimum element value in; represents the binary instance mask matrix of the i-th entity in the j-th frame;

[0015] S13. Perform weighted K-means clustering on the set of mask effective pixel coordinates and output the set of clustering centers after iterative convergence Finally, generate a multi-scale control point set covering the key motion areas where N is the total number of entities, represents the set of 2D control point coordinates of the i-th entity in the j-th frame; the weight function of the weighted K-means clustering adopts a spatial Gaussian distribution where μ is the mask centroid, controls the spatial attenuation rate; the key motion areas include edges, joints, significantly deformed areas, and dynamic texture areas.

[0016] Specifically, step S2 specifically includes the following steps:

[0017] S21. Input the multi-scale control point set into the depth estimation network DepthAnythingV2, and extract the dense depth map D of the original image I through its encoder-decoder structure where N is the total number of entities; represents the set of 2D control point coordinates of the i-th entity in the j-th frame;

[0018] S22. Assign a depth value to each control point through the bilinear interpolation formula and then construct the 3D trajectory points of the i-th entity in the j-th frame and through temporal smoothing constraints to eliminate depth jump noise and finally output the global 3D trajectory set where represents the 3D trajectory matrix of the i-th entity within F frames; F is the total number of frames in the video.

[0019] Specifically, step S3 specifically includes the following steps:

[0020] S31. Input the global 3D trajectory set and the temporal direction control instruction input by the user into the time-frequency analysis module together, and perform multi-scale decomposition on the trajectory signal through discrete wavelet transform to obtain the low-frequency approximation component and the high-frequency detail component satisfying where represents the 3D trajectory matrix of the i-th entity within F frames; is the unit direction vector of the i-th entity in the j-th frame; L is the decomposition level;

[0021] S32. Based on the temporal direction control instruction U, perform direction weighting adjustment on the high-frequency detail component , specifically: apply the adaptive direction gain to the l-th layer detail component and at the same time suppress noise through the threshold function to reconstruct the optimized trajectory and finally output the optimized 3D trajectory set after time-frequency optimization where ⊙ represents element-wise multiplication, is the layer-related gain coefficient.

[0022] Specifically, step S4 specifically includes the following steps:

[0023] S41. Input the entity instance mask set the global 3D trajectory set and the optimized 3D trajectory set T′={T′ i} into the multi-scale control feature fusion network together; the multi-scale control feature fusion network consists of a cascaded encoder-decoder structure and a gated cross-scale attention module;

[0024] S42. Based on each entity trajectory in T′, calculate the per-frame dense optical flow field O i through the optical flow generation module, and the value of the per-frame dense optical flow field O i is determined by the projection transformation of the optimized 3D trajectory point in the camera coordinate system; where is the camera's internal parameter, \(R\in SO(3)\), is the external parameter; the optimized 3D trajectory points

[0025] S43. Build a multi-scale feature pyramid: Extract hierarchical features from the original image \(I\) through the ResNet-50 backbone network and downsample the optical flow field \(O\) i to the corresponding scale by bilinear interpolation to generate and fuse it into the hierarchical features in each scale \(l\) through a gated cross-scale attention module to obtain multi-scale features

[0026] The formula of the gated cross-scale attention module is as follows:

[0027]

[0028] where is the learnable weight matrix, represents the convolution operation, [·] is channel concatenation, and ⊙ is element-wise multiplication; \(C\) l is the number of channels of the \(l\)-th layer feature map extracted by the ResNet-50 backbone network, and its value is determined by the network structure; represents the downsampled optical flow field of the \(i\)-th entity at the \(l\)-th scale;

[0029] S44. Concatenate the multi-scale features with the upsampled features of the decoder, and output a spatially aligned multi-scale control signal after fusion by a convolution kernel which satisfies the spatial alignment constraint where \(D\) is the dimension of the control signal; the convolution kernel is a 3×3 convolution kernel; MLP is a multi-layer perceptron; \(S\) j (x,y) represents the motion control intensity weight at the spatial coordinates (x,y) of the \(j\)-th frame.

[0030] Specifically, step S5 specifically includes the following steps:

[0031] S51. Input the multi-scale control signal and the original image into the improved model based on the Stable Video Diffusion framework. First, compress the input image \(I\) into a latent representation through a 3D-VAE encoder At the same time, map the multi-scale control signal \(S\) to a conditional embedding through a spatio-temporal convolutional network During the diffusion denoising process, the latent variable \(z\) t is iteratively updated through an improved UNet architecture, and its time-dependent residual block calculation form is:

[0032]

[0033] Among them, AdaGN(z t ,t) is the embedding vector injected into time step t, representing the adaptive group normalization layer; CrossAttn(z t ,c) is the cross-modal attention module, used to calculate z t The c interaction weight And weighted fusion as conditional features; h = H / 8; w = w / 8; d = 4; t∈[1,T]; Represents the output of the denoising network; Conv 3D It is a 3D convolutional layer that jointly models motion continuity in both the spatial dimension (H×W) and the temporal dimension (implicit stacking between frames);

[0034] S52, introduce annealing sampling strategy: in the denoising step t∈[T,T c ], and when t∈[T c ,1], the weight of the multi-scale control signal S is gradually attenuated γ(t)=min(1,(T c -t) / (T c -1)) to eliminate artifacts caused by over-constraints; finally, the potential representation is stabilized by the frequency domain stabilization module Post-processing outputs a temporally smoothed video latent representation sequence is a learnable frequency domain filter; T end represents the termination time step of the annealing sampling strategy; T c Represents the starting time step of the annealing sampling strategy.

[0035] Specifically, the method further includes:

[0036] S6. Video latent representation sequence Each frame image is divided into local blocks and decomposed into amplitude spectrum M and phase spectrum P through FFT transform. The amplitude spectrum is adaptively smoothed and optimized based on the cross-frame consistency constraint and combined with the phase spectrum P. It is restored to the time domain image block through inverse FFT, and the complete video frame is reconstructed through weighted overlap method to output the final video sequence X. f ' inal .

[0037] Specifically, step S6 specifically includes the following steps:

[0038] S61, video potential representation sequence Input to the frequency domain stabilization module, the video flicker and motion artifacts are eliminated by jointly removing the time-frequency image, and then the video potential representation sequence is transformed into Each frame of the image is divided into local blocks and decomposed into amplitude spectrum M and phase spectrum P through FFT transformation;

[0039] S62. Perform adaptive smoothing filtering on the amplitude spectrum based on cross-frame frequency-domain consistency constraints: Calculate the neighborhood mean using a sliding window, and fuse local and global frequency-domain features through learnable spectral attention weights to generate an optimized amplitude spectrum M′ that suppresses high-frequency noise. Combine the optimized amplitude spectrum M′ with the phase spectrum P, restore the time-domain image block through inverse FFT, and reconstruct the complete video frame through weighted overlap to output the final video sequence X f ′ inal 。

[0040] Specifically, step S11 specifically includes the following steps:

[0041] S111. Input the original image into the backbone network of the pre-trained MaskR-CNN, and extract multi-scale feature maps through convolutional layers and the Feature Pyramid Network (FPN) where l represents the feature pyramid level; the backbone network of the pre-trained Mask R-CNN model is ResNet-101-FPN, and the downsampling ratio is 2 l ;

[0042] S112. Based on the feature pyramid, the Region Proposal Network generates candidate regions through a sliding window and filters overlapping regions through non-maximum suppression, retaining the top N largest candidate regions; where K is the initial number of candidates, (x n , y n ) is the center coordinate of the candidate region, (w n , h n ) is the width and height of the candidate region; the threshold θ NMS of non-maximum suppression is taken as 0.7;

[0043] S113. Use the RoIAlign layer to map the N largest candidate regions to the multi-scale feature maps, extract region features F of a fixed size roi and input them into the fully connected classification branch and regression branch. After two NMS screenings, obtain the final set of entity candidate boxes

[0044] S114. For each final set of entity candidate boxes the segmentation head generates a binary mask through a transposed convolutional layer and per-pixel Sigmoid activation and upsamples it to the original resolution H×W through bilinear interpolation to generate a set of multi-entity instance masks aligned with the input image

[0045] Second aspect, the present invention provides a trajectory control video generation device based on depth information and time-frequency optimization, which is based on the trajectory control video generation method based on depth information and time-frequency optimization described in the above first aspect, and includes the following units:

[0046] A control point generation unit, configured to perform multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determine the number of control points according to the entity area, and extract a multi-scale control point set C covering the key motion area through weighted clustering; the multi-scale control point set C is a 2D control point set;

[0047] A global 3D trajectory generation unit, configured to extract an image depth map D through a depth estimation network based on the multi-scale control point set C, and map the multi-scale control point set C combined with the depth value to a global 3D trajectory set T;

[0048] A 3D trajectory optimization unit, configured to perform time-frequency decomposition on the global 3D trajectory set T by using discrete wavelet transform to obtain a low-frequency approximation component and a high-frequency detail component, and combine the user direction instruction U and the adaptive direction gain Adjust the high-frequency detail component to optimize the global 3D trajectory set T to obtain an optimized 3D trajectory set T';

[0049] A multi-scale control signal generation unit, configured to input the multi-entity instance mask set M, the global 3D trajectory set T, and the optimized 3D trajectory set T' into a multi-scale fusion network to generate an entity-level optical flow field O i and multi-scale features and fuse the entity-level optical flow field O through a gated cross-scale attention mechanism i and multi-scale features to generate a multi-scale control signal S;

[0050] A video latent representation sequence generation unit, configured to input the multi-scale control signal S and the original image I into an improved Stable Video Diffusion model, and generate a video latent representation sequence through annealing sampling and cross-modal attention during the latent space diffusion process of the improved Stable Video Diffusion model

[0051] A video denoising unit, configured to Segment each frame of the image into local blocks, decompose them into an amplitude spectrum M and a phase spectrum P through FFT transformation, adaptively smooth and optimize the amplitude spectrum based on cross-frame consistency constraints, combine it with the phase spectrum P, restore it to a time-domain image block through inverse FFT, and reconstruct a complete video frame through weighted overlap to output the final video sequence X f ′ inal 。

[0052] The present invention optimizes the 3D trajectory through multi-entity segmentation, depth estimation, and time-frequency decomposition, combines user instructions, and uses a multi-scale fusion network to generate control signals. Finally, these signals and the original image are input into an improved Stable VideoDiffusion model to generate a sequence of video latent representations, solving the problems of insufficient dynamic entity motion control accuracy and poor cross-frame consistency in existing video generation methods. Through a 3D trajectory modeling and time-frequency joint optimization mechanism guided by depth information, the motion smoothness, spatial authenticity, and time-frequency stability of the generated video are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0054] Figure 1 Schematic diagram of a trajectory control video generation method based on depth information and time-frequency optimization provided by an embodiment of the present invention;

[0055] Figure 2 Schematic diagram of a trajectory control video generation device based on depth information and time-frequency optimization provided by an embodiment of the present invention;

[0056] Figure 3 Schematic diagram of a trajectory control video generation device based on depth information and time-frequency optimization provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The present invention can be explained in detail through the following embodiments. The purpose of providing the present invention is to protect all technical improvements within the scope of the present invention. In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "multiple" and "a plurality" mean two or more, unless otherwise specifically defined.

[0058] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail in combination with the drawings and specific embodiments.

[0059] Embodiment 1

[0060] Refer to Figure 1, this embodiment provides a trajectory control video generation method based on depth information and time-frequency optimization, including the following steps:

[0061] S1. Perform multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determine the number of control points according to the entity area, and extract a multi-scale control point set C covering the key motion area through weighted clustering; the multi-scale control point set C is a 2D control point set;

[0062] The original image I is the image corresponding to the video frame in the video sequence; the total number of frames in the video sequence is F (i.e., the length of the time dimension);

[0063] Extract multi-entity control points of the input image I through Mask R-CNN instance segmentation and weighted K-means clustering, and generate a control point set covering the key motion area by combining Gaussian spatial weights;

[0064] Step S1 specifically includes the following steps:

[0065] S11. Input the original image into the pre-trained Mask R-CNN instance segmentation network to generate a multi-entity instance mask set where H is the image height, W is the image width; 3 is the RGB three channels; the entity instance mask element represents the pixel-level segmentation result of the i-th entity in the j-th frame, where N is the total number of detected entities;

[0066] Input the original image into the pre-trained Mask R-CNN instance segmentation network, and the specific steps for generating the multi-entity instance mask set include the following:

[0067] S111. Input the original image into the backbone network of the pre-trained MaskR-CNN, and extract multi-scale feature maps through the convolutional layer and the Feature Pyramid Network (FPN) where l represents the feature pyramid level; the backbone network of the pre-trained Mask R-CNN model is ResNet-101-FPN, and the downsampling ratio is 2 l ;

[0068] S112. Based on the feature pyramid, the Region Proposal Network (RPN) generates candidate regions through a sliding window and filters overlapping regions through non-maximum suppression (NMS), and retains the top N largest candidate regions; where K is the initial number of candidates, and K takes 2000 in this embodiment, and N takes 1000; (x n , y n ) is the center coordinate of the candidate region, (wn , h n ) is the width and height of the candidate region; the threshold θ of non-maximum suppression NMS Take 0.7;

[0069] Based on Faster R-CNN, Mask R-CNN adds a Mask prediction branch and an ROIAlign layer, and uses ResNet+FPN as the Backbone network for feature extraction, realizing multi-task processing of object detection, classification, and pixel-level segmentation; Mask R-CNN mainly includes three main sub-networks: the backbone network, the RPN network, and the head network.

[0070] The Backbone network is used for feature extraction of images. The Head network includes two main branches: one is the branch for classification and bounding box regression, and the other is the branch for generating segmentation masks.

[0071] The Region Proposal Network (RPN) is a neural network structure for object detection tasks. Its main function is to quickly generate a set of potential target regions (proposals) in the input image (original image), and these regions may be candidate boxes (candidate regions) containing the target.

[0072] The main features and working principles of RPN include:

[0073] Shared convolutional features: RPN shares the same set of convolutional feature maps with the backbone network of the object detection network, which can reduce redundant calculations and improve efficiency.

[0074] Sliding window mechanism: RPN generates multiple anchors with fixed ratios and scales at each position on the convolutional feature map by applying a sliding window, and these anchors serve as potential candidate regions.

[0075] Binary classification task: For each anchor, RPN performs a binary classification task, that is, to determine whether the anchor is foreground (containing the target) or background.

[0076] Bounding box regression: In addition to the classification task, RPN also performs bounding box regression to adjust the position and size of the anchor to make it more tightly enclose the target object.

[0077] Non-maximum suppression (NMS): The generated candidate regions usually overlap. RPN will adopt a non-maximum suppression strategy to remove redundant candidate regions and retain the regions most likely to contain the target.

[0078] End-to-end training: The RPN is trained end-to-end with the object detection network, and the network parameters are optimized through backpropagation.

[0079] Non-Maximum Suppression (NMS) is an algorithm step used in object detection, mainly used to remove redundant detection results and retain the most likely bounding boxes. In object detection tasks, multiple candidate boxes (candidate regions) are usually obtained. These candidate boxes may overlap or be very close, so a method is needed to select the best box.

[0080] The basic idea of non-maximum suppression is as follows:

[0081] Sorting: First, all candidate boxes are sorted according to their scores, usually based on classification confidence or other relevant scoring metrics.

[0082] Select the box with the highest score: Select the candidate box with the highest score as the reference box.

[0083] Calculate the Intersection over Union (IoU): For each remaining candidate box, calculate the ratio of the overlapping part of it and the current reference box to the sum of the areas of the two boxes (i.e., the IoU).

[0084] Suppress low-score boxes: If the IoU of a certain candidate box and the reference box exceeds the set threshold θ NMS , it is considered that their overlap is too high, and at this time, the box with a lower score is removed.

[0085] Iterative processing: Take the next unremoved box with the highest score as the new reference box, and repeat the above process until all boxes are processed.

[0086] S113. Use the RoIAlign layer to map the N largest candidate regions to the multi-scale feature map, and extract the region features F with a fixed size roi Input to the fully connected classification branch and regression branch, and after two NMS screenings, the final set of entity candidate boxes is obtained

[0087] The classification branch is used to output the entity category probability The regression branch is used to output the bounding box offset The fixed size is 7×7×256, and the region features The threshold θ of the secondary NMS NMS Take 0.5; C class is the number of categories in the pre-training dataset;

[0088] S114. For each final set of entity candidate boxes The segmentation head generates a binary mask through the deconvolution layer and per-pixel Sigmoid activation And upsample it to the original resolution H×W through bilinear interpolation to generate an instance-level mask aligned with the input image

[0089] The RoIAlign (Region of Interest Align) layer is an important component in Mask R-CNN, which is used to accurately map the candidate region (Region of Interest, RoI) from the original image space to the corresponding region on the feature map. This is the prior art and will not be elaborated here.

[0090] S12. For each entity instance mask element Calculate its effective pixel area And dynamically determine the number of control points based on the area ratio And when the change rate of the entity mask area Force to set where α is an empirical parameter for balancing the control point density; Denote taking The maximum element value in Denote taking The minimum element value in; Denote the binary instance mask matrix of the i-th entity in the j-th frame, and its value is defined as: when the pixel coordinate (x, y) belongs to the entity segmentation region Otherwise it is 0, which is used to identify the pixel-level coverage range of the entity in the image.

[0091] In this embodiment, α takes 50, and it can also be modified according to actual needs; when the change rate of the dynamic area of the video (such as the optical flow gradient) exceeds the threshold (10), setting ≥3 control points can construct a quadratic or higher-order polynomial trajectory. By increasing the sampling point density and curve flexibility, it can ensure accurate parametric modeling of high-dynamic motions such as fast deformation and rotation, and avoid motion distortion caused by underfitting. Therefore, when the change rate of the entity mask area Force to set To ensure the dynamic motion capture ability;

[0092] S13. Perform weighted K-means clustering on the set of mask effective pixel coordinates After iterative convergence, output the set of clustering centers Finally, generate a multi-scale control point set covering the key motion regions where N is the total number of entities, Denote the set of 2D control point coordinates of the i-th entity in the j-th frame; the weight function of the weighted K-means clustering adopts a spatial Gaussian distribution where μ is the centroid of the mask, Control the spatial attenuation rate; the key motion regions include edges, joints, significantly deformed regions, and dynamic texture regions.

[0093] Multi-scale control point set It is required to be evenly distributed on the entity appearance and cover key motion regions such as edges and joints, serving as the basic input for subsequent 3D trajectory construction. The multi-scale control point set C is a 2D control point set.

[0094] The weighted K-means clustering algorithm is an improved version of the K-means clustering algorithm. It deals with the influence degree of different samples in the dataset on the clustering result by introducing weights. In traditional K-means clustering, the distance from each sample point to the cluster center is equally weighted, while in weighted K-means clustering, different sample points can be assigned different weights according to their importance or reliability.

[0095] The steps of weighted K-means clustering usually include:

[0096] Initialize the cluster centers and weights: First, select some initial cluster centers and assign an initial weight to each sample point.

[0097] Calculate the weighted distance: For each sample point, calculate its weighted distance to each cluster center. The calculation method of the weighted distance can be a simple product (i.e., weight multiplied by distance), or a more complex functional form.

[0098] Update the cluster centers and weights: According to the results of the weighted distance, assign each sample point to the nearest cluster center, and update the position and weight of the cluster center. This step may involve recalculating the mean of the cluster center and adjusting the weights of the sample points.

[0099] Iterative optimization: Repeat the above steps until the cluster centers no longer change significantly or reach the preset convergence condition.

[0100] Weighted K-means clustering is applied in many fields, such as machine learning, data analysis, image processing, etc. It can effectively handle sample points with different importance, thereby improving the accuracy and robustness of clustering.

[0101] In step S13, for the set of mask valid pixel coordinates Perform weighted K-means clustering, and output the set of cluster centers after iterative convergence Specifically, it includes the following steps:

[0102] S131. Based on each entity mask The set of valid pixel coordinates When initializing the clustering centers, a density-driven strategy is adopted. First, the probability density function of the spatial distribution of pixel coordinates is calculated. The initial centroids are selected through kernel density estimation. where is the kernel bandwidth parameter; is the local density maximum point; x′, y′ represent the normalized coordinate values after geometric correction or spatio-temporal transformation (such as the new coordinates obtained through affine transformation or optical flow estimation), which are used for feature alignment or motion trajectory correction, and their range is usually [-1, 1] or [0, 1], depending on the normalization method after transformation;

[0103] Kernel Density Estimation (KDE) is a non-parametric statistical method used to estimate the probability density function. It is a method for estimating the probability density function of a random variable, which is a prior art and will not be elaborated here.

[0104] S132. In the iterative optimization stage, the assignment step in the t-th iteration assigns each pixel (x, y) to the nearest neighbor clustering center. The distance metric fuses the spatial coordinates and Gaussian weights:

[0105]

[0106] where the weight function is the masked centroid, controls the weight decay range;

[0107] The update step recalculates the weighted centroids of each cluster:

[0108]

[0109] where is the set of pixels in the k-th cluster;

[0110] The iteration termination condition is the centroid offset or reaching the maximum number of iterations T max ;

[0111] Finally, the set of clustering centers is output where final represents the value of t at the end of iteration, ensuring that it evenly covers the entity's apparent features in the masked space and is preferentially distributed in areas with high edge curvature and motion sensitivity (such as limb joints and turning points of object contours). In addition to edges and joints, key motion areas also include areas with significant deformation (such as the bent parts of non-rigid objects) and dynamic texture areas (such as cloth folds and fluid surface fluctuations), especially in non-rigid entities or complex motion scenarios, additional coverage is required.

[0112] In this example, ò = 1 pixel, and the maximum number of iterations Tmax = 100;

[0113] S2. Extract the image depth map D from the depth estimation network based on the multi-scale control point set C, and map the multi-scale control point set C combined with the depth values to the global 3D trajectory set T;

[0114] Step S2 specifically includes the following steps:

[0115] S21. Input the multi-scale control point set into the depth estimation network DepthAnythingV2, and extract the dense depth map D of the original image I through its encoder-decoder structure ; where N is the total number of entities, represents the 2D control point coordinate set of the i-th entity in the j-th frame;

[0116] The encoder-decoder structure is used to extract the dense depth map D of the original image I;

[0117] Step S21 specifically includes the following steps:

[0118] S211. Extract the multi-scale feature map of the original image through the encoder where D l = {256, 512, 1024, 2048} corresponds to the number of feature channels at different levels, the encoder is a VisionTransformer encoder, version ViT-L / 16; the downsampling ratio is 2 l ; l is the number of network layers;

[0119] The encoder is a deep learning model that has been pre-trained on a large dataset (such as ImageNet). These models (such as ResNet, VGG, etc.) learn rich image features through training on large-scale datasets. The original image can extract multi-level spatial features through the encoder, and then adapt to the video generation task through lightweight fine-tuning (only optimizing some high-level parameters). At the same time, combined with data augmentation in the time dimension (such as random frame sampling, temporal flipping) and hybrid loss (reconstruction loss + adversarial loss + temporal consistency regularization term) for joint training to ensure that the feature representation takes into account both static semantics and dynamic relevance.

[0120] The Depth Estimation Network DepthAnythingV2 utilizes a pre-trained Vision Transformer (ViT-L / 16) encoder to extract multi-scale feature maps of the input image. Here, Vision Transformer (ViT-L / 16) is a component in the DepthAnythingV2 network, responsible for the task of feature extraction;

[0121] Vision Transformer (ViT-L / 16): Vision Transformer is a deep learning model for image classification. It is based on the Transformer architecture, originally used for natural language processing tasks. The "L" in ViT-L / 16 represents "large", i.e., the large model version, and "16" indicates that the input image is divided into 16x16 image patches. This pre-trained model can capture high-level semantic information in the image.

[0122] DepthAnythingV2 is a network for depth estimation that uses the pre-trained ViT-L / 16 as its feature extractor. The goal of depth estimation is to predict the depth information of each pixel in the image from a single or multiple images, i.e., to recover three-dimensional spatial information from a single two-dimensional image.

[0123] The DepthAnythingV2 network uses ViT-L / 16 to process the input image. ViT-L / 16 divides the input image into image patches of a fixed size and extracts features through its Transformer structure. These features contain multi-scale information of the image, which helps with subsequent depth estimation.

[0124] Different layers of ViT-L / 16 will output feature maps of different scales. The lower layers may capture detailed information, while the higher layers capture more abstract semantic information. These multi-scale feature maps are very important for the depth estimation task because different objects and structures in the scene may appear at different scales.

[0125] Downsampling refers to reducing the resolution of the feature map. In ViT-L / 16, as the number of network layers increases, the resolution of the feature map gradually decreases, which is usually achieved through convolutional operations between layers. The downsampling ratio refers to the reduction ratio of the feature map relative to the original image. For example, if the original image size is 256x256 and the feature map size is 64x64, then the downsampling ratio is 4;

[0126] S212. The decoder gradually upsamples the feature map through cascaded transposed convolution modules to obtain enhanced features The regression head uses convolutional operations and the Sigmoid activation function to process the enhanced features Mapped to the dense depth map D of the original image I, satisfying in is a learnable parameter, σ is the Sigmoid activation function; D(x,y) represents the normalized relative depth value of the dense depth map D at the pixel coordinate (x,y), with a value range of [0,1], where 0 represents the nearest and 1 represents the farthest;

[0127] The decoder gradually upsamples the feature map through cascaded transposed convolution modules. Specifically, each transposed convolution module first increases the resolution of the feature map by 2 times, concatenates it with the corresponding layer features of the encoder, and then refines it through a 3×3 convolution residual block, and finally outputs the enhanced feature. The regression head enhances the features through convolution operation and Sigmoid activation function. Mapped to normalized relative depth map That is, the dense depth map D of the original image I satisfies in is a learnable parameter, σ is the Sigmoid activation function, and the depth value range is [0,1].

[0128] The Cascaded Transposed Convolution Module is a structural design for progressive upsampling in deep learning, which is usually used in image generation, super-resolution reconstruction or segmentation tasks. Its core idea is to gradually restore high-resolution details from low-resolution features through the cascade of multiple layers of transposed convolution.

[0129] The transposed convolution module belongs to the decoder component. The decoder contains multiple transposed convolution modules. First, the feature map resolution is increased by 2 times, and then it is concatenated with the corresponding layer features of the encoder and refined by a 3×3 convolution residual block, and finally the enhanced features are output. The regression head enhances the features through convolution operation and Sigmoid activation function. The ViT-L / 16 encoder only extracts global features through patch embedding and self-attention layers, while the transposed convolution module is located at the decoder end (such as the U-Net architecture), which is responsible for upsampling the low-resolution feature map by 2 times (restoring spatial details) and fusing it with the encoder jump connection features to achieve step-by-step resolution improvement. The division of labor between the two is: the encoder reduces the dimension to extract semantic information, and the decoder transposed convolution reconstructs high-resolution output.

[0130] S22, use the bilinear interpolation formula for each control point Assigning Depth Values Then construct the 3D trajectory point of the i-th entity in the j-frame And through the timing smoothing constraint Eliminate the depth jump noise and finally output the global 3D trajectory set where represents the 3D trajectory matrix of the i-th entity within F frames; F is the total number of frames in the video;

[0131] The bilinear interpolation formula is shown as follows:

[0132]

[0133] where w(m,n) is the interpolation weight function; D(m,n) represents the discrete sampling value of the density map D at the integer coordinates (m,n) (such as the pixel value of the m-th row and n-th column of the feature map), and D(x,y) is the estimated value at the continuous coordinates (x,y) calculated by bilinear interpolation, which is obtained by weighted averaging the D(m,n) values of the adjacent four grid points (m,n), (m+1,n), (m,n+1), (m+1,n+1), and its weight is determined by the fractional parts of x and y, thus realizing the smooth transition from discrete to continuous.

[0134] This step constructs the multi-entity 3D trajectory based on the DepthAnythingV2 depth estimation network and fuses the temporal smoothing constraint to eliminate the depth jump noise;

[0135] S3. Perform time-frequency decomposition on the global 3D trajectory set T using discrete wavelet transform to obtain the low-frequency approximation component and the high-frequency detail component, and combine the user direction instruction U and the adaptive direction gain Adjust the high-frequency detail component to optimize the global 3D trajectory set T to obtain the optimized 3D trajectory set

[0136] Step S3 specifically includes the following steps:

[0137] S31. Input the global 3D trajectory set and the temporal direction control instruction input by the user into the time-frequency analysis module, and perform multi-scale decomposition on the trajectory signal through discrete wavelet transform to obtain the low-frequency approximation component and the high-frequency detail component satisfying where represents the 3D trajectory matrix of the i-th entity within F frames; is the unit direction vector of the i-th entity at the j-th frame; L is the decomposition layer number, and by default L = 3;

[0138] Specifically, performing multi-scale decomposition on the trajectory signal through discrete wavelet transform is: applying the Daubechies-4 wavelet basis function ψ(t) to decompose the three-dimensional components of each entity trajectory T i respectively; the three-dimensional components include the X-axis, Y-axis, and Z-axis;

[0139] The Discrete Wavelet Transform (DWT) is an important time-frequency analysis tool that decomposes a signal into wavelet components of different scales (frequencies).

[0140] Daubechies-4 is a specific wavelet basis function proposed by Ingrid Daubechies and is one of the Daubechies wavelet series.

[0141] The time-frequency analysis module consists of three parts: the discrete wavelet transform (DWT) decomposition layer, the direction gain adjustment unit, and the signal reconstruction algorithm. It decomposes the trajectory signal into a low-frequency approximation component (representing the global motion trend) and a high-frequency detail component (representing local motion noise and details) through a wavelet basis function (such as Daubechies-4), adaptively weights the high-frequency components in combination with the user's direction instruction (such as enhancing the motion amplitude in a specific direction and suppressing irrelevant noise), and finally reconstructs the optimized trajectory signal through the inverse wavelet transform to achieve the time-frequency joint optimization of the motion path.

[0142] S32. Based on the user direction instruction U, perform direction weighting adjustment on the high-frequency detail component Specifically: Apply an adaptive direction gain to the l-th layer detail component while suppressing noise through a threshold function to reconstruct the optimized trajectory Finally, output the optimized 3D trajectory set after time-frequency optimization where ⊙ represents element-wise multiplication; τ l = 0.05 × 2 l-L ; is the layer-related gain coefficient, λ l = 0.1 × 2 L-l ; λ l is used to balance the loss contributions of different decoder levels (such as shallow details and deep semantics), and usually takes values that decrease with the increase in the level (such as λ = 1.0 for the shallow layer and λ = 0.2 for the deep layer), which is determined by grid search or heuristic rules based on feature sensitivity to optimize the multi-scale feature alignment effect.

[0143] S4. Input the multi-entity instance mask set M, the global 3D trajectory set T, and the optimized 3D trajectory set T' into the multi-scale fusion network to generate the entity-level optical flow field O i and the multi-scale features and fuse the entity-level optical flow field O i with the multi-scale features Generate a multi-scale control signal S; the frame-by-frame control signal S integrates multi-modal information of entity motion, segmentation, and image content;

[0144] Step S4 specifically includes the following steps:

[0145] S41. Input the entity instance mask set the global 3D trajectory set and the optimized 3D trajectory set T′ = {T′ i} into the multi-scale control feature fusion network together; the multi-scale control feature fusion network consists of a cascaded encoder-decoder structure and a gated cross-scale attention module;

[0146] S42. Based on each entity trajectory in T′ calculate the frame-by-frame dense optical flow field O through the optical flow generation module i , the value of the frame-by-frame dense optical flow field O i is determined by the projection transformation of the optimized 3D trajectory points in the camera coordinate system; where is the camera internal parameter, R ∈ SO(3), is the external parameter; the optimized 3D trajectory points

[0147] are directly generated based on depth estimation and fuse the time-frequency domain motion constraint and the user control intention.

[0148] represents the optimized 3D trajectory matrix of the i-th entity in the j-th frame (output of step S3); T′ i is the specific trajectory point (such as the position coordinate) in this trajectory matrix, and the two notations are equivalent and contextually consistent.

[0149] The optical flow generation module is: where represents the displacement vector of the pixel (x, y) of the i-th entity in the j-th frame;

[0150] S43. Construct a multi-scale feature pyramid: Extract hierarchical features from the original image I through the ResNet-50 backbone network and downsample the optical flow field O i to the corresponding scale through bilinear interpolation to generate and fuse it into the hierarchical features through the gated cross-scale attention module at each scale level l to obtain the multi-scale features

[0151] The formula of the gated cross-scale attention module is:

[0152] ​

[0153] wherein is a learnable weight matrix, represents a convolution operation, [·] is channel concatenation, and ⊙ is element-wise multiplication; C l is the number of channels of the feature map of the l-th layer extracted by the ResNet-50 backbone network, and its value is determined by the network structure, C l ={256, 512, 1024, 2048}; represents the downsampled optical flow field of the i-th entity at the l-th layer scale;

[0154] S44. Through skip connections, multi-scale features are concatenated with the upsampled features of the decoder, and after fusion by the convolution kernel, a spatially aligned multi-scale control signal is output which satisfies the spatial alignment constraint where D is the dimension of the control signal; the convolution kernel is a 3×3 convolution kernel; MLP is a Multi-Layer Perceptron, and its role is to perform non-linear transformation and dimension compression on multi-scale features to generate cross-scale gating weights for dynamically fusing motion control signals at different resolutions; S j (x, y) represents the motion control intensity weight at the spatial coordinates (x, y) of the j-th frame, which is calculated by normalizing the optical flow field amplitude with Softmax and the control point distribution density, and is used to characterize the contribution weight of this position to trajectory optimization;

[0155] In this embodiment, D = 64;

[0156] S5. Input the multi-scale control signal S and the original image I into the improved StableVideoDiffusion model, and the improved StableVideoDiffusion model generates a sequence of video latent representations through annealing sampling and cross-modal attention during the latent space diffusion process

[0157] Specifically, the improved StableVideoDiffusion model introduces an annealing sampling strategy and a cross-modal spatio-temporal attention module during the latent space diffusion process. The specific improvements include: embedding a cross-modal attention mechanism in the UNet architecture, and dynamically calculating the spatio-temporal correlation weight α t between the latent variable z t,j and the multi-scale control signal conditional embedding c(x, y) to achieve fine-grained alignment between the control signal and the diffusion process;

[0158] Specifically, the original image I is compressed into a latent space representation z0 by a 3D-VAE encoder, and the per-frame control signal S is mapped to a conditional embedding c; the denoising latent variable z is iteratively processed through an improved UNet architecture t , fusing the control signal and temporal features, and adopting an annealing strategy to dynamically adjust the control signal weight to balance the generation freedom and constraint strength. The denoised latent variable is optimized for temporal coherence by a frequency domain stabilization module, and finally, a sequence of video latent representations is output by a 3D-VAE decoder The frequency domain stabilization module includes FFT frequency domain filtering and inverse transformation; where H and W are the frame height and width, and F is the total number of frames

[0159] The improved Stable Video Diffusion model is a model based on the Stable Video Diffusion framework

[0160] The Stable Video Diffusion (SVD) framework by default uses a 2D encoder / decoder to process video frame sequences, while the improved Stable Video Diffusion model replaces its encoder / decoder with a 3D-VAE (introducing three-dimensional convolutional kernels) to simultaneously model spatial features and temporal continuity, thus more efficiently extracting and reconstructing the spatio-temporal correlation of the video. The specific improvements include: 1) using 3D-VAE to enhance inter-frame dynamic modeling; 2) incorporating temporal conditional residual blocks and cross-modal attention mechanisms into UNet; 3) expanding 3D convolutional layers to strengthen temporal consistency. These modifications optimize the generation quality and control accuracy of high-dynamic motion (such as rapid deformation and multi-entity interaction)

[0161] The diffusion process of the Stable Video Diffusion framework is an iterative denoising method for generating video sequences, which is based on deep learning models, especially conditional generation models. The following are the general steps of the diffusion process

[0162] Initialization: First, start with the original video sequence, where the video frames are clear and ordered

[0163] Gradually add noise: During the diffusion process, noise is gradually added to the video frames, and this process is usually divided into multiple time steps. At each step, the amount of added noise gradually increases until the video frames become almost indistinguishable from pure random noise

[0164] Latent representation: At each time step, the mixture of the video frame and noise is converted into a latent space representation (usually in the form of a vector or matrix), which captures the content and noise level of the current frame

[0165] Conditional Embedding: During the denoising process, an additional conditional signal (such as text description, key frames, audio, etc.) may be required. This signal is converted into a conditional embedding through an embedding network to guide the denoising process.

[0166] Denoising Iteration: Starting from the maximum noise level, denoising is performed step by step in reverse. At each step, a pre-trained neural network (such as the U-Net architecture) is used to predict the noise at the current noise level and remove it from the latent representation. This network typically takes the latent representation at the current time step and the conditional embedding as inputs.

[0167] Generating Video Frames: After completing the denoising for all time steps, the final latent representation is converted back to the video frame domain to generate a clear video sequence.

[0168] Post-Processing: The generated video may require post-processing, such as adjusting brightness and contrast, to improve the visual effect.

[0169] The key to this process lies in training a neural network that can accurately predict noise so that high-quality video content can be gradually restored during the denoising process. Through this method, the Stable Video Diffusion framework can generate coherent and realistic video sequences given some conditions (such as text, audio, or other video frames).

[0170] Step S5 specifically includes the following steps:

[0171] S51. Input the multi-scale control signal and the original image into the improved model based on the Stable Video Diffusion framework. First, the input image I is compressed into a latent representation by the 3D-VAE encoder At the same time, the multi-scale control signal S is mapped into a conditional embedding through the spatio-temporal convolutional network During the diffusion denoising process, the latent variable z t is iteratively updated through the improved UNet architecture, and its time-dependent residual block calculation form is:

[0172]

[0173] where AdaGN(z t ,t) is the embedding vector injected at time step t, representing the adaptive group normalization layer; CrossAttn(z t ,c) is the cross-modal attention module used to calculate the c interaction weights of z t and weighted fusion into conditional features; h = H / 8; w = w / 8; d = 4; t ∈ [1,T]; ​Denotes the output of the denoising network, whose input is the current noise latent variable z t , the time step t (controlling the denoising stage), and the conditional signal c (such as the optical flow trajectory), used to predict the current noise to achieve progressive denoising; Conv 3D is a three-dimensional convolutional layer that jointly models the motion continuity in the spatial dimension (H×W) and the temporal dimension (implicit frame-by-frame stacking) to ensure the spatio-temporal consistency of the cross-frame generation results.

[0174] The latent variable z t is iteratively generated by the forward diffusion process. By injecting Gaussian noise ò into the initial latent variable z0 (extracted by the VAE encoder) at each time step t, according to the preset noise schedule β t calculate where α t = 1 - β t , is the cumulative product of α t , that is

[0175] 3D-VAE (Variational Autoencoder for 3D data) is a variational autoencoder specifically designed for processing three-dimensional data. It consists of two parts: an encoder and a decoder;

[0176] The task of the encoder is to convert high-dimensional three-dimensional input data (such as three-dimensional shapes represented by point clouds, meshes, or voxelizations) into a low-dimensional latent space (latent variable) representation.

[0177] The task of the decoder is to reconstruct the original high-dimensional three-dimensional data from the low-dimensional latent space (latent variable) representation.

[0178] The improved UNet architecture includes time-dependent residual blocks (by embedding the conditional parameter of the time step t) and cross-modal spatio-temporal attention modules. The two together form the core of the improvement. The former realizes the time coherence modeling of motion dynamics, and the latter completes the fine-grained alignment of the control signal and the latent variable.

[0179] The improved UNet architecture mainly improves the denoising effect in three aspects: First, introduce time-dependent residual blocks. By embedding the time step t into the parameters of the adaptive normalization layer, the network dynamically perceives the characteristics of different denoising stages and distinguishes the phased tasks of noise elimination and detail reconstruction; Second, design a cross-modal spatio-temporal attention module. Use the optical flow field and control point trajectory as conditional signals, and the latent variable z tPerform cross-modal interaction on the Query to achieve refined guidance of the motion direction and amplitude; finally, expand the 3D convolutional kernel in the deep feature fusion layer, and explicitly capture the inter-frame motion continuity through spatio-temporal joint modeling (spatial dimension + time dimension) to avoid the temporal jitter problem caused by independent generation of each frame. These improvements jointly enhance the adaptability of the model to complex motion patterns (such as rapid acceleration and multi-entity interaction), and improve the spatio-temporal consistency of the generated results through the efficient alignment of conditional signals and latent variables.

[0180] S52. Introduce an annealing sampling strategy: At the denoising step \(t\in[T c ,T end \), use the complete conditional embedding \(c\), while at \(t\in[1,T c \), gradually decay the weight \(\gamma(t)=\min(1,(T c -t) / (T c -1))\) of the multi-scale control signal \(S\) to eliminate artifacts caused by over-constraint; the annealing starts from \(T c \) to the final step \(T end \); \(T end \) represents the termination time step of the annealing sampling strategy; \(T c \) represents the starting time step of the annealing sampling strategy, usually taking values in the last 20% - 30% of the total number of steps of the annealing sampling strategy. For example, if the total number of steps is 1000, \(T c = 800\);

[0181] Finally, post-process the latent representation through the frequency domain stabilization module to output a temporally smoothed video latent representation sequence where \(\mathcal{F}\) is a learnable frequency domain filter.

[0182] The above-mentioned finally post-process the latent representation through the frequency domain stabilization module to output a temporally smoothed video sequence Specifically: Perform a fast Fourier transform on each frame of features to obtain the frequency domain components Apply the spectral attention mechanism Then restore the time domain features through the inverse FFT Finally, output a temporally smoothed video latent representation sequence through the 3D-VAE decoder

[0183] S6. Divide each frame of the video latent representation sequence into local blocks and decompose them into the amplitude spectrum \(M\) and phase spectrum \(P\) through FFT transformation. Then, adaptively smooth and optimize the amplitude spectrum based on the cross-frame consistency constraint and combine it with the phase spectrum \(P\). After restoring it to the time domain image block through the inverse FFT, reconstruct the complete video frame through the weighted superposition method to output the final video sequence \(X'\)final .

[0184] Use the frequency-domain stability module to perform time-frequency energy consistency constraint on the video latent representation sequence to suppress flicker artifacts. Perform frequency-domain stability enhancement on the video sequence generated in step S5. Based on the cross-frame consistency constraint, adaptively smooth and optimize the amplitude spectrum. The amplitude spectrum M is used to characterize the frequency-domain energy;

[0185] Step S6 specifically includes the following steps:

[0186] S61. Input the video latent representation sequence into the frequency-domain stability module, eliminate video flicker and motion artifacts through time-frequency joint, and then divide each frame image of the video latent representation sequence into local blocks and decompose them into amplitude spectrum M and phase spectrum P through FFT transformation;

[0187] Specifically, first perform window overlapping block processing on each frame image to obtain a set of local image blocks For each image block B j,m,n perform two-dimensional fast Fourier transform (2D-FFT) to obtain the frequency-domain components and calculate its amplitude spectrum M j,m,n (u, v) = ‖F j,m,n (u, v)‖2 and phase spectrum P j,m,n (u, v) = ∠F j,m,n (u, v); where represents the generated image of the jth frame; the window size k of the window overlapping block processing is 32, the window overlapping rate is 50%, and the step size S is 16;

[0188] In this embodiment, the unit of size is pixel, and it can also be set according to actual needs.

[0189] Each frame image is evenly divided into window blocks with a size of W×W (for example, W = 32 pixels), and overlapping regions P (pixels) are set in the horizontal and vertical directions between adjacent windows. The sliding step size S = W - P (in this example, S = 24) to ensure that the windows cover the entire image and the overlapping regions are fused by weighted average to avoid boundary artifacts. After block division, each window is independently input into the encoder to extract local features, and finally they are stitched together in order to form a complete feature map for subsequent fine correction of the motion trajectory.

[0190] ​S62. Based on the cross-frame frequency domain consistency constraint, perform adaptive smoothing filtering on the amplitude spectrum: calculate the neighborhood mean using a sliding window, and fuse local and global frequency domain features through learnable spectral attention weights to generate an optimized amplitude spectrum M' that suppresses high-frequency noise. Combine the optimized amplitude spectrum M' with the phase spectrum P, restore the time-domain image block through inverse FFT, and reconstruct the complete video frame through weighted overlap addition to output the final video sequence X'. final 。

[0191] Specifically, through the sliding window calculate the neighborhood mean and use parametric spectral attention weights to fuse local and global frequency domain features to obtain the optimized amplitude spectrum Combine the optimized amplitude spectrum M' j,m,n with the original phase spectrum P j,m,n to form Restore the time-domain image block B' through inverse FFT j,m,n = IFFT(F' j,m,n ), and reconstruct the complete frame X' using weighted overlap addition j , and finally output the final video sequence with frequency domain smoothing

[0192] Among them, is a learnable parameter, T represents the length of the temporal context window used in the motion trajectory optimization process (for example, T = 3 means considering the adjacent temporal information of the current frame and its previous and next frames), which is used to constrain the trajectory smoothness through local temporal correlation; in this embodiment, T = 3;

[0193] The weighted overlap-add method (WOLA) is a signal processing technique used to recombine blocks after block processing of signals (such as audio signals or image frames) to restore the original continuous signal. WOLA is an extension of the overlap-add method. It introduces a weight factor when recombining blocks to improve the transition smoothness at the block edges. This is prior art and will not be elaborated here.

[0194] In this embodiment, through multi-entity segmentation, depth estimation, and time-frequency decomposition, combined with user instructions to optimize the 3D trajectory, and use a multi-scale fusion network to generate control signals; finally, these signals and the original image are input into an improved Stable VideoDiffusion model to generate a video latent representation sequence, which solves the problems of insufficient dynamic entity motion control accuracy and poor cross-frame consistency in existing video generation methods. Through the 3D trajectory modeling guided by depth information and the time-frequency joint optimization mechanism, the motion smoothness, spatial authenticity, and time-frequency stability of the generated video are significantly improved.

[0195] Example 2

[0196] Reference Figure 2 , this embodiment provides a trajectory control video generation device based on depth information and time-frequency optimization, including the following units:

[0197] The control point generation unit is used to perform multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determine the number of control points according to the entity area, and extract a multi-scale control point set C covering the key motion area through weighted clustering; the multi-scale control point set C is a 2D control point set;

[0198] The global 3D trajectory generation unit is used to extract the image depth map D through the depth estimation network based on the multi-scale control point set C, and map the multi-scale control point set C combined with the depth value to the global 3D trajectory set T;

[0199] The 3D trajectory optimization unit is used to perform time-frequency decomposition on the global 3D trajectory set T by using discrete wavelet transform to obtain the low-frequency approximation component and the high-frequency detail component, and combine the user direction instruction U and the adaptive direction gain Adjust the high-frequency detail component to optimize the global 3D trajectory set T to obtain the optimized 3D trajectory set T';

[0200] The multi-scale control signal generation unit is used to input the multi-entity instance mask set M, the global 3D trajectory set T and the optimized 3D trajectory set T' into the multi-scale fusion network to generate the entity-level optical flow field O i And multi-scale features And fuse the entity-level optical flow field O through the gated cross-scale attention mechanism i And multi-scale features Generate a multi-scale control signal S;

[0201] The video latent representation sequence generation unit is used to input the multi-scale control signal S and the original image I into the improved Stable Video Diffusion model, and the video latent representation sequence is generated through annealing sampling and cross-modal attention during the latent space diffusion process of the improved Stable Video Diffusion model

[0202] The video denoising unit is used for the video latent representation sequence Each frame of the image is segmented into local blocks and decomposed into the amplitude spectrum M and the phase spectrum P through FFT transform, and the amplitude spectrum is adaptively smoothed and optimized based on the cross-frame consistency constraint and then combined with the phase spectrum P, and restored to the time-domain image block through inverse FFT, and the complete video frame is reconstructed through the weighted overlap method to output the final video sequence X f ′ inal .

[0203] Embodiment 3

[0204] Reference Figure 3 , Figure 3 FIG. 9 is a schematic structural diagram of the trajectory control video generation device based on depth information and time-frequency optimization according to this embodiment. The trajectory control video generation device 20 based on depth information and time-frequency optimization according to this embodiment includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module / unit in the above device embodiments are implemented.

[0205] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the trajectory control video generation device 20 based on depth information and time-frequency optimization. For example, the computer program may be divided into the respective modules in Embodiment 2, and the specific functions of each module may refer to the working process of the device described in the above embodiment, which will not be elaborated here.

[0206] The trajectory control video generation device 20 based on depth information and time-frequency optimization may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the trajectory control video generation device 20 based on depth information and time-frequency optimization, and does not constitute a limitation on the trajectory control video generation device 20 based on depth information and time-frequency optimization. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, the trajectory control video generation device 20 based on depth information and time-frequency optimization may further include an input / output device, a network access device, a bus, etc.

[0207] The processor 21 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 21 is the control center of the trajectory control video generation device 20 based on depth information and time-frequency optimization, and connects various parts of the entire trajectory control video generation device 20 based on depth information and time-frequency optimization through various interfaces and lines.

[0208] The memory 22 can be used to store the computer programs and / or modules. The processor 21 realizes various functions of the trajectory control video generation device 20 based on depth information and time-frequency optimization by running or executing the computer programs and / or modules stored in the memory 22, and by calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0209] Among them, if the module / unit integrated with the trajectory control video generation device 20 based on depth information and time-frequency optimization is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0210] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.

[0211] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process multiple processes and / or blocksFigure 1 a device with functions specified in one or more boxes

[0212] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 a box or multiple boxes

[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 a box or multiple boxes

[0214] The parts not detailed in the present invention are prior art. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and all changes falling within the meaning and scope of the equivalent elements are intended to be included in the present invention.

Claims

1. A trajectory control video generation method based on depth information and time-frequency optimization, characterized in that: The following steps are involved: S1, performing multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determining the number of control points according to the area of ​​the entity region to extract a multi-scale control point set C covering the key motion area through weighted clustering; the multi-scale control point set C is a 2D control point set; S2, extracting the image depth map D through the depth estimation network based on the multi-scale control point set C, and mapping the multi-scale control point set C combined with the depth value into a global 3D trajectory set T; S3, use discrete wavelet transform to perform time-frequency decomposition on the global 3D trajectory set T to obtain low-frequency approximate components and high-frequency detail components, combined with the user direction command U and the adaptive direction gain D′ l i Adjust the high-frequency detail components to optimize the global 3D trajectory set T to obtain the optimized 3D trajectory set T′; S4: Input the multi-entity instance mask set M, the global 3D trajectory set T and the optimized 3D trajectory set T′ into the multi-scale fusion network to generate the entity-level optical flow field O. i With multi-scale features And the entity-level optical flow field O is fused through the gated cross-scale attention mechanism i With multi-scale features generating a multi-scale control signal S; S5. Input the multi-scale control signal S and the original image I into the improved Stable Video Diffusion model. In the latent space diffusion process of the improved Stable Video Diffusion model, a video potential representation sequence is generated through annealing sampling and cross-modal attention.

2. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11, the original image Input to the pre-trained Mask R-CNN instance segmentation network to generate a set of multi-entity instance masks Among them, H is the image height, W is the image width; 3 is the RGB three-channel; entity instance mask element represents the pixel-level segmentation result of the i-th entity in the j-th frame, where i∈[1,N] represents the entity index and N is the total number of entities; S12. Mask elements for each entity instance Calculate its effective pixel area And dynamically determine the number of control points based on area ratio And when the entity mask area change rate Forced to set Where α is the empirical parameter of the equilibrium control point density; Indicates taking The largest element value in Indicates taking The smallest element value in ; represents the binary instance mask matrix of the i-th entity in the j-th frame; S13, effective pixel coordinate set for mask Perform weighted K-means clustering and output the cluster center set after iterative convergence Finally, a set of multi-scale control points covering the key motion area is generated. Where N is the total number of entities, represents the 2D control point coordinate set of the i-th entity in the j-th frame; the weight function of the weighted K-means clustering adopts spatial Gaussian distribution Among them, μ is the mask centroid, Controlling the spatial attenuation rate; the key motion areas include edges, joints, areas with significant deformation, and dynamic texture areas.

3. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21. Set the multi-scale control points Input to the depth estimation network DepthAnythingV2, through its encoder-decoder structure Extract a dense depth map D of the original image I; where N is the total number of entities; Represents the 2D control point coordinate set of the i-th entity in the j-th frame; S22, use the bilinear interpolation formula for each control point Assigning Depth Values Then construct the 3D trajectory point of the i-th entity in the j-frame And through the timing smoothing constraint Eliminate depth jump noise and finally output a global 3D trajectory set in Represents the 3D trajectory matrix of the i-th entity in the F-frame; F is the total number of video frames.

4. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 2, characterized in that: Step S3 specifically includes the following steps: S31, the global 3D trajectory collection Timing direction control instructions input by the user The signals are input into the time-frequency analysis module, and the trajectory signals are decomposed into multiple scales through discrete wavelet transform to obtain the low-frequency approximate components. and high frequency detail components satisfy in represents the 3D trajectory matrix of the i-th entity in the F frame; is the unit direction vector of the i-th entity in the j-th frame; L is the number of decomposition levels; S32, based on the timing direction control instruction U, the high frequency detail component Perform directional weighted adjustment, specifically: for the l-th layer detail component Apply adaptive directional gain At the same time, through the threshold function D′ l ' i =SoftThreshold(D′ l i ,τ l ) Suppress noise and reconstruct the optimized trajectory The final output is the optimized 3D trajectory set after time-frequency optimization Among them, ⊙ represents element-by-element multiplication, is the layer-dependent gain coefficient.

5. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 1, characterized in that: Step S4 specifically includes the following steps: S41. Mask the entity instance collection Global 3D trajectory collection And optimize the 3D trajectory set T′={T′ i } are jointly input into a multi-scale control feature fusion network; the multi-scale control feature fusion network is composed of a cascaded encoder-decoder structure and a gated cross-scale attention module; S42, based on each entity trajectory in T' The optical flow generation module calculates the frame-by-frame dense optical flow field O i , the frame-by-frame dense optical flow field O i The value of is determined by optimizing the 3D trajectory points Projection transformation in camera coordinate system Determine; among them is the camera internal parameter, R∈SO(3), is an external parameter; the optimized 3D trajectory point S43. Construct a multi-scale feature pyramid: Extract hierarchical features from the original image I through the ResNet-50 backbone network And the optical flow field O i Generate by downsampling to the corresponding scale through bilinear interpolation And in each scale l, the gated cross-scale attention module is used to fuse the hierarchical features to obtain multi-scale features The formula of the gated cross-scale attention module is as follows: Where W g,l , W a,l , is the learnable weight matrix, represents the convolution operation, [·] is channel concatenation, and ⊙ is element-by-element multiplication; C l is the number of channels of the l-th layer feature map extracted by the ResNet-50 backbone network, and its value is determined by the network structure; Represents the downsampled optical flow field of the i-th entity at the l-th scale; S44, multi-scale features through skip connections It is concatenated with the decoder upsampled features and outputs a spatially aligned multi-scale control signal after convolution kernel fusion. It satisfies the spatial alignment constraint Where D is the control signal dimension; the convolution kernel is a 3×3 convolution kernel; MLP is a multi-layer perceptron; S j (x, y) represents the motion control strength weight of the j-th frame at the spatial coordinate (x, y).

6. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 1, characterized in that: Step S5 specifically includes the following steps: S51, multi-scale control signal With the original image Input to the improved model based on the Stable Video Diffusion framework, first the input image I is compressed into a potential representation through the 3D-VAE encoder At the same time, the multi-scale control signal S is mapped into a conditional embedding through the spatiotemporal convolutional network. In the diffusion denoising process, the latent variable z t Through the iterative update of the improved UNet architecture, the time-dependent residual block calculation form is: ò θ (z t ,t,c)=Conv 3D (AdaGN(z t ,t)+CrossAttn(z t ,c)), Among them, AdaGN(z t ,t) is the embedding vector injected into time step t, representing the adaptive group normalization layer; CrossAttn(z t ,c) is the cross-modal attention module, used to calculate z t The c interaction weight And weighted fusion as conditional features; h = H / 8; w = w / 8; d = 4; t∈[1,T]; θ (z t ,t,c) represents the output of the denoising network; Conv 3D It is a 3D convolutional layer that jointly models motion continuity in both the spatial dimension (H×W) and the temporal dimension (implicit stacking between frames); S52, introduce annealing sampling strategy: in the denoising step t∈[T,T c ], and in t∈[1,T c ], the weight of the multi-scale control signal S is gradually attenuated γ(t) = min(1,(T c -t) / (T c -1)) to eliminate artifacts caused by over-constraints; finally, the potential representation is stabilized by the frequency domain stabilization module Post-processing outputs a temporally smoothed video latent representation sequence is a learnable frequency domain filter; T end represents the termination time step of the annealing sampling strategy; T c Represents the starting time step of the annealing sampling strategy.

7. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 1, characterized in that: The method further comprises: S6. Video latent representation sequence Each frame image is divided into local blocks and decomposed into amplitude spectrum M and phase spectrum P through FFT transform. The amplitude spectrum is adaptively smoothed and optimized based on the cross-frame consistency constraint and combined with the phase spectrum P. It is restored to the time domain image block through inverse FFT, and the complete video frame is reconstructed through weighted overlap method to output the final video sequence X. f ' inal .

8. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 7, characterized in that: Step S6 specifically includes the following steps: S61, video potential representation sequence Input to the frequency domain stabilization module, the video flicker and motion artifacts are eliminated by jointly removing the time-frequency image, and then the video potential representation sequence is transformed into Each frame of the image is divided into local blocks and decomposed into amplitude spectrum M and phase spectrum P through FFT transformation; S62. Based on the cross-frame frequency domain consistency constraint, the amplitude spectrum is adaptively smoothed: the neighborhood mean is calculated using a sliding window, and the local and global frequency domain features are fused through learnable spectral attention weights to generate an optimized amplitude spectrum M′ that suppresses high-frequency noise. The optimized amplitude spectrum M′ is combined with the phase spectrum P, and the time domain image block is restored through inverse FFT. The complete video frame is reconstructed through the weighted overlap method to output the final video sequence X f ' inal .

9. The method for generating trajectory control video based on depth information and time-frequency optimization according to claim 2, characterized in that: Step S11 specifically includes the following steps: S111, the original image Input to the pre-trained MaskR-CNN backbone network, and extract multi-scale feature maps through convolutional layers and feature pyramid networks (FPN) Where l represents the feature pyramid level; the backbone network of the pre-trained Mask R-CNN model is ResNet-101-FPN, and the downsampling ratio is 2 l ; S112. Based on the feature pyramid, the region proposal network generates candidate regions through a sliding window The overlapping areas are filtered by non-maximum suppression, and the top N largest candidate areas are retained; where K is the initial number of candidates, (x n ,y n ) is the center coordinate of the candidate region, (w n ,h n ) is the width and height of the candidate region; the threshold value θ for non-maximum suppression NMS Take 0.7; S113, use the RoIAlign layer to map the N largest candidate regions to the multi-scale feature map, and extract the fixed-size regional features F roi The input is sent to the fully connected classification branch and regression branch, and the final entity candidate box set is obtained after two NMS screenings. S114. For each final entity candidate box set The segmentation head generates a binary mask through a deconvolution layer and pixel-wise sigmoid activation. Then upsample to the original resolution H×W through bilinear interpolation to generate a set of multi-entity instance masks aligned with the input image.

10. A trajectory control video generation device based on depth information and time-frequency optimization, characterized in that: The following units are included: A control point generation unit is used to perform multi-entity instance segmentation on the original image I to generate a multi-entity instance mask set M, and dynamically determine the number of control points according to the area of ​​the entity region to extract a multi-scale control point set C covering the key motion area through weighted clustering; the multi-scale control point set C is a 2D control point set; A global 3D trajectory generation unit extracts an image depth map D based on a multi-scale control point set C through a depth estimation network, and maps the multi-scale control point set C combined with the depth value into a global 3D trajectory set T; The 3D trajectory optimization unit is used to use discrete wavelet transform to perform time-frequency decomposition on the global 3D trajectory set T to obtain low-frequency approximate components and high-frequency detail components, combined with the user direction command U and the adaptive direction gain D′ l i Adjust the high-frequency detail components to optimize the global 3D trajectory set T to obtain the optimized 3D trajectory set T′; The multi-scale control signal generation unit is used to input the multi-entity instance mask set M, the global 3D trajectory set T and the optimized 3D trajectory set T′ into the multi-scale fusion network to generate the entity-level optical flow field O i With multi-scale features And the entity-level optical flow field O is fused through the gated cross-scale attention mechanism i With multi-scale features generating a multi-scale control signal S; A video latent representation sequence generation unit is used to input the multi-scale control signal S and the original image I into the improved Stable Video Diffusion model, and the video latent representation sequence is generated by annealing sampling and cross-modal attention in the latent space diffusion process of the improved Stable Video Diffusion model. Video denoising unit, used to transform the video potential representation sequence Each frame image is divided into local blocks and decomposed into amplitude spectrum M and phase spectrum P through FFT transform. The amplitude spectrum is adaptively smoothed and optimized based on the cross-frame consistency constraint and combined with the phase spectrum P. It is restored to the time domain image block through inverse FFT, and the complete video frame is reconstructed through weighted overlap method to output the final video sequence X. f ' inal .

Citation Information

Patent Citations

  • Real-time video quality optimization and enhancement method based on deep learning

    CN119418254A

  • Liveness detection method and apparatus, electronic device, and storage medium

    US20230290187A1

Cited By

  • Video processing method and device, readable storage medium and program product

    CN120877176A

  • Video stream image optimization method and device

    CN121010766A

  • Video stream image optimization method and device

    CN121010766B

  • Image-to-video generation method and device, computer equipment and storage medium

    CN121037649A

  • Single wind measurement laser radar three-dimensional wind field inversion method based on deep learning

    CN121385839A