A highly controllable long-timing video generation method

CN122891784APending Publication Date: 2026-10-09STATE GRID ANHUI ULTRA HIGH VOLTAGE CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610837038.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0007]本发明要解决的技术问题是提供一种高可控长时序视频生成方法,解决了多模态数据融合不充分的问题,显著提升了控制精度,并有效规避长时序生成过程中的误差累积,显著提升长时序视频内容的一致性

Benefits of technology

(1)、本发明构建了相机动作单元CAU体系,将多模态运镜控制数据拆解为48个原子相机动作单元,并映射为CAU时序激活序列,实现了自然语言运镜风格描述文本、关键帧运镜参数和专业运镜文件的统一语义表示,完整保留多模态运镜控制数据的所有细节特征。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122891784A_ABST
    Figure CN122891784A_ABST
Patent Text Reader

Abstract

The application discloses a high-controllable long-time sequence video generation method, which comprises the following steps: firstly, acquiring multi-modal input data and performing data processing; then, constructing a multi-modal fusion generation network based on a double-flow shared attention mechanism to generate a fusion condition vector; finally, constructing a long-time sequence video generation network, generating a preview video based on the constraint of the fusion condition vector, and then generating a target high-definition video in sections. The application solves the problem of insufficient multi-modal data fusion, significantly improves the control accuracy, effectively avoids error accumulation in the long-time sequence generation process, and significantly improves the consistency of long-time sequence video content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video generation technology, specifically a highly controllable long-time-series video generation method. Background Technology

[0002] With the rapid iteration of generative artificial intelligence technology, video generation technology has become one of the core research directions in the field of AIGC (AI Generative Content). It can automatically generate video content that meets user requirements based on the control conditions input by the user, greatly reducing the threshold and cost of digital content creation and promoting technological innovation in fields such as film and television production and virtual human applications.

[0003] Current controllable video generation technologies still face a series of core challenges in their journey towards practical application. Firstly, regarding the uniformity and accuracy of control conditions, existing solutions lack a unified underlying semantic representation framework for heterogeneous modal inputs such as text descriptions, keyframe parameters, and professional animation files. Each modal control signal is typically encoded independently and mapped to a latent space, resulting in significant semantic gaps. This leads to deviations or loss of detail in the cross-modal conversion and fusion of complex user intentions, making frame-level precise control difficult to achieve.

[0004] Secondly, in terms of multimodal feature deep fusion mechanisms, mainstream methods mostly employ simple operations such as feature stitching, addition, or cross-attention, failing to achieve fine-grained, structured interaction between visual content tokens and diverse control tokens. This shallow fusion approach is prone to feature collapse or control attenuation, resulting in insufficient alignment accuracy between generated content and complex control conditions (such as temporal actions and fine camera movements), manifesting as character identity drift, blurred action details, or camera movement becoming disconnected from scene content.

[0005] Third, global consistency and error accumulation in long-term generation remain persistent bottlenecks. Most autoregressive or one-time generation models lack the ability to plan for long sequences. During generation, local errors propagate and amplify with each frame, eventually leading to inconsistencies in video content, subject pose shifts, or scene layout drift. Especially in scenarios requiring closed-loop roaming or long-take narratives, existing methods struggle to ensure temporal coherence and visual closure, severely impacting professional applications.

[0006] Fourth, existing technologies lack sufficient control granularity and user interface flexibility, failing to cover the full range of needs from rapid creative sketches to film-level precision rendering. Solutions often fall into a dilemma: either providing only high-level text-based control at the expense of controllability and accuracy, or requiring users to provide complete, film-industry-grade animation data, creating an excessively high professional barrier. There is a lack of an intelligent mechanism that can automatically adapt the control level based on user input, achieving a balance between freedom and ease of use. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a highly controllable long time-series video generation method, which solves the problem of insufficient multimodal data fusion, significantly improves control accuracy, effectively avoids error accumulation in the long time-series generation process, and significantly improves the consistency of long time-series video content.

[0008] The technical solution of this invention is as follows: A highly controllable long-time-series video generation method specifically includes the following steps: (1) Acquiring multimodal input data: Multimodal input data includes the input reference image, character detail control data and multimodal camera movement control data; (2) Data processing: Convert the reference image into global visual features, convert the character detail control data into a unified control semantic feature sequence, and convert the multimodal camera movement control data into a CAU feature sequence; (3) Construct a multimodal fusion generative network based on a dual-stream shared attention mechanism to generate fusion condition vectors; (4) Construct a long-time video generation network. Based on the constraints of the fusion condition vector, first generate a preview video, and then refine and generate the target high-definition video segment by segment.

[0009] The character detail control data includes natural language descriptions of character facial expressions and action details, as well as standardized text descriptions based on the character control unit AU. The character control unit AU includes multiple AU sub-units, each AU sub-unit corresponding to a standardized text description of a specific character facial expression or action detail. The standardized text descriptions based on the character control unit AU are input into one or more AU sub-units in the character control unit AU to realize the standardized text descriptions of character facial expressions and action details. The multimodal camera control data includes natural language camera style description text, keyframe camera parameters, and professional camera movement files.

[0010] The data processing specifically includes the following steps: S21. Convert the reference image into global visual features, specifically: the input reference image is... The resolution of the reference image is First, the reference image is divided into sections of a fixed size. A sequence of non-overlapping image patches, each image patch being of size [size missing]. This partitioning method is aligned with the input structure of the visual encoder, ensuring that each patch corresponds to a visual token. Then, the pre-trained visual encoder is used to extract spatial visual features, which are used to encode each image patch into a feature vector of fixed dimensions, as shown in the following formula (1): (1); In equation (1), Represents a visual encoder. Representing the One image patch, Representing the The visual features corresponding to each image patch constitute a spatial visual feature sequence. Then, the spatial visual feature sequence Each of them Dimensionality reduction is performed using linear projection to obtain global visual features. ; S22. Convert the character detail control data into a unified control semantic feature sequence. Specifically, input the natural language description of the character's facial expressions and action details, as well as the standardized text description based on the character control unit AU, into the pre-trained CLIP text encoder. The natural language description of the character's facial expressions and action details is encoded to obtain the initial text embedding features. The standardized text description based on the character control unit (AU) is encoded to obtain an embedding vector. Initial text embedding features The initial text compression embedding features are then obtained through custom pooling. Compress and embed features from the initial text and embedding vector The concatenation yields the text embedding features. Then embed the text features After processing by a lightweight adaptive network, unified control semantic features are obtained. The lightweight adaptable network is a lightweight network containing two fully connected layers. The processing procedure is shown in the following formula (2): (2); In equation (2), and These represent the weight matrix and bias of the first fully connected layer, respectively. and These represent the weight matrix and bias of the second fully connected layer, respectively. Represents the activation function of the Gaussian error linear unit; S23. Convert the multimodal camera movement control data into CAU feature sequences. Specifically: First, a camera action unit (CAU) system is constructed, which includes... Each atomic camera action unit corresponds to an ID and a structured natural language description of the camera movement control operation. Then, the multimodal camera movement control data is mapped to a CAU temporal activation sequence. , This represents the total number of frames in the target high-definition video to be generated. This represents the total number of action units in the atomic camera, and finally, the CAU temporal activation sequence is calculated. After being concatenated with structured natural language and encoded by a text encoder, the CAU feature sequence is obtained. CAU characteristic sequence After compressing the CAU dimension using custom pooling, the CAU compressed feature sequence is obtained. , Not less than 2 and less than Integers; The custom pooling specifically involves introducing n learnable query vectors. The input features are aggregated into a fixed-length feature matrix through a cross-attention mechanism, as shown in equation (3) below: (3); In equation (3), when the input features For initial text embedding features At that time, the output is a fixed-length feature matrix. For initial text compression embedding features When input features CAU characteristic sequence At that time, the output fixed-length feature matrix CAU compressed feature sequence ; As a query As keys and values; This represents the cross-attention mechanism.

[0011] The process of mapping multimodal camera movement control data into CAU time-series activation sequences. Specifically: S231. The text describing the camera movement style in natural language is encoded into text features using a pre-trained text encoder. Then, the text features are projected through a projection layer. The mapping is to the CAU time-series activation sequence, as shown in equation (4) below: (4); In equation (4), and These are the weight matrix and bias of the projection layer, respectively. Text features Dimensions Represents the Sigmoid activation function; When the input natural language camera movement style description text is a camera movement style description that already exists in the CAU style library, the corresponding CAU combination template is directly called from the pre-trained CAU style library to automatically generate the corresponding CAU temporal activation sequence. S232. Keyframe camera movement parameters include camera parameters for N keyframes, where each keyframe's camera parameters are mapped to a corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. , see the following formula (5) for details: (5); In equation (5), To normalize the time step, The time of the intermediate frame that needs to be calculated now. This represents the time of the i-th keyframe. This represents the time of the (i+1)th keyframe. These are third-order Bessel basis functions; CAU activation vectors for keyframes; CAU activation vectors for all intermediate frames. Composing a complete frame-by-frame CAU temporal activation sequence ; S233. Analyze the professional camera movement files, extract the camera parameters frame by frame, and map the camera parameters of each frame to the corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. CAU activation vectors of all intermediate frames Composing a complete frame-by-frame CAU temporal activation sequence .

[0012] The Each atomic camera action unit includes six types of camera action units: translation atomic camera action units, rotation atomic camera action units, scaling atomic camera action units, motion curve atomic camera action units, camera movement atomic camera action units, and parametric atomic camera action units.

[0013] The CAU compressed feature sequence In the time dimension Press up to increase speed Sampling was performed to obtain the original preview stage CAU feature sequence. Then, the original preview stage CAU feature sequence Expanding along the time dimension and concatenating them, we obtain the CAU feature sequence for the preview stage. Simultaneously, CAU compresses the feature sequence. The segments are divided into L frames of refined length, then expanded along the time dimension, and spliced ​​together to obtain the CAU segment feature sequence for each segment. .

[0014] The construction of a multimodal fusion generative network based on a two-stream shared attention mechanism to generate fusion condition vectors specifically includes the following steps: S31. Unify the control of semantic features and preview stage CAU feature sequences Concatenated into semantic control features Then global visual features and semantic control features The corresponding query, key, and value vectors are calculated, as shown in equations (6) and (7) below: (6); (7); In equations (6) and (7), , and Representing global visual features The query weight matrix, key weight matrix, and value weight matrix; , and Representing global visual features The corresponding query, key, and value vectors; , and They represent semantic control features respectively The query weight matrix, key weight matrix, and value weight matrix; , and They represent semantic control features respectively The corresponding query, key, and value vectors; S32, Global visual features and semantic control features The query, key, and value vectors are horizontally concatenated to form a fused query vector. fusion type key vector and fusion-type value vector , see the following formula (8) for details: (8); In equation (8), Represents horizontal splicing; S33, For fused query vectors fusion type key vector and fusion-type value vector Differentiated rotational position encoding is performed, as shown in the following formula (9): (9); In equation (9), for , Represents 2D spatial coordinates. 2D Rotational Position Encoding (2DRoPE) is used to capture the spatial relationships and texture distribution structure of the image; for and , Represents 1D temporal position coordinates. 1D Rotation Position Encoding (1D RoPE) is used to adapt the temporal relationships and logical associations of semantic sequences; This represents the fused query vector after rotational position encoding; This represents the fused key vector after rotational position encoding; S34. Perform multi-head attention calculation, as shown in the following formula (10): (10); In equation (10), Represents the dimension of the key vector. As a scaling factor; Represents the Softmax normalization function; Output features representing multi-head attention; Then, the output features of multi-head attention are decomposed into global visual interaction features. Unified control of semantic interaction features CAU interaction feature sequence during the preview phase ; S35, Incorporate global visual interaction features Unified control of semantic interaction features CAU interaction feature sequence during the preview phase Perform joint conditional coding, as shown in the following formula (11): (11); In equation (11), Represents horizontal splicing; and These represent the weight matrix and bias of the learnable fusion projection, respectively; This represents the joint condition vector.

[0015] The long-time-series video generation network generates the target high-definition video, specifically including the following steps: S41, Joint condition vector The input is fed into the preview video generation network of the long-time sequence video generation network to generate a low frame rate, high speed preview video, as shown in the following formula (12): (12); In equation (12), Represents the network that generates preview videos; This represents the generated preview video; the frame rate of the preview video is [a percentage] of the target high-definition video frame rate. , The downsampling factor; S42. The visibility mask generation module of the long-time-series video generation network divides the target high-definition video into n consecutive refined segments, each refined segment having a length of L frames. For the k-th refined segment, a binary visibility mask is generated. binary visibility mask Each element in Corresponding to one frame of the preview video, This represents the total number of frames in the preview video. , This represents the starting position of the window for the k-th retouched clip in the preview video. The window size representing the binary visibility mask, i.e., the frame number corresponding to the k-th retouched clip in the preview video; The preview condition construction module of the long-time-series video generation network first constructs the preview video. Encoded into a latent representation via a 3D VAE encoder. , This represents the total number of frames in the preview video. It is the number of channels. and Representing the height and width of the potential space respectively; then the binary visibility mask downsampling to After the same length of time and Element-wise multiplication yields the preview conditions after masking. , see the following formula (13) for details: (13); In equation (13), Represents downsampling, This represents element-wise multiplication; S43. Preview conditions after masking the kth retouched segment CAU segmented feature sequences The input is fed into the refinement video generation network of the long-sequence video generation network to generate the k-th refined segment. , see the following formula (14) for details: (14); In equation (14), Represents a network for generating high-quality retouched videos; Finally, the long-time sequence video generation network splices all the generated refined segments in time sequence to obtain the complete target high-definition video. , This represents the total number of frames in the target high-definition video. and These represent the height and width of the target high-definition video, respectively. This represents the number of channels in the target high-definition video.

[0016] The preview video generation network and the refinement video generation network are based on the same RFM generation network, and the generation target of the RFM generation network is a sequence of video frames. That is, preview video Or target high-definition video , Video frame sequence For the image data of the t-th frame, define a continuous time stream of length 1. For any time step... Gradually interpolate the values ​​to the actual data to obtain an intermediate state. , see the following formula (15) for details: (15); In equation (15), Initial noise, The goal of generating the RFM generation network; The RFM generator network is trained using the RFM loss function. See the following formula (16): (16); In equation (16), The expectation operator represents the expectation over time steps. Initial noise and video frame sequence The distribution takes the expected value; Represents the velocity field predicted by the RFM generator network; Represents the joint condition vector; This represents the square of the L2 norm.

[0017] The total loss function used in training the overall model corresponding to the highly controllable long-time-series video generation method is... See the following formula (17) for details: (17); In equation (17), The RFM loss function is represented by equation (16); The consistency loss is calculated by the following formula (18); This represents the L2 regularization term, which penalizes the sum of squares of the learnable parameters in the overall model to prevent overfitting; and These are the balancing weight coefficients for each corresponding loss term; (18); In equation (18), This represents the total number of frames in the preview video or the target high-definition video. This represents the amount of inter-frame feature changes in the preview video or the target high-definition video; This represents a hyperparameter used to balance the scale difference between inter-frame feature variations in the target preview video or target high-definition video and CAU feature sequence variations; Representative CAU feature sequence The amount of inter-frame variation; The features representing the t-th frame of the preview video or the target high-definition video; The features representing the (t-1)th frame of the preview video or the target high-definition video; Represents the L2 norm. The square of the L2 norm; Image data representing the t-th frame of the preview video or the target high-definition video; Represents a visual encoder.

[0018] Advantages of this invention: (1) The present invention constructs a camera action unit (CAU) system, decomposes multimodal camera movement control data into 48 atomic camera action units, and maps them into CAU temporal activation sequences, realizing a unified semantic representation of natural language camera movement style description text, keyframe camera movement parameters and professional camera movement files, and fully preserving all detailed features of multimodal camera movement control data.

[0019] (2) This invention constructs three levels of camera control granularity for three types of multimodal camera control data: natural language camera style description text, keyframe camera parameters, and professional camera files. These granularities are CAU style library calibration, keyframe level, and frame-by-frame precision level, respectively adapting to the needs of rapid creation, flexible adjustment, and precise replication, and greatly expanding the applicability of the technology.

[0020] (3) The present invention constructs a multimodal fusion generation network based on the dual-stream shared attention mechanism to generate a fusion condition vector. Through the dual-stream shared attention mechanism and rotation position coding, the token-level deep fusion of dual-stream features, namely global visual features and semantic control features, is realized, so that the control features can be accurately applied to the corresponding area of ​​the generated video content. This effectively solves the problem of insufficient multimodal data fusion and significantly improves the control accuracy.

[0021] (4) The present invention constructs a long time-series video generation network. Based on the constraint of fusion condition vector, a preview video is generated first, and then the target high-definition video is generated segment by segment. This effectively avoids the accumulation of errors in the long time-series generation process. At the same time, consistency loss is introduced to ensure that the changes between each frame in the generated preview video or target high-definition video conform to the camera motion preset degree. It supports highly controllable long time-series videos up to minutes in length and significantly improves the consistency of long time-series video content. Attached Figure Description

[0022] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] See Figure 1 A highly controllable long-time-series video generation method, specifically including the following steps: (1) Acquiring multimodal input data: Multimodal input data includes the input reference image, character detail control data and multimodal camera movement control data; The character detail control data includes natural language descriptions of character facial expressions and action details, as well as standardized text descriptions based on the character control unit AU. The character control unit AU includes fifteen AU sub-units (see Table 1 below). Each AU sub-unit corresponds to a standardized text description of a specific character facial expression or action detail. The standardized text description based on the character control unit AU is input into one or more AU sub-units in the character control unit AU to realize the standardized text description of character facial expressions and action details. Table 1

[0025] Multimodal camera movement control data includes natural language camera movement style description text, keyframe camera movement parameters, and professional camera movement files; (2) Data processing, which specifically includes the following steps: S21. Convert the reference image into global visual features, specifically: the input reference image is... The resolution of the reference image is First, the reference image is divided into sections of a fixed size. A sequence of non-overlapping image patches, each image patch being of size [size missing]. This partitioning method is aligned with the input structure of the visual encoder, ensuring that each patch corresponds to a visual token. Then, the pre-trained visual encoder is used to extract spatial visual features, which are used to encode each image patch into a feature vector of fixed dimensions, as shown in the following formula (1): (1); In equation (1), Represents a visual encoder. Representing the One image patch, Representing the The visual features corresponding to each image patch, and the visual features corresponding to all image patches, form a spatial visual feature sequence. Then, the spatial visual feature sequence Each of them Dimensionality reduction is performed using linear projection to obtain global visual features. ; S22. Convert the character detail control data into a unified control semantic feature sequence. Specifically, input the natural language description of the character's facial expressions and action details, as well as the standardized text description based on the character control unit AU, into the pre-trained CLIP text encoder. The natural language description of the character's facial expressions and action details is encoded to obtain the initial text embedding features. The standardized text description based on the character control unit (AU) is encoded to obtain an embedding vector. Initial text embedding features The initial text compression embedding features are then obtained through custom pooling. Compress and embed features from the initial text and embedding vector The concatenation yields the text embedding features. Then embed the text features After processing by a lightweight adaptive network, unified control semantic features are obtained. The lightweight adaptable network is a lightweight network containing two fully connected layers. The processing procedure is shown in the following formula (2): (2); In equation (2), and These represent the weight matrix and bias of the first fully connected layer, respectively. and These represent the weight matrix and bias of the second fully connected layer, respectively. Represents the activation function of the Gaussian error linear unit; During training, freezing the backbone parameters of other parts and making only small-scale fine-tuning of the parameters of the lightweight adaptation network can significantly reduce adaptation costs and support rapid personalization to adapt to the different needs of different users. S23. Convert the multimodal camera movement control data into CAU feature sequences. Specifically, a camera action unit (CAU) system is constructed, which includes forty-eight atomic camera action units (see Table 2 below: translation-type atomic camera action units (CAU01-CAU06), rotation-type atomic camera action units (CAU07-CAU12), scaling-type atomic camera action units (CAU13-CAU14), motion curve-type atomic camera action units (CAU15-CAU24), camera movement-type atomic camera action units (CAU25-CAU40), and parametric-type atomic camera action units (CAU41-CAU48)). Each atomic camera action unit corresponds to an ID and a structured natural language description of the camera movement control operation. Then, the multimodal camera movement control data is mapped to CAU temporal activation sequences. , This represents the total number of frames in the target high-definition video to be generated, and finally the CAU temporal activation sequence. After being concatenated with structured natural language (e.g.: After being concatenated with structured natural language, it becomes: translated in the positive X-axis direction, intensity 0.13; ...; image stability coefficient 0.02). After being encoded by a text encoder, the CAU feature sequence is obtained. CAU characteristic sequence After compressing the CAU dimension using custom pooling, the CAU compressed feature sequence is obtained. ; CAU compressed feature sequence In the time dimension Press up to increase speed Sampling was performed to obtain the original preview stage CAU feature sequence. Then, the original preview stage CAU feature sequence Expanding along the time dimension and concatenating them, we obtain the CAU feature sequence for the preview stage. Simultaneously, CAU compresses the feature sequence. The segments are divided into L frames of refined length, then expanded along the time dimension, and spliced ​​together to obtain the CAU segment feature sequence for each segment. ; Table 2

[0026]

[0027]

[0028] Custom pooling specifically involves introducing n learnable query vectors. The input features are aggregated into a fixed-length feature matrix through a cross-attention mechanism, as shown in equation (3) below: (3); In equation (3), when the input features For initial text embedding features At that time, the output fixed-length feature matrix For initial text compression embedding features When input features CAU characteristic sequence At that time, the output is a fixed-length feature matrix. CAU compressed feature sequence ; As a query As keys and values; Represents the cross-attention mechanism; In this process, multimodal camera movement control data is mapped to CAU temporal activation sequences. Specifically: S231. The text describing the camera movement style in natural language is encoded into text features using a pre-trained text encoder. Then, the text features are projected through a projection layer. The mapping is to the CAU time-series activation sequence, as shown in equation (4) below: (4); In equation (4), and These are the weight matrix and bias of the projection layer, respectively. Text features Dimensions Represents the Sigmoid activation function; When the input natural language camera movement style description text is a camera movement style description that already exists in the CAU style library, such as "cinematic handheld camera movement" or "drone aerial photography surround", the corresponding CAU combination template is directly called from the pre-trained CAU style library to automatically generate the corresponding CAU temporal activation sequence. S232. Keyframe camera movement parameters include camera parameters for N keyframes, where each keyframe's camera parameters are mapped to a corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. , see the following formula (4) for details: (4); In equation (4), To normalize the time step, The time of the intermediate frame that needs to be calculated now. This represents the time of the i-th keyframe. This represents the time of the (i+1)th keyframe. These are third-order Bessel basis functions; CAU activation vectors for keyframes; CAU activation vectors for all intermediate frames. Composing a complete frame-by-frame CAU temporal activation sequence ; S233. Analyze the professional camera movement files, extract the camera parameters frame by frame, and map the camera parameters of each frame to the corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. CAU activation vectors of all intermediate frames Composing a complete frame-by-frame CAU temporal activation sequence ; (3) Construct a multimodal fusion generative network based on a two-stream shared attention mechanism to generate fusion condition vectors, which specifically includes the following steps: S31. Unify the control of semantic features and preview stage CAU feature sequences Concatenated into semantic control features Then global visual features and semantic control features The corresponding query, key, and value vectors are calculated, as shown in equations (5) and (6) below: (5); (6); In equations (5) and (6), , and Representing global visual features The query weight matrix, key weight matrix, and value weight matrix; , and Representing global visual features The corresponding query, key, and value vectors; , and Representing semantic control features The query weight matrix, key weight matrix, and value weight matrix; , and Representing semantic control features The corresponding query, key, and value vectors; S32, Global visual features and semantic control features The query, key, and value vectors are horizontally concatenated to form a fused query vector. fusion type key vector and fusion-type value vector , see the following formula (7) for details: (7); In equation (7), Represents horizontal splicing; S33, For fused query vectors fusion type key vector and fusion-type value vector Differentiated rotational position encoding is performed, as shown in the following formula (8): (8); In equation (8), for , Represents 2D spatial coordinates. 2D Rotational Position Encoding (2DRoPE) is used to capture the spatial relationships and texture distribution structure of the image; for and , Represents 1D temporal position coordinates. 1D Rotation Position Encoding (1D RoPE) is used to adapt the temporal relationships and logical associations of semantic sequences; This represents the fused query vector after rotational position encoding; This represents the fused key vector after rotational position encoding; S34. Perform multi-head attention calculation, as shown in the following formula (9): (9); In equation (9), Represents the dimension of the key vector. As a scaling factor; Represents the Softmax normalization function; Output features representing multi-head attention; Then, the output features of multi-head attention are decomposed into global visual interaction features. Unified control of semantic interaction features CAU interaction feature sequence during the preview phase ; The dual-stream shared attention mechanism enables bidirectional fine-grained interaction between each image patch and each semantic token, achieving precise cross-modal alignment; S35, Incorporate global visual interaction features Unified control of semantic interaction features CAU interaction feature sequence during the preview phase Perform joint conditional coding, as shown in the following formula (10): (10); In equation (10), Represents horizontal splicing; and These represent the weight matrix and bias of the learnable fusion projection, respectively; Represents the joint condition vector; (4) Construct a long-time video generation network. Based on the constraints of the fusion conditional vector, first generate a preview video, and then refine and generate the target high-definition video segment by segment. The specific steps include: S41, Joint condition vector The input is fed into the preview video generation network of the long-time sequence video generation network to generate a low frame rate, high speed preview video, as shown in the following formula (11): (11); In equation (11), Represents the network that generates preview videos; This represents the generated preview video; the frame rate of the preview video is [a percentage] of the target high-definition video frame rate. , The downsampling factor; S42. The visibility mask generation module of the long-time-series video generation network divides the target high-definition video into n consecutive refined segments, each refined segment having a length of L frames. For the k-th refined segment, a binary visibility mask is generated. binary visibility mask Each element in Corresponding to one frame of the preview video, This represents the total number of frames in the preview video. , This represents the starting position of the window for the k-th retouched clip in the preview video. The window size representing the binary visibility mask, i.e., the frame number corresponding to the k-th retouched clip in the preview video; The preview condition construction module of the long-time-series video generation network first constructs the preview video. Encoded into a latent representation via a 3D VAE encoder. , This represents the total number of frames in the preview video. It is the number of channels. and Representing the height and width of the potential space respectively; then the binary visibility mask downsampling to After the same length of time and Element-wise multiplication yields the preview conditions after masking. , see the following formula (12) for details: (12); In equation (12), Represents downsampling, This represents element-wise multiplication; S43. Preview conditions after masking the kth retouched segment CAU segmented feature sequences The input is fed into the refinement video generation network of the long-sequence video generation network to generate the k-th refined segment. , see the following formula (13) for details: (13); In equation (13), Represents a network for generating high-quality retouched videos; Finally, the long-time sequence video generation network splices all the generated refined segments in time sequence to obtain the complete target high-definition video. (Default resolution 720×1440) This represents the total number of frames in the target high-definition video. and These represent the height and width of the target high-definition video, respectively. The number of channels representing the target high-definition video; The preview video generation network and the refinement video generation network are based on the same RFM generation network, and the target of the RFM generation network is a sequence of video frames. That is, preview video Or target high-definition video , Video frame sequence For the image data of the t-th frame, define a continuous time stream of length 1. For any time step... Gradually interpolate the values ​​to the actual data to obtain an intermediate state. , see the following formula (14) for details: (14); In equation (14), Initial noise, The goal of generating the RFM generation network; The RFM generation network is trained using the RFM loss function to drive the network to learn the optimal transmission path, thereby improving the clarity and realism of the generated video content. The RFM loss function... See the following formula (15): (15); In equation (15), The expectation operator represents the expectation over time steps. Initial noise and video frame sequence The distribution takes the expected value; Represents the velocity field predicted by the RFM generator network; Represents the joint condition vector; This represents the square of the L2 norm.

[0029] The overall loss function used in training the overall model of the highly controllable long-term video generation method See the following formula (16) for details: (16); In equation (15), The RFM loss function is calculated from equation (15); The consistency loss is calculated by the following formula (17); This represents the L2 regularization term, which penalizes the sum of squares of the learnable parameters in the overall model to prevent overfitting; and These are the balancing weight coefficients for each loss term; (17); In equation (17), This represents the total number of frames in the preview video or the target high-definition video. This represents the amount of inter-frame feature changes in the preview video or the target high-definition video; This represents a hyperparameter used to balance the scale difference between inter-frame feature variations in the target preview video or target high-definition video and CAU feature sequence variations; Representative CAU feature sequence The amount of inter-frame variation; The features representing the t-th frame of the preview video or the target high-definition video; The features representing the (t-1)th frame of the preview video or the target high-definition video; Represents the L2 norm. The square of the L2 norm; Image data representing the t-th frame of the preview video or the target high-definition video; Represents a visual encoder.

[0030] The overall model corresponding to this invention is implemented using the PyTorch deep learning framework and deployed on a server environment equipped with high-performance GPUs. Both the visual encoder and the text encoder use CLIP models pre-trained on large-scale image and text datasets.

[0031] During the training phase, target high-definition videos were generated based on a large-scale video and image dataset, following the steps outlined in this invention, and the AdamW optimizer was used for parameter updates. The initial learning rate was set to 1e-4, with cosine annealing used for learning rate decay. The batch size was set to 8, and the total number of training epochs was set to 100.

[0032] The evaluation metrics for the overall model corresponding to this invention include FID (Fréchet Inception Distance), CLIP similarity, and camera movement control error. For samples in the test set, FID is used to measure the realism of the generated video, CLIP similarity is used to measure the matching degree between the generated content and the control conditions, and camera movement control error is used to measure the fit between the camera movement trajectory of the generated video and the user input.

[0033] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A highly controllable long-time-series video generation method, characterized in that: Specifically, it includes the following steps: (1) Acquiring multimodal input data: Multimodal input data includes the input reference image, character detail control data and multimodal camera movement control data; (2) Data processing: Convert the reference image into global visual features, convert the character detail control data into a unified control semantic feature sequence, and convert the multimodal camera movement control data into a CAU feature sequence; (3) Construct a multimodal fusion generative network based on a dual-stream shared attention mechanism to generate fusion condition vectors; (4) Construct a long-time video generation network. Based on the constraints of the fusion condition vector, first generate a preview video, and then refine and generate the target high-definition video segment by segment.

2. The highly controllable long-time-series video generation method according to claim 1, characterized in that: The character detail control data includes natural language descriptions of character facial expressions and action details, as well as standardized text descriptions based on the character control unit AU. The character control unit AU includes multiple AU sub-units, each AU sub-unit corresponding to a standardized text description of a specific character facial expression or action detail. The standardized text descriptions based on the character control unit AU are input into one or more AU sub-units in the character control unit AU to realize the standardized text descriptions of character facial expressions and action details. The multimodal camera control data includes natural language camera style description text, keyframe camera parameters, and professional camera movement files.

3. The highly controllable long-time-series video generation method according to claim 2, characterized in that: The data processing specifically includes the following steps: S21. Convert the reference image into global visual features, specifically: the input reference image is... The resolution of the reference image is First, the reference image is divided into sections of a fixed size. A sequence of non-overlapping image patches, each image patch being of size [size missing]. This partitioning method is aligned with the input structure of the visual encoder, ensuring that each patch corresponds to a visual token. Then, the pre-trained visual encoder is used to extract spatial visual features, which are used to encode each image patch into a feature vector of fixed dimensions, as shown in the following formula (1): (1); In equation (1), Represents a visual encoder. Representing the One image patch, Representing the The visual features corresponding to each image patch constitute a spatial visual feature sequence. Then, the spatial visual feature sequence Each of them Dimensionality reduction is performed using linear projection to obtain global visual features. ; S22. Convert the character detail control data into a unified control semantic feature sequence. Specifically, input the natural language description of the character's facial expressions and action details, as well as the standardized text description based on the character control unit AU, into the pre-trained CLIP text encoder. The natural language description of the character's facial expressions and action details is encoded to obtain the initial text embedding features. The standardized text description based on the character control unit (AU) is encoded to obtain an embedding vector. Initial text embedding features The initial text compression embedding features are then obtained through custom pooling. Compress and embed features from the initial text and embedding vector The concatenation yields the text embedding features. Then embed the text features After processing by a lightweight adaptive network, unified control semantic features are obtained. The lightweight adaptable network is a lightweight network containing two fully connected layers. The processing procedure is shown in the following formula (2): (2); In equation (2), and These represent the weight matrix and bias of the first fully connected layer, respectively. and These represent the weight matrix and bias of the second fully connected layer, respectively. Represents the activation function of the Gaussian error linear unit; S23. Convert the multimodal camera movement control data into CAU feature sequences. Specifically: First, a camera action unit (CAU) system is constructed, which includes... Each atomic camera action unit corresponds to an ID and a structured natural language description of the camera movement control operation. Then, the multimodal camera movement control data is mapped to a CAU temporal activation sequence. , This represents the total number of frames in the target high-definition video to be generated. This represents the total number of action units in the atomic camera, and finally, the CAU temporal activation sequence is calculated. After being concatenated with structured natural language and encoded by a text encoder, the CAU feature sequence is obtained. CAU characteristic sequence After compressing the CAU dimension using custom pooling, the CAU compressed feature sequence is obtained. , Not less than 2 and less than Integers; The custom pooling specifically involves introducing n learnable query vectors. The input features are aggregated into a fixed-length feature matrix through a cross-attention mechanism, as shown in equation (3) below: (3); In equation (3), when the input features For initial text embedding features At that time, the output fixed-length feature matrix For initial text compression embedding features When input features CAU characteristic sequence At that time, the output fixed-length feature matrix CAU compressed feature sequence ; As a query As keys and values; This represents the cross-attention mechanism.

4. The highly controllable long-time-series video generation method according to claim 3, characterized in that: The process of mapping multimodal camera movement control data into CAU time-series activation sequences. Specifically: S231. The text describing the camera movement style in natural language is encoded into text features using a pre-trained text encoder. Then, the text features are projected through a projection layer. The mapping is to the CAU time-series activation sequence, as shown in equation (4) below: (4); In equation (4), and These are the weight matrix and bias of the projection layer, respectively. Text features Dimensions Represents the Sigmoid activation function; When the input natural language camera movement style description text is a camera movement style description that already exists in the CAU style library, the corresponding CAU combination template is directly called from the pre-trained CAU style library to automatically generate the corresponding CAU temporal activation sequence. S232. Keyframe camera movement parameters include camera parameters for N keyframes, where each keyframe's camera parameters are mapped to a corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. , see the following formula (5) for details: (5); In equation (5), To normalize the time step, The time of the intermediate frame that needs to be calculated now. This represents the time of the i-th keyframe. This represents the time of the (i+1)th keyframe. These are third-order Bessel basis functions; CAU activation vectors for keyframes; CAU activation vectors for all intermediate frames. Composing a complete frame-by-frame CAU temporal activation sequence ; S233. Analyze the professional camera movement files, extract the camera parameters frame by frame, and map the camera parameters of each frame to the corresponding CAU activation vector. Then, a third-order Bezier interpolation algorithm is used to generate smooth intermediate frame CAU activation vectors between adjacent keyframes. CAU activation vectors of all intermediate frames Composing a complete frame-by-frame CAU temporal activation sequence .

5. The highly controllable long-time-series video generation method according to claim 3, characterized in that: The Each atomic camera action unit includes six types of camera action units: translation atomic camera action units, rotation atomic camera action units, scaling atomic camera action units, motion curve atomic camera action units, camera movement atomic camera action units, and parametric atomic camera action units.

6. The highly controllable long-time-series video generation method according to claim 3, characterized in that: The CAU compressed feature sequence In the time dimension Press up to increase speed Sampling was performed to obtain the original preview stage CAU feature sequence. Then, the original preview stage CAU feature sequence Expanding along the time dimension and concatenating them, we obtain the CAU feature sequence for the preview stage. Simultaneously, CAU compresses the feature sequence. The segments are divided into L frames of refined length, then expanded along the time dimension, and spliced ​​together to obtain the CAU segment feature sequence for each segment. .

7. The highly controllable long-time-series video generation method according to claim 6, characterized in that: The construction of a multimodal fusion generative network based on a two-stream shared attention mechanism to generate fusion condition vectors specifically includes the following steps: S31. Unify the control of semantic features and preview stage CAU feature sequences Concatenated into semantic control features Then global visual features and semantic control features The corresponding query, key, and value vectors are calculated, as shown in equations (6) and (7) below: (6); (7); In equations (6) and (7), , and Representing global visual features The query weight matrix, key weight matrix, and value weight matrix; , and Representing global visual features The corresponding query, key, and value vectors; , and They represent semantic control features respectively The query weight matrix, key weight matrix, and value weight matrix; , and They represent semantic control features respectively The corresponding query, key, and value vectors; S32, Global visual features and semantic control features The query, key, and value vectors are horizontally concatenated to form a fused query vector. fusion type key vector and fusion-type value vector , see the following formula (8) for details: (8); In equation (8), Represents horizontal splicing; S33, For fused query vectors fusion type key vector and fusion-type value vector Differentiated rotational position encoding is performed, as shown in the following formula (9): (9); In equation (9), for , Represents 2D spatial coordinates. 2D Rotational Position Encoding (2D RoPE) is used to capture the spatial relationships and texture distribution structure of the image; for and , Represents 1D temporal position coordinates. 1D Rotation Position Encoding (1D RoPE) is used to adapt the temporal relationships and logical associations of semantic sequences; This represents the fused query vector after rotational position encoding; This represents the fused key vector after rotational position encoding; S34. Perform multi-head attention calculation, as shown in the following formula (10): (10); In equation (10), Represents the dimension of the key vector. As a scaling factor; Represents the Softmax normalization function; Output features representing multi-head attention; Then, the output features of multi-head attention are decomposed into global visual interaction features. Unified control of semantic interaction features CAU interaction feature sequence during the preview phase ; S35, Incorporate global visual interaction features Unified control of semantic interaction features CAU interaction feature sequence during the preview phase Perform joint conditional coding, as shown in the following formula (11): (11); In equation (11), Represents horizontal splicing; and These represent the weight matrix and bias of the learnable fusion projection, respectively; This represents the joint condition vector.

8. The highly controllable long-time-series video generation method according to claim 7, characterized in that: The long-time-series video generation network generates the target high-definition video, specifically including the following steps: S41, Joint condition vector The input is fed into the preview video generation network of the long-time sequence video generation network to generate a low frame rate, high speed preview video, as shown in the following formula (12): (12); In equation (12), Represents the network that generates preview videos; This represents the generated preview video; the frame rate of the preview video is [a percentage] of the target high-definition video frame rate. , The downsampling factor; S42. The visibility mask generation module of the long-time-series video generation network divides the target high-definition video into n consecutive refined segments, each refined segment having a length of L frames. For the k-th refined segment, a binary visibility mask is generated. binary visibility mask Each element in Corresponding to one frame of the preview video, This represents the total number of frames in the preview video. , This represents the starting position of the window for the k-th retouched clip in the preview video. The window size representing the binary visibility mask, i.e., the frame number corresponding to the k-th retouched clip in the preview video; The preview condition construction module of the long-sequence video generation network first constructs the preview video. Encoded into a latent representation via a 3D VAE encoder. , This represents the total number of frames in the preview video. It is the number of channels. and Representing the height and width of the potential space respectively; then the binary visibility mask downsampling to After the same length of time and Element-wise multiplication yields the preview conditions after masking. , see the following formula (13) for details: (13); In equation (13), Represents downsampling, This represents element-wise multiplication; S43. Preview conditions after masking the kth retouched segment CAU segmented feature sequences The input is fed into the refinement video generation network of the long-sequence video generation network to generate the k-th refined segment. , see the following formula (14) for details: (14); In equation (14), Represents a network for generating high-quality retouched videos; Finally, the long-time sequence video generation network splices all the generated refined segments in time sequence to obtain the complete target high-definition video. , This represents the total number of frames in the target high-definition video. and These represent the height and width of the target high-definition video, respectively. This represents the number of channels in the target high-definition video.

9. The method for generating highly controllable long-term time-series video according to claim 8, characterized in that: The preview video generation network and the refinement video generation network are based on the same RFM generation network, and the generation target of the RFM generation network is a sequence of video frames. That is, preview video Or target high-definition video , Video frame sequence For the image data of the t-th frame, define a continuous time stream of length 1. For any time step... Gradually interpolate the values ​​to the actual data to obtain an intermediate state. , see the following formula (15) for details: (15); In equation (15), Initial noise, The goal of generating the RFM generation network; The RFM generator network is trained using the RFM loss function. See the following formula (16): (16); In equation (16), The expectation operator represents the expectation over time steps. Initial noise and video frame sequence The distribution takes the expected value; Represents the velocity field predicted by the RFM generator network; Represents the joint condition vector; This represents the square of the L2 norm.

10. A highly controllable long-time-series video generation method according to claim 9, characterized in that: The total loss function used in training the overall model corresponding to the highly controllable long-time-series video generation method is... See the following formula (17) for details: (17); In equation (17), The RFM loss function is represented by equation (16); The consistency loss is calculated by the following formula (18); This represents the L2 regularization term, which penalizes the sum of squares of the learnable parameters in the overall model to prevent overfitting; and These are the balancing weight coefficients for each corresponding loss term; (18); In equation (18), This represents the total number of frames in the preview video or the target high-definition video. This represents the amount of inter-frame feature change in the preview video or the target high-definition video; This represents a hyperparameter used to balance the scale difference between inter-frame feature variations in the target preview video or target high-definition video and CAU feature sequence variations; Representative CAU feature sequence The amount of inter-frame variation; The features representing the t-th frame of the preview video or the target high-definition video; The features representing the (t-1)th frame of the preview video or the target high-definition video; Represents the L2 norm. The square of the L2 norm; Image data representing the t-th frame of the preview video or the target high-definition video; Represents a visual encoder.