Video generation method and system based on visual hierarchical autoregression

Through the video generation method of visual hierarchical autoregression, video generation is generated using text encoding and hierarchical autoregression models, the problem of high computational cost of diffusion model is solved, efficient and flexible video generation is achieved, and the generation quality and fluency are ensured.

CN120390126APending Publication Date: 2025-07-29TERMINUS GENERAL TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510581427.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing diffusion model has serious problems in computing power consumption and time-consuming in video generation, resulting in high computing costs and slow generation speed, making it difficult to meet real-time needs.

Method used

The video generation method based on visual hierarchical autoregression is adopted. The description text data is converted into text token sequences through text encoder, and the hierarchical autoregression model is input to generate video tokens sequences of each scale, and high-quality video data is generated through feature encoding and decoder. The multi-scale pyramid structure and VQ-VAE method are combined to extract spatiotemporal features and optimize the generation process.

Benefits of technology

It improves the efficiency and quality of video generation, reduces manual intervention, realizes real-time monitoring and early warning, and can dynamically adjust the generation quality based on computing resources to ensure the smoothness and natural transition of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390126A_ABST
    Figure CN120390126A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and system based on visual hierarchical autoregression. The method is applied to the technical field of computer vision, and comprises the following steps: acquiring description text data, and encoding the description text data according to a preset text encoder to obtain a text token sequence of the description text data; inputting the text token sequence into a preset hierarchical autoregression model to obtain video token sequences of various scales; determining feature coding data according to the video token sequence of each scale and a preset coding mode; and outputting the feature coding data to a decoder to obtain video data. The scheme can improve the production efficiency and reduce manual intervention. Customized video contents are generated according to different requirements, and the generation quality is dynamically adjusted to adapt to different computing resources. The layered autoregression model and the feature coding technology ensure the fluency and natural transition of video generation, optimize resource use and ensure high-quality output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a video generation method and system based on visual hierarchical autoregression. Background Art

[0002] With the rapid development of generative artificial intelligence technology, video generation, as a core direction of multimodal content generation, has shown great application potential in fields such as film and television entertainment, digital advertising, and virtual creation. Different from text or image generation, video generation needs to meet the requirements of spatio-temporal continuity, physical consistency, and motion coherence at the same time, and its technical complexity increases exponentially. For example, to generate a 3-second physically realistic short film, it is necessary to process the dynamic correlations, object motion trajectories, and light and shadow interactions between thousands of frames, which poses a double challenge to the spatio-temporal modeling ability and computational efficiency of the model.

[0003] Today's video generation methods use diffusion models for generation. Represented by the DiT architecture, videos are generated through an iterative process of gradually denoising. The model starts from a random noise state, predicts and eliminates the previous noise at each step, and finally converges to a clear result.

[0004] However, although diffusion models have high generation quality, the problems of high computing power consumption and time consumption restrict their practical applications, and multiple-step iterations lead to high computational costs and slow generation speeds. Summary of the Invention

[0005] To solve the deficiencies of the prior art, the present disclosure provides a video generation method and system based on visual hierarchical autoregression. The present disclosure solves the problems that although today's diffusion models have high generation quality, the problems of high computing power consumption and time consumption restrict their practical applications, and multiple-step iterations lead to high computational costs and slow generation speeds.

[0006] According to a first aspect of the present disclosure, there is provided a video generation method based on visual hierarchical autoregression, including: obtaining descriptive text data, encoding the descriptive text data according to a preset text encoder to obtain a text token sequence of the descriptive text data;

[0007] Inputting the text token sequence into a preset hierarchical autoregression model to obtain video token sequences at each scale;

[0008] Determining feature encoding data according to the video token sequences at each scale and a preset encoding method;

[0009] Outputting the feature encoding data to a decoder to obtain video data.

[0010] According to a second aspect of the present disclosure, there is provided a video generation system based on visual hierarchical autoregression for performing the method as described in the first aspect, including: a text encoding module for obtaining descriptive text data, encoding the descriptive text data according to a preset text encoder to obtain a text token sequence of the descriptive text data;

[0011] a hierarchical autoregressive generation module for inputting the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale;

[0012] a feature encoding module for determining feature encoding data according to the video token sequences at each scale and a preset encoding method;

[0013] a decoding module for outputting the feature encoding data to a decoder to obtain video data.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, which includes: a memory and a processor, where a computer program is stored on the memory, and the processor implements the method as described above when executing the program.

[0015] In a video generation method and system based on visual hierarchical autoregression provided as above, the embodiments of the present disclosure can achieve real-time monitoring and early warning, timely discover problems and respond quickly. Through multi-dimensional evaluation, the score of each dimension can be used as a basis for improvement, which helps the workshop to formulate digital improvement strategies. The scoring model automatically calculates the maturity score, reduces manual intervention, and improves the evaluation efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 shows a schematic flowchart of a video generation method based on visual hierarchical autoregression according to an embodiment of the present disclosure;

[0018] Figure 2 shows a schematic flowchart of a video generation method based on visual hierarchical autoregression according to an embodiment of the present disclosure;

[0019] Figure 3 shows a schematic block diagram of a video generation system based on visual hierarchical autoregression according to an embodiment of the present disclosure;

[0020] Figure 4 The block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0022] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them. It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more. It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it can generally be understood as one or more. In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after. It should also be understood that the present disclosure emphasizes the differences between various embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be described in detail one by one.

[0023] At the same time, it should be understood that for the sake of description, the sizes of the various parts shown in the drawings are not drawn in actual proportional relationships. The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present disclosure and its application or use. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification. It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0025] Figure 1 Schematic diagram of a video generation method based on visual hierarchical autoregression provided by an embodiment of the present disclosure. The method of the embodiment of the present disclosure aims to achieve precise detection of large and small targets in pictures.

[0026] S101, obtain descriptive text data, and encode the descriptive text data according to a preset text encoder to obtain a text token sequence of the descriptive text data.

[0027] The descriptive text data may refer to a literal description of the video content, which is content expressed in natural language. It may include the subject: the main elements shown in the video, such as people, objects, scenes, etc. Background: the environment or background setting of the video, such as a city, natural scenery, indoor scene, etc. Shot: the perspective, shooting method, etc. of the video, which may involve perspective changes, camera movement, shot switching, etc. Style: the visual style of the video, such as an artistic style, color tone, lighting effect, etc. Color tone: the color temperature, saturation, etc. of the video, which may affect the overall atmosphere of the video.

[0028] A text encoder may be a model or method that converts natural language text into a numerical representation that can be processed by a computer. Generally, the preset text encoder refers to a trained natural language processing (NLP) model that can convert text input (such as text describing a video) into a feature vector or token sequence. Common preset text encoders include: T5 (Text-to-Text Transfer Transformer), BERT (Bidirectional Encoder Representations from Transformers).

[0029] In NLP, a token is the basic unit of text, usually a word or subword. A text token sequence refers to a sequence obtained by splitting the descriptive text data into words or subwords. For example, the sentence "In a sunny forest" can be decomposed into multiple tokens: ["In", "a", "sunny", "forest", "in"]. By using a preset text encoder, the input text is disassembled into tokens, and each token is converted into a numerical vector to form a token sequence. This token sequence can be used as the basic input for generating video content.

[0030] The descriptive text can be written manually or generated by other systems or tools. Specifically, it can include manual writing: writing the descriptive text manually, for example: "In a sunny forest, a deer is walking slowly, and the trees and grass are swaying with the gentle breeze." Automatic generation: If there is existing video or picture content, image annotation or video analysis models can be used to automatically generate the descriptive text. For example, an image recognition model can be used to analyze the scenes in the video and automatically generate a natural language description. Text generation: In some applications, the descriptive text can be automatically generated based on context information. For example, according to the user's selection or input, the system can automatically generate a related video description. The obtained descriptive text data is used as input and input into a preset text encoder for encoding. An open-source library (such as Transformers from HuggingFace) can be used to load the preset text encoder, input the descriptive text, and obtain a token sequence. The text encoder converts the input descriptive text into a series of tokens, which are numerical representations that the model can understand and process. For example, in T5, the token sequence is a tensor containing multiple integers, and each integer corresponds to a token in the descriptive text.

[0031] The training process of the preset text encoder includes:

[0032] When training a text encoder, it is first necessary to collect and preprocess a large amount of text data. Common corpora include: Wikipedia: Provides a large amount of encyclopedic knowledge. BooksCorpus: Includes texts from various books. CommonCrawl: Contains a large amount of web page data on the Internet. WebText: Includes web page content obtained from Reddit. This data needs to be cleaned and processed to remove noise (such as HTML tags, advertisements, duplicate content, etc.) and standardize the text. The goal of pre-training is to let the model learn the basic features of the language, such as vocabulary, syntax, context, etc. The pre-training objective can be sequence-to-sequence learning (used in T5, BART, etc.): T5 is a typical sequence-to-sequence (seq2seq) model that uses a text generation objective and attempts to convert an input text into another output text. Text generation objective: T5 converts the task into a text-to-text form. For example, "Translate: I like you" -> "Translation: I love you". Through such input and output, the model can learn how to generate appropriate outputs in different tasks. The pre-training process usually uses a large amount of computing resources to train the model through GPUs or TPUs. The steps of the training process are that the model parameters (such as weights) are randomly initialized or initialized using a certain preset rule. Then, the text data is input into the model for training. The input text (such as a sentence, paragraph, or document) is input into the model, and the model performs calculations through multiple Transformer layers to generate prediction results. For each prediction result, calculate the gap with the true label (i.e., the actual next word or the masked word). The commonly used loss function is CrossEntropyLoss. Calculate the gradients through the backpropagation algorithm and update the weights of the model using an optimizer (such as the Adam optimizer). The goal of backpropagation is to minimize the loss function so that the model can accurately predict the next word or the structure of a sentence. This process is repeated millions of times (usually taking several weeks) to ensure that the model can learn the statistical characteristics and deep semantic understanding of the language from a large amount of data. After pre-training is completed, the model is fine-tuned on specific tasks so that the model can be better applied to actual tasks. For example, in tasks such as text classification, sentiment analysis, and question-answering systems, the model will be fine-tuned according to the labels of the tasks. For example, in a text generation task, the model will generate the corresponding output text based on the given input text. During fine-tuning, a relatively small learning rate is usually used and training is performed on a dataset in a specific domain. During the training process, it is necessary to evaluate the performance of the model to ensure that it can also perform well on tasks outside the training data. Common evaluation metrics include accuracy, perplexity: commonly used for language models, and the lower the perplexity, the better the model.BLEU (for machine translation tasks), ROUGE (for text summarization tasks).

[0033] S102, input the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale.

[0034] The preset hierarchical autoregressive model can be an autoregressive-based model, based on the standard decoder-only Transformer architecture, which can gradually generate data or encode data. This model gradually generates representations of multiple scales (e.g., frame level, sequence level, etc.) of video data based on the input text.

[0035] Video token sequences can represent the content of each frame of a video or a certain video segment as discrete tokens. They are usually mapped to a fixed-dimensional space through an embedding process and passed to downstream models (e.g., decoders or generators) in sequence. For example, video token sequences can include visual features of each frame (e.g., the encoded representation of an image), actions or events occurring in the video (e.g., "jumping" or "running"), and scene information of the video (e.g., "beach" or "city street").

[0036] The generated text token sequence can be input into a hierarchical autoregressive model for processing. The autoregressive model will gradually generate video token sequences at different scales. The model will first generate rough features of the video (such as background, main scene) at a lower level, and then refine the content at a higher level (such as specific actions, object recognition, etc.). This step is completed through a hierarchical recursive generation process. Level 1: The model may generate an overview of the entire video, such as the background or scene. Level 2: The model then generates actions or main objects in the video (e.g., dog, running, etc.). Level 3: The model may further generate details, such as the detailed actions or positions of the dog. During the generation process of each layer of the hierarchical autoregressive model, it conditions on the results generated by the previous layer. For example, after generating a rough video feature at a low resolution, the model generates more refined content based on this feature. Suppose the goal is to generate a continuous video sequence with actions: The first layer generates general scene information and may output some rough tokens, such as "dog", "park". The second layer generates specific behaviors (e.g., "running") and determines the details of these actions based on the tokens generated by the first layer. The third layer refines the details of each action and may generate more fine-grained tokens, such as the running posture of the dog. Finally, the hierarchical autoregressive model will output video token sequences at multiple scales. For example: The tokens output by the first layer may be low-level representations (such as background) describing the scene. The tokens output by the second layer may be mid-level representations for a specific action (such as running, jumping). The tokens output by the third layer may be specific frame-level details (such as the specific actions of the dog).

[0037] The training steps of the preset hierarchical autoregressive model are as follows:

[0038] 1. Prepare "video-text" data pairs, where the text is a description of the video content;

[0039] 2. Extract frames from the video data of each pair of samples, and divide every 4 consecutive frames into a group;

[0040] 3. Use the trained "multi-scale video encoder" to encode each group of videos to obtain the token sequence of each group of video frames; these tokens are used as data labels for self-supervised training with the video tokens generated by the model;

[0041] 4. Use a trained text encoder (T5 encoder, and the parameters of the text encoder remain unchanged during training) to encode the text data to obtain the feature embedding of the text;

[0042] 5. The model starts predicting from only text input and first predicts to generate the lowest resolution video tokens;

[0043] 6. Then continue to recursively predict the video tokens at the next resolution level;

[0044] 7. Compare the prediction results at each resolution with the ground truth tokens and use the cross - entropy loss function;

[0045] 8. When a set of video predictions for the video is completed, the technique proceeds to train the next set of video frames. The process is the same as steps 5 - 7, except that the input to the model includes: text encoding, and the video encodings of other previously generated sets;

[0046] Continue this process until all the video data in the sample is trained, and then proceed to train the next "video - text" pair of data.

[0047] Based on the above technical solution, optionally, input the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale, including:

[0048] Input the text token sequence into a preset hierarchical autoregressive model. The preset hierarchical autoregressive model gradually generates video token sequences at each scale and generates the video token sequence at the next scale based on the currently generated video token sequence until a preset termination condition is reached, thereby obtaining video token sequences at each scale.

[0049] In this solution, the preset termination condition can refer to when the model needs to determine when to stop the generation process during the generation of the video token sequence. It can include generating a complete token sequence: that is, at each scale level, the model has generated a sufficient number of tokens, or the generated tokens reach a preset number of frames or video length. Reaching the maximum number of iterations: Set a maximum number of generations or a maximum generation time to ensure that the model does not generate tokens indefinitely and prevent the generation process from getting out of control. The generated quality meets the standard: For example, use a quality metric (such as the smoothness of the generated video, content consistency, etc.) to determine whether the generation has met the predetermined quality standard. If the generated result is good enough, the model stops generating. The loss function of the model converges: By calculating the value of the loss function (such as temporal consistency loss, quality loss, etc.), if the loss no longer decreases significantly, it indicates that the generation is completed and meets the requirements, and the generation can be terminated.

[0050] The text description can be converted into a token sequence (such as the T5 model) through a preset text encoder, which provides preliminary semantic information for video generation. The hierarchical autoregressive model will gradually generate the token sequence of the video according to the input text token sequence. First, low-resolution video tokens are generated, and then the resolution is gradually increased to generate more detailed and higher-quality tokens. During the generation process of each layer of resolution, the model needs to check the preset termination conditions when generating video tokens at each scale. If the generated video tokens reach the set number of frames, or the generated tokens are sufficient to meet the duration or quality requirements of the video, the generation process can be terminated. If the maximum number of iterations is reached, the model will also stop. Alternatively, when the model detects that the loss function (such as quality loss, temporal consistency loss, etc.) converges, the generation process automatically ends. If the termination conditions are not met, the model will continue to generate the video token sequence of the next scale until the video generation is finally completed. Once the termination conditions are met, the video token sequences of all scales will be output and can be handed over to the subsequent decoder for decoding to generate the final video.

[0051] In this solution, by setting reasonable termination conditions, it is possible to ensure that the generation process is both efficient and meets the predetermined quality and duration requirements, avoiding excessive calculations and improving the fluency and consistency of generation.

[0052] S103, determine the feature encoding data according to the video token sequences of each scale and the preset encoding method.

[0053] A preset encoding method usually refers to a set of predefined rules or methods in a model for converting data (such as a video tokens sequence) into a feature representation. Specifically, these encoding methods define how to map discrete tokens (such as video features per frame) into a high-dimensional space for subsequent processing and generation. The design of the encoding method depends on the goals and task requirements of video generation. Common encoding methods include: Embedding encoding: This is the most common method, which converts discrete tokens (such as words, frame-level features) into a continuous vector. Through word embeddings (such as Word2Vec, GloVe) or positional embeddings, etc., the tokens will be mapped into a high-dimensional vector space. Each token corresponds to a dense vector representation, which contains the semantic information of the token. Convolutional Neural Network (CNN) encoding: For image or video data, convolutional neural networks are usually used to extract features. Video frames are converted into a set of feature maps, and embedding vectors are obtained after convolutional processing. These feature maps or feature embeddings represent the spatial information of the image. Positional Encoding: In sequence data processing, positional encoding is used to provide the order information of each element in the sequence. In video generation, positional encoding needs to be added to the tokens of each frame or each action to maintain temporal and spatial consistency. Multiscale Encoding: For video generation, multiscale encoding processes features at different levels or resolutions simultaneously to ensure that the model can handle information at different granularities (such as rough background, detailed actions, etc.).

[0054] Feature-encoded data can refer to the high-dimensional feature vectors or tensors mapped from the original video tokens sequence through a preset encoding method. These features represent the key information of the video content and convert the spatio-temporal information of the video (such as pixel information per frame, action details, etc.) into a format suitable for calculation and processing. Specifically, the feature-encoded data can include the feature vectors of video frames: After being processed by the encoder, the video information of each frame is converted into a feature vector representation, which contains the spatial structure, content, and actions of the frame. Temporal features: Encoding in the time dimension to ensure the temporal consistency of the video. The feature-encoded data may also include the relationship information between frames to ensure the smoothness and coherence of the video. Multiscale features: Feature representations at different levels (or scales), which may include coarse-grained and fine-grained information. For example, the initial description of the video content may be represented by coarse-grained features (such as the background), while the subsequent generated details may be represented by more refined features (such as actions, specific manifestations of objects).

[0055] The video token sequence of each frame can be converted into a corresponding feature representation through a preset encoding method (such as embedding encoding, CNN encoding, etc.). Perform positional encoding on each token sequence: If chronological information is required, add positional encoding to ensure that the model understands the order of video frames. Embedding representation: Each token is mapped to a high-dimensional vector space through pre-trained word embeddings (such as Word2Vec) to represent the semantics of the token. Convolutional encoding: If it is video frame data, spatial features may be extracted through a convolutional neural network to further enrich the representation of video frames. Through the above steps, the video tokens of each frame are converted into high-dimensional feature representations. For the token sequences generated by each layer, they can be encoded separately, and the encoding results of different layers can be combined. Combining multi-scale features, positional encoding, action encoding, etc., the finally generated feature encoding data will contain information such as temporal consistency, spatial structure, and action. Organize all the processed feature encoding data into a high-dimensional tensor or vector, and these feature encoding data will be used as the input of the model for subsequent decoding, generation, or further optimization.

[0056] Based on the above technical solution, optionally, determining the feature encoding data according to the video token sequences of each scale and the preset encoding method includes:

[0057] Convert the video token sequences of each scale into multi-frame image data, and perform 3D convolution operations on the multi-frame image data based on the VQ-VAE method to obtain spatio-temporal feature representations;

[0058] Encode the spatio-temporal feature representations according to the multi-scale pyramid structure to obtain the feature encoding data.

[0059] In this solution, the multi-frame image data can be extracted from the generated video token sequences and represent different frames of the video. Each frame is usually an image containing the visual information of the video content. By decoding the video tokens of each frame into image data, multi-frame image data can be obtained. Specifically, the composition of the image data: Each frame of image data is usually composed of pixel values (RGB values), representing information such as the color, brightness, shape, and texture of the image. Multi-frame image data means image data at multiple time points generated from the video, and these images together form different time slices of the video content.

[0060] Spatio-temporal feature representation can refer to the feature representation in a video that contains both spatial information (the spatial structure of the image content, such as objects, people, scenes, etc.) and temporal information (the motion, changes, etc. between video frames). A video not only contains image content but also undergoes dynamic changes over time. Therefore, spatio-temporal feature representation combines the information of both: Spatial features: Reflect the content of a single frame image, such as the position, shape, color of an object, etc. Temporal features: Reflect the dynamic changes between video frames, capturing information such as motion and transitions between different frames.

[0061] VQ-VAE (Vector Quantized Variational Autoencoder) is a quantization-based variational autoencoder model, aiming to learn the latent space representation of data by mapping the input data to discrete vector representations (i.e., "quantization"). The core idea of VQ-VAE is that by discretizing (quantizing) the input space, the model can learn more structured latent representations. Different from the traditional variational autoencoder (VAE), VQ-VAE uses discrete vectors in the latent space instead of continuous representations. For multi-frame image data, VQ-VAE can use a 3D convolutional network to capture spatio-temporal features. In this process, the spatio-temporal information of the video is extracted to generate high-quality spatio-temporal feature representations.

[0062] The multi-scale pyramid structure can be a structure for feature extraction at multiple scales, usually used to capture features at different levels in image or video data. It helps the model learn local and global information simultaneously at multiple scales by processing data at different resolutions. The multi-scale structure usually includes multiple levels from low resolution to high resolution, which can capture detailed and global information respectively. By learning features at different scales, the model can understand the image structure at different levels. In image processing, the pyramid structure refers to multiple successively reduced images from the original image to a low-resolution version. Each layer is usually a downsampling of the previous layer, gradually extracting more abstract features. When encoding spatio-temporal features, using the multi-scale pyramid structure can effectively capture different levels of information from details to the global. Different scales can reflect different levels of changes in the video. The low scale captures the general dynamic changes and object positions, while the high scale captures more detailed local details.

[0063] The tokens generated from the video can be decoded into image frames to obtain multi-frame image data in the video. These images will serve as the input for subsequent processing by the model. Apply the VQ-VAE method to the multi-frame image data. Through the encoder of VQ-VAE, map the image data to a discrete latent space representation. Use 3D convolutional operations to capture spatio-temporal features. 3D convolution can consider both the spatial and temporal dimensions between video frames simultaneously, thereby extracting spatio-temporal feature representations. Pass the spatio-temporal feature representation into a multi-scale pyramid structure, and the model processes the spatio-temporal features at different resolutions. Each scale level learns spatio-temporal information at different levels, enabling the model to have a good understanding ability in both details and overall. Through the encoding of the multi-scale pyramid structure, finally obtain feature encoding data. These feature encoding data include the spatio-temporal information in the video and will be used for subsequent generation or decoding steps. For example, based on the VQ-VAE method, after performing 3D convolutional operations on multi-frame images, then perform encoding. During the encoding process, use a multi-scale pyramid structure to sequentially encode and generate feature encodings of three scales: 1x1, 3x3, and 5x5.

[0064] In this solution, through 3D convolutional operations, it is possible to consider both the spatial information (image content of each frame) and temporal information (dynamic changes between frames) of the video simultaneously. This processing method helps to capture the spatio-temporal relationship in the video, avoids ignoring the temporal changes between frames, and ensures the coherence and naturalness of video generation. VQ-VAE discretizes the image representation through a quantization method, which can more efficiently extract the latent features of the image, especially when dealing with large-scale video data. This feature learning enables the model to generate more diverse and high-quality video content. The multi-scale pyramid structure can extract features from different resolutions and scales, capturing both local details (such as textures, object contours) and understanding the global structure (such as scenes, backgrounds). This multi-level feature extraction helps the model obtain more comprehensive and detailed spatio-temporal information.

[0065] S104, output the feature encoding data to a decoder to obtain video data.

[0066] The decoder can be a neural network module or algorithm used to convert the input feature-encoded data (such as feature vectors or feature tensors) into visual outputs (such as videos, images, text, etc.). In video generation tasks, the role of the decoder is to convert the feature-encoded data obtained from the video tokens sequence back into the original video frame data or generate new and expected video content. Technologies commonly adopted by decoders include: Transposed convolutional neural network: Restores the low-dimensional feature map to the high-dimensional image space through transposed convolution (upsampling) to generate video frames. Autoregressive generative model: For example, uses LSTM, GRU, or Transformer models to gradually generate frame sequences, maintaining temporal and spatial coherence. Generative adversarial network: Utilizes adversarial training between the generator and the discriminator to generate realistic video frame data. Variational autoencoder: Generates realistic video frames through the reparameterization trick while being able to control the diversity of the generated content.

[0067] Video data can refer to the image data of consecutive frames output by the decoder, usually a three-dimensional tensor or a video file. Each frame of the video data can be regarded as a two-dimensional image, which is composed of the pixel values of each frame. These video frames can be converted into standard video formats (such as MP4, AVI, etc.) for playback or used for subsequent analysis and processing. The composition of video data is as follows: Frame data: Each video frame is a two-dimensional matrix representing the pixels of the image. Time axis: The video consists of consecutive frames, and the time information is determined by the order of each frame.

[0068] The feature encoding data is obtained from the video tokens sequence through a preset encoding method and contains the spatio-temporal information and structural features of the video. After passing through the encoding stage of the model, this data is transmitted to the decoder. The decoder receives the feature encoding data from the encoder. According to the model architecture, the decoder generates the pixel data of the video frames based on these features. Specifically, deconvolution network (or upsampling): The decoder may use deconvolution operations to gradually map the low-dimensional feature vectors back to a higher-dimensional space to generate the final video frames. Autoregressive generation process: If an autoregressive model (such as LSTM, GRU) is used, the decoder generates the video data frame by frame, depending on the content generated in the previous frame. GAN decoder: The decoder of the generative adversarial network generates images from the feature encoding and continuously optimizes the generation quality through the adversarial process. The decoder generates a series of video frames based on the feature encoding data. These video frames are gradually generated and combined through the decoding process to finally form a complete video. During the generation of each frame, the decoder uses the previous frame and the current features to generate a new frame to maintain the spatio-temporal consistency of the video. The generated video data can exist in the form of an image sequence or be packed through video encoding methods (such as encoded into the MP4 format, etc.). The final output is a complete video data, represented as a series of consecutive frames with a reasonable spatio-temporal structure. This video data can be saved as a video file or continue to be used for subsequent processing and analysis.

[0069] In the embodiments of this application, descriptive text data is obtained, and the descriptive text data is encoded according to a preset text encoder to obtain the text token sequence of the descriptive text data; the text token sequence is input into a preset hierarchical autoregressive model to obtain video tokens sequences of each scale; feature encoding data is determined according to the video tokens sequences of each scale and a preset encoding method; the feature encoding data is output to the decoder to obtain video data. Through the above video generation method based on visual hierarchical autoregression, through the automated generation process of converting text descriptions into video content, production efficiency can be improved and manual intervention can be reduced. It is not only flexible and efficient, capable of generating customized video content according to different requirements, but also can dynamically adjust the generation quality to adapt to different computing resources. At the same time, the hierarchical autoregressive model and feature encoding technology ensure the smoothness and natural transition of video generation, optimize resource utilization, and guarantee high-quality output.

[0070] Figure 2 It is a schematic flowchart of the video generation method based on visual hierarchical autoregression provided by the embodiments of this disclosure. The method may include the following steps:

[0071] S201. Obtain the descriptive text data, and encode the descriptive text data according to a preset text encoder to obtain the text token sequence of the descriptive text data.

[0072] S202. Input the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences of each scale.

[0073] S203. Determine the feature encoding data according to the video token sequences of each scale and a preset encoding method.

[0074] S204. Output the feature encoding data to a decoder to obtain video data.

[0075] S205. Determine the pixel information of each frame of the video data and the current video frame rate of the video data, obtain the current available computing resources, the preset maximum available computing resources, and the preset maximum frame rate. Calculate the video generation quality loss according to the pixel information, the current video frame rate, the current available computing resources, the preset maximum available computing resources, the preset maximum frame rate, and a preset quality loss calculation formula.

[0076] The pixel information can be the color and brightness values of each pixel in each frame of the video. Each frame of the video can be regarded as a two-dimensional matrix, and each element in the matrix represents the color and brightness information of a pixel point. For example, the pixel information of a color video frame usually includes the values of the RGB three channels, and each channel varies between 0 and 255. For the video generation task, the pixel information is the basis for calculating the video quality because it directly determines the image quality of each frame.

[0077] The frame rate can refer to the number of frames displayed per second, and the unit is usually frames per second (FPS). The frame rate determines the smoothness of the video. A higher frame rate usually means smoother video playback. Common video frame rates are 24FPS, 30FPS, 60FPS, etc.

[0078] The current available computing resources can be the hardware resources currently available in the system during the video generation process, including CPU, GPU, memory, hard disk, etc. The limitation of computing resources will directly affect the speed and quality of video generation, especially when generating high-resolution and high-frame-rate videos.

[0079] The preset maximum available computing resources can refer to the maximum value of the computing resources that the video generation system can use under the most ideal conditions. This usually refers to the maximum performance potential of the system hardware, such as the maximum GPU memory, the maximum CPU frequency, etc. The performance of the system is usually limited by these preset maximum values.

[0080] The preset maximum frame rate can refer to the maximum frame rate that a video generation system can support under ideal conditions. This value depends on the computing resources and the resolution of the video. With sufficient computing resources, the system can maintain a higher frame rate.

[0081] The video generation quality loss can be measured by calculating the gap between the quality of the generated video and the ideal quality. When calculating the video quality, factors such as the resolution, frame rate, pixel information, and computing resources of the video are usually considered. The greater the quality loss, the worse the quality of the generated video. The video generation quality loss is usually used to dynamically adjust the generation parameters (such as frame rate, resolution, etc.) to optimize the generation quality.

[0082] Obtaining the pixel information of each frame of the video data can be extracted from the video file through an image processing algorithm. Then, determine the frame rate of the video, that is, how many frames are displayed per second. The frame rate of the video is usually provided by video encoding or configuration parameters. Obtain the currently available computing resources. The current state of the computing resources can be obtained by monitoring the usage of hardware (such as GPU, CPU, memory). Obtain the preset maximum computing resources and preset maximum frame rate of the system. This information is usually provided by the hardware configuration or preset system parameters. Then, substitute the above data into the preset quality loss calculation formula to obtain the video generation quality loss.

[0083] In this embodiment, various parameters of video generation can be dynamically adjusted according to the actual computing resources of the system, so as to ensure the quality and smoothness of the video.

[0084] Based on the above technical solution, optionally, after calculating the video generation quality loss, the method further includes:

[0085] If the video generation quality loss exceeds the preset quality loss threshold, input the video generation quality loss, the current video frame rate, the currently available computing resources, the preset maximum available computing resources, and the preset maximum frame rate into a preset video optimization model to obtain a video optimization plan, and optimize the video data according to the video optimization plan.

[0086] In this solution, the preset quality loss threshold can be a standard set by the system for judging whether the video generation quality reaches an acceptable range. If the calculated video generation quality loss exceeds this threshold, the system will take optimization measures (such as adjusting the frame rate, resolution, or computing resources, etc.) to improve the quality of the generated video. The threshold is usually obtained through empirical or historical data analysis and reflects the maximum quality loss that the system can tolerate.

[0087] The preset video optimization model can be a machine learning or optimization model used to predict and optimize various aspects of video generation based on input parameters such as video generation quality loss, current video frame rate, current available computing resources, preset maximum computing resources, and maximum frame rate. These models can use methods such as regression analysis, reinforcement learning, and deep learning to determine how to adjust the frame rate, resolution, computing resources, etc. to achieve a balance between video quality and generation efficiency.

[0088] The video optimization plan can be the result generated by the video optimization model based on the input data, used to guide how to adjust the parameters in the video generation process to optimize the video quality. The optimization plan usually includes adjustments to multiple generation parameters. For example, frame rate adjustment: increasing or decreasing the frame rate of the video. Resolution adjustment: adjusting the video resolution according to the computing resources and generation quality loss. Computing resource allocation: reallocating computing resources to ensure the video generation quality, such as allocating more GPU resources or reducing memory occupancy.

[0089] The calculated video generation quality loss can be compared with the preset quality loss threshold. If the quality loss exceeds the threshold, the video generation quality loss, current video frame rate, current available computing resources, preset maximum available computing resources, and preset maximum frame rate are input into the preset video optimization model, and the preset video optimization model (such as a deep learning model, reinforcement learning model, etc.) is used to process the input data to generate an optimization plan. The optimization model will predict the best parameter adjustment plan based on the input data, such as adjusting the frame rate, resolution, computing resources, etc., to minimize the video generation quality loss. Specific adjustments are made according to the optimization plan. For example, frame rate adjustment: increasing or decreasing the video frame rate according to the optimization plan to improve the generation quality or enhance the computing efficiency. Resolution adjustment: adjusting the video resolution according to the availability of computing resources (for example, reducing the resolution to reduce the consumption of computing resources). Optimizing computing resource allocation: such as reallocating computing resources (for example, increasing GPU usage or adjusting memory usage) to improve the video generation quality. According to the optimized plan, the video is regenerated to ensure that the video quality loss is reduced to an acceptable range and the video generation efficiency is improved.

[0090] The training process of the preset video optimization model is as follows:

[0091] Collect historical data from different video generation tasks. This data typically includes the following: Input data for the video generation process: including pixel information of the video, the current video frame rate, the currently available computing resources, the preset maximum available computing resources, the preset maximum frame rate, etc. Optimization solutions: including changes in video generation parameters before and after optimization, such as frame rate, resolution, computing resource allocation, etc. The collected data needs to be organized into a dataset suitable for training the model. Usually, the dataset should include input feature data: such as video generation quality loss, frame rate, currently available computing resources, preset maximum available computing resources, preset maximum frame rate. Output target (label) data: This part of the data is the optimization target of the model and can be the final optimization solution of the system. The task of label annotation is to correspond the historical data with the actual video generation quality loss, generation effect, etc. For example, if the generated video has excessive quality loss, the optimization solution may be to reduce the video frame rate, decrease the resolution, or reallocate computing resources. Annotation target: When training the model, the target output of each data point, that is, the optimization solution, needs to be annotated. This usually requires manual analysis of the data or the use of existing heuristic methods to generate optimization strategies. Once the dataset is prepared and annotated, the video optimization model can be trained. The training process usually includes the following steps: Select the model architecture: Model type: Common models include deep learning-based regression models, decision trees, reinforcement learning models, etc. In this case, deep neural networks (such as multi-layer perceptrons, convolutional neural networks, LSTMs, etc.) are often used for such tasks, especially when dealing with complex input data. Data partitioning: Training set: Used to train the model, usually accounting for 70%-80% of the dataset. Validation set: Used to evaluate the performance of the model during training and adjust the hyperparameters of the model. Test set: Used to finally evaluate the performance of the model and check whether the model can effectively generalize to new data. Loss function design: The core of model training is the loss function, which determines the direction of model optimization. For video optimization tasks, multiple loss functions can be designed: Mean squared error loss function: Commonly used in regression tasks and can be used to measure the difference between the video generation quality loss and the target optimization solution. Categorical cross-entropy loss: If the optimization solution is discretized into different categories, the cross-entropy loss function can be used for training. Reinforcement learning loss function: If a reinforcement learning model is used, the loss function may be adjusted according to the rewards or punishments after the model makes decisions. Model training process: Forward propagation: Through the input feature data, the model performs forward calculations to obtain the predicted optimization solution. Backward propagation: Calculate the gradient of the loss function with respect to the model parameters and adjust the model parameters through the backward propagation algorithm to reduce the loss. Iterative training: Optimize the model through multiple iterations to gradually reduce the loss until the model converges. Adjustment and optimization: During the training process, hyperparameters (such as learning rate, number of training epochs, batch size, etc.) can be adjusted, and the validation set can be used to select the optimal model.Then, use the test set to evaluate the performance of the model, including calculating metrics such as accuracy and mean squared error. Once the model passes the evaluation and meets the performance requirements, it can be deployed to an actual video generation system.

[0092] In this solution, an intelligent and automated optimization process is introduced, which can not only improve the quality of video generation and user experience, but also efficiently utilize computing resources while ensuring the video quality, thereby ensuring the stability and flexibility of the system.

[0093] Based on the above technical solution, optionally, the preset formula for calculating the quality loss is:

[0094]

[0095] where L quality is the quality loss of video generation; N is the total number of video frames; i is the video frame index; V frame (i) is the pixel information of the video data of the i-th frame; is the sum of the squares of the pixels of the video data of the i-th frame; C resource is the currently available computing resource; C max is the preset maximum available computing resource; α is the preset video quality adjustment factor; FrameRate is the current video frame rate; F max is the preset maximum frame rate.

[0096] In this solution, the value of α can be obtained through experimental tuning, and users can adjust it according to the actual situation to gradually find the most suitable value of α. For example, use methods such as cross-validation to evaluate the impact of different values of α on the balance between video quality and frame rate.

[0097] Based on the above technical solution, optionally, after obtaining the video data, the method further includes:

[0098] Calculate the temporal consistency loss of consecutive frames of the video data according to the current video frame rate of each frame of the video data and the preset formula for calculating the temporal consistency loss;

[0099] If there are consecutive frames with a temporal consistency loss greater than the preset temporal consistency loss threshold, process the consecutive frames with a temporal consistency loss greater than the preset temporal consistency loss threshold according to the preset processing mechanism.

[0100] In this solution, consecutive frames can be adjacent frames arranged in chronological order in the video. For example, the i-th frame and the (i - 1)-th frame form a pair of consecutive frames. The time interval between them is usually fixed and equal to the reciprocal of the frame rate. Consecutive frames are used to measure the coherence of motion or changes in the video.

[0101] The temporal consistency loss can be an indicator used to measure the changes or inconsistencies between adjacent frames. If the content changes too much between adjacent frames, it may affect the smoothness and naturalness of the video, resulting in discontinuous frame jumps or distortions. The goal is to minimize the temporal consistency loss so that the transition between adjacent frames is smoother and more coherent.

[0102] The preset temporal consistency loss threshold can be a set threshold used to determine which consecutive frames have an excessive temporal consistency loss. When the temporal consistency loss between consecutive frames exceeds this threshold, it means that there is a problem with the smoothness of video generation and further optimization is required.

[0103] The preset processing mechanism can be a scheme for automatically adjusting or correcting those consecutive frames with excessive temporal consistency loss during video generation. Specifically, it can include frame interpolation: inserting new frames between two consecutive frames to reduce content mutations and incoherence. Motion estimation and compensation: smoothing the transition of the picture by estimating the motion vectors between adjacent frames. Frame reconstruction: regenerating or correcting those frames that do not meet the consistency requirements to ensure video quality.

[0104] The current video frame rate can be substituted into the preset temporal consistency loss calculation formula to obtain the temporal consistency loss of consecutive frames of video data, and the calculated temporal consistency loss is compared with the preset threshold. If the temporal consistency loss of a certain frame pair is greater than the threshold, it is considered that there is a large inconsistency in this consecutive frame. For those consecutive frames with a temporal consistency loss exceeding the threshold, they are corrected according to the preset processing mechanism. For example, frame interpolation: inserting additional intermediate frames to make the changes between frames smoother. Motion estimation and compensation: adjusting the inconsistent parts to a smooth transition by calculating the motion vectors between adjacent frames. Frame reconstruction: ensuring the natural and smooth transition with surrounding frames by regenerating or optimizing these frames. After processing the frame pairs with excessive consistency loss, the consecutive frames of the video are rechecked until the temporal consistency loss of all frame pairs is below the threshold to ensure the smoothness and natural transition of the video.

[0105] In this solution, the natural transition between consecutive frames during video generation can be ensured by dynamically calculating and adjusting the temporal consistency loss, thereby improving the smoothness and visual consistency of the video.

[0106] Based on the above technical solution, optionally, the preset temporal consistency loss calculation formula is:

[0107]

[0108] where, L temporal is the temporal consistency loss; N is the total number of video frames; i is the video frame number index; Δt is the time difference between consecutive frames; Vframe V(i) is the pixel information of the video data of the i-th frame; frame V(i - 1) is the pixel information of the video data of the (i - 1)-th frame.

[0109] In this solution, △t can be the time interval between two adjacent frames in the video. Usually, each frame of the video is displayed at a constant time interval. Therefore, the time difference between consecutive frames is usually fixed and equal to the reciprocal of the frame rate.

[0110] Figure 3 FIG. is a schematic block diagram of a video generation system based on visual hierarchical autoregression provided by an embodiment of the present disclosure. It is characterized in that the system includes:

[0111] A text encoding module 301, configured to obtain description text data, and encode the description text data according to a preset text encoder to obtain a text token sequence of the description text data;

[0112] A hierarchical autoregressive generation module 302, configured to input the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences of each scale;

[0113] A feature encoding module 303, configured to determine feature encoding data according to the video token sequences of each scale and a preset encoding method;

[0114] A decoding module 304, configured to output the feature encoding data to a decoder to obtain video data.

[0115] Figure 4 FIG. shows a schematic block diagram of an electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0116] The electronic device 400 includes a computing unit 401, which can execute various appropriate actions and processes according to a computer program stored in the ROM 402 or a computer program loaded from the storage unit 408 into the RAM 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The I / O interface 405 is also connected to the bus 404.

[0117] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0118] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the above-described video generation method based on visual hierarchical autoregression. For example, in some embodiments, the above-described video generation method based on visual hierarchical autoregression can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the above-described video generation method based on visual hierarchical autoregression can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the above-described video generation method based on visual hierarchical autoregression in any other suitable manner (e.g., by means of firmware).

[0119] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0120] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0122] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0123] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0124] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0125] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0126] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video generation method based on visual hierarchical autoregression, characterized in that The method includes: Obtain descriptive text data, and encode the descriptive text data according to a preset text encoder to obtain a text token sequence of the descriptive text data; Input the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale; Determine feature encoding data according to the video token sequences at each scale and a preset encoding method; Output the feature encoding data to a decoder to obtain video data.

2. The method according to claim 1, characterized in that, Wherein, Inputting the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale includes: Input the text token sequence into a preset hierarchical autoregressive model. The preset hierarchical autoregressive model generates video token sequences at each scale step by step, and generates video token sequences at the next scale according to the currently generated video token sequences until a preset termination condition is reached, thereby obtaining video token sequences at each scale.

3. The method according to claim 1, characterized in that Wherein, Determining feature encoding data according to the video token sequences at each scale and a preset encoding method includes: Convert the video token sequences at each scale into multi-frame image data, and perform 3D convolution operations on the multi-frame image data based on the VQ-VAE method to obtain spatio-temporal feature representations; Encode the spatio-temporal feature representations according to a multi-scale pyramid structure to obtain feature encoding data.

4. The method according to claim 1, characterized in that, Wherein, After obtaining the video data, the method further includes: Determine the pixel information of each frame of the video data and the current video frame rate of the video data, obtain the current available computing resources, a preset maximum available computing resource, and a preset maximum frame rate, and calculate the video generation quality loss according to the pixel information, the current video frame rate, the current available computing resources, the preset maximum available computing resource, the preset maximum frame rate, and a preset quality loss calculation formula.

5. The method according to claim 4, characterized in that, Wherein, After calculating the video generation quality loss, the method further includes: If the video generation quality loss exceeds a preset quality loss threshold, input the video generation quality loss, the current video frame rate, the current available computing resources, the preset maximum available computing resource, and the preset maximum frame rate into a preset video optimization model to obtain a video optimization plan, and optimize the video data according to the video optimization plan.

6. The method according to claim 4, wherein Wherein, The preset quality loss calculation formula is: Among them, L quality is the video generation quality loss; N is the total number of video frames; i is the video frame index; V frame (i) is the pixel information of the video data of the i-th frame; is the sum of squares of pixels of the video data of the i-th frame; C resource is the currently available computing resource; C max is the preset maximum available computing resource; α is the preset video quality adjustment factor; FrameRate is the current video frame rate; F max is the preset maximum frame rate.

7. The method according to claim 4, wherein Wherein, After obtaining the video data, the method further includes: Calculate the temporal consistency loss of consecutive frames of the video data according to the current video frame rate of each frame of the video data and a preset temporal consistency loss calculation formula; If there are consecutive frames with a temporal consistency loss greater than a preset temporal consistency loss threshold, process the consecutive frames with a temporal consistency loss greater than the preset temporal consistency loss threshold according to a preset processing mechanism.

8. The method according to claim 7, wherein Wherein, The preset temporal consistency loss calculation formula is: Among them, L temporal is the temporal consistency loss; N is the total number of video frames; i is the video frame index; Δt is the time difference between consecutive frames; V frame (i) is the pixel information of the video data of the i-th frame; V frame (i - 1) is the pixel information of the video data of the (i - 1)-th frame.

9. A video generation system based on visual hierarchical autoregression for performing the method according to any one of claims 1-8, characterized in that, The system includes: A text encoding module, configured to obtain descriptive text data, and encode the descriptive text data according to a preset text encoder to obtain a text token sequence of the descriptive text data; A hierarchical autoregressive generation module for inputting the text token sequence into a preset hierarchical autoregressive model to obtain video token sequences at each scale; A feature encoding module for determining feature encoding data according to the video token sequences at each scale and a preset encoding method; A decoding module for outputting the feature encoding data to a decoder to obtain video data.

10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

Citation Information

Cited By

  • Automatic video generation system and method based on AI Agent multi-mode cooperative control

    CN121126084A

  • An automated video generation system and method with AIAgentic multimodal cooperative control

    CN121126084B