Video generation method and device and electronic equipment

By extracting the features and related information of multiple sub-signals, structured text description labels and description content are generated, which solves the controllability and accuracy problems of AI models when processing diverse conditional instructions and improves the quality of video generation.

CN120935426APending Publication Date: 2025-11-11NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510789120.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing AI models struggle to accurately capture all content when dealing with diverse and complex conditional instructions, resulting in poor controllability and accuracy in video generation and low video quality.

Method used

By acquiring multiple sub-signals (text instructions, spatial structure information, object action information, object appearance information, and camera control information), signal features and their associated information are extracted to generate structured text description labels and description content, which are then input into the video generation model to output the target video.

Benefits of technology

It improves the controllability and accuracy of video generation, resulting in videos that highly match user intent and enhance video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935426A_ABST
    Figure CN120935426A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device and electronic equipment. The method comprises the following steps: acquiring a control signal; extracting signal features corresponding to the sub-signals and feature association information among the signal features of the multiple sub-signals; on the basis of the signal features and the feature association information, generating description content corresponding to a preset text description tag; wherein the text description label comprises multiple of global scene description, object description, background description, camera description, style description and behavior description; and inputting the text description label and the description content into a preset video generation model, and outputting a target video. According to the mode, the generated video is highly matched with the intention of the user, so that the controllability and the accuracy of video generation are improved, and the video generation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a video generation method, apparatus, and electronic device. Background Technology

[0002] When generating video content using an AI (Artificial Intelligence) model, conditional instructions need to be input into the AI ​​model to guide its output. When there are many conditional instructions, or when their formats are complex, the video content generated by the AI ​​model may not satisfy all of them, resulting in poor controllability and accuracy in video generation, and consequently, lower video quality. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a video generation method, apparatus and electronic device to improve the controllability and accuracy of video generation and improve the quality of video generation.

[0004] In a first aspect, embodiments of the present invention provide a video generation method, the method comprising: acquiring control signals; wherein the control signals include multiple sub-signals among text instructions, spatial structure information, object action information, object appearance information, and camera control information; extracting signal features corresponding to the sub-signals, and feature association information between the signal features of the multiple sub-signals; wherein the feature association information is used to indicate the semantic association relationship and / or constraint relationship between the signal features of the multiple sub-signals; generating description content corresponding to preset text description tags based on the signal features and feature association information; wherein the text description tags include multiple types among global scene description, object description, background description, camera description, style description, and behavior description; inputting the text description tags and description content into a preset video generation model, and outputting a target video.

[0005] Secondly, embodiments of the present invention provide a video generation apparatus, comprising: a signal acquisition module for acquiring control signals; wherein the control signals include multiple sub-signals among text instructions, spatial structure information, object action information, object appearance information, and camera control information; a feature extraction module for extracting signal features corresponding to the sub-signals, and feature association information between the signal features of the multiple sub-signals; wherein the feature association information is used to indicate semantic association relationships and / or constraint relationships between the signal features of the multiple sub-signals; a content generation module for generating description content corresponding to preset text description tags based on the signal features and feature association information; wherein the text description tags include multiple types among global scene description, object description, background description, camera description, style description, and behavior description; and a video output module for inputting the text description tags and description content into a preset video generation model and outputting a target video.

[0006] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-described video generation method.

[0007] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the above-described video generation method.

[0008] The embodiments of the present invention bring the following beneficial effects:

[0009] The aforementioned video generation method, apparatus, and electronic device acquire control signals, which include multiple sub-signals such as text commands, spatial structure information, object action information, object appearance information, and camera control information. They extract signal features corresponding to the sub-signals, as well as feature association information between the signal features of the multiple sub-signals. The feature association information indicates semantic relationships and / or constraint relationships between the signal features of the multiple sub-signals. Based on the signal features and feature association information, they generate description content corresponding to preset text description tags. These text description tags include multiple types such as global scene description, object description, background description, camera description, style description, and behavior description. The text description tags and description content are input into a preset video generation model to output the target video.

[0010] In the above method, after extracting signal features and feature relationships from multiple sub-signals, text description labels and corresponding description content are generated. These text description labels and corresponding description content are structured text information that can clearly and comprehensively express the user's intent. The video generation model can more easily understand and follow the instructions to generate videos, resulting in a high degree of matching between the generated videos and the user's intent. This improves the controllability and accuracy of video generation and enhances the quality of video generation.

[0011] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0013] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0014] Figure 1 A flowchart of a video generation method provided in an embodiment of the present invention;

[0015] Figure 2 This is a schematic diagram illustrating the training method of the feature extractor provided in an embodiment of the present invention;

[0016] Figure 3 A schematic diagram illustrating the training method of the semantic parsing model provided in an embodiment of the present invention;

[0017] Figure 4 A schematic diagram of a video generation device provided in an embodiment of the present invention;

[0018] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] In the field of video creation, such as the production of game videos, promotional videos, and cutscenes, users often need to input diverse and complex conditional instructions into AI models, enabling the AI ​​models to generate video content according to the requirements of these instructions. Conditional instructions may include: text descriptions (such as "a warrior is running in the forest"), specific character images (such as character concept art), desired character actions (such as a piece of motion capture data or a sequence of skeletal animations), camera movements (such as camera movement scripts or camera trajectory data), and specific scene atmosphere or art style (such as a reference image or style description).

[0021] Current AI models, when understanding and integrating conditional instructions with different content (such as appearance, actions, camera angles, and styles) or different formats (such as text, images, and coordinate data), cannot accurately capture all the content of the conditional instructions. They may attempt to satisfy one instruction condition but overlook another. For example, a video might match the scene described in the text of a conditional instruction, but the character's actions in the video might not match the specified action instructions. Similarly, a character's appearance might match the description in a conditional instruction, but the camera movement might not match the specified camera movement instructions. This results in poor controllability and accuracy in video generation, leading to lower video quality.

[0022] Based on this, the video generation method, apparatus, and electronic device provided in this embodiment of the invention can be applied to the process of generating videos through AI models.

[0023] like Figure 1 As shown, the video generation method includes the following steps:

[0024] Step S102: Acquire control signals; wherein, the control signals include multiple sub-signals from text commands, spatial structure information, object motion information, object appearance information, and camera control information;

[0025] This control signal can be edited by the user according to their needs. It is used to indicate the video content of the target video to be generated. The aforementioned text instructions, spatial structure information, object motion information, object appearance information, and camera control information are five sub-signals. The control signal can include all of these five sub-signals, or any two, three, or four of them.

[0026] The aforementioned text instructions are typically in text format, meaning instructions described in natural language, such as Chinese or English. Text instructions may include requirements for the target video to be generated, or supplementary descriptions for other sub-signals. For example, a text instruction might be "Show a swordsman practicing martial arts in a bamboo forest at night, requiring an ink painting style, with the camera following the swordsman."

[0027] The aforementioned spatial structure information is usually in image format. The spatial structure information may include one or more images, which may include background environment, video style, and other content. It may also include depth information, which indicates the geometric structure, layout position, outline, and occlusion relationship of each object in the video.

[0028] The aforementioned object motion information can be in coordinate format, specifically including a multi-frame pose sequence of the object. Each frame of the pose sequence includes the object's key points and their corresponding coordinates. For example, when the object is a human body, the key points are the joints of the human body. This object motion information can also be in image or video format; that is, based on the aforementioned pose sequence, the key points are connected to obtain a pose diagram.

[0029] The aforementioned object appearance information can be in image format, and the object can be a person, animal, still life, etc.; different objects correspond to different appearance images. The aforementioned camera control information can be in tabular or numerical format, and can specifically include multi-frame control data. In each frame of control data, data such as the virtual camera's position and orientation are included.

[0030] Step S104: Extract the signal features corresponding to the sub-signals, and the feature association information between the signal features of multiple sub-signals; wherein, the feature association information is used to indicate the semantic association relationship and / or constraint relationship between the signal features of multiple sub-signals;

[0031] Signal features of sub-signals can be extracted using a feature extractor. Different sub-signals can use the same or different feature extractors. Considering that different sub-signals may have different formats or contents, using different feature extractors for different sub-signals can make the signal features more accurate. The feature extractors used for different sub-signals can be trained separately using sample signals that match the sub-signals, thereby improving the accuracy and comprehensiveness of the feature extractor in extracting features specific to a particular sub-signal.

[0032] Extracting feature association information between signal features of multiple sub-signals can be achieved through large language models or other AI models; for example, the attention mechanism in a large language model can be used to extract feature association information between signal features of different sub-signals.

[0033] The feature association information can include only semantic association, only constraint association, or both semantic association and constraint association. Semantic association can be understood as the correlation between semantics in different sub-signals. For example, if the text instruction includes 'run', the text instruction 'run' will have a high attention weight with the action information of the object's legs in the object action information, indicating that the action information of the object's legs is highly semantically related to the text instruction 'run'.

[0034] A constraint relationship can be understood as the possible restrictive relationship between different sub-signals; for example, if a text instruction includes 'wearing a red hat' and an image of the object's appearance information also includes a red hat, then the text instruction 'wearing a red hat' and the red hat in the image of the object's appearance information are strongly correlated, forming a constraint relationship.

[0035] Step S106: Based on signal features and feature association information, generate description content corresponding to preset text description tags; wherein, the text description tags include multiple types of global scene description, object description, background description, camera description, style description, and behavior description;

[0036] In this embodiment, the extracted signal features and feature association information are not directly input into the video generation model, but are used to generate description content in the same format; text description tags are pre-set, and the description content corresponding to each text description tag is generated and filled in by the signal features and feature association information.

[0037] The generation of descriptive content can be achieved through a large language model or other AI models. The descriptive content corresponding to the global scene description mentioned above can include general information such as the scene, time, and location of the target video to be generated; the descriptive content corresponding to the object description mentioned above can include the appearance, attributes, and state of objects in the video; the descriptive content corresponding to the background description mentioned above can include the environment, atmosphere, and secondary objects in the video.

[0038] The descriptions corresponding to the camera descriptions mentioned above may include the initial position, viewpoint, motion trajectory, and operation mode of the virtual camera; the descriptions corresponding to the style descriptions mentioned above may include the visual style of the video (such as realistic style, cartoon style, oil painting style, etc.), color tone, lighting effects, emotional atmosphere, etc.; the descriptions corresponding to the behavior descriptions mentioned above may include object actions and interactive behaviors arranged in chronological order.

[0039] The aforementioned text description labels and their corresponding descriptions can be textualized and structured data. Some descriptions can also be associated with data such as images and pose sequence information. When signal features and feature association information are input into the model, the model needs to predict the description content corresponding to each text description label.

[0040] Step S108: Input the text description tags and description content into the preset video generation model and output the target video.

[0041] The text description tags and description content are in a structured text format, which is highly compatible with most video generation models. The text description tags and description content can be directly input into the video generation model, or input into the video generation model after simple adaptation processing.

[0042] This video generation model can meet the requirements of this embodiment without training, using existing video generation models. This video generation model can be a SOTA (State-Of-The-Art, currently the highest level model) model.

[0043] The conditional encoder (such as a text encoder) in the video generation model converts the text description tags and description content into the internal data format of the video generation model, and then performs the denoising process or autoregressive process in the video generation model to gradually generate video frames of the target video.

[0044] In the aforementioned video generation method, control signals are first acquired; these control signals include multiple sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information; signal features corresponding to the sub-signals are extracted, as well as feature association information between the signal features of the multiple sub-signals; the feature association information is used to indicate the semantic association and / or constraint relationship between the signal features of the multiple sub-signals; based on the signal features and feature association information, description content corresponding to preset text description tags is generated; the text description tags include multiple types from global scene description, object description, background description, camera description, style description, and behavior description; the text description tags and description content are input into a preset video generation model to output the target video.

[0045] In the above method, after extracting signal features and feature relationships from multiple sub-signals, text description labels and corresponding description content are generated. These text description labels and corresponding description content are structured text information that can clearly and comprehensively express the user's intent. The video generation model can more easily understand and follow the instructions to generate videos, resulting in a high degree of matching between the generated videos and the user's intent. This improves the controllability and accuracy of video generation and enhances the quality of video generation.

[0046] In one implementation, multiple sub-signals are input into feature extractors corresponding to each sub-signal, and the outputs signal features corresponding to each sub-signal. A corresponding feature extractor can be set for each sub-signal, and this feature extractor is pre-trained using the sample signal corresponding to that sub-signal. For example, a motion extractor can be set for object action information, and a camera feature extractor can be set for camera control information.

[0047] Alternatively, feature extractors can be set according to the format of the sub-signal. For example, an image feature extractor can be determined for an image-formatted sub-signal, and this image feature extractor is pre-trained using images; the spatial structure information and object appearance information in the sub-signal can be in image format. Similarly, a video feature extractor can be determined for a video-formatted sub-signal, and this video feature extractor is pre-trained using videos, etc. The spatial structure information and object motion information in the sub-signal can be in video format.

[0048] In the above method, feature extractors are set up separately for sub-signals of different formats or types, so that the feature extractors can focus more on extracting single signal features, thereby obtaining more complete and accurate features for various sub-signals, which is conducive to improving the controllability and accuracy of video generation and improving the quality of video generation.

[0049] Specifically, if the control signal includes a first sub-signal in image format, the first sub-signal is input into the image feature extractor; the image feature extractor divides the first sub-signal into multiple image blocks, generates a first projection vector corresponding to each image block, and sets the first block position information corresponding to the first projection vector; wherein, the first block position information is used to: indicate the spatial position of the image block corresponding to the first projection vector in the first sub-signal; the first projection vector and the first block position information are encoded to obtain the signal features corresponding to the first sub-signal.

[0050] The first sub-signal can be spatial structure information, object appearance information, etc.; it should be noted that the spatial structure information can be in image format, video format, or both image and video formats. The image feature extractor can be a model based on a hierarchical visual transformer, such as the ViT (Visual Transformer) model, or other models.

[0051] The image feature extractor segments the first sub-signal to obtain multiple non-overlapping image blocks; each image block generates a corresponding first projection vector through a preset projection method, such as linear projection. The aforementioned first block location information can specifically be a spatial location code of the image block.

[0052] The first projection vector and the first block location information are input into the encoder in the image feature extractor, which outputs the signal features corresponding to the first sub-signal. This encoder typically includes a multi-head attention mechanism and a feedforward network.

[0053] If the control signal includes a second sub-signal in video format, the second sub-signal is input into the video feature extractor. The video feature extractor divides the video frames in the second sub-signal into multiple frame blocks, generates a second projection vector corresponding to each frame block, and sets the second block position information corresponding to the second projection vector and the temporal position information of the video frame to which the frame block belongs. The second block position information is used to indicate the spatial position of the frame block corresponding to the second projection vector in the video frame. The second projection vector, the second block position information, and the temporal position information are encoded to obtain the signal features corresponding to the second sub-signal.

[0054] The second sub-signal can be spatial structure information, object action information, etc.; the video feature extractor can be a model based on a hierarchical visual converter, such as the ViT model, or other models.

[0055] The video feature extractor segments each video frame in the second sub-signal to obtain multiple non-overlapping frame blocks. Each frame block generates a corresponding second projection vector through a preset projection method, such as linear projection. Specifically, the aforementioned second block location information can be the spatial location encoding of the frame block within its corresponding video frame. Furthermore, the aforementioned temporal location information can be the sequential encoding of the video frames within the second sub-signal.

[0056] The aforementioned second projection vector and second block location information are input to the encoder in the video feature extractor, which outputs the signal features corresponding to the second sub-signal. This encoder typically includes a multi-head attention mechanism and a feedforward network. It should be noted that the video feature extractor can be the aforementioned image feature extractor, or it can be a feature extractor trained separately based on sample videos.

[0057] If the control signal includes object motion information, and the object motion information is in coordinate format, the object motion information is input into the motion extractor, and the motion features corresponding to the object motion information are output; or, if the control signal includes object motion information, and the object motion information is in video format, the object motion information is input into the video feature extractor, and the motion features of the object motion information are output.

[0058] When the object's motion information is in coordinate format, it includes a multi-frame pose sequence of the object. Each frame's pose sequence includes keypoints and their corresponding coordinates. In this case, the motion extractor can be a Spatio-Temporal Graph Convolutional Network (ST-GCN) or a Pose Transformer. ST-GCN utilizes the graph structure of a human or animal skeleton for convolutional operations to capture the spatial dependencies and temporal dynamics between keypoints. The Pose Transformer, on the other hand, directly applies a self-attention mechanism to the keypoint sequence to learn spatiotemporal relationships.

[0059] When the object's motion information is in image or video format, the motion extractor can be either the aforementioned image feature extractor or video feature extractor.

[0060] If the control signal includes camera control information, the camera control information is input into the camera feature extractor, which outputs the camera operation features of the camera control information. This camera feature extractor can be a lightweight spatiotemporal feature extraction network, such as a small 3D-CNN (3D-Convolutional Neural Networks) model, or a simplified visual transducer.

[0061] Camera control information includes multi-frame control data. Each frame's control data includes information such as the virtual camera's position and orientation. The camera feature extractor divides each frame's control data into multiple data blocks and sets the spatial location information of each data block within the control data, as well as the temporal location information of the control data to which that data block belongs. Spatiotemporal features of the data blocks are extracted using convolutional layers or attention layers, thereby extracting the virtual camera's translation, rotation, scaling, and other motion patterns.

[0062] Furthermore, the signal features of various sub-signals are arranged in a preset order, and feature markers corresponding to the signal features are set; wherein, the feature markers are used to indicate the start and end positions of the sub-signals and / or signal features corresponding to the signal features.

[0063] The preset order can be: signal features of text commands first, followed by spatial structure information, object motion information, object appearance information, and finally camera control information; of course, other arrangements are also possible.

[0064] The feature marker can indicate only the type of sub-signal corresponding to the signal feature, or only the start and end positions of the signal feature of each sub-signal, or simultaneously indicate both the type of sub-signal corresponding to the signal feature and the start and end positions of the signal feature of the sub-signal.

[0065] In one example, <|camera_start|>F_camera<|camera_end|>; where <|camera_start|> and <|camera_end|> are feature markers located at the left and right ends of the signal feature F_camera, enclosing the signal feature. The word 'camera' in the feature markers indicates that this signal feature corresponds to camera control information; the positions of <|camera_start|> and <|camera_end|> indicate the start and end positions of the signal feature corresponding to the camera control information.

[0066] By setting feature tags, the signal features of multiple sub-signals can be arranged into a sequence, which facilitates subsequent feature processing.

[0067] In one approach, if the control signal includes a text instruction and a third sub-signal other than the text instruction, the signal features of the third sub-signal are converted into a preset text feature space to obtain the converted signal features of the third sub-signal; wherein, the text feature space is the feature space where the signal features of the text instruction are located.

[0068] This text feature space is related to the model that extracts the signal features of the text instruction. For example, text instructions are usually extracted using a large language model; therefore, the feature space containing the signal features of the text instruction is the word embedding space of the large language model.

[0069] A multilayer perceptron or Q-Former model can be used to transform the signal features of the third sub-signal into a preset text feature space. The transformed signal features of the third sub-signal are in the same feature space as the signal features of the text instruction, thereby achieving seamless interaction between signal features and facilitating the subsequent extraction of feature association information.

[0070] In practical implementation, a pre-defined multi-head attention mechanism is used to extract feature association information between the signal features of any two sub-signals. The multi-head attention mechanism in the large language model can calculate the correlation degree between the signal features of any two sub-signals, thereby achieving deep interaction and information fusion between text features and image features, video features, action features, and camera operation features. By calculating the correlation degree between signal features, the large language model learns and understands the semantic associations, complementarities, and potential constraints between different signal features.

[0071] Furthermore, the signal features with feature correlation information are fused to obtain the fused signal features. Feature correlation information may exist between any two sub-signals. Specifically, this feature correlation information can be attention weights. The larger the attention weights, the greater the correlation between the signal features of the two corresponding sub-signals.

[0072] In one example, the signal features of the fourth and fifth sub-signals have feature association information, namely attention weights. In the specific fusion process, the signal features of the fourth and fifth sub-signals are weighted and summed through attention weights to obtain the signal features of the fused fourth sub-signal and the signal features of the fused fifth sub-signal.

[0073] By fusing signal features, the signal features of each sub-signal absorb and integrate the signal features of other sub-signals, making it easier for the model to understand the correlation, dependence and potential conflict between the sub-signals, and to understand the user's overall creative intent, thereby improving the quality of video generation.

[0074] Specifically, the fused signal features are input into a pre-defined prediction model, which then outputs descriptive content corresponding to pre-defined text labels. This prediction model can be a Sequential Autoregressive Prediction (SAP) model or other types of prediction models. The prediction model can be guided to begin generating descriptive content by a specific start prompt or control representative upon receiving the fused signal features.

[0075] The prediction model predicts the description content for each text description tag one by one. For example, the prediction model is first prompted to generate a description of the global scene, and then prompted to generate a description of the object, and so on. The description content of the prediction model may also need to be decoded, for example, by using beam search or Top-K / Top-P sampling methods, to make the text of the description content more fluent and accurate, to make the language of the description text as precise and detailed as possible, and to eliminate ambiguity as much as possible.

[0076] The text description tags and their corresponding descriptions are highly structured and information-rich text sequences. They use precise natural language to describe in detail various aspects of the target video, such as the overall scene narrative, details and states of key objects, background environmental elements, camera angles and movements, visual style and lighting atmosphere, and the timing of key actions. This accurately and comprehensively describes the user's video creation intent and improves the quality of video generation.

[0077] In actual implementation, such as Figure 2 As shown, the feature extractor described above is trained in the following manner:

[0078] Step S202: Input the sample sub-signal to the feature extractor and output the sample features corresponding to the sample sub-signal; wherein, the sample sub-signal is set with sample text description;

[0079] Feature extractors can be trained for all the aforementioned sub-signals, or only for a subset of the sub-signals, such as motion extractors and camera feature extractors. Different feature extractors use different sample sub-signals.

[0080] The purpose of training the feature extractor is to ensure that the features output by the feature extractor, after being processed by the spatial transformation module, closely approximate the feature space of the signal features of the text instruction. Therefore, each sample sub-signal is assigned a corresponding sample text description.

[0081] Step S204: Input the sample features into the spatial transformation model and output the text features corresponding to the sample features; the spatial transformation model can be a multilayer perceptron or a Q-Former model.

[0082] Step S206: Determine a first loss value based on text features and sample text descriptions, and adjust the parameters in the feature extractor and / or spatial transformation model based on the first loss value.

[0083] The first loss value indicates the similarity between the text features and the sample text description. Specifically, the first loss value can be calculated using mean squared error or cross-entropy loss. During the training of the feature extractor, you can adjust only the parameters in the feature extractor or adjust the parameters in both the feature extractor and the spatial transformation model.

[0084] Furthermore, the signal features of multiple sub-signals are input into the semantic parsing model. The semantic parsing model extracts the feature association information between the signal features of multiple sub-signals. Based on the signal features and feature association information, the description content corresponding to the preset text description label is generated.

[0085] Among them, such as Figure 3 As shown, this semantic parsing model is trained in the following way:

[0086] Step S302: Input the sample sub-signal in the sample control signal into the feature extractor and output the sample signal features corresponding to the sample sub-signal; wherein, the sample control signal is set with sample description content.

[0087] After the feature extractor is trained, the semantic parsing model is trained. The purpose of training the semantic parsing model is to make the final description output by the semantic parsing model more accurate and comprehensive.

[0088] Step S304: Input the sample signal features into the semantic parsing model and output the predicted description content;

[0089] After the sample signal features are input into the semantic parsing model, the semantic parsing model extracts the feature association information between the sample signal features of various sample sub-signals. Based on the sample signal features and feature association information, it outputs the predicted description content corresponding to the preset text description label.

[0090] Step S306: Determine the second loss value based on the predicted description content and the sample description content, and adjust the parameters in the semantic parsing model based on the second loss value.

[0091] The second loss value indicates the degree of similarity between the predicted description and the sample description. Specifically, the second loss value can be calculated using the autoregressive language model loss function or the cross-entropy loss function.

[0092] This semantic parsing model can also employ a progressive hybrid training approach. Specifically, the sample control signals are pre-divided into multiple signal groups; wherein, the sample control signals in the first group of multiple signal groups include: any type of sample sub-signal other than text instructions; for example, object action information or camera control information,

[0093] The sample control signals in the second group of multiple signal groups include sample sub-signals from text commands, spatial structure information, object motion information, object appearance information, and camera control information; for example, each sample control signal includes two or three sample sub-signals. The sample control signals in the third group of multiple signal groups include sample sub-signals from text commands, spatial structure information, object motion information, object appearance information, and camera control information; in this third group, the sample control signals include all sample sub-signals.

[0094] Multiple sets of signals are sequentially input into the feature extractor, and the output sample feature signals are then input into the semantic parsing model for training. First, in the initial training phase, the sample control signals from the first set are input into the semantic parsing model; then, in the middle training phase, the sample control signals from the second set are input into the semantic parsing model; finally, in the later training phase, the sample control signals from the third set are input into the semantic parsing model.

[0095] Each of the aforementioned signal groups is assigned a corresponding loss weight, which is used to determine the second loss value for that signal group. The loss weight for each group of sample control signals can be pre-set. After multiple groups of sample control signals are input into the semantic parsing model, the loss values ​​for each group of sample control signals are weighted and summed using these loss weights to obtain the final second loss value. This second loss value is then used to adjust the parameters in the semantic parsing model, thereby training the semantic parsing model.

[0096] Furthermore, at least some of the text instructions in the sample control signals are deleted according to a preset probability. During training, some or all of the text instructions in the sample control signals are randomly deleted, and the deleted text instructions can be replaced with empty strings. In this way, the semantic parsing model can be guided to infer information and generate descriptive content based on sub-signals other than text instructions, thereby increasing the semantic parsing model's ability to understand non-text signals and its robustness.

[0097] The video generation method provided in this embodiment has the following advantages:

[0098] (1) Extremely high controllability: By accurately parsing multimodal and multi-structured control signals into text instructions of the same structure, the ability to control the details of the generated video content is greatly enhanced; whether it is the specific actions of the character, the complex camera arrangement, or the subtle style requirements, they can be more accurately reflected.

[0099] (2) Improve video quality and consistency: The description content input to the semantic parsing model is a structured semantic descriptor, which is richer in information and less ambiguous, and helps to generate higher quality, more coherent content and more consistent details in the video.

[0100] (3) Strong compatibility and flexibility: The semantic parsing model can be used in combination with a variety of different video generation models. Developers can choose the existing video generation model that best suits the project needs without modifying or retraining these models themselves. It has a plug-and-play feature, which greatly reduces integration costs and technical barriers and improves video generation efficiency.

[0101] (4) Improve content creation efficiency: In game development, for example, when creating promotional videos or cutscenes, staff can quickly express their ideas by combining various methods such as text, reference images, motion data, and camera movement requirements. This embodiment allows for accurate understanding of these mixed control signals and the generation of video drafts or final products that meet expectations. This significantly reduces the large amount of work involved in repetitive communication, manual adjustments, and post-editing in traditional processes, accelerating content iteration and output. For example, it can quickly generate multiple versions of videos featuring the same character using different skills in different scenes and with different camera movements.

[0102] See Figure 4 The diagram shows a video generation apparatus, which includes:

[0103] The signal acquisition module 40 is used to acquire control signals; wherein, the control signals include multiple sub-signals from text commands, spatial structure information, object motion information, object appearance information, and camera control information;

[0104] The feature extraction module 42 is used to extract the signal features corresponding to the sub-signals, as well as the feature association information between the signal features of multiple sub-signals; wherein, the feature association information is used to indicate the semantic association relationship and / or constraint relationship between the signal features of multiple sub-signals;

[0105] The content generation module 44 is used to generate description content corresponding to preset text description tags based on signal features and feature association information; wherein, the text description tags include multiple types of global scene description, object description, background description, camera description, style description and behavior description;

[0106] The video output module 46 is used to input text description tags and description content into a preset video generation model and output the target video.

[0107] The aforementioned video generation device acquires control signals, which include multiple sub-signals such as text commands, spatial structure information, object action information, object appearance information, and camera control information. It extracts signal features corresponding to the sub-signals, as well as feature association information between the signal features of the multiple sub-signals. The feature association information indicates the semantic relationships and / or constraints between the signal features of the multiple sub-signals. Based on the signal features and feature association information, it generates description content corresponding to preset text description tags. These text description tags include multiple types such as global scene description, object description, background description, camera description, style description, and behavior description. The text description tags and description content are input into a preset video generation model to output the target video.

[0108] In the above method, after extracting signal features and feature relationships from multiple sub-signals, text description labels and corresponding description content are generated. These text description labels and corresponding description content are structured text information that can clearly and comprehensively express the user's intent. The video generation model can more easily understand and follow the instructions to generate videos, resulting in a high degree of matching between the generated videos and the user's intent. This improves the controllability and accuracy of video generation and enhances the quality of video generation.

[0109] The aforementioned feature extraction module is used to: input multiple sub-signals into the feature extractors corresponding to the sub-signals respectively, and output the signal features corresponding to the sub-signals.

[0110] The aforementioned feature extraction module is used to: if the control signal includes a first sub-signal in image format, input the first sub-signal into the image feature extractor; divide the first sub-signal into multiple image blocks using the image feature extractor, generate a first projection vector corresponding to each image block, and set the first block position information corresponding to the first projection vector; wherein, the first block position information is used to: indicate the spatial position of the image block corresponding to the first projection vector in the first sub-signal; encode the first projection vector and the first block position information to obtain the signal features corresponding to the first sub-signal.

[0111] The aforementioned feature extraction module is used to: if the control signal includes a second sub-signal in video format, input the second sub-signal into the video feature extractor; divide the video frames in the second sub-signal into multiple frame blocks using the video feature extractor, generate a second projection vector corresponding to the frame block, and set the second block position information corresponding to the second projection vector and the time position information of the video frame to which the frame block belongs; wherein, the second block position information is used to: indicate the spatial position of the frame block corresponding to the second projection vector in the video frame; encode the second projection vector, the second block position information, and the time position information to obtain the signal features corresponding to the second sub-signal.

[0112] The aforementioned feature extraction module is used to: if the control signal includes object motion information and the object motion information is in coordinate format, input the object motion information into the motion extractor and output the motion features corresponding to the object motion information; or, if the control signal includes object motion information and the object motion information is in video format, input the object motion information into the video feature extractor and output the motion features of the object motion information.

[0113] The aforementioned feature extraction module is used to: if the control signal includes camera control information, input the camera control information into the camera feature extractor and output the camera operation features of the camera control information.

[0114] The above-mentioned device also includes a feature sorting module, used to: arrange the signal features of multiple sub-signals in a preset order, and set feature markers corresponding to the signal features; wherein, the feature markers are used to indicate: the start and end positions of the sub-signals and / or signal features corresponding to the signal features.

[0115] The aforementioned device further includes a spatial conversion module, used to: if the control signal includes a text instruction and a third sub-signal other than the text instruction, convert the signal features of the third sub-signal to a preset text feature space to obtain the signal features of the converted third sub-signal; wherein, the text feature space is the feature space where the signal features of the text instruction are located.

[0116] The aforementioned feature extraction module is used to: extract feature correlation information between signal features of any two sub-signals through a preset multi-head attention mechanism.

[0117] The aforementioned device also includes a feature fusion module, used to fuse signal features with feature association information to obtain fused signal features.

[0118] The above content generation module is used to: input the fused signal features into a preset prediction model, and output the description content corresponding to the preset text description label through the prediction model.

[0119] The aforementioned device further includes a first training module, used to train a feature extractor in the following manner: inputting a sample sub-signal into the feature extractor and outputting sample features corresponding to the sample sub-signal; wherein the sample sub-signal is provided with a sample text description; inputting the sample features into a spatial transformation model and outputting text features corresponding to the sample features; determining a first loss value based on the text features and the sample text description, and adjusting the parameters in the feature extractor and / or the spatial transformation model based on the first loss value.

[0120] The aforementioned video output module is used to: input the signal features of multiple sub-signals into a semantic parsing model, extract feature association information between the signal features of multiple sub-signals through the semantic parsing model, and generate description content corresponding to preset text description labels based on the signal features and feature association information; wherein, the semantic parsing model is trained in the following way: inputting the sample sub-signals in the sample control signal into the feature extractor, and outputting the sample signal features corresponding to the sample sub-signals; wherein, the sample control signal is correspondingly set with sample description content; inputting the sample signal features into the semantic parsing model, and outputting the predicted description content; determining a second loss value based on the predicted description content and the sample description content, and adjusting the parameters in the semantic parsing model based on the second loss value.

[0121] The aforementioned sample control signals are pre-divided into multiple signal groups. The first group of signal groups includes any sample sub-signal except for text commands. The second group includes a subset of sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. The third group includes sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. These signal groups are sequentially input to a feature extractor, and the output sample feature signals are then input to a semantic parsing model for training.

[0122] Each of the aforementioned signal groups is assigned a loss weight, which is used to determine the second loss value corresponding to the signal group.

[0123] The aforementioned device also includes a sample control module, used to delete at least a portion of the text instructions in the sample control signal according to a preset probability.

[0124] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-described video generation method. This electronic device can be a server or a terminal device.

[0125] See Figure 5As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100. The processor 100 executes the computer-executable instructions to implement the video generation method described above.

[0126] Furthermore, Figure 5 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 100, the communication interface 103 and the memory 101 connected via the bus 102.

[0127] The memory 101 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0128] Processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 100 or by instructions in software form. Processor 100 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0129] The processor in the aforementioned electronic device, by executing computer-executable instructions, can perform the following operations in the aforementioned video generation method:

[0130] A video generation method includes: acquiring control signals; wherein the control signals include multiple sub-signals from text instructions, spatial structure information, object action information, object appearance information, and camera control information; extracting signal features corresponding to the sub-signals, as well as feature association information between the signal features of the multiple sub-signals; wherein the feature association information is used to indicate the semantic association and / or constraint relationship between the signal features of the multiple sub-signals; generating description content corresponding to preset text description tags based on the signal features and feature association information; wherein the text description tags include multiple types from global scene description, object description, background description, camera description, style description, and behavior description; and inputting the text description tags and description content into a preset video generation model to output a target video.

[0131] Multiple sub-signals are input into the feature extractors corresponding to the sub-signals respectively, and the signal features corresponding to the sub-signals are output.

[0132] If the control signal includes a first sub-signal in image format, the first sub-signal is input into the image feature extractor; the image feature extractor divides the first sub-signal into multiple image blocks, generates a first projection vector corresponding to each image block, and sets the first block position information corresponding to the first projection vector; wherein, the first block position information is used to: indicate the spatial position of the image block corresponding to the first projection vector in the first sub-signal; the first projection vector and the first block position information are encoded to obtain the signal features corresponding to the first sub-signal.

[0133] If the control signal includes a second sub-signal in video format, the second sub-signal is input into the video feature extractor. The video feature extractor divides the video frames in the second sub-signal into multiple frame blocks, generates a second projection vector corresponding to each frame block, and sets the second block position information corresponding to the second projection vector and the temporal position information of the video frame to which the frame block belongs. The second block position information is used to indicate the spatial position of the frame block corresponding to the second projection vector in the video frame. The second projection vector, the second block position information, and the temporal position information are encoded to obtain the signal features corresponding to the second sub-signal.

[0134] If the control signal includes object motion information, and the object motion information is in coordinate format, the object motion information is input into the motion extractor, and the motion features corresponding to the object motion information are output; or, if the control signal includes object motion information, and the object motion information is in video format, the object motion information is input into the video feature extractor, and the motion features of the object motion information are output.

[0135] If the control signal includes camera control information, the camera control information is input into the camera feature extractor, which outputs the camera operation features of the camera control information.

[0136] The signal features of multiple sub-signals are arranged in a preset order, and feature markers corresponding to the signal features are set; wherein, the feature markers are used to indicate the start and end positions of the sub-signals and / or signal features corresponding to the signal features.

[0137] If the control signal includes a text instruction and a third sub-signal other than the text instruction, the signal features of the third sub-signal are converted into a preset text feature space to obtain the signal features of the converted third sub-signal; wherein, the text feature space is the feature space where the signal features of the text instruction are located.

[0138] By using a pre-defined multi-head attention mechanism, feature correlation information between the signal features of any two sub-signals is extracted.

[0139] The signal features with feature association information are fused to obtain the fused signal features.

[0140] The fused signal features are input into a preset prediction model, which then outputs the description content corresponding to the preset text description label.

[0141] The feature extractor is trained as follows: a sample sub-signal is input into the feature extractor, and the sample features corresponding to the sample sub-signal are output; wherein, a sample text description is set for the sample sub-signal; the sample features are input into the spatial transformation model, and the text features corresponding to the sample features are output; a first loss value is determined based on the text features and the sample text description, and the parameters in the feature extractor and / or the spatial transformation model are adjusted based on the first loss value.

[0142] The signal features of multiple sub-signals are input into a semantic parsing model. The semantic parsing model extracts the feature association information between the signal features of the multiple sub-signals. Based on the signal features and feature association information, the description content corresponding to the preset text description label is generated. The semantic parsing model is trained as follows: the sample sub-signals in the sample control signal are input into the feature extractor, and the sample signal features corresponding to the sample sub-signals are output. The sample control signal is set with sample description content. The sample signal features are input into the semantic parsing model, and the predicted description content is output. A second loss value is determined based on the predicted description content and the sample description content. The parameters in the semantic parsing model are adjusted based on the second loss value.

[0143] The sample control signals are pre-divided into multiple signal groups. The first group of signal groups includes any sample sub-signal except for text commands. The second group includes some sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. The third group includes sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. These signal groups are sequentially input to a feature extractor, and the output sample feature signals are then input to a semantic parsing model for training.

[0144] Multiple signal groups are assigned corresponding loss weights, which are used to determine the second loss value for each signal group.

[0145] According to a preset probability, at least a portion of the text instructions in the sample control signal are deleted.

[0146] In the above method, after extracting signal features and feature relationships from multiple sub-signals, text description labels and corresponding description content are generated. These text description labels and corresponding description content are structured text information that can clearly and comprehensively express the user's intent. The video generation model can more easily understand and follow the instructions to generate videos, resulting in a high degree of matching between the generated videos and the user's intent. This improves the controllability and accuracy of video generation and enhances the quality of video generation.

[0147] This embodiment also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above-described video generation method.

[0148] The computer-executable instructions stored in the aforementioned computer-readable storage medium can, by executing the aforementioned computer-executable instructions, perform the following operations in the aforementioned video generation method:

[0149] A video generation method includes: acquiring control signals; wherein the control signals include multiple sub-signals from text instructions, spatial structure information, object action information, object appearance information, and camera control information; extracting signal features corresponding to the sub-signals, as well as feature association information between the signal features of the multiple sub-signals; wherein the feature association information is used to indicate the semantic association and / or constraint relationship between the signal features of the multiple sub-signals; generating description content corresponding to preset text description tags based on the signal features and feature association information; wherein the text description tags include multiple types from global scene description, object description, background description, camera description, style description, and behavior description; and inputting the text description tags and description content into a preset video generation model to output a target video.

[0150] Multiple sub-signals are input into the feature extractors corresponding to the sub-signals respectively, and the signal features corresponding to the sub-signals are output.

[0151] If the control signal includes a first sub-signal in image format, the first sub-signal is input into the image feature extractor; the image feature extractor divides the first sub-signal into multiple image blocks, generates a first projection vector corresponding to each image block, and sets the first block position information corresponding to the first projection vector; wherein, the first block position information is used to: indicate the spatial position of the image block corresponding to the first projection vector in the first sub-signal; the first projection vector and the first block position information are encoded to obtain the signal features corresponding to the first sub-signal.

[0152] If the control signal includes a second sub-signal in video format, the second sub-signal is input into the video feature extractor. The video feature extractor divides the video frames in the second sub-signal into multiple frame blocks, generates a second projection vector corresponding to each frame block, and sets the second block position information corresponding to the second projection vector and the temporal position information of the video frame to which the frame block belongs. The second block position information is used to indicate the spatial position of the frame block corresponding to the second projection vector in the video frame. The second projection vector, the second block position information, and the temporal position information are encoded to obtain the signal features corresponding to the second sub-signal.

[0153] If the control signal includes object motion information, and the object motion information is in coordinate format, the object motion information is input into the motion extractor, and the motion features corresponding to the object motion information are output; or, if the control signal includes object motion information, and the object motion information is in video format, the object motion information is input into the video feature extractor, and the motion features of the object motion information are output.

[0154] If the control signal includes camera control information, the camera control information is input into the camera feature extractor, which outputs the camera operation features of the camera control information.

[0155] The signal features of multiple sub-signals are arranged in a preset order, and feature markers corresponding to the signal features are set; wherein, the feature markers are used to indicate the start and end positions of the sub-signals and / or signal features corresponding to the signal features.

[0156] If the control signal includes a text instruction and a third sub-signal other than the text instruction, the signal features of the third sub-signal are converted into a preset text feature space to obtain the signal features of the converted third sub-signal; wherein, the text feature space is the feature space where the signal features of the text instruction are located.

[0157] By using a pre-defined multi-head attention mechanism, feature correlation information between the signal features of any two sub-signals is extracted.

[0158] The signal features with feature association information are fused to obtain the fused signal features.

[0159] The fused signal features are input into a preset prediction model, which then outputs the description content corresponding to the preset text description label.

[0160] The feature extractor is trained as follows: a sample sub-signal is input into the feature extractor, and the sample features corresponding to the sample sub-signal are output; wherein, a sample text description is set for the sample sub-signal; the sample features are input into the spatial transformation model, and the text features corresponding to the sample features are output; a first loss value is determined based on the text features and the sample text description, and the parameters in the feature extractor and / or the spatial transformation model are adjusted based on the first loss value.

[0161] The signal features of multiple sub-signals are input into a semantic parsing model. The semantic parsing model extracts the feature association information between the signal features of the multiple sub-signals. Based on the signal features and feature association information, the description content corresponding to the preset text description label is generated. The semantic parsing model is trained as follows: the sample sub-signals in the sample control signal are input into the feature extractor, and the sample signal features corresponding to the sample sub-signals are output. The sample control signal is set with sample description content. The sample signal features are input into the semantic parsing model, and the predicted description content is output. A second loss value is determined based on the predicted description content and the sample description content. The parameters in the semantic parsing model are adjusted based on the second loss value.

[0162] The sample control signals are pre-divided into multiple signal groups. The first group of signal groups includes any sample sub-signal except for text commands. The second group includes some sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. The third group includes sample sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information. These signal groups are sequentially input to a feature extractor, and the output sample feature signals are then input to a semantic parsing model for training.

[0163] Multiple signal groups are assigned corresponding loss weights, which are used to determine the second loss value for each signal group.

[0164] According to a preset probability, at least a portion of the text instructions in the sample control signal are deleted.

[0165] In the above method, after extracting signal features and feature relationships from multiple sub-signals, text description labels and corresponding description content are generated. These text description labels and corresponding description content are structured text information that can clearly and comprehensively express the user's intent. The video generation model can more easily understand and follow the instructions to generate videos, resulting in a high degree of matching between the generated videos and the user's intent. This improves the controllability and accuracy of video generation and enhances the quality of video generation.

[0166] The computer program products of the video generation method, apparatus, and electronic device provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0167] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0168] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0169] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0170] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0171] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A video generation method, characterized in that, The method includes: Acquire control signals; wherein, the control signals include multiple sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information; Extract the signal features corresponding to the sub-signal, as well as the feature association information between the signal features of the various sub-signals; wherein, the feature association information is used to indicate the semantic association relationship and / or constraint relationship between the signal features of the various sub-signals; Based on the signal features and the feature association information, a description content corresponding to a preset text description tag is generated; wherein, the text description tag includes multiple of the following: global scene description, object description, background description, camera description, style description, and behavior description; The text description tag and the description content are input into a preset video generation model to output the target video.

2. The method according to claim 1, characterized in that, The step of extracting the signal features corresponding to the sub-signal includes: Each of the sub-signals is input into a feature extractor corresponding to the sub-signal, and the signal features corresponding to the sub-signal are output.

3. The method according to claim 2, characterized in that, The steps of inputting multiple sub-signals into feature extractors corresponding to the sub-signals respectively, and outputting signal features corresponding to the sub-signals, include: If the control signal includes a first sub-signal in image format, the first sub-signal is input to the image feature extractor; The image feature extractor divides the first sub-signal into multiple image blocks, generates a first projection vector corresponding to each image block, and sets the first block position information corresponding to the first projection vector; wherein, the first block position information is used to: indicate the spatial position of the image block corresponding to the first projection vector in the first sub-signal. The first projection vector and the first block location information are encoded to obtain the signal features corresponding to the first sub-signal.

4. The method according to claim 2, characterized in that, The steps of inputting multiple sub-signals into feature extractors corresponding to the sub-signals respectively, and outputting signal features corresponding to the sub-signals, include: If the control signal includes a second sub-signal in video format, the second sub-signal is input to the video feature extractor; The video feature extractor divides the video frame in the second sub-signal into multiple frame blocks, generates a second projection vector corresponding to the frame block, and sets the second block position information corresponding to the second projection vector and the time position information of the video frame to which the frame block belongs; wherein, the second block position information is used to indicate the spatial position of the frame block corresponding to the second projection vector in the video frame. The second projection vector, the second block location information, and the time location information are encoded to obtain the signal features corresponding to the second sub-signal.

5. The method according to claim 2, characterized in that, The steps of inputting multiple sub-signals into feature extractors corresponding to the sub-signals respectively, and outputting signal features corresponding to the sub-signals, include: If the control signal includes object motion information, and the object motion information is in coordinate format, the object motion information is input into the motion extractor, and the motion features corresponding to the object motion information are output. Alternatively, if the control signal includes object motion information, and the object motion information is in video format, the object motion information is input into a video feature extractor, and the motion features of the object motion information are output.

6. The method according to claim 2, characterized in that, The steps of inputting multiple sub-signals into feature extractors corresponding to the sub-signals respectively, and outputting signal features corresponding to the sub-signals, include: If the control signal includes camera control information, the camera control information is input to the camera feature extractor, and the camera operation features of the camera control information are output.

7. The method according to claim 1, characterized in that, Before the step of extracting feature correlation information among the signal features of the various sub-signals, the method further includes: The signal features of the various sub-signals are arranged in a preset order, and feature markers corresponding to the signal features are set; wherein, the feature markers are used to indicate: the start and end positions of the sub-signals corresponding to the signal features and / or the signal features themselves.

8. The method according to claim 1, characterized in that, Before the step of extracting feature correlation information among the signal features of the various sub-signals, the method further includes: If the control signal includes a text instruction and a third sub-signal other than the text instruction, the signal features of the third sub-signal are converted into a preset text feature space to obtain the converted signal features of the third sub-signal; wherein, the text feature space is the feature space where the signal features of the text instruction are located.

9. The method according to claim 1, characterized in that, The step of extracting feature correlation information among signal features of multiple sub-signals includes: By using a pre-defined multi-head attention mechanism, feature correlation information between the signal features of any two sub-signals is extracted.

10. The method according to claim 1, characterized in that, After the step of extracting feature correlation information among the signal features of the various sub-signals, the method further includes: The signal features with the aforementioned feature association information are fused to obtain the fused signal features.

11. The method according to claim 1, characterized in that, The step of generating description content corresponding to a preset text description tag based on the signal features and the feature association information includes: The fused signal features are input into a preset prediction model, and the prediction model outputs the description content corresponding to the preset text description label.

12. The method according to claim 2, characterized in that, The feature extractor is trained in the following manner: The sample sub-signal is input to the feature extractor, and the sample features corresponding to the sample sub-signal are output; wherein, the sample sub-signal is associated with a sample text description. The sample features are input into the spatial transformation model, and the text features corresponding to the sample features are output. A first loss value is determined based on the text features and the sample text description, and the parameters in the feature extractor and / or the spatial transformation model are adjusted based on the first loss value.

13. The method according to claim 1, characterized in that, The step of extracting feature association information among signal features of multiple sub-signals, and generating description content corresponding to preset text description tags based on the signal features and the feature association information, includes: The signal features of multiple sub-signals are input into a semantic parsing model. The semantic parsing model extracts the feature association information between the signal features of multiple sub-signals. Based on the signal features and the feature association information, the description content corresponding to the preset text description tag is generated. The semantic parsing model is trained in the following manner: The sample sub-signal in the sample control signal is input into the feature extractor, and the sample signal feature corresponding to the sample sub-signal is output; wherein, the sample control signal is configured with sample description content. The sample signal features are input into the semantic parsing model, and the predicted description content is output. A second loss value is determined based on the predicted description content and the sample description content, and the parameters in the semantic parsing model are adjusted based on the second loss value.

14. The method according to claim 13, characterized in that, The sample control signal is pre-divided into multiple signal groups; Among them, the sample control signal in the first group of the multiple signal groups includes: any sample sub-signal other than text instructions; The sample control signals in the second group of the multiple signal groups include: partial sample sub-signals from text commands, spatial structure information, object motion information, object appearance information, and camera control information; The sample control signals in the third group of the multiple signal groups include: sample sub-signals of text commands, spatial structure information, object motion information, object appearance information, and camera control information; Multiple sets of signals are sequentially input into the feature extractor, and the output sample feature signals are then input into the semantic parsing model to train the semantic parsing model.

15. The method according to claim 14, characterized in that, Each of the signal groups is assigned a loss weight, which is used to determine the second loss value corresponding to the signal group.

16. The method according to claim 13, characterized in that, Before the step of inputting the sample sub-signal from the sample control signal to the feature extractor and outputting the sample signal features corresponding to the sample sub-signal, the method further includes: At least a portion of the text instructions in the sample control signal are deleted according to a preset probability.

17. A video generation apparatus, characterized in that, The device includes: The signal acquisition module is used to acquire control signals; wherein, the control signals include multiple sub-signals from text commands, spatial structure information, object action information, object appearance information, and camera control information; The feature extraction module is used to extract the signal features corresponding to the sub-signal, as well as the feature association information between the signal features of the various sub-signals; wherein, the feature association information is used to indicate the semantic association relationship and / or constraint relationship between the signal features of the various sub-signals; The content generation module is used to generate description content corresponding to preset text description tags based on the signal features and the feature association information; wherein, the text description tags include multiple types of global scene description, object description, background description, camera description, style description and behavior description; The video output module is used to input the text description tags and the description content into a preset video generation model and output the target video.

18. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the video generation method according to any one of claims 1-16.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the video generation method according to any one of claims 1-16.