Video generation, method of training a model and system
By jointly training a semantic feature generation network and a video generation network, and combining video description information and semantic extension instructions, the problem of fluctuating video generation quality was solved, and high-quality and stable video generation was achieved.
Patent Information
- Application Number
- CN202411702267.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-11-25
AI Technical Summary
In existing video generation technologies, video quality is easily affected by fluctuations in the quality of the video description information input by the user, resulting in unstable video quality.
By jointly training a semantic feature generation network and a video generation network, and combining video description information and N semantic extension instructions, target semantic features are generated, thereby achieving stability and high-quality output of the video generation model.
It improves the performance and generation effect of the video generation model, ensures the richness of the generated video content and the stability of its quality, and reduces the dependence on the quality of user input description information.
Smart Images

Figure CN119653201B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a video generation and model training method and system. Background Technology
[0002] Text-to-video technology can convert text into video content and is widely used in advertising, content creation, education, and product demonstrations. Based on user-input descriptive text, this technology can generate videos with specific styles and durations, significantly lowering the barrier to entry and cost of video production, and improving the efficiency and diversity of content creation. However, the quality of this video generation method largely depends on the user-input descriptive information, resulting in significant fluctuations in output video quality and an inability to guarantee consistent quality.
[0003] Therefore, there is a need to provide a method that can stably output high-quality video.
[0004] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the filing date of this specification, nor does it imply that it can be considered prior art in this specification. Summary of the Invention
[0005] This specification provides a method and system for video generation and model training, which can stably generate high-quality videos based on video description information.
[0006] In a first aspect, this specification provides a video generation method, comprising: obtaining video description information and N semantic expansion instructions, wherein N is an integer greater than or equal to 1; obtaining a pre-trained video generation model, wherein the video generation model includes a semantic feature generation network and a video generation network, wherein, during the training process of the video generation model, the semantic feature generation network and the video generation network are jointly trained; extracting target semantic features from the video description information and the N semantic expansion instructions through the semantic feature generation network, and generating a video based on the target semantic features through the video generation network to obtain a target video that semantically matches the video description information, wherein the target semantic features represent the semantic features after semantic expansion of the video description information through the N semantic expansion instructions; and outputting the target video.
[0007] Secondly, this specification also provides a training method for a video generation model. The method includes: obtaining a sample set; the sample set includes multiple sample pairs, each sample pair including a sample video and sample description information; obtaining a video generation model to be trained, the video generation model including a semantic feature generation network and a video generation network; and performing multiple iterative training on the video generation model using the sample set to obtain a trained video generation model, wherein each iteration includes: extracting semantic features from the sample description information and N semantic extension instructions through the semantic feature generation network to obtain sample semantic features, the sample semantic features representing the semantic features after semantic extension of the sample description information through the N semantic extension instructions; generating a predicted video based on the sample semantic features through the video generation network, and updating the parameters of the semantic feature generation network and the video generation network with the training objective of minimizing the difference between the predicted video and the sample video.
[0008] Thirdly, this specification also provides a video generation system, including at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for performing a video generation method; the at least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and executes the video generation method described in any of the first aspects according to the instructions of the at least one instruction set.
[0009] Fourthly, this specification also provides a training system for a video generation model, including at least one storage medium and at least one processor, wherein the at least one storage medium stores at least one instruction set for training a video generation model; the at least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set during operation and executes the training method for the video generation model according to the instructions of the at least one instruction set: any of the above-mentioned second aspects.
[0010] Fifthly, this specification also provides a computer-readable non-volatile storage medium, wherein the computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the method as described in any one of the first or second aspects above.
[0011] As can be seen from the above technical solutions, the video generation method provided in this specification acquires video description information and N semantic extension instructions, and obtains a pre-trained video generation model. The semantic feature generation network and the video generation network included in the video generation model are jointly trained during the training process, thereby achieving balanced learning of visual and linguistic elements within a unified framework. The jointly trained semantic feature generation network can generate target semantic features that better meet the needs of the video generation network, and the jointly trained video generation network can more accurately generate high-quality target videos based on the target semantic features, thus improving the overall performance of the video generation model and the video generation effect. Furthermore, the target semantic features in this specification are comprehensively determined based on video description information and N semantic extension instructions. By introducing N semantic extension instructions to supplement and expand the video description information, comprehensive and detailed feature acquisition at the semantic level is achieved. Therefore, the feature information contained in the target semantic features is not limited to the user-input video description information, but can integrate more semantic elements contained in the semantic extension instructions, improving the richness and accuracy of the target semantic feature expression. The target video generated based on such target semantic features has richer video content and more stable video content quality.
[0012] The video generation, model training methods, and other system functions provided in this specification will be partially listed in the following description. The inventive aspects of the video generation, model training methods, and system provided in this specification can be fully explained through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A schematic diagram of an application scenario provided according to an embodiment of this specification is shown;
[0015] Figure 2 A schematic diagram of the hardware structure of a computing device provided according to some embodiments of this specification is shown;
[0016] Figure 3 A flowchart illustrating a training method for a video generation model according to an embodiment of this specification is shown.
[0017] Figure 4A flowchart illustrating a training method for a video generation model according to another embodiment of this specification is shown.
[0018] Figure 5 A flowchart illustrating a process for determining the semantic features of a sample according to an embodiment of this specification is shown.
[0019] Figure 6 A flowchart illustrating a video generation method according to an embodiment of this specification is shown; and
[0020] Figure 7 A flowchart illustrating a video generation method according to another embodiment of this specification is shown. Detailed Implementation
[0021] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.
[0022] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.
[0023] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0024] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0025] In this specification, "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X can include only one of A, B, and C, or any combination of A, B, and C, as well as other possible content / elements. Any combination of A, B, and C can be A, B, C, AB, AC, BC, or ABC.
[0026] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0027] It should be noted that the user data obtained in this manual is authorized by the user and does not involve user privacy.
[0028] For ease of description, the terms that will appear later in this manual will be explained first.
[0029] Large Language Models (LLMs): These are commonly used in the field of artificial intelligence in Natural Language Processing (NLP), specifically referring to large machine learning models with a large number of parameters and computational resources. Large models are named for their huge number of parameters and complex network structures. They have powerful feature representation and feature understanding capabilities, and are better able to capture patterns and regularities in data when dealing with complex tasks. They are designed and trained to better understand and generate natural language.
[0030] Diffusion models are a class of generative networks that excel in image generation, audio synthesis, and video generation. They aim to generate new data samples by learning the distribution of data. Commonly used in image generation, they can produce extremely realistic and high-quality images with rich detail and natural textures. The working principle of a diffusion network includes a forward process and a reverse process. The forward process, also known as the noise addition process, gradually transforms the input data into noise. The reverse process, also known as the denoising process, gradually recovers the input data from the noise.
[0031] In text-based video generation (generating videos based on descriptive text), the user typically inputs a video description, and the video generation model then generates a video corresponding to that description. However, this method suffers from a high degree of correlation between the generated video quality and the user-inputted description. The more descriptive information the user provides, the better the quality of the generated video; conversely, the less descriptive information, the worse the quality.
[0032] For example, if a user inputs a video description like, "A woman with her family is driving a red car along a vast road, with a river flowing beside it and an airplane flying overhead," this description is clearly rich, and the video generated based on it will also be rich in content. However, if the user inputs a video description like, "A car is driving on the road," this description is obviously too simplistic. The video generation model cannot extract more information from this description, and the generated video will also be monotonous.
[0033] In other words, in existing technologies, the quality of the generated video content primarily depends on the quality of the video description information input by the user. The more detailed and richer the user-input video description information, the better the video quality. Conversely, the less detailed the user-input video description information, the worse the video quality. This method of video generation results in significant fluctuations in the quality of the output video content, leading to instability in the generated video quality.
[0034] The video generation method provided in this specification obtains a pre-trained video generation model. The semantic feature generation network and the video generation network included in the pre-trained video generation model are jointly trained during the training of the video generation model. Therefore, the semantic feature generation network can generate target semantic features that better match the needs of the video generation network, and the video generation network can more accurately generate high-quality target videos based on the target semantic features, improving the overall performance of the video generation model and the generation effect of the target video. Furthermore, after obtaining video description information and N semantic extension instructions, the semantic feature generation network can perform semantic extraction on the video description information and the N semantic extension instructions to obtain target semantic features, and then the video generation network can generate the target video based on these target semantic features. This approach supplements the target semantic features through the N semantic extension instructions. Since the target semantic features are determined based on the video description information and the N semantic extension instructions, even simple video description information can be semantically extended through the semantic extension instructions to uncover more potential details and elements, resulting in more comprehensive and richer target semantic features. Consequently, the content of the target video generated based on the target semantic features will be richer. This setup provides users with a more reliable video generation method, eliminating concerns about the richness of the video description text. For different users inputting video description text of varying quality, the video generation method provided in this manual consistently yields a target video with relatively guaranteed content quality.
[0035] It should be noted that the above description of application scenarios is only one of the many use cases provided in this specification. Those skilled in the art should understand that when the video generation, model training methods, and systems provided in this specification are applied to other use cases, their implementation methods and technical effects are similar.
[0036] Figure 1 A schematic diagram of an application scenario 100 provided according to an embodiment of this specification is shown.
[0037] like Figure 1 As shown, the application scenario 100 may include a training system 130 (hereinafter referred to as training system 130) for a video generation model and a video generation system 150. The application scenario 100 involves two execution phases.
[0038] In the first stage, the training system 130 can iteratively train the video generation model to be trained based on multiple sample pairs in the sample set, in order to optimize the parameters of the semantic feature generation network and the video generation network included in the video generation model. Each sample pair contains sample video and sample description information. The trained video generation model is deployed in the video generation system 150.
[0039] In the second stage, after obtaining video description information and N semantic extension instructions, the video generation system 150 can extract the semantics of the video description information and N semantic extension instructions based on the semantic feature generation network in the video generation model to obtain target semantic features. Then, the video generation network in the video generation model generates a target video that semantically matches the video description information and outputs it.
[0040] The training system 130 may be a computing system with a certain computing capability. The training system 130 can execute the training method of the video generation model provided in this specification, thereby training the video generation model to be trained based on a sample set, and obtaining the trained video generation model. The training system 130 may store data or instructions for executing the training method of the video generation model described in this specification, and may execute or be used to execute the data or instructions. The training system 130 may include hardware devices with data information processing capabilities and the necessary programs required to drive the hardware devices.
[0041] The training system 130 can be a single computing device or a cluster system composed of multiple computing devices. The data or instructions stored in the training system 130 for executing the training method of the video generation model described in this specification can adopt any form of system architecture. For example, layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc.
[0042] In some embodiments, the training system 130 may first obtain a sample set comprising multiple sample pairs. Each sample pair includes a sample video and sample description information. Then, the training system 130 iteratively trains the video generation model based on the aforementioned sample set to obtain a trained video generation model. During each iteration, the semantic feature generation network in the video generation model extracts semantic features from the sample description information and N semantic extension instructions. The video generation network in the video generation model then generates a predicted video based on these sample semantic features. Finally, the parameters of the semantic feature generation network and the video generation network are updated with the training objective of minimizing the difference between the predicted video and the sample video.
[0043] The video generation model trained using the aforementioned training system 130 achieves a level of semantic feature generation network within the video generation model where the differences between semantic features obtained from samples fluctuate within a preset range when extracting semantic information from different sample descriptions. Furthermore, the video generation network in the video generation model can stably output a predicted video that minimizes the difference from the sample video based on the sample semantic features.
[0044] The video generation system 150 can be a computing system with a certain computing capability. The video generation system 150 can execute the video generation method provided in this specification, thereby generating and outputting a target video based on the obtained video description information and N semantic extension instructions. The video generation system 150 can store data or instructions for executing the video generation method described in this specification, and can execute or be used to execute the data or instructions. The video generation system 150 may include hardware devices with data information processing capabilities and the necessary programs required to drive the hardware devices.
[0045] The video generation system 150 can be a single computing device or a cluster system composed of multiple computing devices. The data or instructions stored in the video generation system 150 for executing the video generation method described in this specification can adopt any form of system architecture. For example, layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc.
[0046] In some embodiments, the video generation system 150 may first obtain a pre-trained video generation model, video description information, and N semantic extension instructions. Then, the video generation system 150 may use the semantic feature generation network in the aforementioned video generation model to perform semantic extraction on the video description information and the N semantic extension instructions, obtaining target semantic features. The video generation system 150 may then use the video generation network in the aforementioned video generation model to generate a video based on the target semantic features, obtaining a target video that semantically matches the video description information. Finally, the video generation system 150 may output the target video.
[0047] The above method ensures that the target video output by the video generation system 150 is semantically consistent with the video description information. Compared to methods that directly determine and output the video based on the video description information, this method ensures that the differences between the target semantic features output by the semantic feature generation network for different video description information fluctuate within a preset range. In other words, the differences between the target semantic features corresponding to richly descriptive video descriptions and those with simple descriptions are within a preset range. This allows even very simple video descriptions to output target semantic features containing rich features, greatly reducing the impact of video description information on the quality of the generated target video, thereby stabilizing the performance of the video generation model and, consequently, the quality of the video generated by the model.
[0048] The training system 130 and the video generation system 150 may correspond to the same computing system or to different computing systems; this specification does not limit this.
[0049] Figure 2A schematic diagram of the hardware structure of a computing device 200 according to some embodiments of this specification is shown. The training system 130 and the video generation system 150 may have, for example... Figure 2 The structure of the computing device 200 shown.
[0050] like Figure 2 As shown, the computing device 200 includes at least one storage medium 230 and at least one processor 220. In some embodiments, the computing device 200 may further include an internal communication bus 210. In some embodiments, the computing device 200 may further include a communication port 250. In some embodiments, the computing device 200 may further include I / O components 260.
[0051] The internal communication bus 210 can connect different system components, including storage medium 230 and processor 220. I / O component 260 supports input / output between computing device 200 and other components.
[0052] Communication port 250 is used for data communication between computing device 200 and the outside world. For example, computing device 200 can connect to a network through communication port 250.
[0053] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 236. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set is computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc., that execute the training methods for the video generation methods or video generation models provided in this specification.
[0054] At least one processor 220 is communicatively connected to at least one storage medium 230 via an internal communication bus 210. The at least one processor 220 is used to execute at least one instruction set. When the system 130 is running, the at least one processor 220 reads at least one instruction set and executes the video generation method or video generation model training method provided in this specification according to the instructions of the at least one instruction set.
[0055] Processor 220 can execute all the steps included in the training method of the video generation method or video generation model. Processor 220 can be in the form of one or more processors. Processor 220 can issue execution instructions. Processor 220 may include one or more hardware processors, such as microcontrollers, microprocessors, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), central processing units (CPUs), graphics processing units (GPUs), physical processing units (PPUs), microcontroller units, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), advanced RISC machines (ARMs), programmable logic devices (PLDs), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0056] For illustrative purposes only, only one processor 220 is shown in the accompanying drawings of the computing device 200. However, it should be noted that the computing device 200 may also include multiple processors. Therefore, the operation and / or method steps disclosed herein may be executed by a single processor or by multiple processors in combination, as described herein. For example, if processor 220 of the computing device 200 in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., a first processor executes step A, a second processor executes step B, or the first and second processors jointly execute steps A and B).
[0057] Figure 3 A flowchart illustrating a training method for a video generation model according to an embodiment of this specification is shown. Figure 4 A flowchart illustrating a training method for a video generation model according to another embodiment of this specification is shown; the training method P300 for this video generation model can be executed by the training system 130. Figure 3 and Figure 4 As shown, the method P300 provided in this specification may include S310-S350, wherein:
[0058] S310: Obtain the sample set; the sample set includes multiple sample pairs, each sample pair including sample video and sample description information.
[0059] To improve the performance of the video generation model, sample diversity needs to be ensured during training. Therefore, the sample set can include a large number of sample videos of different types and scenes. For example, the sample set can include: landscape videos, people videos, animal videos, videos of different styles (cartoon style, oil painting style, retro style, realistic style, etc.), architectural videos, etc. Sample description information is used to describe the content of the sample videos. Sample description information can be simple descriptions of the sample videos, such as "autumn scenery"; "children flying kites"; "fishing," or it can be rich and complex descriptions of the sample videos, such as: "Generating oil painting style yellow maple leaves slowly falling on the asphalt road, with some pedestrians and cyclists nearby"; "A five-year-old boy and his parents are flying a swallow kite in the park, a happy family of three"; "Two kittens are fishing, one kitten is very serious, and the other kitten is busy catching butterflies." It should be understood that the above embodiments are only for rational illustration. The types of sample videos included in the specific sample set, as well as the types of sample description information, can be flexibly adjusted according to user needs or actual applications, and are not limited to those given in the above embodiments.
[0060] S330: Obtain the video generation model to be trained, which includes a semantic feature generation network and a video generation network.
[0061] In the embodiments of this specification, the semantic feature generation network and the video generation network are incorporated as part of the video generation model. These two networks are jointly trained as the video model iterates, achieving balanced learning of visual and linguistic elements within a unified framework. This optimizes the video generation model holistically, avoiding compatibility and performance degradation issues that might arise from training the semantic feature generation network and video generation network separately and then combining them. During joint training, the parameters of both the video generation network and the semantic feature generation network can be continuously adjusted along with the iterations of the video generation model. This allows the video generation network and the semantic feature generation network to better adapt to each other, thereby improving the performance of the video generation model when faced with various video descriptions and video generation tasks. The jointly trained semantic feature generation network and video generation network can then cooperate and work collaboratively.
[0062] S350: The video generation model is trained iteratively using a sample set to obtain the trained video generation model.
[0063] In each iteration, the training system 130 performs similar steps based on each sample pair. The following describes the specific execution of S351-S353 by the training system 130 based on sample pair 1, taking sample pair 1 as an example (sample pair 1 includes sample video 1 and sample description information 1). The steps performed by the training system 130 based on other sample pairs besides sample pair 1 are similar and will not be repeated here.
[0064] S351: The semantic features of the sample are obtained by semantically extracting the sample description information 1 and N semantic extension instructions through the semantic feature generation network. The sample semantic features represent the semantic features after semantic extension of the sample description information 1 through N semantic extension instructions.
[0065] In the embodiments of this specification, N semantic expansion instructions are used to provide a series of instruction prompts. Semantic expansion instructions are pre-defined feature information with specific directional or guiding functions, serving as a reference benchmark to provide an initial guiding direction for subsequent processing. For example, semantic expansion instructions may include any form such as: "Please describe the type of motion in the video," "Please describe the relationships between the characters in the video," "Please describe the theme color in the video," "Please describe the environmental conditions in the video," "Please describe the weather conditions in the video," etc., and are not limited to those given in the above embodiments.
[0066] Figure 5 A flowchart illustrating a process for determining sample semantic features according to an embodiment of this specification is shown, such as... Figure 5 As shown, step S351 also includes:
[0067] S3511: Use a large model to extract semantic information from sample description information 1 to obtain the semantic features of the first sample.
[0068] In the embodiments of this specification, the process of obtaining the first sample semantic features can be as follows: A large model is used to semantically extract the sample description information 1 from multiple dimensions, obtaining sample semantic sub-features corresponding to multiple dimensions. Extracting sample semantic sub-features from multiple dimensions allows the extracted sample semantic sub-features to cover multiple aspects of the sample description information. Compared to single-dimensional feature extraction, the sample semantic sub-features extracted by the above-mentioned multi-dimensional extraction method are richer and more comprehensive. Subsequently, the sample semantic sub-features corresponding to multiple dimensions are concatenated to obtain concatenated sample features. Then, the concatenated sample features are convolved to obtain the first sample semantic features.
[0069] By convolutional processing of concatenated features, it is possible to better extract, fuse, and optimize semantic sub-features corresponding to multiple dimensions of the sample. This allows for more thorough feature interactions between semantic sub-features across multiple dimensions, resulting in more comprehensive, stable, and representative first-sample semantic features. Furthermore, concatenated features may have high dimensionality, while convolutional processing can map high-dimensional features to a relatively low-dimensional space. This leads to more compact and easier-to-understand first-sample semantic features, which helps reduce subsequent computation and improve model efficiency.
[0070] In the embodiments of this specification, the video generation model further includes convolutional layers. These convolutional layers process the concatenated features of the samples to obtain the semantic features of the first sample. To reduce the impact of the convolutional layers on the backbone networks (semantic feature generation network and video generation network) and improve the model's convergence speed, the initial state of the convolutional layers can be set to 0, and the parameters of the convolutional layers gradually stabilize as the model is trained. During the training process of the video generation model, the convolutional layers, semantic feature generation network, and video generation network are jointly trained to continuously adjust the parameters of the convolutional layers, which remain unchanged after training.
[0071] Each dimension's sample semantic sub-features can include question features and answer features. Question features are generated by the large model based on sample description information. In other words, the large model directly extracts feature information that reflects the basic situation and key elements of the video from the sample description information; this information constitutes the question features corresponding to the sample description information. Question features can be understood as an extraction of the inherent features of the sample description information itself, reflecting the demand information within the sample description information.
[0072] The answer features are determined by the large model using an autoregressive algorithm. When processing sample description information, the large model continuously predicts the next possible word based on the existing text sequence within the sample description. By repeatedly performing this word prediction operation and collecting relevant feature information during the prediction process, the answer features corresponding to the sample description information are ultimately obtained. In other words, the answer features are features captured by the large model during the dynamic generation of subsequent text based on the sample description information.
[0073] For example, when the sample description information is "cat fishing", the possible question features and answer features are as follows: The above sample description information may include core description dimension, content detail dimension and value dimension, and each dimension has its corresponding question features and answer features.
[0074] For example, the question feature corresponding to the core description dimension might include: "What is the main story of the kitten fishing?" This question feature reflects the need for information on the story's outline in the video, clearly pointing to the core description of the story "The Kitten Fishing." The corresponding answer feature might be: "The kitten and its mother went fishing by the river. At first, the kitten was distracted, catching butterflies and dragonflies, and in the end, it didn't catch a single fish. Later, the kitten listened to its mother and focused on fishing, finally catching a fish." This answer feature provides the complete story content to satisfy the question feature's need for an outline.
[0075] Regarding the detail dimension, question features might include: "What animals did the kitten encounter while fishing?" This indicates interest in specific plot points and characters, reflecting a need to uncover story details. A corresponding answer feature might be: "The kitten encountered butterflies and dragonflies while fishing. It put down its fishing rod to chase them, which delayed its fishing." This answer feature addresses the question's focus on detail by listing specific animals and the kitten's interactions with them.
[0076] Regarding the value dimension, question features might include: "What inspiration can children draw from the story of the kitten fishing?" This reflects a focus on the educational significance of the story, representing a question feature that moves from content comprehension to value extraction. The corresponding answer feature might be: "This story teaches children that they need to concentrate on what they do and not be distracted, just like the kitten initially couldn't catch fish because it wasn't focused, but later it was able to catch fish after concentrating." This answer feature extracts educational value from the story content to meet the question feature's need for inspiration. It should be understood that the above examples are merely illustrative. The specific circumstances of multiple dimensions, and the question and answer features that may be included under each dimension, are determined based on the specific details of the video description information and the knowledge reserves of the large model, and are not limited to those given in the above examples.
[0077] S3513: Use a large model to extract semantics from N semantic extension instructions to obtain the second sample semantic features corresponding to the N semantic extension instructions.
[0078] Semantic extraction of semantic extension instructions is similar to text analysis, extracting key semantic information from the instructions. The resulting second sample semantic features focus on reflecting the semantic characteristics of the semantic extension instructions themselves, including the instruction's goal (e.g., extracting certain information), the objects involved (e.g., people or scenes in a video), and the actions (e.g., analysis or recognition).
[0079] Taking the semantic extension instruction "analyze the dialogue topic between characters in the video" as an example, the large model will identify key elements such as "video," "characters," "dialogue topic," and "analysis." The large model may combine these key elements into a second semantic feature that can represent the semantic extension instruction through word vector representation or other semantic representation methods.
[0080] S3515: Perform feature fusion on the semantic features of the first sample and the semantic features of the second sample corresponding to N semantic extension instructions to obtain the sample semantic features.
[0081] The first sample semantic features are extracted from the sample description information, covering the basic content of the predicted video that the user expects to generate, such as the semantics of scenes, people, and events. The second sample semantic features come from semantic extension instructions, which contain specific processing or mining requirements for the predicted video, such as extracting emotional details and identifying specific objects. By fusing the first and second sample semantic features, the basic content of the predicted video and the specific analytical requirements can be combined, thereby achieving a more comprehensive understanding of video semantics.
[0082] In the embodiments of this specification, the semantic feature generation network further includes a semantic stabilizer, which is used to fuse the semantic features of the first sample and the semantic features of the second sample. The fused sample semantic features contain more semantic dimensions, including both the content dimension of the subsequently generated predicted video itself as expected in the sample description information, and the dimension for analyzing and mining the predicted video. This allows the sample semantic features to express the information of the video more richly. For example, in the sample description information, the semantic sub-feature about the scene might be "wedding scene". After fusing the above sample semantic sub-feature with the second sample semantic sub-feature about sentiment classification, a more detailed semantic feature "wedding with a strong emotional atmosphere" can be obtained.
[0083] In the embodiments of this specification, in order to further constrain the first sample semantic features corresponding to the sample description text through the semantic stabilizer, it is also necessary to obtain the third sample semantic features corresponding to N semantic expansion instructions during the process of determining the sample semantic features. For each semantic expansion instruction, the second and third sample semantic features corresponding to the semantic expansion instruction are fused to obtain the comprehensive sample semantic features corresponding to the semantic expansion instruction. The first sample semantic features and the comprehensive sample semantic features corresponding to each of the N semantic expansion instructions are concatenated to obtain the sample semantic features.
[0084] The third sample semantic features corresponding to each of the N semantic expansion instructions are learned during the training of the video generation model. These third sample semantic features can be weighted features corresponding to each semantic expansion instruction. In the embodiments of this specification, the semantic stabilizer is used to first fuse the second and third sample semantic features corresponding to each of the N semantic expansion instructions to obtain a comprehensive sample semantic feature corresponding to the semantic expansion instruction. Then, the comprehensive sample semantic feature is fused with the first sample semantic feature to obtain a stable sample semantic feature.
[0085] The above method for determining the comprehensive semantic features of samples relies on the fact that the third sample semantic features are learned during the training of the video generation model. These features represent a relatively stable and optimized set of features, possessing a certain degree of universality and representativeness. Fusing the second and third sample semantic features corresponding to the semantic extension instructions allows them to complement each other, resulting in richer and more accurate comprehensive semantic features. Furthermore, the stability of the third sample semantic features enhances the stability of the comprehensive semantic features.
[0086] The method of concatenating the semantic features of the first sample and the comprehensive semantic features of the samples corresponding to the N semantic extension instructions is used to obtain the sample semantic features. By concatenating the features from two different sources (the first sample semantic features come from the sample description information; the comprehensive semantic features of the samples come from the N semantic extension instructions and the third sample semantic features), a stable sample semantic feature is formed.
[0087] The first sample semantic feature is obtained by processing the sample video description text, containing rich details extracted from it, such as specific object attributes, the sequence of actions, and spatial relationships within the scene. This allows for a comprehensive and detailed depiction of the video's semantic content at the textual level. The comprehensive sample semantic feature, guided by the second sample semantic features corresponding to each of the N semantic extension instructions, is formed by combining the third sample semantic feature. The comprehensive sample semantic feature constrains and supplements the first sample semantic feature. By concatenating the comprehensive sample semantic feature with the first sample semantic feature, a comprehensive integration of multi-dimensional and multi-level semantic information can be achieved, enabling the sample semantic features to cover various semantic features required for predictive video generation. These include, but are not limited to, macro-level thematic style requirements, general semantic patterns, and specific textual details. This provides a more sufficient and complete semantic foundation for subsequently generating high-quality, content-rich, and logically consistent predictive videos.
[0088] S353: The predicted video is generated based on the semantic features of the sample through the video generation network. The training objective is to minimize the difference between the predicted video and the sample video 1. The parameters of the semantic feature generation network and the video generation network are updated.
[0089] Because diffusion models possess powerful semantic understanding capabilities, diverse generation abilities, and stable progressive generation, they can be included in video generation networks in some embodiments of this specification. The diffusion model generates video based on the acquired sample semantic features, obtaining sample videos corresponding to the sample description information. Of course, in other possible embodiments, the video generation network may also include other models for video generation, such as generative adversarial networks (GANs), Transformer models, etc. The specific model selection can be flexibly adjusted according to the needs of the application scenario and is not limited to the embodiments described above.
[0090] To achieve a balance between visual and semantic elements and improve the performance of the video generation model, the embodiments in this specification also extract visual features from sample video 1 in sample pair 1. In some possible embodiments, a variational autoencoder (VAE) can be used to extract the visual features of sample video 1. Of course, other feature extraction methods can also be selected, such as visual feature extraction based on convolutional neural networks or optical flow methods, etc., and this specification does not impose any limitations here. Subsequently, the extracted visual features corresponding to sample video 1 are fused with the sample semantic features corresponding to sample video description text 1, and the fused features are input into the video generation model so that the video generation model outputs a prediction of video 1 based on the fused features. This achieves the fusion of visual and semantic features during the training process.
[0091] After outputting predicted video 1, the training system 130 determines the difference between predicted video 1 and sample video 1, and compares this difference with a preset threshold. If the difference is less than the preset threshold, the error of the current video generation model is considered to be within a preset range, and the video generation model is considered to have been successfully trained, so training of the video generation model stops. Otherwise, iterative training of the video generation model continues based on other sample pairs in the sample set besides sample pair 1, until the difference between the predicted video and the sample video is minimized.
[0092] Because the training process integrates visual and semantic features, and video generation is based on the fused features, it achieves a comprehensive consideration of both visual and semantic features during training. This allows for balanced learning of visual and linguistic elements within a unified framework. The jointly trained semantic feature generation network can generate sample semantic features that better meet the needs of the video generation network. Conversely, the jointly trained video generation network can more accurately generate high-quality predicted videos based on the target semantic features, thereby improving the overall performance and generation quality of the video generation model.
[0093] Those skilled in the art will understand that, through the aforementioned model training process, the video generation model is trained such that, for different sample description information, the differences between the sample semantic features output by the semantic feature generation network fluctuate within a preset range. In other words, regardless of whether the sample description information is simple or complex, the differences between the sample semantic features output by the semantic feature generation network remain within the preset range. This reduces the dependence of the quality of the video generated by the video generation model on the video description information. Even when the video description information is relatively simple, based on the semantic expansion capabilities learned during training, the video generation model can output target semantic features with rich and comprehensive semantics, thereby generating high-quality videos based on these target semantic features.
[0094] After the video generation model is trained, the parameters in the semantic feature generation network and the video generation network will remain unchanged. Further adjustments to these parameters will only be made during the next training iteration of the video generation model.
[0095] In embodiments where the video generation model also includes convolutional layers, once the video generation model is trained, it is also determined that the relevant parameters in the convolutional layers have been trained. Further adjustments to the relevant parameters in the convolutional layers will only be made during the next training iteration of the video generation model.
[0096] In summary, the training method and system for the video generation model provided in this specification acquire the video generation model to be trained and a sample set including multiple sample pairs. The video generation model generates predicted videos based on the sample description information in the multiple sample pairs. The predicted videos corresponding to the multiple sample pairs are then compared with the corresponding sample videos for each pair, with the goal of minimizing the difference between the predicted videos and the sample videos. The video generation model is trained iteratively multiple times until a trained video generation model is obtained. The semantic feature generation network and the video generation network included in the video generation model are jointly trained during the training process. When the trained video generation model generates a target video based on the user-input video description information, it ensures the richness of the target video content and improves the stability of the target video content quality.
[0097] Figure 6 A schematic flowchart of a video generation method according to an embodiment of this specification is shown; Figure 7 A flowchart illustrating a video generation method P500 according to another embodiment of this specification is shown. This video generation method P500 can be executed by a video generation system 150. Figure 6 and Figure 7 As shown, the method P500 provided in this specification may include S510-S570, wherein:
[0098] S510: Obtain video description information and N semantic extension instructions, where N is an integer greater than or equal to 1.
[0099] In the embodiments of this specification, the video description information is the video description information input by the user. The input method can be either the video description text information directly input by the user, or the video description information obtained by converting the user's input speech information into text. The N semantic expansion instructions can be N semantic expansion instructions input by the user, or N semantic expansion instructions obtained from a preset instruction library, or a combination of some semantic expansion instructions obtained from a preset instruction library and some user-input semantic expansion instructions, to constitute the N semantic expansion instructions. It should be understood that the above embodiments are merely illustrative examples, and the specific methods for obtaining the video description information and semantic expansion instructions can be flexibly adjusted according to the current application scenario and user needs, and are not limited to those given in the above embodiments.
[0100] S530: Obtain a pre-trained video generation model, which includes a semantic feature generation network and a video generation network. During the training of the video generation model, the semantic feature generation network and the video generation network are jointly trained.
[0101] The semantic feature generation network and the video generation network are jointly trained during the training of the video generation model. The trained semantic feature generation network and video generation network achieve balanced learning of visual and linguistic elements within a unified framework. This allows the semantic feature generation network and video generation network to better adapt to each other, thereby improving the performance of the video generation model when faced with various video descriptions and video generation tasks. It also enables the semantic feature generation network and video generation network to cooperate and work collaboratively. The video generation model trained in this way, in subsequent video generation processes, neither overemphasizes semantics at the expense of visual presentation, nor focuses solely on video generation while weakening semantic integration. This ensures that the target video generated by the video generation model accurately conveys rich semantic content and has excellent visual performance, improving the overall quality of the generated target video in terms of both content and form.
[0102] S550: The target semantic features are obtained by semantically extracting video description information and N semantic extension instructions through a semantic feature generation network, and then the target video is generated based on the target semantic features through a video generation network to obtain a target video that matches the semantics of the video description information. The target semantic features represent the semantic features after semantic extension of the video description information through N semantic extension instructions.
[0103] In the embodiments of this specification, N semantic expansion instructions are used to provide a series of instruction prompts. Semantic expansion instructions are pre-defined feature information with specific directional or guiding functions, serving as a reference benchmark to provide an initial guiding direction for subsequent processing. For example, semantic expansion instructions may include any form such as: "Please describe the type of motion in the video," "Please describe the relationships between characters in the video," "Please describe the emotional tone in the video," "Please describe the atmosphere of the environment in the video," "Please describe the emotions of the characters in the video," etc.
[0104] In some possible embodiments, the semantic feature generation network may include a large model and a semantic stabilizer. In determining the target semantic features, the large model is used to extract semantics from the video description information to obtain the first semantic feature. Then, the large model is used to extract semantics from N semantic extension instructions to obtain the second semantic features corresponding to the N semantic extension instructions. Subsequently, the semantic stabilizer fuses the first semantic feature and the second semantic features corresponding to the N semantic extension instructions to obtain the target semantic feature.
[0105] In other words, the method provided in this specification introduces N semantic extension instructions and fuses the second semantic features corresponding to these N semantic extension instructions with the first semantic features corresponding to the video description information to obtain the target semantic features. This achieves the effect of supplementing the video description information with N semantic extension instructions. This allows both simple and complex video description information to be transformed into relatively stable target semantic features, reducing the requirements for user-input video description information. It enables the semantic generation network to better handle video description information of varying levels of detail; even if the user-input video description information is simple, the semantic feature generation network can still extract target semantic features with richer semantic information to a certain extent. The video quality of the target video generated based on target semantic features containing rich semantic information will also be more stable. This ensures the stability of the quality of the target video generated by the video generation model based on the video description information.
[0106] S570: Output target video.
[0107] The target video is output to the user, who can choose to download or save it. Alternatively, if the user is not satisfied with the output target video, they can re-enter the video description information to regenerate and output the target video.
[0108] It is understandable that the technical details and beneficial effects described in the video generation method P500 and the training method P300 of the video generation model can be referenced each other, and will not be repeated here.
[0109] In summary, the video generation method and system provided in this specification, because the semantic feature generation network and the video generation network included in the video generation model are jointly trained during the training process of the video generation model, enable the semantic feature generation network to generate target semantic features that are more closely aligned with the needs of the video generation network. This allows the video generation network to more accurately generate high-quality target videos based on the target semantic features, thereby improving the overall performance and generation effect of the video generation model. Furthermore, the target semantic features in this specification are determined comprehensively based on video description information and N semantic extension instructions. By introducing N semantic extension instructions to supplement and expand the video description information, comprehensive and detailed feature acquisition at the semantic level is achieved. Therefore, the feature information contained in the target semantic features is not limited to the user-input video description information but can integrate more semantic elements contained in the semantic extension instructions, enhancing the richness of the target semantic feature expression.
[0110] To further improve the accuracy of the target semantic features, this specification first determines the first semantic feature of the video description information, and then determines the comprehensive semantic feature corresponding to N semantic extension instructions. Subsequently, the target semantic feature is determined based on the first semantic feature and the comprehensive semantic feature. Specifically, the semantic stabilizer first combines the determined comprehensive semantic feature of the N semantic extension instructions based on the second and third semantic features corresponding to each of the N semantic extension instructions. Then, the semantic stabilizer fuses the comprehensive semantic feature and the first semantic feature to obtain the target semantic feature. Since the third semantic feature is learned during the training of the video generation model, the comprehensive semantic feature obtained in the above manner will be more stable. The stability of the target semantic feature obtained based on such stable comprehensive semantic features and the first semantic feature is also significantly improved. The target video generated based on these target semantic features not only ensures the richness of the target video content but also guarantees the stability of the video content quality.
[0111] This specification, in another aspect, provides a computer-readable non-transitory storage medium storing at least one set of instructions executable for performing a video generation method or a video generation model training method. When the at least one set of instructions is executed by a processor, it instructs the processor to implement the steps of the video generation method P300 or the video generation model training method P500 described herein. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on system 130, the program code causes system 130 to perform the steps of the video generation model training method P300 described herein. The program product for implementing the above method may employ a portable compact disk read-only memory (CD-ROM) containing program code and may run on system 130. When the program product is run on system 150, the program code causes system 150 to perform the steps of the video generation method P500 described herein. The program product for implementing the above method may employ a portable compact disk read-only memory (CD-ROM) containing program code and may run on system 150. However, the program product described herein is not limited to this. In this specification, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system. The program product can take any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. A computer-readable storage medium can include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar programming languages.
[0112] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0113] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure may be presented by way of example only and may not be restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.
[0114] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.
[0115] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.
[0116] Every patent, patent application, publication of a patent application, and other material, such as articles, books, specifications, publications, documents, and literature (excluding any related historical examination documents), cited in this disclosure is incorporated herein for all purposes, including, for example, in the specification and claims of this disclosure. However, in the event of any inconsistency or conflict between the descriptions, definitions, and / or terms used in the foregoing and those used in this disclosure, the descriptions, definitions, and / or terms used in this disclosure shall prevail.
[0117] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.
Claims
1. A video generation method, comprising: obtaining video description information and N semantic expansion instructions, the N being an integer greater than or equal to 1; obtaining a pre-trained video generation model, the video generation model comprising a semantic feature generation network and a video generation network, wherein, in a training process of the video generation model, the semantic feature generation network and the video generation network are jointly trained; performing semantic extraction on the video description information and the N semantic expansion instructions by the semantic feature generation network to obtain target semantic features, and performing video generation based on the target semantic features by the video generation network to obtain a target video consistent with the semantics of the video description information, wherein the target semantic features represent semantic features after semantic expansion of the video description information by the N semantic expansion instructions; and outputting the target video.
2. The method of claim 1, wherein, The semantic extraction on the video description information and the N semantic expansion instructions to obtain target semantic features comprises: performing semantic extraction on the video description information by a large model to obtain first semantic features; performing semantic extraction on the N semantic expansion instructions by the large model to obtain second semantic features corresponding to the N semantic expansion instructions; and performing feature fusion on the first semantic features and the second semantic features corresponding to the N semantic expansion instructions to obtain the target semantic features.
3. The method of claim 2, wherein, The feature fusion on the first semantic features and the second semantic features corresponding to the N semantic expansion instructions to obtain the target semantic features comprises: obtaining third semantic features corresponding to the N semantic expansion instructions; for each semantic expansion instruction, fusing the second semantic feature and the third semantic feature corresponding to the semantic expansion instruction to obtain a comprehensive semantic feature corresponding to the semantic expansion instruction; and concatenating the first semantic features and the comprehensive semantic features corresponding to the N semantic expansion instructions respectively to obtain the target semantic features.
4. The method of claim 3, wherein, The third semantic features corresponding to the N semantic expansion instructions are learned in the training process of the video generation model.
5. The method of claim 2, wherein, The semantic extraction on the video description information by the large model to obtain first semantic features comprises: performing semantic extraction on the video description information from multiple dimensions by the large model to obtain semantic sub-features corresponding to the multiple dimensions; concatenating the semantic sub-features corresponding to the multiple dimensions to obtain a concatenated feature; and performing convolution processing on the concatenated feature to obtain the first semantic features.
6. The method of claim 5, wherein, Each dimension corresponds to a semantic sub-feature including a question feature and an answer feature, the question feature being generated by the large model based on the video description information, and the answer feature being determined by the large model through a self-recurrence algorithm.
7. The method of claim 1, wherein, The video generation model is trained such that, for different video description information, the target semantic features output by the semantic feature generation network fluctuate within a preset range.
8. The method of claim 1, wherein, The obtaining of the video description information and the N semantic expansion instructions comprises: obtaining the video description information input by the user, and the N semantic expansion instructions input by the user; or obtaining the video description information input by the user, and obtaining the N semantic expansion instructions from a preset instruction library.
9. A training method of a video generation model, the method comprising: obtaining a sample set; the sample set comprising a plurality of sample pairs, each sample pair comprising a sample video and sample description information; obtaining a video generation model to be trained, the video generation model comprising a semantic feature generation network and a video generation network; and iteratively training the video generation model using the sample set to obtain a trained video generation model, wherein each iteration process comprises: performing semantic extraction on the sample description information and N semantic expansion instructions by the semantic feature generation network to obtain a sample semantic feature, the sample semantic feature representing a semantic feature after semantic expansion of the sample description information by the N semantic expansion instructions, and performing video generation based on the sample semantic feature by the video generation network to obtain a predicted video, taking minimizing the difference between the predicted video and the sample video as a training target, and updating parameters of the semantic feature generation network and the video generation network.
10. The method of claim 9, wherein, The semantic extraction on the sample description information and N semantic expansion instructions to obtain a sample semantic feature comprises: performing semantic extraction on the sample description information by a large model to obtain a first sample semantic feature; performing semantic extraction on the N semantic expansion instructions by the large model to obtain second sample semantic features corresponding to the N semantic expansion instructions; and performing feature fusion on the first sample semantic feature and the second sample semantic features corresponding to the N semantic expansion instructions to obtain the sample semantic feature.
11. The method of claim 10, wherein, The feature fusion on the first sample semantic feature and the second sample semantic features corresponding to the N semantic expansion instructions to obtain the sample semantic feature comprises: obtaining third sample semantic features corresponding to the N semantic expansion instructions; for each semantic expansion instruction, fusing the second sample semantic feature and the third sample semantic feature corresponding to the semantic expansion instruction to obtain a sample comprehensive semantic feature corresponding to the semantic expansion instruction; and concatenating the first sample semantic feature and the sample comprehensive semantic features corresponding to the N semantic expansion instructions to obtain the sample semantic feature.
12. The method of claim 10, wherein, The semantic extraction on the sample description information by the large model to obtain a first sample semantic feature comprises: performing semantic extraction on the sample description information by the large model from multiple dimensions to obtain sample semantic sub-features corresponding to the multiple dimensions; concatenating the sample semantic sub-features corresponding to the multiple dimensions to obtain a sample concatenated feature; and performing convolution processing on the sample concatenated feature to obtain the first sample semantic feature.
13. The method of claim 12, wherein, Each dimension corresponds to a sample semantic sub-feature, including a sample question feature and a sample answer feature, the sample question feature is generated by the large model based on the sample description information, and the sample answer feature is determined by the large model through a self-recurrence algorithm.
14. The method of claim 9, wherein, The video generation model is trained to fluctuate within a preset range between sample semantic features output by the semantic feature generation network for different sample description information.
15. A video generation system, comprising: at least one storage medium storing at least one instruction set for video generation; and at least one processor in communication connection with the at least one storage medium, wherein the at least one processor reads the at least one instruction set when running and executes the method of any one of claims 1-8 according to the indication of the at least one instruction set.
16. A training system for a video generation model, comprising: at least one storage medium storing at least one instruction set for training a video generation model; and at least one processor in communication connection with the at least one storage medium, wherein the at least one processor reads the at least one instruction set when running and executes the method of any one of claims 9-14 according to the indication of the at least one instruction set.
Citation Information
Patent Citations
Training method of image editing model and image editing method and device
CN116363261A
Video generation method and device, medium and computer equipment
CN117095085A