Video generation method, computing device, computer readable storage medium, and computer program product
By receiving video generation instructions and adjustment instructions, using noise graph groups and prompt text to generate target videos in the video generation model, the problem of inaccurate video editing in the prior art is solved and the high accuracy of video editing is achieved.
Patent Information
- Application Number
- PCT/IB2024/063338
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2024-12-31
- Publication Date
- 2025-08-14
AI Technical Summary
In the prior art, video generation methods based on text editing are difficult to achieve accurate modification of the original video, resulting in a large difference between the generated video and the original video, limiting the modification accuracy of the video content.
The initial prompt text is generated by receiving the video generation instructions, and the reference noise graph group and the initial prompt text are input to the first video generation model, the target video feature information collection is obtained, and the target video is generated by combining the video adjustment instructions, and the target video generation model is input to the second video generation model to generate the target video to ensure the accuracy of video editing.
The generated target video is achieved to maintain a small difference from the original video, improving the accuracy and accuracy of video editing.
Smart Images

Figure IB2024063338_14082025_PF_FP_ABST
Abstract
Description
[0001]TECHNICAL FIELD: Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a video generation method. Background: With the rapid development of generative artificial intelligence (AI) technology, text-based video editing has become one of the most notable and widely used technologies in this field. This technology uses text as input to edit videos, aiming to enhance the visual appeal and creativity of video content. In particular, with the advancement of text-to-image generation models, text-to-video generation technology has also significantly improved. Currently, when users use text that slightly modifies the description of an original video to regenerate a video based on the original video, the resulting video often differs significantly from the original video, limiting content modification based on the original video. Therefore, to address this issue, a video generation method is needed that enables precise editing of original videos using text instructions. SUMMARY: In view of this, embodiments of the present disclosure provide a video generation method, a video generation method for a cloud server. One or more embodiments of the present disclosure relate to a video generation apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art. According to a first aspect of an embodiment of the present disclosure, a video generation method is provided, comprising: receiving a video generation instruction and generating initial prompt text according to the video generation instruction; obtaining a reference noise map group, inputting the reference noise map group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model during the generation of the reference video; receiving a video adjustment instruction and generating target prompt text according to the video adjustment instruction and the initial prompt text; and inputting the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model.According to a second aspect of an embodiment of the present disclosure, a video generation method for a cloud server is provided, comprising: receiving a video generation instruction sent by a terminal device, and generating initial prompt text based on the video generation instruction; obtaining a reference noise map group, inputting the reference noise map group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model during the generation of the reference video; receiving a video adjustment instruction sent by a terminal device, and generating target prompt text based on the video adjustment instruction and the initial prompt text; inputting the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model, obtaining a target video generated by the second video generation model, and returning the target video to the terminal device. According to a third aspect of an embodiment of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions. When executed by the processor, the computer-executable instructions implement the aforementioned video generation method and the steps of the video generation method for a cloud server. According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, storing computer-executable instructions. When executed by a processor, the instructions implement the aforementioned video generation method and the steps of the video generation method applied to a cloud server. According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program / instructions. When executed by a processor, the computer program / instructions implement the aforementioned video generation method and the steps of the video generation method applied to a cloud server. One embodiment of the present disclosure implements receiving a video generation instruction and generating initial prompt text based on the video generation instruction; obtaining a reference noise map group, inputting the reference noise map group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model during the generation of the reference video; receiving a video adjustment instruction and generating a target prompt text based on the video adjustment instruction and the initial prompt text; and inputting the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model.Applying the solution of the embodiments of the present disclosure, a reference video is generated using initial prompt text, and the target video feature information set generated during the reference video generation process is used to generate a target video based on the target prompt text. This ensures that the target video generated using the target prompt text, which is slightly different from the initial prompt text, maintains minimal differences from the reference video, thereby improving the accuracy of video editing. BRIEF DESCRIPTION OF THE DRAWINGS FIG1 is a flow chart of a video generation method according to an embodiment of the present disclosure; FIG2 is a schematic diagram of a process for generating and editing a video according to an embodiment of the present disclosure; FIG3 is a schematic diagram of a process for editing an original video according to an embodiment of the present disclosure; FIG4 is a schematic diagram of a process for editing a masked video according to an embodiment of the present disclosure; FIG5 is a flow chart of a video generation method applied to a cloud server according to an embodiment of the present disclosure; FIG6 is an architecture diagram of a video generation system according to an embodiment of the present disclosure; FIG7 is a flow chart of a process for editing a user's original video according to an embodiment of the present disclosure; FIG8 is a schematic diagram of the structure of a video generation device according to an embodiment of the present disclosure; and FIG9 is a block diagram of the structure of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS The following description sets forth numerous specific details to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed below. The terminology used in one or more embodiments of the present disclosure is for the purpose of describing specific embodiments only and is not intended to limit the present disclosure. As used in one or more embodiments of the present disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be understood that while the terms "first," "second," and so on may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another. For example, "first" could be referred to as "second," and similarly, "second" could be referred to as "first," without departing from the scope of one or more embodiments of the present disclosure. Depending on the context, the term "if" as used herein could be interpreted as "when," "when," or "in response to determining."Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models can also be called cornerstone models / foundation models. They are pre-trained using large-scale unlabeled corpora to produce pre-trained models with parameters exceeding 100 million. Such models are adaptable to a wide range of downstream tasks and have good generalization capabilities. Examples include large language models (LLMs) and multi-modal pre-training models. In practical applications, large models only require a small number of samples to fine-tune the pre-trained model and can be applied to various tasks. Large models can be widely used in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Key application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. First, the terms used in one or more embodiments of this disclosure are explained. Deterministic Denoising Diffusion Implicit Models (DDIM): This is a generative model based on a diffusion model, characterized by its deterministic sample generation method. Unlike traditional random diffusion models, the output of each step in the denoising process is deterministic, rather than random. This approach not only speeds up the sampling process but also enables more accurate control of the quality of the generated samples.DDIM gradually recovers clear data from noisy data through a series of inverse diffusion steps and is widely used in fields such as image generation and super-resolution. Stable Diffusion is an advanced deep learning model specifically designed for generating high-quality images. It combines diffusion models with deep learning techniques to generate images by gradually removing noise, thereby achieving the generation of complex and high-resolution images. The key advantage of Stable Diffusion is its ability to stably generate images with rich detail and realism while maintaining high computational efficiency. This has attracted widespread attention in fields such as image synthesis, artistic creation, and media editing. Variational Autoencoder (VAE) is a deep learning model used for unsupervised learning tasks, particularly data generation. It consists of two parts: an encoder and a decoder. The encoder transforms the input data into a representation in a latent space, while the decoder reconstructs the data from this latent representation. VAEs are trained by minimizing the reconstruction error and the difference between the latent representation and a prior distribution (typically a Gaussian distribution). This enables VAEs to both efficiently generate data and learn complex data distributions. VAEs are widely used in fields such as image generation, denoising, and style transfer. This disclosure provides a video generation method, a video generation method applied to a cloud server. This disclosure also relates to a video generation apparatus, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments. Referring to Figure 1, a flowchart of a video generation method according to one embodiment of this disclosure is shown, specifically comprising the following steps. Step 102: Receive a video generation instruction and generate initial prompt text based on the instruction. In practical applications, the video generation instruction is an instruction for generating a reference video, and the initial prompt text is a text describing the reference video. Specifically, a video generation instruction can be understood as an instruction for generating a reference video that a user wants to modify. For example, a user uploading the original video (Video 1) on the video modification service page and then clicking "Modify with this video" sends a video generation instruction to the server. Alternatively, a user entering "Generate a video with a focus on rabbits" and clicking "Send" sends a video generation instruction to the server, and so on. This disclosure does not impose any restrictions on this.The initial prompt text can be understood as standardized text describing the target video that the user wants to modify. For example, based on the original video uploaded by the user, Video 1, a description of the video might be generated as "a video of a rabbit eating watermelon." Another example is that the user input text "Generate a video focusing on rabbits for me" might be parsed to generate a standardized description as "a photo video of a rabbit." This disclosure does not impose any restrictions on this. By receiving a video generation instruction, a standardized initial prompt text can be generated based on the video generation instruction, making the video generated by the video generation model more accurate. The standardized initial prompt text also allows subsequent adjustments to the video to be implemented using another standardized text, thereby improving the accuracy of video editing. Considering that a user may first generate a video in text form and then modify the video generated based on the initial text, the video generation instruction carries initial text information. Furthermore, receiving the video generation instruction and generating initial prompt text based on the video generation instruction includes: parsing the video generation instruction to obtain the initial text information; generating the initial prompt text based on the initial text information; accordingly, obtaining a reference noise map group includes: generating the reference noise map group according to a preset noise map group generation rule; accordingly, receiving a video adjustment instruction includes: receiving the video adjustment instruction corresponding to the reference video. In actual applications, the initial text information is text information directly input by the user, the reference noise map group is a set of noise maps used to generate a reference video, the noise map group generation rule is a rule for generating a noise map group when the video generation instruction sent by the user carries the initial text information, the reference video is a video generated based on the reference noise map group and the initial prompt text, and the video adjustment instruction is an instruction for adjusting the reference video. Specifically, the initial text can be understood as non-standard information directly input by the user, for example, "Generate a video with a focus on rabbits," "Give me a video of a jumping rabbit," and so on. If the video generation instruction carries initial text information, the obtained initial prompt text can be understood as text information that standardizes the initial text sent by the user and highlights the key points of the initial text. For example, if the initial text sent by the user is "Generate a video with a focus on rabbits," the initial prompt text determined after parsing the initial text may be "A video of a rabbit portrait." This initial prompt text highlights the focus on rabbits in the initial text and interprets "focus on rabbits" as "rabbit portraits," making the generated video more accurate.When the video generation instruction carries initial text information, the obtained reference noise map group can be understood as a set of random noise images, such as a set of Gaussian noise maps, a set of uniform noise maps, a set of speckle noise maps, and so on, though this disclosure imposes no restrictions on this. The noise map group generation rule can be understood as a rule for generating a set of random noise maps, such as a Gaussian noise map generation rule or a uniform noise map generation rule. The reference video can be understood as a video generated using the initial prompt text that meets the description of the initial prompt text. For example, if the initial prompt text is "a photo video of a rabbit," the generated reference video is Video A, which is a video of a rabbit resting in the grass, with only rabbits on the grass. The video generation instruction sent by the user carries initial text information, which can be understood as requiring the user to first generate a video using the initial text information and then edit the generated video. Referring to Figure 2, a schematic flow diagram of video generation and editing, provided in one embodiment of the present disclosure, is shown. In this process, an initial text message sent by a user is received and parsed into a standardized initial prompt text 204. A set of Gaussian noise reference noise map groups 202 is generated using preset rules. The initial prompt text 204 and the reference noise map group 202 are then input together into a cross-feature extraction unit of a first feature extraction layer 2062 of a first video generation model 206 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. The cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key positions in the noise map group. The cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information that includes spatiotemporal features in the set of reference noise maps. The cross-feature information, self-attention features, and spatiotemporal feature information are then input together into the first video generation layer 2064 to obtain a reference video 208.After receiving the reference video 208, the user issues a modification command to the reference video 208, sending a video adjustment instruction. Subsequently, based on the adjustment text in the video adjustment instruction, a target prompt text 216 with slight variations from the initial prompt text 204 is generated. The same reference noise map group 202 and target prompt text 216 are then input into the second video generation model 218. Since the second video generation model 218 has the same structure and parameters as the first video generation model 206, the reference noise map group 202 and target prompt text 216 undergo the same operations in the second feature extraction layer 2182 as the reference noise map group 202 and the initial prompt text 204 in the first feature extraction layer 2062. Subsequently, the question feature information and key value feature information in the self-attention features generated based on the target prompt text 216 and the reference noise map group 202 are replaced with the reference question feature information 210 and reference key value feature information 212 generated based on the initial prompt text 204 and the reference noise map group 202. The spatiotemporal feature information generated based on the target prompt text 216 and the reference noise map group 202 is then replaced with the reference spatiotemporal feature information 214 generated based on the initial prompt text 204 and the reference noise map group 202. The replaced feature information is then input into the second video generation layer 2184, obtaining the target video 220 obtained by the user modifying the reference video 208. By receiving the video generation instruction carrying the initial text information, a video is first generated in the form of text, and then the video is modified based on the video generated based on the initial text, thereby enabling user editing of the video. Considering that users may directly upload an original video and then modify it based on the original video, the video generation instruction carries initial video information. Furthermore, receiving the video generation instruction and generating initial prompt text based on the video generation instruction includes: parsing the video generation instruction to obtain the initial video information; generating the initial prompt text based on the initial video information, wherein the initial prompt text describes the initial video information; accordingly, obtaining a reference noise map group includes: generating the reference noise map group corresponding to the initial video information based on the initial video information; accordingly, receiving the video adjustment instruction includes: receiving the video adjustment instruction corresponding to the initial video information. In actual applications, the initial video information is the video uploaded by the user. Specifically, the initial video information can be understood as the video to be edited by the user.When the video generation instruction carries the initial video information, the obtained reference noise map group can be understood as a set of noise maps generated by performing noise processing on each frame in the initial video information. Specifically, the noise processing method for the frames in the initial video information can be reverse DDIM (Deterministic Denoising Diffusion Implicit Models), etc., which is not limited in this disclosure. When the video generation instruction carries the initial video information, the obtained initial prompt text can be understood as a summary of the initial video information, obtaining a text describing the initial video. For example, if the video uploaded by the user is a video of a rabbit eating watermelon on a table, the initial prompt text corresponding to the video is "A video of a rabbit eating watermelon on a table." By using the set of noise maps obtained from the initial video information and the text describing the initial video information, the reference video generated based on the above data can be made nearly identical to the initial video. The feature values generated when generating the reference video can also be understood as the feature values of the initial video information. The target video subsequently generated using these feature values will also be similar to the initial video. The video generation instruction sent by the user carries initial video information. This can be understood as indicating that the user needs to edit the original video. However, since the original video has not been generated using the first video generation model, it is necessary to first use the first video generation model to generate a reference video that is substantially identical to the original video to obtain some features of the original video. Referring to FIG3 , FIG3 shows a schematic flow diagram of an original video editing process provided by one embodiment of the present disclosure. In this process, initial video information 302 sent by a user is received, and noise processing is then performed on each frame in the initial video information 302 to generate a set of reference noise map groups 304. The initial video information 302 is parsed into initial prompt text 306 describing the initial video information. The initial prompt text 306 and the reference noise map group 304 are then input together into a cross-feature extraction unit in a first feature extraction layer 3082 in a first video generation model 308 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. This cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key locations in the noise map group. This cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information that includes spatiotemporal features in the set of reference noise maps. The cross-feature information, self-attention features, and spatiotemporal feature information are then input together into the first video generation layer 3084 to obtain a reference video 310.The user issues a modification command to the initial video information 302, sending a video adjustment instruction. Subsequently, based on the adjustment text in the video adjustment instruction, a target prompt text 318 with slight variations from the initial prompt text 306 is generated. The same reference noise map group 304 and target prompt text 318 are then input into the second video generation model 320. Since the second video generation model 320 has the same structure and parameters as the first video generation model 308, the reference noise map group 304 and target prompt text 318 undergo the same operations in the second feature extraction layer 3202 as the reference noise map group 304 and the initial prompt text 306 in the first feature extraction layer 3082. Subsequently, the question feature information and key value feature information in the self-attention features generated based on the target prompt text 318 and the reference noise map group 304 are replaced with the reference question feature information 312 and reference key value feature information 314 generated based on the initial prompt text 306 and the reference noise map group 304. The spatiotemporal feature information generated based on the target prompt text 318 and the reference noise map group 304 is then replaced with the reference spatiotemporal feature information 316 generated based on the initial prompt text 306 and the reference noise map group 304. This replaced feature information is then input into the second video generation layer 3204, obtaining a target video 322 generated by the user modifying the reference video 310. By receiving a video generation instruction carrying the initial video information, the user modifies the video based on the original video, thereby enabling the user to edit the original video. Step 104: Obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, and obtain a target video feature information set generated by the first video generation model during the processing of the reference noise map group and the initial prompt text. In practical applications, the first video is used as a model to extract partial features of the video to be modified, and the target video feature information set is the partial features of the video to be modified. Specifically, an initial prompt text describing the video to be modified is input into the first video generation model. The first video generation model can extract some features of the video to be modified by processing the initial prompt text. By using the extracted features to generate a target video, the generated target video can be made more similar to the unmodified parts of the video to be modified, thereby improving the accuracy of the video modification.Furthermore, inputting the reference noise pattern group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated by the first video generation model during processing of the reference noise pattern group and the initial prompt text, includes: inputting the reference noise pattern group and the initial prompt text into the first video generation model to generate a reference video; and obtaining a target video feature information set generated by the first video generation model during the generation of the reference video. In practical applications, the reference video is the video to be modified. In other words, the first video generation model can be further understood as a model for extracting partial features of the reference video, and the reference video can be understood as the video to be modified. Inputting the initial prompt text describing the video to be modified into the first video generation model to generate the reference video can be understood as extracting partial features of the reference video during the reference video generation process. By using these extracted partial features to generate the target video, the generated target video can be made more similar to the unmodified portions of the reference video, thereby improving the accuracy of the video modification. The target video feature information set can be understood as partial features of the reference video. The target video feature set may include one or more features of the reference video, such as temporal attention features representing the temporal features of each entity in the reference video, reference spatial attention features representing the spatial features of each entity in the reference video, reference spatiotemporal attention features representing the temporal and spatial fusion features of each entity in the reference video, reference self-attention features representing the relationship between each pixel in the reference video, and reference cross-attention features representing the association between the reference video and the initial prompt text. This disclosure does not impose any restrictions on this. In one embodiment provided in this disclosure, the initial prompt text "a video of a rabbit portrait" and a set of Gaussian noise images are input into a first video generation model. The obtained reference video is a video of a white rabbit resting in the grass. During the generation of this video, spatiotemporal attention features representing the temporal and spatial fusion features of each entity in the video and self-attention features representing the relationship between each pixel in the video are obtained. By extracting features corresponding to the reference video generated during the reference video generation process and then using the acquired features to generate the target video, the generated target video can be made more similar to the unmodified parts of the reference video.Considering that video generation requires the automatic discovery and extraction of useful information or features from the prompt text and noise image for further processing and analysis, the first video generation model includes a first feature extraction layer and a first video generation layer. Furthermore, inputting the reference noise image group and the initial prompt text into the first video generation model to generate a reference video includes: inputting the reference noise image group and the initial prompt text into the first feature extraction layer to obtain a reference attention feature information set output by the first feature extraction layer; and inputting the reference attention feature information set into the first video generation layer to obtain a reference video output by the first video generation layer. In practical applications, the first feature extraction layer is used to extract features of the reference video, the first video generation layer is used to generate a deep learning network layer for the reference video based on the features extracted by the first feature extraction layer, and the reference attention feature information set is input into the video generation layer to obtain the reference video. Specifically, the first feature extraction layer can be understood as a computational layer that generates features of the reference video. This layer can be a feature extraction layer with or without a neural network structure. This disclosure does not impose any restrictions on the results of this layer. For example, based on the input initial prompt text and a set of noise maps, self-attention features representing the relationship between pixels in the generated video, cross-attention features representing the relationship between the generated video and the initial prompt text, etc. can be extracted. This disclosure does not impose any restrictions on this. The first video generation layer can be understood as a deep learning network layer that generates a corresponding video based on the extracted video features. For example, DDIM, Stable Diffusion, and other network models can be used to generate the video. This disclosure does not impose any restrictions on this. The reference attention feature information set can be understood as all features of the reference video to be generated. It includes all features of the reference video extracted by the first feature extraction layer, such as temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, etc. This disclosure does not impose any restrictions on this. Through the first feature extraction layer, all features of the reference video are extracted. Features that the user will not modify can be obtained from the extracted features and used to generate the target video, so that the generated target video can be more similar to the unmodified parts of the reference video.Furthermore, obtaining a target video feature information set generated by the first video generation model during the process of generating the reference video includes: obtaining a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature information set includes at least one type of attention feature information; and determining a target video feature information set based on the reference attention feature information set, wherein the target video feature information set includes at least one type of attention feature information. Specifically, the target video feature information set can be understood as a selection of features that are not modified by the user from the reference attention feature information set representing all features of the reference video. By obtaining features that are not modified by the user and using them to generate the target video, the generated target video can be made more similar to the unmodified portions of the reference video. Preferably, considering that users typically do not modify the spatiotemporal relationships between entities in the reference video, or the relationships between elements in the reference video, the reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information. Furthermore, determining the target video feature information set based on the reference attention feature information set includes: determining reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information based on the reference attention feature information set; and determining the target video feature information set based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information. In practical applications, the reference self-attention feature information is a feature value representing the relationship between pixels in the reference video, and the reference spatiotemporal self-attention feature information is a feature value representing the temporal and spatial relationships between entities in the reference video. Specifically, the reference self-attention feature information can be understood as the characteristics of the correlation between each pixel in the reference video itself. For example, the reference video is of a rabbit resting on the grass, where the rabbit's two eyes are pixels a and pixel b, and pixel c is in the sky of the video. Then, there is a strong correlation between pixels a and b, and pixel c has a weak correlation with either pixel a or pixel b. This is reflected in the feature matrix as 0.9 for points <a,b> and <b,a>, and 0.1 for points <c,a> and <c,b>. The larger the number, the stronger the correlation.The reference spatiotemporal self-attention feature information can be understood as the spatiotemporal relationships between entities in the reference video. For example, in a reference video of a rabbit resting on grass, the temporal relationship of the rabbit itself is strongly correlated because it changes over time. Since the rabbit rests on the grass, the grass beneath it moves as it moves, so the spatial relationship between the rabbit and the grass is strongly correlated. However, since the rabbit's movement does not affect the sky, and the sky is not time-dependent, the spatial and temporal relationships between the rabbit and the sky are weak. In other words, the reference spatiotemporal self-attention feature information can also be understood as the motion characteristics of entities in the reference video. By obtaining the correlation features between the pixels in the video itself (which the user does not modify), as well as the motion characteristics of entities in the video, and using these features to generate the target video, the generated target video can be made more similar to the unmodified parts of the reference video. Considering that the inclusion of initial prompt text and reference noise map feature information during the generation of a reference video can ensure consistency between the background of the target video and the reference video during the subsequent generation of the target video, the reference self-attention feature information includes reference question feature information, reference key value feature information, and reference value feature information, and the reference spatiotemporal self-attention feature information includes reference spatiotemporal feature information. Furthermore, determining a target video feature information set based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information includes: determining reference question feature information and reference key value feature information based on the reference self-attention feature information; determining reference spatiotemporal feature information based on the reference spatiotemporal self-attention feature information; and determining the reference question feature information, the reference key value feature information, and the reference spatiotemporal feature information as target feature information. In practical applications, the reference question feature information and the reference key value feature information are feature information used to calculate the relationship between pixels in the reference video, and the reference spatiotemporal feature information is feature information representing temporal and spatial fusion features between entities in the reference video. Specifically, the reference question feature information can be understood as feature information that uses each pixel in the reference video as the question, and the reference key value feature information can be understood as feature information that uses each pixel in the reference video as the key value. The dot product of the question feature information and the key value feature information is performed to determine the relationship between each pixel in the reference video. This is then added to the answer value feature information for each pixel in the reference video to strengthen the relationship between each pixel in the reference video. This allows the acquisition of reference self-attention feature information representing the relationship between each pixel in the reference video.Reference spatiotemporal feature information can be understood as the temporal and spatial relationships between entities in the reference video, or as feature information representing the motion characteristics of the reference video. By obtaining problematic feature information and key value feature information for videos that users typically don't modify, we can avoid the problem of users not seeing changes reflected when modifying the relationships between pixels in the video. Furthermore, by obtaining the motion characteristics of entities in the video and using these features to generate the target video, the generated target video can be made more similar to the unmodified portions of the reference video. Step 106: Receive a video adjustment instruction and generate target prompt text based on the video adjustment instruction and the initial prompt text. In practical applications, a video adjustment instruction is an instruction to adjust a video. A video adjustment instruction can be understood as a user-sent instruction to adjust the video to be adjusted. For example, if a user wants to adjust an uploaded initial video, the video adjustment instruction is an instruction to adjust the initial video. If a user wants to adjust a reference video generated based on text, the adjustment instruction is an instruction to adjust the reference video, and so on. The target prompt text is the prompt text used to generate the target video. The target prompt text can be understood as text with slight modifications to the initial text. For example, the initial prompt text is "A photo video of a rabbit," and the slightly modified target prompt text is "A photo video of a black rabbit." In one embodiment provided by the present disclosure, a video adjustment instruction containing a user-edited command "Make the rabbit black" is received. The target prompt text "A photo video of a black rabbit" is generated by combining the initial prompt text with the original prompt text. By combining the target prompt text with the initial prompt text, the initial prompt text is slightly modified according to the user's requirements, making the generated video more similar to the unmodified parts of the original reference video. Step 108: Input the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model. In practical applications, the second video generation model is the target of generating the target video, and the target video is the video generated based on the target prompt text. Specifically, the second video generation model can be understood as a video generation model with the same structure and parameters as the first video generation model, and is used to generate a target video by combining the reference noise pattern group and the target prompt text. The target video can be understood as a video obtained by modifying some entities or features in the reference video. For example, the reference video before modification is a white rabbit sleeping on the grass, and the modified target video is a black rabbit sleeping on the grass. The sleeping posture of the black rabbit and the weather in the sky are the same after modification.By processing the same reference noise pattern set using a second video generation model with the same parameters as the first video generation model, the modified target video can be effectively made somewhat consistent with the pre-modified reference video. Subsequently, by replacing some of the features used to generate the target video with some of the features of the reference video generated during the reference video generation process, the generated target video can be made more similar to the unmodified portions of the original reference video. Considering the need to generate two similar videos, the second video generation model also includes a second feature extraction layer and a second video generation layer, and the parameters of the second video generation layer are the same as those of the first video generation layer of the first video generation model. Furthermore, inputting the reference noise pattern set, the target prompt text, and the target video feature information set into the second video generation model to obtain a target video generated by the second video generation model includes: inputting the reference noise pattern set, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set; and inputting the target attention feature information set into the second video generation layer to obtain a target video output by the second video generation layer. In practical applications, the second feature extraction layer is used to extract features of the target video. The second video generation layer is used to generate a deep learning network layer for the target video based on the features extracted by the second feature extraction layer and some features extracted from the first feature extraction layer. The target attention feature information set is input into the video generation layer to obtain the target video. Specifically, the second feature extraction layer can be understood as a computational layer that generates features of the target video. This layer can be a feature extraction layer with or without a neural network structure. This disclosure does not impose any restrictions on the results of this layer. If this layer is a feature extraction layer with a neural network structure, the parameters of the neural network structure in this layer are the same as those of the first feature extraction layer. For example, based on the input target prompt text and a set of noise maps, self-attention features of the relationship between each pixel in the generated video, cross-attention features of the relationship between the generated video and the target prompt text, etc. are extracted. This disclosure does not impose any restrictions on this. The second video generation layer can be understood as a deep learning network layer that generates the corresponding video based on the extracted video features. The parameters of this neural network layer are the same as those of the first video generation layer described above. For example, DDIM, Stable Diffusion (stable diffusion model), etc. can generate a network model for the video, and this disclosure does not impose any restrictions on this.The target attention feature information set can be understood as all features of the target video to be generated. This includes features extracted by the second feature extraction layer and then replaced with corresponding features in the target video feature set. For example, temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, and so on are obtained. This disclosure imposes no limitations on this. By extracting some features of the target video through the second feature extraction layer, the user-unmodifiable features are replaced with corresponding features obtained by the second feature extraction layer. Generating the target video using this replaced feature set can make the generated target video more similar to the unmodified portions of the reference video. Furthermore, inputting the reference noise pattern group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set includes: inputting the reference noise pattern group and the target prompt text into the second feature extraction layer to obtain an intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one type of attention feature information; and generating a target attention feature information set based on the intermediate attention feature information set and the target video feature information set. In practical applications, the intermediate attention feature information set is the set of all features output by the second feature extraction layer. Specifically, the intermediate attention feature information set can be understood as features output by the second feature extraction layer and possessed by the video generated solely based on the target prompt text, such as temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, etc., which are not limited in this disclosure. It should be noted that, considering the consistency of the generated target video with the original reference video in the unmodified portions, the question feature information and key value question feature information used in calculating the target video's self-attention features can be replaced with the reference question feature information and key value question feature information of the reference video from the previous model. By replacing the aforementioned user-unmodified features (the target video feature set) with the corresponding features obtained through the second feature extraction layer (features of a video generated solely based on the target prompt text), and then generating the target video based on this replaced feature set, the generated target video can be made more similar to the unmodified portions of the reference video than two videos generated by simply processing the same noise image set and slightly different prompt text using a model with the same parameters.Furthermore, generating a target attention feature information set based on the intermediate attention feature information set and the target video feature information set includes: determining target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set; determining to-be-replaced attention feature information in the intermediate attention feature information set, wherein the to-be-replaced attention feature information is the attention feature information corresponding to the target attention feature information; and replacing each to-be-replaced attention feature information with the corresponding target attention feature information to generate the target attention feature information set. In practical applications, the target attention feature information is the feature information in the target video feature information set, and the to-be-replaced attention feature information is the feature information corresponding to the target attention feature information. Specifically, the to-be-replaced attention feature information can be understood as partial features of a video generated solely based on the target prompt text. These partial features are typically not modified by the user, and include, for example, self-attention features representing the relationship between pixels in the video, spatiotemporal attention features representing the motion relationship between entities in the video, and so on. This disclosure does not impose any limitations on this. By finding the feature information to be replaced corresponding to each feature information in the target video feature information set and replacing the corresponding feature information to be replaced, and then generating a target video based on the replaced feature information set, the generated target video can be made more similar to the unmodified parts of the reference video than two videos generated by simply processing the same noise image group and slightly different prompt text through a model with the same parameters. Taking into account that, in the process of generating a reference video, information including initial prompt text and reference noise map features is included, so that the background of the target video can be kept consistent with that of the reference video in the subsequent process of generating a target video, the intermediate attention feature information set includes intermediate question feature information, intermediate key value feature information and intermediate spatiotemporal feature information, and the target video feature information set includes reference question feature information, reference key value feature information and reference spatiotemporal feature information; further, determining the replacement attention feature information in the intermediate attention feature information set includes: when the target attention feature information is the reference question feature information, determining that the intermediate question feature information corresponding to the reference question feature information is the question feature information to be replaced; when the target attention feature information is the reference key value feature information, determining that the intermediate key value feature information corresponding to the reference key value feature information is the key value feature information to be replaced; when the target attention feature information is the reference spatiotemporal feature information, determining that the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is the spatiotemporal feature information to be replaced.In practical applications, the intermediate question feature information and intermediate key value feature information are feature information used to calculate the relationship between each pixel in the video generated based on the target prompt text, the question feature information to be replaced and the key value feature information to be replaced are feature information used to replace the reference question feature information and reference key value feature information, and the intermediate spatiotemporal feature information is feature information representing the temporal and spatial fusion features between entities in the video generated based on the target prompt text. The spatiotemporal feature information to be replaced is feature information used to replace the reference spatiotemporal feature information. Specifically, the intermediate question feature information can be understood as feature information that uses each pixel of the video generated based on the target prompt text as the question, and the intermediate key value feature information can be understood as feature information that uses each pixel in the video generated based on the target prompt text as the key value. The feature information used as the question is multiplied with the feature information used as the key value to determine the relationship between each pixel in the video generated based on the target prompt text. The two intermediate question feature information and the intermediate key value feature information are replaced to obtain the relationship between each pixel in the reference video. This is then added to the value feature information that uses each pixel of the video generated based on the target prompt text as the answer. While ensuring that the relationship between each pixel in the generated target video is similar to the relationship between each pixel in the reference video, user modifications to the video (that is, the relationship between each pixel in the video generated based on the target prompt text) can be added. This ensures that the generated target video is more similar to the unmodified portions of the reference video, thereby improving the user experience. Intermediate spatiotemporal feature information can be understood as the temporal and spatial relationships between entities in a video generated based on the target prompt text, or as feature information representing the motion characteristics of a video generated based on the target prompt text. Replacing this intermediate spatiotemporal feature information with reference spatiotemporal feature information representing the motion characteristics of a reference video effectively preserves the motion information of each entity in the reference video in the newly generated target video, thereby effectively making the generated target video more similar to the unmodified portions of the reference video. Furthermore, replacing each set of attention feature information to be replaced with the corresponding set of target attention feature information includes: replacing the question feature information to be replaced with the reference question feature information; replacing the key value feature information to be replaced with the reference key value feature information; and replacing the spatiotemporal feature information to be replaced with the reference spatiotemporal feature information.By replacing each feature information in the target video feature information set with its corresponding feature information to be replaced, and then generating a target video based on the replaced feature information set, the generated target video can be made more similar to the unmodified parts of the reference video than two videos generated by simply processing the same noise image group and slightly different prompt text through a model with the same parameters. Taking into account that the accuracy of video modification can be further improved without considering resource consumption, the second feature extraction layer includes a second cross-feature extraction unit; further, after obtaining the target video generated by the second video generation model, the method also includes: determining the prompt text distinguishing word information according to the initial prompt text and the target prompt text; determining at least one associated word information with the prompt word text distinguishing word information according to the initial prompt text and the prompt text distinguishing word information; obtaining the distinguishing cross-attention feature information corresponding to the prompt word text distinguishing word information and the associated cross-attention feature information corresponding to each associated word information generated by the second cross-feature extraction unit according to the reference noise map group and the target prompt word text; generating a target edited video according to the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video. In practical applications, the second cross-feature extraction unit is an attention unit that obtains the association between the target prompt text and the target video. The prompt text distinguishing word information is the word that distinguishes the target prompt text from the initial prompt text. The associated word information is the word in the initial prompt text that is related to the prompt text distinguishing word information. The distinguishing cross-attention feature information is feature information indicating the association between the prompt text distinguishing word information and the target video. The associated cross-attention feature information is feature information indicating the association between each associated word and the target video. Specifically, the prompt text distinguishing word information determined based on the initial prompt text and the target prompt text can be understood as the difference between the initial prompt text and the target prompt text, and can also be further understood as the word that the user wants to adjust. For example, if the initial prompt text is "a video of a rabbit photo" and the target prompt text is "a video of a black rabbit photo," then the prompt text distinguishing word information corresponding to these two words is "black."The at least one associated word information determined to correspond to the prompt text's distinguishing word information can be understood as the preset number of words in the initial prompt text that have the greatest association with the acquired prompt text's distinguishing word information. It can also be further understood as the preset number of words that are most relevant to the prompt text's distinguishing word information. For example, if the initial prompt text is "a rabbit photo video" and the confirmed prompt text's distinguishing word information is "black," then if the preset number of associated words is 1, then the one associated word associated with the prompt text's distinguishing word information is "rabbit." If the preset number of associated words is 2, then the two associated words currently associated with the prompt text are "rabbit" and "photo." It should be noted that the method for obtaining the prompt text's distinguishing word information and determining the associated words corresponding to the prompt text's distinguishing word information can be by calculating feature vectors and distances for each word, or by organizing the relationships between words using a model, etc., and this disclosure does not impose any limitation on this. The distinguishing cross-attention feature information can be understood as the relationship between the obtained prompt text distinguishing word information and the generated target video, and can also be further understood as the relationship between the prompt text distinguishing word information and various entities in the video. Similarly, the associative cross-attention feature information can be understood as the relationship between the obtained associative word information and the generated target video, and can also be further understood as the relationship between the associative word information and various entities in the video. Referring to Figure 4, Figure 4 shows a schematic flow diagram of an additional mask video editing process provided by one embodiment of the present disclosure. In this process, an initial text message sent by a user is received and parsed into a standardized initial prompt text 404. A set of Gaussian noise reference noise map groups 402 is generated using preset rules. The initial prompt text 404 and the reference noise map group 402 are then input together into a cross-feature extraction unit in a first feature extraction layer 4062 in a first video generation model 406 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. The cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key positions in the noise map group. The cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information that includes spatiotemporal features in the set of reference noise maps. The cross-feature information, self-attention features, and spatiotemporal feature information are then input together into the first video generation layer 4064 to obtain a reference video 408.After receiving the reference video 408, the user issues a modification command to the reference video 408, sending a video adjustment instruction. Subsequently, based on the adjustment text in the video adjustment instruction, a target prompt text 416 is generated that is slightly modified from the initial prompt text 404. The same reference noise pattern group 402 and target prompt text 416 are then input into the second video generation model 418. Since the second video generation model 418 has the same structure and parameters as the first video generation model 406, the reference noise pattern group 402 and target prompt text 416 are processed in the second feature extraction layer 4184 in the same manner as the reference noise pattern group 402 and the initial prompt text 404 in the first feature extraction layer 4062. Subsequently, the question feature information and key value feature information in the self-attention features generated based on the target prompt text 416 and the reference noise pattern group 402 are replaced with the reference question feature information 410 and reference key value feature information 414 generated based on the initial prompt text 404 and the reference noise pattern group 402. The spatiotemporal feature information in the spatiotemporal feature information generated based on the target prompt text 416 and the reference noise map group 402 is then replaced with the reference spatiotemporal feature information 414 generated based on the initial prompt text 404 and the reference noise map group 402. This replaced feature information is then input into the second video generation layer 4184, obtaining the target video 420 that the user has modified based on the reference video 408. Subsequently, the distinguishing cross-attention feature information and the associated cross-attention feature information output by the cross-feature extraction unit in the second feature extraction layer 4182 are obtained and then fused to obtain target editing mask information 422. Furthermore, after noise processing is performed on the reference video 408 and the target video 420, reference noise video 424 and target noise video 426 are obtained, respectively. The target editing mask information 422, the reference noise video 424, and the target noise video 426 are then integrated to obtain the target edited video 428. By extracting the cross-feature information corresponding to the distinguishing words between the target prompt text and the initial prompt text, as well as the cross-feature information corresponding to the associated words of the distinguishing words, and masking the noise video after noise processing of the reference video and the target video based on the extracted cross-feature information, and then generating the target edited video based on the result of the masking process, the similarity between the target video and the unmodified parts of the reference video can be further improved.Furthermore, generating a target edited video based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video, includes: processing the distinguishing cross-attention feature information and each associated cross-attention feature information according to a preset mask generation rule to generate target edited mask information; processing the reference video and the target video according to a preset noise processing rule to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video; and generating the target edited video based on the target edited mask information, the reference noise video, and the target noise video. In practical applications, the mask generation rule is a data processing rule for generating the target edited mask information, the reference noise video is a noise video carrying some features of the reference video, the target noise video is a noise video carrying some features of the target video, and the target edited video is the edited video generated according to the video adjustment instruction sent by the user. Specifically, the mask generation rule can be understood as a rule for setting the area associated with the distinguishing words in the prompt text as a mask. For example, the aforementioned distinguishing cross-attention feature information and each associated cross-attention feature information are added together, and then the areas where the sum is less than a preset threshold are set to zero. In other words, the smaller the value obtained by adding the distinguishing cross-attention feature information and each associated cross-attention feature information, the greater the correlation. Therefore, setting the area less than the preset threshold to zero is equivalent to setting the area associated with the distinguishing words in the prompt text as a mask. It should be noted that the noise processing rules here are the same as those used to convert the initial video into the reference noise map group, and will not be repeated here. In addition, the target edited video may be generated by first combining the acquired mask information with the reference noise video and the target noise video to generate a masked noise video, and then restoring the masked noise video using a mask restoration model to generate the target edited video. The mask restoration model may be, for example, a VAE (Variational Autoencoder) or the like, which restores a masked noise video to a video. The present disclosure does not impose any restrictions on this. In one embodiment provided by the present disclosure, the specific method of combining the acquired mask information with the reference noise video and the target noise video to generate a masked noise video is as shown in Formula (1): x. t = x src * (1 — mask) + x dst* mask Formula (1) Wherein, Xt is the masked noise video containing the mask, x ". is the reference noise video generated according to the reference video, and Xd " is the target noise video generated according to the target video. By extracting the cross-feature information corresponding to the distinguishing words of the target prompt text and the initial prompt text and the cross-feature information corresponding to the associated words of the distinguishing words, and masking the noise video after the reference video and the target video are noise-processed according to the extracted cross-feature information, and then generating the target edited video according to the result after the masking process, the similarity between the target video and the unmodified parts in the reference video can be further improved. Applying the solution of the embodiment of the present disclosure, a reference video is generated by the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate the target video generated based on the target prompt text. On the basis of not needing to retrain the model, it is ensured that the target video generated by the target prompt text that is slightly changed from the initial prompt text maintains a small difference with the reference video, thereby improving the accuracy of video editing. On this basis, the basic video modification method only modifies the self-attention features and spatiotemporal self-attention feature information, avoiding the incompatibility issues caused by modifying the cross-self-attention features. To further ensure consistency between the generated target video and the reference video, the target video and the reference video are further blended by obtaining the difference between the prompt text before and after the modification and the feature mapping of the corresponding distinguishing words at the cross-self-attention layer. This ensures that the target video differs from the reference video only in the modified areas, while the background unrelated to the modified areas is the same as the reference video, further improving the accuracy of video editing. Referring to Figure 5, a flowchart of a video generation method for a cloud server according to an embodiment of the present disclosure is shown, specifically comprising the following steps: Step 502: Receive a video generation instruction sent by a GAL device and generate initial prompt text based on the video generation instruction. Step 504: Obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, and obtain a target video feature information set generated during the processing of the reference noise map group and the initial prompt text.Considering that, when a user sends text, it is necessary to provide feedback to the user on the video generated based on the text so that the user can decide where to edit based on the previous video. Therefore, before obtaining the target video feature information set, the method further includes: inputting the reference noise map group and the initial prompt text into a first video generation model to obtain a reference video generated by the first video generation model; and, if the video generation instruction carries the initial text information, returning the reference video to the end device. If the video generation instruction sent by the user carries the initial text information, it can be understood that the user needs to generate a video based on the sent initial text information and edit the video. Therefore, if the instruction sent by the user carries the initial text, it is necessary to return the reference video generated based on the initial prompt text to the user's end device so that the user can send a video adjustment instruction to the server based on the reference video. Step 506: Receive the video adjustment instruction sent by the end device and generate the target prompt text based on the video adjustment instruction and the initial prompt text. Step 508: Input the reference noise pattern group, the target prompt text, and the target video feature information set into a second video generation model, obtain the target video generated by the second video generation model, and return the target video to the end-Giga1 device. Considering that if the user is still dissatisfied with the modified video, they can further modify the target video, after returning the target video to the end-Giga1 device, the method further includes: receiving video adjustment instructions from the end-Giga1 device for the target video, generating a target adjusted video corresponding to the target video based on the video adjustment instructions; and sending the target adjusted video to the end-Giga1 device. In practical applications, the target adjusted video can be understood as a video generated by the target video based on the video adjustment instructions. By receiving the video adjustment instructions generated by the user based on the target video and further adjusting the target video, the user can further modify the target video, thereby further improving the user experience. The above is an exemplary embodiment of the video generation method applied to a cloud server. It should be noted that the technical solution of the video generation method applied to the cloud server and the technical solution of the video generation method described above belong to the same concept. For details not described in detail in the technical solution of the video generation method applied to the cloud server, please refer to the description of the technical solution of the video generation method described above.Using the solution of the disclosed embodiments, a reference video is generated using initial prompt text. The target video feature information set generated during the reference video generation process is then used to generate a target video based on the target prompt text. This ensures that the target video generated using a slightly modified target prompt text maintains minimal differences from the reference video, thereby improving video editing accuracy. Furthermore, the generated reference video is returned based on the textual instructions input by the user, allowing the user to issue further editing instructions based on the generated reference video. Furthermore, after completing a video edit, the user can edit the target video again, further enhancing the user experience. 6 , which shows an architecture diagram of a video generation system provided by an embodiment of the present disclosure. The video generation system may include a client 100 and a server 200; the client 100 is configured to send a video generation instruction and a video editing instruction to the server 200; the server 200 is configured to receive the video generation instruction sent by the client 100 and generate an initial prompt text according to the video generation instruction; obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, obtain a reference video generated by the first video generation model, and obtain target video feature values generated by the first video generation model in the process of generating the reference video; if the video generation instruction is an initial text instruction, send the reference video to the client 100; then receive a video adjustment instruction sent by the client 100, and generate a target prompt text according to the video adjustment instruction and the initial prompt text; input the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model; and send the target video to the client 100; the client 100 It is also used to receive the target video and reference video sent by the server 200. Using the solution of the embodiments of the present disclosure, a reference video is generated using the initial prompt text. The target video feature information set generated during the reference video generation process is then used to generate a target video based on the target prompt text. This ensures that the target video generated using the target prompt text, which is slightly different from the initial prompt text, maintains minimal differences from the reference video, thereby improving the accuracy of video editing. The video generation system can include multiple clients 100 and a server 200. The client 100 can be referred to as a client G1 device, and the server 200 can be referred to as a cloud G1 device.Multiple clients 100 can establish a communication connection through the server 200. In a video generation scenario, the server 200 is used to provide video generation services between the multiple clients 100. The multiple clients 100 can act as senders or receivers, communicating through the server 200. Users can interact with the server 200 through the clients 100 to receive data from other clients 100 or send data to other clients 100. In the video generation scenario, users can publish data streams to the server 200 through the clients 100. The server 200 generates target videos and reference videos based on the data streams and pushes the target videos and reference videos to other clients with whom communication has been established. The connection between the clients 100 and the server 200 is established through a network. The network provides the medium for the communication link between the clients 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to undergo encoding, transcoding, compression, and other processing before being published to the server 200. The client 100 can be a browser, an application (APP), a web application such as an H5 (HyperText Markup Languages, version 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be developed based on a software development kit (SDK) for the corresponding service provided by the server 200, such as a real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and rely on the device or certain apps in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, or a personal computer. Electronic devices may also typically be configured with various other types of applications, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social networking platform software, etc. The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support background training for models used on clients, and servers that process data sent by clients.It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data, and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. It is worth noting that the video generation method provided in the embodiments of the present disclosure is generally executed by the server. However, in other embodiments of the present disclosure, the client can also have similar functions to the server and thus execute the video generation method provided in the embodiments of the present disclosure. In other embodiments, the video generation method provided in the embodiments of the present disclosure can also be executed jointly by the client and the server. The following, with reference to FIG7 , uses the application of the video generation method provided in the present disclosure to edit a user's existing video as an example to further illustrate the video generation method. FIG7 shows a flowchart of a process for editing a user's original video, provided by one embodiment of the present disclosure. The process specifically includes the following steps: Step 702: Receive Video 1 sent by the user and generate an initial prompt text "A photo video of a rabbit" based on Video 1. Step 704: Perform DDIM inversion on all frames in Video 1 to obtain a set of potential noise images. Step 706: Input the potential noise image set and the initial prompt text "A photo video of a rabbit" into a first video generation model, and obtain the reference question feature value and reference key feature value output by the self-attention layer of the first self-attention module in the video generation model, as well as the reference spatiotemporal key feature value output by the spatiotemporal self-attention layer. Step 708: Receive an adjustment instruction from the user, "Change the rabbit in Video 1 to black." Step 710: Generate the adjustment prompt text "A photo video of a black rabbit" based on the adjustment instruction. Step 712: Input the above-mentioned potential noise image group and the adjustment prompt text "A photo video of a black rabbit" into the second video generation model, and replace the above-mentioned reference question feature value, reference key value feature value, and reference spatiotemporal key value feature value with the intermediate question feature value and intermediate key value feature value output by the self-attention layer in the second self-attention module in the second video generation model, as well as the intermediate spatiotemporal key value feature value output by the spatiotemporal self-attention layer, to obtain the target self-attention feature value after replacement.Step 714: Input the target self-attention feature value into the diffusion module of the second video generation model to obtain the target video. Applying the solution of the disclosed embodiment, a reference video is generated using the initial prompt text. The target video feature information set generated during the reference video generation process is then used to generate the target video based on the target prompt text. This ensures that the target video generated using a slightly modified target prompt text maintains minimal differences from the reference video without requiring model retraining, thereby improving video editing accuracy. Furthermore, the basic video modification method only modifies the self-attention and spatiotemporal self-attention feature information, avoiding the mismatch caused by modifying the cross-self-attention feature. To further ensure consistency between the generated target video and the reference video, the difference between the prompt text before and after modification and the feature mapping of the corresponding distinguishing words at the cross-self-attention layer are obtained. Based on this, the target video and the reference video are further blended. This ensures that the target video differs from the reference video only in the modified areas, while the background unrelated to the modified areas is the same as the reference video, further improving video editing accuracy. Corresponding to the above-mentioned method embodiments, the present disclosure also provides an embodiment of a video generation device. FIG8 shows a schematic structural diagram of a video generation device provided by one embodiment of the present disclosure. As shown in FIG8 , the device includes: a generation instruction receiving module 802, configured to receive a video generation instruction and generate initial prompt text based on the video generation instruction; a video feature value acquisition module 804, configured to obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, and obtain target video feature values generated during processing of the reference noise map group and the initial prompt text; an adjustment instruction receiving module 806, configured to receive a video adjustment instruction and generate target prompt text based on the video adjustment instruction and the initial prompt text; and a target video generation module 808, configured to input the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model.Optionally, the video generation instruction carries initial text information; the generation instruction receiving module 802 is further configured to: parse the video generation instruction to obtain the initial text information; and generate the initial prompt text based on the initial text information. Accordingly, obtaining the reference noise map group includes: generating the reference noise map group according to a preset noise map group generation rule; and accordingly, receiving the video adjustment instruction includes: receiving the video adjustment instruction corresponding to the reference video. Optionally, the video generation instruction carries initial video information; the generation instruction receiving module 802 is further configured to: parse the video generation instruction to obtain the initial video information; and generate the initial prompt text based on the initial video information, wherein the initial prompt text describes the initial video information. Accordingly, obtaining the reference noise map group includes: generating the reference noise map group corresponding to the initial video information based on the initial video information; and accordingly, receiving the video adjustment instruction includes: receiving the video adjustment instruction corresponding to the initial video information. Optionally, the video feature value acquisition module 804 is further configured to: input the reference noise pattern group and the initial prompt text into a first video generation model to generate a reference video; and obtain a target video feature information set generated by the first video generation model during the process of generating the reference video. Optionally, the first video generation model includes a first feature extraction layer and a first video generation layer; the video feature value acquisition module 804 is further configured to: input the reference noise pattern group and the initial prompt text into the first feature extraction layer to obtain a reference attention feature information set output by the first feature extraction layer; input the reference attention feature information set into the first video generation layer to obtain a reference video output by the first video generation layer. Optionally, the video feature value acquisition module 804 is further configured to: obtain a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature information set includes at least one type of attention feature information; and determine a target video feature information set based on the reference attention feature information set, wherein the target video feature information set includes at least one type of attention feature information.Optionally, the reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information; the video feature value acquisition module 804 is further configured to: determine reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information based on the reference attention feature information set; and determine a target video feature information set based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information. Optionally, the reference self-attention feature information includes reference question feature information, reference key value feature information, and reference value feature information, and the reference spatiotemporal self-attention feature information includes reference spatiotemporal feature information; the video feature value acquisition module 804 is further configured to: determine reference question feature information and reference key value feature information based on the reference self-attention feature information; determine reference spatiotemporal feature information based on the reference spatiotemporal self-attention feature information; and determine the reference question feature information, the reference key value feature information, and the reference spatiotemporal feature information as target feature information. Optionally, the second video generation model includes a second feature extraction layer and a second video generation layer, and the parameters of the second video generation layer are the same as the parameters of the first video generation layer of the first video generation model. The target video generation module 808 is further configured to: input the reference noise pattern group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set; input the target attention feature information set into the second video generation layer to obtain a target video output by the second video generation layer. Optionally, the target video generation module 808 is further configured to: input the reference noise pattern group and the target prompt text into the second feature extraction layer to obtain an intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one type of attention feature information; and generate a target attention feature information set based on the intermediate attention feature information set and the target video feature information set. Optionally, the target video generation module 808 is further configured to: determine target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set; determine the attention feature information to be replaced in the intermediate attention feature information set, wherein the attention feature information to be replaced is the attention feature information corresponding to the target attention feature information; replace each attention feature information to be replaced with the corresponding target attention feature information to generate a target attention feature information set.Optionally, the intermediate attention feature information set includes intermediate question feature information, intermediate key-value feature information, and intermediate spatiotemporal feature information, and the target video feature information set includes reference question feature information, reference key-value feature information, and reference spatiotemporal feature information. The target video generation module 808 is further configured to: if the target attention feature information is the reference question feature information, determine that the intermediate question feature information corresponding to the reference question feature information is the question feature information to be replaced; if the target attention feature information is the reference key-value feature information, determine that the intermediate key-value feature information corresponding to the reference key-value feature information is the key-value feature information to be replaced; if the target attention feature information is the reference spatiotemporal feature information, determine that the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is the spatiotemporal feature information to be replaced. Optionally, the target video generation module 808 is further configured to: replace the question feature information to be replaced with the reference question feature information; replace the key-value feature information to be replaced with the reference key-value feature information; and replace the spatiotemporal feature information to be replaced with the reference spatiotemporal feature information. Optionally, the second feature extraction layer includes a second cross-feature extraction unit; optionally, the video generation device further includes a mask module, which is configured to: determine the prompt text distinguishing word information based on the initial prompt text and the target prompt text; determine at least one associated word information with the prompt word text distinguishing word information based on the initial prompt text and the prompt text distinguishing word information; obtain the distinguishing cross-attention feature information corresponding to the prompt word text distinguishing word information and the associated cross-attention feature information corresponding to each associated word information generated by the second cross-feature extraction unit based on the reference noise map group and the target prompt word text; generate a target edited video based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video. The optional mask module is further configured to: process the distinguishing cross-attention feature information and each associated cross-attention feature information according to preset mask generation rules to generate target editing mask information; process the reference video and the target video according to preset noise processing rules to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video; and generate a target editing video based on the target editing mask information, the reference noise video, and the target noise video. The above is a schematic diagram of a video generation device according to this embodiment.It should be noted that the technical solutions of this video generation device and the aforementioned video generation method share the same concept. For details not described in detail in the technical solution of the video generation device, reference can be made to the description of the technical solution of the aforementioned video generation method. Applying the solution of the disclosed embodiment, a reference video is generated using initial prompt text, and the target video feature information set generated during the reference video generation process is used to generate a target video based on the target prompt text. This ensures that the target video generated using slightly modified target prompt text maintains minimal differences from the reference video without requiring model retraining, thereby improving video editing accuracy. Furthermore, the basic video modification method only modifies self-attention features and spatiotemporal self-attention features, avoiding the incompatibility issues caused by modifying cross-self-attention features. To further ensure consistency between the generated target video and the reference video, the difference between the prompt text before and after modification is obtained, and the feature maps of the corresponding distinguishing words at the cross-self-attention layer are obtained. Based on this, the target video and the reference video are further blended. This ensures that the target video differs from the reference video only in the modified areas, while the background unrelated to the modified areas is the same as the reference video, further improving the accuracy of video editing. Figure 9 shows a block diagram of a computing device 900 according to one embodiment of the present disclosure. The components of computing device 900 include, but are not limited to, a memory 910 and a processor 920. Processor 920 and memory 910 are connected via a bus 930. A database 950 is used to store data. Computing device 900 also includes an access device 940, which enables computing device 900 to communicate via one or more networks 960. Examples of these networks include Public Switched Telephone Network (PSTN) > Local Area Network (LAN) > Wide Area Network (WAN) > Personal Area Network (PAN) or a combination of communication networks such as the Internet.The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface. In one embodiment of the present disclosure, the aforementioned components of the computing device 900 and other components not shown in FIG. 9 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 9 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed. The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), A wearable computing device (e.g., a smartwatch, smart glasses, etc.) or other mobile device, or a stationary computing device such as a desktop computer or personal computer (PC). The computing device 900 may also be a mobile or stationary server. The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned video generation method or the video generation method applied to a cloud server. The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned video generation method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned video generation method or the video generation method applied to a cloud server. An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by the processor, implement the steps of the aforementioned video generation method or the video generation method applied to a cloud server. The above is a schematic diagram of a computer-readable storage medium according to this embodiment.It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the video generation method or the video generation method applied to a cloud server described above. Any details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the video generation method or the video generation method applied to a cloud server described above. One embodiment of the present disclosure also provides a computer program product, including a computer program / instructions. When executed by a processor, the computer program / instructions implement the steps of the video generation method or the video generation method applied to a cloud server described above. The above is an illustrative embodiment of a computer program of this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solution of the video generation method or the video generation method applied to a cloud server described above. Any details not described in detail in the technical solution of the computer program can be found in the description of the technical solution of the video generation method or the video generation method applied to a cloud server described above. The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals. It should be noted that for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present disclosure are not limited by the order of the actions described, as certain steps may be performed in a different order or simultaneously according to the embodiments of the present disclosure.Secondly, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions and modules described are not necessarily required for the embodiments of the present disclosure. In the above embodiments, the descriptions of each embodiment are given with emphasis. For portions not described in detail in a particular embodiment, reference should be made to the relevant descriptions of other embodiments. The preferred embodiments disclosed above are merely provided to illustrate the present disclosure. The alternative embodiments do not describe all details in detail, nor do they limit the invention to the specific implementations described. Obviously, many modifications and variations are possible based on the content of the embodiments of the present disclosure. These embodiments are selected and described in detail in this disclosure to better explain the principles and practical applications of the embodiments of the present disclosure, thereby enabling those skilled in the art to better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
Claims 1. A video generation method, comprising: receiving a video generation instruction, and generating an initial prompt text according to the video generation instruction; Obtaining a reference noise image group, inputting the reference noise image group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of processing the reference noise image group and the initial prompt text; receiving a video adjustment instruction, and generating a target prompt text according to the video adjustment instruction and the initial prompt text; The reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model to obtain a target video generated by the second video generation model.
2. The method of claim 1, wherein Inputting the reference noise image group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of processing the reference noise image group and the initial prompt text, including: inputting the reference noise image group and the initial prompt text into the first video generation model to generate a reference video; obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video.
3. The method of claim 2, wherein: The first video generation model includes a first feature extraction layer and a first video generation layer; Inputting the reference noise image group and the initial prompt text into a first video generation model to generate a reference video, including: inputting the reference noise image group and the initial prompt text into the first feature extraction layer, and obtaining a reference attention feature information set output by the first feature extraction layer; The reference attention feature information set is input into the first video generation layer to obtain the reference video output by the first video generation layer.
4. The method of claim 3, wherein: Obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video, including: obtaining a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature 29 The feature information set includes at least one attention feature information; and a target video feature information set is determined according to the reference attention feature information set, wherein the target video feature information set includes at least one attention feature information.
5. The method of claim 4, wherein: The reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information; Determining a target video feature information set according to the reference attention feature information set, comprising: determining reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information according to the reference attention feature information set; Determine a target video feature information set based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information.
6. The method of claim 5, wherein: The reference self-attention feature information includes reference question feature information, reference key value feature information and reference value feature information, and the reference spatiotemporal self-attention feature information includes reference spatiotemporal feature information; According to the reference self-attention feature information and the reference spatiotemporal self-attention feature information, a target video feature information set is determined, including: determining reference question feature information and reference key value feature information according to the reference self-attention feature information; determining reference spatiotemporal feature information according to the reference spatiotemporal self-attention feature information; determining the reference question feature information, the reference key value feature information and the reference spatiotemporal feature information as target feature information.
7. The method of claim 1, wherein The second video generation model includes a second feature extraction layer and a second video generation layer, and the parameters of the second video generation layer are the same as the parameters of the first video generation layer of the first video generation model; the reference noise map group, the target prompt text and the target video feature information set are input into the second video generation model to obtain the target video generated by the second video generation model, including: inputting the reference noise map group, the target prompt text and the target video feature information set into the second feature extraction layer to obtain the target attention feature information set; inputting the target attention feature information set into the second video generation layer to obtain the target video output by the second video generation layer. 30 8. The method according to claim 7, wherein: Inputting the reference noise map group, the target prompt text and the target video feature information set into the second feature extraction layer to obtain the target attention feature information set, including: inputting the reference noise map group and the target prompt text into the second feature extraction layer to obtain the intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one kind of attention feature information; generating a target attention feature information set according to the intermediate attention feature information set and the target video feature information set.
9. The method of claim 8, wherein According to the intermediate attention feature information set and the target video feature information set, a target attention feature information set is generated, including: determining target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set; determining attention feature information to be replaced in the intermediate attention feature information set, wherein the attention feature information to be replaced is the attention feature information corresponding to the target attention feature information; replacing each attention feature information to be replaced with the corresponding target attention feature information to generate a target attention feature information set.
10. The method of claim 9, wherein: The intermediate attention feature information set includes intermediate question feature information, intermediate key value feature information, and intermediate spatiotemporal feature information; the target video feature information set includes reference question feature information, reference key value feature information, and reference spatiotemporal feature information; Determining the attention feature information to be replaced in the intermediate attention feature information set includes: when the target attention feature information is the reference question feature information, determining the intermediate question feature information corresponding to the reference question feature information as the question feature information to be replaced; When the target attention feature information is the reference key-value feature information, the intermediate key-value feature information corresponding to the reference key-value feature information is determined to be the key-value feature information to be replaced; when the target attention feature information is the reference spatiotemporal feature information, the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is determined to be the spatiotemporal feature information to be replaced.
11. The method according to claim 10, wherein: Replacing each to-be-replaced attention feature information with a corresponding target attention feature information set includes: replacing the to-be-replaced question feature information with the reference question feature information; The key value feature information to be replaced is replaced with the reference key value feature information; and the spatiotemporal feature information to be replaced is replaced with the reference spatiotemporal feature information.
12. The method of claim 7, wherein: The second feature extraction layer includes a second cross-feature extraction unit; after obtaining the target video generated by the second video generation model, the method further includes: determining the prompt text distinguishing word information based on the initial prompt text and the target prompt text; determining at least one associated word information with the prompt word text distinguishing word information based on the initial prompt text and the prompt text distinguishing word information; obtaining the distinguishing cross-attention feature information corresponding to the prompt word text distinguishing word information and the associated cross-attention feature information corresponding to each associated word information generated by the second cross-feature extraction unit based on the reference noise map group and the target prompt word text; generating a target edited video based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video.
13. The method according to claim 12, wherein: Generating a target editing video based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video, including: processing the distinguishing cross-attention feature information and each associated cross-attention feature information according to a preset mask generation rule to generate target editing mask information; Processing the reference video and the target video according to a preset noise processing rule to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video; A target edited video is generated according to the target edit mask information, the reference noise video, and the target noise video.
14. The method according to claim 1, wherein: The video generation instruction carries initial text information; receiving the video generation instruction and generating initial prompt text according to the video generation instruction, including: parsing the video generation instruction to obtain the initial text information; generating the initial prompt text based on the initial text information; correspondingly, obtaining the reference noise map group, including: The reference noise map group is generated by presetting a noise map group generation rule; accordingly, receiving a video adjustment instruction includes: receiving the video adjustment instruction corresponding to the reference video.
15. The method according to claim 1, wherein: The video generation instruction carries initial video information; Receiving a video generation instruction and generating an initial prompt text according to the video generation instruction includes: parsing the video generation instruction to obtain the initial video information; generating the initial prompt text based on the initial video information, wherein the initial prompt text is used to describe the initial video information; correspondingly, obtaining a reference noise map group includes: generating the reference noise map group corresponding to the initial video information based on the initial video information; correspondingly, receiving a video adjustment instruction includes: receiving the video adjustment instruction corresponding to the initial video information.
16. A video generation method, applied to a cloud server, comprising: The method includes: receiving a video generation instruction sent by a Gamma-1 device at the receiving end, and generating an initial prompt text based on the video generation instruction; obtaining a reference noise pattern group, inputting the reference noise pattern group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated during processing of the reference noise pattern group and the initial prompt text; receiving a video adjustment instruction sent by the Gamma-1 device at the receiving end, and generating a target prompt text based on the video adjustment instruction and the initial prompt text; The reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model, a target video generated by the second video generation model is obtained, and the target video is returned to the terminal device.
17. The method according to claim 16, wherein: Before obtaining the target video feature information set, the method further includes: inputting the reference noise map group and the initial prompt text into a first video generation model to obtain a reference video generated by the first video generation model; and returning the reference video to the terminal if the video generation instruction carries initial text information. 33 Example equipment.
18. The method of claim 16, wherein: After returning the target video to the end Gigabit 1 device, the method further includes: receiving a video adjustment instruction for the target video from the end Gigabit 1 device, and generating a target adjusted video corresponding to the target video according to the video adjustment instruction; and sending the target adjusted video to the end Gigabit 1 device.
19. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1-15 or 16-18 are implemented.
20. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15 or 16 to 18.
21. A computer program product comprising a computer program / instructions, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15 or 16 to 18. 34
Citation Information
Patent Citations
Video generation method and server
CN116233491A
Video reconstruction model training method, video reconstruction method, device and equipment
CN116757970A
Video generation model training method and device, equipment and storage medium
CN117499711A