Video generation method, computing device, computer readable storage medium and computer program product
By receiving video generation instructions and adjustment instructions, and using reference noise map groups and initial prompt text to generate target videos, the problem of inaccurate video editing in the prior art is solved, and the accuracy of video editing is improved.
Patent Information
- Application Number
- CN202410172465.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, when users use slightly changed text to edit the original video, the generated video often has a large difference from the original video, resulting in inaccurate content modification.
The initial prompt text is generated by receiving the video generation instructions, and the first video generation model is inputted using the reference noise map group and the initial prompt text, and the target video feature information set is obtained, and the target video text is generated in combination with the video adjustment instructions, and input it to the second video generation model to generate the target video to ensure the accuracy of video editing.
It is achieved that when the initial prompt text changes slightly, the generated target video and the reference video are kept small, which improves the accuracy of video editing.
Smart Images

Figure CN120448588A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a video generation method. Background Art
[0002] With the rapid development of generative AI, text-based video editing has become one of the most compelling and widely used techniques in this field. This technique uses text as input to edit videos, aiming to enhance the visual appeal and creativity of video content. In particular, with the advancement of text-to-image generation models, text-to-video generation technology has also significantly improved.
[0003] Currently, when users use slightly modified text to regenerate a video based on the original video description, the resulting video often differs significantly from the original, limiting the ability to modify the original video content. Therefore, to address this issue, a video generation method is needed that allows precise editing of the original video using text instructions. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a video generation method, a video generation method applied to a cloud server. One or more embodiments of this specification also relate to a video generation apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a video generation method is provided, including:
[0006] Receive a video generation instruction, and generate an initial prompt text according to the video generation instruction;
[0007] Obtaining a reference noise image group, inputting the reference noise image group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video;
[0008] receiving a video adjustment instruction, and generating a target prompt text according to the video adjustment instruction and the initial prompt text;
[0009] The reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model to obtain a target video generated by the second video generation model.
[0010] According to a second aspect of the embodiments of this specification, a video generation method applied to a cloud server is provided, including:
[0011] The receiving end-side device generates a video generation instruction, and generates an initial prompt text according to the video generation instruction;
[0012] Obtaining a reference noise image group, inputting the reference noise image group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video;
[0013] receiving a video adjustment instruction sent by a terminal-side device, and generating a target prompt text according to the video adjustment instruction and the initial prompt text;
[0014] The reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model, a target video generated by the second video generation model is obtained, and the target video is returned to the end-side device. According to a third aspect of the embodiment of this specification, a computing device is provided, comprising:
[0015] memory and processor;
[0016] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the above-mentioned video generation method and the steps of the video generation method applied to the cloud server are implemented.
[0017] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned video generation method and the video generation method applied to the cloud server are implemented.
[0018] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned video generation method and the steps of the video generation method applied to a cloud server.
[0019] One embodiment of the present specification implements receiving a video generation instruction and generating an initial prompt text according to the video generation instruction; obtaining a reference noise map group, inputting the reference noise map group and the initial prompt text into a first video generation model, obtaining a reference video generated by the first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video; receiving a video adjustment instruction, and generating a target prompt text according to the video adjustment instruction and the initial prompt text; inputting the reference noise map group, the target prompt text and the target video feature information set into a second video generation model, and obtaining a target video generated by the second video generation model.
[0020] By applying the solution of the embodiments of this specification, a reference video is generated through the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate a target video based on the target prompt text, thereby ensuring that the target video generated by the target prompt text that is slightly different from the initial prompt text maintains a small difference from the reference video, thereby improving the accuracy of video editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a video generation method provided by one embodiment of this specification;
[0022] Figure 2 This is a flowchart of a method for generating video editing provided by an embodiment of this specification;
[0023] Figure 3 This is a flowchart of an embodiment of the present disclosure for editing an original video.
[0024] Figure 4 This is a flowchart of an additional mask video editing process provided by an embodiment of this specification;
[0025] Figure 5 This is a flow chart of a video generation method applied to a cloud server provided by one embodiment of this specification;
[0026] Figure 6 This is an architecture diagram of a video generation system provided by one embodiment of this specification;
[0027] Figure 7 This is a flowchart of a process for editing a user's original video provided by an embodiment of this specification;
[0028] Figure 8 This is a structural diagram of a video generation device provided by an embodiment of this specification;
[0029] Figure 9 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0030] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0031] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0032] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0033] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0034] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.
[0035] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0036] First, the terms involved in one or more embodiments of this specification are explained.
[0037] Deterministic Denoising Diffusion Implicit Models (DDIM): A generative model based on a diffusion model, characterized by its deterministic approach to sample generation. Unlike traditional random diffusion models, the output of each step in the denoising process is deterministic, rather than random. This approach not only speeds up the sampling process but also enables more accurate control of the quality of generated samples. DDIM gradually recovers clear data from noisy data through a series of inverse diffusion steps and is widely used in fields such as image generation and super-resolution.
[0038] Stable Diffusion is an advanced deep learning model specifically designed for generating high-quality images. It combines diffusion models with deep learning techniques to generate images by gradually removing noise, thereby enabling the generation of complex and high-resolution images. The key advantage of Stable Diffusion is its ability to stably generate images with rich detail and realism while maintaining high computational efficiency. This has attracted widespread attention in fields such as image synthesis, artistic creation, and media editing.
[0039] Variational Autoencoder (VAE): A deep learning model used for unsupervised learning tasks, particularly data generation. It consists of two parts: an encoder and a decoder. The encoder transforms the input data into a representation in a latent space, while the decoder reconstructs the data from this latent representation. VAEs are trained by minimizing the reconstruction error and the difference between the latent representation and a prior distribution (typically a Gaussian distribution). This enables VAEs to both efficiently generate data and learn complex data distributions. VAEs are widely used in image generation, denoising, style transfer, and other fields.
[0040] In this specification, a video generation method is provided, a video generation method applied to a cloud server. This specification also involves a video generation device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0041] See also Figure 1 , Figure 1 A flowchart of a video generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0042] Step 102: Receive a video generation instruction, and generate an initial prompt text according to the video generation instruction.
[0043] In practical applications, the video generation instruction is an instruction for generating a reference video, and the initial prompt text is text describing the reference video. Specifically, the video generation instruction can be understood as an instruction for generating the reference video that the user wants to modify. For example, after a user uploads the original video: Video 1 on the video modification service page and clicks "Modify with this video," the instruction sent to the server is a video generation instruction. Alternatively, the instruction sent to the server after a user enters "Generate a video for me with the focus on rabbits" and clicks "Send" is another example. This specification does not impose any restrictions on this.
[0044] The initial prompt text can be understood as a standardized text used to describe the target video that the user wants to modify. For example, based on the original video uploaded by the user: Video 1, a description of the video is generated as "a video of a rabbit eating watermelon"; for example, the text input by the user "Generate a video for me that focuses on rabbits" is parsed to generate a standardized description as "a photo video of a rabbit", etc. This manual does not impose any restrictions on this.
[0045] By receiving video generation instructions, a standardized initial prompt text can be generated according to the video generation instructions, which can make the video generated by the video generation model more accurate. The standardized initial prompt text can also enable subsequent adjustments to the video to be achieved through another standardized text, thereby improving the accuracy of video editing.
[0046] Considering that the user will first generate a video in the form of text and then modify the video based on the video generated according to the initial text, the video generation instruction carries the initial text information;
[0047] Furthermore, receiving a video generation instruction and generating an initial prompt text according to the video generation instruction include:
[0048] Parsing the video generation instruction to obtain the initial text information;
[0049] generating the initial prompt text based on the initial text information;
[0050] Accordingly, a reference noise map group is obtained, including:
[0051] Generating the reference noise map group by presetting a noise map group generation rule;
[0052] Accordingly, receiving a video adjustment instruction includes:
[0053] The video adjustment instruction corresponding to the reference video is received.
[0054] In actual applications, the initial text information is text information directly input by the user, the reference noise map group is a group of noise maps used to generate a reference video, the noise map group generation rule is a rule for generating a noise map group when the video generation instruction sent by the user carries the initial text information, the reference video is a video generated based on the reference noise map group and the initial prompt text, and the video adjustment instruction is an instruction for adjusting the reference video.
[0055] Specifically, the initial text can be understood as non-standard information directly input by the user, for example, "Generate a video for me that focuses on rabbits", "Give me a video of a rabbit jumping", etc. In the case where the video generation instruction carries initial text information, the initial prompt text obtained can be understood as text information that standardizes the initial text sent by the user and highlights the key points in the initial text. For example, if the initial text sent by the user is "Generate a video for me that focuses on rabbits", the initial prompt text determined after parsing the initial text is "A video of a rabbit photo". This initial prompt text highlights the focus of the rabbit in the initial text and understands "focus on rabbits" as "rabbit photos", making the generated video more accurate.
[0056] In the case where the video generation instruction carries initial text information, the reference noise map group obtained can be understood as a group of random noise images, such as a group of Gaussian noise maps, a group of uniform noise maps, a group of speckle noise maps, etc., and this specification does not impose any restrictions on this. The noise map group generation rule can be understood as a rule for generating a group of random noise maps, such as a Gaussian noise map generation rule, a uniform noise map generation rule, etc. The reference video can be understood as a video generated by the initial prompt text, which meets the description of the initial prompt text. For example, when the initial prompt text is "a photo video of a rabbit", the reference video generated is Video A, which is a video of a rabbit resting on the grass, and there is only a rabbit on the grass.
[0057] The video generation instruction sent by the user carries the initial text information. It can be understood that the user needs to generate a video based on the initial text information first, and then edit the generated video. Figure 2 , Figure 2 A flowchart for generating video editing provided by an embodiment of the present specification is shown. The flowchart includes receiving an initial text message sent by a user, parsing the initial text message into a standardized initial prompt text 204, and generating a set of Gaussian noise reference noise map groups 202 using preset rules. The initial prompt text 204 and the reference noise map group 202 are then input together into a cross-feature extraction unit of a first feature extraction layer 2062 of a first video generation model 206 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. The cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key positions in the noise map group. The cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information containing spatiotemporal features in the reference noise map group. The cross-feature information, self-attention features, and spatiotemporal feature information are then input together into the first video generation layer 2064 to obtain a reference video 208.
[0058] After receiving the reference video 208, the user gives a modification command to the reference video 208 and sends a video adjustment instruction. Then, based on the adjustment text in the video adjustment instruction, the target prompt text 216 with slight changes from the initial prompt text 204 is generated. Then, the same reference noise map group 202 and target prompt text 216 are input into the second video generation model 218. Since the second video generation model 218 has the same structure and the same parameters as the first video generation model 206, the reference noise map group 202 and the target prompt text 216 in the second feature extraction layer 2182 have the same operation as the reference noise map group 202 and the initial prompt text 204 in the first feature extraction layer 2062. Then, the question feature information and key value feature information in the self-attention feature generated according to the target prompt text 216 and the reference noise map group 202 are replaced with the reference question feature information 210 and the reference key value feature information 212 generated according to the initial prompt text 204 and the reference noise map group 202, and then the spatiotemporal feature information in the spatiotemporal feature information generated according to the target prompt text 216 and the reference noise map group 202 is replaced with the reference spatiotemporal feature information 214 generated according to the initial prompt text 204 and the reference noise map group 202, and the replaced feature information is input into the second video generation layer 2184 to obtain the target video 220 obtained by the user by modifying the reference video 208.
[0059] By receiving a video generation instruction carrying initial text information, a video is first generated in the form of text, and then the video is modified based on the video generated according to the initial text, thereby enabling the user to edit the video.
[0060] Considering that users may directly upload an original video and modify the video based on the original video, the video generation instruction carries the initial video information;
[0061] Furthermore, receiving a video generation instruction and generating an initial prompt text according to the video generation instruction include:
[0062] Parsing the video generation instruction to obtain the initial video information;
[0063] generating the initial prompt text based on the initial video information, wherein the initial prompt text is used to describe the initial video information;
[0064] Accordingly, a reference noise map group is obtained, including:
[0065] Based on the initial video information, generating the reference noise map group corresponding to the initial video information;
[0066] Accordingly, receiving a video adjustment instruction includes:
[0067] The video adjustment instruction corresponding to the initial video information is received.
[0068] In practical applications, the initial video information is the video uploaded by the user. Specifically, the initial video information can be understood as the video that the user wants to edit.
[0069] When the video generation instruction carries the initial video information, the reference noise map group obtained can be understood as a group of noise maps generated by performing noise processing on each frame in the initial video information. Specifically, the noise processing method for the frames in the initial video information can be a reverse DDIM (Deterministic Denoising Diffusion Implicit Models), etc. This specification does not impose any restrictions on this.
[0070] When the video generation instruction carries the initial video information, the initial prompt text obtained can be understood as summarizing the initial video information, obtaining a text describing the initial video. For example, if the video uploaded by the user is a video of a rabbit eating watermelon on a table, the initial prompt text corresponding to the video is "a video of a rabbit eating watermelon on a table."
[0071] By obtaining a set of noise maps from the initial video information and a descriptive text for the initial video information, the reference video generated based on the above data can be made almost identical to the initial video. The characteristic value generated when generating the reference video can also be understood as the characteristic value of the initial video information. The target video subsequently generated using the characteristic value will not be much different from the initial video.
[0072] The video generation instruction sent by the user carries the initial video information. It can be understood that the user needs to edit the original video. However, since the original video has not been generated by the first video generation model, it is necessary to first use the first video generation model to generate a reference video that is basically the same as the original video to obtain some features in the original video. Figure 3 , Figure 3A schematic diagram of a process for editing an original video provided by an embodiment of the present specification is shown. The process includes receiving initial video information 302 sent by a user, performing noise processing on each frame in the initial video information 302, generating a set of reference noise map groups 304, parsing the initial video information 302 into initial prompt text 306 describing the initial video information, and then inputting the initial prompt text 306 and the reference noise map group 304 together into a cross-feature extraction unit in a first feature extraction layer 3082 in a first video generation model 308 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. The cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key positions in the noise map group. The cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information containing the spatiotemporal features in the reference noise map group. The cross-feature information, self-attention features, and spatiotemporal feature information are then input into the first video generation layer 3084 to obtain a reference video 310.
[0073] The user gives a modification command to the initial video information 302 and sends a video adjustment instruction. Then, based on the adjustment text in the video adjustment instruction, a target prompt text 318 with a slight change from the initial prompt text 306 is generated. Then, the same reference noise map group 304 and target prompt text 318 are input into the second video generation model 320. Since the second video generation model 320 has the same structure and the same parameters as the first video generation model 308, the reference noise map group 304 and the target prompt text 318 have the same operation in the second feature extraction layer 3202 as the reference noise map group 304 and the initial prompt text 306 in the first feature extraction layer 3082. The problem feature information and key value feature information in the self-attention features generated by the target prompt text 318 and the reference noise map group 304 are replaced by the reference problem feature information 312 and the reference key value feature information 314 generated according to the initial prompt text 306 and the reference noise map group 304, and then the spatiotemporal feature information in the spatiotemporal feature information generated according to the target prompt text 318 and the reference noise map group 304 is replaced by the reference spatiotemporal feature information 316 generated according to the initial prompt text 306 and the reference noise map group 304, and the replaced feature information is input into the second video generation layer 3204 to obtain the target video 322 obtained by the user through modification based on the reference video 310.
[0074] By receiving a video generation instruction carrying initial video information, the user modifies the video based on the original video, thereby realizing the user's editing of the original video.
[0075] Step 104: Obtain a reference noise image group, input the reference noise image group and the initial prompt text into a first video generation model, and obtain a target video feature information set generated by the first video generation model in the process of processing the reference noise image group and the initial prompt text.
[0076] In practical applications, the first video is a model for extracting partial features of the video to be modified, and the target video feature information set is partial features of the video to be modified.
[0077] Specifically, the initial prompt text describing the video to be modified is input into the first video generation model. The first video generation model can extract some features of the video to be modified by processing the above initial prompt text. By using the extracted partial features to generate the target video, the generated target video can be made more similar to the unmodified parts of the video to be modified, thereby improving the accuracy of the video modification.
[0078] Furthermore, the reference noise image group and the initial prompt text are input into a first video generation model, and a target video feature information set generated by the first video generation model in the process of processing the reference noise image group and the initial prompt text is obtained, including:
[0079] Inputting the reference noise image group and the initial prompt text into a first video generation model to generate a reference video;
[0080] Obtain a target video feature information set generated by the first video generation model in the process of generating the reference video.
[0081] In practical applications, the reference video is the video to be modified.
[0082] In other words, the first video generation model can be further understood as a model for extracting partial features of a reference video, and the reference video can be understood as the video to be modified. Inputting the initial prompt text describing the video to be modified into the first video generation model to generate the reference video can be understood as extracting partial features of the reference video during the reference video generation process. By using these extracted partial features to generate the target video, the generated target video can be made more similar to the unmodified parts of the reference video, thereby improving the accuracy of the video modification.
[0083] The target video feature information set can be understood as partial features of the reference video. The target video feature set may include one or more features of the reference video, such as temporal attention features representing the temporal features of each entity in the reference video, reference spatial attention features representing the spatial features of each entity in the reference video, reference spatiotemporal attention features representing the temporal and spatial fusion features of each entity in the reference video, reference self-attention features representing the relationship between each pixel in the reference video, reference cross-attention features representing the association relationship between the reference video and the initial prompt text, etc. This specification does not impose any restrictions on this.
[0084] In one embodiment provided in this specification, an initial prompt text "a rabbit photo video" and a set of Gaussian noise images are input into a first video generation model, and the reference video obtained is a video of a white rabbit resting on the grass. In the process of generating the video, spatiotemporal attention features representing the temporal and spatial fusion features of each entity in the video and self-attention features representing the relationship between each pixel in the video are obtained.
[0085] By extracting features corresponding to the reference video generated during the generation of the reference video, and then using the acquired features to generate the target video, the generated target video can be made more similar to the unmodified parts of the reference video.
[0086] Considering that generating a video requires automatically discovering and extracting useful information or features from prompt text and noisy images to facilitate further processing and analysis, the first video generation model includes a first feature extraction layer and a first video generation layer;
[0087] Furthermore, the reference noise image group and the initial prompt text are input into a first video generation model to generate a reference video, including:
[0088] Inputting the reference noise image group and the initial prompt text into the first feature extraction layer, obtaining a reference attention feature information set output by the first feature extraction layer;
[0089] The reference attention feature information set is input into the first video generation layer to obtain the reference video output by the first video generation layer.
[0090] In actual applications, the first feature extraction layer is used to extract the features of the reference video, the first video generation layer is used to generate a deep learning network layer of the reference video based on the features extracted by the first feature extraction layer, and the reference attention feature information set is used to input into the video generation layer to obtain the reference video.
[0091] Specifically, the first feature extraction layer can be understood as a computing layer that generates the features of the reference video. This layer can be a feature extraction layer with a neural network structure or a feature extraction layer without a neural network structure. This specification does not impose any restrictions on the results of this layer. For example, based on the input initial prompt text and a set of noise maps, the self-attention features of the relationship between each pixel in the video to be generated, the cross-attention features of the association relationship between the video to be generated and the initial prompt text, etc. are extracted. This specification does not impose any restrictions on this.
[0092] The first video generation layer can be understood as a deep learning network layer that generates the corresponding video based on the extracted video features. For example, DDIM, Stabilized Diffusion Model (stable diffusion model), etc. can generate a network model for the video. This specification does not impose any restrictions on this.
[0093] The reference attention feature information set can be understood as all the features of the reference video to be generated, which includes all the features about the reference video extracted by the first feature extraction layer, such as temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, etc. This specification does not impose any restrictions on this.
[0094] Through the first feature extraction layer, all features of the reference video are extracted. Features that the user will not modify can be obtained from the extracted features and used to generate a target video, so that the generated target video can be more similar to the unmodified parts of the reference video.
[0095] Furthermore, obtaining a target video feature information set generated by the first video generation model in the process of generating the reference video includes:
[0096] Obtaining a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature information set includes at least one type of attention feature information;
[0097] A target video feature information set is determined based on the reference attention feature information set, wherein the target video feature information set includes at least one type of attention feature information.
[0098] Specifically, the target video feature information set can be understood as the features selected from the reference attention feature information set representing all features of the reference video, representing features that the user would not modify. By obtaining these features that the user would not modify and using them to generate the target video, the generated target video can be made more similar to the unmodified parts of the reference video.
[0099] Preferably, considering that users usually do not modify the spatiotemporal relationship between entities in the reference video, as well as the relationship between elements in the reference video, the reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information;
[0100] Furthermore, determining a target video feature information set based on the reference attention feature information set includes:
[0101] Determine reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information according to the reference attention feature information set;
[0102] A target video feature information set is determined based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information.
[0103] In practical applications, the reference self-attention feature information is the feature value representing the relationship between each pixel in the reference video, and the reference spatiotemporal self-attention feature information is the feature value representing the temporal and spatial relationships between entities in the reference video.
[0104] Specifically, the reference self-attention feature information can be understood as the characteristics of the correlation relationship between each pixel in the reference video itself. For example, the reference video is a rabbit resting on the grass, where the rabbit's two eyes are pixels a and pixel b respectively, and pixel c is in the sky of the video. Then the correlation relationship between pixel a and pixel b is considered strong, and pixel c has a weak correlation with pixel a or pixel b, which is reflected in the feature matrix as point<a,b> point<b,a> Both are 0.9, point<c,a> ,point<c,b> The value is 0.1, and the larger the number, the stronger the correlation.
[0105] The reference spatiotemporal self-attention feature information can be understood as the spatiotemporal relationship between the time of each entity in the reference video. For example, in the reference video of a rabbit resting on the grass, since the rabbit itself changes over time, the temporal relationship of the rabbit itself is strongly correlated. Since the rabbit is resting on the grass, as the rabbit moves, the grass under the rabbit also moves, so the spatial relationship between the rabbit and the grass under the rabbit in the grass is strongly correlated. However, since the rabbit's movement does not affect the sky, and the sky is not related to time, the spatial and temporal relationships between the rabbit and the sky are both weakly correlated. In other words, the reference spatiotemporal self-attention feature information can also be understood as the motion characteristics of the entities in the reference video.
[0106] By obtaining the features of the association between the pixels of the video itself that the user will not modify, as well as the motion features of the entities in the video, and using the above features to generate a target video, the generated target video can be made more similar to the unmodified parts of the above reference video.
[0107] Taking into account that, in the process of generating the reference video, the information of the initial prompt text and the reference noise map features is included, so that the background of the target video can be kept consistent with that of the reference video in the subsequent process of generating the target video, the reference self-attention feature information includes reference question feature information, reference key value feature information and reference value feature information, and the reference spatiotemporal self-attention feature information includes reference spatiotemporal feature information;
[0108] Furthermore, determining a target video feature information set based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information includes:
[0109] Determining reference question feature information and reference key value feature information based on the reference self-attention feature information;
[0110] Determining reference spatiotemporal feature information based on the reference spatiotemporal self-attention feature information;
[0111] The reference question feature information, the reference key value feature information, and the reference spatiotemporal feature information are determined to be target feature information.
[0112] In practical applications, the reference question feature information and the reference key value feature information are feature information used to calculate the relationship between each pixel in the reference video, and the reference spatiotemporal feature information is feature information representing the temporal and spatial fusion features between entities in the reference video.
[0113] Specifically, the reference question feature information can be understood as the feature information that uses each pixel in the reference video as the question, and the reference key value feature information can be understood as the feature information that uses each pixel in the reference video as the key value. The feature information used as the question is dot-multiplied with the feature information used as the key value to determine the relationship between each pixel in the reference video. Subsequently, this is added to the value feature information used as the answer for each pixel in the reference video to strengthen the relationship between each pixel in the reference video. In this way, reference self-attention feature information representing the relationship between each pixel in the reference video can be obtained.
[0114] The reference spatiotemporal feature information can be understood as the temporal and spatial relationship between various entities in the reference video, and can also be understood as feature information representing the motion characteristics of the reference video.
[0115] By obtaining problem feature information and key value feature information of videos that users usually do not modify, the problem that the modification is not reflected when the user modifies the relationship between pixels in the video can be avoided. The motion features of the entities in the video are obtained, and the above features are used to generate the target video, so that the generated target video can be more similar to the unmodified parts of the above reference video.
[0116] Step 106: Receive a video adjustment instruction, and generate a target prompt text according to the video adjustment instruction and the initial prompt text.
[0117] In practical applications, a video adjustment instruction is an instruction for adjusting a video. A video adjustment instruction can be understood as an instruction sent by a user to adjust the video to be adjusted. For example, if a user wants to adjust an uploaded initial video, the video adjustment instruction is an instruction to adjust the initial video. If a user wants to adjust a reference video generated based on text, the adjustment instruction is an instruction to adjust the reference video, and so on.
[0118] The target prompt text is the prompt text used to generate the target video. The target prompt text can be understood as a text that is slightly changed from the initial text. For example, the initial prompt text is "a photo video of a rabbit", and the slightly changed target prompt text is "a photo video of a black rabbit".
[0119] In one embodiment provided in this specification, a video adjustment instruction containing a user edit of "turn the rabbit black" is received, and the target prompt text "a photo video of a black rabbit" is generated in combination with the previous initial prompt text "a photo video of a rabbit".
[0120] By combining the target prompt text generated by the initial prompt text, the initial prompt text is slightly changed according to the needs put forward by the user, so that the generated video is more similar to the unmodified parts of the original reference video.
[0121] Step 108: Input the reference noise image group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model.
[0122] In practical applications, the second video generation model is the target of generating a target video, and the target video is generated according to the target prompt text.
[0123] Specifically, the second video generation model can be understood as a video generation model with the same structure and parameters as the first video generation model, and is used to generate the target video in combination with the reference noise image group and the target prompt text. The target video can be understood as a video that has been modified with some entities or features in the reference video. For example, the reference video before modification is a white rabbit sleeping on the grass, and the modified target video is a black rabbit sleeping on the grass. The sleeping posture of the black rabbit and the weather in the sky are the same after modification.
[0124] By processing the same reference noise map group with a second video generation model using the same parameters as the first video generation model, the modified target video can be effectively made consistent with the reference video before modification. Subsequently, by replacing some features of the reference video generated in the process of generating the reference video with some features used to generate the target video, the generated target video can be made more similar to the unmodified parts of the original reference video.
[0125] Considering that two similar videos are to be generated, the second video generation model also includes a second feature extraction layer and a second video generation layer, and the parameters of the second video generation layer are the same as the parameters of the first video generation layer of the first video generation model;
[0126] Furthermore, the reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model to obtain a target video generated by the second video generation model, including:
[0127] Inputting the reference noise image group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set;
[0128] The target attention feature information set is input into the second video generation layer to obtain the target video output by the second video generation layer.
[0129] In actual applications, the second feature extraction layer is used to extract the features of the target video, and the second video generation layer is used to generate a deep learning network layer of the target video based on the features extracted by the second feature extraction layer and some features extracted from the first feature extraction layer. The target attention feature information set is used to input into the video generation layer to obtain the target video.
[0130] Specifically, the second feature extraction layer can be understood as a computing layer that generates the features of the target video. This layer can be a feature extraction layer with a neural network structure or a feature extraction layer without a neural network structure. This specification does not impose any restrictions on the results of this layer. If this layer is a feature extraction layer with a neural network structure, then the parameters of the common neural network structure in this layer are the same as those of the first feature extraction layer mentioned above. For example, based on the input target prompt text and a set of noise maps, the self-attention features of the relationship between each pixel in the video to be generated, the cross-attention features of the association relationship between the video to be generated and the target prompt text, etc. are extracted. This specification does not impose any restrictions on this.
[0131] The second video generation layer can be understood as a deep learning network layer that generates the corresponding video based on the extracted video features, and the parameters of the neural network layer are the same as those of the above-mentioned first video generation layer. For example, DDIM, StabilDiffusion (stable diffusion model), etc. can generate a network model for the video. This manual does not impose any restrictions on this.
[0132] The target attention feature information set can be understood as all the features of the target video to be generated. It includes the features extracted by the second feature extraction layer, and then the features in the above-mentioned target video feature set are replaced with their corresponding features. The features obtained include, for example, temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, etc. This specification does not impose any restrictions on this.
[0133] Through the second feature extraction layer, some features of the target video are extracted, and then the above-mentioned features that the user will not modify are replaced with the corresponding features obtained through the second feature extraction layer. The target video is generated by the replaced feature set, which can make the generated target video more similar to the unmodified parts of the above-mentioned reference video.
[0134] Furthermore, the reference noise image group, the target prompt text, and the target video feature information set are input into the second feature extraction layer to obtain a target attention feature information set, including:
[0135] Inputting the reference noise image group and the target prompt text into the second feature extraction layer, and obtaining an intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one type of attention feature information;
[0136] A target attention feature information set is generated based on the intermediate attention feature information set and the target video feature information set.
[0137] In practical applications, the intermediate attention feature information set is the set of all features output by the second feature extraction layer. Specifically, the intermediate attention feature information set can be understood as the features of the video generated solely based on the target prompt text, output by the second feature extraction layer, such as temporal attention features, spatial attention features, spatiotemporal attention features, self-attention features, cross-attention features, etc. This specification does not impose any restrictions on this.
[0138] It should be noted that, considering the consistency of the generated target video and the original reference video in the unmodified part, when calculating the self-attention features of the target video, the question feature information and key value question feature information of the self-attention features can be replaced with the reference question feature information and key value question feature information of the reference video in the previous model.
[0139] By replacing the above-mentioned features that the user will not modify (target video feature set) with the corresponding features obtained through the second feature extraction layer (features of the video generated simply based on the target prompt text), and then generating the target video based on the replaced feature set, compared with only processing the same noise map group and two videos generated by slightly different prompt texts through a model with the same parameters, the generated target video can be made more similar to the unmodified parts of the above-mentioned reference video.
[0140] Furthermore, generating a target attention feature information set based on the intermediate attention feature information set and the target video feature information set includes:
[0141] Determining target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set;
[0142] Determining the attention feature information to be replaced in the intermediate attention feature information set, wherein the attention feature information to be replaced is the attention feature information corresponding to the target attention feature information;
[0143] Replace each piece of attention feature information to be replaced with the corresponding target attention feature information to generate a target attention feature information set.
[0144] In practical applications, the target attention feature information is feature information in the target video feature information set, and the attention feature information to be replaced is feature information corresponding to the target attention feature information.
[0145] Specifically, the attention feature information to be replaced can be understood as some features of the video generated simply based on the target prompt text. These features are usually not modified by the user. For example, the self-attention features that represent the relationship between pixels in the video, the spatiotemporal attention features that represent the movement relationship of entities in the video, etc. This manual does not impose any restrictions on this.
[0146] By finding the feature information to be replaced corresponding to each feature information in the target video feature information set and replacing the corresponding feature information to be replaced, and then generating a target video based on the replaced feature information set, compared with two videos generated by only processing the same noise map group and slightly different prompt texts through a model with the same parameters, the generated target video can be made more similar to the unmodified parts of the above-mentioned reference video.
[0147] Taking into account that, in the process of generating the reference video, the information of the initial prompt text and the reference noise map features is included, so that the background of the target video can be kept consistent with that of the reference video in the subsequent process of generating the target video, the intermediate attention feature information set includes intermediate question feature information, intermediate key value feature information, and intermediate spatiotemporal feature information, and the target video feature information set includes reference question feature information, reference key value feature information, and reference spatiotemporal feature information;
[0148] Furthermore, determining the replacement attention feature information in the intermediate attention feature information set includes:
[0149] In a case where the target attention feature information is the reference question feature information, determining that the intermediate question feature information corresponding to the reference question feature information is the question feature information to be replaced;
[0150] In a case where the target attention feature information is the reference key-value feature information, determining that the intermediate key-value feature information corresponding to the reference key-value feature information is the key-value feature information to be replaced;
[0151] In a case where the target attention feature information is reference spatiotemporal feature information, the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is determined to be the spatiotemporal feature information to be replaced.
[0152] In actual applications, the intermediate question feature information and the intermediate key value feature information are feature information used to calculate the relationship between each pixel in the video generated based on the target prompt text, the question feature information to be replaced and the key value feature information to be replaced are feature information used to replace the reference question feature information and the reference key value feature information, the intermediate spatiotemporal feature information is feature information representing the temporal and spatial fusion features between entities in the video generated based on the target prompt text, and the spatiotemporal feature information to be replaced is feature information used to replace the reference spatiotemporal feature information.
[0153] Specifically, the intermediate question feature information can be understood as taking each pixel of the video generated based on the target prompt text as the feature information of the question, and the intermediate key value feature information can be understood as taking each pixel in the video generated based on the target prompt text as the feature information of the key value. By performing dot multiplication on the feature information as the question and the feature information as the key value, the relationship between each pixel in the video generated based on the target prompt text can be determined. By replacing the above two intermediate question feature information and intermediate key value feature information, the relationship between each pixel in the reference video can be obtained, and then added to the value feature information of each pixel of the video generated based on the target prompt text as the answer. On the basis of ensuring that the relationship between each pixel in the generated target video is similar to the relationship between each pixel in the reference video, the user's modification to the video (that is, the relationship between each pixel in the video generated based on the target prompt text) can be added. This achieves the addition of some user modifications to the video on the basis of ensuring that the generated target video is more similar to the unmodified parts of the above reference video, thereby improving the user experience.
[0154] The intermediate spatiotemporal feature information can be understood as the temporal and spatial relationship between various entities in the video generated based on the target prompt text, or it can be understood as feature information representing the motion features of the video generated based on the target prompt text. Replacing the above-mentioned intermediate spatiotemporal feature information with the reference spatiotemporal feature information representing the motion features of the reference video can effectively retain the motion information of each entity in the reference video in the newly generated target video, thereby effectively making the generated target video more similar to the unmodified parts of the above-mentioned reference video.
[0155] Furthermore, each to-be-replaced attention feature information is replaced with the corresponding target attention feature information set, including:
[0156] Replacing the to-be-replaced question characteristic information with the reference question characteristic information;
[0157] Replacing the key value feature information to be replaced with the reference key value feature information;
[0158] The to-be-replaced spatiotemporal feature information is replaced with the reference spatiotemporal feature information.
[0159] By replacing each feature information in the target video feature information set with its corresponding feature information to be replaced, and then generating a target video based on the replaced feature information set, the generated target video can be made more similar to the unmodified parts of the above-mentioned reference video, compared with two videos generated by only processing the same noise map group and slightly different prompt texts through a model with the same parameters.
[0160] Taking into account that the accuracy of video modification can be further improved without considering resource consumption, the second feature extraction layer includes a second cross feature extraction unit;
[0161] Furthermore, after obtaining the target video generated by the second video generation model, the method further includes:
[0162] Determining prompt text distinguishing word information according to the initial prompt text and the target prompt text;
[0163] Determining at least one associated word information with the prompt text distinguishing word information based on the initial prompt text and the prompt text distinguishing word information;
[0164] Obtaining, by the second cross-feature extraction unit, distinguishing cross-attention feature information corresponding to the distinguishing word information of the prompt word text and associated cross-attention feature information corresponding to each associated word information, generated according to the reference noise pattern group and the target prompt word text;
[0165] A target edited video is generated based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video.
[0166] In actual applications, the second cross-feature extraction unit is an attention unit that obtains the association between the target prompt text and the target video. The prompt text distinguishing word information is the word that distinguishes the target prompt word text and the initial prompt word text. The associated word information is the word related to the prompt text distinguishing word information in the initial prompt text. The distinguishing cross-attention feature information is the feature information that represents the association between the prompt text distinguishing word information and the target video. The associated cross-attention feature information is the feature information that represents the association between each associated word and the target video.
[0167] Specifically, based on the initial prompt text and the target prompt text, the determined prompt text distinguishing word information can be understood as the difference between the initial prompt text and the target prompt text, and can also be further understood as the word that the user wants to adjust. For example, the initial prompt text is "a rabbit photo video" and the target prompt text is "a black rabbit photo video", then the prompt text distinguishing word information corresponding to the two is "black".
[0168] The at least one associated word information determined to correspond to the prompt text distinguishing word information can be understood as the preset number of words that have the greatest association relationship with the acquired prompt text distinguishing word information in the initial prompt text, and can also be further understood as the preset number of words that are most relevant to the prompt text distinguishing word information. For example, the initial prompt text is "a rabbit photo video", and the confirmed prompt text distinguishing word information is "black". Then, if the preset number of associated words is 1, then one associated word associated with the prompt text distinguishing word information is "rabbit". If the preset number of associated words is 2, then the two associated words associated with the prompt text this time are "rabbit" and "photo".
[0169] It should be noted that the method of obtaining the prompt text distinguishing word information and the method of determining the associated words corresponding to the prompt text distinguishing word information can be to calculate the feature vector of each word and calculate the distance, or to organize the relationship between each word through a model, etc. to obtain the relationship between words. This manual does not impose any restrictions on this.
[0170] The distinguishing cross-attention feature information can be understood as the relationship between the obtained prompt text distinguishing word information and the generated target video, and can also be further understood as the relationship between the prompt text distinguishing word information and each entity in the video. Similarly, the associative cross-attention feature information can be understood as the relationship between the obtained associative word information and the generated target video, and can also be further understood as the relationship between the associative word information and each entity in the video.
[0171] refer to Figure 4 , Figure 4 A flowchart for mask-adding video editing provided by an embodiment of the present specification is shown. The flowchart includes receiving an initial text message sent by a user, parsing the initial text message into a standardized initial prompt text 404, and generating a set of Gaussian noise reference noise map groups 402 using preset rules. The initial prompt text 404 and the reference noise map group 402 are then input together into a cross-feature extraction unit in a first feature extraction layer 4062 in a first video generation model 406 to obtain cross-feature information that incorporates noise map group features and initial prompt text features. The cross-feature information is then input into a self-attention extraction unit for self-attention extraction to obtain self-attention features that enhance key positions in the noise map group. The cross-feature information is then input into a spatiotemporal self-attention extraction unit to obtain spatiotemporal feature information containing spatiotemporal features in the reference noise map group. The cross-feature information, self-attention features, and spatiotemporal feature information are then input together into the first video generation layer 4064 to obtain a reference video 408.
[0172] After receiving the reference video 408, the user gives a modification command to the reference video 408 and sends a video adjustment instruction. Then, based on the adjustment text in the video adjustment instruction, the target prompt text 416 with slight changes from the initial prompt text 404 is generated. Then, the same reference noise map group 402 and target prompt text 416 are input into the second video generation model 418. Since the second video generation model 418 has the same structure and the same parameters as the first video generation model 406, the reference noise map group 402 and the target prompt text 416 in the second feature extraction layer 4184 have the same operation as the reference noise map group 402 and the initial prompt text 404 in the first feature extraction layer 4062. Then, the question feature information and key value feature information in the self-attention feature generated according to the target prompt text 416 and the reference noise map group 402 are replaced with the reference question feature information 410 and the reference key value feature information 414 generated according to the initial prompt text 404 and the reference noise map group 402, and then the spatiotemporal feature information in the spatiotemporal feature information generated according to the target prompt text 416 and the reference noise map group 402 is replaced with the reference spatiotemporal feature information 414 generated according to the initial prompt text 404 and the reference noise map group 402, and the replaced feature information is input into the second video generation layer 4184 to obtain the target video 420 obtained by the user by modifying the reference video 408.
[0173] Subsequently, the distinguishing cross-attention feature information and the associated cross-attention feature information output by the cross-feature extraction unit in the second feature extraction layer 4182 are obtained, and then the distinguishing cross-attention feature information and the associated cross-attention feature information are fused to obtain the target editing mask information 422. In addition, after noise processing is performed on the reference video 408 and the target video 420, a reference noise video 424 and a target noise video 426 are obtained, respectively. Subsequently, the target editing mask information 422, the reference noise video 424, and the target noise video 426 are integrated to obtain the target editing video 428.
[0174] By extracting the cross-feature information corresponding to the distinguishing words between the target prompt text and the initial prompt text, as well as the cross-feature information corresponding to the associated words corresponding to the distinguishing words, and masking the noise video after noise processing of the reference video and the target video according to the extracted cross-feature information, and then generating the target edited video according to the result after mask processing, the similarity between the target video and the unmodified parts of the reference video can be further improved.
[0175] Furthermore, generating a target edited video according to the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video, includes:
[0176] Processing the distinguishing cross-attention feature information and each associated cross-attention feature information according to a preset mask generation rule to generate target editing mask information;
[0177] Processing the reference video and the target video according to a preset noise processing rule to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video;
[0178] A target edited video is generated according to the target edit mask information, the reference noise video and the target noise video.
[0179] In practical applications, the mask generation rule is a data processing rule for generating target editing mask information, the reference noise video is a noise video carrying some features of the reference video, the target noise video is a noise video carrying some features of the target video, and the target editing video is the edited video generated according to the video adjustment instructions sent by the user.
[0180] Specifically, the mask generation rule can be understood as a generation rule for setting the area related to the distinguishing words in the prompt text as a mask. For example, the above-mentioned distinguishing cross-attention feature information and each associated cross-attention feature information are added, and then the area where the sum is less than the preset threshold is set to 0. In other words, the smaller the value obtained by adding the above-mentioned distinguishing cross-attention feature information and each associated cross-attention feature information, the greater the correlation. Therefore, setting the area less than the preset threshold to 0 is to set the area related to the distinguishing words in the prompt text as a mask.
[0181] It should be noted that the noise processing rules here are the same as the noise processing rules for converting the initial video into the reference noise image group, and will not be repeated here. In addition, the method of generating the target edited video can be to first combine the obtained mask information with the reference noise video and the target noise video to generate a masked noise video containing a mask, and then restore the mask noise video through a mask restoration model to generate the target edited video. The above-mentioned mask restoration model can be, for example, a VAE (Variational Autoencoder) or the like, which restores the masked noise video to a video model. This specification does not impose any restrictions on this.
[0182] In one embodiment provided in this specification, a specific method of combining the acquired mask information with the reference noise video and the target noise video to generate a masked noise video containing a mask is shown in Formula 1:
[0183] x t =x src *(1-mask)+x dst *mask...Formula 1
[0184] Among them, x t is the mask noise video containing the mask, x stc is the reference noise video generated based on the reference video, x dst is the target noise video generated based on the target video.
[0185] By extracting the cross-feature information corresponding to the distinguishing words between the target prompt text and the initial prompt text, as well as the cross-feature information corresponding to the associated words corresponding to the distinguishing words, and masking the noise video after noise processing of the reference video and the target video according to the extracted cross-feature information, and then generating the target edited video according to the result after mask processing, the similarity between the target video and the unmodified parts of the reference video can be further improved.
[0186] By applying the solution of the embodiments of this specification, a reference video is generated through the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate a target video based on the target prompt text. Without the need to retrain the model, it is ensured that the target video generated by the target prompt text that is slightly different from the initial prompt text maintains a small difference from the reference video, thereby improving the accuracy of video editing.
[0187] On this basis, the basic video modification method only modifies the self-attention features and spatiotemporal self-attention feature information, avoiding the incompatibility problem caused by modifying the cross-self-attention features. To further ensure the consistency of the generated target video and the reference video, by obtaining the difference between the prompt text before and after the modification and the feature mapping of the corresponding distinguishing words at the cross-self-attention layer, the target video and the reference video are further mixed based on this. The target video only shows differences from the reference video in the modified areas, and the background unrelated to the modified areas is the same as the reference video, further improving the accuracy of video editing.
[0188] See also Figure 5 , Figure 5 A flowchart of a video generation method applied to a cloud server provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0189] Step 502: Receive the video generation instruction sent by the terminal-side device, and generate an initial prompt text according to the video generation instruction.
[0190] Step 504: Obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, and obtain a target video feature information set generated in the process of processing the reference noise map group and the initial prompt text.
[0191] Considering that, in the case where a user sends a text, a video generated based on the text sent by the user needs to be fed back to the user so that the user can decide where to edit based on the previous video, before obtaining the target video feature information set, the method further includes:
[0192] Inputting the reference noise image group and the initial prompt text into a first video generation model to obtain a reference video generated by the first video generation model;
[0193] In a case where the video generation instruction carries initial text information, the reference video is returned to the terminal-side device.
[0194] When the video generation instruction sent by the user carries an initial text, it can be understood that the user needs to generate a video through the initial text information sent and edit the video. Therefore, when the instruction sent by the user carries an initial text, the reference video generated according to the initial prompt text needs to be returned to the user's terminal device so that the user can send a video adjustment instruction to the server based on the reference video.
[0195] Step 506: Receive the video adjustment instruction sent by the terminal-side device, and generate a target prompt text according to the video adjustment instruction and the initial prompt text.
[0196] Step 508: Input the reference noise image group, the target prompt text, and the target video feature information set into the second video generation model, obtain the target video generated by the second video generation model, and return the target video to the terminal device.
[0197] Considering that if the user is still not satisfied with the modified video, the user can further modify the target video, after returning the target video to the terminal device, the method further includes:
[0198] The receiving end-side device receives a video adjustment instruction for the target video, and generates a target adjusted video corresponding to the target video according to the video adjustment instruction;
[0199] The target adjusted video is sent to the terminal side device.
[0200] In practical applications, the target adjustment video can be understood as a video generated by the target video according to the video adjustment instruction. By receiving the video adjustment instruction generated by the user according to the target video and further adjusting the target video, the user can further modify the target video, thereby further improving the user experience.
[0201] The above is a schematic diagram of the video generation method for a cloud server according to this embodiment. It should be noted that the technical solution of this video generation method for a cloud server shares the same concept as the technical solution of the aforementioned video generation method. For details not described in detail in the technical solution of the video generation method for a cloud server, please refer to the description of the technical solution of the aforementioned video generation method.
[0202] By applying the solution of the embodiments of this specification, a reference video is generated using initial prompt text, and the target video feature information set generated during the reference video generation process is used to generate a target video based on the target prompt text. This ensures that the target video generated using a target prompt text that is slightly different from the initial prompt text maintains minimal differences from the reference video, thereby improving the accuracy of video editing. Furthermore, the generated reference video is returned based on the textual instructions input by the user, allowing the end-side user to issue the next editing instruction based on the generated reference video. Furthermore, after the user completes a video edit, they can edit the target video again, further improving the user experience.
[0203] See also Figure 6 , Figure 6 1 shows an architecture diagram of a video generation system provided by an embodiment of the present specification. The video generation system may include a client 100 and a server 200.
[0204] The client 100 is used to send video generation instructions and video editing instructions to the server 200;
[0205] The server 200 is configured to receive a video generation instruction sent by the client 100, and generate an initial prompt text according to the video generation instruction; obtain a reference noise map group, input the reference noise map group and the initial prompt text into a first video generation model, obtain a reference video generated by the first video generation model, and obtain a target video feature value generated by the first video generation model in the process of generating the reference video; send the reference video to the client 100 when the video generation instruction is an initial text instruction; then receive a video adjustment instruction sent by the client 100, and generate a target prompt text according to the video adjustment instruction and the initial prompt text; input the reference noise map group, the target prompt text, and the target video feature information set into a second video generation model, obtain a target video generated by the second video generation model; and send the target video to the client 100;
[0206] The client 100 is further configured to receive the target video and the reference video sent by the server 200 .
[0207] By applying the solution of the embodiments of this specification, a reference video is generated through the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate a target video based on the target prompt text, thereby ensuring that the target video generated by the target prompt text that is slightly different from the initial prompt text maintains a small difference from the reference video, thereby improving the accuracy of video editing.
[0208] The video generation system may include multiple clients 100 and a server 200. The clients 100 may be referred to as end-side devices, and the server 200 may be referred to as cloud-side devices. The multiple clients 100 may establish a communication connection through the server 200. In a video generation scenario, the server 200 provides video generation services between the multiple clients 100. The multiple clients 100 may act as either senders or receivers, communicating through the server 200.
[0209] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the video generation scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 can generate a target video and a reference video based on the data stream, and push the target video and reference video to other clients with which communication has been established.
[0210] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
[0211] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0212] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that support background training for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that integrates a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0213] It is worth noting that the video generation method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server to execute the video generation method provided in the embodiments of this specification. In other embodiments, the video generation method provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0214] The following combined Figure 7 , taking the application of the video generation method provided in this specification in editing the user's original video as an example, the video generation method is further explained. Figure 7A flowchart of a processing process for editing a user's original video provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0215] Step 702: Receive the video 1 sent by the user, and generate an initial prompt text "a photo video of a rabbit" based on the video 1.
[0216] Step 704: Perform DDIM inversion on all frames in video 1 to obtain a set of potential noise images.
[0217] Step 706: Input the above-mentioned potential noise image group and the initial prompt text "A photo video of a rabbit" into the first video generation model, and obtain the reference question eigenvalue and reference key value eigenvalue output by the self-attention layer in the first self-attention module in the video generation model, as well as the reference spatiotemporal key value eigenvalue output by the spatiotemporal self-attention layer.
[0218] Step 708: Receive the adjustment instruction sent by the user "Change the rabbit in video 1 to black".
[0219] Step 710: Generate an adjustment prompt text “a photo video of a black rabbit” according to the above adjustment instruction.
[0220] Step 712: Input the above-mentioned potential noise image group and the adjustment prompt text "A photo video of a black rabbit" into the second video generation model, and replace the above-mentioned reference question feature value, reference key value feature value and reference spatiotemporal key value feature value with the intermediate question feature value and intermediate key value feature value output by the self-attention layer in the second self-attention module in the second video generation model, and the intermediate spatiotemporal key value feature value output by the spatiotemporal self-attention layer to obtain the target self-attention feature value after replacement.
[0221] Step 714: Input the target self-attention feature value into the diffusion module of the second video generation model to obtain the target video.
[0222] By applying the solution of the embodiments of this specification, a reference video is generated through the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate a target video based on the target prompt text. Without the need to retrain the model, it is ensured that the target video generated by the target prompt text that is slightly different from the initial prompt text maintains a small difference from the reference video, thereby improving the accuracy of video editing.
[0223] On this basis, the basic video modification method only modifies the self-attention features and spatiotemporal self-attention feature information, avoiding the incompatibility problem caused by modifying the cross-self-attention features. On the premise of further ensuring the consistency of the generated target video and the reference video, by obtaining the difference between the prompt text before and after the modification, and obtaining the feature mapping of the corresponding distinguishing words on the cross-self-attention layer, the target video and the reference video are further mixed on this basis, so that the target video only shows a difference from the reference video in the modified area, and is the same as the reference video in the background unrelated to the modified area, thereby further improving the accuracy of video editing. Corresponding to the above method embodiment, this specification also provides a video generation device embodiment, Figure 8 A schematic structural diagram of a video generating device provided by an embodiment of this specification is shown.
[0224] like Figure 8 As shown, the device includes:
[0225] The generation instruction receiving module 802 is configured to receive a video generation instruction and generate an initial prompt text according to the video generation instruction;
[0226] The video feature value acquisition module 804 is configured to obtain a reference noise image group, input the reference noise image group and the initial prompt text into a first video generation model, and obtain target video feature values generated in the process of processing the reference noise image group and the initial prompt text;
[0227] The adjustment instruction receiving module 806 is configured to receive a video adjustment instruction and generate a target prompt text according to the video adjustment instruction and the initial prompt text;
[0228] The target video generation module 808 is configured to input the reference noise image group, the target prompt text and the target video feature information set into the second video generation model to obtain the target video generated by the second video generation model.
[0229] Optionally, the video generation instruction carries initial text information;
[0230] The generation instruction receiving module 802 is further configured to:
[0231] Parsing the video generation instruction to obtain the initial text information;
[0232] generating the initial prompt text based on the initial text information;
[0233] Accordingly, a reference noise map group is obtained, including:
[0234] Generating the reference noise map group by presetting a noise map group generation rule;
[0235] Accordingly, receiving a video adjustment instruction includes:
[0236] The video adjustment instruction corresponding to the reference video is received.
[0237] Optionally, the video generation instruction carries initial video information;
[0238] The generation instruction receiving module 802 is further configured to:
[0239] Parsing the video generation instruction to obtain the initial video information;
[0240] generating the initial prompt text based on the initial video information, wherein the initial prompt text is used to describe the initial video information;
[0241] Accordingly, a reference noise map group is obtained, including:
[0242] Based on the initial video information, generating the reference noise map group corresponding to the initial video information;
[0243] Accordingly, receiving a video adjustment instruction includes:
[0244] The video adjustment instruction corresponding to the initial video information is received.
[0245] Optionally, the video feature value acquisition module 804 is further configured to:
[0246] Inputting the reference noise image group and the initial prompt text into a first video generation model to generate a reference video;
[0247] Obtain a target video feature information set generated by the first video generation model in the process of generating the reference video.
[0248] Optionally, the first video generation model includes a first feature extraction layer and a first video generation layer;
[0249] The video feature value acquisition module 804 is further configured to:
[0250] Inputting the reference noise image group and the initial prompt text into the first feature extraction layer, obtaining a reference attention feature information set output by the first feature extraction layer;
[0251] The reference attention feature information set is input into the first video generation layer to obtain the reference video output by the first video generation layer.
[0252] Optionally, the video feature value acquisition module 804 is further configured to:
[0253] Obtaining a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature information set includes at least one type of attention feature information;
[0254] A target video feature information set is determined based on the reference attention feature information set, wherein the target video feature information set includes at least one type of attention feature information.
[0255] Optionally, the reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information;
[0256] The video feature value acquisition module 804 is further configured to:
[0257] Determine reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information according to the reference attention feature information set;
[0258] A target video feature information set is determined based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information.
[0259] Optionally, the reference self-attention feature information includes reference question feature information, reference key value feature information and reference value feature information, and the reference spatiotemporal self-attention feature information includes reference spatiotemporal feature information;
[0260] The video feature value acquisition module 804 is further configured to:
[0261] Determining reference question feature information and reference key value feature information based on the reference self-attention feature information;
[0262] Determining reference spatiotemporal feature information based on the reference spatiotemporal self-attention feature information;
[0263] The reference question feature information, the reference key value feature information, and the reference spatiotemporal feature information are determined to be target feature information.
[0264] Optionally, the second video generation model includes a second feature extraction layer and a second video generation layer, and parameters of the second video generation layer are the same as parameters of the first video generation layer of the first video generation model;
[0265] The target video generation module 808 is further configured to:
[0266] Inputting the reference noise image group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set;
[0267] The target attention feature information set is input into the second video generation layer to obtain the target video output by the second video generation layer.
[0268] Optionally, the target video generation module 808 is further configured to:
[0269] Inputting the reference noise image group and the target prompt text into a second feature extraction layer, and obtaining an intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one type of attention feature information;
[0270] A target attention feature information set is generated based on the intermediate attention feature information set and the target video feature information set.
[0271] Optionally, the target video generation module 808 is further configured to:
[0272] Determining target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set;
[0273] Determining the attention feature information to be replaced in the intermediate attention feature information set, wherein the attention feature information to be replaced is the attention feature information corresponding to the target attention feature information;
[0274] Replace each piece of attention feature information to be replaced with the corresponding target attention feature information to generate a target attention feature information set.
[0275] Optionally, the intermediate attention feature information set includes intermediate question feature information, intermediate key value feature information, and intermediate spatiotemporal feature information, and the target video feature information set includes reference question feature information, reference key value feature information, and reference spatiotemporal feature information;
[0276] The target video generation module 808 is further configured to:
[0277] In a case where the target attention feature information is the reference question feature information, determining that the intermediate question feature information corresponding to the reference question feature information is the question feature information to be replaced;
[0278] In a case where the target attention feature information is the reference key-value feature information, determining that the intermediate key-value feature information corresponding to the reference key-value feature information is the key-value feature information to be replaced;
[0279] In a case where the target attention feature information is reference spatiotemporal feature information, the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is determined to be the spatiotemporal feature information to be replaced.
[0280] Optionally, the target video generation module 808 is further configured to:
[0281] Replacing the to-be-replaced question characteristic information with the reference question characteristic information;
[0282] Replacing the key value feature information to be replaced with the reference key value feature information;
[0283] The to-be-replaced spatiotemporal feature information is replaced with the reference spatiotemporal feature information.
[0284] Optionally, the second feature extraction layer includes a second cross feature extraction unit;
[0285] Optionally, the video generating device further includes a mask module configured to:
[0286] Determining prompt text distinguishing word information according to the initial prompt text and the target prompt text;
[0287] Determining at least one associated word information with the prompt text distinguishing word information based on the initial prompt text and the prompt text distinguishing word information;
[0288] Obtaining, by the second cross-feature extraction unit, distinguishing cross-attention feature information corresponding to the distinguishing word information of the prompt word text and associated cross-attention feature information corresponding to each associated word information, generated according to the reference noise pattern group and the target prompt word text;
[0289] A target edited video is generated based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video.
[0290] Optionally, the mask module is further configured to:
[0291] Processing the distinguishing cross-attention feature information and each associated cross-attention feature information according to a preset mask generation rule to generate target editing mask information;
[0292] Processing the reference video and the target video according to a preset noise processing rule to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video;
[0293] A target edited video is generated according to the target edit mask information, the reference noise video and the target noise video.
[0294] The above is a schematic diagram of a video generation device according to this embodiment. It should be noted that the technical solution of the video generation device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the video generation device, please refer to the description of the technical solution of the above-mentioned video generation method.
[0295] By applying the solution of the embodiments of this specification, a reference video is generated through the initial prompt text, and the target video feature information set generated in the process of generating the reference video is used to generate a target video based on the target prompt text. Without the need to retrain the model, it is ensured that the target video generated by the target prompt text that is slightly different from the initial prompt text maintains a small difference from the reference video, thereby improving the accuracy of video editing.
[0296] On this basis, the basic video modification method only modifies the self-attention features and spatiotemporal self-attention feature information, avoiding the incompatibility problem caused by modifying the cross-self-attention features. To further ensure the consistency of the generated target video and the reference video, by obtaining the difference between the prompt text before and after the modification and the feature mapping of the corresponding distinguishing words at the cross-self-attention layer, the target video and the reference video are further mixed based on this. The target video only shows differences from the reference video in the modified areas, and the background unrelated to the modified areas is the same as the reference video, further improving the accuracy of video editing.
[0297] Figure 9 The block diagram of a computing device 900 according to one embodiment of the present disclosure is shown. Components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0298] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0299] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 9 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0300] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.
[0301] The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned video generation method or the video generation method applied to the cloud server.
[0302] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the aforementioned video generation method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned video generation method or the video generation method applied to the cloud server.
[0303] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned video generation method or the video generation method applied to a cloud server.
[0304] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the aforementioned video generation method or the technical solution of the video generation method applied to a cloud server. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned video generation method or the technical solution of the video generation method applied to a cloud server.
[0305] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned video generation method or the video generation method applied to a cloud server.
[0306] The above is an illustrative solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the aforementioned video generation method or the technical solution of the video generation method applied to a cloud server. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the aforementioned video generation method or the technical solution of the video generation method applied to a cloud server.
[0307] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0308] The computer instructions include computer program codes, which may be in source code form, object code form, executable files, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0309] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0310] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0311] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A video generation method, comprising: Receive a video generation instruction, and generate an initial prompt text according to the video generation instruction; Obtaining a reference noise image group, inputting the reference noise image group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated by the first video generation model in the process of processing the reference noise image group and the initial prompt text; receiving a video adjustment instruction, and generating a target prompt text according to the video adjustment instruction and the initial prompt text; The reference noise image group, the target prompt text, and the target video feature information set are input into a second video generation model to obtain a target video generated by the second video generation model.
2. The method according to claim 1, wherein the step of inputting the reference noise pattern group and the initial prompt text into a first video generation model and obtaining a target video feature information set generated by the first video generation model during processing of the reference noise pattern group and the initial prompt text comprises: Inputting the reference noise image group and the initial prompt text into a first video generation model to generate a reference video; Obtain a target video feature information set generated by the first video generation model in the process of generating the reference video.
3. The method of claim 2, wherein the first video generation model comprises a first feature extraction layer and a first video generation layer; Inputting the reference noise image group and the initial prompt text into a first video generation model to generate a reference video, comprising: Inputting the reference noise image group and the initial prompt text into the first feature extraction layer, obtaining a reference attention feature information set output by the first feature extraction layer; The reference attention feature information set is input into the first video generation layer to obtain the reference video output by the first video generation layer.
4. The method according to claim 3, wherein obtaining the target video feature information set generated by the first video generation model in the process of generating the reference video comprises: Obtaining a reference attention feature information set output by the first feature extraction layer, wherein the reference attention feature information set includes at least one type of attention feature information; A target video feature information set is determined based on the reference attention feature information set, wherein the target video feature information set includes at least one type of attention feature information.
5. The method of claim 4, wherein the reference attention feature information set includes reference self-attention feature information and reference spatiotemporal self-attention feature information; Determining a target video feature information set according to the reference attention feature information set includes: Determine reference self-attention layer feature information and reference spatiotemporal self-attention layer feature information according to the reference attention feature information set; A target video feature information set is determined based on the reference self-attention feature information and the reference spatiotemporal self-attention feature information.
6. The method of claim 5, wherein the reference self-attention feature information comprises reference question feature information, reference key value feature information, and reference value feature information, and the reference spatiotemporal self-attention feature information comprises reference spatiotemporal feature information; Determining a target video feature information set according to the reference self-attention feature information and the reference spatiotemporal self-attention feature information, including: Determining reference question feature information and reference key value feature information based on the reference self-attention feature information; Determining reference spatiotemporal feature information based on the reference spatiotemporal self-attention feature information; The reference question feature information, the reference key value feature information, and the reference spatiotemporal feature information are determined to be target feature information.
7. The method of claim 1 , wherein the second video generation model comprises a second feature extraction layer and a second video generation layer, and parameters of the second video generation layer are the same as parameters of the first video generation layer of the first video generation model; Inputting the reference noise image group, the target prompt text, and the target video feature information set into a second video generation model to obtain a target video generated by the second video generation model includes: Inputting the reference noise image group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain a target attention feature information set; The target attention feature information set is input into the second video generation layer to obtain the target video output by the second video generation layer.
8. The method of claim 7, wherein the step of inputting the reference noise image group, the target prompt text, and the target video feature information set into the second feature extraction layer to obtain the target attention feature information set comprises: Inputting the reference noise image group and the target prompt text into the second feature extraction layer, and obtaining an intermediate attention feature information set output by the second feature extraction layer, wherein the intermediate attention feature information set includes at least one type of attention feature information; A target attention feature information set is generated based on the intermediate attention feature information set and the target video feature information set.
9. The method of claim 8, wherein generating a target attention feature information set based on the intermediate attention feature information set and the target video feature information set comprises: Determining target attention feature information, wherein the target attention feature information is any one of the attention feature information in the target video feature information set; Determining the attention feature information to be replaced in the intermediate attention feature information set, wherein the attention feature information to be replaced is the attention feature information corresponding to the target attention feature information; Replace each piece of attention feature information to be replaced with the corresponding target attention feature information to generate a target attention feature information set.
10. The method of claim 9, wherein the intermediate attention feature information set includes intermediate question feature information, intermediate key value feature information, and intermediate spatiotemporal feature information, and the target video feature information set includes reference question feature information, reference key value feature information, and reference spatiotemporal feature information; Determining the attention feature information to be replaced in the intermediate attention feature information set includes: In a case where the target attention feature information is the reference question feature information, determining that the intermediate question feature information corresponding to the reference question feature information is the question feature information to be replaced; In a case where the target attention feature information is the reference key-value feature information, determining that the intermediate key-value feature information corresponding to the reference key-value feature information is the key-value feature information to be replaced; In a case where the target attention feature information is reference spatiotemporal feature information, the intermediate spatiotemporal feature information corresponding to the reference spatiotemporal feature information is determined to be the spatiotemporal feature information to be replaced.
11. The method according to claim 10, wherein each set of to-be-replaced attention feature information is replaced with a corresponding set of target attention feature information, comprising: Replacing the to-be-replaced question characteristic information with the reference question characteristic information; Replacing the key value feature information to be replaced with the reference key value feature information; The to-be-replaced spatiotemporal feature information is replaced with the reference spatiotemporal feature information.
12. The method according to claim 7, wherein the second feature extraction layer comprises a second cross feature extraction unit; After obtaining the target video generated by the second video generation model, the method further includes: Determining prompt text distinguishing word information according to the initial prompt text and the target prompt text; Determining at least one associated word information with the prompt text distinguishing word information based on the initial prompt text and the prompt text distinguishing word information; Obtaining, by the second cross-feature extraction unit, distinguishing cross-attention feature information corresponding to the distinguishing word information of the prompt word text and associated cross-attention feature information corresponding to each associated word information, generated according to the reference noise pattern group and the target prompt word text; A target edited video is generated based on the distinguishing cross-attention feature information and each associated cross-attention feature information, as well as the reference video and the target video.
13. The method of claim 12, wherein generating a target edited video based on the distinguishing cross-attention feature information and each associated cross-attention feature information, the reference video, and the target video comprises: Processing the distinguishing cross-attention feature information and each associated cross-attention feature information according to a preset mask generation rule to generate target editing mask information; Processing the reference video and the target video according to a preset noise processing rule to generate a reference noise video corresponding to the reference video and a target noise video corresponding to the target video; A target edited video is generated according to the target edit mask information, the reference noise video and the target noise video.
14. The method according to claim 1, wherein the video generation instruction carries initial text information; Receiving a video generation instruction and generating an initial prompt text according to the video generation instruction, including: Parsing the video generation instruction to obtain the initial text information; generating the initial prompt text based on the initial text information; Accordingly, a reference noise map group is obtained, including: Generating the reference noise map group by presetting a noise map group generation rule; Accordingly, receiving a video adjustment instruction includes: The video adjustment instruction corresponding to the reference video is received.
15. The method according to claim 1, wherein the video generation instruction carries initial video information; Receiving a video generation instruction and generating an initial prompt text according to the video generation instruction, including: Parsing the video generation instruction to obtain the initial video information; generating the initial prompt text based on the initial video information, wherein the initial prompt text is used to describe the initial video information; Accordingly, a reference noise map group is obtained, including: Based on the initial video information, generating the reference noise map group corresponding to the initial video information; Accordingly, receiving a video adjustment instruction includes: The video adjustment instruction corresponding to the initial video information is received.
16. A video generation method, applied to a cloud server, comprising: The receiving end-side device generates a video generation instruction, and generates an initial prompt text according to the video generation instruction; Obtaining a reference noise image group, inputting the reference noise image group and the initial prompt text into a first video generation model, and obtaining a target video feature information set generated in the process of processing the reference noise image group and the initial prompt text; receiving a video adjustment instruction sent by a terminal-side device, and generating a target prompt text according to the video adjustment instruction and the initial prompt text; The reference noise image group, the target prompt text and the target video feature information set are input into the second video generation model, the target video generated by the second video generation model is obtained, and the target video is returned to the terminal side device.
17. The method according to claim 16, before obtaining the target video feature information set, the method further comprises: Inputting the reference noise image group and the initial prompt text into a first video generation model to obtain a reference video generated by the first video generation model; In a case where the video generation instruction carries initial text information, the reference video is returned to the terminal-side device.
18. The method according to claim 16, after returning the target video to the end-side device, further comprising: The receiving end-side device receives a video adjustment instruction for the target video, and generates a target adjusted video corresponding to the target video according to the video adjustment instruction; The target adjusted video is sent to the terminal side device.
19. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1-15 or 16-18 are implemented.
20. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15 or 16 to 18.
21. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15 or 16 to 18.
Citation Information
Patent Citations
Training and sorting method and device of sorting model, electronic equipment and storage medium
CN113392266A
Image processing method, device and equipment and computer readable storage medium
CN116704221A
Video editing method and device, electronic equipment and storage medium
CN116980541A
Image generation method and device, equipment and storage medium
CN117058276A