Video generation model training method and device, electronic equipment and storage medium

By using the video generation model's first frame image and special effect prompt word processing, the video generation model is trained, and the problem of low generation efficiency caused by user description is solved, and efficient and accurate generation of special effect videos is achieved.

CN120358394AActive Publication Date: 2025-07-22BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510853773.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the prior art, the training method of the video generation model relies on the user to describe the special effects by himself, resulting in low efficiency in the generation of special effects videos, making it difficult for users to quickly and accurately input complex description information.

Method used

By obtaining the first frame image and special effect prompt words of the sample special effect video, using the video generation model for processing and training, generating special effect videos, and calculating the loss through the differences between the special effect features and the sample special effect video, adjusting the parameters of the text encoder or special effect encoder to achieve automatic generation of special effect videos.

Benefits of technology

It improves the generation efficiency of special effects videos, ensures that the generated special effects videos are accurately in line with the sample special effects videos, reduces semantic ambiguity, and reduces the needs of user descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358394A_ABST
    Figure CN120358394A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device of a video generation model, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining respective sample special effect videos and special effect cues of a plurality of special effects, wherein the special effect cues are used for indicating special effects in a video to be generated by a video generation model; for any special effect in the plurality of special effects, through the video generation model, processing a first frame image in a sample special effect video of the special effect and a special effect prompt word of the special effect to obtain a first special effect video, the first frame image being used for indicating a subject in the video to be generated by the video generation model; and training the video generation model based on the first special effect video and the sample special effect video corresponding to the plurality of special effects. According to the technical scheme, the generation efficiency of the special effect video is improved on the basis of guaranteeing the generation effect of the special effect video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly relates to a training method, apparatus, electronic device, and storage medium for a video generation model. Background Art

[0002] With the rapid development of artificial intelligence technology, the method of creating special effects videos based on artificial intelligence (AI) technology has become increasingly popular. Among them, a trained video generation model is usually used to produce special effects videos. The performance of the video generation model often depends on the training method of video generation.

[0003] In related technologies, the commonly used training method is: input an original video (base video), a description information about special effects, and region prompts into the video generation model, and the video generation model generates special effects in the specified region prompted in the original video according to the description information, so as to obtain a special effects video; the model training is carried out with the goal that the special effects video is more natural and conforms to the visual laws of the real world.

[0004] However, for the video generation model trained by the above technical solution, during the model usage, it requires the user to describe the special effects by themselves to generate the corresponding special effects video. However, the description information about special effects is usually relatively complex, and it is difficult for ordinary users to quickly input reasonable and accurate description information, resulting in low efficiency of generating special effects videos. Summary of the Invention

[0005] The present disclosure provides a training method, apparatus, electronic device, and storage medium for a video generation model, which improves the generation efficiency of special effects videos on the basis of ensuring the generation effect of special effects videos. The technical solution of the present disclosure is as follows.

[0006] According to one aspect of the embodiments of the present disclosure, a training method for a video generation model is provided, including: Obtaining sample special effects videos and special effects prompt words of multiple special effects respectively, where the special effects prompt words are used to indicate the special effects in the videos to be generated by the video generation model; For any one of the multiple special effects, through the video generation model, processing the first frame image in the sample special effects video of the special effect and the special effects prompt word of the special effect to obtain a first special effects video, where the first frame image is used to indicate the main body in the video to be generated by the video generation model; Training the video generation model based on the first special effects videos and sample special effects videos corresponding to the multiple special effects.

[0007] According to another aspect of the embodiments of the present disclosure, a training apparatus for a video generation model is provided, including: An acquisition module, configured to acquire sample special effect videos and special effect prompt words for each of multiple special effects, where the special effect prompt words are used to indicate special effects in a video to be generated by a video generation model; A first processing module, configured to, for any one of the multiple special effects, process a first frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect through the video generation model to obtain a first special effect video, where the first frame image is used to indicate a main body in the video to be generated by the video generation model; A training module, configured to train the video generation model based on the first special effect videos and the sample special effect videos corresponding to the multiple special effects.

[0008] In some embodiments, the first processing module includes: A first encoding unit, configured to, for any one of the multiple special effects, encode a first frame image in the sample special effect video of the special effect through an image encoder in the video generation model to obtain a first image encoding feature; A second encoding unit, configured to encode the special effect prompt word of the special effect through a text encoder in the video generation model to obtain a text encoding feature; A processing unit, configured to process the first image encoding feature and the text encoding feature through a video generation network in the video generation model to obtain the first special effect video.

[0009] In some embodiments, the processing unit is configured to process the first image encoding feature and the text encoding feature through a video generation network in the video generation model to obtain a first video feature, where the first video feature is used to indicate the first special effect video; The training module is configured to, for any one of the multiple special effects, determine a first loss based on the first video feature and a sample video feature of the sample special effect video of the special effect, where the first loss is used to represent the difference between the first video feature and the sample video feature; and train the video generation model with the goal of minimizing the first loss.

[0010] In some embodiments, the training module is configured to freeze parameters in other structures of the video generation model except for the text encoder and adjust parameters of the text encoder in the video generation model.

[0011] In some embodiments, the apparatus further includes: The special effect encoding module is configured to perform, for any one of the multiple special effects, encoding the sample video features of the sample special effect video of the special effect through the special effect encoder in the video generation model to obtain special effect features, where the special effect features are used to indicate the attributes of the special effect in the sample special effect video; The first processing module is further configured to perform processing the special effect features, the first frame image in the sample special effect video of the special effect, and the special effect prompt words of the special effect through the video generation model to obtain a second special effect video; The training module is further configured to perform re-training the video generation model based on the second special effect videos and the sample special effect videos corresponding to the multiple special effects.

[0012] In some embodiments, the first processing module is configured to perform, for any one of the multiple special effects, processing the sample video features through the attention layer in the special effect encoder to obtain the query feature, key feature, and value feature of the attention layer, where the query feature and the key feature are used to indicate the position of the special effect in the sample special effect video, and the value feature is used to indicate the attributes of the special effect; fusing the query feature, the key feature, and the value feature to obtain the special effect features.

[0013] In some embodiments, the first processing module is configured to perform processing the special effect features, the first image encoding feature corresponding to the first frame image, and the text encoding feature corresponding to the special effect prompt words through the video generation network in the video generation model to obtain a second video feature, where the second video feature is used to indicate the second special effect video; The training module is configured to perform, for any one of the multiple special effects, determining a second loss based on the second video feature and the sample video features of the sample special effect video of the special effect, where the second loss is used to represent the difference between the second video feature and the sample video features; training the video generation model with the goal of minimizing the second loss.

[0014] In some embodiments, it is characterized in that the training module is configured to perform freezing the parameters in other structures in the video generation model except the special effect encoder and adjusting the parameters in the special effect encoder in the video generation model.

[0015] In some embodiments, the video generation model further includes a plurality of special effect adapters, and the plurality of special effect adapters correspond one-to-one to each network layer in the video generation network in the video generation model, and the video generation network is used to generate special effect videos; The device further includes: A second processing module, configured to perform, for any one of the multiple special effects, encoding any frame image in the sample special effect video of the special effect through an image encoder in the video generation model to obtain a second image encoding feature; for a special effect adapter corresponding to the i-th network layer in the video generation network, processing a third video feature and the second image encoding feature through the special effect adapter to obtain a fourth video feature, where the third video feature is the input of the i-th network layer, and the third video feature is used to indicate a special effect video including the special effect generated by the (i - 1)-th network layer; fusing the fourth video feature with the output of the i-th network layer to obtain the input of the (i + 1)-th network layer; and training multiple special effect adapters in the video generation model based on the target video feature output by the last network layer in the video generation network and the sample special effect video.

[0016] In some embodiments, the first processing module is configured to perform, for a special effect adapter corresponding to the i-th network layer in the video generation network, broadcasting the second image encoding feature to image features corresponding to each image frame in the third video feature through the special effect adapter to obtain the third video feature.

[0017] In some embodiments, the multiple special effects include interactive special effects, and the interactive special effects are used to indicate special effects in which both the main body and the environment in the video change.

[0018] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the above-mentioned training method of the video generation model.

[0019] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enabling the electronic device to execute the above-mentioned training method of the video generation model.

[0020] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, implementing the above-mentioned training method of the video generation model.

[0021] An embodiment of the present disclosure provides a method for training a video generation model. By using the video generation model, for the first-frame image and the special effect prompt word of the sample special effect video of each special effect, the first special effect video of each special effect is generated. Through the first special effect video of each special effect and the corresponding sample special effect video, the video generation model is trained. This not only makes the special effect videos generated by the video generation model more and more in line with the sample special effect videos, that is, the generated special effect videos are more and more accurate, but also realizes the association between the special effect prompt words and the special effect videos, and the video generation model can learn this association relationship. As a result, the video generation model can accurately determine the style of the special effect through the special effect prompt word. Compared with the user's custom description related to the special effect, in this solution, the user does not need to consider the description of the special effect, and the semantic ambiguity related to the special effect can be reduced through the special effect prompt word. Therefore, a special effect video that meets the requirements can be automatically generated based on the special effect prompt word, improving the generation efficiency of the special effect video while ensuring the generation effect of the special effect video.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0024] Figure 1 is a schematic diagram of an implementation environment of a method for training a video generation model shown according to an exemplary embodiment.

[0025] Figure 2 is a flowchart of a method for training a video generation model shown according to an exemplary embodiment.

[0026] Figure 3 is a flowchart of another method for training a video generation model shown according to an exemplary embodiment.

[0027] Figure 4 is a schematic diagram of a special effect classification shown according to an exemplary embodiment.

[0028] Figure 5 is a schematic diagram of training a text encoder shown according to an exemplary embodiment.

[0029] Figure 6 is a schematic diagram of training a special effect encoder shown according to an exemplary embodiment.

[0030] Figure 7 is a schematic diagram of training a special effect adapter shown according to an exemplary embodiment.

[0031] Figure 8 It is a schematic diagram of a special effect adapter shown according to an exemplary embodiment.

[0032] Figure 9 It is a schematic diagram of a model inference stage shown according to an exemplary embodiment.

[0033] Figure 10 It is a block diagram of a training device for a video generation model shown according to an exemplary embodiment.

[0034] Figure 11 It is a block diagram of a terminal shown according to an exemplary embodiment.

[0035] Figure 12 It is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners

[0036] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0038] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the respective sample special effect videos and special effect prompt words of multiple special effects involved in the present disclosure are obtained under full authorization.

[0039] Figure 1 It is a schematic diagram of an implementation environment of a training method for a video generation model shown according to an exemplary embodiment. Taking the electronic device being provided as a server as an example, see Figure 1 , this implementation environment specifically includes: a terminal 101 and a server 102.

[0040] The terminal 101 is at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop portable computer. An application program is installed and running on the terminal 101. The application program can be a multimedia application program, a clip application program, a game application program, a social application program, etc. The embodiments of the present disclosure do not limit this. The user can log in to the application program through the terminal 101 to obtain the services provided by the application program. The terminal 101 can be connected to the server 102 through a wireless network or a wired network, and thus can send a sample special effect video and a special effect prompt word to the server 102, and the server 102 trains a video generation model through the sample special effect video and the special effect prompt word.

[0041] The terminal 101 generally refers to one of multiple terminals. In this embodiment, the terminal 101 is used as an example for illustration. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals can be several, or the above terminals are dozens or hundreds, or a larger number. The embodiments of the present disclosure do not limit the number and device type of the terminals.

[0042] The server 102 is at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 can be connected to the terminal 101 and other terminals through a wireless network or a wired network. The server 102 can receive the sample special effect video and the special effect prompt word sent by the terminal 101, input the sample special effect video and the special effect prompt word into the video generation model, and train the video generation model. In some embodiments, the number of the above servers can be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 also includes other functional servers to provide more comprehensive and diverse services.

[0043] Figure 2 is a flowchart of a method for training a video generation model shown according to an exemplary embodiment. See Figure 2 , and the method for training the video generation model is applied to a server and includes the following steps.

[0044] In step 201, the server obtains the sample special effect video and the special effect prompt word of each of multiple special effects. The special effect prompt word is used to indicate the special effect in the video to be generated by the video generation model.

[0045] In the embodiments of the present disclosure, the multiple special effects may include a petrification special effect, a lightning special effect, an explosion special effect, etc. The embodiments of the present disclosure do not limit the styles of the multiple special effects. For any special effect, the server obtains the sample special effect video and the special effect prompt word of the special effect. The sample special effect video of each special effect is used to present the style of the special effect. The special effect prompt word of each special effect is used to indicate the special effect, which is equivalent to a special effect identifier.

[0046] In step 202, for any one of the multiple special effects, the server processes the first frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect through a video generation model to obtain a first special effect video, where the first frame image is used to indicate the main body in the video to be generated by the video generation model.

[0047] In the embodiments of the present disclosure, for any one of the multiple special effects, the server inputs the first frame image in the sample special effect video of the special effect and the special effect prompt word into the video generation model, and processes the first frame image and the special effect prompt word through the video generation model to generate a first special effect video. The first special effect video is a special effect video automatically generated by the video generation model based on the first frame image and the special effect prompt word. During the process of generating the first special effect video, the video generation model determines the main body in the special effect video to be generated from the first frame image, and determines the special effect in the special effect video to be generated from the special effect prompt word.

[0048] Among them, the main body refers to the main carrier for transmitting information in the video. The main body can be a person or an object in the video, etc., and the embodiments of the present disclosure do not limit this.

[0049] In step 203, the server trains the video generation model based on the first special effect videos and the sample special effect videos corresponding to the multiple special effects.

[0050] In the embodiments of the present disclosure, for any one of the multiple special effects, the server can determine the model loss of the video generation model based on the first special effect video corresponding to the special effect and the sample special effect video of the special effect. The model loss is used to represent the difference between the first special effect video generated by the model and the sample special effect video. Then, the server trains the video generation model with the goal of minimizing the model loss, so that the special effect videos generated by the video generation model are getting closer and closer to the sample special effect videos.

[0051] An embodiment of the present disclosure provides a method for training a video generation model. By using the video generation model, for the first-frame image and the special effect prompt word of the sample special effect video of each special effect, a first special effect video of each special effect is generated. By using the first special effect video of each special effect and the corresponding sample special effect video, the video generation model is trained. This not only makes the special effect video generated by the video generation model more and more conform to the sample special effect video, that is, the generated special effect video is more and more accurate, but also realizes the association between the special effect prompt word and the special effect video, and the video generation model can learn this association relationship. Thus, the video generation model can accurately determine the style of the special effect through the special effect prompt word. Compared with the user's custom description related to the special effect, in this solution, the user does not need to consider the description of the special effect, and the semantic ambiguity related to the special effect can be reduced through the special effect prompt word. Therefore, a special effect video that meets the requirements can be automatically generated based on the special effect prompt word, improving the generation efficiency of the special effect video while ensuring the generation effect of the special effect video.

[0052] In some embodiments, for any one of the multiple special effects, processing the first-frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect through the video generation model to obtain a first special effect video includes: For any one of the multiple special effects, encoding the first-frame image in the sample special effect video of the special effect through the image encoder in the video generation model to obtain a first image encoding feature; Encoding the special effect prompt word of the special effect through the text encoder in the video generation model to obtain a text encoding feature; Processing the first image encoding feature and the text encoding feature through the video generation network in the video generation model to obtain the first special effect video.

[0053] In the solution provided by the embodiment of the present disclosure, the video generation model includes an image encoder, a text encoder, and a video generation network. The first-frame image in the sample special effect video and the special effect prompt word are respectively encoded through the image encoder and the text encoder to obtain a first image encoding feature and a text encoding feature. Then, the video generation network generates a first special effect video based on the first image encoding feature and the text encoding feature, making the picture in the first special effect video as consistent as possible with the sample special effect video, which is conducive to achieving the desired training effect with fewer training times.

[0054] In some embodiments, processing the first image encoding feature and the text encoding feature through the video generation network in the video generation model to obtain the first special effect video includes: Process the first image encoding feature and the text encoding feature through a video generation network in the video generation model to obtain a first video feature, where the first video feature is used to indicate the first special effect video; Training the video generation model based on the first special effect videos corresponding to the multiple special effects and the sample special effect videos includes: For any one of the multiple special effects, determine a first loss based on the first video feature and the sample video feature of the sample special effect video of the special effect, where the first loss is used to represent the difference between the first video feature and the sample video feature; Train the video generation model with the goal of minimizing the first loss.

[0055] The solution provided by the embodiments of the present disclosure calculates the difference between the sample special effect video and the generated first special effect video at the level of video features and trains the video model. Without converting the video features into the video itself, it can improve the calculation efficiency of the first loss, thereby facilitating the improvement of the training efficiency of the video generation model.

[0056] In some embodiments, training the video generation model includes: Freeze the parameters in other structures of the video generation model except for the text encoder, and adjust the parameters of the text encoder in the video generation model.

[0057] The solution provided by the embodiments of the present disclosure only trains the text encoder in the video generation model. Not only can the text encoder learn the relationship between the special effect prompt words and the special effect itself, enabling the text encoder to accurately understand the special effect to be generated, so as to extract more accurate features of the special effect prompt words and provide guarantee for accurately generating the special effect video subsequently, but also only adjust the parameters of the text encoder, and the number of adjusted parameters is small, which is conducive to improving the training efficiency.

[0058] In some embodiments, the method further includes: For any one of the multiple special effects, encode the sample video feature of the sample special effect video of the special effect through the special effect encoder in the video generation model to obtain a special effect feature, where the special effect feature is used to indicate the attribute of the special effect in the sample special effect video; Retrain the video generation model through the special effect feature, the first frame image in the sample special effect video of the special effect, and the special effect prompt word of the special effect.

[0059] The solution provided by the embodiments of the present disclosure encodes the sample video features of the sample special effect video through a special effect encoder, so as to obtain the features of the special effect in the sample special effect video. Then, based on the first frame image in the sample special effect video and the special effect prompt words of the special effect, together with the special effect features, a second special effect video is generated. By training the video generation model again with the second special effect video and the sample special effect video, since the special effect features are the special effect information obtained from the visual angle and the special effect prompt words are the special effect information obtained from the language angle, and the second special effect video is generated based on the special effect features and the special effect prompt words, the special effect in the second special effect video is more in line with the special effect in the sample special effect video, that is, the special effect is more accurate. On this basis, the model is trained, so that the video generation model can generate special effects from the visual angle and the language angle, which is beneficial to improving the generation accuracy of the video generation model.

[0060] In some embodiments, for any one of the multiple special effects, encoding the sample video features of the sample special effect video of the special effect through the special effect encoder in the video generation model to obtain special effect features includes: For any one of the multiple special effects, processing the sample video features through the attention layer in the special effect encoder to obtain the query feature, key feature, and value feature of the attention layer. The query feature and the key feature are used to indicate the position of the special effect in the sample special effect video, and the value feature is used to indicate the attribute of the special effect; Fusing the query feature, the key feature, and the value feature to obtain the special effect features.

[0061] The solution provided by the embodiments of the present disclosure processes the sample video features through the attention layer to obtain the query feature, key feature, and value feature of the attention layer, and fuses these three features to obtain the special effect features. Since the query feature and the key feature are used to indicate the position of the special effect in the sample special effect video and the value feature is used to indicate the attribute of the special effect, the obtained special effect features can accurately reflect the position and style of the special effect in the sample special effect video, that is, the accuracy of the special effect features can be improved.

[0062] In some embodiments, processing the special effect features, the first frame image in the sample special effect video of the special effect, and the special effect prompt words of the special effect through the video generation model to obtain a second special effect video includes: Processing the special effect features, the first image encoding feature corresponding to the first frame image, and the text encoding feature corresponding to the special effect prompt words through the video generation network in the video generation model to obtain a second video feature, where the second video feature is used to indicate the second special effect video; Retraining the video generation model based on the second special effect video and the sample special effect video corresponding to the multiple special effects includes: For any one of the multiple special effects, based on the second video feature and the sample video feature of the sample special effect video of the special effect, determine a second loss, where the second loss is used to represent the difference between the second video feature and the sample video feature; Taking minimizing the second loss as the goal, train the video generation model.

[0063] The solution provided by the embodiments of the present disclosure calculates the difference between the sample special effect video and the generated second special effect video at the level of video features, and trains the video model. Without converting the video features into the video itself, it can improve the calculation efficiency of the second loss, thereby facilitating the improvement of the training efficiency of the video generation model.

[0064] In some embodiments, the retraining of the video generation model includes: Freeze the parameters in other structures of the video generation model except for the special effect encoder, and adjust the parameters in the special effect encoder of the video generation model.

[0065] The solution provided by the embodiments of the present disclosure only trains the special effect encoder in the video generation model. Not only can the special effect encoder learn the features of the special effect, enabling the special effect encoder to accurately understand the special effect to be generated, so as to extract more accurate special effect features, providing a guarantee for accurately generating the special effect video later, but also only adjust the parameters of the special effect encoder, and the number of adjusted parameters is small, which is conducive to improving the training efficiency.

[0066] In some embodiments, the video generation model further includes a plurality of special effect adapters, and the plurality of special effect adapters correspond one-to-one to each network layer in the video generation network in the video generation model, and the video generation network is used to generate special effect videos; The method further includes: For any one of the multiple special effects, encode any frame image in the sample special effect video of the special effect through the image encoder in the video generation model to obtain a second image encoding feature; For the special effect adapter corresponding to the i-th network layer in the video generation network, process the third video feature and the second image encoding feature through the special effect adapter to obtain a fourth video feature, where the third video feature is the input of the i-th network layer, and the third video feature is used to indicate the special effect video including the special effect generated by the (i - 1)-th network layer; Fuse the fourth video feature with the output of the i-th network layer to obtain the input of the (i + 1)-th network layer; Train multiple special effect adapters in the video generation model based on the target video features output by the last network layer in the video generation network and the sample special effect video.

[0067] The solution provided by the embodiments of the present disclosure can decouple the video subject and the special effects in the sample special effect video by configuring adapters on each network layer in the video generation network, learn the physical motion laws of the special effects and migrate them to any input picture, accurately capture and reproduce the physical motion laws of the special effects, ensure the visual consistency of the generated video, and accurately reflect the characteristics of the special effects.

[0068] In some embodiments, for the special effect adapter corresponding to the i-th network layer in the video generation network, processing the third video feature and the second image encoding feature through the special effect adapter to obtain a fourth video feature includes: For the special effect adapter corresponding to the i-th network layer in the video generation network, through the special effect adapter, broadcast the second image encoding feature to the image features corresponding to each image frame in the third video feature to obtain the third video feature.

[0069] The solution provided by the embodiments of the present disclosure can reflect the video subject with the second image encoding feature. By broadcasting the second image encoding feature to all images, it can be used as an appearance guide during training, enabling the model to move the learned subject according to the expected motion pattern in the existing video, thus decoupling the complete special effect video from the random video frame spatial information, capturing more purely the physical visual effects that do not exist in real life, and further enhancing the performance of the special effect generation result after being injected into the generation model.

[0070] In some embodiments, the multiple special effects include interactive special effects, and the interactive special effects are used to indicate special effects in which both the subject and the environment in the video change.

[0071] The solution provided by the embodiments of the present disclosure includes interactive special effects in which both the subject and the environment change in the sample special effect video, facilitating the video generation model to learn the interactive special effects, so that it can generate special effects with interactions between the special effects and the subject, enriching the ability of the video generation model to generate special effects.

[0072] The above Figure 2 shows only the basic process of the present disclosure. Based on a specific implementation manner below, the solution provided by the present disclosure will be further elaborated. Figure 3 is a flowchart of another method for training a video generation model shown according to an exemplary embodiment. Taking the electronic device being provided as a server as an example, see Figure 3 , the method includes the following steps.

[0073] In step 301, the server obtains the sample special effect videos and special effect prompt words of each of multiple special effects, and the special effect prompt words are used to indicate the special effects in the video to be generated by the video generation model.

[0074] In the embodiments of the present disclosure, the multiple special effects may include petrification special effects, lightning special effects, explosion special effects, etc., and the embodiments of the present disclosure do not limit the styles of the multiple special effects. For any special effect, the server obtains the sample special effect video and special effect prompt word of this special effect. The sample special effect video of each special effect is used to present the style of this special effect. The special effect prompt word of each special effect is used to indicate this special effect, which is equivalent to a special effect identifier.

[0075] In some embodiments, the above-mentioned multiple special effects may include at least one of main body special effects, background special effects, and interaction special effects. The main body special effect refers to a special effect used to indicate that the main body in the video changes. The background special effect is used to indicate the change of the background (environment) in the video. The interaction special effect is used to indicate that both the main body and the environment (background) in the video change. In the solution provided by the embodiments of the present disclosure, the sample special effect video includes the interaction special effect in which both the main body and the environment change, which is convenient for the video generation model to learn the interaction special effect, so that it can generate a special effect in which there is an interaction between the special effect and the main body, enriching the ability of the video generation model to generate special effects.

[0076] Among them, each type of special effect may also include multiple sub-types. For example, the main body special effect may include special effects of main body physical changes, special effects of main body material changes, special effects of main body transformation, etc. Among them, the special effect of main body physical change refers to a special effect in which the form of the main body changes. The special effect of main body transformation refers to a special effect in which the main body transforms into a preset character. The background special effect may include real special effects and virtual special effects, etc. Among them, the background in the real special effect actually exists in real life. The background in the virtual special effect does not exist in real life but is fictional. The interaction special effect may include action following special effects, main body static special effects, and main body dynamic special effects. Among them, the action following special effect refers to a special effect in which the display position of the special effect changes following the action of the main body. In the main body static special effect, the main body is in a static state. For example, in the special effect of "hair growing longer", the display position of the head as the main body does not change. In the main body dynamic special effect, the main body is in a dynamic state. For example, in the special effect of "pinching the shoulder", the shoulder as the main body deforms in response to the action of pinching the shoulder.

[0077] For example, according to the types of special effects, this solution evenly constructs a special effect data set of about 50 types, each type of special effect contains 10 special effect videos, and the total amount of the special effect data set exceeds 500, which is used for subsequent training of the special effect video generation model. See Figure 4 , Figure 4 is a schematic diagram of a special effect classification shown according to an exemplary embodiment.

[0078] In step 302, for any one of the multiple special effects, the server processes the first-frame image in the sample special-effect video of the special effect and the special-effect prompt word of the special effect through a video generation model to obtain a first special-effect video. The first-frame image is used to indicate the main body in the video to be generated by the video generation model.

[0079] In the embodiments of the present disclosure, for any one of the multiple special effects, the server inputs the first-frame image in the sample special-effect video of the special effect and the special-effect prompt word into the video generation model. The video generation model determines the main body in the special-effect video to be generated from the first-frame image in the sample special-effect video of the special effect, and determines the special effect in the special-effect video to be generated from the special-effect prompt word, so as to generate a first special-effect video based on the main body and the special effect.

[0080] In some embodiments, the video generation model includes an image encoder, a text encoder, and a video generation network. The image encoder is used to extract the encoded features of the first-frame image in the sample special-effect video. The text encoder is used to extract the encoded features of the special-effect prompt word. The video generation network is used to generate a special-effect video based on the encoded features of the first-frame image and the encoded features of the special-effect prompt word. Accordingly, the process by which the server generates the first special-effect video based on the video generation model includes: for any one of the multiple special effects, the server encodes the first-frame image in the sample special-effect video of the special effect through the image encoder in the video generation model to obtain a first image encoded feature. The server encodes the special-effect prompt word of the special effect through the text encoder in the video generation model to obtain a text encoded feature. Then, the server processes the first image encoded feature and the text encoded feature through the video generation network in the video generation model to obtain a first special-effect video.

[0081] In the solution provided by the embodiments of the present disclosure, the video generation model includes an image encoder, a text encoder, and a video generation network. The first-frame image in the sample special-effect video and the special-effect prompt word are respectively encoded by the image encoder and the text encoder to obtain a first image encoded feature and a text encoded feature. Then, the video generation network generates a first special-effect video based on the first image encoded feature and the text encoded feature, so that the picture in the first special-effect video can conform to the sample special-effect video as much as possible, which is conducive to achieving the desired training effect with fewer training times.

[0082] Among them, in the process of generating the first special-effect video, the server can process the first image encoded feature and the text encoded feature through the video generation network in the video generation model to obtain a first video feature, and the first video feature is used to indicate the first special-effect video. That is to say, the server only needs to generate the video feature for subsequent model training, without having to convert to the video level.

[0083] In step 303, the server trains a video generation model based on the first special effect videos corresponding to multiple special effects and the sample special effect videos.

[0084] In the embodiments of the present disclosure, for any one of the multiple special effects, the server can determine the model loss of the video generation model based on the first special effect video corresponding to the special effect and the sample special effect video of the special effect. The model loss is used to represent the difference between the first special effect video generated by the video generation model and the sample special effect video. Then, the server trains the video generation model with the goal of minimizing the model loss, so that the special effect video generated by the video generation model is closer and closer to the sample special effect video.

[0085] Among them, in the process of calculating the model loss, the server can calculate the model loss based on the video features of the first special effect video and the video features of the sample special effect video. Correspondingly, for any one of the multiple special effects, the server determines the first loss based on the first video features and the sample video features of the sample special effect video of the special effect. The first loss is used to represent the difference between the first video features and the sample video features. Then, the server trains the video generation model with the goal of minimizing the first loss. Among them, the video features of the sample special effect video can be extracted by a video encoder independent of the video generation model. The embodiments of the present disclosure do not limit the calculation method of the first loss. The solution provided by the embodiments of the present disclosure calculates the difference between the sample special effect video and the generated first special effect video at the level of video features and trains the video model, without converting the video features into the video itself, which can improve the calculation efficiency of the first loss, thereby facilitating the improvement of the training efficiency of the video generation model.

[0086] In the above training process of the video generation model, the server can adjust all the parameters in the video generation model (including the parameters of the image encoder, the parameters of the text encoder, and the parameters of the video generation network) at the same time; or, the server can also only adjust the parameters of the text encoder. The embodiments of the present disclosure do not limit this.

[0087] In some embodiments, in the above training process, the server only adjusts the parameters of the text encoder. Correspondingly, the server freezes the parameters in other structures of the video generation model except the text encoder and adjusts the parameters of the text encoder in the video generation model. That is to say, the above training process is essentially a training process of the text encoder. The solution provided by the embodiments of the present disclosure only trains the text encoder in the video generation model, which not only enables the text encoder to learn the relationship between the special effect prompt words and the special effect itself, enables the text encoder to accurately understand the special effect to be generated, so as to extract more accurate features of the special effect prompt words, providing a guarantee for accurately generating special effect videos in the future, but also only adjusts the parameters of the text encoder, and the number of adjusted parameters is small, which is conducive to improving the training efficiency.

[0088] For example, Figure 5 is a schematic diagram showing the training of a text encoder according to an exemplary embodiment. Refer to Figure 5 , in the above training process, for any special effect, the server inputs the special effect prompt word of the special effect into the text encoder in the video generation model , inputs the first frame image in the sample special effect video of the special effect into the image encoder in the video generation model, extracts the text encoding features of the special effect prompt word through the text encoder , and extracts the first image encoding features of the first frame image through the image encoder; then, the text encoding features of the special effect prompt word and the first image encoding features of the first frame image enter the video generation network in the video generation model, and are processed by the video generation network to obtain the video features of the first special effect video. The server can also input the sample special effect video of the special effect into the video encoder, and extract the video features of the sample special effect video through the video encoder. Then, the server determines the first loss based on the video features of the first special effect video and the sample video features of the sample special effect video. Then, the server freezes components such as the video generation network, the image encoder, and the video encoder, and trains only the text encoder with the special effect prompt word as the input with the goal of minimizing the first loss.

[0089] Among them, in the above training process, there can be multiple sample special effect videos for each special effect. The subjects in different sample special effect videos of the same special effect are different. That is to say, the server can perform model training through multiple sample special effect videos of the same special effect, so that the video generation model can distinguish the special effect and the subject, accurately learn the features of the special effect, and improve the accuracy of the video generation model.

[0090] After training the video generation model for multiple special effects, the server can match the special effect video with the special effect prompt word , which can solve the problem that the special effect video is difficult to be described by text and the video generation network cannot fully understand and cause generation errors. Introducing the special effect prompt word ( ) as the core carrier of text description in the special effect video generation task can effectively solve the problems of description ambiguity and generation instability existing in traditional text-driven generation, and bring multiple technical advantages. The special effect prompt word can establish a standardized description term library by refining the visual essential features of the special effect (such as "dynamic particle halo", "fluid distortion", etc.), and can accurately anchor the physical properties and dynamic laws of the special effect, reducing semantic ambiguity compared with free text description.

[0091] In step 304, for any one of the multiple special effects, the server encodes the sample video features of the sample special effect video of the special effect through the special effect encoder in the video generation network to obtain special effect features, and the special effect features are used to indicate the attributes of the special effect in the sample special effect video.

[0092] In the embodiments of the present disclosure, the video generation model further includes a special effect encoder. The special effect encoder is used to extract the special effect features in the sample special effect video. For any one of the multiple special effects, the server inputs the sample video features of the sample special effect video of the special effect into the special effect encoder in the video generation model, and encodes the sample video features through the special effect encoder to obtain the special effect features of the special effect. The special effect features can reflect the features or styles of the special effect at the visual level. The embodiments of the present disclosure do not limit the extraction method of the special effect features.

[0093] In some embodiments, the server can extract the special effect features of the special effect through the attention mechanism. Correspondingly, step 304 includes: for any one of the multiple special effects, the server processes the sample video features through the attention layer in the special effect encoder to obtain the query feature, key feature, and value feature of the attention layer. Among them, the query feature and the key feature are used to indicate the position of the special effect in the sample special effect video. The value feature is used to indicate the attributes of the special effect. Then, the server fuses the query feature, key feature, and value feature to obtain the special effect features. Among them, the query feature, key feature, and value feature refer to the matrices corresponding to Q, K, and V in the attention mechanism.

[0094] The solution provided by the embodiments of the present disclosure uses the attention layer to process the sample video features to obtain the query feature, key feature, and value feature of the attention layer, and fuses these three features to obtain the special effect features. Since the query feature and the key feature are used to indicate the position of the special effect in the sample special effect video, and the value feature is used to indicate the attributes of the special effect, the obtained special effect features can accurately reflect the position and style of the special effect in the style special effect video, that is, the accuracy of the special effect features can be improved.

[0095] In step 305, the server processes the special effect features, the first frame image in the sample special effect video of the special effect, and the special effect prompt word of the special effect through the video generation network to obtain the second special effect video.

[0096] In an embodiment of the present disclosure, the server processes the special effect feature of any special effect, the first image encoding feature of the first frame image in the sample special effect video of the special effect, and the text encoding feature corresponding to the special effect prompt word through the video generation network in the video generation model to obtain a second special effect video. The video generation model determines the subject in the special effect video to be generated from the first frame image in the sample special effect video of the special effect, and determines the special effect in the special effect video to be generated from the special effect prompt word and the special effect, so as to generate the second special effect video based on the subject and the special effect. That is, the video generation model determines the special effect to be generated from the text description level and the image vision level.

[0097] In some embodiments, the server processes the special effect feature, the first image encoding feature corresponding to the first frame image, and the text encoding feature corresponding to the special effect prompt word through the video generation network in the video generation network to obtain a second video feature. The second video feature is used to indicate the second special effect video. That is, the server only needs to generate the video feature for subsequent model training without converting it to the video level.

[0098] In step 306, the server retrains the video generation model based on the second special effect videos corresponding to multiple special effects and the sample special effect videos.

[0099] In an embodiment of the present disclosure, for any special effect among multiple special effects, the server can re-determine the model loss of the video generation model based on the second special effect video corresponding to the special effect and the sample special effect video of the special effect. The model loss is used to represent the difference between the second special effect video generated by the video generation model and the sample special effect video. Then, the server retrains the video generation model with the goal of minimizing the model loss, so that the special effect video generated by the video generation model is closer and closer to the sample special effect video.

[0100] Among them, in the process of calculating the model loss, the server can calculate the model loss based on the video feature of the second special effect video and the video feature of the sample special effect video. Correspondingly, for any special effect among multiple special effects, the server determines a second loss based on the second video feature and the sample video feature of the sample special effect video of the special effect. The second loss is used to represent the difference between the second video feature and the sample video feature; the server retrains the video generation model with the goal of minimizing the second loss. The solution provided by the embodiment of the present disclosure calculates the difference between the sample special effect video and the generated second special effect video at the video feature level and trains the video model without converting the video feature into the video itself, which can improve the calculation efficiency of the second loss, thereby facilitating the improvement of the training efficiency of the video generation model.

[0101] During the above training process of the video generation model, the server can adjust all the parameters in the video generation model at the same time; or, the server can also only adjust the parameters of the special effect encoder. The embodiments of the present disclosure do not limit this.

[0102] In some embodiments, during the above training process, the server only adjusts the parameters of the special effect encoder. Accordingly, the server freezes the parameters in other structures in the video generation network except the special effect encoder, and adjusts the parameters in the special effect encoder in the video generation network. That is to say, the training process in steps 304 to 306 is essentially the training process of the special effect encoder. The solution provided by the embodiments of the present disclosure only trains the special effect encoder in the video generation network, which not only enables the special effect encoder to learn the features of the special effects, enables the special effect encoder to accurately understand the special effects to be generated, so as to extract more accurate special effect features, providing guarantee for accurately generating special effect videos subsequently, but also only adjusts the parameters of the special effect encoder, with fewer parameters adjusted, which is beneficial to improving the training efficiency.

[0103] For example, Figure 6 is a schematic diagram of training a special effect encoder shown according to an exemplary embodiment. Refer to Figure 6 , by using the first frame image of the sample special effect video and the special effect prompt word as conditions, the special effect encoder (VFXencoder) is trained. This special effect encoder can extract the sample video features of the sample special effect video and encode them in the form of embedding. For the extracted features, this solution uses the method of linear mapping to convert them into the form of QKV in the attention mechanism. Among them, embedding QK is converted into a one-dimensional vector and directly superimposed on the latent space features of the video generation network to indicate the motion embedding in the time dimension. In addition, embedding V is converted into a two-dimensional vector and also directly superimposed on the latent space features of the video generation network to indicate the motion embedding in the spatial dimension. As Figure 6 shown, then these features will perform the operation of the self-attention mechanism in the time attention layer, and the formula is as follows:

[0104] Among them, Q, K, and V are the query, key, and value matrices obtained from F respectively. dk represents the dimension of the key vector, which serves as a scaling factor to maintain numerical stability in the softmax function. This temporal attention mechanism allows the updated features of each frame to gather information from other frames, enhancing the inter-frame relationship and capturing the temporal continuity necessary for video generation. The purpose of setting embeddingQK as a one-dimensional vector is to extract only the features between frames of the special effect video. Embedding V is set as two-dimensional because there are a large number of deformation changes in the special effect video, and two-dimensional embedding V can better extract the spatial deformation changes in the video.

[0105] In some embodiments, the above video generation network further includes multiple special effect adapters. The multiple special effect adapters correspond one-to-one to each network layer in the video generation network within the video generation network. The video generation network is used to generate special effect videos.

[0106] For any one of the multiple special effects, the server encodes any frame image in the sample special effect video of the special effect through the image encoder in the video generation network to obtain a second image encoding feature. For the special effect adapter corresponding to the i-th network layer in the video generation network, the server processes the third video feature and the second image encoding feature through the special effect adapter to obtain a fourth video feature. The third video feature is the input of the i-th network layer. The third video feature is used to indicate the special effect video containing the special effect generated by the (i - 1)-th network layer. Then, the server fuses the fourth video feature with the output of the i-th network layer to obtain the input of the (i + 1)-th network layer. The server trains the multiple special effect adapters in the video generation network based on the target video feature output by the last network layer in the video generation network and the sample special effect video.

[0107] The solution provided by the embodiments of the present disclosure can decouple the video subject and the special effect in the sample special effect video by configuring adapters on each network layer within the video generation network, learn the physical motion law of the special effect and migrate it to any input picture, accurately capture and reproduce the physical motion law of the special effect, ensure the visual consistency of the generated video, and accurately reflect the characteristics of the special effect.

[0108] Among them, the server calculates a third loss based on the target video feature output by the last network layer in the video generation network and the sample video feature of the sample special effect video. The third loss is used to represent the difference between the target video feature and the sample video feature. Then, the server trains the multiple special effect adapters in the video generation network with the goal of minimizing the third loss. When the third loss is lower than a preset value or the number of parameter adjustment times of the special effect video adapter reaches a preset number of times, the server stops training the multiple special effect adapters.

[0109] For example, Figure 7 is a schematic diagram of a training special effect adapter shown according to an exemplary embodiment. Refer to Figure 7 , in order to further enable the model to understand the special effects, this solution designs an additional special effect adapter. Its input includes a random frame image in the sample special effect video. The purpose of doing this is to learn the physical motion pattern in the sample special effect video and decouple it from the subject undergoing special effect changes in the video.

[0110] In some embodiments, for the special effect adapter corresponding to the i-th network layer in the video generation network, the server broadcasts the second image encoding feature to the image features corresponding to each image frame in the third video feature through the special effect adapter to obtain the third video feature. In the solution provided by the embodiments of the present disclosure, the second image encoding feature can reflect the video subject. By broadcasting the second image encoding feature to all images, it can be used as an appearance guidance during training, enabling the model to make the learned subject move according to the expected motion pattern in the existing video, so that the decoupling between the complete special effect video and the random video spatial information can capture more purely the physical visual effects that do not exist in real life. After being injected into the video generation network, it further enhances the performance of the special effect generation result.

[0111] For example, Figure 8 is a schematic diagram of a special effect adapter shown according to an exemplary embodiment. Refer to Figure 8 , the special effect adapter includes a downsampling layer ( ), a nonlinear layer (Nonlinear), an upsampling layer ( ), and a conditional layer ( ). The downsampling layer is used to reduce the dimension or resolution of the feature map. Through the downsampling operation, the number of channels of the feature map after non-linear transformation can be reduced, thereby reducing the computational complexity while retaining the most important feature information. The nonlinear layer is usually a non-linear activation function, such as ReLU (Rectified Linear Unit). The role of the non-linear activation function is to introduce non-linearity, enabling the model to learn more complex feature representations. Without the non-linear activation function, the neural network will only be able to learn linear transformations, which will greatly limit the expressive ability of the model. The upsampling layer is usually used to increase the dimension or resolution of the feature map. Through the upsampling operation, the number of channels of the input feature map can be increased, thereby providing more feature information for subsequent non-linear transformations. The conditional layer is usually used to fuse the output of the adapter module with the original input. Through the conditional layer, the features extracted by the adapter module can be weighted and fused with the original input features, thereby enhancing the expressive ability of the model. The output of the conditional layer is usually added element-wise to the original input feature map (as shown by the plus sign in Figure 8 ) to achieve feature fusion.

[0112] Among them, the special effect adapter can be customized according to a type of video (e.g., videos representing various dog movements), multiple videos showing the same movement, or even the movement patterns extracted from a single video. Although the special effect adapter can capture movement patterns, during the training process, it inevitably learns the appearance of the subjects in the input videos. To separate spatial and temporal information, we add appearance guidance in the special effect adapter to force it to learn pure movement. Specifically, we add a conditional linear layer with weights to integrate appearance information into the temporal hidden state . Among them, is the hidden state at time t, with a dimension of , is the batch size, is the size of the feature map, is the number of features, is some additional dimension (such as the number of layers). Then, we randomly select a frame from the training videos and obtain its image embedding through the image encoder. Among them, is the image embedding obtained from the image encoder, with a dimension of , is the number of features of the embedding. This image embedding is then broadcast to all frames as the appearance guidance during training. The forward process of the special effect adapter is formulated as:

[0113]

[0114] Among them, is the output of the special effect adapter. is the activation function, usually a non-linear function such as ReLU or sigmoid, used to introduce non-linear characteristics; and These are weight matrices used for further transforming the hidden state, usually used for dimension reduction and dimension increase operations to adapt to different layers or modules of the network. This operation first calculates the matrix multiplication of e and , and then broadcasts the result to the same shape as so that element-wise addition can be performed. During inference, the present invention randomly selects the training images provided by the user as the appearance condition input of the special effect adapter.

[0115] During the model usage stage, the video generation model can generate special effect customization based on the input reference image according to any input reference image and the special effect prompt as a condition. For example, Figure 9It is a schematic diagram of a model inference stage shown according to an exemplary embodiment.

[0116] The embodiments of the present disclosure provide a method for training a video generation model. Through a video generation network, for the first-frame image of the sample special-effect video of each special effect and the special-effect prompt word, a first special-effect video of each special effect is generated. Through the first special-effect video of each special effect and the corresponding sample special-effect video, the video generation network is trained. This not only makes the special-effect videos generated by the video generation network more and more conform to the sample special-effect videos, that is, the generated special-effect videos are more and more accurate, but also realizes the association between the special-effect prompt words and the special-effect videos, and the video generation network can learn this association relationship. Thus, the video generation network can accurately determine the style of the special effect through the special-effect prompt word. Compared with the user's custom description related to the special effect, in this solution, the user does not need to consider the description of the special effect, and the semantic ambiguity related to the special effect can be reduced through the special-effect prompt word. Therefore, a special-effect video that meets the requirements can be automatically generated based on the special-effect prompt word, improving the generation efficiency of the special-effect video while ensuring the generation effect of the special-effect video.

[0117] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated one by one here.

[0118] Figure 10 It is a block diagram of a training device for a video generation model shown according to an exemplary embodiment. Refer to Figure 10 This training device for the video generation model includes an acquisition module 1001, a first processing module 1002, and a training module 1003.

[0119] The acquisition module 1001 is configured to acquire the sample special-effect videos and special-effect prompt words of multiple special effects respectively, and the special-effect prompt words are used to indicate the special effects in the videos to be generated by the video generation network; The first processing module 1002 is configured to, for any one of the multiple special effects, process the first-frame image in the sample special-effect video of the special effect and the special-effect prompt word of the special effect through the video generation network to obtain a first special-effect video, and the first-frame image is used to indicate the main body in the video to be generated by the video generation network; The training module 1003 is configured to train the video generation network based on the first special-effect videos and sample special-effect videos corresponding to the multiple special effects.

[0120] In some embodiments, the first processing module 1002 includes: The first encoding unit is configured to, for any one of the multiple special effects, encode the first-frame image in the sample special-effect video of the special effect through the image encoder in the video generation network to obtain a first image encoding feature; A second encoding unit, configured to perform encoding on the special effect prompt words of the special effect through a text encoder in a video generation network to obtain text encoding features; A processing unit, configured to perform processing on the first image encoding features and the text encoding features through a video generation network in the video generation network to obtain a first special effect video.

[0121] In some embodiments, the processing unit is configured to perform processing on the first image encoding features and the text encoding features through a video generation network in the video generation network to obtain first video features, where the first video features are used to indicate the first special effect video; A training module, configured to perform, for any one of multiple special effects, determining a first loss based on the first video features and the sample video features of the sample special effect video of the special effect, where the first loss is used to represent the difference between the first video features and the sample video features; and training the video generation model with the goal of minimizing the first loss.

[0122] In some embodiments, the training module 1003 is configured to perform freezing the parameters in other structures in the video generation network except for the text encoder, and adjusting the parameters of the text encoder in the video generation network.

[0123] In some embodiments, the apparatus further includes: A special effect encoding module, configured to perform, for any one of multiple special effects, encoding the sample video features of the sample special effect video of the special effect through a special effect encoder in the video generation network to obtain special effect features, where the special effect features are used to indicate the attributes of the special effect in the sample special effect video; The first processing module 1002 is further configured to perform processing on the special effect features, the first frame image in the sample special effect video of the special effect, and the special effect prompt words of the special effect through the video generation network to obtain a second special effect video; The training module 1003 is further configured to perform re-training the video generation network based on the second special effect videos corresponding to multiple special effects and the sample special effect videos.

[0124] In some embodiments, the first processing module is configured to perform, for any one of multiple special effects, processing the sample video features through an attention layer in the special effect encoder to obtain a query feature, a key feature, and a value feature of the attention layer, where the query feature and the key feature are used to indicate the position of the special effect in the sample special effect video, and the value feature is used to indicate the attribute of the special effect; and fusing the query feature, the key feature, and the value feature to obtain special effect features.

[0125] In some embodiments, the first processing module 1002 is configured to process the special effect features, the first image encoding features corresponding to the first frame image, and the text encoding features corresponding to the special effect prompt words through a video generation network in the video generation network to obtain second video features, and the second video features are used to indicate a second special effect video. The training module 1003 is configured to, for any one of multiple special effects, determine a second loss based on the second video features and the sample video features of the sample special effect video of the special effect, where the second loss is used to represent the difference between the second video features and the sample video features; and train the video generation model with the goal of minimizing the second loss.

[0126] In some embodiments, the training module 1003 is configured to freeze the parameters in other structures in the video generation network except for the special effect encoder and adjust the parameters in the special effect encoder in the video generation network.

[0127] In some embodiments, the video generation network further includes a plurality of special effect adapters, and the plurality of special effect adapters correspond one-to-one to each network layer in the video generation network in the video generation network, and the video generation network is used to generate special effect videos. The apparatus further includes: The second processing module is configured to, for any one of multiple special effects, encode any frame image in the sample special effect video of the special effect through an image encoder in the video generation network to obtain second image encoding features; for the special effect adapter corresponding to the i-th network layer in the video generation network, process the third video features and the second image encoding features through the special effect adapter to obtain fourth video features, where the third video features are the input of the i-th network layer, and the third video features are used to indicate the special effect video including the special effect generated by the (i - 1)-th network layer; fuse the fourth video features with the output of the i-th network layer to obtain the input of the (i + 1)-th network layer; and train the plurality of special effect adapters in the video generation network based on the target video features output by the last network layer in the video generation network and the sample special effect video.

[0128] In some embodiments, the first processing module 1002 is configured to, for the special effect adapter corresponding to the i-th network layer in the video generation network, broadcast the second image encoding features to the image features corresponding to each image frame in the third video features through the special effect adapter to obtain the third video features.

[0129] In some embodiments, the multiple special effects include interactive special effects, and the interactive special effects are used to indicate special effects in which both the main body and the environment in the video change.

[0130] An embodiment of the present disclosure provides a training device for a video generation model. Through a video generation network, for the first frame image and special effect prompt words of the sample special effect video of each special effect, a first special effect video of each special effect is generated. Through the first special effect video of each special effect and the corresponding sample special effect video, the video generation network is trained. This not only makes the special effect video generated by the video generation network more and more conform to the sample special effect video, that is, the generated special effect video is more and more accurate, but also realizes the association between the special effect prompt words and the special effect video, and the video generation network can learn this association relationship. Thus, the video generation network can accurately determine the style of the special effect through the special effect prompt words. Compared with the user's custom description related to the special effect, in this solution, the user does not need to consider the description of the special effect, and the semantic ambiguity related to the special effect can be reduced through the special effect prompt words. Therefore, a special effect video that meets the requirements can be automatically generated based on the special effect prompt words, improving the generation efficiency of the special effect video while ensuring the generation effect of the special effect video.

[0131] It should be noted that when the training device for the video generation model provided in the above embodiment trains the video generation model, only the division of the above functional units is used for illustration. In practical applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the training device for the video generation model provided in the above embodiment and the method embodiment for training the video generation model belong to the same concept, and the specific implementation process is detailed in the method embodiment and will not be elaborated here.

[0132] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0133] When the electronic device is provided as a terminal, Figure 11 is a block diagram of a terminal 1100 shown according to an exemplary embodiment. The terminal Figure 11 shows a block diagram of the structure of the terminal 1100 provided in an exemplary embodiment of the present disclosure. The terminal 1100 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio LayerIII), an MP4 (Moving Picture Experts Group AudioLayer IV) player, a notebook computer or a desktop computer. The terminal 1100 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0134] Generally, the terminal 1100 includes: a processor 1101 and a memory 1102.

[0135] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0136] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1101 to implement the training method of the video generation model provided in the method embodiments of the present application.

[0137] In some embodiments, the terminal 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0138] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0139] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1104 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0140] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input as control signals to the processor 1101 for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, which is disposed on the front panel of the terminal 1100; in other embodiments, there may be at least two display screens 1105, which are respectively disposed on different surfaces of the terminal 1100 or are in a foldable design; in other embodiments, the display screen 1105 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 1100. Even further, the display screen 1105 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1105 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0141] The camera module 1106 is used to collect images or videos. In some embodiments, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, to implement functions such as the fusion of the main camera and the depth camera to achieve the background blur function, the fusion of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions, or other fusion shooting functions. In some embodiments, the camera module 1106 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0142] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may further include a headphone jack.

[0143] The power supply 1108 is used to supply power to each component in the terminal 1100. The power supply 1108 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1108 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0144] Those skilled in the art can understand that Figure 10 the structure shown in

[0145] does not limit the terminal 1100, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements. Figure 12 When the electronic device is provided as a server,

[0146] In an exemplary embodiment, a computer-readable storage medium including instructions is further provided, such as the memory 1102 or the memory 1202 including instructions. The above instructions can be executed by the processor 1101 of the terminal 1100 or the processor 1201 of the server 1200 to complete the above method for training the video generation model. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0147] A computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the above method for training the video generation model is implemented.

[0148] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0149] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for a video generation model, characterized in that The method includes: Obtaining sample special effect videos and special effect prompt words for respective multiple special effects, where the special effect prompt words are used to indicate special effects in the videos to be generated by the video generation model; For any one of the multiple special effects, through the video generation model, processing a first frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect to obtain a first special effect video, where the first frame image is used to indicate the main body in the video to be generated by the video generation model; Training the video generation model based on the first special effect videos and the sample special effect videos corresponding to the multiple special effects.

2. The training method of the video generation model according to claim 1, wherein, The step of, for any one of the multiple special effects, through the video generation model, processing a first frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect to obtain a first special effect video includes: For any one of the multiple special effects, through the image encoder in the video generation model, encoding the first frame image in the sample special effect video of the special effect to obtain a first image encoding feature; Encoding the special effect prompt word of the special effect through the text encoder in the video generation model to obtain a text encoding feature; Processing the first image encoding feature and the text encoding feature through the video generation network in the video generation model to obtain the first special effect video.

3. The training method of the video generation model according to claim 2, wherein The step of, through the video generation network in the video generation model, processing the first image encoding feature and the text encoding feature to obtain the first special effect video includes: Processing the first image encoding feature and the text encoding feature through the video generation network in the video generation model to obtain a first video feature, where the first video feature is used to indicate the first special effect video; The step of, based on the first special effect videos and the sample special effect videos corresponding to the multiple special effects, training the video generation model includes: For any one of the multiple special effects, determining a first loss based on the first video feature and the sample video feature of the sample special effect video of the special effect, where the first loss is used to represent the difference between the first video feature and the sample video feature; Training the video generation model with the goal of minimizing the first loss.

4. The training method of the video generation model according to claim 2 or 3, characterized in that The step of training the video generation model includes: Freezing the parameters in other structures of the video generation model except the text encoder, and adjusting the parameters of the text encoder in the video generation model.

5. The training method of the video generation model according to claim 1, wherein The method further includes: For any one of the multiple special effects, encoding the sample video feature of the sample special effect video of the special effect through the special effect encoder in the video generation model to obtain a special effect feature, where the special effect feature is used to indicate the attribute of the special effect in the sample special effect video; Processing the special effect feature, the first frame image in the sample special effect video of the special effect, and the special effect prompt word of the special effect through the video generation model to obtain a second special effect video; Retraining the video generation model based on the second special effect videos and the sample special effect videos corresponding to the multiple special effects.

6. The training method of the video generation model according to claim 5, characterized in that For any one of the multiple special effects, encoding the sample video features of the sample special effect video of the special effect through a special effect encoder in a video generation model to obtain special effect features, including: For any one of the multiple special effects, processing the sample video features through an attention layer in the special effect encoder to obtain query features, key features, and value features of the attention layer, where the query features and the key features are used to indicate the position of the special effect in the sample special effect video, and the value features are used to indicate the attributes of the special effect; Fusing the query features, the key features, and the value features to obtain the special effect features.

7. The training method of the video generation model according to claim 5, wherein Processing the special effect features, the first frame image in the sample special effect video of the special effect, and the special effect prompt word of the special effect through the video generation model to obtain a second special effect video, including: Processing the special effect features, the first image encoding features corresponding to the first frame image, and the text encoding features corresponding to the special effect prompt word through a video generation network in the video generation model to obtain second video features, where the second video features are used to indicate the second special effect video; Retraining the video generation model based on the second special effect videos and the sample special effect videos corresponding to the multiple special effects, including: For any one of the multiple special effects, determining a second loss based on the second video features and the sample video features of the sample special effect video of the special effect, where the second loss is used to represent the difference between the second video features and the sample video features; Training the video generation model with the goal of minimizing the second loss.

8. The training method of the video generation model according to any one of claims 5-7, characterized in that The retraining of the video generation model includes: Freezing the parameters in other structures except the special effect encoder in the video generation model, and adjusting the parameters in the special effect encoder in the video generation model.

9. The training method of the video generation model according to claim 1, characterized in that The video generation model further includes multiple special effect adapters, where the multiple special effect adapters correspond one-to-one to each network layer in the video generation network in the video generation model, and the video generation network is used to generate special effect videos; The method further includes: For any one of the multiple special effects, encoding any frame image in the sample special effect video of the special effect through an image encoder in the video generation model to obtain second image encoding features; For the special effect adapter corresponding to the i-th network layer in the video generation network, processing the third video features and the second image encoding features through the special effect adapter to obtain fourth video features, where the third video features are the input of the i-th network layer, and the third video features are used to indicate the special effect video containing the special effect generated by the (i - 1)-th network layer; Fusing the fourth video features with the output of the i-th network layer to obtain the input of the (i + 1)-th network layer; Training the multiple special effect adapters in the video generation model based on the target video features output by the last network layer in the video generation network and the sample special effect videos.

10. The training method of the video generation model according to claim 9, wherein For the special effect adapter corresponding to the i-th network layer in the video generation network, processing the third video feature and the second image encoding feature through the special effect adapter to obtain a fourth video feature includes: For the special effect adapter corresponding to the i-th network layer in the video generation network, through the special effect adapter, broadcasting the second image encoding feature to the image features corresponding to each image frame in the third video feature to obtain the third video feature.

11. The training method of the video generation model according to claim 1, characterized in that, The multiple special effects include interactive special effects, and the interactive special effects are used to indicate special effects in which both the main body and the environment in the video change.

12. A training device for a video generation model, characterized in that, The device includes: An acquisition module configured to acquire sample special effect videos and special effect prompt words of multiple special effects respectively, where the special effect prompt words are used to indicate special effects in the video to be generated by the video generation model; A first processing module configured to, for any one of the multiple special effects, process the first frame image in the sample special effect video of the special effect and the special effect prompt word of the special effect through the video generation model to obtain a first special effect video, where the first frame image is used to indicate the main body in the video to be generated by the video generation model; A training module configured to train the video generation model based on the first special effect videos and the sample special effect videos corresponding to the multiple special effects.

13. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the training method of the video generation model according to any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method of the video generation model according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the training method of the video generation model according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Video generation model training method and device, equipment and storage medium

    CN117499711A

  • Content generation method and device based on text cue word and image driving and medium

    CN117911584A

  • Video processing method and device, equipment and medium

    CN119011930A

  • Special effect template generation method and device, electronic equipment and storage medium

    CN119946211A

  • Special effect template generation method and device, electronic equipment and storage medium

    CN120104026A