Special effects video generation methods, devices, electronic equipment and storage media
By generating target images and calling video generation models to process them, the problem of poor realism and coherence in the generation of complex special effects videos is solved, thereby improving the realism and coherence of special effects videos.
Patent Information
- Application Number
- CN202411757940.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing technologies suffer from poor realism and coherence when generating videos with complex special effects, especially in scenarios involving hairstyle effects, clothing effects, and animal-to-human transformation effects. The uncontrolled changes in video frames result in discontinuous and unrealistic videos.
By acquiring the image to be processed, generating the target image, and calling the video generation model to process the image to be processed and the target image, a special effects video is generated. The special effects video includes the first frame, the last frame, and the middle frames. The first frame has the same content as the image to be processed, the last frame has the same content as the target image, and the middle frames represent the continuous change process of the target effect.
It improves the realism and coherence of special effects videos, making the content changes in special effects videos smoother and more realistic, and enhancing video quality.
Smart Images

Figure CN119583736B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and in particular to a method, apparatus, electronic device, and storage medium for generating special effects videos. Background Technology
[0002] Currently, in scenarios where video generation is based on artificial intelligence (AI) technology, users can input a static image and a statement describing the video content into the video generation model to generate a corresponding video, thereby achieving the effect of making the image "move".
[0003] However, in special content generation scenarios, such as video generation for hairstyle effects, existing solutions cannot generate realistic and coherent videos with such complex effects due to the special nature of these effects. In other words, the generated videos with complex effects have poor realism and coherence. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for generating special effects videos, thereby overcoming the problem of poor realism and coherence in generating videos with special effects.
[0005] In a first aspect, embodiments of this disclosure provide a method for generating special effects videos, including:
[0006] Obtain an image to be processed; based on the image to be processed, generate a target image, wherein the target image is an image generated after applying a target effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect; call a video generation model to process the image to be processed and the target image to generate an effect video, wherein the effect video includes a first frame, a last frame, and at least one intermediate frame located between the first frame and the last frame, the first frame has the same image content as the image to be processed, the last frame has the same image content as the target image, and the at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
[0007] Secondly, embodiments of this disclosure provide a special effects video generation apparatus, comprising:
[0008] The acquisition module is used to acquire the image to be processed;
[0009] The first generation module is used to generate a target image based on the image to be processed, wherein the target image is an image generated after applying a target effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect;
[0010] The second generation module is used to call a video generation model to process the image to be processed and the target image to generate a special effects video. The special effects video includes a first frame, a last frame, and at least one intermediate frame between the first frame and the last frame. The first frame has the same image content as the image to be processed, and the last frame has the same image content as the target image. The at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
[0011] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0012] The memory stores computer-executed instructions;
[0013] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the special effects video generation method as described in the first aspect and various possible designs of the first aspect.
[0014] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the special effects video generation method described in the first aspect and various possible designs of the first aspect.
[0015] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the special effects video generation method described in the first aspect and various possible designs of the first aspect.
[0016] The special effects video generation method, apparatus, electronic device, and storage medium provided in this embodiment acquire an image to be processed; generate a target image based on the image to be processed, wherein the target image is an image generated after applying a target special effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target special effect; call a video generation model to process the image to be processed and the target image to generate a special effects video, wherein the special effects video includes a first frame, a last frame, and at least one intermediate frame located between the first frame and the last frame, the first frame has the same image content as the image to be processed, the last frame has the same image content as the target image, and the at least one intermediate frame is used to characterize the continuous change process of generating the target effect. First, a target image with special effects is generated based on the image to be processed. Then, a video generation model is called to process the image to be processed and the target image to obtain intermediate video frames that represent the continuous change process of the target effect from the image to the target image. Finally, a special effects video consisting of the first frame, intermediate frames, and last frame is generated. The last frame ensures the realism of the special effects while making the content changes of the special effects video smoother and more realistic, thus improving the video quality of the special effects video. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 An application scenario diagram of the special effects video generation method provided in the embodiments of this disclosure;
[0019] Figure 2 Flowchart of the special effects video generation method provided in the embodiments of this disclosure Figure 1 ;
[0020] Figure 3 A schematic diagram of a target image provided in an embodiment of this disclosure;
[0021] Figure 4 for Figure 2 A flowchart illustrating the specific implementation of step S102 in the illustrated embodiment;
[0022] Figure 5 for Figure 4 A flowchart illustrating the specific implementation of step S1022 in the illustrated embodiment;
[0023] Figure 6 for Figure 2 A flowchart illustrating the specific implementation of step S103 in the illustrated embodiment;
[0024] Figure 7 Flowchart of the special effects video generation method provided in the embodiments of this disclosure Figure 2 ;
[0025] Figure 8 for Figure 7 A flowchart illustrating the specific implementation of step S205 in the illustrated embodiment;
[0026] Figure 9 for Figure 8 A flowchart illustrating the specific implementation of step S2053 in the illustrated embodiment;
[0027] Figure 10 This is a schematic diagram illustrating the process of generating a special effects video according to an embodiment of the present disclosure;
[0028] Figure 11 A structural block diagram of the special effects video generation apparatus provided in the embodiments of this disclosure;
[0029] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;
[0030] Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0033] The application scenarios of the embodiments of this disclosure are explained below:
[0034] Figure 1This diagram illustrates an application scenario of the special effects video generation method provided in this embodiment. This method can be applied to applications (APPs) with special effects video generation capabilities, such as video generation applications. More specifically, it can be applied to application scenarios involving special effects video generation. The executing entity in this embodiment can be a terminal device running the aforementioned application with special effects video generation capabilities, a server deploying the server-side component corresponding to the aforementioned application, or other electronic devices performing similar functions. When the executing entity is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application. When the executing entity is a server, the server-side component of the aforementioned application with special effects video generation capabilities can run partially or entirely on the server, executing the method provided in this embodiment on the server side, while the terminal device runs the client-side component of the application. Communication between the server and the terminal device is based on server-client communication, enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.
[0035] In some embodiments, the terminal device or server can implement the special effects video generation method provided in this disclosure by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run, or mini-programs embedded in any APP, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin; the specific implementation can be configured as needed. Furthermore, in implementing the special effects video generation method provided in this disclosure, the terminal device can execute the method by running computer-executable instructions or computer programs set locally, or by calling computer-executable instructions or computer programs set in an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.
[0036] refer to Figure 1As shown in the diagram, taking a terminal device as an example, the terminal device runs a target application with special effects video generation capabilities. The target application loads a static image input by the user (such as a selfie uploaded by the user). Then, in response to the user's trigger operation (such as clicking the "Generate" button), it calls a video generation model deployed in the cloud or locally to convert the static image into a corresponding dynamic video. During this process, the target application can also receive a descriptive statement input by the user (shown as ABCDEF in the diagram) and control the video content of the generated video according to the descriptive statement, thereby generating a dynamic video that meets the user's preferences and needs. Alternatively, the user can control the content of the corresponding special effects video generated based on the static image by clicking on the special effects cards (special effects controls) provided by the target application.
[0037] In existing technologies, in scenarios involving the generation of special video content, such as special effects videos involving hairstyles, clothing, or animals transforming into humans, the unique nature of these complex effects necessitates accurately representing their appearance and transformation process. However, existing solutions often use static images as a starting point to generate subsequent video frames, which frequently results in uncontrolled changes in the generated video frames. This leads to poor continuity between video frames and significant discrepancies between the final frame and the intended effect, ultimately resulting in disjointed and unrealistic special effects videos.
[0038] This disclosure provides a method for generating special effects videos to solve the above-mentioned problems.
[0039] refer to Figure 2 , Figure 2 Flowchart of the special effects video generation method provided in the embodiments of this disclosure Figure 1 The method described in this embodiment can be applied to terminal devices or servers. This special effects video generation method includes:
[0040] Step S101: Obtain the image to be processed.
[0041] Step S102: Based on the image to be processed, generate a target image, wherein the target image is an image generated after applying a target effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect.
[0042] refer to Figure 1The illustrated application scenario diagram illustrates the method for generating special effects videos using a terminal device as the execution subject. For example, firstly, after running the target application, the terminal device receives an image uploaded by the user, i.e., the image to be processed. This image is the starting point for generating the final special effects video. Specifically, the image to be processed is, for example, a user-uploaded photo containing specific objects, such as people or animals, which are the main content of the image. Then, the terminal device processes the image to be processed using the capabilities provided by the target application to generate a target image. The target image is the image after adding the target special effects to the image to be processed; that is, the target special effects image corresponding to the image to be processed. A target special effect is a special effect used to add a target category of effect to a target object. An object is an object that changes due to the addition of a special effect, such as the changed "hair" in a hairstyle effect, the "person" in a dress-up effect, or the "animal" in an animal-to-human effect, and so on. Figure 3 A schematic diagram of a target image provided in an embodiment of this disclosure, such as... Figure 3 As shown, exemplarily, the image to be processed is an image containing a person object A (the target object). Referring to the figure, the current hairstyle of person object A is, for example, a first hairstyle (short hair). The terminal device processes the image to be processed by calling the corresponding algorithm and model, and generates a target image. The target image still contains person object A, but person object A's hairstyle changes to a second hairstyle (long hair). This second hairstyle is the target effect of the target category generated by the target effect. In other possible implementations, the generated target image changes with the change of video effect type. For example, in the "clothing change effect" scenario (not shown in the figure), the target object is person object B in the image to be processed. Person object B has a first outfit. The terminal device processes the image to be processed by calling the corresponding algorithm and model, and generates a target image. The target image still contains person object B (the target object), but person object B's outfit changes to a second style of outfit. This second outfit is the target effect after applying the target effect to the target object in the image to be processed.
[0043] Furthermore, in other possible scenarios, the target effect can also include other video effects such as "animal to human effect," "human to animal effect," and "object material change effect." Depending on the target effect, the generated target image changes accordingly, and ultimately the content of the effect video generated based on the target image also changes. That is, the solution provided in this embodiment can be applied to the generation of effect videos in different scenarios; the specific settings can be configured as needed, and no limitations are imposed here.
[0044] Furthermore, in one possible implementation, such as Figure 4 As shown, the specific implementation of step S102 includes:
[0045] Step S1021: Obtain the first cue word for the target effect of the target category. The first cue word is used to describe the appearance characteristics of the target effect of the target category.
[0046] Step S1022: Process the image to be processed and the first prompt word using an image generation model to generate the target image.
[0047] For example, in the process of generating the target image, the terminal device calls an image generation model to complete the target image generation step. Specifically, the image generation model is a pre-trained neural network model used to regenerate the input image based on prompt words. The terminal device first receives a first prompt word as input through the target application it is running. This first prompt word describes the appearance features of the target effect of the target category. For example, the content of the first prompt word includes "long curly hair" or "full and dense natural small curls". Then, the image to be processed and the first prompt word are input into the image generation model. Using the capabilities of the image generation model, the image content in the image to be processed is "modified" into an image containing the target effect of the target category described by the first prompt word, i.e., the target image. The above process and the process of achieving the target effect are described. The first prompt word can be preset configuration information, dynamically determined according to the image content of the image to be processed, or it can be information input by the user, which can be set as needed.
[0048] In this embodiment, the appearance characteristics of the target effect are defined by the first prompt word, thereby achieving precise control over the visual effect of the target effect, making the effect more diverse, differentiated and refined, and ultimately improving the realism and visual performance of the generated effect video.
[0049] Furthermore, in one possible implementation, the image generation model includes a first functional unit and a second functional unit. The first functional unit generates a target effect for a target category, and the second functional unit generates a second object that matches the target effect based on the target object. The appearance similarity between the second object and the target object is greater than a similarity threshold. Figure 5 As shown, the specific implementation of step S1022 includes:
[0050] Step S1022-1: Process the image to be processed and the first prompt word through the first functional unit to generate effect feature data representing the target effect.
[0051] Step S1022-2: The second functional unit processes the image to be processed and the effect feature data to generate object feature data representing the second object.
[0052] Step S1022-3: Generate the target image based on the effect feature data and object feature data.
[0053] For example, in one possible implementation, the image generation model consists of a first functional unit and a second functional unit. The first functional unit is used to generate the target effect of the target category. For example, the first functional unit has the ability to edit hairstyles, such as generating feature data corresponding to hairstyles like short hair, curly hair, and long straight blonde hair. The feature data can be a set of pixels or other feature vectors or matrices that need to be further decoded to represent the image. The first functional unit processes the image to be processed and the first prompt word, and generates corresponding feature data based on the appearance features described by the first prompt word. This feature data can realize the image or feature description of the target effect of the target category. The second functional unit is used to generate object feature data representing the second object. That is, the second functional unit has the ability to edit the object, such as generating feature data corresponding to human faces, animal faces, and physical contours in the image to be processed. In other words, the target object in the image to be processed is edited into the second object (i.e., to achieve a "face swap" effect). This makes the second object in the generated target image more consistent with the target effect generated by the first functional unit, improving the image realism. On the other hand, it maintains a high similarity between the target object and the second object, so that most of the appearance features of the target object in the image to be processed can be preserved, improving the consistency between the main object (second object) in the target image after adding the target effect and the main object (target object) in the image to be processed without adding the target effect.
[0054] Furthermore, in one possible implementation, the first prompt word includes a first segmentation, a second segmentation, and a third segmentation. The first segmentation indicates the target object; the second segmentation characterizes the shape features of the target effect; and the third segmentation characterizes the style features of the target effect. That is, the first prompt word consists of three parts: the first segmentation, the second segmentation, and the third segmentation. The first segmentation controls where the image generation model generates the target effect (target object); the second segmentation controls the local appearance (shape features) of the target effect, such as hair shape and color; and the third segmentation characterizes the overall features (style features) of the target effect, such as hair volume, curl curvature, and density. Through these first, second, and third segmentations, more detailed control over the target effect can be achieved, enabling the generated target image and the special effects video generated based on the target image to better meet user needs and have better visual effects.
[0055] Step S103: Call the video generation model to process the image to be processed and the target image to generate a special effects video. The special effects video includes the first frame of the video, the last frame of the video, and at least one intermediate frame of the video located between the first frame of the video and the last frame of the video. The first frame of the video has the same image content as the image to be processed, the last frame of the video has the same image content as the target image, and at least one intermediate frame of the video is used to characterize the continuous change process of the generated target effect.
[0056] For example, after generating the target image, the image content of the image to be processed is used as the first frame content of the special effects video, and the image content of the target image is used as the last frame content of the special effects video. The video generation model is then called to generate the intermediate frame content of the video that continuously changes between the first frame content and the last frame content, thereby obtaining a video that describes the continuous change of the target effect from the first frame content to the last frame content, i.e., the special effects video. The video generation model has the ability to generate images that continuously change between two static images. This ability is obtained through pre-training based on training samples, which will not be elaborated here.
[0057] In one possible implementation, such as Figure 6 As shown, the specific implementation of step S103 includes:
[0058] Step S1031: Obtain the second cue word, which is used to describe the characteristics of the change process of the target effect in the special effects video;
[0059] Step S1032: Call the video generation model, process the image to be processed and the target image based on the second prompt word, and generate a special effects video with change process characteristics.
[0060] For example, before invoking the video generation model, the terminal device first obtains a second prompt word. Similar to the first prompt word, the second prompt word can be preset configuration information or information input by the user, and can be set as needed. The second prompt word describes the changing process characteristics of the target effect in the special effects video. For example, the content of the second prompt word includes: "The hair gradually becomes curly and fluffy, with a natural and gradual smooth dynamic." Through the above, the second prompt word can control the changing process of the special effects video, thereby making the changing process between the image content of the image to be processed and the image content of the target image more accurate, meeting the user's personalized needs, and making the generated special effects video more realistic. It should be noted that the content of the second prompt word in this embodiment is only exemplary; the specific content of the second prompt word can be set as needed, and will not be elaborated here.
[0061] In this embodiment, an image to be processed is acquired; a target image is generated based on the image to be processed, wherein the target image is the image after adding target effects to the image to be processed, the image content of the image to be processed contains a target object, and the target effects are used to add target effects of the target category to the target object; a video generation model is called to process the image to be processed and the target image to generate an effects video, wherein the effects video includes a first frame, a last frame, and at least one intermediate frame located between the first frame and the last frame. The first frame has the same image content as the image to be processed, the last frame has the same image content as the target image, and at least one intermediate frame is used to represent the continuous change process of the generated target effect. By first generating a target image with effects based on the image to be processed, then calling the video generation model to process the image to be processed and the target image, intermediate frames representing the continuous change process of the target effect from the image to the target image are obtained. Finally, an effects video composed of the first frame, intermediate frames, and last frame is generated. The last frame ensures the realism of the effects while making the content changes of the effects video smoother and more realistic, thus improving the video quality of the effects video.
[0062] refer to Figure 7 , Figure 7 Flowchart of the special effects video generation method provided in the embodiments of this disclosure Figure 2 This embodiment is in Figure 2 Based on the illustrated embodiment, step S103 is further refined. In this embodiment, the video generation model includes a variational autoencoder unit and a diffusion transformer unit. The special effects video generation method includes:
[0063] Step S200: Jointly train the variational autoencoder unit and the diffusion converter unit.
[0064] Step S201: Obtain the image to be processed.
[0065] Step S202: Based on the image to be processed, generate a target image, wherein the target image is the image after adding target effects to the image to be processed, the image content of the image to be processed contains a target object, and the target effects are used to add target effects of the target category to the target object.
[0066] Step S203: Process the image to be processed and the target image through the variational autoencoder unit to generate the first feature code corresponding to the image to be processed and the second feature code corresponding to the target image, respectively.
[0067] Step S204: Concatenate the first feature code, the blank frame code, and the second feature code sequentially to obtain the input feature code.
[0068] Step S205: Process the input feature encoding through the diffusion transformer unit to generate special effects video.
[0069] For example, after generating the target image, a video generation model is invoked to process the image to be processed and the target image to generate a special effects video. The video generation model includes a Variational Autoencoder (VAE) unit and a Diffusion Transformer (DFT) unit. The VAE is a generative model that combines the ideas of autoencoders and probabilistic graphical models, primarily used to learn the latent distribution of data and generate new data samples. The DFT is a deep learning model based on diffusion models, mainly used for the generation and transformation of images and videos. In this embodiment, the VAE unit, implemented based on the VAE, is used to encode the image to be processed and the target image, forming corresponding first feature codes and second feature codes corresponding to the target image, respectively. In subsequent steps, it is used to decode the feature codes output by the DFT to restore the corresponding multi-frame images, thereby obtaining the final video. In this embodiment, the diffusion transformer unit, which is based on the diffusion transformer, is used to predict based on the encoded feature code to generate the encoded features corresponding to the continuously changing images of multiple frames. In short, it generates the continuously changing images between the image frame corresponding to the image to be processed and the image frame corresponding to the target image, which is the process of the target effect continuously changing.
[0070] Specifically, in this embodiment, the terminal device first processes the image to be processed and the target image through a variational autoencoder unit, generating a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, that is, encoding the image to be processed and the target image; then, a blank code with the same size as the first feature code and the second feature code is generated, and the first feature code, the blank frame code, and the second feature code are sequentially concatenated to obtain the input feature code; finally, the input feature code is input to a diffusion transformer unit, which performs encoding prediction to generate the encoding features corresponding to one or more frames of continuously changing images between the image frames of the image to be processed and the image frames of the target image. This step performed by the diffusion transformer is equivalent to the process of generating the blank code in the input feature code through the first and second feature codes in the input feature code; finally, the feature code output by the diffusion transformer is decoded into the corresponding video frame by the variational autoencoder, thereby generating the special effects video.
[0071] Furthermore, the variational autoencoder unit includes an encoder subunit and a decoder subunit, and the specific implementation of step S203 includes:
[0072] Step S2031: Encode the image to be processed and the target image through the encoding subunit to generate the first feature code corresponding to the image to be processed and the second feature code corresponding to the target image, respectively.
[0073] Correspondingly, such as Figure 8 As shown, step S205 is implemented in the following ways:
[0074] Step S2051: Obtain the randomly initialized random feature code;
[0075] Step S2052: Input the input feature code and the random feature code into the diffusion converter unit, and use the diffusion converter unit to update the random feature code in multiple rounds based on the input feature code until the preset stopping condition is reached, and then obtain the output feature code.
[0076] Step S2053: Decode the output feature code through the decoding subunit to generate a special effects video.
[0077] For example, firstly, the terminal device obtains a feature encoding array, vector, or matrix through random initialization. This random feature encoding corresponds to N video frames, meaning it will be used to generate N frames of special effects video. Then, using the input feature encoding and random feature encoding generated in the previous steps as input, a diffusion transformer unit is invoked. The diffusion transformer unit updates the random feature encoding multiple times based on the input feature encoding until a preset stopping condition is met, such as the number of update rounds reaching a preset value, or the residual being less than a preset value. The output feature encoding of the diffusion transformer unit is then obtained; this output feature encoding is the result of multiple iterative updates to the random feature encoding. In the above process, a pre-trained diffusion transformer unit is used to perform inference with the first and second feature codes in the input feature codes as effective information, resulting in an output feature code that includes an intermediate frame code that "connects" the first and second feature codes. This intermediate frame code can represent the feature changes from the first feature code to the second feature code. The image generated after decoding the temporal frame code can represent the image changes from the image to be processed to the target image, and conforms to the real growth pattern of hairstyle (hair strands). Therefore, the special effects video generated by the output feature code can coherently and realistically display the appearance process of the target special effects in the video, and has stronger visual expressiveness.
[0078] Furthermore, in one possible implementation, the output feature encoding includes a first feature encoding, a first temporal frame encoding, and a second temporal frame encoding. The first temporal frame encoding is generated based on the first feature encoding and is used to characterize the image content of N continuously changing video frames after the first frame of the special effects video. The second temporal frame encoding is generated based on the first temporal frame encoding and the second feature encoding and is used to characterize the image content of M continuously changing video frames at the end of the special effects video, where N and M are integers greater than 0. For example... Figure 9 As shown, the specific implementation of step S2053 includes:
[0079] Step S2053-1: Decode the first feature code through the decoding subunit to generate the first frame of the video;
[0080] Step S2053-2: Decode the first time frame encoding through the decoding subunit to generate N intermediate video frames;
[0081] Step S2053-3: Decode the second time frame encoding through the decoding subunit to generate M video end frames;
[0082] Step S2053-4: Segment the first frame of the video, N intermediate frames, and M last frames to generate a special effects video.
[0083] For example, Figure 10 This is a schematic diagram illustrating the process of generating a special effects video according to an embodiment of the present disclosure. The following is in conjunction with... Figure 10The above process is described in detail, referring to the diagram. First, the encoding subunit generates input feature codes. Then, the input feature codes and random feature codes are input to the diffusion transformer unit to obtain the output feature codes. The output feature codes include feature code F1 (first feature code), feature code F2 (first temporal frame code), and feature code F3 (second temporal frame code). Next, the decoding subunit directly decodes the first feature code to generate the first frame of the video, i.e., the first frame of the special effects video, shown as P1 in the diagram. In this step, the restoration of the image to be processed is achieved. Next, the decoding subunit decodes feature code F2 to generate N intermediate video frames, where N is, for example, 4. As shown in the diagram, these four intermediate frames are P2, P3, P4, and P5. These are four intermediate frames located after the first frame of the video, representing frames 2 to 5 of the special effects video. Then, the decoding subunit decodes feature code F3 to generate M final video frames, where M is, for example, 4. As shown in the diagram, these four intermediate frames are P6, P7, P8, and P9. These are four final video frames, representing frames 6 to 9 of the special effects video. Finally, these nine video frames are combined to obtain the final special effects video.
[0084] In one possible implementation, this embodiment can further combine prompt words, such as a first prompt word and / or a second prompt word, to control the generation process of the target image and special effects video. For example, the second prompt word can be applied to the variational autoencoder unit and the diffusion transformer unit to control the generation process of the special effects video. The specific usage of the first and second prompt words described above is as follows: Figure 2 The embodiments shown have been described in detail and will not be repeated here.
[0085] In this embodiment, firstly, the trained diffusion transformer unit, based on the first feature encoding, compresses the image features of N intermediate video frames and M final video frames into a single feature encoding (first temporal frame encoding and second temporal frame encoding). Then, combined with the variational autoencoder unit trained in conjunction with the diffusion transformer unit, the first feature encoding corresponding to the first video frame, the first temporal frame encoding corresponding to the N intermediate video frames, and the second temporal frame encoding corresponding to the M final video frames are restored to the first video frame, the N intermediate video frames, and the M final video frames. This achieves the generation of special effects videos, ensuring that the beginning and end of the special effects videos accurately match the initially input video to be edited and the target video after adding special effects. This achieves content process control of the special effects videos, avoiding problems such as uncontrolled changes in the video frames, video content distortion, and significant differences between the final video frame and the expected effect (target special effects). This improves the coherence and realism of the generated special effects videos, enhancing the viewing experience.
[0086] In this embodiment, the implementation of steps 201-S202 is the same as that in this disclosure. Figure 2 The implementation methods of steps S101-S102 in the illustrated embodiment are the same, and will not be described in detail here.
[0087] Corresponding to the special effects video generation method in the above embodiments, Figure 11 This is a structural block diagram of a special effects video generation device provided in an embodiment of this disclosure. The method described in the above embodiments can be executed by this special effects video generation device, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.
[0088] For ease of explanation, only the parts relevant to embodiments of this disclosure are shown. (Refer to...) Figure 11 The special effects video generation device 3 includes:
[0089] Acquisition module 31 is used to acquire the image to be processed;
[0090] The first generation module 32 is used to generate a target image based on the image to be processed, wherein the target image is an image generated after applying a target effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect;
[0091] The second generation module 33 is used to call the video generation model to process the image to be processed and the target image to generate a special effects video. The special effects video includes a video first frame, a video last frame, and at least one video middle frame located between the video first frame and the video last frame. The video first frame has the same image content as the image to be processed, the video last frame has the same image content as the target image, and at least one video middle frame is used to characterize the continuous change process of the generated target effect.
[0092] According to one or more embodiments of this disclosure, the first generation module 32 is specifically configured to: obtain a first prompt word for the target effect of a target category, the first prompt word being used to describe the appearance features of the target effect of the target category; and process the image to be processed and the first prompt word through an image generation model to generate a target image.
[0093] According to one or more embodiments of this disclosure, the image generation model includes a first functional unit and a second functional unit. The first functional unit is used to generate a target effect for a target category, and the second functional unit is used to generate a second object object that matches the target effect based on the target object. The appearance similarity between the second object object and the target object is greater than a similarity threshold. When the first generation module 32 processes the image to be processed and the first prompt word through the image generation model to generate the target image, it is specifically used to: process the image to be processed and the first prompt word through the first functional unit to generate effect feature data representing the target effect; process the image to be processed and the effect feature data through the second functional unit to generate object feature data representing the second object object; and generate the target image based on the effect feature data and the object feature data.
[0094] According to one or more embodiments of this disclosure, the first prompt word includes a first segment, a second segment, and a third segment, wherein the first segment is used to indicate the target object; the second segment is used to characterize the shape features of the target effect; and the third segment is used to characterize the style features of the target effect.
[0095] According to one or more embodiments of this disclosure, the video generation model includes a variational autoencoder unit and a diffusion transformer unit. The second generation module 33 is specifically used for: processing the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively; sequentially concatenating the first feature code, the blank frame code, and the second feature code to obtain the input feature code; and processing the input feature code through the diffusion transformer unit to generate a special effects video.
[0096] According to one or more embodiments of this disclosure, the variational autoencoder unit includes an encoding subunit and a decoding subunit. When the second generation module 33 processes the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, it is specifically used to: encode the image to be processed and the target image through the encoding subunit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively. When the second generation module 33 processes the input feature code through the diffusion transformer unit to generate a special effects video, it is specifically used to: obtain a randomly initialized random feature code; input the input feature code and the random feature code into the diffusion transformer unit, and use the diffusion transformer unit to update the random feature code multiple times based on the input feature code until a preset stopping condition is reached to obtain an output feature code; and decode the output feature code through the decoding subunit to generate a special effects video.
[0097] According to one or more embodiments of this disclosure, the output feature encoding includes a first feature encoding, a first temporal frame encoding, and a second temporal frame encoding, wherein the first temporal frame encoding is generated based on the first feature encoding and is used to characterize the image content of N continuously changing video frames after the first frame of the special effects video; the second temporal frame encoding is generated based on the first temporal frame encoding and the second feature encoding and is used to characterize the image content of M continuously changing video frames at the end of the special effects video.
[0098] According to one or more embodiments of this disclosure, when the second generation module 33 decodes the output feature code through the decoding subunit to generate a special effects video, it is specifically used to: decode the first feature code through the decoding subunit to generate the first video frame; decode the first temporal frame code through the decoding subunit to generate N intermediate video frames; decode the second temporal frame code through the decoding subunit to generate M final video frames; and splice the first video frame, the N intermediate video frames, and the M final video frames to generate a special effects video.
[0099] According to one or more embodiments of this disclosure, when the second generation module 33 calls the video generation model to process the image to be processed and the target image to generate a special effects video, it is specifically used to: obtain a second prompt word, the second prompt word being used to describe the change process characteristics of the target effect in the special effects video; call the video generation model, process the image to be processed and the target image based on the second prompt word, and generate a special effects video with change process characteristics.
[0100] The acquisition module 31, the first generation module 32, and the second generation module 33 are connected sequentially. The special effects video generation device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0101] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 12 As shown, the electronic device 4 includes:
[0102] Processor 41, and memory 42 communicatively connected to processor 41;
[0103] Memory 42 stores instructions executed by the computer;
[0104] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-10 The special effects video generation method in the illustrated embodiment.
[0105] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0106] For relevant instructions, please refer to the corresponding text. Figures 2-10 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.
[0107] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-10 The special effects video generation method provided in any of the corresponding embodiments.
[0108] This disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements this disclosure. Figures 2-10 The special effects video generation method provided in any of the corresponding embodiments.
[0109] To implement the above embodiments, this disclosure also provides an electronic device.
[0110] refer to Figure 13 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 13 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0111] like Figure 13 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0112] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0113] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0114] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0115] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0116] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0117] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0119] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.
[0120] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] In a first aspect, according to one or more embodiments of this disclosure, a method for generating special effects videos is provided, comprising:
[0123] Obtain an image to be processed; based on the image to be processed, generate a target image, wherein the target image is an image after adding a target effect to the image to be processed, the image content of the image to be processed contains a target object, and the target effect is used to add a target effect of a target category to the target object; call a video generation model to process the image to be processed and the target image to generate an effect video, wherein the effect video includes a first frame, a last frame, and at least one intermediate frame located between the first frame and the last frame, the first frame has the same image content as the image to be processed, the last frame has the same image content as the target image, and the at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
[0124] According to one or more embodiments of this disclosure, generating a target image based on the image to be processed includes: obtaining a first prompt word for a target effect of the target category, the first prompt word being used to describe the appearance features of the target effect of the target category; and processing the image to be processed and the first prompt word through an image generation model to generate the target image.
[0125] According to one or more embodiments of this disclosure, the image generation model includes a first functional unit and a second functional unit. The first functional unit is used to generate a target effect for the target category, and the second functional unit is used to generate a second object object matching the target effect based on the target object. The appearance similarity between the second object object and the target object is greater than a similarity threshold. The step of processing the image to be processed and the first prompt word through the image generation model to generate the target image includes: processing the image to be processed and the first prompt word through the first functional unit to generate effect feature data characterizing the target effect; processing the image to be processed and the effect feature data through the second functional unit to generate object feature data characterizing the second object object; and generating the target image based on the effect feature data and the object feature data.
[0126] According to one or more embodiments of this disclosure, the first prompt word includes a first segment, a second segment, and a third segment, wherein the first segment is used to indicate the target object; the second segment is used to characterize the shape features of the target effect; and the third segment is used to characterize the style features of the target effect.
[0127] According to one or more embodiments of this disclosure, the video generation model includes a variational autoencoder unit and a diffusion transformer unit. The step of calling the video generation model to process the image to be processed and the target image to generate a special effects video includes: processing the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively; sequentially concatenating the first feature code, the blank frame code, and the second feature code to obtain an input feature code; and processing the input feature code through the diffusion transformer unit to generate the generated special effects video.
[0128] According to one or more embodiments of this disclosure, the variational autoencoder unit includes an encoding subunit and a decoding subunit. The step of processing the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively, includes: encoding the image to be processed and the target image through the encoding subunit to generate the first feature code corresponding to the image to be processed and the second feature code corresponding to the target image, respectively; and processing the input feature code through a diffusion transformer unit to generate the special effects video, including: obtaining a randomly initialized random feature code; inputting the input feature code and the random feature code into the diffusion transformer unit; using the diffusion transformer unit to update the random feature code multiple times based on the input feature code until a preset stopping condition is reached to obtain an output feature code; and decoding the output feature code through the decoding subunit to generate the special effects video.
[0129] According to one or more embodiments of this disclosure, the output feature encoding includes a first feature encoding, a first temporal frame encoding, and a second temporal frame encoding, wherein the first temporal frame encoding is generated based on the first feature encoding and is used to characterize the image content of N continuously changing video frames after the first frame of the special effects video; the second temporal frame encoding is generated based on the first temporal frame encoding and the second feature encoding and is used to characterize the image content of M continuously changing video frames at the end of the special effects video.
[0130] According to one or more embodiments of this disclosure, the step of decoding the output feature code through the decoding subunit to generate the special effects video includes: decoding the first feature code through the decoding subunit to generate a video first frame; decoding the first temporal frame code through the decoding subunit to generate N video intermediate frames; decoding the second temporal frame code through the decoding subunit to generate M video last frames; and splicing the video first frame, the N video intermediate frames, and the M video last frames to generate the special effects video.
[0131] According to one or more embodiments of this disclosure, the step of calling a video generation model to process the image to be processed and the target image to generate a special effects video includes: obtaining a second prompt word, the second prompt word being used to describe the change process characteristics of the target effect in the special effects video; calling a video generation model to process the image to be processed and the target image based on the second prompt word to generate a special effects video having the change process characteristics.
[0132] Secondly, according to one or more embodiments of this disclosure, a special effects video generation apparatus is provided, comprising:
[0133] The acquisition module is used to acquire the image to be processed;
[0134] The first generation module is used to generate a target image based on the image to be processed, wherein the target image is an image after adding target effects to the image to be processed, the image content of the image to be processed contains a target object, and the target effects are used to add target effects of target category to the target object;
[0135] The second generation module is used to call a video generation model to process the image to be processed and the target image to generate a special effects video. The special effects video includes a first frame, a last frame, and at least one intermediate frame between the first frame and the last frame. The first frame has the same image content as the image to be processed, and the last frame has the same image content as the target image. The at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
[0136] According to one or more embodiments of this disclosure, the first generation module is specifically configured to: obtain a first prompt word for the target effect of the target category, the first prompt word being used to describe the appearance features of the target effect of the target category; and process the image to be processed and the first prompt word through an image generation model to generate the target image.
[0137] According to one or more embodiments of this disclosure, the image generation model includes a first functional unit and a second functional unit. The first functional unit is used to generate a target effect for the target category, and the second functional unit is used to generate a second object object matching the target effect based on the target object. The appearance similarity between the second object object and the target object is greater than a similarity threshold. When the first generation module processes the image to be processed and the first prompt word through the image generation model to generate the target image, it is specifically used to: process the image to be processed and the first prompt word through the first functional unit to generate effect feature data characterizing the target effect; process the image to be processed and the effect feature data through the second functional unit to generate object feature data characterizing the second object object; and generate the target image based on the effect feature data and the object feature data.
[0138] According to one or more embodiments of this disclosure, the first prompt word includes a first segment, a second segment, and a third segment, wherein the first segment is used to indicate the target object; the second segment is used to characterize the shape features of the target effect; and the third segment is used to characterize the style features of the target effect.
[0139] According to one or more embodiments of this disclosure, the video generation model includes a variational autoencoder unit and a diffusion transformer unit. The second generation module is specifically used for: processing the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively; sequentially concatenating the first feature code, the blank frame code, and the second feature code to obtain an input feature code; and processing the input feature code through the diffusion transformer unit to generate the generated special effects video.
[0140] According to one or more embodiments of this disclosure, the variational autoencoder unit includes an encoding subunit and a decoding subunit. When the second generation module processes the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, it is specifically configured to: encode the image to be processed and the target image through the encoding subunit to generate the first feature code corresponding to the image to be processed and the second feature code corresponding to the target image, respectively. When the second generation module processes the input feature code through the diffusion transformer unit to generate the special effects video, it is specifically configured to: obtain a randomly initialized random feature code; input the input feature code and the random feature code into the diffusion transformer unit, and use the diffusion transformer unit to update the random feature code multiple times based on the input feature code until a preset stopping condition is reached to obtain an output feature code; decode the output feature code through the decoding subunit to generate the special effects video.
[0141] According to one or more embodiments of this disclosure, the output feature encoding includes a first feature encoding, a first temporal frame encoding, and a second temporal frame encoding, wherein the first temporal frame encoding is generated based on the first feature encoding and is used to characterize the image content of N continuously changing video frames after the first frame of the special effects video; the second temporal frame encoding is generated based on the first temporal frame encoding and the second feature encoding and is used to characterize the image content of M continuously changing video frames at the end of the special effects video.
[0142] According to one or more embodiments of this disclosure, when the second generation module decodes the output feature code through the decoding subunit to generate the special effects video, it is specifically configured to: decode the first feature code through the decoding subunit to generate the first video frame; decode the first temporal frame code through the decoding subunit to generate N intermediate video frames; decode the second temporal frame code through the decoding subunit to generate M final video frames; and splice the first video frame, the N intermediate video frames, and the M final video frames to generate the special effects video.
[0143] According to one or more embodiments of this disclosure, when the second generation module calls a video generation model to process the image to be processed and the target image to generate a special effects video, it is specifically used to: obtain a second prompt word, the second prompt word being used to describe the change process characteristics of the target effect in the special effects video; call the video generation model, process the image to be processed and the target image based on the second prompt word, and generate a special effects video having the change process characteristics.
[0144] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0145] The memory stores computer-executed instructions;
[0146] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the special effects video generation method as described in the first aspect and various possible designs of the first aspect.
[0147] Fourthly, according to one or more embodiments of this disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the special effects video generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0148] Fifthly, according to one or more embodiments of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the special effects video generation method as described in the first aspect and various possible designs of the first aspect.
[0149] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0150] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0151] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for generating special effects videos, characterized in that, include: Obtain an image to be processed, wherein the image to be processed contains a target object; Based on the image to be processed, a target image is generated, wherein the target image is an image generated after applying a target effect to the target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect; The video generation model is invoked to process the image to be processed and the target image to generate a special effects video. The special effects video includes a first frame, a last frame, and at least one intermediate frame located between the first frame and the last frame. The first frame has the same image content as the image to be processed, and the last frame has the same image content as the target image. The at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
2. The method according to claim 1, characterized in that, The step of generating a target image based on the image to be processed includes: Obtain a first cue word for the target effect of the target category, wherein the first cue word is used to describe the appearance characteristics of the target effect of the target category; The target image is generated by processing the image to be processed and the first prompt word using an image generation model.
3. The method according to claim 2, characterized in that, The image generation model includes a first functional unit and a second functional unit. The first functional unit is used to generate a target effect for the target category. The second functional unit is used to generate a second object object that matches the target effect based on the target object. The appearance similarity between the second object object and the target object is greater than a similarity threshold. The step of processing the image to be processed and the first prompt word through an image generation model to generate the target image includes: The first functional unit processes the image to be processed and the first prompt word to generate effect feature data characterizing the target effect. The second functional unit processes the image to be processed and the effect feature data to generate object feature data characterizing the second object. The target image is generated based on the effect feature data and the object feature data.
4. The method according to claim 2, characterized in that, The first prompt word includes a first segment, a second segment, and a third segment, wherein the first segment is used to indicate the target object; the second segment is used to characterize the shape features of the target effect; and the third segment is used to characterize the style features of the target effect.
5. The method according to claim 1, characterized in that, The video generation model includes a variational autoencoder unit and a diffusion transformer unit. The process of calling the video generation model to process the image to be processed and the target image to generate a special effects video includes: The image to be processed and the target image are processed by a variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively. The first feature code, the blank frame code, and the second feature code are concatenated sequentially to obtain the input feature code; The input feature encoding is processed by a diffusion transformer unit to generate the generated special effects video.
6. The method according to claim 5, characterized in that, The variational autoencoder unit includes an encoding subunit and a decoding subunit. The process of processing the image to be processed and the target image through the variational autoencoder unit to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image includes: The encoding subunit encodes the image to be processed and the target image to generate a first feature code corresponding to the image to be processed and a second feature code corresponding to the target image, respectively. The process of processing the input feature encoding through a diffusion transform unit to generate the generated special effects video includes: Obtain the randomly initialized random feature code; The input feature code and the random feature code are input into the diffusion converter unit. The diffusion converter unit updates the random feature code in multiple rounds based on the input feature code until a preset stopping condition is reached, and then the output feature code is obtained. The output feature encoding is decoded by the decoding subunit to generate the special effects video.
7. The method according to claim 6, characterized in that, The output feature encoding includes a first feature encoding, a first temporal frame encoding, and a second temporal frame encoding, wherein... The first temporal frame encoding is generated based on the first feature encoding and is used to characterize the image content of N video frames that change continuously after the first video frame of the special effects video. The second temporal frame encoding is generated based on the first temporal frame encoding and the second feature encoding, and is used to characterize the image content of the M continuously changing video frames at the end of the special effects video.
8. The method according to claim 7, characterized in that, The step of decoding the output feature code through the decoding subunit to generate the special effects video includes: The first feature code is decoded by the decoding subunit to generate the first frame of the video; The decoding subunit decodes the first time-series frame encoding to generate N intermediate video frames. The second time-series frame encoding is decoded by the decoding subunit to generate M video end frames; The special effects video is generated by splicing the first frame, N middle frames, and M last frames of the video.
9. The method according to claim 1, characterized in that, The process of calling the video generation model to process the image to be processed and the target image to generate a special effects video includes: Obtain a second prompt word, which is used to describe the characteristics of the change process of the target effect in the special effects video; The video generation model is invoked to process the image to be processed and the target image based on the second prompt word, thereby generating a special effects video with the characteristics of the change process.
10. A special effects video generation device, characterized in that, include: The acquisition module is used to acquire the image to be processed; The first generation module is used to generate a target image based on the image to be processed, wherein the target image is an image generated after applying a target effect to a target object in the image to be processed, and the target object in the target image has the target effect produced by the target effect; The second generation module is used to call a video generation model to process the image to be processed and the target image to generate a special effects video. The special effects video includes a first frame, a last frame, and at least one intermediate frame between the first frame and the last frame. The first frame has the same image content as the image to be processed, and the last frame has the same image content as the target image. The at least one intermediate frame is used to characterize the continuous change process of generating the target effect.
11. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the special effects video generation method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the special effects video generation method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the special effects video generation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Special effect graph generation method and device, equipment and storage medium
CN115063335A
Style image generation method and device, equipment and medium
CN116934577A