Animation generation method, device, electronic device, storage medium and program product
Through automated rendering and image fusion processing technology, pre-made animation material frames and customized rendering models are used to solve the problems of slow production speed and poor picture quality in traditional animation, and efficient and good-quality animation generation is achieved.
Patent Information
- Application Number
- CN202411159695.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-22
AI Technical Summary
Traditional animation production methods require a lot of manual intervention, resulting in low animation production speed and poor picture quality after rendering.
By obtaining pre-made animation material frames, combining character rendering models for different character types and preset background rendering models, the animation material is automatically rendered and image fusion processing to generate target animations.
It improves the efficiency and picture quality of animation production, reduces manual intervention, and ensures the consistency of the visual performance of animation characters and the authenticity of the background environment.
Smart Images

Figure CN118674839B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an animation generation method, device, electronic device, storage medium and program product. Background Art
[0002] Animation production mainly includes two important stages: one is the visual exploration stage in the early stage of animation production, and the other is the animation rendering stage.
[0003] In the related art, the rendering stage of animation is the process of rendering a three-dimensional model into a two-dimensional image with rich colors and light and shadow details through three-dimensional modeling software.
[0004] However, the method of rendering animations through 3D modeling software requires a lot of manual intervention to adjust details such as light and shadow to achieve the best visual effect. Therefore, due to limitations in human resources and technical means, the animation production speed is low and the quality of the rendered images is poor. Summary of the invention
[0005] The embodiments of the present application provide an animation generation method, device, electronic device, storage medium and program product, which can improve the production efficiency and picture quality of animation.
[0006] The technical solution of the embodiment of the present application is implemented as follows:
[0007] An embodiment of the present application provides an animation generation method, which includes: obtaining a pre-made animation material frame; the animation material in the animation material frame includes an animation character and an animation background; based on the character type of the animation character, determining a character rendering model corresponding to the animation character; calling the character rendering model to render the animation character in the animation material frame to obtain a character image; calling a preset background rendering model to render the animation background in the animation material frame to obtain a background image; performing image fusion processing on the character image and the background image to obtain a target animation containing the animation material.
[0008] An embodiment of the present application provides an animation generation device, including: a material frame acquisition module, used to acquire a pre-made animation material frame; the animation material in the animation material frame includes an animation character and an animation background; a model determination module, used to determine a character rendering model corresponding to the animation character based on the character type of the animation character; a character rendering module, used to call the character rendering model to render the animation character in the animation material frame to obtain a character image; a background rendering module, used to call a preset background rendering model to render the animation background in the animation material frame to obtain a background image; an animation generation module, used to perform image fusion processing on the character image and the background image to obtain a target animation containing the animation material.
[0009] In the above scheme, the material frame acquisition module is also used to respond to the received animation production request, obtain the animation style type and animation script of the target animation to be generated; for the character type and scene type in the animation script, call the painting model to generate a scene concept map under the scene type based on the animation style type; generate three views of the animation character of the character type in different postures based on the animation style type and the preset skeleton map; generate the animation material frame based on the three views of the animation character in different postures and the scene concept map under the scene type.
[0010] In the above scheme, the material frame acquisition module is also used to create a three-dimensional model of the animation character and the scene type based on the three-view images of the animation character in different postures and the scene concept map under the scene type; generate a three-dimensional animation of each animation scene in the animation script based on the animation script and the three-dimensional model; the three-dimensional animation includes the animation character and the animation background under the scene type; and extract the two-dimensional animation material frame of the corresponding animation scene from the three-dimensional animation of each animation scene.
[0011] In the above scheme, the animation generation device also includes a background rendering model training module, which is used to call multiple painting models to generate pictures of multiple picture style types respectively; wherein each painting model is used to generate multiple pictures of one picture style type; in response to the selection operation of the animation style type among the multiple picture style types, the preset diffusion model is trained using the pictures of the animation style type to obtain the preset background rendering model with the animation style type.
[0012] In the above scheme, the background rendering model training module is also used to use the picture of the animation style type to train the preset diffusion model to obtain an initial rendering model with the animation style type; determine the edge image of the picture of the animation style type; use the edge image to train the preset control model to obtain an edge control model; merge the initial rendering model and the edge control model to obtain the preset background rendering model.
[0013] In the above scheme, the animation generation device also includes a character rendering model training module, which is used to perform close-up processing of different parts of the animation character based on the three views of the animation character in different postures, and obtain multiple training material images of the animation character; use the multiple training material images of the animation character to train the preset character rendering model to be trained, and obtain the character rendering model corresponding to the animation character.
[0014] In the above scheme, the model determination module is also used to segment the animation material frame based on the animation character in the animation material frame to obtain multiple material sub-frames; wherein each material sub-frame includes an animation character; for each material sub-frame, based on the character type of the animation character in the material sub-frame, determine the character rendering model corresponding to the animation character.
[0015] In the above scheme, the pre-made animation material frame includes multiple continuous animation material frames; the character rendering module is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; splice the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; call the character rendering model to render multiple animation characters in the spliced animation material frame to obtain the character image.
[0016] In the above scheme, the character rendering module is also used to call the character rendering model to render multiple animation characters in the spliced animation material frame to obtain an initial character image; perform image segmentation on the initial character image to obtain a first character image corresponding to the previous frame of animation material frame, a second character image corresponding to the current animation material frame, and a third character image corresponding to the next frame of animation material frame; and determine the second character image as the character image.
[0017] In the above scheme, the pre-made animation material frame includes multiple continuous animation material frames; the background rendering module is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; splice the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; call a preset background rendering model to render multiple animation backgrounds in the spliced animation material frame to obtain the background image.
[0018] In the above scheme, the background rendering module is also used to call a preset background rendering model to render multiple animation backgrounds in the spliced animation material frame to obtain an initial background image; perform image segmentation on the initial background image to obtain a first background image corresponding to the previous frame of animation material frame, a second background image corresponding to the current animation material frame, and a third background image corresponding to the next frame of animation material frame; and determine the second background image as the background image.
[0019] In the above scheme, the animation generation module is also used to determine the animation image corresponding to the animation material frame based on the character image and background image corresponding to the animation material frame; determine the edge information and optical flow information of each animation material frame from the i-th frame animation material frame to the i+j-th frame animation material frame in the pre-made multi-frame animation material frame; i and j are both integers greater than 0; based on the edge information and optical flow information of each animation material frame from the i-th frame animation material frame to the i+j-th frame animation material frame, as well as the i-th frame animation image corresponding to the i-th frame animation material frame and the i+j-th frame animation image corresponding to the i+j-th frame animation material frame, determine the i+k-th frame animation image; the value of k is any integer between 1 and j-1; splice the 1st frame animation image to the nth frame animation image into the target animation, n is the number of the multi-frame animation material frames, and n is an integer greater than 1.
[0020] In the above scheme, the animation generation module is also used to extract background features of the background image to obtain image feature information of the background image; extract character features of the character image to obtain image feature information of the character image; generate a background layer of the background image based on the image feature information of the background image; generate a character layer of the character image based on the image feature information of the character image; and superimpose the character layer and the background layer to obtain an animation image corresponding to the animation material frame.
[0021] In the above scheme, the animation material in the animation material frame also includes animation special effects; the animation generation module is also used to determine the special effects rendering model corresponding to the animation special effects based on the special effects type of the animation special effects; call the special effects rendering model to render the animation special effects in the animation material frame to obtain a special effects image; perform image fusion processing on the special effects image, the character image and the background image to obtain a target animation containing the animation material.
[0022] In the above scheme, the pre-produced animation material frame includes multiple continuous animation material frames; the background rendering module is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; encode the previous animation material frame, the animation material frame and the next animation material frame respectively to obtain a first material frame vector corresponding to the previous animation material frame, a second material frame vector corresponding to the current animation material frame and a third material frame vector corresponding to the next animation material frame; fuse the first material frame vector, the second material frame vector and the third material frame vector to obtain a fused material frame vector; call the background rendering model, render the animation background in the animation material frame based on the fused material frame vector, and obtain the background image.
[0023] In the above scheme, the background rendering module is also used to call the background rendering model, perform attention processing on the fused material frame vector, and obtain attention features; based on the fused material frame vector and the attention features, determine the background image of the animation background.
[0024] An embodiment of the present application provides an electronic device, comprising: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the animation generation method provided in the embodiment of the present application.
[0025] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the animation generation method provided in the embodiment of the present application when executed by a processor.
[0026] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the animation generation method provided in the embodiment of the present application is implemented.
[0027] The embodiments of the present application have the following beneficial effects:
[0028] In the process of making animations, by using pre-made animation material frames and combining them with character rendering models customized for different character types, it is possible to ensure that the visual performance of each animated character meets its unique design style and artistic requirements, thereby enhancing the expressiveness and recognizability of the animated character. The animation background is rendered using a preset background rendering model. After obtaining the background image, the character image and the background image are subjected to image fusion processing to obtain the target animation containing the animation material. This not only makes the background environment more realistic and vivid, and enhances the immersion and viewing of the overall picture, but also allows the character rendering model and the background rendering model to perform fast, batch, and semi-automatic animation rendering, thereby improving the efficiency of animation rendering. Therefore, the embodiments of the present application can improve the production efficiency and picture quality of animations. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of the effect of "rendering three to two" in the related technology;
[0030] Figure 2 is a schematic diagram of a user interface provided in an embodiment of the present application;
[0031] Figure 3 is a schematic diagram of the architecture of the animation generation system provided in an embodiment of the present application;
[0032] Figure 4is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0033] Figure 5 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 1 ;
[0034] Figure 6 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 2 ;
[0035] Figure 7 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 3 ;
[0036] Figure 8 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 4 ;
[0037] Fig. 9 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 5 ;
[0038] Fig.10 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 6 ;
[0039] Fig.11 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 7 ;
[0040] Fig.12 This is a schematic diagram of the process of the animation generation method provided in the embodiment of the present application. Figure 8 ;
[0041] Fig.13 is another optional flowchart of the animation generation method provided in the embodiment of the present application;
[0042] Fig.14 is a schematic diagram of the AI-assisted animation process provided in an embodiment of the present application;
[0043] Fig.15 It is a multi-angle different view of a single character provided by the embodiment of the present application;
[0044] Fig.16 This is a comparison diagram of the picture effects before and after rendering provided by the embodiment of the present application;
[0045] Fig.17 is a simplified system flow chart of the animation generation method provided in an embodiment of the present application;
[0046] Fig.18 is a flowchart of an animation generation method provided in an embodiment of the present application;
[0047] Fig.19 is a schematic diagram of multi-frame rendering provided by an embodiment of the present application;
[0048] Fig. 20 is a schematic diagram of the AI frame interpolation technology provided in an embodiment of the present application;
[0049] Fig.21 It is a technical flow chart of AI animation rendering provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0051] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0052] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0053] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0054] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0055] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0056] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0057] 1) Artificial Intelligence Generated Content (AIGC): refers to content generated by artificial intelligence technology, which can include text, images, audio, video and other forms. AIGC technology uses machine learning and deep learning algorithms to automatically generate high-quality content, reducing the time and cost of manual creation. The working principle of AIGC mainly relies on the following key technologies and steps: data collection and preprocessing, model selection and training, generation and optimization, and evaluation and iteration. Data collection and preprocessing: Collect a large amount of training data, which includes text, images, audio, video and other forms, and clean, annotate and format the collected training data to facilitate model training. Model selection and training: Select a suitable machine learning or deep learning model, use the preprocessed training data to train the model, and the model training process is to continuously adjust the model parameters so that the model can generate high-quality content. Commonly used models include Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Transformer models (such as GPT-3), etc. Generation and optimization: A trained model can generate content based on input conditions or random noise. For example, a text generation model can generate articles based on input keywords, and an image generation model can generate images based on input labels. The generated content may need to be further optimized and adjusted to improve quality and meet specific needs. This step can be achieved through post-processing technology or further model fine-tuning. Evaluation and iteration: Evaluate the generated content to ensure its quality and meet expectations. Evaluation can be performed through automated indicators (such as automatic indicators for machine translation quality evaluation (Bilingual Evaluation Understudy, BLEU), indicators for evaluating the quality of generated images (Frechet Inception Distance, FID)) or manual review. Based on the evaluation results, iterate and optimize the model to continuously improve the quality of generated content.
[0058] 2) Stable Diffusion: It is a generative technique based on diffusion models for generating high-quality images. Diffusion models are deep learning models based on probabilistic processes, which are used to generate the final output image from random noise through a step-by-step denoising process based on given initial conditions (such as a noisy image). Stable Diffusion models are able to generate clear and detailed images through a stable diffusion process. Stable Diffusion models usually include the following key parts: Forward Diffusion Process, Reverse Diffusion Process, Noise Prediction Network, and Sampling Process. Forward Diffusion Process: During the training process, this process gradually adds noise to the image to make it pure noise. This is a supervised process because there is a real image as a reference at each step. This process usually includes multiple steps, and each step adds more noise to the image until the image is completely noise. Backward diffusion process: When generating images, this process is the opposite of the forward process, gradually converting pure noise back to images. The backward diffusion process usually requires more steps because each step reduces noise and restores image details. This process is learned by the model during training, so it is able to generate images based on given noise. Noise prediction network: This network is the core of the stable diffusion model. It is responsible for learning how to convert between a given image and noise. During training, this network learns how to predict noise from images so that it can operate in reverse when generating images. Sampling process: When generating images, the sampling process is responsible for sampling noise from the noise distribution and providing it to the backward diffusion process to generate images. The sampling process can be simple random sampling or more complex sampling techniques such as path sampling or fractional guided sampling to generate higher quality images.
[0059] Specifically, the generation process of the stable diffusion model is as follows: Initial noise: First, a noise sample is sampled from the standard normal distribution (or a predefined noise distribution) as the starting point of the generation process. Predict noise: At each step, the noise prediction network receives the current noise sample and the current time step, and predicts the noise distribution at the next moment. Sampling noise: Based on the output of the noise prediction network, a new noise sample is sampled from the predicted noise distribution. Backward diffusion: Using the new noise sample and time step, the noise is gradually converted back to the image through the back diffusion process. Iterative process: Repeat the above steps until the desired image is generated.
[0060] 3) Animation "3D Rendering 2" technology: "3D Rendering 2" is also called cartoon rendering (Cel shading or Toon shading), which is a special rendering style of non-realistic rendering. Figure 1 This is a schematic diagram of the effect of "rendering three to two" in related technologies, such as Figure 1 As shown, the "3D Rendering 2" technology analyzes the plane outline on the basic appearance of the three-dimensional object, so that the object has a three-dimensional perspective while also presenting a two-dimensional effect. Generally speaking, "3D Rendering 2" is modeling through 3D technology, and then rendering the 3D model into a 2D (2D) color block effect. Compared with traditional full 2D hand-painted, "3D Rendering 2" can save the production burden of characters or objects that appear repeatedly, reduce costs and production cycles, and also have a free camera movement method. Commonly used "3D Rendering 2" production software includes Maya, Unreal Engine (Unity Engine, UE) and 3D Max.
[0061] 4) Dreambooth: Dreambooth is a personalized content generation technology based on deep learning. It generates models through training so that they can generate content that meets specific needs or specific styles. The core of Dreambooth is to use a large amount of training data and complex neural network structures to capture and generate specific visual or text features. Its working principle mainly includes the following four parts: Data collection: Collect a large amount of image training data related to the target style or theme. Model training: Use this data to train the Stable Diffusion image generation model. Personalized adjustment: After the initial training is completed, personalized adjustments can be made to make the model more suitable for a specific style or theme. This step usually requires a small amount of additional data and fine-tuning process. Content generation: In the end, the trained and adjusted model can generate content that meets specific needs. For example, if you want to combine C style and D style, you can obtain multiple C style pictures and multiple D style pictures as training sets to train Dreambooth.
[0062] 5) "Low-Rank Adaptation" model (Lora): Lora is a method for efficiently fine-tuning large pre-trained models. It introduces low-rank matrices to adapt to specific tasks, thereby reducing the number of parameters that need to be updated during fine-tuning. This method is particularly suitable for resource-constrained environments because it significantly reduces computing and storage costs. Its principle is mainly to decompose the weight matrix of the pre-trained model into two commensurate low-rank matrices. During the fine-tuning process, the weight matrix 𝑊 of the pre-trained model remains unchanged, and only the low-rank matrices 𝐴 and 𝐵 are updated, which greatly reduces the number of parameters for model training.
[0063] 6) Web User Interface (WebUI): It is a web interface of the stable diffusion model, written using an open source library (Gradio), in which users can generate artificial intelligence (AI) images and videos, and adjust parameters to make corresponding changes to the generated results. Figure 2 Schematic diagram of the user interface provided by the embodiment of the present application. Figure 2 As shown, taking text conversion to image as an example, the user can input text in the text input box 101 of the user interface, set the width and height of the generated image in the parameter setting window 105, select the image style in the style setting bar 103, click the generate button 102, and convert the text into the corresponding image, which can be displayed in the image display window 104.
[0064] 7) Edge information: refers to the location information of the boundary of an object in an image. In image processing, an edge refers to an area in an image where the grayscale value changes dramatically, usually indicating the outline of an object or the boundary between different areas. Edge detection is an important task in computer vision, used to identify the outline and structure of objects in an image. Edge information is very useful for subsequent image processing and analysis.
[0065] 8) Optical flow information refers to the direction and speed of movement of objects or pixels in an image between consecutive frames. Optical flow calculation is another important task in computer vision, which is used to analyze the movement of objects in video sequences. Optical flow information is very important for understanding and reconstructing the movement between consecutive frames.
[0066] 9) Controlnet: It is a deep learning network structure that controls the output of the generative model so that it can generate a specific type of image according to the input control signal. Controlnet is mainly used for image generation tasks, such as image-to-image conversion, image restoration, image synthesis, etc. The core idea of the controlnet is to guide the learning process of the generator by introducing an additional control branch, so that the generator can generate more accurate and controllable images according to the control signal.
[0067] In the related art, the animation production process is mainly based on manual operation. On the one hand, in the early stage of animation production, the animators and animation directors conduct preliminary visual exploration. According to the theme, style, and director's expectations of the picture effect of the animation, the animators draw reference pictures and manuscripts, and discuss and negotiate with the director. During the negotiation process, it is necessary to draw reference pictures many times to evaluate different picture effects and finally determine the visual effect of the animation. For the characters in the animation, it is also necessary to draw manuscripts and sketches many times, and conduct detailed discussions on details such as facial features, clothing, and accessories, including drawing three views of the characters and pictures of accessories, etc., to finally determine the image of the animation characters. On the other hand, in the process of "rendering three to two" in animation, a lot of manual operation is required. The animation rendering team needs to use professional painting software, such as Maya, 3D Max, etc., to render the initial version of the 3D picture into a 2D animation with rich colors and light and shadow details through the painting software, so as to obtain the best picture effect. Each animation clip of about one minute takes several hours to render completely. At the same time, some of the light and shadow details need to be manually adjusted and optimized by professional animation rendering personnel. The whole process requires a lot of manpower costs. And the final rendering effect will be affected by the manual operation process. The production efficiency and speed are very slow. Therefore, the relevant technology has the following technical problems: 1. The traditional "three-render-two" process requires a lot of manpower support. Each animation clip needs to be rendered by professional animation production software. The repetitive workload is large and highly dependent on manual operation. Therefore, the generated animation effect will also be affected by the subjective operation of the animation rendering personnel, and the final effect is not uniform. In the traditional production process, it is often necessary to find a special team of more than ten people and work continuously for several months to complete the rendering process of an animation, which is time-consuming, labor-intensive and inefficient. 2. The visual exploration process in the early stage of animation production is complicated, and it is necessary to repeatedly draw the three views of the relevant characters for comparison. At the same time, the drawing of some picture styles is difficult to complete in a short time. It requires a high level of painting skills for the painter, and it is necessary to accurately depict the effects that the director and producer want. It is greatly affected by the subjective influence of the painter, so the efficiency of animation generation is reduced.
[0068] Based on the problems existing in the related technologies, the embodiments of the present application provide an animation generation method, device, electronic device, storage medium and program product, which can improve the production efficiency and picture quality of animation.
[0069] In the animation generation method provided in the embodiment of the present application, first, a pre-made animation material frame is obtained; the animation material in the animation material frame includes an animation character and an animation background; then, based on the character type of the animation character, a character rendering model corresponding to the animation character is determined; then, the character rendering model is called to render the animation character in the animation material frame to obtain a character image; and a preset background rendering model is called to render the animation background in the animation material frame to obtain a background image; finally, image fusion processing is performed on the character image and the background image to obtain a target animation containing the animation material.
[0070] The following describes an exemplary application of the animation generation device provided in the embodiment of the present application. The animation generation device provided in the embodiment of the present application is an electronic device for implementing the animation generation method. The electronic device provided in the embodiment of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals, and can also be implemented as servers. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application. Below, the exemplary application of the animation generation device when it is implemented as a server or a terminal will be described.
[0071] See also Figure 3 , Figure 3 It is a schematic diagram of the architecture of the animation generation system provided in the embodiment of the present application. To realize animation generation, an animation generation application can be provided. For example, the animation generation application can be an application dedicated to animation generation, or it can be a functional module in other applications (such as an animation generation module in a game application, etc.). The animation generation system 100 provided in the embodiment of the present application at least includes a terminal 400, a network 300 and a server 200, wherein the server 200 is a server of the animation generation application. The server 200 can constitute the animation generation device of the embodiment of the present application, that is, the animation generation method of the embodiment of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0072] See also Figure 3, the user can perform interactive operations on the client of the animation generation application through the terminal 400, and the interactive operations can be, for example, inputting pre-made animation material frames, clicking to start generating animations, etc. After receiving the interactive operations of the user, the client can encapsulate the pre-made animation material frames into the animation generation request through the terminal, and send the animation generation request to the server 200 through the network 300. After receiving the animation generation request, the server 200 obtains the pre-made animation material frames in response to the animation generation request; the animation material in the animation material frame includes an animation character and an animation background; the server 200 determines the character rendering model corresponding to the animation character based on the character type of the animation character; the server 200 calls the character rendering model to render the animation character in the animation material frame to obtain a character image; the server 200 calls the preset background rendering model to render the animation background in the animation material frame to obtain a background image; the server 200 performs image fusion processing on the character image and the background image to obtain a target animation containing the animation material. After generating the target animation, the server 200 can also send the target animation to the terminal 400 to show the target animation to the user.
[0073] In some embodiments, the terminal 400 itself can also perform the animation generation method of the embodiment of the present application, that is, after the terminal 400 receives the interactive operation input by the user through the client, the terminal 400 obtains the pre-made animation material frame; the terminal 400 determines the character rendering model corresponding to the animation character based on the character type of the animation character; the terminal 400 calls the character rendering model to render the animation character in the animation material frame to obtain the character image; the terminal 400 calls the preset background rendering model to render the animation background in the animation material frame to obtain the background image; the terminal 400 performs image fusion processing on the character image and the background image to obtain the target animation containing the animation material. After determining the target animation, the target animation is displayed on the client interface of the terminal 400.
[0074] See also Figure 4 , Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, Figure 4 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 4 Various buses are labeled as bus system 440 .
[0075] The processor 410 may be an integrated circuit chip having signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0076] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0077] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, and the like. The memory 450 may optionally include one or more storage devices physically located away from the processor 410. The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory. In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures, or subsets or supersets thereof, as exemplified below.
[0078] An operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic businesses and processing hardware-based tasks; a network communication module 452, for reaching other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.; a presentation module 453, for enabling information to be presented (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., display screens, speakers, etc.) associated with a user interface 430; an input processing module 454, for detecting one or more user inputs or interactions from one of the one or more input devices 432 and translating the detected inputs or interactions.
[0079] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 4 An animation generating device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a material frame acquisition module 4551, a model determination module 4552, a character rendering module 4553, a background rendering module 4554 and an animation generating module 4555. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0080] In other embodiments, the animation generating device provided in the embodiments of the present application can be implemented in hardware. As an example, the animation generating device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the animation generating method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (Application Specific Integrated Circuit, ASIC), digital signal processors (Digital Signal Processor, DSP), programmable logic devices (Programmable Logic Device, PLD), complex programmable logic devices (Complex Programmable Logic Device, CPLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other electronic components.
[0081] It should be noted that, based on the understanding of the following, those skilled in the art can apply the animation generation method provided in the embodiments of the present application to any scene that uses artificial intelligence technology to generate animation, such as movie and television animation production scenes, game development scenes, live interactive scenes, education and training scenes, cultural protection and reconstruction scenes, medical health visualization scenes, experimental simulations, etc. The following is an example of using artificial intelligence technology to produce a short animation scene.
[0082] Figure 5 This is a flow chart of the animation generation method provided by the embodiment of the present application. Figure 5 The steps shown are described in detail. As mentioned above, the electronic device that implements the animation generation method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution subject of each step will not be repeatedly described below. Figure 5 As shown, the animation generation method is described by taking the execution subject of the server as an example, and the method includes the following steps S101 to S105:
[0083] Step S101, obtaining pre-made animation material frames.
[0084] The animation material in the animation material frame includes an animation character and an animation background.
[0085] Here, the pre-made animation material frame refers to a two-dimensional static image prepared in advance in the animation production process. The pre-made animation material frame may include at least one frame of animation material frame. The animation material frame usually represents the decomposition picture of each key moment or specific action in the animation. The animation material frame can be drawn by hand or generated by computer graphics technology. Exemplarily, the user can export the pre-made animation material frame through three-dimensional modeling software. Animation material is a variety of visual and audio elements used in the animation production process. Animation material can include but is not limited to animation characters and animation backgrounds. Animation characters refer to characters that appear in the animation, which can be people, animals or other creatures, or even non-biological entities such as robots or magic items. Animation characters usually have unique personalities, appearances and actions, and have the function of promoting the plot in the animation. Animation background is the environment setting in the animation, including but not limited to places, buildings, natural landscapes, weather conditions, etc. Animation background is used to provide an environment for activities and interactions for animation characters, making the animation more realistic and three-dimensional.
[0086] In some embodiments, see Figure 6 , Figure 6 It is shown that obtaining the pre-made animation material frame in step S101 can be achieved by the following steps S1011 to S1014:
[0087] Step S1011, in response to the received animation production request, obtaining the animation style type and animation script of the target animation to be generated.
[0088] Here, the user can input the animation script, start the animation production, and other operations on the client of the animation generation application to generate an animation production request. The animation production request may include the animation style type and animation script of the target animation to be generated. The animation style type is the overall visual style of the target animation, such as hand-painted style, cartoon style, pixel art style, or even the style of a specific painter. The animation script is the basis of the animation production process and is a textual expression of the target animation, including a detailed description of the animation characters, dialogues, actions, scenes, and plots. Exemplarily, the animation script may include the following components: scene description: background setting of each scene, including location, time, environmental atmosphere, etc.; character action: body movements and expressions of the animation characters, as well as interaction with the environment, etc.; dialogue: dialogue between animation characters, including instructions for tone and speed; internal notes: notes of the director or producer, which may include lens angles, camera movements, special effects prompts, etc.; sound effects and music: selection of background music, and specific sound effect descriptions, etc. In response to the received animation production request, the animation style type and animation script of the target animation to be generated can be obtained from the animation production request.
[0089] Step S1012, for the character type and scene type in the animation script, call the drawing model to generate a scene concept map under the scene type based on the animation style type.
[0090] Here, the character type is an identifier used to distinguish the animation characters, and each animation character in the animation script corresponds to a character type. For example, the animation character "Little Fox" corresponds to the character type "1", the animation character "Little Boy" corresponds to the character type "2", etc. The scene type is the different locations or environmental settings that appear in the animation script, such as natural environments: forests, mountains, beaches, etc., urban landscapes: bustling cities, old neighborhoods, etc., building structures: palaces, castles, schools, etc., special places: laboratories, other worlds, etc. The painting model is a pre-trained model that uses AIGC technology to generate images. It can be a machine learning or deep learning model, such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Transformer models (such as GPT-3), etc. The scene concept map is a sketch or illustration generated by visualizing the scene described in the animation script, which aims to show the basic layout, atmosphere, color matching, lighting effects and other elements of the scene.
[0091] In the embodiment of the present application, for the scene type in the animation script, the text description of the scene type can be obtained from the animation script, and the painting model is called based on the animation style type and the text description of the scene type to generate a scene concept map of the scene type. Exemplarily, for the scene type "park" in the animation script, assuming that the animation style type is ink painting style, the text description of the "park" can be obtained from the animation script, and the text description of the "park" and the ink painting style are used as parameters to input into the pre-trained painting model, and the painting model generates a "park" scene concept map in ink painting style.
[0092] Step S1013, generating three views of the animation character of the character type in different postures based on the animation style type and the preset skeleton diagram.
[0093] Here, the preset skeleton diagram is a framework composed of joints, which defines the joint position and range of motion of the animated character, and is used to control the movement of the body parts of the animated character in different postures. The skeleton diagram usually includes the joint connection points of the main parts such as the head, torso, and limbs. The three views refer to the front view, side view, and top view of the animated character. For the character type in the animation script, the text description of the character type can be obtained from the animation script, and the drawing model is called to design the specific appearance of the animated character based on the animation style type and the text description of the character type, including clothing, hairstyle, facial features, etc. The designed character appearance is bound to the skeleton diagram to ensure that the various parts of the animated character can change with the movement of the skeleton. The drawing model creates a series of different postures for the animated character according to the action design in the animation script, including basic actions such as standing, walking, and jumping. For each posture, the front view, side view, and top view of the animated character are generated to fully display the appearance and action details of the animated character in that posture.
[0094] Step S1014, generating animation material frames based on the three views of the animation character in different postures and the scene concept map under the scene type.
[0095] Here, the animation scenes can be determined based on the story script and scene concept map in the animation script, and each animation scene includes multiple storyboards. Storyboard refers to the conversion of the animation story script into a series of continuous pictures, each picture represents a shot or scene. Animation material frame refers to the basic unit that constitutes the animation, and each frame is a static picture. In animation, the continuous playback of these animation material frames will produce a dynamic effect, so that the audience can see continuous action. The number of frames played in one second is called the frame rate, and common frame rates include 24fps (24 frames per second), 30fps, etc. Each shot will be decomposed into multiple continuous animation material frames. For example, if there is an action of an animation character from standing to running in a shot, then this action will be decomposed into multiple continuous animation material frames, each animation material frame showing a momentary state of the animation character. Multiple continuous animation material frames may include frames where the animation character stands still, frames where the animation character starts to prepare to run, frames where the animation character takes the first step, multiple frames where the animation character is running, showing different stages of running, and frames where the animation character ends running.
[0096] After determining the scene in each animation material frame, the position of the animation character and the action sequence, for each animation material frame, the scene concept map corresponding to the scene in the animation material frame and the three views of the animation character are combined to generate the animation material frame of the shot.
[0097] The embodiment of the present application automatically generates animation material frames based on the animation style type and animation script through a drawing model, which reduces the time consumption of manual drawing and design, improves the overall efficiency of animation production, and ensures the consistency of the visual style of the animation material frames through a unified animation style type, thereby improving the overall quality of the animation.
[0098] In an embodiment of the present application, in step S1014, generating animation material frames based on the three-view images of the animation character in different postures and the scene concept map under the scene type can be achieved in the following way: first, creating a three-dimensional model of the animation character and the scene type based on the three-view images of the animation character in different postures and the scene concept map under the scene type; then, generating a three-dimensional animation for each animation scene in the animation script based on the animation script and the three-dimensional model; the three-dimensional animation includes the animation character and the animation background under the scene type; finally, extracting the two-dimensional animation material frames of the corresponding animation scene from the three-dimensional animation of each animation scene.
[0099] Here, a three-dimensional model of an animated character can be created according to the three-view images of the animated character in different postures by a three-dimensional modeling software, and a three-dimensional model of a scene can be created according to a scene concept map under a scene type. The animated characters and scene types contained in each animation material frame in each animation scene, as well as the actions of the animated characters, are determined from the animation script. For each animation material frame, a three-dimensional model corresponding to a suitable character posture is selected in a three-dimensional space based on the animated character and the action of the animated character by a three-dimensional modeling software, and the three-dimensional model corresponding to the selected character posture is placed in the three-dimensional model of the scene type to obtain the three-dimensional model of the animation material frame. Based on the animation script, the three-dimensional models of multiple animation material frames in the animation scene are rendered into a three-dimensional animation of the animation scene. The three-dimensional animation includes an animated character and an animation background under a scene type. Key frames are extracted from the three-dimensional animation of each animation scene as animation material frames. These key frames can be key moments of character actions, key moments of scene changes, etc. The extracted key frames are rendered into two-dimensional images to form animation material frames.
[0100] The embodiments of the present application greatly improve the efficiency of animation production by automatically creating three-dimensional models, generating three-dimensional animations, and extracting two-dimensional animation material frames, and ensure the consistency of the visual style of the animation material frames, while ensuring the smoothness of character movements and scene changes.
[0101] Step S102: determining a character rendering model corresponding to the animation character based on the character type of the animation character.
[0102] Here, each animated character corresponds to a character type, and each character type corresponds to a character rendering model. The character rendering model corresponding to each character type can be pre-trained and stored in a character rendering model library. The character rendering model can be a machine learning model, such as a Lora model. The character rendering model is used to render the animated character and enrich the character details of the animated character. After determining the character type of the animated character, the character rendering model corresponding to the character type is obtained from the preset character rendering model library based on the character type as the character rendering model corresponding to the animated character.
[0103] In some embodiments, determining the character rendering model corresponding to the animated character based on the character type of the animated character in step S102 can be achieved in the following way: first, segmenting the animation material frame based on the animated character in the animation material frame to obtain multiple material sub-frames; wherein each material sub-frame includes an animated character; then, for each material sub-frame, determining the character rendering model corresponding to the animated character based on the character type of the animated character in the material sub-frame.
[0104] Here, when the animation material frame includes multiple animation characters, an image segmentation algorithm can be used to identify the animation characters in the animation material frame. The animation material frame is segmented to obtain multiple material sub-frames, each of which includes an animation character. For each material sub-frame, the character type of the animation character in the material sub-frame, and the specific implementation process of determining the character rendering model corresponding to the animation character can refer to step S102 in the above embodiment, which will not be repeated here.
[0105] Alternatively, when the animation material frame includes multiple animation characters, the animation material frame can also be segmented based on a preset segmentation method to obtain multiple material sub-frames. For example, the preset segmentation method is to segment the animation material frame from the middle line to obtain a material sub-frame of the left half image and a material sub-frame of the right half image.
[0106] The embodiments of the present application reduce the time for model customization and adjustment and improve the efficiency of animation production by identifying animation characters, automatically segmenting animation material frames and selecting the character rendering model that is most suitable as the basis for rendering.
[0107] Step S103, calling the character rendering model to render the animation character in the animation material frame to obtain a character image.
[0108] Here, the character rendering model is a Lora model, and the character rendering model is a branch network mounted on a stable diffusion model. After the animation material frame is input into the stable diffusion model, the animated character in the animation material frame can be subjected to multiple rounds of iterative denoising based on the weight parameters of the character rendering model, and finally a rendered character image is obtained. The embodiment of the present application does not further limit the rendering process of the character rendering model, and reference can be made to the inverse diffusion process of the stable diffusion model. Compared with the animated character in the original animation material frame, the rendered character image has a significantly improved overall picture effect and richer details, such as more makeup details for eye makeup, more painting strokes and shadows for clothing, and the nose and mouth have changed from basically no to obvious shapes, and the overall character is more three-dimensional.
[0109] In some embodiments, the pre-made animation material frame includes a plurality of consecutive animation material frames. Figure 7 , Figure 7 It is shown that in step S103, the character rendering model is called to render the animation character in the animation material frame to obtain the character image, which can be achieved by the following steps S1031 to S1034:
[0110] Step S1031, determining a current animation material frame from a plurality of animation material frames.
[0111] For example, the pre-made animation material frame may be a plurality of continuous animation material frames within 1 second, for example, when the frame rate is 24fps, the pre-made animation material frame may include 24 animation material frames. The current animation material frame is the animation material frame being rendered by the character rendering model at the current moment, for example, the third animation material frame among the 24 animation material frames.
[0112] Step S1032, obtaining the previous animation material frame and the next animation material frame of the current animation material frame.
[0113] For example, if the current animation material frame is the 3rd animation material frame among 24 animation material frames, then the previous animation material frame is the 2nd animation material frame, and the next animation material frame is the 4th animation material frame.
[0114] Step S1033, splicing the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame.
[0115] For example, the current animation material frame is the 3rd frame animation material frame, the previous frame animation material frame is the 2nd frame animation material frame, and the next frame animation material frame is the 4th frame animation material frame. The 2nd frame animation material frame is placed on the left, the 3rd frame animation material frame is placed in the middle, and the 4th frame animation material frame is placed on the right to stitch the screen to obtain a complete image, and the complete image is determined as the spliced animation material frame.
[0116] Step S1034, calling the character rendering model to render multiple animation characters in the spliced animation material frames to obtain character images.
[0117] Here, the spliced animation material frames are input into the character rendering model, and based on the weight parameters of the character rendering model, multiple rounds of iterative denoising are performed on multiple animated characters in the spliced animation material frames to obtain character images. During the multi-round iterative denoising process, the features of the previous frame, the current frame, and the next frame in the spliced animation material frames are gradually merged to improve the consistency of continuous frame rendering.
[0118] In the embodiment of the present application, since the noise addition and denoising process of the character rendering model acts on the entire image of the spliced animation material frame, the spliced animation material frame is input into the character rendering model. When the character image of the current animation material frame is generated, it will be affected by the previous animation material frame and the next animation material frame, thereby significantly improving the stability of the continuous picture.
[0119] In an embodiment of the present application, calling the character rendering model in step S1034 to render multiple animated characters in the spliced animation material frame to obtain a character image can be achieved in the following way: first, calling the character rendering model to render multiple animated characters in the spliced animation material frame to obtain an initial character image; then, performing image segmentation on the initial character image to obtain a first character image corresponding to a previous frame of animation material, a second character image corresponding to a current animation material frame, and a third character image corresponding to a next frame of animation material; finally, determining the second character image as the character image.
[0120] Here, the character rendering model is called to render multiple animated characters in the spliced animation material frame to obtain an initial character image of the same size as the spliced animation material frame. That is, the initial character image includes a first character image corresponding to the previous animation material frame, a second character image corresponding to the current animation material frame, and a third character image corresponding to the next animation material frame. The position information or size information of the previous animation material frame, the current animation material frame, and the next animation material frame in the spliced animation material frame are obtained, and the initial character image is segmented based on the position information or size information to obtain the first character image corresponding to the previous animation material frame, the second character image corresponding to the current animation material frame, and the third character image corresponding to the next animation material frame. The second character image located in the middle is determined as the character image.
[0121] The embodiment of the present application can improve the stability of continuous images by segmenting the initial character image after rendering and retaining only the middle second character image as the character image.
[0122] Step S104, calling a preset background rendering model to render the animation background in the animation material frame to obtain a background image.
[0123] Here, the background rendering model can be a Dreambooth model. The Dreambooth model is a model obtained by training and fine-tuning the stable diffusion model with animation style types. The character rendering model is a branch network mounted on the background rendering model. After the animation material frame is input into the background rendering model, the animation background in the animation material frame can be subjected to multiple rounds of iterative denoising processing based on the weight parameters of the background rendering model, and finally a rendered background image is obtained. Compared with the animation background in the original animation material frame, the rendered background image has a significantly improved overall picture effect and richer details.
[0124] In some embodiments, the pre-made animation material frame includes a plurality of consecutive animation material frames. Figure 8 , Figure 8It is shown that in step S104, a preset background rendering model is called to render the animation background in the animation material frame to obtain a background image, which can be achieved by the following steps S1041 to S1044:
[0125] Step S1041, determining a current animation material frame from a plurality of animation material frames.
[0126] Here, the specific implementation process of determining the current animation material frame from the multiple animation material frames can refer to step S1031 in the above embodiment, and will not be repeated here.
[0127] Step S1042, obtaining the previous animation material frame and the next animation material frame of the current animation material frame.
[0128] Here, the specific implementation process of obtaining the previous animation material frame and the next animation material frame of the current animation material frame can refer to step S1032 in the above embodiment, which will not be repeated here.
[0129] Step S1043, splicing the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame.
[0130] Here, the specific implementation process of splicing the previous animation material frame, the current animation material frame and the next animation material frame to obtain the spliced animation material frame can refer to step S1033 in the above embodiment, which will not be repeated here.
[0131] Step S1044, calling a preset background rendering model to render multiple animation backgrounds in the spliced animation material frames to obtain a background image.
[0132] Here, the spliced animation material frames are input into the background rendering model, and multiple animation backgrounds of the spliced animation material frames are subjected to multiple rounds of iterative denoising based on the weight parameters of the background rendering model to obtain a background image. During the multiple rounds of iterative denoising process, the features of the previous frame, the current frame, and the next frame of the spliced animation material frame are gradually merged to improve the consistency of continuous frame rendering.
[0133] In the embodiment of the present application, since the noise addition and denoising process of the background rendering model acts on the entire image of the spliced animation material frame, the spliced animation material frame is input into the background rendering model. When the background image of the current animation material frame is generated, it will be affected by the previous frame of animation material frame and the next frame of animation material frame, thereby significantly improving the stability of the continuous picture.
[0134] In an embodiment of the present application, in step S1044, a preset background rendering model is called to render multiple animation backgrounds in the spliced animation material frame to obtain a background image, which can be achieved in the following way: first, a preset background rendering model is called to render multiple animation backgrounds in the spliced animation material frame to obtain an initial background image; then, the initial background image is segmented to obtain a first background image corresponding to a previous frame of animation material, a second background image corresponding to a current animation material frame, and a third background image corresponding to a next frame of animation material; finally, the second background image is determined as the background image.
[0135] Here, the background rendering model is called to render multiple animation backgrounds in the spliced animation material frame to obtain an initial background image of the same size as the spliced animation material frame. That is, the initial background image includes a first background image corresponding to the previous animation material frame, a second background image corresponding to the current animation material frame, and a third background image corresponding to the next animation material frame. The position information or size information of the previous animation material frame, the current animation material frame, and the next animation material frame in the spliced animation material frame are obtained, and the initial background image is segmented based on the position information or size information to obtain the first background image corresponding to the previous animation material frame, the second background image corresponding to the current animation material frame, and the third background image corresponding to the next animation material frame. The second background image located in the middle is determined as the background image.
[0136] The embodiment of the present application can improve the stability of the continuous picture by dividing the initial background image after rendering and retaining only the middle second background image as the background image.
[0137] In some embodiments, the pre-made animation material frame includes a plurality of continuous animation material frames. In step S104, calling a preset background rendering model to render the animation background in the animation material frame to obtain a background image can also be implemented in the following manner: first, determining the current animation material frame from the plurality of animation material frames; and obtaining the previous animation material frame and the next animation material frame of the current animation material frame; then, respectively encoding the previous animation material frame, the animation material frame and the next animation material frame to obtain a first material frame vector corresponding to the previous animation material frame, a second material frame vector corresponding to the current animation material frame and a third material frame vector corresponding to the next animation material frame; then, fusing the first material frame vector, the second material frame vector and the third material frame vector to obtain a fused material frame vector; finally, calling the background rendering model to render the animation background in the animation material frame based on the fused material frame vector to obtain a background image.
[0138] Here, the specific implementation process of determining the current animation material frame from multiple animation material frames can refer to step S1031 in the above embodiment, which will not be repeated here. The specific implementation process of obtaining the previous animation material frame and the next animation material frame of the current animation material frame can refer to step S1032 in the above embodiment, which will not be repeated here. The previous animation material frame, the animation material frame and the next animation material frame are respectively encoded to obtain a first material frame vector corresponding to the previous animation material frame, a second material frame vector corresponding to the current animation material frame and a third material frame vector corresponding to the next animation material frame. The embodiment of the present application does not limit the encoding method. For example, a deep learning model (such as a convolutional neural network) can be used to encode the animation material frames of the previous frame, the current frame and the next frame to obtain a corresponding vector representation. The embodiment of the present application does not limit the specific manner of fusing the first material frame vector, the second material frame vector and the third material frame vector. For example, it can be a weighted average process or a direct splicing process. The fused material frame vector obtained after fusion is input into the background rendering model for processing to obtain a background image.
[0139] The embodiment of the present application automatically encodes continuous animation material frames to obtain vectors, and fuses the vectors, so that the influence of adjacent animation material frames can be taken into account during the rendering of the current animation material frame, thereby improving the efficiency of background rendering and the consistency of background images of continuous frames.
[0140] In an embodiment of the present application, a background rendering model is called to render the animation background in the animation material frame based on the fused material frame vector to obtain a background image. This can be achieved in the following ways: calling the background rendering model to perform attention processing on the fused material frame vector to obtain an attention feature; and determining the background image of the animation background based on the fused material frame vector and the attention feature.
[0141] Here, since the background rendering model is a stable diffusion model, the background rendering model will perform attention processing on the fused material frame vector to obtain attention features. The embodiment of the present application does not limit the specific process of attention processing, and the model mechanism of the stable diffusion model can be referred to. After obtaining the attention feature, the attention feature can be linearly normalized, etc. to obtain predicted noise, which is also a vector representation. Subtract the predicted noise from the fused material frame vector to obtain the input of the background rendering model for the next round of iterative processing, and repeat the above steps until the background rendering model outputs the background image of the animation background.
[0142] The embodiments of the present application greatly improve the efficiency of background rendering through automated encoding, fusion processing and attention mechanism, while taking into account the influence of previous and next frames, ensuring the consistency of the visual style of the background image and improving the accuracy and fluency of background rendering.
[0143] Step S105, performing image fusion processing on the character image and the background image to obtain a target animation containing animation materials.
[0144] Here, for each frame of animation material, the character layer of the character image and the background layer of the background image in the animation material frame can be determined. The character layer and the background layer are superimposed to obtain the animation image of the animation material frame. The animation images of multiple frames of animation material frames are spliced in sequence to obtain an animation material frame sequence of an animation scene. The animation material frame sequence of each animation scene is spliced in sequence according to the requirements of the animation script to obtain a target animation containing animation material.
[0145] In the process of making animation, the embodiment of the present application can ensure that the visual performance of each animated character meets its unique design style and artistic requirements through pre-made animation material frames combined with character rendering models customized for different character types, thereby enhancing the expressiveness and recognition of the animated character. The animation background is rendered using a preset background rendering model. After obtaining the background image, the character image and the background image are subjected to image fusion processing to obtain the target animation containing the animation material. This can not only make the background environment more realistic and vivid, and enhance the immersion and viewing of the overall picture, but also enable the character rendering model and the background rendering model to perform fast, batch, and semi-automatic animation rendering, thereby improving the efficiency of animation rendering. Therefore, the embodiment of the present application can improve the production efficiency and picture quality of the animation.
[0146] In some embodiments, see Fig. 9 , Fig. 9 It is shown that in step S105, the character image and the background image are subjected to image fusion processing to obtain a target animation containing animation materials, which can be achieved by the following steps S1051 to S1054:
[0147] Step S1051, determining the animation image corresponding to the animation material frame based on the character image and the background image corresponding to the animation material frame.
[0148] Here, the character layer of the character image and the background layer of the background image in the animation material frame can be determined, and the character layer and the background layer are superimposed to obtain the animation image of the animation material frame.
[0149] In an embodiment of the present application, determining the animation image corresponding to the animation material frame based on the character image and background image corresponding to the animation material frame in step S1051 can be achieved in the following manner: first, performing background feature extraction on the background image to obtain image feature information of the background image; and performing character feature extraction on the character image to obtain image feature information of the character image; then, based on the image feature information of the background image, generating a background layer of the background image; and based on the image feature information of the character image, generating a character layer of the character image; finally, superimposing the character layer and the background layer to obtain the animation image corresponding to the animation material frame.
[0150] Here, an image processing algorithm (such as a convolutional neural network) can be used to extract features from the background image to obtain image feature information of the background image. Similarly, an image processing algorithm is used to extract features from the character image to obtain image feature information of the character image. The image feature information of the background image can include the position information and size information of the background image in the animation image frame, as well as the color features, texture features, edge features, shape features, spatial layout features and lighting features of the background. Among them, the color features include the main color distribution, hue, saturation, etc. in the background image. The texture features describe the texture details of the surface of the object in the background image, such as roughness, pattern repeatability, etc. The edge features represent the boundary information between different areas in the background image, which helps to distinguish different objects or areas. The shape features describe the shape features of the objects in the background image, such as round, square, etc. The spatial layout features represent the spatial distribution and relative position relationship of the objects in the background image. The lighting features describe the lighting conditions in the background image, such as the position, intensity, direction, etc. of the light source.
[0151] The image feature information of the character image may include the position information and size information of the animated character in the animated image frame, as well as the color features, texture features, edge features, shape features, posture features, expression features and action features of the animated character. The color features of the character include the main color distribution, hue, saturation, etc. in the character image. The texture features of the character describe the texture details of the surface of the object in the character image, such as the texture of clothing, the texture of the skin, etc. The edge features of the character represent the boundary information between different regions in the character image, which helps to distinguish different objects or regions. The shape features of the character describe the shape features of the object in the character image, such as the body outline and facial features of the character. The posture features represent the posture information of the character in the character image, such as standing, walking, jumping, etc. The expression features describe the expression information of the character in the character image, such as smiling, frowning, etc. The action features represent the action information of the character in the character image, such as gestures and eyes.
[0152] A background layer can be generated based on the image feature information of the background image. The background layer retains the complete information of the background image, including color, texture, etc. A character layer can be generated based on the image feature information of the character image. The character layer retains the complete information of the character image, including the shape, action, etc. of the character. Ensure that the format of the background layer and the character layer is consistent for subsequent overlay operations. Use the image overlay algorithm to overlay the character layer on the background layer to form a complete animation image.
[0153] The embodiments of the present application greatly improve the efficiency of animation image generation through automated feature extraction and layer generation.
[0154] Step S1052 , determining edge information and optical flow information of each animation material frame from the i-th animation material frame to the i+j-th animation material frame in the pre-made multi-frame animation material frame.
[0155] Wherein, i and j are both integers greater than 0.
[0156] Here, at least two key frames can be selected from the pre-made multi-frame animation material frames, and the key frames can be rendered based on the character rendering model and the background rendering model to obtain the animation images corresponding to the key frames. The i-th animation material frame and the i+j-th animation material frame are key frames. After rendering to obtain the i-th animation image corresponding to the i-th animation material frame and the i+j-th animation image corresponding to the i+j-th animation material frame, an edge algorithm (such as a canny algorithm) can be used to determine the edge information of each animation material frame from the i-th animation material frame to the i+j-th animation material frame. Edge information refers to the position information of the boundaries of the animation characters and objects in the background in the animation image, which is used to identify the outlines and shapes of the animation characters and objects. An optical flow algorithm (such as a GMFlow algorithm) can be used to determine the optical flow information of each animation material frame from the i-th animation material frame to the i+j-th animation material frame. Optical flow information refers to the movement direction and speed information of pixels between consecutive frames in the animation image.
[0157] Edge algorithms usually include the following steps: Grayscale conversion: convert the color animation image into a grayscale image. Noise filtering: use filters (such as Gaussian filters) to remove noise in grayscale images. Gradient calculation: calculate the gradient amplitude and direction of each pixel in the grayscale image. Non-maximum suppression: retain the pixels with the local maximum gradient amplitude as edge points. Double threshold detection: use high and low thresholds to further screen the edge points to obtain the final edge information. Optical flow algorithms usually include the following steps: Brightness constant assumption: assume that the brightness of the same pixel between two consecutive frames of animation images remains unchanged. Gradient calculation: calculate the gradient amplitude and direction of each pixel in the animation image. Optical flow equation solution: solve the optical flow equation based on the gradient information and the brightness constant assumption to obtain the optical flow vector of each pixel. Optical flow vector field: generate an optical flow vector field based on the obtained optical flow vector to represent the movement of the character or object between consecutive frames of animation images.
[0158] For example, assuming that there are 24 frames of animation material, the 1st frame, the 4th frame, the 6th frame, ..., the 24th frame are selected as key frames, and when i is 1 and j is 3, for the 1st frame of animation material frame and the 4th frame of animation material frame, the 1st frame of animation image corresponding to the 1st frame of animation material frame and the 4th frame of animation image corresponding to the 4th frame of animation material frame can be generated first. Then, the canny algorithm is used to generate edge information of the 1st frame of animation material frame, the 2nd frame of animation material frame, the 3rd frame of animation material frame and the 4th frame of animation material frame respectively, and the GMFlow optical flow algorithm is used to generate optical flow information of the 1st frame of animation material frame, the 2nd frame of animation material frame, the 3rd frame of animation material frame and the 4th frame of animation material frame respectively.
[0159] Step S1053, determining the i+kth frame animation image based on the edge information and optical flow information of each animation material frame from the i-th frame animation material frame to the i+j-th frame animation material frame, as well as the i-th frame animation image corresponding to the i-th frame animation material frame and the i+j-th frame animation image corresponding to the i+j-th frame animation material frame.
[0160] The value of k is any integer between 1 and j-1.
[0161] For example, when i is 1 and j is 3, k=1, 2, the second frame of the animation image and the second frame of the animation image can be determined based on the edge information and optical flow information of the first frame of the animation material frame, the second frame of the animation material frame, the third frame of the animation material frame and the fourth frame of the animation material frame, as well as the first frame of the animation image and the fourth frame of the animation image. The motion of the character or object in the second frame of the animation image is predicted using the optical flow information, and the outline and shape of the character or object in the second frame of the animation image is predicted using the edge information. The second frame of the animation image is generated based on the first frame of the animation image and the fourth frame of the animation image, as well as the motion of the character or object in the second frame of the animation image, and the outline and shape of the character or object in the second frame of the animation image. Similarly, the third frame of the animation image can be generated. In the above prediction process, weight parameters can also be preset to make the second frame of the animation image closer to the first frame of the animation image, and the third frame of the animation image closer to the fourth frame of the animation image. Exemplarily, in the process of generating the second frame of the animation image, a weight parameter of 0.8 is added to the first frame of the animation image, and a weight parameter of 0.2 is added to the second frame of the animation image.
[0162] Step S1054, splicing the first frame of animation image to the nth frame of animation image into a target animation.
[0163] Wherein, n is the number of frames of the multi-frame animation material.
[0164] Here, n is an integer greater than 1. For example, if n is 24, the first frame of the animation image to the 24th frame of the animation image can be sequentially spliced to obtain the target animation.
[0165] The embodiment of the present application performs model rendering processing on only part of the animation material frames in multiple frames of animation material frames, and performs frame supplementation prediction on the remaining animation material frames through edge information and optical flow information to obtain animation images, thereby improving the efficiency of animation production.
[0166] In some embodiments, the animation material in the animation material frame also includes animation special effects. In step S105, image fusion processing is performed on the character image and the background image to obtain a target animation containing the animation material, which can also be achieved in the following manner: first, based on the special effect type of the animation special effect, a special effect rendering model corresponding to the animation special effect is determined; then, the special effect rendering model is called to render the animation special effect in the animation material frame to obtain a special effect image; finally, image fusion processing is performed on the special effect image, the character image and the background image to obtain a target animation containing the animation material.
[0167] Here, animation special effects refer to various visual effects added to animations to enhance the expressiveness and visual impact of animations. The special effect types of animation special effects can be classified according to the expression form and purpose of animation special effects, and the special effect types include but are not limited to explosion special effects, flame special effects, magic special effects, etc. For each frame of animation material, if there is an animation special effect in the animation material frame, the special effect type of the animation special effect in the animation material frame can be obtained. Based on the special effect type of the animation special effect, the special effect rendering model corresponding to the animation special effect is determined. The special effect rendering model is a Lora model that can be mounted on the background rendering model. Each special effect type corresponds to a special effect rendering model. After the animation material frame is input into the special effect rendering model, the animation special effects in the animation material frame can be subjected to multiple rounds of iterative denoising based on the weight parameters of the special effect rendering model, and finally a rendered special effect image is obtained. The special effect layer of the special effect image in the animation material frame, the character layer of the character image, and the background layer of the background image can be determined. The special effect layer, the character layer, and the background layer are superimposed to obtain the animation image of the animation material frame. The animation images of multiple frames of animation material frames are spliced in order to obtain an animation material frame sequence of an animation scene. The animation material frame sequence of each animation scene is spliced in order according to the requirements of the animation script to obtain a target animation containing animation materials.
[0168] The embodiment of the present application renders animation special effects through a special effects rendering model, thereby improving the quality of animation images, and greatly improves the efficiency of animation image generation through automated special effects rendering and image fusion processing.
[0169] In some embodiments, Fig.10 is another flowchart of the animation generation method provided in the embodiment of the present application. Fig.10 As shown, the animation generation method provided in the embodiment of the present application further includes the following steps S201 to S202:
[0170] Step S201, calling multiple painting models to generate pictures of multiple picture styles respectively.
[0171] Each painting model is used to generate multiple pictures of a certain picture style.
[0172] Here, the painting model can be a machine learning model or a deep learning model. Each painting model corresponds to a picture style type. Picture style types include but are not limited to hand-painted style, cartoon style, pixel art style, and even the style of a specific painter. For each painting model, you can set parameters according to the requirements of the painting model, such as image size, color mode, etc., enter text describing the picture, and call the painting model to batch generate multiple pictures of the picture style type corresponding to the text.
[0173] Step S202 , in response to the selection operation of the animation style type among the multiple picture style types, a preset diffusion model is trained using a picture of the animation style type to obtain a preset background rendering model with the animation style type.
[0174] For example, the various picture style types include hand-painted style, cartoon style, and pixel art style. Obtain multiple pictures of hand-painted style, multiple pictures of cartoon style, and multiple pictures of pixel art style. The user can select multiple picture style types on the client of the animation generation application, and select an animation style type from the multiple picture style types. For example, the selected animation style types are hand-painted style and cartoon style. Then, multiple pictures of hand-painted style and multiple pictures of cartoon style are used to train the preset diffusion model. The preset diffusion model is a stable diffusion model. The training process of the preset diffusion model is briefly described as follows: First, forward diffusion is performed: for each picture, a series of intermediate states x1, x2, ..., x are generated through the forward diffusion process. n . Then, add noise: add Gaussian noise to each intermediate state. Noise prediction: use the noise prediction network to predict noise. Loss calculation: calculate the mean square error (MSE) between the predicted noise and the actual noise. Parameter update: use the gradient descent method or other optimization algorithms to update the model parameters. Iterative training: repeat the above steps until the loss function converges or reaches the preset training rounds to obtain a preset background rendering model with an animation style type. When multiple animation style types are selected, the background image output by the trained background rendering model is the picture style of the fusion of multiple animation style types. When an animation style type is selected, the background image output by the trained background rendering model is the animation style type.
[0175] In the animation preparation stage, the embodiment of the present application automatically calls multiple painting models to generate pictures of various picture styles, and uses pictures of selected animation style types to train the background rendering model, which saves manual operations in the animation preparation stage and improves the efficiency of animation production.
[0176] In some embodiments, see Fig.11 , Fig.11 It is shown that in step S202, a preset diffusion model is trained using an animation style picture to obtain a preset background rendering model with an animation style, which can be achieved by the following steps S2021 to S2024:
[0177] Step S2021, using an animation-style picture to train a preset diffusion model to obtain an initial rendering model with an animation-style type.
[0178] Here, the specific process of using animation-style pictures to train the preset diffusion model to obtain the initial rendering model with the animation style type can refer to the training process of the preset diffusion model in the above embodiment, which will not be repeated here.
[0179] Step S2022, determining the edge image of the animation style type picture.
[0180] Here, an edge algorithm can be used to process the animation-style image to obtain an edge image.
[0181] Step S2023: Use the edge image to train the preset control model to obtain the edge control model.
[0182] Here, multiple edge images can be used to train the preset control model to obtain the edge control model. The preset control model can be a Controlnet network structure. The embodiment of the present application does not limit the training process of the control model. The training process of the Lora model can be referred to, and will not be repeated here.
[0183] Step S2024: The initial rendering model and the edge control model are merged to obtain a preset background rendering model.
[0184] Here, after the edge control model is obtained, the edge control model is mounted on the initial rendering model as a branch network to obtain a preset background rendering model.
[0185] The embodiment of the present application adds an edge control model to increase edge control during background rendering, thereby improving picture control capabilities and improving the quality of animated images.
[0186] In some embodiments, Fig.12 is another flowchart of the animation generation method provided in the embodiment of the present application. Fig.12 As shown, the animation generation method provided in the embodiment of the present application further includes the following steps S301 to S302:
[0187] Step S301, based on the three views of the animated character in different postures, close-up processing of different parts of the animated character is performed to obtain multiple training material images of the animated character.
[0188] For example, 40 training material images can be collected for each animated character. For each animated character, 10 images of face close-up, close-up above the chest, close-up above the body, close-up above the knees, etc. are collected to obtain 40 training material images for each animated character.
[0189] Step S302, using a plurality of training material images of the animation character, training the preset character rendering model to be trained, and obtaining the character rendering model corresponding to the animation character.
[0190] Here, the preset character rendering model to be trained is the Lora model. The Lora model can be mounted on the background rendering model as a branch network, and the weight parameters of the background rendering model can be frozen, that is, during the training process, the weight parameters of the background rendering model are kept unchanged, and the weight parameters of the character rendering model to be trained are iteratively updated until the loss function converges to obtain the character rendering model corresponding to the animated character.
[0191] The embodiment of the present application trains the character rendering model to be trained through multiple close-up training material images of different parts, and obtains the character rendering model corresponding to the animated character, which can improve the accuracy of the trained character rendering model, thereby improving the rendering effect of the character image and improving the image quality.
[0192] Fig.13 is another optional flow chart of the animation generation method provided in the embodiment of the present application, such as Fig.13 As shown, the method includes the following steps S401 to S410:
[0193] Step S401: The terminal receives an interactive operation from a user.
[0194] Here, the user's interactive operation may be an operation of inputting a pre-made animation material frame operation, a click to start generating an animation operation, and the like.
[0195] Step S402: the terminal generates an animation generation request in response to the interactive operation.
[0196] Here, after receiving the user's interactive operation, the terminal may encapsulate the pre-made animation material frame into the animation generation request.
[0197] Step S403: the terminal sends an animation generation request to the server.
[0198] Step S404: the server obtains pre-made animation material frames in response to the animation generation request.
[0199] The animation material in the animation material frame includes an animation character and an animation background.
[0200] Here, the specific process of obtaining the pre-made animation material frame can refer to step S101 in the above embodiment, and will not be repeated here.
[0201] Step S405: The server determines a character rendering model corresponding to the animation character based on the character type of the animation character.
[0202] Here, the specific process of determining the character rendering model corresponding to the animation character based on the character type of the animation character can refer to step S102 in the above embodiment, which will not be repeated here.
[0203] Step S406: The server calls the character rendering model to render the animation character in the animation material frame to obtain a character image.
[0204] Here, the specific process of calling the character rendering model to render the animation character in the animation material frame and obtaining the character image can refer to step S103 in the above embodiment, which will not be repeated here.
[0205] Step S407: the server calls a preset background rendering model to render the animation background in the animation material frame to obtain a background image.
[0206] Here, the specific process of calling the preset background rendering model to render the animation background in the animation material frame and obtaining the background image can refer to step S104 in the above embodiment, which will not be repeated here.
[0207] Step S408: the server performs image fusion processing on the character image and the background image to obtain a target animation containing animation materials.
[0208] Here, the specific process of performing image fusion processing on the character image and the background image to obtain the target animation containing the animation material can refer to step S105 in the above embodiment, which will not be repeated here.
[0209] Step S409: the server sends the target animation to the terminal.
[0210] Step S410: the terminal displays the target animation on the current interface.
[0211] The embodiment of the present application uses pre-made animation material frames combined with character rendering models customized for different character types to ensure that the visual performance of each animated character meets its unique design style and artistic requirements, thereby enhancing the expressiveness and recognition of the animated character. The animation background is rendered using a preset background rendering model. After obtaining the background image, the character image and the background image are subjected to image fusion processing to obtain a target animation containing animation materials. This not only makes the background environment more realistic and vivid, and enhances the immersion and viewing of the overall picture, but also allows the character rendering model and the background rendering model to perform fast, batch, and semi-automatic animation rendering, thereby improving animation rendering efficiency. Therefore, the embodiment of the present application can improve the production efficiency and picture quality of animation.
[0212] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0213] It should be noted that the animation generation method provided in the embodiment of the present application is applied to any scenario that uses artificial intelligence technology to generate animation, such as film and television animation production scenarios, game development scenarios, live interactive scenarios, education and training scenarios, cultural protection and reconstruction scenarios, medical and health visualization scenarios, experimental simulations, etc.
[0214] In the traditional animation production process, the labor cost and time cost are high. The cost of a single minute of animation is between 120,000 and 150,000 yuan. Due to the limitations of manpower and machine resources, even if the time requirements of the animation project are urgent, it is difficult to find a quick way to improve efficiency. Based on this, the embodiment of the present application proposes an animation generation method, which is a full-process assisted animation production method based on a stable diffusion model (Stable Diffusion) and an AI painting model. On the basis of not changing the traditional animation production process in the industry, artificial intelligence technology is introduced into the animation production process to assist the efficient production and production of animation, improve the picture quality, and improve the production efficiency of the animation industry.
[0215] The embodiment of the present application solves the following technical problems: In the early stage of animation production, a stable diffusion model (StableDiffusion) is used for visual exploration and character design. With the characteristics of diverse styles and high drawing efficiency of AI painting models, a variety of different picture styles and visual exploration pictures are generated to help directors quickly determine the main picture style of the animation. At the same time, with conditional control technologies such as Controlnet control network, hand-drawn sketches, design drafts of relevant characters, three-view drawings and other materials are generated, which can improve production efficiency and picture quality. In other words, the embodiment of the present application uses AIGC technology to efficiently generate picture style maps. Different style pictures can be quickly generated for the same scene for selection. Similarly, for character design, AIGC technology can be used to quickly customize character images, generate character three-view drawings, manuscripts and other materials. At the same time, details can be adjusted and optimized in real time. Greatly improve the communication and confirmation efficiency between animation producers and directors. Help directors and producers quickly determine the overall animation production direction.
[0216] In the process of animation three-to-two rendering, AI painting technology based on the stable diffusion model (Stable Diffusion) is used to replace the traditional industrial software rendering. By training the overall style of the picture, the Dreambooth fine-tuning technology (corresponding to the background rendering model in the above embodiment) is used to ensure that the style meets the director's requirements. For each animated character, a dedicated Lora model (corresponding to the character rendering model in the above embodiment) is trained to determine the facial features and clothing details of the character. At the same time, in order to solve the stability problem of continuous picture generation. The embodiment of the present application proposes a technical solution of multi-frame rendering and AI interpolation to solve the stability problem of continuous picture generation, and combined with the training of the algorithm model, it effectively solves the problem of picture jitter and detail changes in AI video generation. Multi-frame rendering refers to splicing the pictures of the previous and next two frames on the left and right sides of the picture during the generation of the current frame; because the noise and denoising process of the stable diffusion model acts on the entire image. Therefore, the picture generation of the current frame will be affected by the content of the previous and next frames, thereby improving the stability of the continuous picture. AI interpolation refers to first selecting a key frame to generate through a stable diffusion model, and between key frames, a special AI model is used to supplement. Compared with the complete image generation process of the AI painting model, AI frame interpolation pays more attention to the migration of picture style and has higher picture stability.
[0217] That is to say, the animation "three-render-two" process of the embodiment of the present application is completely replaced by an AI painting model based on Stable Diffusion, and combined with the technology to improve the stability of continuous screen generation, it can perform fast, batch, and semi-automatic screen rendering. At the same time, after determining the model parameters, the server-side rendering project code can be used to perform parallel batch processing, significantly improving the efficiency of animation production; and because a set of rendering project codes can run on multiple servers, the speed of animation rendering will no longer be limited by machine resources. At the same time, it can reduce the subjective influence of manual operation on the final effect and achieve the unification of animation rendering effects.
[0218] The embodiment of the present application also uses the latest stable diffusion model, combined with the control information introduction of the control network (Controlnet) structure, and multi-frame rendering and AI frame filling technology to significantly improve the efficiency of animation generation and picture effects. Compared with traditional painting software, it can bring a refreshing effect to the audience, give full play to the director's imagination and artistic creativity, and achieve some amazing effects that traditional processes cannot achieve.
[0219] The animation generation method provided in the embodiment of the present application is described in detail below.
[0220] In the embodiments of the present application, the use of AIGC technology to assist in animation generation is divided into several parts; first, in the early stage of animation production, on the one hand, based on the stable diffusion model, the corresponding AI painting model is trained. Based on the animation theme, visual exploration is carried out, including different picture styles, visual directions, etc. On the other hand, with the help of the AI painting model's ability to quickly output pictures, combined with the introduction of control information from the Controlnet control network, the design of characters, scenes, and props is carried out, and the design drafts and three-view drawings of related characters are generated. In addition, in the process of rendering traditional animations from 3D pictures to 2D, the embodiments of the present application use AI painting technology based on a stable diffusion model to replace traditional industrial software rendering. Using the AI painting model, the animation pictures are redrawn with high quality; at the same time, the generation stability of continuous pictures is improved in conjunction with technologies such as multi-frame rendering. The embodiments of the present application not only improve the rendering efficiency, but also further improve the effect and quality of the generated pictures.
[0221] On the product side, the embodiment of the present application realizes the full process of auxiliary animation production through AIGC technology. Fig.14 It is a schematic diagram of the AI-assisted animation process provided by the embodiment of the present application. Among them, after the introduction of AIGC technology, AI full-process assisted animation 600 includes step 601 early visual exploration, step 602 based on AI painting to explore the picture style, step 603 character scene prop design, step 604 using AI painting to generate character design drafts, three views, etc., step 605 initial scene framework design, step 606 animation scene construction, step 607 character asset production and action binding, step 608 script writing, step 609 formal storyboard and dynamic script writing, step 610: 3D animation production, step 611 AI three-rendering two-screen production, step 612 rendering model advance training and step 613 animation post-processing and production, etc. The characters may include character 1, character 2, etc. 3D animation production can divide the animation into the first animation, the second animation, the third animation, etc. according to the animation scenes for production. The three-rendering two-screen production is rendered according to the divided animation scenes. In step 611 of the embodiment of the present application, when performing AI three-in-one rendering and two-in-two animation production, technical research can be first conducted to obtain training materials for the rendering model. The trained rendering model can be used for rendering to improve picture stability, rendering effect, and efficiency. Then, step 612 is performed to train the rendering model in advance to obtain the rendering model, and the three-in-one rendering and two-in-two animations such as the first scene and the second scene are rendered based on the rendering model. The embodiment of the present application adopts AI painting technology based on Stable Diffusion, which completely replaces the traditional manual animation rendering operation, and realizes the full process of animation production assistance and acceleration.
[0222] Here, the process of exploring the picture style based on AI painting is as follows: users can use the user interface (WebUI) of the AI drawing tool to select a variety of popular AI painting models to draw pictures in different scenes and explore the expressiveness of different painting styles for the picture. After the user preliminarily determines the AI painting model to be used, according to the director's understanding and expectations of the picture, select the picture of the corresponding style to fine-tune the Dreambooth model used, and finally achieve the desired effect. Using AI painting to generate a character design draft, the process of three views is as follows: With the help of AIGC technology, generate three views and other materials of the animation character, determine the details of the character such as clothing, hairstyle, facial features, etc., and make convenient and quick adjustments. Fig.15 This is a multi-angle view of a single character provided by the embodiment of the present application. Fig.15 As shown, through AIGC technology, the skeleton diagram of the corresponding posture can be controlled through the Controlnet control network structure to generate postures, and multi-angle pictures of the character can be drawn at one time, which greatly improves the efficiency of the character design stage.
[0223] The process of AI three-rendering and two-animation production is as follows: Use the trained algorithm model (the Dreambooth model (corresponding to the background rendering model in the above embodiment) with the Lora model (corresponding to the character rendering model in the above embodiment) and Controlnet (corresponding to the control model in the above embodiment)) to quickly, batch-wise, and semi-automatically render the initial version of the 3D-generated image (corresponding to the animation material frame in the above embodiment) using AIGC technology. Improve the color and details of the image and optimize the image style. The overall rendering process is carried out according to animation scenes (for example, one scene includes 30-60s animation and 24 frames of animation images per second). After the batch rendering is complete, the final film editing is completed through video editing software. Fig.16 is a comparison diagram of the screen effects before and after rendering provided by the embodiment of the present application. Fig.16 After AI rendering, the overall picture effect has been significantly improved. At the same time, in terms of character details: the eye makeup adds more makeup details, the clothing can also add painting strokes and shadows, the nose and mouth have obvious shapes, and the overall character is more three-dimensional.
[0224] Fig.17 is a simplified system flow chart of the animation generation method provided in the embodiment of the present application. Fig.17, the embodiment of the present application divides the animation production process into three parts: animation production pre-preparation 700, 3D animation production 701 and AI animation rendering process 702. The main work of the embodiment of the present application is in the animation production pre-preparation 700 and the AI animation rendering process 702. In the animation production pre-preparation stage 700, AIGC drawing technology is used to assist in style exploration and character design. Visual exploration and character design can be performed. Among them, visual exploration is achieved through step 704 AI style exploration, and based on the determined style, the picture style model is trained, that is, the Dreambooth model with an exclusive picture style is trained, that is, the stable diffusion model. Character design is achieved through step 705 AI multi-role material generation to obtain materials for multiple animation characters. Based on the determined animation character's material, the LoRA model corresponding to each animation character is trained separately, that is Fig.17 The multi-role model (Finetune) in is used in the rendering stage. The 3D animation production 701 includes the script writing storyboard and the scene animation production, and the initial version of the 3D animation sequence fragment (corresponding to the animation material frame in the above embodiment) is obtained. In the AI animation rendering process 702, based on the initial version of the 3D animation sequence fragment and the algorithm model (Dreambooth model and LoRA model) trained in the preparation stage, the AI "three renderings and two renderings" rendering process can be performed. First, step 706 is performed to render the key frames in each scene: character fragment rendering, background material rendering, and special effect material rendering to obtain the rendered animation picture. Then step 707 AI algorithm interpolation is performed, and step 708 the picture is layered and merged to obtain the target animation output as the final result. The embodiment of the present application can improve the generation effect and generation speed of the picture with multi-frame rendering and AI algorithm interpolation technology. Finally, the characters, background, special effects, etc. are combined to output the final result as the target animation. During the rendering process, it can be checked whether the image of the key frame obtained after rendering is of unqualified quality. If the quality is unqualified, it can be returned to the model for reprocessing.
[0225] Take the rendering process as an example: multiple frames of animation material for an animation scene can be input into the Dreambooth model, and the Dreambooth model is mounted with lora models of different characters. Assuming 1s24 frames of animation material frames, select a few frames as key frames for model rendering, such as the 3rd and 6th frames. For single-character materials, the corresponding lora model performs character rendering, and the Dreambooth model performs background rendering. Other special effects lora models perform special effects rendering, and only a small number of materials need special effects rendering. For multi-character materials, the animation material frames are divided into left and right, and the left character corresponds to lora model 1, and the right model corresponds to lora model 2 for rendering. After rendering, the characters, backgrounds, etc. are layered and merged (automatically executed by the painting software, such as directly superimposing layers), and the animation images of the 3rd and 6th frames are obtained. At this time, the AI algorithm needs to be used to fill in the 4th and 5th frames, and finally the 1-24 frames of animation images are obtained, and the target animation of the animation scene is obtained.
[0226] Fig.18 It is a flow chart of the animation generation method provided in the embodiment of the present application. In the embodiment of the present application, the animation material frames obtained after the animation production pre-preparation 700 and the 3D animation production 701 stages are input into the stable diffusion model of the AI animation rendering process 702 stage. The stable diffusion model is the Dreambooth model, which is mounted by the Lora model and the ControlNet model. The stable diffusion model includes an image encoder, an image generator, and an image decoder. In the embodiment of the present application, the capabilities of the AI painting model are used in combination with targeted model fine-tuning training to improve the final rendering effect. At the same time, the Controlnet model is used to improve the picture control capability. At the same time, the multi-frame rendering technology and AI frame supplementation technology are used to improve the stability of continuous picture generation. The following is a detailed description of each module.
[0227] Fine-tuning training based on the AI painting model. The stable diffusion model is used as the basis for training. As a type of generative model, the stable diffusion model generates high-quality images from random noise through a gradual denoising process. Dreambooth is selected as the stable diffusion model for fine-tuning technology of picture style. After training, the Dreambooth model can generate images of a specific style or a specific theme. The training process of the Dreambooth model is briefly described as follows: Forward diffusion: For each training image x0, a series of intermediate states x1, x2, ..., x are generated through the forward diffusion process. n. Noise addition: Add Gaussian noise to each intermediate state. Noise prediction: Use the noise prediction network to predict noise. Loss calculation: Calculate the mean square error (MSE) between the predicted noise and the actual noise. Parameter update: Update the model parameters using gradient descent or other optimization algorithms. Iterative training: Repeat the above steps until the loss function converges or reaches the preset training rounds. Exemplarily, in order to make the painting style of the AI model meet the needs of the director, more than 1,000 public paintings of several designated artists can be collected as training data to train the Dreambooth model, so that the overall output style of the Dreambooth model is close to the style of these artists.
[0228] For animated characters, the LoRA fine-tuning technology can be used. Compared with the Dreambooth model, the LoRA model can reduce the number of parameters for large-scale training and improve training efficiency. For example, about 40 portraits are collected for each animated character, including four parts, 10 pictures each of face close-up, chest and above close-up, half-body close-up, and knee and above close-up, as training materials. A dedicated LoRA model is trained for each animated character, and the training process is similar to that of the Dreambooth model mentioned above. After fine-tuning training, the LoRA model contains the image characteristics and details of the character, so after mounting the LoRA model, the AI painting model can generate high-quality pictures for specific characters and restore the character details.
[0229] The embodiment of the present application also achieves improved stability of continuous images. Since the stable diffusion model is a process from noise to a complete image, each step of random diffusion will lead the image to a different generation result, so it is difficult to generate a stable continuous image. The embodiment of the present application greatly improves the stability of continuous images by fine-tuning the Dreambooth model, coordinating multi-frame rendering and AI frame supplementation technology. Multi-frame rendering: In the process of generating the current frame, the pictures of the previous and next two frames are spliced on the left and right sides of the picture. Because the noise addition and denoising process of the Dreambooth model acts on the entire image. Therefore, when the current frame is generated, it will be affected by the content of the previous and next frames, which significantly improves the stability of the continuous picture. Fig.19It is a schematic diagram of multi-frame rendering provided by an embodiment of the present application. The three pictures are spliced into a complete picture and sent to the Dreambooth model for AI transfer painting, and the current frame in the middle is the target key frame picture to be generated. Exemplarily, the image editing (Inpainting) function of the user interface (WebUI) can be used to perform AI generation only for the mask part of the large image, that is, the target key frame. The left and right of the key frame picture are the previous frame picture and the next frame picture, and the previous frame picture and the next frame picture are input into the Dreambooth model as information of similar conditions. In the process of adding noise to the Dreambooth model, the features of the three pictures will gradually merge, thereby improving the consistency of AI rendering of consecutive frames. After the rendering is completed, the output large picture is segmented, and only the target picture in the middle is retained. The above process can all be implemented in code, and will not affect the use process of the animation generation application.
[0230] The embodiment of the present application also proposes an AI frame interpolation technology. Compared with generating each frame of animation image for a stable diffusion model, AI frame interpolation focuses on the migration of picture style, so the picture stability is higher. Fig. 20 Schematic diagram of the AI frame interpolation technology provided in the embodiment of the present application. When migrating the style of the original image to the target image, three corresponding conditional images need to be provided: the mask information, position information, and edge information of the original image. The image style migration is completed based on these three conditional images. Fig. 20 In the figure, (a) is the image, (b) is the mask information, (c) is the position information, (d) is the edge information, (e) and (f) are the images generated by interpolation. The original image is the initial version of the animation material frame, and the target image is the rendered image.
[0231] In the embodiment of the present application, the canny algorithm is used to generate the edge information of the original image, and the GMFlow optical flow algorithm is used as another guiding condition to generate the optical flow information of the original image. The style transfer effect is achieved based on the edge information and optical flow information, that is, AI frame filling. The generation of a single frame filling picture only takes about 3 seconds, which greatly improves the efficiency of animation production. Exemplary, for example, if the animation material frames of the 3rd and 6th frames are known, the edge information and optical flow information of each frame can be calculated based on the animation material frames of each frame in the 3rd-6th frames, and the rendered 3rd frame animation image and the 6th frame animation image, as well as the edge information and optical flow information of each frame are used as input to automatically generate the 4th frame and the 5th frame animation image. At the same time, the AI frame filling algorithm module supports server-side operation and concurrent multi-threaded processing. It can perform AI frame filling generation for multiple pictures at the same time, greatly improving rendering efficiency.
[0232] Fig.21 It is a technical flow chart of AI animation rendering provided by the embodiment of the present application. In the early stage of animation production, the embodiment of the present application simultaneously conducts technical research and determines the overall technical solution and rendering process, such as performing step 801 to organize multi-role materials, step 802 to train multi-role models, and then performing step 803 to align specific parameters based on test clips, and step 804 to organize batch rendering scripts by recording parameters for different roles. That is, train fine-tuning models of many roles in advance, determine rendering detail parameters, complete code development work for related modules, etc. After the animation production is completed, AI rendering production starts immediately, including material organization, model training, rendering parameter alignment, and the overall rendering stage. Among them, the embodiment of the present application also introduces the latent consistency model (LCM) technology, and adds the consistency constraint distillation on the basis of the stable diffusion model, which can reduce the sampling steps of the original stable diffusion model of 20-30 time steps (steps) to within 5 steps, while speeding up the speed and improving the rendering effect. The embodiment of the present application also introduces AI frame interpolation technology to improve rendering efficiency and greatly shorten the time of the rendering stage. The embodiment of the present application also writes an engineered AI rendering module code, which processes different animation clips in parallel on separate graphics cards, improves the resource utilization of a single cloud server, and speeds up the rendering speed. For example, for a 10-minute animation short, the entire rendering process can be completed by two engineers within two weeks, which means that the amount of work that would take a professional team of about ten people several months in traditional technology can be completed in 28 man-days.
[0233] In addition, the embodiments of the present application can also optimize the structure of the stable diffusion model, such as extracting the information of the previous and next sequence frames as a coding vector and inputting it into the stable diffusion model to provide stronger control capabilities (the animation material frame is first encoded into a vector, and then exchanged with the information of the original model through cross-attention calculation, which is similar to the principle of multi-frame rendering). The AI rendering code script written in the embodiments of the present application can be further optimized into a complete service, providing a program interface (API) or an operation interface, which is convenient for non-algorithm programmers to use.
[0234] In summary, the embodiment of the present application proposes a complete set of AI-assisted animation production solutions without changing the traditional animation production process. Therefore, it can be applied to all animation production processes, helping existing animation producers to introduce AIGC technology to improve efficiency and optimize processes. For the key "three-rendering-two" process in animation production, the efficient rendering capability provided by the AI painting large model can change the dependence of traditional processes on manpower and time, and provide better drawing efficiency and quality. In the early design stage of the animation, the embodiment of the present application uses the AI painting model to quickly produce design drawings of different styles, helping the director team to quickly determine the overall style and character image; in the mid-to-late stage of picture rendering, the trained AI model is used in conjunction with continuous picture stability improvement technology to greatly improve rendering efficiency and quality. At the same time, the AI rendering effect is amazing and stable. After the parameters are determined, it is not interfered by manual operation, reducing the labor cost of the project. The embodiment of the present application uses artificial intelligence algorithms to use AI models and project codes for parallel processing during rendering, which significantly improves efficiency, and reduces the production cycle by 40%-50% under AI addition. In terms of animation production costs, the traditional animation production process costs between 120,000 and 150,000 yuan per minute. With the help of AI technology, the cost per minute is saved to 50,000 to 80,000 yuan, a reduction of more than 50%.
[0235] It is understandable that in the embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0236] The following is a description of an exemplary structure of the animation generating device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 4 As shown, the software modules in the animation generating device 455 stored in the memory 450 may include: a material frame acquiring module 4551, used to acquire pre-made animation material frames; the animation materials in the animation material frames include animation characters and animation backgrounds; a model determining module 4552, used to determine the character rendering model corresponding to the animation character based on the character type of the animation character; a character rendering module 4553, used to call the character rendering model to render the animation character in the animation material frame to obtain a character image; a background rendering module 4554, used to call a preset background rendering model to render the animation background in the animation material frame to obtain a background image; an animation generating module 4555, used to perform image fusion processing on the character image and the background image to obtain a target animation containing the animation material.
[0237] In some embodiments, the material frame acquisition module 4551 is also used to respond to the received animation production request, obtain the animation style type and animation script of the target animation to be generated; for the character type and scene type in the animation script, call the drawing model to generate a scene concept map under the scene type based on the animation style type; generate three views of the animated character of the character type in different postures based on the animation style type and the preset skeleton map; generate animation material frames based on the three views of the animated character in different postures and the scene concept map under the scene type.
[0238] In some embodiments, the material frame acquisition module 4551 is also used to create three-dimensional models of animated characters and scene types based on the three-view images of the animated characters in different postures and the scene concept maps under the scene types; generate three-dimensional animation for each animation scene in the animation script based on the animation script and the three-dimensional model; the three-dimensional animation includes animated characters and animation backgrounds under the scene types; and extract two-dimensional animation material frames of the corresponding animation scene from the three-dimensional animation of each animation scene.
[0239] In some embodiments, the animation generation device 455 also includes a background rendering model training module, which is used to call multiple painting models to generate pictures of multiple picture style types respectively; wherein each painting model is used to generate multiple pictures of a picture style type; in response to the selection operation of the animation style type among the multiple picture style types, the preset diffusion model is trained using the pictures of the animation style type to obtain a preset background rendering model with the animation style type.
[0240] In some embodiments, the background rendering model training module is also used to train a preset diffusion model using an animation style type picture to obtain an initial rendering model with an animation style type; determine an edge image of the animation style type picture; use the edge image to train a preset control model to obtain an edge control model; and fuse the initial rendering model and the edge control model to obtain a preset background rendering model.
[0241] In some embodiments, the animation generation device 455 also includes a character rendering model training module, which is used to perform close-up processing of different parts of the animation character based on the three views of the animation character in different postures, and obtain multiple training material images of the animation character; use the multiple training material images of the animation character to train the preset character rendering model to be trained, and obtain the character rendering model corresponding to the animation character.
[0242] In some embodiments, the model determination module 4552 is also used to segment the animation material frame based on the animation character in the animation material frame to obtain multiple material sub-frames; wherein each material sub-frame includes an animation character; for each material sub-frame, based on the character type of the animation character in the material sub-frame, determine the character rendering model corresponding to the animation character.
[0243] In some embodiments, the pre-made animation material frames include multiple consecutive animation material frames; the character rendering module 4553 is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; splice the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; call the character rendering model to render multiple animation characters in the spliced animation material frame to obtain a character image.
[0244] In some embodiments, the character rendering module 4553 is also used to call the character rendering model to render multiple animation characters in the spliced animation material frames to obtain an initial character image; perform image segmentation on the initial character image to obtain a first character image corresponding to a previous frame of animation material, a second character image corresponding to a current animation material frame, and a third character image corresponding to a next frame of animation material; and determine the second character image as the character image.
[0245] In some embodiments, the pre-made animation material frames include multiple consecutive animation material frames; the background rendering module 4554 is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; splice the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; call the preset background rendering model to render the multiple animation backgrounds in the spliced animation material frame to obtain a background image.
[0246] In some embodiments, the background rendering module 4554 is also used to call a preset background rendering model to render multiple animation backgrounds in the spliced animation material frames to obtain an initial background image; perform image segmentation on the initial background image to obtain a first background image corresponding to the previous frame of animation material, a second background image corresponding to the current animation material frame, and a third background image corresponding to the next frame of animation material; and determine the second background image as the background image.
[0247] In some embodiments, the animation generation module 4555 is also used to determine the animation image corresponding to the animation material frame based on the character image and background image corresponding to the animation material frame; determine the edge information and optical flow information of each animation material frame from the i-th frame animation material frame to the i+j-th frame animation material frame in the pre-produced multi-frame animation material frame; i and j are both integers greater than 0; determine the i+k-th frame animation image based on the edge information and optical flow information of each animation material frame from the i-th frame animation material frame to the i+j-th frame animation material frame, as well as the i-th frame animation image corresponding to the i-th frame animation material frame and the i+j-th frame animation image corresponding to the i+j-th frame animation material frame; the value of k is any integer between 1 and j-1; splice the 1st frame animation image to the nth frame animation image into the target animation, n is the number of multi-frame animation material frames, and n is an integer greater than 1.
[0248] In some embodiments, the animation generation module 4555 is also used to extract background features of the background image to obtain image feature information of the background image; to extract character features of the character image to obtain image feature information of the character image; to generate a background layer of the background image based on the image feature information of the background image; to generate a character layer of the character image based on the image feature information of the character image; and to superimpose the character layer and the background layer to obtain an animation image corresponding to the animation material frame.
[0249] In some embodiments, the animation material in the animation material frame also includes animation special effects; the animation generation module 4555 is also used to determine the special effects rendering model corresponding to the animation special effects based on the special effects type of the animation special effects; call the special effects rendering model to render the animation special effects in the animation material frame to obtain a special effects image; perform image fusion processing on the special effects image, character image and background image to obtain a target animation containing the animation material.
[0250] In some embodiments, the pre-made animation material frame includes multiple consecutive animation material frames; the background rendering module 4554 is also used to determine the current animation material frame from the multiple animation material frames; obtain the previous animation material frame and the next animation material frame of the current animation material frame; encode the previous animation material frame, the animation material frame and the next animation material frame respectively to obtain a first material frame vector corresponding to the previous animation material frame, a second material frame vector corresponding to the current animation material frame and a third material frame vector corresponding to the next animation material frame; fuse the first material frame vector, the second material frame vector and the third material frame vector to obtain a fused material frame vector; call the background rendering model to render the animation background in the animation material frame based on the fused material frame vector to obtain a background image.
[0251] In some embodiments, the background rendering module 4554 is further used to call the background rendering model, perform attention processing on the fused material frame vector, and obtain attention features; based on the fused material frame vector and the attention features, determine the background image of the animation background.
[0252] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the animation generation method described in the embodiment of the present application.
[0253] The present application embodiment provides a computer-readable storage medium in which computer executable instructions or computer programs are stored. When the computer executable instructions or computer programs are executed by a processor, the processor will execute the animation generation method provided by the present application embodiment, for example, Figure 5 The animation generation method shown.
[0254] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0255] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0256] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0257] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0258] In summary, the multi-frame rendering technology and AI frame-filling algorithm proposed in the embodiments of the present application are used to solve the problem of poor stability of AI drawing when drawing continuous pictures. The AI painting large model fine-tuning training technology Dreambooth and LoRA model are used to train corresponding models for painting styles and characters respectively, thereby improving the final output quality. Without changing the traditional animation production process, a complete process of using AIGC technology to assist in animation production is proposed to reduce costs and increase efficiency. An engineered AI rendering module code is written so that the entire rendering process can be executed concurrently on the server side, realizing semi-automatic operation and improving animation production efficiency.
[0259] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. An animation generation method, characterized in that: The method comprises: Acquire a pre-made animation material frame; the animation material in the animation material frame includes an animation character and an animation background; Based on the character type of the animation character, determining a character rendering model corresponding to the animation character; the character type is an identifier used to distinguish the animation character, and each animation character corresponds to a character type; Calling the character rendering model to render the animation character in the animation material frame to obtain a character image; Calling a preset background rendering model to render the animation background in the animation material frame to obtain a background image; wherein the preset background rendering model is trained by the following steps: calling multiple painting models to generate pictures of multiple picture style types respectively; wherein each painting model is used to generate multiple pictures of one picture style type; in response to a selection operation of an animation style type among the multiple picture style types, using the picture of the animation style type to train a preset diffusion model to obtain an initial rendering model with the animation style type; determining an edge image of the picture of the animation style type; using the edge image to train a preset control model to obtain an edge control model; fusing the initial rendering model with the edge control model to obtain the preset background rendering model; Performing image fusion processing on the character image and the background image to obtain a target animation containing the animation material; The image fusion processing of the character image and the background image to obtain the target animation containing the animation material includes: determining the animation image of the animation material frame based on the character image and the background image in the animation material frame, and splicing the animation images of multiple frames of animation material frames in sequence to obtain the target animation containing the animation material.
2. The method according to claim 1, characterized in that The step of obtaining a pre-made animation material frame includes: In response to the received animation production request, obtaining an animation style type and an animation script of a target animation to be generated; For the character type and scene type in the animation script, calling the drawing model to generate a scene concept map under the scene type based on the animation style type; Generate three views of the animation character of the character type in different postures based on the animation style type and the preset skeleton diagram; The animation material frames are generated based on the three views of the animation character in different postures and the scene concept map of the scene type.
3. The method according to claim 2, characterized in that The generating the animation material frame based on the three views of the animation character in different postures and the scene concept map under the scene type includes: Creating a three-dimensional model of the animation character and the scene type based on the three-view drawing of the animation character in different postures and the scene concept drawing of the scene type; Based on the animation script and the three-dimensional model, a three-dimensional animation of each animation scene in the animation script is generated; the three-dimensional animation includes the animation character and the animation background under the scene type; From the three-dimensional animation of each animation scene, the two-dimensional animation material frames of the corresponding animation scene are extracted.
4. The method according to claim 2, characterized in that: The method further comprises: Based on the three views of the animated character in different postures, close-up processing of different parts of the animated character is performed to obtain a plurality of training material images of the animated character; A plurality of training material images of the animation character are used to train a preset character rendering model to be trained, so as to obtain a character rendering model corresponding to the animation character.
5. The method according to claim 1, characterized in that The step of determining a character rendering model corresponding to the animation character based on the character type of the animation character comprises: Segmenting the animation material frame based on the animation character in the animation material frame to obtain a plurality of material sub-frames; wherein each material sub-frame includes an animation character; For each material sub-frame, a character rendering model corresponding to the animation character is determined based on the character type of the animation character in the material sub-frame.
6. The method according to claim 1, characterized in that The pre-made animation material frames include a plurality of continuous animation material frames; and calling the character rendering model to render the animation character in the animation material frames to obtain the character image includes: Determine a current animation material frame from the plurality of animation material frames; Obtaining a previous animation material frame and a next animation material frame of the current animation material frame; Splicing the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; The character rendering model is called to render a plurality of animation characters in the spliced animation material frame to obtain the character image.
7. The method according to claim 6, characterized in that The calling of the character rendering model to render the multiple animation characters in the spliced animation material frame to obtain the character image includes: Calling the character rendering model to render multiple animation characters in the spliced animation material frame to obtain an initial character image; Performing image segmentation on the initial character image to obtain a first character image corresponding to the previous animation material frame, a second character image corresponding to the current animation material frame, and a third character image corresponding to the next animation material frame; The second character image is determined as the character image.
8. The method according to claim 1, characterized in that The pre-made animation material frames include a plurality of continuous animation material frames; The calling of a preset background rendering model to render the animation background in the animation material frame to obtain a background image includes: Determine a current animation material frame from the plurality of animation material frames; Obtaining a previous animation material frame and a next animation material frame of the current animation material frame; Splicing the previous animation material frame, the current animation material frame and the next animation material frame to obtain a spliced animation material frame; A preset background rendering model is called to render multiple animation backgrounds in the spliced animation material frame to obtain the background image.
9. The method according to claim 8, characterized in that The calling of a preset background rendering model to render a plurality of animation backgrounds in the spliced animation material frame to obtain the background image includes: Calling a preset background rendering model to render multiple animation backgrounds in the spliced animation material frame to obtain an initial background image; Performing image segmentation on the initial background image to obtain a first background image corresponding to the previous animation material frame, a second background image corresponding to the current animation material frame, and a third background image corresponding to the next animation material frame; The second background image is determined as the background image.
10. The method according to claim 1, characterized in that The performing image fusion processing on the character image and the background image to obtain a target animation containing the animation material includes: Determining the animation image corresponding to the animation material frame based on the character image and the background image corresponding to the animation material frame; Determine edge information and optical flow information of each animation material frame from the i-th animation material frame to the i+j-th animation material frame in the pre-made multi-frame animation material frame; i and j are both integers greater than 0; Based on the edge information and optical flow information of each animation material frame from the i-th animation material frame to the i+j-th animation material frame, the i-th animation image corresponding to the i-th animation material frame, and the i+j-th animation image corresponding to the i+j-th animation material frame, determining the i+k-th animation image; the value of k is any integer between 1 and j-1; The first frame of animation image to the nth frame of animation image are spliced into the target animation, where n is the number of frames of the multi-frame animation material and n is an integer greater than 1.
11. The method according to claim 10, characterized in that The step of determining the animation image corresponding to the animation material frame based on the character image and the background image corresponding to the animation material frame comprises: Extracting background features from the background image to obtain image feature information of the background image; Extracting character features from the character image to obtain image feature information of the character image; Based on the image feature information of the background image, generating a background layer of the background image; Based on the image feature information of the character image, generating a character layer of the character image; The character layer and the background layer are superimposed to obtain an animation image corresponding to the animation material frame.
12. The method according to claim 1, characterized in that The animation material in the animation material frame also includes animation special effects; The performing image fusion processing on the character image and the background image to obtain a target animation containing the animation material includes: Based on the special effect type of the animation special effect, determining a special effect rendering model corresponding to the animation special effect; Calling the special effect rendering model to render the animation special effects in the animation material frame to obtain a special effect image; The special effect image, the character image and the background image are subjected to image fusion processing to obtain a target animation containing the animation material.
13. The method according to claim 1, characterized in that The pre-made animation material frames include a plurality of continuous animation material frames; The calling of a preset background rendering model to render the animation background in the animation material frame to obtain a background image includes: Determine a current animation material frame from the plurality of animation material frames; Obtaining a previous animation material frame and a next animation material frame of the current animation material frame; Encoding the previous animation material frame, the animation material frame and the next animation material frame respectively to obtain a first material frame vector corresponding to the previous animation material frame, a second material frame vector corresponding to the current animation material frame and a third material frame vector corresponding to the next animation material frame; performing fusion processing on the first material frame vector, the second material frame vector and the third material frame vector to obtain a fused material frame vector; The background rendering model is called to render the animation background in the animation material frame based on the fused material frame vector to obtain the background image.
14. The method according to claim 13, characterized in that The calling of the background rendering model to render the animation background in the animation material frame based on the fused material frame vector to obtain the background image includes: Calling the background rendering model, performing attention processing on the fused material frame vector, and obtaining an attention feature; Based on the fused material frame vector and the attention feature, a background image of the animation background is determined.
15. An animation generating device, characterized in that: The device comprises: A material frame acquisition module is used to acquire pre-made animation material frames; the animation materials in the animation material frames include animation characters and animation backgrounds; A model determination module, used to determine a character rendering model corresponding to the animation character based on the character type of the animation character; the character type is an identifier used to distinguish the animation character, and each animation character corresponds to a character type; A character rendering module, used for calling the character rendering model to render the animation character in the animation material frame to obtain a character image; A background rendering module, used for calling a preset background rendering model to render the animation background in the animation material frame to obtain a background image; A background rendering model training module is used to call multiple painting models to generate pictures of multiple picture style types respectively; wherein each painting model is used to generate multiple pictures of one picture style type; in response to a selection operation of an animation style type among the multiple picture style types, a preset diffusion model is trained using the picture of the animation style type to obtain an initial rendering model of the animation style type; an edge image of the picture of the animation style type is determined; a preset control model is trained using the edge image to obtain an edge control model; the initial rendering model and the edge control model are merged to obtain the preset background rendering model; The animation generation module is used to perform image fusion processing on the character image and the background image to obtain a target animation containing the animation material, and is also used to determine the animation image of the animation material frame based on the character image and the background image in the animation material frame, and to splice the animation images of multiple frames of animation material frames in sequence to obtain a target animation containing the animation material.
16. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; The processor is used to implement the animation generation method described in any one of claims 1 to 14 when executing the computer executable instructions or computer programs stored in the memory.
17. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the animation generation method according to any one of claims 1 to 14 is implemented.
18. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the animation generation method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
Rendering method of spatial model, electronic equipment and readable storage medium
CN114463468A
Dynamic cartoon generation method and device, storage medium and electronic equipment
CN117252966A
Animation display method, device and equipment and computer readable storage medium
CN117876540A