Image generation method and device, equipment and medium

By obtaining the color fill and line image collection of the target character, as well as the multi-frame image of three-dimensional angles, and directly generating multi-angle color fill images using the animation generation model, the problem of low efficiency of three-dimensional animation generation in the existing technology is solved, and more efficient animation generation is achieved.

CN120451349APending Publication Date: 2025-08-08BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510360169.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When converting two-dimensional animations into three-dimensional animations, the prior art requires fine modeling and binding, resulting in low efficiency in generating three-dimensional animations.

Method used

A first image set including the color fill image and line image of the target character, and a second image set of multiple frames corresponding to the multiple three-dimensional angles of the target character are obtained, and a request is sent to the target animation generation model, and a multi-frame video frame is generated using the model, and a multi-angle color fill image of the target character is included.

Benefits of technology

By guiding the three-dimensional animation generation model, the multi-angle color fill image of the target character is directly generated, which improves the animation generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451349A_ABST
    Figure CN120451349A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an image generation method and device, equipment and a medium, and the method comprises the steps: obtaining a first image set, the first image set comprises at least one frame of first image, and the at least one frame of first image comprises a color filling image containing a target role, and / or a line image containing the target role; a second image set is obtained, and the second image set comprises multiple frames of second images corresponding to the multiple three-dimensional angles of the target role; a first animation generation request is sent to a target animation generation model, the first animation generation request carries a first image set and a second image set, and the target animation generation model learns in advance to generate multiple video frames according to the input image set in response to the animation generation request; multiple video frames output by the target animation generation model are obtained, and the multiple video frames comprise multi-angle color filling images of the target role. According to the technical scheme, the animation generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer application technology, and in particular to an image generation method, apparatus, device, and medium. Background Art

[0002] Currently, when generating a three-dimensional animation, a two-dimensional animation representing the animation plot is first obtained, and then the two-dimensional animation is converted into a three-dimensional animation.

[0003] In related technologies, converting a 2D character animation into a 3D animation requires sophisticated modeling and rigging. Specifically, the 3D model of the character is constructed, the 3D model is adjusted based on the angle of the character in each 2D image frame of the 2D animation, and an image of the 3D model at the corresponding angle is obtained. The 2D image is then replaced with the image at the corresponding angle to generate the 3D animation. This results in low 3D animation generation efficiency. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image generation method, apparatus, device and medium.

[0005] An embodiment of the present disclosure provides an image generation method, which includes: acquiring a first image set, wherein the first image set contains at least one frame of first image, and the at least one frame of first image includes: a color-filled image of a target character, and / or a line image of the target character; acquiring a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character; sending a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model pre-learns to generate multiple video frames based on the input image set in response to the animation generation request; acquiring multiple video frames output by the target animation generation model, wherein the multiple video frames include color-filled images of the target character from multiple angles.

[0006] An embodiment of the present disclosure also provides an image generating device, which includes: a first acquisition module for acquiring a first image set, wherein the first image set contains at least one frame of first image, and the at least one frame of first image includes: a color-filled image containing a target character, and / or a line image containing the target character; a second acquisition module for acquiring a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character; a request processing module for sending a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model pre-learns to generate multiple video frames based on the input image set in response to the animation generation request; a third acquisition module for acquiring multiple video frames output by the target animation generation model, wherein the multiple video frames contain color-filled images of the target character from multiple angles.

[0007] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the image generation method provided by the embodiment of the present disclosure.

[0008] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the image generation method provided by the embodiment of the present disclosure.

[0009] The embodiment of the present disclosure further provides a computer-readable storage medium, which is used to store multiple video frames generated after executing the above-mentioned image generation method.

[0010] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art:

[0011] The animation generation scheme proposed in the embodiment of the present disclosure obtains a first image set, wherein the first image set contains at least one frame of the first image, and the at least one frame of the first image includes: a color-filled image containing a target character, and / or a line image containing a target character, and then obtains a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character, and sends a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model pre-learns to generate multiple frames of video frames based on the input image set in response to the animation generation request, and obtains multiple frames of video frames output by the target animation generation model, wherein the multiple frames of video frames include color-filled images of the target character from multiple angles. In this technical solution, the first image set is input into the target animation generation model, and the multiple frames of the second image corresponding to the multiple angles of the three-dimensional animation model of the target character are used as a guide to directly generate color-filled images of the target character from multiple angles, which can further improve the efficiency of animation generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0013] Figure 1 A schematic diagram of a flow chart of an image generation method provided in an embodiment of the present disclosure;

[0014] Figure 2 A schematic diagram of an animation generation scene provided by an embodiment of the present disclosure;

[0015] Figure 3 A schematic structural diagram of an image generating device provided in an embodiment of the present disclosure;

[0016] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0019] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] In order to solve the above problems, the embodiments of the present disclosure provide an image generation method, which is introduced below in conjunction with specific embodiments.

[0024] Figure 1 This is a flow chart of an image generation method provided by an embodiment of the present disclosure. The method can be executed by an image generation device, wherein the device can be implemented using software and / or hardware and can generally be integrated into an electronic device. Figure 1 As shown, the method includes:

[0025] Step 101: Acquire a first image set, wherein the first image set includes at least one first image frame, and the at least one first image frame includes: a color-filled image including a target character, and / or a line image including a target character.

[0026] Among them, the first image set contains multiple frames of first images, and each frame of the first image only needs to contain partial image information of the target character. The image information includes but is not limited to the line drawing image information, color filling image information, etc. of the target character. That is, the first image is not required to contain a comprehensive image of the target character, but only needs to contain partial image information, which initially improves the efficiency of animation generation.

[0027] In some possible embodiments, the multiple frames of first images included in the first image set include: when the multiple frames of first images include a color-filled image of a target character, the background of the first image is filled with a preset color, wherein the image information of the target character contained in each frame of the first image is the color-filled image information of the target character, thereby obtaining multiple frames of two-dimensional images drawn by relevant artists, and in order to avoid the influence of the image content contained in the background, the background colors of the multiple frames of two-dimensional images are replaced with preset colors to obtain multiple frames of first images, wherein the color filled in the color-filled image can be partially filled or fully filled, and the filled color can be coarse-grained filling (for example, only the background color of the hair area is filled, but the detailed color of the hair part is not filled), or it can be fine-grained filling.

[0028] In some possible embodiments, when at least one first image frame includes a line drawing of a target character, the first image includes a line drawing outline of the target character. The line drawing outline is used to outline the contour, depict internal structure, or express three-dimensionality in the form of lines. The line color and line thickness corresponding to the line drawing outline are not restricted. That is, in this example, the coloring process is not required, and the first image may only include the line drawing outline of the target character. The line drawing outline can also be obtained from a color-filled image containing the target character, or can be drawn by relevant artists.

[0029] Based on the above description, it can be seen that the first image set only needs to include partial image information of the target character, which greatly reduces the workload of generating the first image and helps to further improve the efficiency of animation generation.

[0030] Step 102: Acquire a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character.

[0031] In some possible embodiments, the preset database may include model images of multiple three-dimensional angles of pre-generated three-dimensional animation models of various characters, wherein the three-dimensional animation model may be a three-dimensional model of the target character pre-built through techniques such as skeleton construction and mapping, or may be a Lora model of the character, etc. The Lora model is obtained by fine-tuning a general 3D modeling or motion capture model to optimize it specifically for animated characters. As a result, the generation efficiency of the Lora model is higher and the algorithm consumption is smaller. When training the Lora model, the training of the three-dimensional model of the animated character can be carried out, etc.

[0032] Therefore, in this embodiment, the second image set can be obtained by reading the preset database. The second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character.

[0033] Step 103: Send a first animation generation request to the target animation generation model, wherein the first animation generation request carries a first image set and a second image set, and the target animation generation model is pre-learned to generate multiple video frames according to the input image set in response to the animation generation request.

[0034] Among them, the target animation generation model can be understood as any artificial intelligence model in the existing technology that can be used for animation generation, including but not limited to SDXL tools, etc., wherein SDXL can integrate AnimateDiff plug-ins, etc. In this embodiment, the target animation generation model is pre-learned to generate an animation video stream according to an input image set in response to an animation generation request, wherein the input image set may include the above-mentioned first image set and second image set.

[0035] In an embodiment of the present disclosure, a first animation generation request is sent to a target animation generation model, wherein the first animation generation request carries a first image set and a second image set, that is, the first image set is used as a guide and combined with the second image set to generate multiple video frames with better quality.

[0036] During the actual execution process, the target animation generation model can ensure the consistency of the generated target character based on the second image set, and this consistency includes color consistency, contour consistency, etc. The first image set can serve as a guide, and the target animation generation model determines the generated animation plot from the image content contained in the first image set, and generates a corresponding animation video stream based on the animation plot. The target animation generation model can also assist in determining the fill color and / or contour information of the target character based on the first image set. Among them, the animation plot is used to indicate the action sequence completed by the target character. For example, when the action plot is "kicking a ball", the animation plot is used to indicate the target character to complete the action sequence of "kicking a ball". For another example, when the action plot is "running", the animation plot is used to indicate the target character to complete the action sequence of "running".

[0037] In one embodiment of the present disclosure, animation plot description text information can also be obtained, wherein the animation plot description text information can be written by relevant technical personnel, and the animation plot description text information can also be carried in the first animation generation request, wherein the target animation generation model also pre-learns to generate multiple video frames in response to the animation generation request based on the input image set and the animation plot description text information, wherein the multiple video frames output by the target animation generation model include multiple video frames corresponding to the action sequence. In this embodiment, the animation plot description text is used to indicate the action sequence of the target character, for example, if the animation plot description text is "the target character completes the action of kicking the ball", then the corresponding animation plot description text is used to indicate that the target character completes the action sequence of "kicking the ball", and for another example, if the animation plot description text is "the target character completes the action of kicking the ball to score a goal", then the corresponding animation plot description text is used to indicate that the target character completes the action sequence of "kicking the ball to score a goal".

[0038] Step 104 : obtaining a plurality of video frames output by the target animation generation model, wherein the plurality of video frames include color-filled images of the target character from multiple angles.

[0039] In an embodiment of the present disclosure, a multi-angle color-filled image of a target character output by a target animation generation model is obtained, and the multi-angle is a multi-three-dimensional angle, wherein the number of the multi-angle color-filled images may be greater than the number of the input first images, that is, the target animation generation model may generate a plurality of video frames reflecting the animation plot based on at least one frame of the input first image, wherein the video frame may include video frames corresponding to the transition plot between different first images, etc., wherein the animation plot of the plurality of video frames may be determined according to the text information describing the animation plot, or may be obtained by learning the image content of the plurality of first images, etc.

[0040] It should be emphasized that since the video frames are generated by the target animation generation model combined with multiple frames of second images corresponding to multiple three-dimensional angles, the multiple frames of second images also include the target characters from multiple angles. The multiple frames of second images including the target characters from multiple angles can be regarded as the initial three-dimensional animation containing the target characters.

[0041] In one embodiment of the present disclosure, in order to further enhance the animation effect, after obtaining multiple frames of second images, a second animation generation request carrying multiple frames of second images can be sent to a target animation generation model to obtain updated multiple frames of second images output by the target animation generation model, that is, to generate the updated multiple frames of second images again, thereby enhancing the 3D effect of the initial three-dimensional animation.

[0042] During the actual execution process, the updated three-dimensional animation can be generated in a loop multiple times. Each time the loop is completed, the target animation generation model can generate a related second image in combination with the multi-frame second images corresponding to the multiple angles of the three-dimensional animation model. The newly generated second image can include the transition image between the original second images, and the second image contains more comprehensive image information of the target character from various angles, thereby improving the 3D effect.

[0043] In one embodiment of the present disclosure, in order to further improve the animation generation effect, after generating multiple video frames, it is also possible to identify whether each video frame in the multiple video frames needs to be regenerated, that is, to perform effect detection on a single video frame, and when there is a video frame that needs to be regenerated in the multiple video frames, an image repair request is sent to the target animation generation model, wherein the image repair request carries the video frame that needs to be regenerated, wherein the target animation generation model pre-learns to obtain an updated image output by the target animation generation model in response to the image repair request and regenerates the updated image according to the input image, and replaces the video frame that needs to be regenerated according to the updated video frame, and the image quality of the regenerated video frame is higher than the image quality of the input video frame that needs to be regenerated, that is, in an embodiment, the target animation generation model also has the function of repairing the image, for example, the target animation generation model includes multiple neural network controlnets for enhanced image generation, and the neural network controlnet based on enhanced image generation can enhance the effect of the input video frame.

[0044] It should be noted that the method of identifying whether each video frame in multiple video frames needs to be regenerated is different in different scenarios. Examples are as follows:

[0045] In some possible embodiments, an image quality parameter value of each video frame is identified, and video frames whose image quality parameter values are determined to be less than a preset image quality parameter threshold are video frames that need to be regenerated. The image quality parameters may include image clarity, contrast, resolution, etc.

[0046] In some possible embodiments, it is determined whether the moving distance between each limb key point position of the target character contained in each frame of video and the same limb key point position in an adjacent video frame is less than a preset distance threshold, and it is identified whether the number of limb key point positions less than the preset distance threshold in each frame of video is greater than a preset number threshold, wherein the video frame in which the number of limb key point positions less than the preset distance threshold is greater than the preset number threshold is the video frame that needs to be regenerated.

[0047] That is, when there is a video frame in which the number of limb key point positions less than the preset distance threshold is greater than the preset number threshold, it is considered that the character action data contained in the video frame and the adjacent video frame are incoherent, that is, the action between the adjacent video frames is incoherent, wherein the adjacent video frames can be the video frames adjacent to the front of each video frame, or the video frames adjacent to the back of each video frame. For example, the action type corresponding to the character action data of the i-th video frame is walking, and the action type corresponding to the character action data of the i+1-th video frame is jumping. Since the character action type has suddenly changed from the i-th video frame to the i+1-th video frame, the number of limb key point positions less than the preset distance threshold between the i-th video frame and the i+1-th video frame is greater than the preset number threshold, and the action between the i-th video frame and the i+1-th frame is incoherent, and therefore, it is considered that the i-th video frame needs to be regenerated, wherein the regenerated video frame can be one frame or multiple frames.

[0048] In some possible embodiments, it is identified whether the action type of the target character in each video frame does not belong to a preset action type set, wherein the video frame whose action type of the target character does not belong to the preset action type set is a video frame that needs to be regenerated.

[0049] Among them, the preset action type set includes the preset action types that can be performed by the target character. For example, when the target character is a child, the preset action type set includes: crawling, jumping, riding a bicycle, running, etc. When the target character is an elderly person, the preset action type set includes: using a cane, walking, etc.

[0050] It should also be noted here that after each generation of multi-frame video frames, the proportion of multi-frame video frames that need to be regenerated in the multi-frame video frames to the total multi-frame video frames is also detected. When the proportion is greater than the preset proportion threshold, it is determined that a second animation generation request carrying multi-frame video frames needs to be sent to the target animation generation model again to obtain the updated multi-frame video frames output by the target animation generation model, until the proportion of multi-frame video frames that need to be regenerated in the generated multi-frame video frames to the total multi-frame video frames is no greater than the preset proportion threshold.

[0051] Furthermore, in some possible examples, a background image corresponding to each video frame is determined, and the background image may be drawn by relevant art personnel, or generated by a target animation generation model, etc. The background images corresponding to different video frames may be the same or different, and the display position of the target character is determined in the background image, wherein the display position may be specified by relevant animation generation personnel, or determined by a relevant model based on the action type of the target character contained in each video frame in the initial animation video stream. For example, if the action type of the target task is running, the empty area in the background image may be determined as the corresponding display position, etc. The target character contained in the corresponding video frame is displayed at the corresponding display position in the background image corresponding to each video frame to generate an animation video stream corresponding to the target character.

[0052] In this embodiment, if other animated characters are also included in the same background image, after displaying the other animated characters in the corresponding background image, the background image (including the other animated characters) can be merged with the same background image (including the target character) to obtain an animation video stream.

[0053] In some possible examples, the background of the video frame may be directly replaced with a preset background image to obtain a corresponding animation video stream.

[0054] In order to enable those skilled in the art to have a clearer understanding of the image generation method of the embodiment of the present disclosure, the following examples are given in conjunction with specific embodiments, wherein, in this embodiment, reference is made to Figure 2 The target character is a "little boy", and the first image set includes video 1 and video 2, wherein video 1 includes multiple frames of color-filled images, wherein the background of the color-filled image is filled with a preset color, and the image information of the target character contained in each frame of the color-filled image is the color-filled image information of the target character, and video 2 includes multiple frames of line draft images corresponding to the multiple frames of color-filled images, wherein the image information of the target character contained in each frame of the line draft image is the line draft image information of the target character contained in the corresponding color-filled image.

[0055] In this embodiment, multiple frames of second images at multiple three-dimensional angles corresponding to the three-dimensional model of the "little boy" are also obtained from a preset database, and a first animation generation request is sent to the target animation generation model, wherein the first animation generation request carries video 1, video 2 and multiple frames of second images at multiple three-dimensional angles, and multiple frames of video frames output by the target animation generation model are obtained.

[0056] In this example, in order to further improve the animation effect, a second animation generation request carrying multiple video frames may be sent to the target animation generation model to obtain the updated multiple video frames output by the target animation generation model.

[0057] After obtaining the updated multi-frame video frames, it is also detected whether each frame of the multi-frame video frames needs to be regenerated. When there are video frames that need to be regenerated in the multi-frame video frames, an image repair request is sent to the target animation generation model to obtain the updated video frames output by the target animation generation model, and the video frames that need to be regenerated are replaced according to the updated video frames to obtain the updated multi-frame video frames again.

[0058] Furthermore, in this embodiment, an animation video stream is generated based on the updated multiple video frames. For example, the background image of each updated video frame is obtained, and the display position of the target character is determined in the background image. The target character contained in the corresponding video frame is displayed at the corresponding display position in the background image corresponding to each video frame to generate an animation video stream corresponding to the target character. Figure 2 In the generated animation video stream, in addition to the "little boy", the "little bear" animation character is also included. The display process of the "little bear" animation character in the background image is the same as that of the "little boy".

[0059] In summary, the image generation method of the embodiment of the present disclosure obtains a first image set, wherein the first image set contains at least one frame of the first image, and the at least one frame of the first image includes: a color-filled image containing a target character, and / or a line image containing a target character, and then obtains a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character, and sends a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model pre-learns to generate multiple frames of video frames based on the input image set in response to the animation generation request, and obtains multiple frames of video frames output by the target animation generation model, wherein the multiple frames of video frames include color-filled images of the target character from multiple angles. In this technical solution, the first image set is input into the target animation generation model, and the multiple frames of the second image corresponding to the multiple angles of the three-dimensional animation model of the target character are used as a guide to directly generate color-filled images of the target character from multiple angles, which can further improve the efficiency of animation generation.

[0060] In order to implement the above embodiment, the present disclosure also proposes an image generation device. The device can be implemented by software and / or hardware, and can generally be integrated into an electronic device to perform animation generation processing. Figure 3 As shown, the device includes: a first acquisition module 310, a second acquisition module 320, a request processing module 330, and a third acquisition module 340, wherein:

[0061] A first acquisition module 310 is configured to acquire a first image set, wherein the first image set includes at least one first image frame, and the at least one first image frame includes: a color-filled image including a target character, and / or a line image including a target character;

[0062] A second acquisition module 320 is configured to acquire a second image set, wherein the second image set includes a plurality of frames of second images corresponding to a plurality of three-dimensional angles of the target character;

[0063] a request processing module 330 configured to send a first animation generation request to a target animation generation model, wherein the first animation generation request carries a first image set and a second image set, and the target animation generation model is pre-learned to generate a plurality of video frames based on the input image set in response to the animation generation request;

[0064] The third acquisition module 340 is configured to acquire multiple video frames output by the target animation generation model, wherein the multiple video frames contain color-filled images of the target character from multiple angles.

[0065] In one embodiment of the present disclosure, when at least one frame of the first image includes a color-filled image of the target character, the background of the first image is filled with a preset color.

[0066] In one embodiment of the present disclosure, when at least one frame of the first image includes a line image of the target character, the first image includes a line drawing outline of the target character.

[0067] In one embodiment of the present disclosure, the first animation generation request further includes: animation plot description text information, wherein the animation plot description text information is used to indicate an action sequence of the target character;

[0068] The target animation generation model is also pre-learned to generate multiple video frames based on the input image set and animation plot description text information in response to the animation generation request, wherein,

[0069] The multi-frame video frames output by the target animation generation model include multi-frame video frames corresponding to the action sequence.

[0070] In one embodiment of the present disclosure, it further includes: a first update processing module, which is used to send a second animation generation request carrying multiple video frames to the target animation generation model, and obtain the updated multiple video frames output by the target animation generation model.

[0071] In one embodiment of the present disclosure, the system further includes: a second update processing module configured to:

[0072] Identify whether each video frame in the plurality of video frames needs to be regenerated;

[0073] When there is a video frame that needs to be regenerated in the multiple video frames, an image restoration request is sent to the target animation generation model, wherein the image restoration request carries the video frame that needs to be regenerated, wherein,

[0074] The target animation generation model is pre-learned to regenerate an updated image based on an input image in response to an image restoration request;

[0075] Obtain the updated video frames output by the target animation generation model, and replace the video frames that need to be regenerated according to the updated video frames.

[0076] In one embodiment of the present disclosure, the second update processing module is used to determine whether the movement distance between each limb key point position of the target character contained in each video frame and the same limb key point position in the adjacent video frame is less than a preset distance threshold,

[0077] Identifying whether the number of limb key point positions less than a preset distance threshold in each video frame is greater than a preset number threshold, wherein the video frame in which the number of limb key point positions less than the preset distance threshold is greater than the preset number threshold is a video frame that needs to be regenerated; and / or,

[0078] It is identified whether the action type of the target character in each video frame does not belong to a preset action type set, wherein the video frame whose action type of the target character does not belong to the preset action type set is a video frame that needs to be regenerated.

[0079] In one embodiment of the present disclosure, the present invention further includes an animation generation module for:

[0080] Get the background image corresponding to each video frame;

[0081] Determine the display position of the target character in the background image;

[0082] The target character included in the video frame corresponding to each frame of the background image is displayed at the corresponding display position in each frame of the background image to generate an animation video stream corresponding to the target character.

[0083] In order to implement the above embodiments, the present disclosure further proposes a computer program product, including a computer program / instruction, which implements the image generation method in the above embodiments when executed by a processor.

[0084] In order to implement the above embodiments, the present disclosure further proposes a computer-readable storage medium, which stores a computer program for executing the above image generation method.

[0085] In order to implement the above embodiment, the present disclosure further proposes a computer-readable storage medium, which is used to store multiple video frames generated after executing the above image generation method.

[0086] In order to implement the above embodiments, the present disclosure also provides an electronic device.

[0087] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0088] The following specific reference Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0089] like Figure 4 As shown, the electronic device 400 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a memory 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processor 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0090] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0091] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 409, or installed from the memory 408, or installed from the ROM 402. When the computer program is executed by the processor 401, the above-mentioned functions defined in the image generation method of the embodiment of the present disclosure are performed.

[0092] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0093] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0094] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0095] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device executes the image generating method.

[0096] The electronic device may write computer program code for performing the operations of the present disclosure in one or more programming languages, or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0098] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0099] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0100] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0101] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0102] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0103] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An image generation method, characterized in that: include: Acquire a first image set, wherein the first image set includes at least one first image frame, and the at least one first image frame includes: a color-filled image including a target character, and / or a line image including the target character; Acquire a second image set, wherein the second image set includes multiple frames of second images corresponding to multiple three-dimensional angles of the target character; Sending a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model is pre-learned to generate multiple video frames based on the input image set in response to the animation generation request; A plurality of video frames output by the target animation generation model is obtained, wherein the plurality of video frames contain multi-angle color-filled images of the target character.

2. The method according to claim 1, wherein When the at least one frame of the first image includes a color-filled image of the target character, the background of the first image is filled with a preset color.

3. The method according to claim 1, wherein When the at least one frame of the first image includes the line image of the target character, the first image includes the line drawing outline of the target character.

4. The method according to claim 1, wherein The first animation generation request further includes: animation plot description text information, wherein the animation plot description text information is used to indicate the action sequence of the target character; The target animation generation model is also pre-learned to generate multiple video frames based on an input image set and animation plot description text information in response to an animation generation request, wherein: The multiple video frames output by the target animation generation model include multiple video frames corresponding to the action sequence.

5. The method according to claim 1, wherein The method further comprises: A second animation generation request carrying the multiple video frames is sent to the target animation generation model to obtain the updated multiple video frames output by the target animation generation model.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Identifying whether each of the multiple video frames needs to be regenerated; When there is a video frame that needs to be regenerated in the multiple video frames, an image restoration request is sent to the target animation generation model, wherein the image restoration request carries the video frame that needs to be regenerated, wherein, The target animation generation model is pre-learned to regenerate an updated image based on an input image in response to an image restoration request; An updated video frame output by the target animation generation model is obtained, and the video frame that needs to be regenerated is replaced according to the updated video frame.

7. The method according to claim 6, wherein The identifying whether each of the multiple video frames needs to be regenerated includes: Determine whether the movement distance between each limb key point position of the target character contained in each video frame and the same limb key point position in an adjacent video frame is less than a preset distance threshold, Identifying whether the number of limb key point positions less than the preset distance threshold in each video frame is greater than a preset number threshold, wherein the video frame in which the number of limb key point positions less than the preset distance threshold is greater than the preset number threshold is the video frame that needs to be regenerated; and / or, Identify whether the action type of the target character in each frame of the video frame does not belong to a preset action type set, wherein the video frame whose action type of the target character does not belong to the preset action type set is the video frame that needs to be regenerated.

8. The method according to claim 1, wherein The method further comprises: Obtaining a background image corresponding to each video frame; Determining a display position of the target character in the background image; The target character included in the video frame corresponding to each frame of the background image is displayed at the corresponding display position in each frame of the background image to generate an animation video stream corresponding to the target character.

9. An image generating device, characterized in that: include: A first acquisition module is configured to acquire a first image set, wherein the first image set includes at least one first image frame, and the at least one first image frame includes: a color-filled image including a target character, and / or a line image including the target character; A second acquisition module is configured to acquire a second image set, wherein the second image set includes a plurality of frames of second images corresponding to a plurality of three-dimensional angles of the target character; a request processing module, configured to send a first animation generation request to a target animation generation model, wherein the first animation generation request carries the first image set and the second image set, and the target animation generation model is pre-learned to generate multiple video frames based on the input image set in response to the animation generation request; The third acquisition module is configured to acquire multiple video frames output by the target animation generation model, wherein the multiple video frames include color-filled images of the target character from multiple angles.

10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the image generation method described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used to execute the image generation method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store multiple video frames generated after executing the image generation method described in any one of claims 1 to 8.

Citation Information

Cited By

  • Long video generation method and system based on dynamic global local memory mechanism

    CN120976355A

  • Long video generation method and system based on dynamic global-local memory mechanism

    CN120976355B

  • Image processing method, electronic equipment, readable storage medium and program product

    CN122023555A