Method for creating animation

The animation creation system uses generative AI to generate frame images from rough line drawings, addressing the challenge of manual transcription in existing methods, enabling efficient and consistent animation production.

JP2026000807APending Publication Date: 2026-01-06BALUS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024098374
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing animation creation methods, such as those described in Patent Document 1, require manual transcription of user-provided images into predetermined movements, lacking a mechanism for easily creating new animations.

Method used

An animation creation system utilizing generative AI models to generate multiple frame images from rough line drawings, combining them with background images to create animations, including features like character appearance and style settings, enabling easy creation of new animations.

Benefits of technology

Facilitates the easy creation of new animations by automating the process, reducing manual effort and enhancing consistency across scenes, allowing for efficient production of high-quality animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026000807000001_ABST
    Figure 2026000807000001_ABST
Patent Text Reader

Abstract

To provide a mechanism for easily creating a new animation.SOLUTION: An animation creation method includes acquiring a plurality of pieces of rough line drawing information, generating a plurality of pieces of frame information about a character by one or a plurality of image generation models in which information about the character is set on the basis of the plurality of pieces of rough line drawing information, and generating moving image information on the basis of the plurality of pieces of frame information. Generating the plurality of pieces of frame information about the character includes generating a plurality of pieces of intermediate image information about the character by a third image generation model in which information about the character is set based on the plurality of pieces of rough line drawing information, generating a plurality of pieces of line drawing information about the character by a fourth image generation model that generates a line drawing based on the plurality of pieces of intermediate image information about the character, and generating a plurality of pieces of frame information by the second image generation model based on the plurality of pieces of line drawing information about the character.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to an animation creation method. [Background technology]

[0002] JP 2023-002280 A (Patent Document 1) is a background technology in this technical field. This publication states that "a moving image data creation device includes: a storage unit that stores format data including a background image and character movement information for each of a plurality of scenes of a work; an image processing unit that receives user images from a user terminal, cuts out a person image from the user image, animates the person image based on the movement information, combines the animated person image with the background image, and generates scene editing data corresponding to each of the plurality of scenes; and a combining unit that combines a plurality of the scene editing data to generate work data" (see Abstract). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-141418 Summary of the Invention [Problem to be solved by the invention]

[0004] Patent Document 1 discloses a technique for creating a picture book-style video in which a user-provided image of a person moves. However, the technique in Patent Document 1 applies the provided image of a person to a pre-prepared picture book-style video work and transcribes predetermined movements, and does not disclose a mechanism for easily creating new animations. Therefore, the present technology provides a mechanism for easily creating new animations. [Means for solving the problem]

[0005] In order to solve the above problems, for example, the configurations described in the claims are adopted. The present application includes a number of configurations and methods for solving the above problems, and examples thereof include: Acquire multiple rough line drawings, generating a plurality of frame information for the character using one or a plurality of image generation models in which character information is set based on the plurality of pieces of rough line drawing information; generating video information based on the plurality of frame information; The present invention provides an animation creation method, including: [Effects of the Invention]

[0006] According to the present technology, it is possible to provide a mechanism that allows new animations to be created easily. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]

[0007] [Figure 1] Figure 1 shows an example of the overall configuration of an animation creation system. [Figure 2] FIG. 2 shows an example of the hardware configuration of the editing terminal 101. [Figure 3] FIG. 3 shows an example of the hardware configuration of the server 102. [Figure 4] FIG. 4 shows an example of the hardware configuration of the user terminal 103. [Figure 5] Figure 5 shows an example of an animation creation flow. [Figure 6] FIG. 6 shows an example of the image generation model setting flow. [Figure 7] FIG. 7 shows an example of a flow for obtaining rough line drawing information. [Figure 8] FIG. 8 shows an example of a flow of generating frame information. [Figure 9] FIG. 9 shows another example of the frame information generation flow. [Figure 10]FIG. 10 is a schematic diagram illustrating an image generation model and its input and output information according to one embodiment. [Figure 11] FIG. 11 shows an example of a flow for generating background image information. [Figure 12] FIG. 12 shows an example of an image generation model setting screen. [Figure 13] FIG. 13 shows an example of a rough line drawing information generation screen. [Figure 14] FIG. 14 shows an example of an image information generation screen. [Figure 15] 15(A) to 15(C) are diagrams showing intermediate image information generated in another embodiment. [Figure 16] 16(A) and 16(B) are schematic diagrams showing other methods of generating intermediate image information. [Figure 17] 17(A) to 17(C) show examples of diagrams created in the frame information generation step in one embodiment. [Figure 18] 18(A) to 18(C) are schematic diagrams showing the generation of image information in one embodiment. [Figure 19] FIG. 19 is an example of an image coloring screen. [Figure 20] FIG. 20 is a schematic diagram showing an image coloring method according to another embodiment. [Figure 21] FIG. 21 shows another example of the background image information generation flow. [Figure 22] FIG. 22 shows an example of a background image information generation screen. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, an animation creation system and an animation creation method according to the present technology will be described with reference to the drawings. Note that in each drawing, components having the same functions may be designated by reference numerals and may not be described in detail.

[0009] [Animation Creation System] FIG. 1 is a diagram illustrating an example of the configuration of an animation creation system 1 according to an embodiment. The animation creation system 1 according to the present technology is a system for creating animation data while reducing the manpower and effort required of creators. The animation creation system 1 can also support creators in creating animation data.

[0010] Animation in this technology is a technology that allows humans to perceive continuous movement (i.e., apparent motion) from multiple still images. This technology targets animation realized by a series of non-realistic still images, rather than moving images made up of images (e.g., live-action images) of actual people or scenes. Such still images may be still images in the style of paintings drawn by hand or using computer graphics (CG) technology, etc.

[0011] Furthermore, animation data refers to data (information) for representing a plurality of still images that realize an animation (hereinafter, sometimes simply referred to as "anime"). Animation data includes a combination (set) of information about a series of still images that change minutely, and by presenting (for example, projecting) a series of still images successively (i.e., continuously) at short time intervals based on this animation data, it is possible to make a user perceive apparent motion. Such animation data may be stored in various storage media such as film or memory. In this embodiment, the present technology will be described using an example of creating 2D animation. However, the present technology can also be used when creating 3D animation.

[0012] As shown in Fig. 1, the animation creation system 1 is configured with one or more editing terminals 101. The animation creation system 1 may additionally include, for example, one or more servers 102 and one or more user terminals 103, each of which is independent of the other. The editing terminals 101, the servers 102, and the user terminals 103 are each capable of transmitting and receiving information to and from each other via a wired or wireless network. The editing terminal 101 may be separate from the server 102, or may be configured integrally with the server 102.

[0013] Each server 102 and terminal (hereinafter, the editing terminal 101, the server 102, and the user terminal 103 may be simply referred to as a "terminal") of the animation creation system 1 may be, for example, a portable terminal (mobile terminal) such as a smartphone, tablet, mobile phone, or personal digital assistant (PDA), or may be a wearable terminal such as glasses (including goggles), a wristwatch, or clothing. Each terminal may also be a stationary or portable computer, or a server located on the cloud or a network. From the perspective of functionality, each terminal may be a VR (Virtual Reality) terminal, an AR (Augmented Reality) terminal, or an MR (Mixed Reality) terminal. Alternatively, each terminal may be a combination of multiple such terminals. For example, a combination of one smartphone and one wearable terminal may logically function as a single terminal. Each terminal may also be an information processing terminal other than those listed above.

[0014] Each terminal and server 102 of the animation creation system 1 may optionally include a processor that executes an operating system, applications, programs, etc., a main storage device such as RAM (Random Access Memory), an auxiliary storage device such as an IC card, hard disk drive, SSD (Solid State Drive), or flash memory, a communication control unit such as a network card, wireless communication module, or mobile communication module, input devices such as a touch panel, keyboard, mouse, audio input device, motion controller, or input device that detects motion by capturing images from a camera unit, input devices such as sensors like GPS, gyro sensor, or acceleration sensor, and output devices such as a monitor or display. Note that the output device may also be a device or terminal that transmits information to be output to an external device such as a monitor or display, printer, audio output device, or oscillator.

[0015] The main memory stores various programs and applications (software modules), and the processor executes these programs and applications to realize each functional element of the overall system. Each module may be implemented as an independent program or application, or as a subprogram or function within a single integrated program or application. Each module may also be implemented as hardware (hardware modules) by integrating circuits or using a microcomputer.

[0016] Furthermore, each module may be implemented by a single processor or multiple processors. Furthermore, each module may be provided in a single terminal (including a management server) or may be provided separately in two or more terminals (including a management server) interconnected via a network. Furthermore, each module may be provided in each or any one or more of two or more terminals (including servers) interconnected via a network. It is also contemplated that some modules may be implemented in a different country from the other modules.

[0017] In this specification, each module is described as the entity (subject) that performs the processing, but in reality, the processing is performed by a processor executing programs, applications, etc. to realize each module.

[0018] The auxiliary storage device stores various databases (DB). A "database" is, for example, a set of data that has been organized and collected so that it can accommodate any data manipulation (e.g., extraction, addition, deletion, overwriting, etc.) from a processor or an external computer. The auxiliary storage device is a functional element (storage unit) that stores one or more sets of data. The implementation method of the database is not limited to this example and may be, for example, a database management system, spreadsheet software, or text files such as XML or JSON. Some or all of this information may be stored in a relational database or a non-relational database. The database may be provided independently of the processor, while being connectable to the processor, etc.

[0019] [Editing device] 2 illustrates an example of the hardware configuration of the editing terminal 101. The editing terminal 101 is the main element for creating animation data. The editing terminal 101 is typically a terminal used by a creator who creates animation data using this animation creation system 1. The editing terminal 101 is configured, for example, by a personal computer or the like.

[0020] The editing terminal 101 includes a main memory device 201 and an auxiliary memory device 202. The editing terminal 101 also includes a processor 203, an input device 204, an output device 205, and a communication control unit 206 as described above.

[0021] The main memory device 201 stores programs and applications such as an animation creation module 211, a model setting module 212, a rough line drawing information acquisition module 213, a frame information acquisition module 214, a background image information generation module 215, and a synthesis module 216. Each functional element of the editing terminal 101 is realized by the processor 203 executing these programs and applications stored in the main memory device 201.

[0022] The auxiliary storage device 202 stores information necessary for the operation of the animation creation system 1. The auxiliary storage device 202 stores, for example, an image generation model 221, model setting information 222, original rough line drawing information 223, rough line drawing information 224, intermediate image information 225, line drawing information 226, coloring information 227, frame information 228, background image information 229, and video information 230. Details of this information will be described later.

[0023] Each functional element of the editing terminal 101 will now be briefly described. The animation creation module 211 comprehensively controls the basic operations of the editing terminal 101. The animation creation module 211 executes animation creation processing by causing modules such as a model setting module 212, a rough line drawing information acquisition module 213, a frame information acquisition module 214, a background image information generation module 215, and a synthesis module 216 to operate in cooperation with each other. The animation creation module 211 also cooperates with the server 102, for example, and stores the created animation data in the server 102.

[0024] The model setting module 212 sets the features of the image generated by the image generation model. Various methods for setting the features of the image generated by the image generation model are considered. Specifically, the model setting module 212 loads a data file related to the feature quantities of a plurality of images having a certain feature, which are generated in the process of learning the images, into the image generation model, or sets tuning information for parameters in the image generation model. This makes it possible, for example, to impart predetermined features to the images generated by the image generation model and ensure consistency. Furthermore, the model setting module 212 introduces, for example, one or more additional learning layers (e.g., adaptation layers) into the image generation model.

[0025] The model setting module 212, for example, changes the additional learning layer to be introduced, causing one image generation model to function as multiple image generation models that each output images with different features. In this embodiment, five different image generation models are stored: a first image generation model M1, a second image generation model M2, a third image generation model M3, a fourth image generation model M4, and an interpolated image generation model Mip. This simplifies the configuration of the image generation model while enabling the generation of a variety of images. For example, it is possible to generate images of multiple characters by depicting their distinct features.

[0026] The rough line drawing information acquisition module 213 acquires multiple pieces of rough line drawing information. The rough line drawing information acquisition module 213 generates multiple pieces of rough line drawing information using, for example, an interpolated image generation model Mip (described later) based on the original rough line drawing information 223. The rough line drawing information will be described later. The frame information acquisition module 214 generates a plurality of frames of information for the character using one or more image generation models based on a plurality of pieces of rough line drawing information. The frame information will be described later.

[0027] The background image information generation module 215 generates a background image using an image generation model. The background image information generation module 215 generates a background image suitable for creating the target animation, for example, using one or more image generation models. For example, the functions of both the frame information acquisition module 214 and the background image information generation module 215 may be implemented by a single image generation module (not shown).

[0028] The compositing module 216 generates moving image information based on the generated images. The compositing module 216 generates frame image information, for example, by compositing frame information (characters) with background image information. As described above, frame image information is digitally recorded information for each frame image that makes up an animation. The compositing module 216 also creates moving image information as animation data based on multiple pieces of frame image information.

[0029] In this technology, "digitally recording" means recording in an electronic, magnetic, or other format that cannot be recognized by human perception.

[0030] Each piece of information stored in the auxiliary storage device 202 will be briefly described below. The image generation model 221 is a generative artificial intelligence (generative AI) trained to output (generate) images in response to various images or prompts. The image generation model 221 generates a plurality of frame images (still images) that constitute an animation. The image generation model 221 of this embodiment is a generative AI configured to output new image information corresponding to each image information based on at least various image information as input. The image generation model 221 may be, for example, a unimodal generative AI whose input is limited to image information, or a multimodal generative AI whose input is not limited to image information. The image generation model 221 may be configured to output an image using, for example, text information and image information (including still image information and video information) as input.

[0031] The image generation model 221 may be, for example, a deep generative model that combines deep learning and a generative model. Examples of deep generative models for generating images include a variational autoencoder (VAE), a generative adversarial network (GAN), a flow-based model, a diffusion model, and a latent diffusion model (LDM).

[0032] In this embodiment, the image generation model 221 is configured to generate images with a plurality of different characteristics. For example, by introducing the same or different additional layers, the image generation model 221 can have five different functions: an interpolation image generation model Mip, a first image generation model M1, a second image generation model M2, a third image generation model M3, or a fourth image generation model M4. Functional implementations of these image generation models will be described later.

[0033] The model setting information 222 is information about prompts, additional layers, parameters, etc. for realizing each function of the image generation model 221. The model setting information 222 may include, for example, overall information and partial information, which will be described later. The model setting information 222 may also include text information used as input when setting the image generation model 221.

[0034] The text information is information used as input to the image generation model 221. The text information is information that expresses in language the characteristics of characters, objects, backgrounds, etc. in animations or frame images. The text information may be information that expresses in language the characteristics of characters, objects, backgrounds, etc. in a positive way, or information that expresses in language the characteristics of characters, objects, backgrounds, etc. in a negative way.

[0035] The original rough line drawing information 223 is information (data) that digitally records the original rough line drawing. In the present technology, the original rough line drawing refers to a roughly drawn line drawing that indicates the layout of a character and the general posture of the character at important timings in the character's movement when drawing the character in animation (for example, see FIG. 13).

[0036] The original rough line drawing can be an image with a precision (roughness) corresponding to an image called a "first original drawing" or "rough original" that is the basis for a so-called "storyboard" or "original drawing" during the animation creation process. Typically, a plurality of original rough line drawings (e.g., about 2 to 5) are prepared for one scene in an animation. The original rough line drawing may be an image composed of lines that characteristically represent the general posture of a character or the like, so as to be suitable for generating frame information, etc., described later, using an image generation model. The rough line drawing may characteristically represent the outline or contour lines that represent the general posture of a character or the like. The rough line drawing is typically, for example, a color-reduced image with a more limited number of colors. For example, the rough line drawing is composed of a monochrome image (e.g., a binary image such as black and white or white-brown) or a grayscale image (e.g., a grayscale image representing shades of white to black, or a sepia image representing shades of white to brown). The rough line drawing is typically configured as an image that does not include color information (for example, a grayscale image that includes only brightness information). The original rough line drawing information 223 can be stored in association with, for example, each scene and other conditions.

[0037] In this embodiment, the original drawing rough line drawing information 223 can typically be manually prepared. For example, the multiple pieces of original drawing rough line drawing information 223 can be prepared by digitizing a storyboard that has been roughly drawn by a person on a paper storyboard sheet and extracting the rough draft of each screen as electronic information. Alternatively, the multiple pieces of original drawing rough line drawing information 223 can be prepared by using computer graphics technology to extract the electronic information of the rough draft of each screen from electronic information of a storyboard that has been roughly drawn by a person on a storyboard sheet on a PC.

[0038] The rough line drawing information 224 is information (data) in which a rough line drawing is digitally recorded. In the present technology, a rough line drawing is information generated by temporally consecutively interpolating between a plurality of pieces of original rough line drawing information 223. As the rough line drawings, for example, at least a number of images corresponding to each frame image (still image) constituting an animation are prepared. The rough line drawings may be, for example, line drawings roughly drawn to indicate the layout of a character, the general posture of the character, and the like for each frame image (still image) constituting an animation. The rough line drawing information 224 can be stored in association with, for example, each scene or other conditions.

[0039] The intermediate image information 225 is information that is created additively. The intermediate image information 225 is digitally recorded information of an intermediate image that is created based on the rough line drawing information 224 in order to create the frame information 228. The intermediate image may be an image of a character in the layout and posture indicated by the rough line drawing. This intermediate image may be a monochrome image or a grayscale image, similar to the rough line drawing, or may be a color image that includes color information, unlike the rough line drawing. Preferably, the intermediate image is a color image. The intermediate image information 225 can be stored in association with, for example, each scene or other conditions.

[0040] The line drawing information 226 is information (data) that digitally records line drawings. In the present technology, line drawings refer to images drawn with lines that represent the outline, contour, and the like of characters in each frame image (still image) that constitutes an animation. This line drawing corresponds, for example, to the "uncolored line drawing layer of a character" in each frame image (still image) that constitutes an animation. This line drawing is an image drawn with relatively high precision, known as a "second original drawing" or "second original (clear copy of the original drawing)." This line drawing, like the rough line drawing, is created as a monochrome image or a grayscale image. The line drawing information 226 can be stored, for example, in association with each scene or other conditions.

[0041] The coloring information 227 is information about one or more colors that are determined for each part of the target character. The coloring information 227 is referenced, for example, when creating frame information 228 based on the line drawing information 226. The coloring information 227 can be stored, for example, in association with a scene or other conditions.

[0042] The frame information 228 is information (data) that digitally records a frame image. In the present technology, a frame image refers to an image of a character in each frame image (still image) that constitutes an animation. This frame image corresponds, for example, to a "character image layer" in each frame image (still image) that constitutes an animation. The characters in the frame image are given a predetermined color based on the coloring information 227. The frame information 228 can be stored in association with, for example, each scene or other conditions.

[0043] The background image information 229 is information (data) that digitally records a background image. In the present technology, a background image refers to an image of the background in each frame image (still image) that constitutes an animation. This background image corresponds to, for example, a "background image layer" in each frame image (still image) that constitutes an animation. The background in the background image is given a predetermined color. The background image information 229 can be stored in association with, for example, each scene or other conditions.

[0044] The video information 230 is digitally recorded information of a part or all of an animation. In this specification, the video information 230 may be referred to as animation data. Based on this video information 230, a computer (e.g., processor 203) can display, for example, a part or all of the animation on an output device 205 such as a display. The video information 230 can be stored in association with, for example, each scene or other conditions.

[0045] [server] 3 illustrates an example of the hardware configuration of the server 102. The server 102 is an additional element of the animation creation system 1, and is configured, for example, by a computer server or the like located on a cloud. The server 102 is typically a terminal used by an administrator who manages and operates the animation creation system 1, and can cooperate with the editing terminal 101 to create video information 230 as necessary.

[0046] The server 102 includes a main storage device 301 and an auxiliary storage device 302. The server 102 also includes a processor 303, an input device 304, an output device 305, and a communication control unit 306 as described above.

[0047] The main memory device 301 stores programs and applications such as an animation creation module 311, a model setting module 312, a rough line drawing information acquisition module 313, a frame information acquisition module 314, a background image information generation module 315, and a compositing module 316. These modules in the server 102 may have the same functions as the modules in the editing terminal 101. These functional elements of the server 102 are realized by the processor 303 executing these programs and applications stored in the main memory device 301.

[0048] The auxiliary storage device 302 can store, for example, an image generation model 321, model setting information 322, original rough line drawing information 323, rough line drawing information 324, intermediate image information 325, line drawing information 326, coloring information 327, frame information 328, background image information 329, and video information 330. These pieces of information may be the same as the corresponding pieces of information 221 to 230 in the editing terminal 101.

[0049] The animation creation module 311 has the same functions as the animation creation module 211 of the editing terminal 101, and creates animation data (video information 330). The animation creation module 311 can also output the created animation to the user terminal 103. The animation creation module 311, for example, works in cooperation with a viewing module 412 of the user terminal 103 to display animation on a display (an example of an output device 405; the same applies below) of the user terminal 103 based on the created video information 330. The animation creation process of the animation creation system 1 may be executed by the editing terminal 101 or by the server 102.

[0050] [User device] FIG. 4 illustrates an example of the hardware configuration of the user terminal 103. The user terminal 103 is an element that outputs created animations. The user terminal 103 is typically a terminal used by a user who uses the animation creation system 1 to view animations. The user terminal 103 is configured, for example, by a personal computer or a smartphone. The user terminal 103 may also be, for example, a terminal with a head-mounted display, a head-mounted AR terminal, a terminal with AR glasses, or the like.

[0051] The user terminal 103 includes a main memory device 401 and an auxiliary memory device 402. The user terminal 103 also includes the above-described processor 403, input device 404, output device 405, camera 406, and communication control unit 407. The camera 406 may be an imaging device built into a smartphone or the like, or may be an imaging device built into a wearable device such as AR glasses or a head-mounted display. The output device 405 may be a display device (display) included in the smartphone or the like, or may be a glasses-type or goggle-type display device in AR glasses or a head-mounted display.

[0052] The main memory device 401 stores programs and applications such as a user management module 411 and a viewing module 412. Each functional element of the user terminal 103 is realized by the processor 403 executing these programs and applications stored in the main memory device 401.

[0053] The auxiliary storage device 402 stores information necessary for the operation of the animation creation system 1. The auxiliary storage device 402 stores, for example, video information 430. The video information 430 may be a part or all of the video information 230 stored in the auxiliary storage device 402 of the server 102.

[0054] The user management module 411 comprehensively manages the basic operations of the user terminal 103. The user management module 411, for example, executes the processing required for connecting to the server 102. For example, the user management module 411 cooperates with the animation creation module 311 of the server 102 to output (display) the homepage, login page, title page, etc. of the anime viewing site provided on the web by the server 102 to an output device 405 such as a display of the user terminal 103.

[0055] The viewing module 412 executes the processing required for the user to view the animation on the user terminal 103. The viewing module 412, for example, works in conjunction with the animation creation module 311 of the server 102 to display the animation on a display.

[0056] [How to create animation] <Embodiment 1> Next, a method for creating an animation using the animation creation system 1 will be described. The creator executes an animation creation program stored in the auxiliary storage device 202 of the editing terminal 101, for example. This starts an animation creation application, and the animation creation module 211 is realized. The animation creation module 211 displays a screen of the animation creation application on the output device 205, such as a display. When the animation creation module 211 receives an instruction from the creator via the screen of the animation creation application to start creating an animation, it starts creating the animation in cooperation with, for example, the modules 212 to 216.

[0057] Figure 5 shows an example of an animation creation flow. In the animation creation method of this embodiment, the animation creation module 211 cooperates with other modules to execute the following processes: The model setting module 212 sets an image generation model (S510); The rough line drawing information acquisition module 213 acquires multiple pieces of rough line drawing information (S520); The frame information acquisition module 214 generates frame information (S530); The background image information generation module 215 generates background image information (S540); The compositing module 216 generates multiple pieces of frame image information representing an image in which the background and the character are combined (S550), and creates video information based on the multiple pieces of frame image information (S560).

[0058] The animation creation module 211 executes the animation creation process in cooperation with a model setting module 212, a rough line drawing information acquisition module 213, a frame information acquisition module 214, a background image information generation module 215, and a synthesis module 216, but the animation creation module 211 may be configured to execute all or part of the processing of each of these modules itself.

[0059] Each step of the animation creation method will be described below with reference to the drawings as appropriate. 10 is a schematic diagram showing an image generation model and its input / output information according to one embodiment. In the first embodiment, an interpolation image generation model Mip, a second image generation model M2, a third image generation model M3, and a fourth image generation model M4 are used to generate video information 230 from original rough line drawing information 223 via rough line drawing information 224, intermediate image information 225, line drawing information 226, and frame information 228.

[0060] First, the image generation model setting step (S510) will be described. Fig. 6 shows an example of an image generation model setting flow 600. Fig. 12 shows an example of a setting screen 1200 for an image generation model.

[0061] An image generation model generates one or more output images in response to a prompt (sometimes referred to as input information, input text information, or text information) based on, for example, input image information. While there is no problem generating a still image of a single scene, there are challenges in generating consistent images across a single work or scene, such as an anime that spans several minutes to several hours. Specifically, there are problems with fixing the animation style or the general characteristics of each character across a single work, or fixing the color and decoration of characters across multiple frames of images that are chronologically consecutive.

[0062] Therefore, the model setting module 212 sets conditions related to image generation of the image generation model. The model setting module 212 sets conditions related to features such as the art style, the outer shape of a character or object, and parts of a character or object. Specifically, the model setting module 212 sets conditions related to image generation of the image generation model, for example, by executing the following model setting flow. The model setting module 212 displays an image generation model setting screen 1200 shown in FIG. 12 on the display of the editing terminal 101, accepts instructions from the creator via this image generation model setting screen, and executes the following model setting flow in accordance with the accepted instructions.

[0063] In this specification, "artistic style" typically refers to the atmospheric characteristics that appear commonly in images generated by an image generation model, and refers to, for example, the characteristics and tendencies such as the drawing technique, touch, line tone, color, color scheme, and texture that appear in the picture, as well as combinations of these. Furthermore, "appearance" typically refers to the character's body contours such as face, limbs, and hair, as well as the skeleton, fleshiness, body balance, physique, and contours of accessories such as clothing and necklaces, and combinations of these.

[0064] In the model setting step, the model setting module 212 executes the following image generation model setting flow 600. The image generation model setting flow 600 includes acquiring an image generation model (S610), setting overall information, which is information about the character's appearance and / or style, in the image generation model (S620), setting partial information about some of the character's features in the image generation model (S630), and setting to generate a line drawing and / or a colored drawing (S640). Note that the order in which steps S620 to S640 are performed does not matter.

[0065] Specifically, the model setting module 212 first acquires an image generation model (S610). The image generation model acquired by the model setting module 212 is, for example, a denoising diffusion probabilistic model (DDPM) configured to generate an image based on image information. DDPMs are trained to minimize the difference between a denoised image and the original image when estimating data from noise by performing an inverse process of converting a noise-denoised image based on the process of gradually adding Gaussian-distributed noise to an image until the image becomes completely noisy. DDPMs are trained on a dataset consisting of a combination of images and text, and are configured to take into account the input text and / or images during noise removal and generate an image that matches the input. DDPMs are composed of an artificial neural network with a large number of parameters (typically tens of millions or more, e.g., hundreds of millions or even tens of billions or more). The model setting module 212 can acquire the image generation model by, for example, copying (reading) an image generation model stored in the server 102. The acquired image generation model is stored as an image generation model 221 in the auxiliary storage device 202, for example.

[0066] The model setting module 212 then sets overall information, which is information about at least one of the character's appearance and style, to the image generation model (S620). The overall information is, for example, information about the feature amounts of an image generated when the image generation model is further trained. The overall information can be, for example, information about the feature amounts of an image generated when the image generation model is further trained preliminarily using an image of a character with a predetermined appearance (for example, physique) and / or style. The overall information is stored, for example, as an overall information file in the model setting information 222 of the auxiliary storage device 202, 302 (which may be one or both of the auxiliary storage device 202 and the auxiliary storage device 302; the same applies below).

[0067] The model setting module 212, for example, displays an overall information file selection field 1210 on the image generation model setting screen 1200. When the overall information file selection field 1210 is selected, the model setting module 212 displays a file selection dialog and accepts the selection of the overall information file to be set by the creator via the file selection dialog. The model setting module 212 sets parameter values ​​for the image generation model, for example, based on the overall information recorded in the overall information file. As a result, the model setting module 212 obtains an image generation model in which parameters for generating an image with a predetermined shape and / or art style are fixed for, for example, an acquired image generation model (typically, DDPM). The model setting module 212 can impart consistency to the shape and / or art style of the character in the generated image by using, for example, such an image generation model. In this embodiment, the model setting module 212 sets overall information for the image generation model, including information representing a 2D anime image, as information related to the art style.

[0068] The model setting module 212 can set one overall information file or two or more overall information files for an image generation model. For example, the model setting module 212 displays an Add button 1220 on the image generation model setting screen 1200, and when the Add button 1220 is selected, displays an additional overall information file selection field (not shown) below the first overall information file selection field 1210. The model setting module 212 then accepts the selection of an overall information file to be additionally set via the additional overall information file selection field. Although not specifically shown, when two or more pieces of overall information are set, the model setting module 212 may be configured to accept a designation of the degree of reflection of each piece of overall information.

[0069] Furthermore, the model setting module 212 can set, for example, a general information file for a first character or a general information file for a second character different from the first character for the image generation model. In other words, the model setting module 212 may be configured to be able to switch the settings of the image generation model so as to output images of different characters.

[0070] The model setting module 212 also sets partial information related to some characteristics of the character in the image generation model (S630). The partial information is, for example, information about the feature quantities of an image generated when the image generation model, which has learned the above-mentioned external shape and / or art style, is further trained on some characteristics of the character. The partial information may be, for example, information about the feature quantities of an image generated when the above-mentioned image generation model is further trained preliminarily using images of a character that represent at least one of the individual characteristics of a predetermined part (e.g., facial features, facial expression, hairstyle, contours, body shape, etc.), appearance (e.g., clothing, etc.), and action (e.g., speaking, singing, standing up, walking, running, playing a specific sport, playing a specific instrument, etc.). The partial information may be, for example, information about feature quantities that characterize the above-mentioned external shape and / or art style at a more detailed level (e.g., the tone of the outline or perimeter of the character, etc.). This partial information is stored, for example, as a partial information file, in the model setting information 222 of the auxiliary storage device 202, 302.

[0071] The model setting module 212, for example, displays a partial information file selection field 1230 on the image generation model setting screen 1200. When the partial information file selection field 1230 is selected, the model setting module 212 displays a file selection dialog and accepts the creator's selection of the partial information file to be set via the file selection dialog. The model setting module 212, for example, sets a very small number of parameter values ​​for the image generation model based on the partial information recorded in the partial information file. This significantly reduces the number of parameters for the image generation model, enabling control of feature amounts while improving image quality. As a result, it is possible to generate images with desired features without significantly increasing the capacity of the image generation model itself. In particular, it is possible to fix some of the features of a character while maintaining the external shape and / or style set by the overall information file.

[0072] The model setting module 212 can set one piece of partial information or two or more pieces of partial information for an image generation model. For example, the model setting module 212 displays an Add button 1240 on the image generation model setting screen 1200, and when the Add button 1240 is selected, displays an additional partial information file selection field (not shown) below the first partial information file selection field 1230. The model setting module 212 then accepts the selection of a partial information file to be additionally set via the additional partial information file selection field.

[0073] Here, partial information can be prepared for each different part of the character. Furthermore, partial information can be prepared for each of a plurality of different features for a certain part of the character. Therefore, the model setting module 212 can set multiple pieces of partial information by changing the combination based on a combination of multiple pieces of partial information. As a result, for example, it is possible to generate an image in which some of the features of the character are changed. Although not specifically shown, when setting two or more pieces of partial information, the model setting module 212 may be configured to accept a designation of the degree of reflection of each piece of partial information. In this embodiment, the model setting module 212 can set partial information by introducing, for example, a trained model or setting information regarding facial features that identify the character and a trained model or setting information regarding clothing into the image generation model.

[0074] The model setting module 212 displays, for example, a preview window 1270 and a preview button 1280 on the image generation model setting screen 1200. When the preview button 1280 is selected, the model setting module 212 displays an example of an image generated by the image generation model for which the overall information and partial information selected in the overall information file selection field 1210 and partial information file selection field 1230 have been set. This allows the creator to check whether the selected overall information file and partial information file are appropriate while viewing the preview window 1270.

[0075] The model setting module 212 further sets line drawing information for the image generation model so that the generated image is generated as an uncolored line drawing and / or a colored line drawing (S640). The line drawing information is, for example, information about features that extract contour lines from the image to be generated and control the line tone of the extracted contour lines (thickness, fineness of extraction, shape / softness of the line ends, etc.). The line drawing information may be, for example, information about the features of an image generated when the image generation model is preliminarily trained using anime line drawings (e.g., uncolored or colored storyboards, original drawings, etc.) rather than simple line drawings. The line drawing information may be, for example, information about feature values ​​(e.g., line tone of the contour or perimeter line of a character, etc.) that characterize the contour lines, outlines, and / or art style (thickness and its transition, trajectory, curvature, corner shape, etc.) of a character. This line drawing information can be stored, for example, as a line drawing information file in the model setting information 222 of the auxiliary storage device 202, 302.

[0076] Furthermore, the model setting module 212 can acquire text information, which is information in text format for conditioning image generation by the image generation model, and set it for the image generation model. This text information can be instruction information for making the image information generated by the image generation model suitable for an anime frame image. The model setting module 212 acquires, for example, text information related to character features. As text information, the model setting module 212 can acquire positive instruction information that instructs on elements that are desired to be included in the image to be generated. Furthermore, the model setting module 212 can acquire negative instruction information that instructs on elements that are not desired to be included in the image to be generated.

[0077] The model setting module 212 displays, for example, a positive instruction information field 1250 for accepting input of positive instruction information from the creator and a negative instruction information field 1260 for accepting input of negative instruction information from the creator on the image generation model setting screen 1200. The model setting module 212, for example, acquires text information entered by the creator in the positive instruction information field 1250 as positive instruction information. The model setting module 212 also acquires text information entered by the creator in the negative instruction information field 1260 as negative instruction information. The model setting module 212 can accept text information at any time via, for example, the positive instruction information field 1250 and the negative instruction information field 1260.

[0078] For example, the model setting module 212 accepts the following as positive instruction information: "...long-sleeved shirt with sleeves half rolled up, pendant in yellow gold, dark green flared skirt with brown belt, knee-length flared skirt..." written in the positive instruction information field 1250. Furthermore, the model setting module 212 accepts, for example, "multiple people, . . . short skirt" written in the negative instruction information column 1260 as negative instruction information.

[0079] The model setting module 212 may be configured to accept, as text information, a combination of multiple pieces of instruction information that have been prepared in advance, as exemplified by "NegativePromptSet_V5" in the negative instruction information field 1260. For example, by preparing a combination of multiple pieces of instruction information in advance for each animation or each character, it is possible to prevent, for example, missing instruction information or forgetting to delete unnecessary instruction information, and it is advantageous for enabling stable generation of images for characters. The model setting module 212 can store the text information set in this way in the model setting information 222 in the auxiliary storage device 202, 302, for example.

[0080] The model setting module 212 can also set an image completion function that, when two input images are given, predicts and generates one or more intermediate images that naturally connect the two input images. During image interpolation, for example, multiple interpolation candidates can be estimated by repeatedly introducing latent vectors containing noise from the two input images and text prompts and posture information extracted from the two input images at various noise levels during a noise removal process. An optimal interpolated image can be output by extracting an image with high similarity to a prompt that describes a target characteristic (e.g., posture). For example, spherical linear interpolation can be used to interpolate the latent space and text embedding, and linear interpolation can be used to interpolate posture. Information for implementing the image completion function can be stored in the model setting information 222, for example, as an image completion function file.

[0081] The model setting module 212 can also set a part coloring function that identifies parts of an object in a first image and generates a second image in which each part is colored with a specified color. The part coloring function may include, for example, a part detection layer that uses an image recognition function to learn to identify parts such as a person's head (and further, hair, face, eyes, nose, and mouth), upper limbs, torso, lower limbs, and attachments (clothing, hat, footwear, glasses, and accessories) based on the positions (including relative positions) and sizes of their feature points. A coloring layer or the like may also be included that colors each part using a specified color as a reference, while adding shading according to the surface shape of the part. Information for implementing such a part coloring function can be stored in the model setting information 222, for example, as a part coloring function file or the like.

[0082] By following the image generation model setting flow 600 described above, an image generation model in which character information is set can be obtained (S650). Image generation models M1 to M4, Mip that generate various images can be functionally realized by combining overall information, partial information, line drawing information, text information, etc., set in the image generation models. Models set in this way may be stored, for example, together with process information indicating in which processing step they are used. Furthermore, each of these models may be stored, for example, as information indicating what model setting information 222 was used to set it.

[0083] The model setting module 212 can realize the functions of the second image generation model M2 by setting, for example, overall information about the target animation, overall information and / or partial information about a specific character, colored line drawing information, and a part coloring function for the image generation model. That is, the second image generation model M2 is trained to identify parts of an object in a first image and generate a second image in which each part is colored with a specified color, and generates the plurality of frame information in which each part is colored with a pre-specified color. The second image generation model M2 can be used, for example, to generate frame information 228 from line drawing information 226.

[0084] The model setting module 212 can also implement the functionality of the third image generation model M3 by setting, for example, overall information about the target animation, overall information and / or partial information about a specific character, and colored line drawing information for the image generation model. That is, the third image generation model M3 is an image generation model trained to generate a second image in which the outline or shape of an object in a first image is displayed using lines. Overall information, which is information about the character's outer shape and style, and partial information about some of the character's features are set for the image generation model. The third image generation model M3 then generates, for example, multiple pieces of colored intermediate image information 225 about the character. The third image generation model M3 can be used, for example, to generate multiple pieces of colored intermediate image information 225.

[0085] The model setting module 212 can realize the function of the fourth image generation model M4 by setting, for example, overall information about the target animation, overall information and / or partial information about a specific character, and uncolored line drawing information for the image generation model. That is, the fourth image generation model M4 is an image generation model trained to display the outline or shape of an object in a first image using lines and generate a second image that displays only the lines. Overall information, which is information about the character's outer shape and style, and partial information about some of the character's features are set for the image generation model. The fourth image generation model M4 then generates the plurality of line drawing information for the character with reduced color. The fourth image generation model M4 can be used, for example, to generate uncolored line drawing information 226.

[0086] The model setting module 212 can realize the functions of the interpolated image generation model Mip by setting, for example, overall information about the target animation, overall information and / or partial information about a specific character, uncolored line drawing information, and an image completion function for the image generation model. That is, when multiple pieces of original rough line drawing information are input, the interpolated image generation model Mip generates multiple pieces of rough line drawing information by temporally consecutively interpolating between the multiple pieces of original rough line drawing information. The interpolated image generation model Mip can be used, for example, to generate the rough line drawing information 224. Note that the overall information and / or partial information about a specific character does not necessarily need to be set in the interpolated image generation model Mip, and can be omitted.

[0087] The model setting module 212 may set the first image generation model M1 and the fourth image generation model M4 to the same settings, or may set them to have different parameters and prompts. Furthermore, the model setting module 212 may set the second image generation model M2 and the third image generation model M3 to the same settings, or may set them to have different parameters and prompts.

[0088] This allows the functions of the second image generation model M2, the third image generation model M3, and the fourth image generation model M4 to be set. As a result, for example, a highly versatile image generation model can be obtained that can generate images with features from various perspectives using DDPM as a base model.

[0089] Note that the model setting module 212 displays, for example, a preview window 1270 and a preview button 1280 on the image generation model setting screen 1200. Then, for example, when the model setting module 212 receives selection of the preview button 1280 by the creator, it displays an image generated by the image generation model set at that time. That is, for example, it displays an image generated when specified original rough line drawing information 223 (not shown) or the like is applied to the image generation model. As shown in the preview window 1270, the model setting module 212 displays an image representing a character in which overall information, partial information, line drawing information, and text information are reflected for the image generation model, for example, for a character based on the original rough line drawing information 223.

[0090] For example, in response to the instruction "...long-sleeved shirt with sleeves half rolled up, pendant in yellow gold, dark green flared skirt with brown belt, knee-length flared skirt..." written in the positive instruction information column 1250, a character is generated in the preview window 1270 that reflects the positive instruction information, that is, a long-sleeved shirt with sleeves half rolled up, a yellow gold pendant, a dark green flared skirt with a brown belt, and a knee-length flared skirt.

[0091] Furthermore, due to the instruction "multiple people, ...short skirt" written in the negative instruction information column 1260, multiple characters and no characters wearing short skirts are generated in the preview window 1270. In other words, characters that are unlikely to reflect negative instruction information are generated. This allows the model setting module 212 to generate image information representing an image of a character that reflects positive instruction information and is less likely to reflect negative instruction information.

[0092] The size, aspect ratio, resolution, etc. of the image that the model setting module 212 causes the image generation model to generate may correspond to the size, shape (for example, an aspect ratio of 16:9), resolution, etc. of the target animation cut or frame, but are not limited to this. The model setting module 212 can cause the image generation model to generate images with, for example, any size, shape, and resolution. The model setting module 212 can also set the resolution of the image that the image generation model generates to, for example, either a standard resolution, an arbitrary resolution other than the standard resolution, or both.

[0093] Furthermore, the model setting information set by the model setting module 212 is not necessarily optimal for the target anime character. Therefore, the model setting module 212 may be configured to enable editing of the model setting information 222 on the image generation model setting screen 1200, for example, and to accept editing of the model setting information 222 by the creator. In this case, when the preview button 1280 is selected again, the model setting module 212 can generate an image by reflecting the corrected model setting information 222 in the image generation model, and display the generated image in the preview window 1270.

[0094] Next, the rough line drawing information acquisition step (S520) will be described. Although not limited to this, in this embodiment, a case will be described in which continuous movements of characters in an animation are acquired based on a plurality of pieces of original rough line drawing information 223 equivalent to a storyboard. Fig. 7 shows an example of a rough line drawing information acquisition flow 700. Fig. 13 shows an example of a rough line drawing information generation screen 1300.

[0095] The rough line drawing information acquisition module 213, for example, displays a rough line drawing information generation screen 1300 on the display of the editing terminal 101, accepts instructions from the creator via this rough line drawing information generation screen 1300, and executes the following rough line drawing information acquisition flow 700 based on the accepted instructions.

[0096] The rough line drawing information acquisition flow 700 includes acquiring a plurality of pieces of original rough line drawing information (S710), and generating a plurality of pieces of rough line drawing information using an interpolated image generation model based on the plurality of pieces of original rough line drawing information (S720).

[0097] 13 on the display of the editing terminal 101. The rough line drawing information acquisition module 213 displays, for example, an input information field 1310, a rough line drawing information display field 1320, an interpolation adjustment field 1330, an intermediate image adjustment field 1340, an intermediate image information display field 1350, a moving image display adjustment field 1360, and a moving image display field 1370 on the rough line drawing information generation screen 1300. The input information field 1310 is, for example, a file selection dialog.

[0098] The rough line drawing information acquisition module 213 acquires a plurality of pieces of original rough line drawing information by, for example, accepting a selection of a plurality of pieces of original rough line drawing information 223 by the creator (S710). The rough line drawing information acquisition module 213 displays, for example, an original rough line drawing based on the acquired original rough line drawing information 223 in the rough line drawing information display field 1320. Here, the original rough line drawing is a roughly drawn outline representing the movement of a person in a scene of climbing a staircase landing. The original rough line drawing is made up of five images representing the person's initial posture in this scene (cut), their final posture, and postures of the person at three points in time that represent characteristic arm movements between them.

[0099] For example, the creator plans to compose this scene with 12 frames and generate a total of 13 frames of images to provide one extra frame for editing, and inputs these conditions as interpolation conditions into the interpolation adjustment field 1330. The rough line drawing information acquisition module 213 acquires the interpolation conditions via the interpolation adjustment field 1330, and inputs these interpolation conditions and a plurality of pieces of original rough line drawing information 223 into the interpolated image generation model Mip, thereby acquiring a plurality of pieces of rough line drawing information 224 (S720).

[0100] The rough line drawing information acquisition module 213 displays, for example, the generated 13 pieces of rough line drawing information 224 in an intermediate image information display field 1350. The rough line drawing information acquisition module 213 is configured to be able to receive, for example, an instruction to adjust the image of the rough line drawing information 224 via the intermediate image adjustment field 1340. When the rough line drawing information acquisition module 213 receives an instruction to adjust the image of the rough line drawing information 224, it displays, in the intermediate image information display field 1350, the rough line drawing information 224 after adjustment in accordance with the instruction.

[0101] Furthermore, the rough line drawing information acquisition module 213 is configured to be able to display the generated plurality of pieces of rough line drawing information 224 frame by frame (i.e., video playback) at a predetermined frame rate in the video display adjustment field 1360. The creator can, for example, check the movements of the person in the rough line drawing displayed in the video display adjustment field 1360 and adjust the interpolation conditions as necessary. The rough line drawing information acquisition module 213 stores the generated rough line drawing information 224 in, for example, the auxiliary storage device 202 as rough line drawing information 224 .

[0102] The frame information generating step (S530) will be described. First, a description will be given of a case where the line drawing information 226 is generated directly without using the intermediate image information 225. Fig. 9 shows an example of a frame information generation flow 900. The frame information acquisition module 214 performs the following frame information generation flow 900.

[0103] The frame information generation flow 900 includes: A plurality of pieces of intermediate image information 225 about the character are generated by a model (third image generation model) that outputs a color image with contours based on the plurality of pieces of rough line drawing information 224 (S910). A plurality of pieces of line drawing information about the character are generated using a model (fourth image generation model) that generates a line drawing image of only contour lines based on a plurality of pieces of intermediate image information about the character (S920); A plurality of frame information pieces are generated based on a plurality of line drawing information pieces about the character by a model for coloring the line drawing (a second image generation model) (S930). This includes:

[0104] FIG. 14 shows an example of an image information generation screen 1400. 14 on the display of the editing terminal 101. The frame information acquisition module 214 displays, for example, an entire model information column 1410, a partial model information column 1420, a line drawing information column 1430, an input image information column 1440, an adjustment column 1450, a rough line drawing information display column 1460, and an intermediate image information display column 1470 on the image information generation screen 1400. The entire model information column 1410, the partial model information column 1420, the line drawing information column 1430, and the input image information column 1440 are, for example, file selection dialogs.

[0105] The frame information acquisition module 214 displays the overall model information, partial model information, and line drawing information set in step S510 in the overall model information column 1410, the partial model information column 1420, and the line drawing information column 1430, respectively, in the form of, for example, file names. The frame information acquisition module 214 also displays the rough line drawing information 224 acquired in step S520 in the input image information column 1440, in the form of, for example, file names. This allows the creator to confirm what information is used to generate the intermediate image information 225. Furthermore, as necessary, changes to the information to be set in the image generation model can be accepted from the creator via the overall model information column 1410, the partial model information column 1420, the line drawing information column 1430, the input image information column 1440, etc., and the information to be set in the image generation model can be changed based on such changes. In this embodiment, a third image generation model M3 is set up to create a colored line drawing of a specific character.

[0106] The frame information acquisition module 214 inputs (applies) the multiple pieces of rough line drawing information 224 specified in the input image information column 1440 to the third image generation model M3, thereby generating intermediate image information 225 corresponding to each piece of rough line drawing information 224 (S910).

[0107] The frame information acquisition module 214 generates an image (intermediate image) of a character that commonly reflects the overall and partial characteristics, such as the style and / or external shape, set in the third image generation model M3 for all rough images. The frame information acquisition module 214 generates consistent intermediate images that commonly display, for example, the character's face, hairstyle, clothing, texture, coloring (including shading), contour lines, and perimeter lines. The frame information acquisition module 214 also generates an intermediate image of a character that reflects, at a more detailed level, the facial expression, posture, pose, and small items (umbrella) that are characteristically shown in the rough line drawing. The frame information acquisition module 214 generates an intermediate image that depicts such a character with the same angle of view and position as shown in the rough line drawing.

[0108] 15(A) to 15(C) are diagrams showing intermediate image information generated in another embodiment. For reference, FIGS. 15(B) and 15(C) respectively show intermediate image information generated when intermediate image information is generated from the same (A) rough line drawing information 224 as in FIG. 14 using an image generation model in which information about another character 1 or character 2 is set. In FIGS. 15(B) and 15(C), too, from a very roughly drawn rough line drawing, an image of a character equipped with the facial expression, posture, pose, and small item (umbrella) characteristically shown in the rough line drawing can be generated with the same angle of view and arrangement as shown in the rough line drawing. However, the depicted person is another character 1 or character 2 set in the image generation model, and is drawn with different contours, external shapes, etc.

[0109] 16(A) and 16(B) are schematic diagrams showing a method for generating intermediate image information in another embodiment. The frame information acquisition module 214 may be configured to accept masking processing indicating that an image should not be generated for a specific region in the rough line drawing information 224. The blacked-out region in the rough line drawing information 224 in FIG. 16(A) is the masked region. In this example, the background portion of the rough line drawing information 224 is masked. By applying such masked rough line drawing information 224 to the third image generation model M3, intermediate image information 225 can be generated, as shown in FIG. 16(B), which consists of only the character and the small item (umbrella) carried by the character, without a background.

[0110] Next, the frame information acquisition module 214 inputs the generated multiple pieces of intermediate image information into the fourth image generation model M4 to generate multiple pieces of line drawing information 226 for the character (S920). The setting of the fourth image generation model M4 and the specification of the multiple pieces of intermediate image information as input image information can be performed in the same manner as in step S910. This makes it possible to obtain uncolored line drawing information 226 for the character in each frame image.

[0111] 17(A) to 17(C) show examples of drawings created in the frame information generation process in one embodiment. In the frame information generation process, Fig. 17 shows a rough line drawing (A) recorded in rough line drawing information 224, an intermediate image (B) recorded in intermediate image information 225, and a line drawing (C) recorded in line drawing information 226.

[0112] The rough line drawing (A) uses an extremely small number of lines to distinctively express the angle and expression of the character's face. The rough line drawing (A) does not contain color information. The lines in the rough line drawing (A) are generally of a uniform thickness. Even if the thickness of the lines varies, the rough line drawing (A) is made up of several types of lines (for example, two or three types). In this rough line drawing (A), the pupils are drawn large below the bangs, and the mouth and nose are not drawn. The area where the bangs contain hair is shown, but the individual hairs are not drawn.

[0113] The intermediate image (B) generated based on the rough line drawing depicts the character at the same angle and size, based on the characteristics of the character set in the image generation model, in accordance with the characteristics shown in the rough line drawing. By incorporating color information, intermediate image (B) expresses the character's three-dimensional shape, shading, etc. in a more detailed, anime-like manner. When viewed alone, intermediate image (B) appears to be almost identical to a finished frame image. In this intermediate image (B), the character's eyes are drawn large and in detail below the bangs. The hair is expressed with shades of color to create a sense of three-dimensionality, and the shape of the shading expresses numerous strands of hair. The nose and cheeks, which are not present in the rough line drawing (A), are drawn modestly in their original positions to balance the eyes. Additionally, ears, which are not present in the rough line drawing (A), are drawn in their original positions to balance the eyes. The mouth is not drawn because its original position would not fit on the screen.

[0114] The line drawing (C) has the same angle of view and size as the intermediate image (B), and represents the outlines and shapes of the characters in the intermediate image (B) as a line drawing without color information. In addition to the outlines and shapes represented in the intermediate image (B), the line drawing (C) includes lines that are represented by color information. The line drawing (C) is made up of lines of various thicknesses. Compared to the rough line drawing (A), the line drawing (C) includes a much larger number of lines that reflect the characteristics of the character (for example, five or more times as many lines). The line drawing (C) uses a wider variety of line thicknesses than the rough line drawing (A) (for example, two or more times as many, five or more times as many). The line drawing (C) has a finished look similar to the original drawing without color, for example.

[0115] Next, the frame information acquisition module 214 generates a plurality of frame information 228 by inputting a plurality of line drawing information 226 about the character to a second image generation model M2 for which coloring information 227 is specified (S930). The second image generation model M2 functions as a coloring model.

[0116] FIG. 19 is an example of an image coloring screen. 19 on the display of the editing terminal 101. The frame information acquisition module 214 displays, for example, an image coloring screen 1900 shown in Fig. 19. On the image coloring screen 1900, the frame information acquisition module 214 displays, for example, an input information field 1910, a coloring information field 1920, a color sample field 1930, a character display field 1940, a part detection field 1950, a line drawing display field 1960, and a frame display field 1970. The input information field 1910 and the coloring information field 1920 are, for example, file selection dialogs.

[0117] The frame information acquisition module 214 receives, for example, designation of line drawing information 226 relating to a plurality of line drawings to be colored and coloring information 227 indicating the coloring conditions thereof via an input information field 1910 and a coloring information field 1920. In this way, the frame information acquisition module 214 acquires information on a plurality of rough original line drawings and color information.

[0118] The color information is information about the color to be used to color a character (or an object) drawn in a line drawing, and can be set for each part of the character, for example. The color information is information that specifies the basic color for each part of the character, for example. The color information may be determined in advance, or may be determined and / or changed on the image coloring screen 1900. Although not limited to this, the color information may be associated with model information (e.g., a 3D model or a 2D model) of the target character or object. Here, the coloring information stores colors corresponding to six parts of the character: hair, skin, clothing 1 (a cut and sewn top), clothing 2 (a skirt), clothing 3 (stockings), and shoes.

[0119] On the image coloring screen 1900, the frame information acquisition module 214 displays the colors recorded in the coloring information in, for example, the color sample field 1930. For example, the frame information acquisition module 214 displays the colors stored for the character's hair, skin, clothing 1 (cut and sewn top), clothing 2 (skirt), clothing 3 (stockings), and shoes in the top six boxes of the color sample field 1930, in that order. The bottom four boxes display a basic color, indicating that no color has been recorded. On the image coloring screen 1900, the frame information acquisition module 214 can accept the specification of a new coloring range (part) and color. For example, when the sleeve portion of clothing 1 (cut and sewn top) is specified as a new part (clothing 4), the frame information acquisition module 214 is configured to accept the sleeve portion as the new part (clothing 4) and reset the original range of clothing 1 to the remaining clothing portion of the cut and sewn top. The frame information acquisition module 214 is configured to be able to accept, for example, a designation of a color for the new part, garment 4 (sleeve portion of the cut-and-sew), that is the same as or different from garment 1 (bodice portion of the cut-and-sew). For color information of the new part, the frame information acquisition module 214 may accept, for example, input of a color code using text. Alternatively, the frame information acquisition module 214 may be configured to, for example, display a color palette, accept a color selection by the creator on the displayed color palette, and accept the selected color as color information for the new part.

[0120] Furthermore, when the coloring information 227 is associated with model information (for example, a three-dimensional model or a two-dimensional model) of a target character (which may be an object), the frame information acquisition module 214 may be configured to display the corresponding character in the character display field 1940, for example, based on the model information. Furthermore, the frame information acquisition module 214 may be configured to display each part of the character by coloring it with a color designated for that part, based on the coloring information 227. The frame information acquisition module 214 may be configured to display a color with shading or the like by reflecting unevenness, shading, or the like based on the model information on the designated color, rather than solidly coloring each part of the character with the designated color.

[0121] For example, when a box in the color sample column 1930 is selected, the frame information acquisition module 214 may be configured to accept changes to the coloring information 227, such as changes to the color specification conditions for each part or settings for specifying a new part color. When the coloring information 227 is changed, the frame information acquisition module 214 changes the coloring of the character parts displayed in the character display column 1940 according to the changed coloring information 227.

[0122] The frame information acquisition module 214 displays, for example, a line drawing stored in the line drawing information 226 (for example, a line drawing corresponding to the first frame image) in a line drawing display field 1960. When the frame information acquisition module 214 receives an image generation instruction by, for example, a creator selecting an image generation button, it inputs the specified line drawing information 226 to the second image generation model M2 and generates frame information 228.

[0123] Here, the second image generation model M2 detects the character's parts in the line drawing and colors the detected parts mainly in the color specified by the coloring information 227. The second image generation model M2 reflects the unevenness and shading of each character part based on the overall information and partial information, and colors the parts mainly in the corresponding color. The frame information acquisition module 214 displays, for example, the parts that have been detected and colored in the part detection field 1950. The frame information acquisition module 214 also displays the colored frame (for example, the frame corresponding to the first frame image) in the frame display field 1970.

[0124] In this way, it is possible to prevent the consistency of the coloring of a character in the same scene from being lost by coloring each part based on the coloring information 227. Furthermore, by using the same coloring information 227 in multiple scenes, it is possible to maintain consistency in the coloring of a character even in different scenes.

[0125] In coloring using an image generation model, the color to be colored has traditionally been specified using prompts. However, specifying a color using prompts has had the drawback of making it difficult to specifically specify the combination of the color to be colored and its color range. When specifying a color using prompts, for example, the color is specified by a text combination of a rough part name and a rough color name, making it difficult to maintain color consistency across multiple frame images. This is true even when an image generation model pre-trained using colored images colored with a specified color as training data is used. Furthermore, an image generation model that has learned a predetermined coloring style has also been used to generate a colored image with a higher resolution and a predetermined coloring style based on a manually colored image that is roughly colored by hand from a line drawing. However, this method requires each frame image to be roughly colored manually, resulting in a drawback of limited flexibility. In contrast, the present technology uses a second image generation model M2 to color using the color explicitly specified in the coloring information 227. Furthermore, since the coloring information 227 does not involve a change in shape, the degree of change from the designated color is small, and the image can be colored in the intended color.

[0126] Furthermore, if the image input to the second image generation model M2 is a colored image, a process of overlaying a line drawing on the color information is required, which may result in an unstable colored image. In contrast, by inputting the image to the second image generation model M2 as a colorless line drawing before coloring, the degree of discrimination of lines, i.e., part regions (boundaries), can be improved, and the lines of the input image can be accurately extracted in the output image, allowing the intended coloring to be performed accurately.

[0127] For example, when the frame information acquisition module 214 receives a selection of the play button in the frame display field 1970, it may be configured to continuously display the frames generated in the frame display field 1970 at a predetermined frame rate based on the plurality of frame information 228. This allows the creator to check the movement of the colored characters for a predetermined scene. The frame information acquisition module 214 stores the generated frame information 228 in the auxiliary storage device 202, for example.

[0128] The background image information generating step (S540) will be described. [Generating background image information 1] FIG. 11 shows an example of a background image information generation flow 1100. In the background image information generation process, background image information for each scene is created. Here, the background image information generation module 215 executes the following steps. That is, in the background image information generation process, for example, original rough line drawing information including the background is acquired (S1110), a background image portion is extracted from the original rough line drawing information including the background (S1120), the extracted background image portion and text information are input to a fifth image generation model (background image generation model) (S1130), multiple pieces of background image information are generated (S1240), and background image information that satisfies conditions is extracted from the multiple pieces of generated background image information (S1250).

[0129] The background image information generation module 215 acquires, for example, original rough line drawing information including the background (S1110). Here, the original rough line drawing including the background refers to a roughly drawn line drawing that indicates the general configuration of the background. Furthermore, the background is not limited to the scenery drawn behind the character, but may also be a concept that includes, for example, things drawn in front of the character (including the foreground, overlying backgrounds such as bushes, background materials such as clouds and rain, etc.). Original rough line drawing information including the background is typically prepared for each scene.

[0130] The original rough line drawing information including the background may be, for example, information about an original rough line drawing of a background created by a creator. The original rough line drawing including the background may be an image with an accuracy (roughness) corresponding to an image known as a "background original drawing," "layout," or the like. Furthermore, the original rough line drawing including the background may be, for example, an image indicating the visual characteristics and arrangement of elements such as "background," "characters," and "other decorations" that make up a "frame image." For example, if the above-described original rough line drawing information 223 or rough line drawing information 224 includes a background, the original rough line drawing information including the background may be this original rough line drawing information 223 or rough line drawing information 224. One or more pieces of original rough line drawing information including a background may be acquired.

[0131] Next, the background image information generation module 215 extracts the background image portion from the original rough line drawing information including the background (S1120). For example, if the original rough line drawing information including a background is the original rough line drawing information 223 or the rough line drawing information 224 and the image includes a character, the background image information generation module 215 extracts the background image portion excluding the image portion of the character. The original rough line drawing information 223 may be masked on the image portion of the character, as opposed to the example shown in FIG. 16(A), and the unmasked portion may be extracted as the background image portion. This configuration can suppress misalignment when superimposing a character image and a background image in the subsequent frame image information generation step S550. Furthermore, for example, in a scene in which the camera viewpoint moves, the movements of the background and the character can be better aligned even though information about the background and information about the character are generated separately.

[0132] Furthermore, the background image information generation module 215 may create a plurality of pieces of background image information from, for example, one piece of original rough line drawing information. For example, the background image information generation module 215 may separately extract, from one piece of original rough line drawing, a background arranged behind a first character (for example, a first background) and a background arranged in front of the first character (for example, a second background). For example, the background image information generation module 215 may separately extract, from a single piece of original rough line drawing, a background arranged behind the first character (for example, a first background), a background arranged in front of the first character and behind a second character (for example, a second background), and a background arranged in front of the second character (for example, a third background). The plurality of pieces of separately extracted background image information are associated with, for example, each scene.

[0133] The background image information generation module 215 may extract the entire background represented by the original drawing rough line drawing information including the background, or may extract a portion of the background represented by the original drawing rough line drawing information including the background. When extracting the entire background, the background image information generation module 215 may omit step S1120 of extracting this background image portion. For example, when acquiring information about a background original drawing (e.g., a separate background sheet) that contains only the background and does not include characters as the original drawing rough line drawing information including the background, step S1120 can be omitted.

[0134] Then, the background image information generation module 215 inputs the extracted background image portion and text information into the fifth image generation model (S1130). The text information is information used as input to the fifth image generation model. As described above, the text information is information that verbally describes the type and characteristics of the background and elements included in the background (sky, trees, buildings, rooms, furniture, small items such as vehicles, etc.). Examples of such text information include a school classroom, a school hallway, a brick building, and a bench under a large tree in the park. The text information may be information that verbally expresses the characteristics of the background in a positive manner (e.g., long branches, white clouds, etc.) or information that verbally expresses the characteristics of the background in a negative manner (e.g., no leaves, not a single cloud, etc.). The text information may include, for example, color information, information instructing the drawing style of the background (e.g., identification information for identifying various shapes of drawing brushes), etc. The background image information generation module 215 may be configured to accept and set color information about the main color of each background element, for example, using the image coloring screen 1900 shown in FIG. 19.

[0135] The fifth image generation model may be, for example, a model in which overall information about the style of the target animation is set in the image generation model. Alternatively, the fifth image generation model may be, for example, a model in which a function for creating a highly accurate colored background image from a roughly drawn background line drawing is set in the image generation model. The fifth image generation model may be realized by a single image generation model, or may be realized by combining two or more image generation models.

[0136] The function of creating a highly accurate colored background image from a roughly drawn background line drawing can be realized, for example, by setting an additional learning layer that is trained to input images of various landscapes (including photographs and illustrations) and output their contour lines and characteristic lines, or a function that colors areas segmented by the contour lines and characteristic lines with a predetermined drawing style, although this is not limited to this. Also, the function of creating a highly accurate background image from a roughly drawn background line drawing may use, for example, the third or fourth image generation model M3 or M4 in which overall information and / or partial information about a specific character is not set.

[0137] As a result, the background image information generation module 215 generates multiple pieces of background image information as output from the fifth image generation model (S1140). This makes it possible to obtain background image information, which is image information corresponding to a highly accurate colored background image, from the original rough line drawing information corresponding to the roughly drawn background line drawing. The background image information generation module 215 can generate background image information corresponding to a background image with a style common to the character by using the fifth image generation model in which overall information related to the style is set.

[0138] The size of the image generated by the fifth image generation model may or may not match the shape (e.g., 16:9) and size of the target animation cut or frame. The size of the image generated by the fifth image generation model can be generated, for example, larger than the target animation cut or frame. The resolution of the image generated by the fifth image generation model can also be generated, for example, at either or both standard resolution and any resolution other than standard resolution.

[0139] Furthermore, in the background image information generated by the fifth image generation model, the background image may be a background image with a style that leaves outlines, or a background image with a style that is colored so that outlines are not left. The number of background image information to be generated is not particularly limited, and may be, for example, about 2 to 10 pieces.

[0140] Next, the background image information generating module 215 extracts background image information that satisfies a condition from among the plurality of generated background image information (S1250). The background image information generation module 215 displays the generated background images in a background candidate display field or the like on the display based on the generated background image information (see, for example, FIG. 22). The background image information generation module 215, for example, accepts a selection by the creator of an image suitable as a background image from the multiple background images displayed in the background candidate display field or the like, and extracts image information representing the selected background image. At this time, the background image information generation module 215 may, for example, display the selected background image in a selected image display field or the like in a distinguished manner. The background image information generation module 215, for example, stores the extracted image information in the auxiliary storage device 202 or the like as background image information.

[0141] The background image information generation module 215 reads the conditions for determining background images and the AI ​​model, and inputs multiple background images generated into it, so that it can display only background images that meet the conditions, or highlight background images that are likely to meet the conditions, thereby distinguishing them from other background images.

[0142] Furthermore, here, background image information generation module 215 may execute step S1120 of extracting a background image portion after, or instead of, steps S1130 to S1140 of generating background image information using the fifth image generation model, either additionally or before steps S1130 to S1140. That is, background image information generation module 215 may extract a portion of the background image portion from the background image generated by the fifth image generation model, and reconstruct the background image information.

[0143] For example, the background image information generation module 215 may separately extract a background of a first size and a background of a second size that is smaller or larger than the first size from original rough line drawing information including one background. Furthermore, for example, the background image information generation module 215 may separately extract a first background and a second background that is moved in any direction within 360° from the first background from original rough line drawing information including one background. By arbitrarily combining these techniques, it is possible to realize camerawork such as gradually enlarging (tracking up or zooming up) or reducing (tracking back or zooming out) the background around a predetermined position (viewpoint), or moving the center of the background (viewpoint) along a predetermined trajectory (e.g., panoramic viewing (PAN) or tilt) the background.

[0144] Furthermore, when the overlapping relationship (front-to-back relationship within a frame) between an object in a background image and a character changes, the background image information generation module 215 may divide one or more objects into different images depending on the distance from the viewpoint (camera), for example. That is, the background image information generation module 215 may generate a plurality of pieces of background image information by dividing an object into layers. The background image information generation module 215 stores the plurality of pieces of background image information divided into layers in the auxiliary storage device 202 or the like, for example, in association with each scene.

[0145] The frame image information generating step (S550) and the moving image information generating step (S560) will be described. The compositing module 216 generates multiple pieces of frame image information. Based on the frame information for the characters created in step S540 and the background image information created in step S550, the compositing module 216 generates multiple pieces of frame image information representing an image in which a background and a character are combined. For example, the compositing module 216 overlays a character image on a background image. For example, based on storyboard information, the compositing module 216 generates multiple pieces of frame image information representing an image in which a plurality of backgrounds divided into layers and characters are appropriately combined. For example, the compositing module 216 overlays the background image and the character image while relatively changing the angle, magnification (size), and position as necessary, and compares the result with the display in the storyboard cut display field 2225.

[0146] The compositing module 216 then creates frame image information representing an image in which the background and character are combined at a predetermined aspect ratio (e.g., 16:9) at an angle, magnification, and position of the character image such that the display of the character superimposed on the background of the selected image display field 2240 matches the display in the storyboard cut display field 2225. The compositing module 216 then generates video information based on the multiple frame image information created in step S550. This allows an animation to be created.

[0147] The composition module 216 may be configured to create frame image information, for example, by combining a background image and a character image, larger than a predetermined frame image size, and then extract from this large frame image a plurality of frame image information pieces of a predetermined size at positions and angles of view corresponding to the moving image where the camera viewpoint moves. This allows for efficient creation of moving images. Furthermore, if the camera viewpoint moves relative to the background image or the character moves relative to the storyboard cut, the position and movement path of the character in the 3D model, as well as the camera work (movement path of the camera's virtual viewpoint), etc. can be specified in advance at the storyboard creation stage, and the background image from the camera's virtual viewpoint and the character image can be geometrically combined.

[0148] <Action and effect> According to the animation creation method configured as above, multiple pieces of rough line drawing information are acquired, and based on the multiple pieces of rough line drawing information, multiple frames of information about the character are generated using one or multiple image generation models in which character information is set, and video information is generated based on the multiple frames of information. This makes it possible to easily create animation while suppressing variations in the coloring of the character, for example.

[0149] Furthermore, in the animation creation method described above, generating multiple frames of information about a character includes generating multiple lines of information about the character using a first image generation model in which character information is set based on multiple pieces of rough line drawing information, and generating multiple frames of information using a second image generation model based on the multiple pieces of line drawing information about the character. This allows, for example, for a character, line drawing information with "original" level accuracy to be easily created based on rough line drawing information with "rough" level accuracy. In the past, the manual process of creating "original" from "rough" required skilled techniques, and the process of creating highly accurate original drawings was inevitably highly dependent on the individual. In contrast, the animation creation system 1 allows line drawing information at the "original" level (line drawing in the present technology) to be obtained from "rough" (rough line drawing in the present technology) without the need for human labor.

[0150] Furthermore, in the above animation creation method, generating a plurality of frame information for a character includes generating a plurality of intermediate image information for the character using a third image generation model in which character information is set, based on a plurality of rough line drawing information, generating a plurality of line drawing information for the character using a fourth image generation model that generates line drawings, based on the plurality of intermediate image information for the character, and generating the plurality of frame information for the character using a second image generation model, based on the plurality of line drawing information for the character. According to the findings of the inventors, when using an image generation model to generate an image with the accuracy of an "original drawing" from an image with the accuracy of a "rough drawing," it is possible to obtain line drawing information closer to the target original drawing by inputting the generated image (intermediate image) into the image generation model again, rather than just inputting the rough line drawing information into the image generation model once to generate the image. In other words, it is preferable to process the image using the image generation model two or more times. This allows for the acquisition of more accurate line drawing information (corresponding to the original drawing).

[0151] Furthermore, the animation creation method can include acquiring one or more unset image generation models in which character information has not been set, setting overall information, which is information about the character's appearance and style, and partial information about some of the character's features, in the unset one or more image generation models, and acquiring one or more image generation models in which character information has been set. This makes it possible to create animation that has consistency in the style, character features, etc. throughout the entire process while using image generation models.

[0152] In the animation creation method, the overall information includes information representing a two-dimensional anime image as information about the style, and one or more image generation models generate a plurality of frames of information relating to the two-dimensional anime-style image of the character. Generally used image generation models are essentially deeply trained to represent objects such as people and objects as photorealistic images (e.g., photographic images). With the above configuration, it is possible to depict objects such as backgrounds and objects throughout the animation as, for example, two-dimensional images that reflect a specific style and idiom specific to animation.

[0153] In the animation creation method, the partial information includes information representing at least one of the personality and behavior of the character, and one or more image generation models generate multiple frames of information relating to images that reflect at least one of the personality and behavior of the character. With this configuration, it is possible to generate images that reflect the characteristics of each character throughout the entire animation.

[0154] In the above animation creation method, The third image generation model is an image generation model trained to generate a second image in which the outline or shape of an object in the first image is displayed by lines, and to which overall information, which is information about the character's outer shape and style, and partial information about some of the character's features are set, The fourth image generation model is an image generation model trained to generate a second image that displays the outline or outer shape of an object in a first image using lines and does not display anything other than lines, and to which overall information, which is information about the outer shape and style of the character, and partial information about some features of the character are set, Generating multiple frames of information for a character includes: generating the plurality of colored intermediate image information of the character using a third image generation model based on the plurality of rough line drawing information; generating the plurality of pieces of color-reduced line drawing information of the character using a fourth image generation model based on the plurality of pieces of colored intermediate image information; This includes:

[0155] According to the inventors' investigations, the accuracy of the next line drawing can be improved by generating intermediate images based on multiple rough line drawing information as colored images. Also, it is preferable to generate line drawings with accuracy corresponding to the "original" as color-reduced images and then use a coloring image generation model to color all the line drawings at once, as this allows for consistency in the coloring aspect.

[0156] In the animation creation method, the second image generation model is an image generation model trained to identify parts of an object in the first image and generate a second image in which each part is colored with a designated color, and generating a plurality of frames of information includes generating the plurality of frames of information in which each part is colored with a pre-designated color by the second image generation model based on a plurality of line drawing information about a character. This makes it possible to reduce variations in coloring of characters, etc. in each frame image, for example, for the same scene or multiple different scenes.

[0157] In the animation creation method described above, the acquisition of multiple pieces of rough line drawing information is achieved by inputting multiple pieces of original rough line drawing information into an interpolation image generation model and generating multiple pieces of rough line drawing information by temporally continuously interpolating between the multiple pieces of original rough line drawing information. This makes it possible to easily create animation of scenes corresponding to a storyboard, for example, based on "several rough drawings" in the storyboard.

[0158] The animation creation system 1 configured as described above allows for the creation of animations of new stories using an image generation model, reducing the amount of manpower and effort required by the creator. Furthermore, by setting overall information and partial information for the image generation model, the feature quantities of the images generated by the image generation model can be fixedly controlled. For example, by setting art style information corresponding to information representing a 2D animated image as overall information, image information representing a given 2D animated image style can be generated throughout a single work. Furthermore, by setting partial information for the image generation model that allows each character to be identified, individual characters can be drawn stably and distinctly while maintaining a common art style for all characters. When using hand-drawn illustrations, the creator's individuality (habits) may be apparent in the illustrations. However, the animation creation system 1 allows for the creation of animations while suppressing unintended image variations (habits).

[0159] <Modification> [Generating background image information 2] The background image information generating step (S540) is not limited to the above example, and can be performed by other methods. FIG. 21 shows another example of a background image information generation flow 2100. In the background image information generating step, the background image information generating module 215 may execute the following steps. That is, for example, the background image information generating step acquires three-dimensional information about the background (S2110), cuts out an image with a field of view that matches the storyboard from the three-dimensional information (S2120), inputs the cut-out image and text information into a second image generation model (S2130), generates multiple pieces of background image information (S2140), and extracts background image information that satisfies conditions from the multiple pieces of generated background image information (S2150). Steps S2140 and S2150 are similar to the above example, so repeated description will be omitted.

[0160] The background image information generation module 215 generates background image information based on, for example, three-dimensional information about the background. The three-dimensional information about the background is information for displaying a 3D model representing a background scene, and a scene when any region of this 3D model is viewed from any viewpoint can be extracted as a consistent background. The background image information generation module 215 extracts a background image from the 3D model in accordance with predetermined storyboard information. The storyboard information is information representing a storyboard consisting of a table illustrating each scene, which is prepared before the animation is created.

[0161] FIG. 22 shows an example of a background image information generation screen 2200. The background image information generation module 215 displays, for example, a three-dimensional information file selection field 2210 and a storyboard information file selection field 2220 on the background image information generation screen 2200, and as already explained, receives designation of a three-dimensional information file and a storyboard information file related to the background to be acquired via these file selection fields 2210, 2220. Based on the received three-dimensional information file and storyboard information file, the background image information generation module 215 displays a 3D model related to the background in a 3D model display field 2215 and a storyboard for one cut in a storyboard cut display field 2225.

[0162] Next, when the background image information generation module 215 receives the creator's selection of the generate button 2250, it compares the displays in the 3D model display field 2215 and the storyboard cut display field 2225 while changing the angle, magnification, position, etc. of the display axis of the 3D model, and cuts out the image displayed in the 3D model display field 2215 at the angle, magnification, and position of the 3D model where the display in the 3D model display field 2215 best matches the display in the storyboard cut display field 2225.

[0163] The background image information generation module 215 inputs the extracted image and text information that provides instructions for appropriately coloring the image into, for example, a fifth image generation model, and generates multiple pieces of background image information. While not limited to this, the fifth image generation model may be, for example, an image generation model similar to the image generation model used to generate image information for the character, and may be, for example, one in which overall information is set in DDPM. This allows background image information for the background image to be generated in a style that is common to the character. The background image information generation module 215 displays the generated multiple background images in the background candidate display field 2230 based on the generated multiple pieces of background image information.

[0164] Background image information generation module 215 accepts a creator's selection of an image suitable as a background image from among the multiple background images displayed in background candidate display field 2230, displays the selected background image in selected image display field 2240, and extracts image information representing the selected background image. For example, when background image information generation module 215 accepts the creator's selection of save button 2260, it stores the extracted image information in auxiliary storage device 202 or the like as background image information.

[0165] Subsequently, the background image information generation module 215 executes a step of generating a plurality of pieces of background image information (S2140) and a step of extracting background image information that satisfies a condition from the plurality of pieces of generated background image information (S2150). This method also makes it possible to generate background image information suitably.

[0166] <Embodiment 2> In the above-described first embodiment, the frame information acquisition module 214 uses the third image generation model M3 and the fourth image generation model M4 to generate intermediate image information 225 from the rough line drawing information 224, and then generates line drawing information 226. However, it is not always necessary to generate the intermediate image information 225.

[0167] In the second embodiment, the frame information acquisition module 214 uses a first image generation model M1 to generate line drawing information 226 directly from rough line drawing information 224, as shown in Fig. 10. The other configurations and effects can be implemented in the same way as in the first embodiment, and therefore will not be described again.

[0168] In this case, the model setting module 212 sets the first image generation model M1 in the image generation model setting step (S510) of the first embodiment. That is, the model setting module 212 can realize the function of the first image generation model M1 by setting, for example, overall information about the target animation, overall information and / or partial information about a specific character, and uncolored line drawing information for the image generation model. The first image generation model M1 has character information set therein, and generates multiple pieces of line drawing information about the character. The first image generation model M1 can be used, for example, to generate the line drawing information 226.

[0169] FIG. 8 shows an example of a flow of generating frame information. In the second embodiment, in the frame information generating step (S530), the frame information acquisition module 214 executes the following frame information generating flow 800.

[0170] The frame information generation flow 800 includes generating a plurality of pieces of line drawing information 226 about a character using an image generation model (first image generation model M1) that outputs a line drawing from a rough line drawing based on a plurality of pieces of rough line drawing information 224 (S810), and generating a plurality of pieces of frame information 228 using a model (second image generation model M2) that colors the line drawing based on the plurality of pieces of line drawing information 226 about the character (S820). Step S820 can be performed in the same manner as step S920 in the first embodiment, for example.

[0171] FIG. 18A is a schematic diagram showing the generation of image information. When the rough line drawing information 224 is input to the first image generation model M1, line drawing information 226 as shown in Fig. 18(A) is generated. This line drawing may more strongly reflect the characteristic lines drawn in the rough line drawing than the line drawing of Fig. 17(C) generated via an intermediate image, for example. The first image generation model M1 generates a line drawing that includes more characteristic lines, for example.

[0172] Therefore, when creating an animation with the style shown in FIG. 18(A), it is advisable to create frame information in accordance with the frame information generation flow 800. On the other hand, if you want to obtain line drawings with a more delicate touch depending on the style of the target animation, more detailed tuning of the first image generation model M1 may be required. Therefore, it is recommended to create frame information according to the frame information generation flow 900 shown in the first embodiment, for example.

[0173] <Modification> In the above embodiment, the first image generation model M1 and the fourth image generation model M4 were set to generate color-reduced (e.g., color-free) line drawing information 226. However, the first image generation model M1 and the fourth image generation model M4 may be set to generate colored line drawing information 226. Although not limited to this, in this case, for example, the third image generation model M3 may be used as the first image generation model M1 and the fourth image generation model M4. This allows the configuration of the editing terminal 101 to be simplified.

[0174] In the above embodiment, the third image generation model M3 was configured to generate color-reduced intermediate image information 225. However, the third image generation model M3 may also be configured to generate color-reduced (e.g., color-free) intermediate image information 225. When the rough line drawing information 224 is input to the third image generation model M3 configured to generate a color-reduced image and then input to the fourth image generation model M4 configured to generate a color-reduced image, line drawing information 226 as shown in FIG. 18(B) is generated. The resulting line drawing may tend to have fewer delicate lines representing hair and eyes than the line drawing of FIG. 17(C) generated via a color intermediate image, for example. When creating an animation with the style shown in FIG. 18(B), it is recommended to create frame information using the third image generation model M3 configured to generate a color-reduced image. In this case, for example, the third image generation model M3 and the fourth image generation model M4 may be the same image generation model configured to generate a color-reduced image. This allows the configuration of the editing terminal 101 to be simplified.

[0175] In the above embodiment, the rough line drawing information 224 and the original drawing rough line drawing information 223 are prepared for a character in a cut that represents an area according to the storyboard. For example, as shown in FIG. 17(A), if a close-up of a character's face is drawn in the storyboard, the rough line drawing information 224 and the original drawing rough line drawing information 223 represent a close-up of the character's face. However, the rough line drawing information 224 and the original drawing rough line drawing information 223 may be information about an image that represents an area larger than the area represented in the storyboard. Furthermore, the rough line drawing information 224 and the original drawing rough line drawing information 223 may be information about an image that represents an area larger than the area represented in the storyboard, and may include information indicating the area represented in the storyboard.

[0176] FIG. 18C is a diagram showing an example of the contents of another original rough line drawing information 223, for example. The original rough line drawing is the area indicated by the dashed line in FIG. 18(C) and includes part of a character's part (e.g., the head) shown in the storyboard. In such a case, the original rough line drawing information 223 may be information representing an image including the entire character's part (e.g., the head) while including part of the character's part (e.g., the head) at the same angle of view as the original rough line drawing, as shown in FIG. 18(C). In this case, the original rough line drawing information 223 may also be provided with information indicating the area of ​​the original rough line drawing (e.g., position information for the upper left and lower right corners of the dashed line). This configuration can improve the quality of the rough line drawing information 224, intermediate image information 225, line drawing information 226, frame information 228, etc. generated by the image generation model.

[0177] In this case, it is also advisable to add information indicating the area of ​​the original rough line drawing (for example, position information of the upper left and lower right corners of the dash-dot line) to the rough line drawing information 224, intermediate image information 225, line drawing information 226, and frame information 228 generated by the image generation model. For example, if information indicating the area of ​​the original rough line drawing is added to the frame information 228, the synthesis module 216 is configured to cut out the image represented by the frame information 228 according to the information indicating the area of ​​the original rough line drawing, and synthesize it with the background image information 229 to generate a frame image. This makes it possible to create a highly accurate animation even for a scene that only shows part of a character.

[0178] In the above embodiment, the coloring information 227 is, for example, information about the color of a character's parts and is associated with the character's parts. This coloring information 227 may be associated with model information (e.g., a three-dimensional model or a two-dimensional model) of the character (which may be an object) that is the subject of the coloring information 227.

[0179] FIG. 20 is a schematic diagram showing an image coloring method according to another embodiment. For example, if coloring information 227 is associated with a three-dimensional model of an (A) character, the frame information acquisition module 214 may be configured to use the second image generation model M2 to generate colored (C) frame information 228 based on the (B) line drawing information 226 and in accordance with the coloring information 227 associated with the (A) three-dimensional model. Also, the frame information acquisition module 214 may be configured to use the second image generation model M2 to generate colored (C) frame information 228 based on the (B) line drawing information 226 and in accordance with the (A) three-dimensional model including the coloring information 227.

[0180] In the above embodiment, the frame information acquisition module 214 is configured to generate an image of a region, etc., that the image generation model generates is basically configured to generate an image of the same region as the region represented by the input image information. However, the region, size, aspect ratio, resolution, etc., of the image generated by the image generation model are not limited to the region, size, shape (e.g., aspect ratio 16:9), resolution, etc., of the target animation cut or frame. The frame information acquisition module 214 can generate an image of, for example, any region, size, shape, or resolution. Furthermore, the frame information acquisition module 214 can set the resolution of the image generated by the image generation model to, for example, either a standard resolution or any resolution other than the standard resolution, or both.

[0181] In the above embodiment, the animation creation method includes a step (S510) of setting an image generation model, and for example, the model setting module 212 sets overall information and partial information for the image generation model. However, if a set image generation model in which this information has already been set is available, it can be understood that the image generation model has been set by obtaining such a set image generation model.

[0182] In the above embodiment, when setting the overall information and partial information, the model setting module 212 applied a prepared overall information file and partial information file to the image generation model. However, the model setting module 212 may set the overall information by training the image generation model using images of a character with a predetermined shape and / or style. Alternatively, the model setting module 212 may set the partial information by introducing a trainable model and setting information into the image generation model and training the model to further learn the characteristics of a predetermined part of the character.

[0183] In the above embodiment, the model setting module 212 sets overall features and partial features for the image generation model in order to generate image information of a character having predetermined overall features and partial features. However, the model setting module 212 is not limited to generating image information of a character, and can set overall features and partial features for the image generation model when generating image information of other images such as backgrounds, for example. This makes it possible to create image information of a consistent style in the image generation model throughout the creation of a single animation.

[0184] In the above embodiment, the model setting module 212 and the background image information generation module 215 each generate frame information and background image information for a character based on separate images, etc., and the compositing module 216 combines these to generate frame image information. However, the generation of frame information by the model setting module 212 and the generation of background image information by the background image information generation module 215 may be performed based on the same image information or video information. For example, storyboard information depicting characters and backgrounds for animation is acquired as the original rough line drawing information 223. Then, the model setting module 212 generates frame information for the character using a first image generation model based on a portion corresponding to one of the storyboard information. Furthermore, the background image information generation module 215 generates background image information using a fifth image generation model based on a portion of the source video corresponding to the same target scene (cut).

[0185] With this configuration, the composition module 216 can generate frame image information with a composition corresponding to the original source video by combining the generated character image information and background image information. As a result, even if the character image information and background image information are generated using different image generation models, the relative three-dimensional positions, angles, and overlaps can be accurately reproduced during composition, making it possible to create accurate video information more quickly and easily.

[0186] In the above embodiment, each screen of the animation creation application is depicted as a separate screen. However, these screens and any combination of two or more of the functions shown on these screens may be configured to be displayed on the same page. Furthermore, each screen may be configured so that, for example, the animation creation module 211 and each of the modules 212 to 216 cooperate with each other to display an appropriate screen on the display of the editing terminal 101 according to the progress of the animation creation method. For example, the processing functions executed by each module may be represented by a node-based UI.

[0187] Furthermore, the animation creation module 211 and each of the modules 212 to 216 may cooperate with each other to, for example, display tab icons representing each screen above each screen, and when this tab icon is selected by the creator, display the screen corresponding to the selected tab icon on the display of the editing terminal 101.

[0188] In the above embodiment, when the creator selects the file selection field, for example, the animation creation system 1 displays a file selection dialog and accepts the designation of an arbitrary information file. However, the method of accepting the designation of an information file by the creator is not limited to this example. For example, when an information file is dragged and dropped onto the preview window 1270, the rough line drawing information display field 1320, the rough line drawing information display field 1460, the character display field 1940, or the like, the information file can be accepted as the information file to be designated.

[0189] Note that the configuration of each screen and the configuration and functions of the buttons provided on each screen are not limited to the examples disclosed in the above embodiment. For example, in the above embodiment, after the model setting module 212 accepts text information input by the creator into the positive instruction information field 1250 and the negative instruction information field 1260 on the image generation model setting screen 1200, when the model setting module 212 accepts selection of the preview button 1280 by the creator, the model setting module 212 causes the image generation model to generate image information of a character. However, for example, the model setting module 212 may be configured to, for example, cause the image generation model to generate image information about an image that reflects the content of the accepted text information each time it accepts input of text information by the creator into the positive instruction information field 1250 and the negative instruction information field 1260. Furthermore, the model setting module 212 may be configured to display the generated image in the preview window 1270 each time the image generation model generates an image.

[0190] In the above embodiment, the background image information generation module 215 generates background image information representing an image related to a background using the fifth image generation model. However, the background image information generation module 215 may generate background image information using, for example, computer graphics without using an image generation model.

[0191] In the above embodiment, the editing terminal 101 executes the animation creation flow under the creator's sequential instructions. However, the creator's sequential instructions are not necessarily required to execute the animation creation flow. The creator's instructions may be recorded in advance as operating conditions, etc. In this case, the editing terminal 101 may be caused to sequentially execute some or all of the steps S510 to S550 of the animation creation flow in accordance with such recording. This allows the animation to be created automatically. For example, the animation can be automatically created based on one or more prepared storyboards.

[0192] In the above embodiment, the rough line drawing information 224 is prepared by complementing the original rough line drawing information 223 with the interpolated image generation model Mip. However, the rough line drawing information 224 may be obtained from information prepared in advance in the auxiliary storage device 202 or the like.

[0193] In the above embodiment, the present technology has been described using an example of creating 2D animation, but animations created using the present technology are not limited to this example, and can be, for example, animations for computer games, AR animations, VR animations, etc., without particular limitations. Note that the present technology is preferably applied to animations made up of still images in a style similar to hand-drawn illustrations, known as so-called "anime," in that the advantages of the present technology are clearly exhibited.

[0194] Each module disclosed in the above embodiments may be configured by combining multiple sub-modules. Furthermore, some or all of the operations or functions performed by one module may be performed or realized by another module. In the above embodiment, one or more functional elements realized by the execution of each module of the user terminal 103 may be realized by the execution of a module of the editing terminal 101 .

[0195] The present technology not only provides an invention related to an animation creation method, but also, in another aspect, provides a program for causing a computer and / or a processor to execute each step (process) of the animation creation method. This program may be composed of a single program or may be composed of two or more subprograms. Furthermore, the program may be a program for causing a computer to execute any one or two or more of the steps and / or substeps of the above method.

[0196] Furthermore, the present technology not only provides an invention relating to an animation creation method, but also, in another aspect, can provide an invention relating to an animation creation device. This animation creation device comprises one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including the following instructions: acquiring a plurality of pieces of rough line drawing information; generating a plurality of frames of information about the character using one or more image generation models in which character information is set based on the plurality of pieces of rough line drawing information; and generating video information based on the plurality of frame information.

[0197] Furthermore, one or more of the one or more editing terminals 101 and one or more user terminals 103 constituting the animation creation system 1 according to the present technology may be installed in different countries. Furthermore, the editing terminal 101 may be realized by one or more computers, any of which may be installed in different countries.

[0198] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0199] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0200] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. The above-described embodiments disclose at least the configurations described in the claims. [Explanation of symbols]

[0201] 1... animation creation system, 101... editing terminal, 102... management server, 103... user terminal

Claims

1. Acquiring multiple rough line drawings, generating a plurality of frame information for the character using one or a plurality of image generation models in which character information is set based on the plurality of pieces of rough line drawing information; generating video information based on the plurality of frame information; How to create animations, including:

2. The acquisition of the plurality of pieces of rough line drawing information includes:

2. The animation creation method according to claim 1, wherein a plurality of pieces of original rough line drawing information are input to an interpolation image generation model, and a plurality of pieces of rough line drawing information are generated and obtained by temporally continuously interpolating between the plurality of pieces of original rough line drawing information.

3. generating a plurality of frames of information for the character, generating a plurality of pieces of line drawing information for the character using a first image generation model in which character information is set, based on the plurality of pieces of rough line drawing information; generating the plurality of frame information based on a plurality of line drawing information about the character using a second image generation model; 3. The animation creating method according to claim 1, further comprising:

4. generating a plurality of frames of information for the character, generating a plurality of pieces of intermediate image information for the character using a third image generation model in which character information is set, based on the plurality of pieces of rough line drawing information; generating a plurality of pieces of line drawing information for the character using a fourth image generation model that generates line drawings based on a plurality of pieces of intermediate image information for the character; generating the plurality of frame information using a second image generation model based on a plurality of line drawing information about the character; 2. The animation creation method of claim 1, comprising:

5. Obtaining one or more unconfigured image generation models in which character information is not configured; Setting overall information, which is information about the character's appearance and style, and partial information, which is information about some of the features of the character, in the one or more unset image generation models; obtaining one or more image generation models in which information about the character is set; The animation creation method according to claim 1 , comprising:

6. the overall information includes information representing a two-dimensional anime image as information about the style of the image; the one or more image generation models generate a plurality of frames of information relating to a two-dimensional anime-style image of the character; The animation creating method according to claim 5.

7. the partial information includes information representing at least one of personality and behavior of the character; the one or more image generation models generate a plurality of frame information related to images that reflect at least one of the personality and the action of the character; The animation creating method according to claim 5.

8. the third image generation model is an image generation model trained to generate a second image in which the outline or outer shape of an object in a first image is displayed by lines, to which overall information, which is information about the outer shape and style of the character, and partial information about some features of the character are set; the fourth image generation model is an image generation model trained to display the outline or outer shape of an object in a first image using lines and generate a second image that displays only the lines, and to which overall information, which is information about the outer shape and style of the character, and partial information about some features of the character are set; generating a plurality of frames of information for the character, generating the plurality of colored intermediate image information of the character using the third image generation model based on the plurality of pieces of rough line drawing information; generating the plurality of pieces of color-reduced line drawing information for the character using the fourth image generation model based on the plurality of pieces of colored intermediate image information; 5. The animation creation method of claim 4, comprising:

9. The second image generation model is an image generation model trained to identify parts of an object in a first image and generate a second image colored with a specified color for each part, generating the plurality of pieces of frame information by using the second image generation model based on a plurality of pieces of line drawing information about the character, the plurality of pieces of frame information being colored with a color designated in advance for each part of the character; 4. The animation creation method of claim 3, comprising:

10. A program for causing a computer to execute each step of the animation creation method according to claim 1.

11. an acquisition unit that acquires a plurality of pieces of rough line drawing information; a line drawing generation unit that generates a plurality of pieces of line drawing information for the character based on the plurality of pieces of rough line drawing information and using one or a plurality of image generation models in which character information is set; a moving image generating unit that generates moving image information based on the plurality of pieces of line drawing information; An animation creation system comprising:

Citation Information

Patent Citations

  • Animation data creation device and method

    JP2022141418A