Animation creation method
The animation creation method uses generative AI models to automate the generation of character and background images, addressing the challenge of manual effort in existing technologies and ensuring consistent animation across scenes.
Patent Information
- Application Number
- PCT/JP2025/013769
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-16
AI Technical Summary
Existing animation creation technologies require significant manual effort and do not provide a straightforward mechanism for creating new animations, especially in picture book-style videos where user-provided character images are applied to pre-prepared works without allowing for easy customization of character movements.
An animation creation method involving an image generation model that sets features, acquires motion and text information, generates character and background images, and combines them to create video information, utilizing generative AI models like variational autoencoders, generative adversarial networks, and diffusion models to automate the animation process.
Enables easy and efficient creation of new animations by automating the generation of consistent character and background images, reducing manual effort and enhancing the creativity and consistency of animations across scenes.
Smart Images

Figure JP2025013769_16102025_PF_FP_ABST
Abstract
Description
How to create animation
[0001] [RELATED APPLICATIONS] This application claims priority to Japanese Patent Application No. 2024-062020 entitled "Animation Creation Method," filed on April 8, 2024, the disclosure of which is incorporated herein by reference in its entirety. The present technology relates to an animation creation method.
[0002] Japanese Patent Application Publication No. 2023-002280 (Patent Document 1) is a background technology in this technical field. This publication states that "a moving image data creation device includes: a storage unit that stores format data including background images and character movement information for each of a plurality of scenes in a work; an image processing unit that receives user images from a user terminal, cuts out person images from the user images, animates the person images based on the movement information, combines the animated person images with the background image, and generates scene editing data corresponding to each of the plurality of scenes; and a combining unit that combines a plurality of the scene editing data to generate work data" (see Abstract).
[0003] Japanese Patent Application Publication No. 2022-141418
[0004] Patent Literature 1 discloses a technology for creating a picture book-style video in which a user-provided character image moves. However, the technology in Patent Literature 1 applies the provided character image to a pre-prepared picture book-style video work and transcribes predetermined movements, but does not disclose a mechanism for easily creating new animations. Therefore, the present technology provides a mechanism for easily creating new animations.
[0005] To solve the above problem, for example, the configuration described in the claims is adopted. The present application includes a plurality of configurations and methods for solving the above problem, and as an example, provides an animation creation method that sets an image generation model, acquires motion information related to object motion, acquires text information related to character features, inputs the motion information and the text information into the image generation model to generate a plurality of image information representing images of the character, generates background image information representing a background image using a second image generation model, generates a plurality of frame image information representing an image obtained by combining the image of the character with the image of the background, and creates video information based on the plurality of frame image information.
[0006] According to the present technology, it is possible to provide a mechanism that allows new animations to be created easily. Problems, configurations, and effects other than those described above will become clear from the following description of the embodiments.
[0007] FIG. 1 shows an example of the overall configuration of an animation creation system. FIG. 2 shows an example of the hardware configuration of an editing terminal 101. FIG. 3 shows an example of the hardware configuration of a server 102. FIG. 4 shows an example of the hardware configuration of a user terminal 103. FIG. 5 shows an example of an animation creation flow. FIG. 6 shows an example of an image generation model setting flow. FIG. 7 shows an example of a motion information acquisition flow. FIG. 8 shows an example of an image information generation flow. FIG. 9 shows an example of a background image information generation flow. FIG. 10 shows an example of an image generation model setting screen. FIG. 11 shows an example of a motion information acquisition screen. FIG. 12 shows an example of an image information generation screen. FIG. 13 shows another example of an image information generation screen. FIG. 14 shows another example of an image information generation screen. FIG. 15 shows another example of an image information generation screen. FIG. 16 shows an example of a background image information generation screen. FIG. 17 shows an example of a storyboard information generation screen. FIG. 18 shows another example of a background image information generation flow. FIG. 19A shows an example of a second animation creation flow. Fig. 19B is a continuation of the example of the second animation creation flow. Fig. 20 shows an example of image association information 2000. Fig. 21 is an example of an additional expression addition flow. Fig. 22 is another example of an additional expression addition flow. Fig. 23 is an example of a color tone correction flow. Fig. 24 is a schematic diagram showing an outline of additional expression addition. Fig. 25 is a schematic diagram showing an outline of additional expression addition. Fig. 26 is a schematic diagram showing an outline of color tone correction.
[0008] Hereinafter, an animation creation system and an animation creation method according to the present technology will be described with reference to the drawings. Note that in each drawing, components having the same functions may be designated by reference numerals and may not be described in detail.
[0009] [Animation Creation System] Fig. 1 is an example of a configuration diagram of an animation creation system 1 according to an embodiment. The animation creation system 1 according to the present technology is a system for creating animation data while reducing the manpower and effort required of creators. The animation creation system 1 can also support creators in creating animation data.
[0010] The term "animation" used in this technology refers to a technology that allows humans to perceive continuous movement (i.e., apparent motion) from multiple still images. This technology targets animations realized from a series of non-realistic still images, rather than moving images made up of images of actual people or scenes (e.g., live-action images). These still images may be hand-drawn or in the style of paintings drawn using computer graphics (CG) technology, etc.
[0011] Furthermore, animation data refers to data (information) for representing multiple still images that realize animation (hereinafter, sometimes simply referred to as "anime"). Animation data includes a combination (set) of information about a series of still images that change minutely, and by presenting (e.g., projecting) a series of still images sequentially at short time intervals based on this animation data, apparent motion can be perceived. Such animation data may be stored in various storage media such as film or memory. In this embodiment, the present technology will be described using an example of creating 2D animation. However, the present technology can also be used when creating 3D animation.
[0012] As shown in Fig. 1, the animation creation system 1 is composed of one or more editing terminals 101. The animation creation system 1 may additionally include, for example, one or more servers 102, one or more user terminals 103, and one or more image capture devices 104, each of which is independent of the other. The editing terminals 101, the server 102, the user terminals 103, and the image capture devices 104 are each capable of transmitting and receiving information to and from each other via a wired or wireless network. The editing terminal 101 may be separate from the server 102, or may be configured integrally with the server 102.
[0013] Each terminal and server 102 (hereinafter, the editing terminal 101, the server 102, and the user terminal 103 may be simply referred to as "terminals") of the animation creation system 1 may be, for example, a mobile terminal such as a smartphone, tablet, mobile phone, or personal digital assistant (PDA), or a wearable terminal such as glasses (including goggles), a wristwatch, or a clothing-type terminal. Each terminal may also be a stationary or mobile computer, or a server located on a cloud or network. From the perspective of functionality, each terminal may be a VR (Virtual Reality) terminal, an AR (Augmented Reality) terminal, or an MR (Mixed Reality) terminal. Alternatively, each terminal may be a combination of multiple of these terminals. For example, a combination of one smartphone and one wearable terminal may logically function as a single terminal. Each terminal may also be other information processing terminals.
[0014] Each terminal and server 102 of the animation creation system 1 may optionally include a processor that executes an operating system, applications, programs, etc., a main storage device such as RAM (Random Access Memory), an auxiliary storage device such as an IC card, hard disk drive, SSD (Solid State Drive), or flash memory, a communication control unit such as a network card, wireless communication module, or mobile communication module, input devices such as a touch panel, keyboard, mouse, audio input device, motion controller, or input device that detects motion by capturing images from a camera unit, input devices such as sensors like GPS, gyro sensor, or acceleration sensor, and output devices such as a monitor or display. Note that the output device may also be a device or terminal that transmits information to be output to an external device such as a monitor or display, printer, audio output device, or oscillator.
[0015] The main memory stores various programs and applications (software modules), and the processor executes these programs and applications to realize each functional element of the overall system. Each module may be implemented as an independent program or application, or as a subprogram or function within a single integrated program or application. Each module may also be implemented as hardware (hardware modules) by integrating circuits or using a microcomputer.
[0016] Furthermore, each module may be implemented by a single processor or multiple processors. Furthermore, each module may be provided in a single terminal (including a management server) or may be provided separately in two or more terminals (including a management server) interconnected via a network. Furthermore, each module may be provided in each or any one or more of two or more terminals (including servers) interconnected via a network. It is also contemplated that some modules may be implemented in a different country from the other modules.
[0017] In this specification, each module is described as the entity (subject) that performs the processing, but in reality, the processing is performed by a processor executing programs, applications, etc. to realize each module.
[0018] The auxiliary storage device stores various databases (DBs). A "database" is, for example, a set of data organized and collected so that it can accommodate arbitrary data operations (e.g., extraction, addition, deletion, overwriting, etc.) from a processor or an external computer. The auxiliary storage device is a functional element (storage unit) that stores one or more sets of data. The implementation method of the database is not limited to this example and may be, for example, a database management system, spreadsheet software, or a text file such as XML or JSON. Some or all of this information may be stored in a relational database or a non-relational database. The database may be provided independently of the processor, while being connectable to the processor, etc.
[0019] [Editing Terminal] Fig. 2 shows an example of the hardware configuration of the editing terminal 101. The editing terminal 101 is the main element for creating animation data. The editing terminal 101 is typically a terminal used by a creator who creates animation data using this animation creation system 1. The editing terminal 101 is configured, for example, by a personal computer or the like.
[0020] The editing terminal 101 includes a main memory device 201 and an auxiliary memory device 202. The editing terminal 101 also includes a processor 203, an input device 204, an output device 205, and a communication control unit 206 as described above.
[0021] The main memory device 201 stores programs and applications such as an animation creation module 211, a model setting module 212, a motion information acquisition module 213, a text information acquisition module 214, a character image information generation module 215, a background image information generation module 216, a synthesis module 217, and an image modification module 218. Each functional element of the editing terminal 101 is realized by the processor 203 executing these programs and applications stored in the main memory device 201.
[0022] The auxiliary storage device 202 stores information necessary for the operation of the animation creation system 1. The auxiliary storage device 202 stores, for example, an image generation model 221, operation information 222, text information 223, image information 224, and video information 225. Details of this information will be described later.
[0023] A brief description of each functional element of the editing terminal 101 follows. The animation creation module 211 comprehensively controls the basic operations of the editing terminal 101. The animation creation module 211 executes animation creation processing by coordinating and operating modules such as a model setting module 212, a motion information acquisition module 213, a text information acquisition module 214, a character image information generation module 215, a background image information generation module 216, a synthesis module 217, and an image modification module 218. The animation creation module 211 also cooperates with the server 102, for example, to store the created animation data in the server 102.
[0024] The model setting module 212 sets the features of the image generated by the image generation model. Various methods for setting the features of the image generated by the image generation model are considered. Specifically, for example, the model setting module 212 loads a data file related to the feature quantities of a plurality of images having a certain feature, which are generated in the process of learning the images, into the image generation model, or sets parameter tuning information in the image generation model. In addition, the model setting module 212 introduces, for example, one or more additional learning layers (e.g., adaptation layers) into the image generation model. This makes it possible, for example, to impart predetermined features to the images generated by the image generation model and ensure consistency.
[0025] The motion information acquisition module 213 acquires motion information relating to the motion of an object. This motion information is reference information for the character's movement when the character is made to move in an animation. The motion information is not limited to this, but may be acquired based on video information or motion capture data, for example.
[0026] The text information acquisition module 214 acquires text information representing instructions or commands to be input to the image generation model. This text information is an instruction for obtaining a desired image (output), and may be, for example, information about the characteristics of the image that you want the image generation model to generate.
[0027] The character image information generation module 215 generates character images using an image generation model. The character image information generation module 215 uses one or more image generation models to generate character images suitable for creating the target animation. The character image information generation module 215 generates image information for, for example, images of main characters and images of mob (other) characters.
[0028] The background image information generation module 216 generates background images using an image generation model. The background image information generation module 216 generates background images suitable for creating the target animation using one or more image generation models. Note that, for example, the functions of both the character image information generation module 215 and the background image information generation module 216 may be configured to be performed by a single image generation module (not shown).
[0029] The compositing module 217 generates image information suitable for animation data based on the generated images. The compositing module 217 generates, for example, frame image information representing an image in which a background and a character are combined. The compositing module 217 also creates video information based on multiple frame image information.
[0030] The image modification module 218 modifies the generated image. For example, the image modification module 218 creates an image in which an additional expression is added to a part of a character. The image modification module 218 also corrects the color tone of the character image based on the color tone of the reference image.
[0031] Each piece of information stored in the auxiliary storage device 202 will be briefly described. The image generation model 221 is a generative artificial intelligence (generative AI) trained to be able to output (generate) an image in response to a prompt as input. The image generation model 221 generates a plurality of frame images (still images) that constitute an animation. The image generation model 221 of this embodiment is a generative AI configured to respond to at least text information as input and output an image related to the text information. The image generation model 221 may be, for example, a unimodal generative AI whose input is limited to text information, or a multimodal generative AI whose input is not limited to text information. The image generation model 221 may be configured to output an image using, for example, text information and image information (including still image information and video information) as input.
[0032] The image generation model 221 may be a deep generative model that combines deep learning and a generative model. Examples of deep generative models for generating images include a variational autoencoder (VAE), a generative adversarial network (GAN), a flow-based model, a diffusion model, and a latent diffusion model.
[0033] The movement information 222 is information used as input to the image generation model 221. The movement information 222 is information representing the movement of an animated character in an animation or a series of postures taken by a character in a frame image. The movement information 222 may be, for example, motion capture information in which the movement of a real person or object is digitally recorded using markers and trackers. The movement information 222 may also be video information in which the movement of a real person or object is recorded as a video.
[0034] The text information 223 is information used as input to the image generation model 221. The text information 223 is information that linguistically expresses the characteristics of characters, objects, backgrounds, etc. in animations or frame images. The text information 223 may be information that linguistically expresses the characteristics of characters, objects, backgrounds, etc. in a positive manner, or may be information that linguistically expresses the characteristics of characters, objects, backgrounds, etc. in a negative manner.
[0035] The image information 224 is digitally recorded information of still images (e.g., part or all of a frame image) of animation characters, objects, backgrounds, etc. Based on this image information 224, a computer (e.g., processor 203) can display, for example, part or all of a frame image on an output device 205 such as a display.
[0036] The video information 225 is digitally recorded information of all or part of an animation. In this specification, the video information 225 may be referred to as animation data. Based on this video information 225, a computer (e.g., the processor 203) can display all or part of the animation on an output device 205 such as a display.
[0037] 3 illustrates an example of the hardware configuration of the server 102. The server 102 is an additional element of the animation creation system 1, and is configured, for example, by a computer server or the like located on a cloud. The server 102 is typically a terminal used by an administrator who manages and operates the animation creation system 1, and can work with the editing terminal 101 to create video information 225 as needed.
[0038] The main memory device 301 of the server 102 stores programs and applications such as an animation creation module 311, a model setting module 312, a motion information acquisition module 313, a text information acquisition module 314, a character image information generation module 315, a background image information generation module 316, a synthesis module 317, and an image modification module 318. Each functional element of the server 102 is realized by the processor 303 executing these programs and applications stored in the main memory device 301. The functions of each module of the server 102 are similar to those of the corresponding modules of the editing terminal 101, and therefore will not be described again.
[0039] The animation creation module 311 of the server 102 has the same functions as the animation creation module 211 of the editing terminal 101, and also outputs the created animation to the user terminal 103. The animation creation module 311 cooperates with, for example, a viewing module 412 of the user terminal 103, and displays animation on a display (an example of the output device 405; the same applies below) of the user terminal 103 based on the created video information 225. The animation creation process of the animation creation system 1 may be executed by the editing terminal 101 or by the server 102.
[0040] [User Terminal] Fig. 4 illustrates an example of the hardware configuration of the user terminal 103. The user terminal 103 is an element that outputs the created animation. The user terminal 103 is typically a terminal used by a user who uses this animation creation system 1 to view animation. The user terminal 103 is configured, for example, by a personal computer or a smartphone. Furthermore, the user terminal 103 may be, for example, a terminal with a head-mounted display, a head-mounted AR terminal, a terminal with AR glasses, or the like.
[0041] The user terminal 103 includes a main memory device 401 and an auxiliary memory device 402. The user terminal 103 also includes the processor 403, input device 404, output device 405, camera 406, and communication control unit 407, as described above. The camera 406 may be an imaging device built into a smartphone or the like, or may be an imaging device built into a wearable device such as AR glasses or a head-mounted display. The output device 405 may be a display device (display) included in the smartphone or the like, or may be a glasses-type or goggle-type display device in AR glasses or a head-mounted display.
[0042] The main memory device 401 stores programs and applications such as a user management module 411 and a viewing module 412. Each functional element of the user terminal 103 is realized by the processor 403 executing these programs and applications stored in the main memory device 401.
[0043] The auxiliary storage device 402 stores information necessary for the operation of the animation creation system 1. For example, the auxiliary storage device 402 stores video information 225. The video information 225 may be a part or all of the video information 225 stored in the auxiliary storage device 402 of the server 102.
[0044] The user management module 411 comprehensively manages the basic operations of the user terminal 103. The user management module 411, for example, executes the processing required for connecting to the server 102. For example, the user management module 411 cooperates with the animation creation module 311 of the server 102 to output (display) the homepage, login page, title page, etc. of the anime viewing site provided on the web by the server 102 to an output device 405 such as a display of the user terminal 103.
[0045] The viewing module 412 executes the processing required for the user to view the animation on the user terminal 103. The viewing module 412, for example, works in conjunction with the animation creation module 311 of the server 102 to display the animation on a display.
[0046] [Animation Creation Method] Next, a method for creating an animation using the animation creation system 1 will be described. The creator executes, for example, an animation creation program stored in the auxiliary storage device 202 of the editing terminal 101. This starts an animation creation application, and the animation creation module 211 is realized. The animation creation module 211 displays a screen of the animation creation application on the output device 205, such as a display. When the animation creation module 211 receives an instruction from the creator via the screen of the animation creation application to start creating an animation, it starts creating the animation in cooperation with, for example, each of the modules 212 to 217.
[0047] FIG. 5 shows an example of an animation creation flow. In the animation creation method of this embodiment, the animation creation module 211 cooperates with other modules to execute the following processes: The model setting module 212 sets an image generation model (S510); The motion information acquisition module 213 acquires motion information related to the object's motion (S520); The text information acquisition module 214 acquires text information related to the character's characteristics (S530); the character image information generation module 215 inputs the motion information and text information into the image generation model and generates multiple image information representing character images (S540); The background image information generation module 216 generates background image information representing images related to the background (S550); The compositing module 217 generates multiple frame image information representing an image obtained by compositing the background and the character (S560), and creates video information based on the multiple frame image information (S570).
[0048] The animation creation module 211 executes the animation creation process in cooperation with the model setting module 212, the motion information acquisition module 213, the text information acquisition module 214, the character image information generation module 215, the background image information generation module 216, and the synthesis module 217, but the animation creation module 211 may be configured to execute all or part of the processing of each of these modules itself.
[0049] Each step of the animation creation method will be described below with reference to the drawings as appropriate. First, the image generation model setting step (S510) will be described. Fig. 6 shows an example of the image generation model setting flow. Fig. 10 shows an example of the setting screen for the image generation model.
[0050] An image generation model generates one or more images in response to an input prompt (sometimes referred to as input information, input text information, or text information). While this is not a problem when generating a still image of a single scene, there are challenges when generating consistent images for a single work, such as an anime series spanning several minutes to several hours. Specifically, it is difficult to fix the animation style or the general characteristics of each character throughout a single work, or to fix the color and decoration of characters across multiple frames of time.
[0051] Therefore, the model setting module 212 sets conditions related to image generation of the image generation model. The model setting module 212 sets conditions related to characteristics such as the art style, the outer shape of a character or object, and parts of a character or object. Specifically, the model setting module 212 sets conditions related to image generation of the image generation model, for example, by executing the following model setting flow. The model setting module 212 displays an image generation model setting screen 1000 shown in FIG. 10 on the display of the editing terminal 101, accepts instructions from the creator via this image generation model setting screen, and executes the following model setting flow in accordance with the accepted instructions.
[0052] In this specification, "artistic style" typically refers to the atmospheric characteristics that appear in common in images generated by an image generation model, such as the drawing technique, touch, line tone, color, color scheme, texture, and other characteristics and tendencies that appear in the picture, as well as combinations of these. Furthermore, "appearance" typically refers to the character's face, limbs, and body contours such as hair, as well as the skeleton, fleshiness, body balance, physique, and contours of accessories such as clothing and necklaces, as well as combinations of these.
[0053] In the model setting process, the model setting module 212 executes the following steps: The model setting module 212 acquires an image generation model (S610), sets overall information, which is information about the character's appearance and / or style, in the image generation model (S620), and sets partial information related to some of the character's features in the image generation model (S630).
[0054] Specifically, the model setting module 212 first acquires an image generation model (S610). The image generation model acquired by the model setting module 212 is, for example, a large language model (LLM) that has been trained in advance to generate one or more images in response to a prompt. The LLM is composed of an artificial neural network with a large number of parameters (typically tens of millions or more, e.g., hundreds of millions or more, or even tens of billions or more). The model setting module 212 can acquire the image generation model by, for example, copying (reading) an image generation model stored in the server 102.
[0055] The model setting module 212 then sets overall information, which is information about at least one of the character's appearance and art style, to the image generation model (S620). The overall information is, for example, information about the feature quantities of an image generated when the image generation model is further trained. The overall information can be, for example, information about the feature quantities of an image generated when the image generation model is further trained preliminarily using images of characters with a predetermined appearance (e.g., physique) and / or art style. The overall information is stored, for example, as an overall information file in the auxiliary storage device 202, 302 (which may be one or both of the auxiliary storage device 202 and the auxiliary storage device 302; the same applies below).
[0056] The model setting module 212 displays, for example, an overall information file selection field 1010 on the image generation model setting screen 1000. When the overall information file selection field 1010 is selected, the model setting module 212 displays a file selection dialog and accepts the creator's selection of the overall information file to be set via the file selection dialog. The model setting module 212 sets parameter values for the image generation model, for example, based on the overall information recorded in the overall information file. As a result, the model setting module 212 obtains, for example, an image generation model (typically, an LLM) as acquired, with fixed parameters for generating an image with a predetermined shape and / or artistic style. By using, for example, such an image generation model, the model setting module 212 can impart consistency to the shape and / or artistic style of the character in the generated image. In this embodiment, the model setting module 212 sets overall information for the image generation model, including information representing a two-dimensional anime image, as artistic style information.
[0057] The model setting module 212 can set one overall information file or two or more overall information files for an image generation model. For example, the model setting module 212 displays an Add button 1020 on the image generation model setting screen 1000, and when the Add button 1020 is selected, displays an additional overall information file selection field 1010 (not shown) below the first overall information file selection field 1010. The model setting module 212 then accepts the selection of an overall information file to be additionally set via the additional overall information file selection field 1010. Although not specifically shown, when two or more pieces of overall information are set, the model setting module 212 may be configured to accept a designation of the degree of reflection of each piece of overall information.
[0058] The model setting module 212 also sets partial information related to some of the character's characteristics in the image generation model (S630). The partial information is, for example, information related to the feature quantities of an image generated when the image generation model, which has learned the above-mentioned external shape and / or art style, is further trained on some of the character's characteristics. The partial information may be, for example, information related to the feature quantities of an image generated when the image generation model is further trained preliminarily using images of a character that represent at least one of a specific part's individuality (e.g., facial features, hairstyle, contours, body shape, etc.), appearance (e.g., clothing, etc.), and action (e.g., speaking, singing, playing a specific sport, playing a specific instrument, etc.). The partial information may also be, for example, information related to feature quantities that characterize the above-mentioned external shape and / or art style at a more detailed level (e.g., the tone of the outline or perimeter of the character, etc.). This partial information is stored, for example, as a partial information file in the auxiliary storage device 202, 302.
[0059] The model setting module 212, for example, displays a partial information file selection field 1030 on the image generation model setting screen 1000. When the partial information file selection field 1030 is selected, the model setting module 212 displays a file selection dialog and accepts the creator's selection of the partial information file to be set via the file selection dialog. The model setting module 212, for example, sets a very small number of parameter values for the image generation model based on the partial information recorded in the partial information file. This significantly reduces the number of parameters for the image generation model, enabling control of feature amounts while improving image quality. As a result, it is possible to generate images with desired features without significantly increasing the capacity of the image generation model itself. In particular, it is possible to fix some of the features of a character while maintaining the outer shape and / or style set by the overall information file.
[0060] The model setting module 212 can set one piece of partial information or two or more pieces of partial information for an image generation model. For example, the model setting module 212 displays an Add button 1040 on the image generation model setting screen 1000, and when the Add button 1040 is selected, displays an additional partial information file selection field 1030 (not shown) below the first partial information file selection field 1030. The model setting module 212 then accepts the selection of a partial information file to be additionally set via the additional partial information file selection field 1030.
[0061] Here, partial information can be prepared for each different part of the character. Furthermore, partial information can be prepared for each of multiple different features for a single part of the character. Therefore, the model setting module 212 can set multiple pieces of partial information by changing the combination based on a combination of multiple pieces of partial information. As a result, for example, it is possible to generate an image in which some of the features of the character are changed. Although not specifically shown, when setting two or more pieces of partial information, the model setting module 212 may be configured to accept a specification of the degree of reflection of each piece of partial information. In this embodiment, the model setting module 212 sets partial information by introducing, for example, a trained model and setting information regarding facial features that identify the character and a trained model and setting information regarding clothing into the image generation model.
[0062] The model setting module 212 displays, for example, a preview window 1070 and a preview button 1080 on the image generation model setting screen 1000. When the preview button 1080 is selected, the model setting module 212 displays an example of an image generated by the image generation model for which the overall information and partial information selected in the overall information file selection field 1010 and partial information file selection field 1030 have been set. This allows the creator to check whether the selected overall information file and partial information file are appropriate while viewing the preview window 1070.
[0063] Next, the motion information acquisition step (S520) will be described. Although not limited to this, in this embodiment, motion information based on the voluntary movement of a real object is used to reflect the continuous movements of a character in an animation in a series of frame images. A case in which the motion information acquisition module 213 acquires this motion information from video information will be described below. FIG. 7 shows an example of a motion information acquisition flow. FIG. 11 shows an example of a motion information acquisition screen 1100. The motion information acquisition module 213 displays the motion information acquisition screen 1100 on the display of the editing terminal 101, receives instructions from the creator via this motion information acquisition screen 1100, and executes the following motion information acquisition flow based on the received instructions.
[0064] In the motion information acquisition step, the motion information acquisition module 213 executes the following steps: The motion information acquisition module 213 captures or acquires a video of an object (S710), inputs the video into a posture estimation model that estimates posture information of an object included in the image from image information (S720), and acquires motion information related to the motion of the object (S730).
[0065] Specifically, the motion information acquisition module 213 acquires, for example, a video of an object (S710). The motion information acquisition module 213 may acquire the video by shooting the video, or may acquire a video that has been shot or created in advance. The object may be, for example, a real person or object. However, the object may also be, for example, an illustrated person or object. When acquiring the video by shooting, the motion information acquisition module 213 captures the movements of the object corresponding to the movements of the character in the animation via an imaging device 104, such as a video camera connected to the editing terminal 101. The motion information acquisition module 213 captures, for example, a video of the behavior, dance, and other movements of a real person (an example of an object) based on the animation script or storyboard. A video file of the captured video is stored, for example, in the auxiliary storage device 202.
[0066] Next, the motion information acquisition module 213 inputs the captured video into the posture estimation model (S720). The posture estimation model is a learning model that estimates the posture information of an object contained in the image from the image information. The posture estimation model is trained by inputting, for example, a still image or a video of the object and outputting key points (feature points, such as the positions of human joints or facial features) and information about combinations of key points whose movements are estimated to be constrained (in other words, connected). By inputting a still image into this posture estimation model, key points and information about their connections are output. Furthermore, by inputting a video into this posture estimation model, key points and information about their connections are output for each of a series of still images that make up the video. The posture estimation model is stored, for example, in the auxiliary storage device 202 or 302.
[0067] Specifically, for example, motion information acquisition module 213 displays video file selection field 1110 on motion information acquisition screen 1100. When video file selection field 1110 is selected, motion information acquisition module 213 displays a file selection dialog and accepts the creator's selection of a video file to be input to the posture estimation model via the file selection dialog. As a result, motion information acquisition module 213 inputs the selected video file into the posture estimation model.
[0068] As a result, the motion information acquisition module 213 acquires information about key points and their connections for a series of still images constituting the video as motion information related to the object's motion (S730). The motion information may be expressed, for example, in a two-dimensional (x, y) coordinate system as 2D pose information, or in a three-dimensional (x, y, z) coordinate system as 3D pose information, or may be information expressed in other ways. The motion information acquisition module 213 stores the acquired motion information, for example, in the auxiliary storage device 202. The motion information acquisition module 213 may set the acquired motion information, for example, in an image generation model.
[0069] Furthermore, the motion information acquisition module 213 displays, for example, a plurality of frame images constituting the captured video in the frame image list display field 1120 based on the selected video file. Then, the motion information acquisition module 213 displays, for example, one of the plurality of frame images displayed in the frame image list display field 1120 in the original frame image display field 1130 of the motion information acquisition screen 1100. For example, the motion information acquisition module 213 displays, for example, in the form of 2D pose information in the capture field 1140, motion information obtained by inputting the frame image displayed in the original frame image display field 1130 into a posture estimation model. The motion information acquisition module 213 displays, for example, the motion information in the form of a so-called stick figure. A stick figure is a diagram that represents the posture of an object by using key points (e.g., feature points) as dots and connecting potentially connectable feature points with sticks.
[0070] The motion information acquisition module 213 may acquire motion information by estimating the posture of the object for all frames included in the video file, or may acquire motion information by estimating the posture of the object for some frame images included in the video file (for example, frame images with the same frame rate as the animation to be created, or frame images with a rate of 1 / 2 to 1 / 3 of the frame rate of the animation to be created, etc.). Furthermore, the display of frame images in the original frame image display field 1130 and the display of corresponding posture information (motion information) in the capture field 1140 may be performed, for example, each time posture information (motion information) for each frame image is acquired. The display of corresponding posture information (motion information) in the capture field 1140 may be performed for only one frame image selected by the creator from among the multiple frame images displayed in the frame image list display field 1120. Alternatively, these may be combined.
[0071] In relation to acquiring motion information, the motion information may include boundary information regarding the boundary of an object, in addition to or independently of the information regarding the key points and their connections. In other words, the motion information acquisition module 213 may be configured to acquire motion information from the perimeter of the object. Furthermore, the object may include a subject performing the action and an attachment worn by the subject. The motion information acquisition module 213 may be configured, for example, to capture or acquire a video of the subject wearing the attachment and acquire boundary information regarding the subject's motion and the boundary (which may be a perimeter, outline, or characteristic line) of the attachment from image information included in the video. Examples of subjects of the action include living organisms such as humans and moving machines such as robots. Examples of attachments include clothing, ornaments such as accessories, portable items such as bags, and moving objects such as transported objects. Note that even if a part of the subject belongs to the subject, an element that the subject cannot freely move, such as hair, may also be considered an attachment. The boundary information can be obtained using, for example, edge detection technology or region division technology using mean shift in computer graphics technology, or a machine learning model trained to estimate the contour line of a subject from an image.
[0072] Furthermore, when acquiring motion information from a video file of an object, the video information may not include background information (e.g., a green screen, etc.), or may include background information (e.g., a two-dimensional background, a three-dimensional background, or a real background, etc.). When background information is included, the motion information acquisition module 213 may be configured, for example, to acquire part or all of the background information as needed. For example, the motion information acquisition module 213 may be configured to accept a user's selection (or instruction) of an area from which background information should be acquired. When the motion information acquisition module 213 accepts a user's selection (or instruction) of an area from which background information should be acquired, it may be configured to also acquire background information of that area from the video information.
[0073] The motion information acquisition module 213 may acquire motion information using motion capture technology. Motion capture technology is a technology for digitally recording the movements of real people or objects. For example, motion is recorded by tracking markers attached to key points on the object over time (optical). Other types of motion capture technology may include inertial sensor technology using an inertial sensor, mechanical technology using a sensor that mechanically measures rotation angle and displacement, magnetic technology using a combination of a magnetic generator and a magnetic sensor, and video technology that analyzes video. For easy creation of the intended animation, it is preferable that the motion information acquired by motion capture technology corresponds to the camera position, camera angle, and camerawork corresponding to the scene of the animation to be created. Such video may be actually captured, or may be prepared by converting the motion capture information into position information corresponding to the video captured virtually with the desired camerawork. Furthermore, instead of capturing video of the object (S710), the motion information acquisition module 213 may acquire pre-acquired motion information.
[0074] Here, the character image information generation module 215 displays, for example, a preview window 1150 and a preview button 1155 on the motion information acquisition screen 1100. Then, the character image information generation module 215 applies the frame image displayed in the original frame image display field 1130 and the posture information (an example of motion information) and boundary information (an example of motion information) of the stick figure displayed in the capture field 1140 to the image generation model set at that time, and displays the generated image in the preview window 1150. As shown in the preview window 1150, the character image information generation module 215 can generate, for the image generation model, an image of a character having a posture, hair outline, finger outline, clothing outline, wrinkles, shading, etc. that matches the posture of the object, hair outline (not shown), finger outline, clothing outline, wrinkles, shading, etc. in the frame image displayed in the original frame image display field 1130.
[0075] The size, aspect ratio, resolution, etc. of the image that the character image information generation module 215 generates for the image generation model are not limited to those corresponding to the size, shape (for example, aspect ratio 16:9), resolution, etc. of the target animation cut or frame. The character image information generation module 215 can, for example, cause the image generation model to generate an image with any size, shape, or resolution. Furthermore, the character image information generation module 215 can also cause the image generation model to generate an image with any resolution, for example, a standard resolution, an arbitrary resolution other than the standard resolution, or both.
[0076] Furthermore, the motion information acquired by the motion information acquisition module 213 is not necessarily optimal for the motion of an animated character. Therefore, the motion information acquisition module 213 may be configured to, for example, display an edit button 1145 in the capture field 1140, and when the edit button 1145 is selected, enable editing of the stick figure (posture information) displayed in the capture field 1140 and accept editing of the stick figure by the creator. In this case, when the preview button 1155 is selected again, the character image information generation module 215 can generate an image by reflecting the corrected posture information of the stick figure in the image generation model and display the generated image in the preview window 1150. Note that, although the motion when the preview button 1155 is selected is described as being executed by the character image information generation module 215, it may also be configured to be executed by the motion information acquisition module 213.
[0077] The text information acquisition step (S530) will now be described. The text information acquisition module 214 acquires text information, which is text-format information used to condition the generation of an image by the image generation model. This text information may be instruction information for making the image information generated by the image generation model suitable for an anime frame image. The text information acquisition module 214 acquires, for example, text information related to the characteristics of a character. The text information acquisition module 214 can acquire, as text information, positive instruction information that instructs on elements that should be included in the image to be generated. The text information acquisition module 214 can also acquire, as text information, negative instruction information that instructs on elements that should not be included in the image to be generated.
[0078] 12 is an example of an image information generation screen 1200. The text information acquisition module 214 displays, for example, a positive instruction information field 1240 for accepting input of positive instruction information from the creator and a negative instruction information field 1250 for accepting input of negative instruction information from the creator on the image information generation screen 1200. The text information acquisition module 214, for example, acquires text information entered by the creator in the positive instruction information field 1240 as positive instruction information. The text information acquisition module 214 also acquires text information entered by the creator in the negative instruction information field 1250 as negative instruction information. The text information acquisition module 214 can accept text information at any time, for example, via the positive instruction information field 1240 and the negative instruction information field 1250.
[0079] The text information acquisition module 214 may be configured to accept, as text information, a combination of multiple pieces of instruction information prepared in advance, as exemplified by "NegativePromptSet_V5" in the negative instruction information field 1250. For example, by preparing multiple combinations of instruction information in advance for each animation or each character, it is possible to prevent, for example, missing instruction information or forgetting to delete unnecessary instruction information, and it is advantageous for stably generating character images.
[0080] The character image information generation step (S540) will now be described. FIG. 8 shows an example of an image information generation flow. The character image information generation module 215 inputs motion information and text information into the image generation model, and generates multiple pieces of image information representing images of the character. The character image information generation module 215 displays, for example, an image information generation screen 1200.
[0081] In the character image information generation process, the character image information generation module 215 executes the following steps: The character image information generation module 215 acquires an image generation model (S810), sets positive instruction information from the text information in the image generation model (S820), sets negative instruction information from the text information in the image generation model (S830), and generates, based on the action information and the positive and negative instruction information, a plurality of pieces of image information representing images of the character that reflect the positive instruction information but are unlikely to reflect the negative instruction information, using the image generation model (S840).
[0082] The character image information generation module 215, for example, acquires the image generation model set in step S510, and sets the positive instruction information from the text information acquired in step S530 as input information for the image generation model, and also sets the negative instruction information as input information for the image generation model. Here, for example, the character image information generation module 215 displays an image information generation screen 1200 shown in Fig. 12 on the display of the editing terminal 101, and displays an overall model information column 1210, a partial model information column 1220, and an action information column 1230 in addition to a positive instruction information column 1240 and a negative instruction information column 1250.
[0083] The character image information generation module 215 displays the overall model information and partial model information set in step S510 in the overall model information column 1210 and the partial model information column 1220, respectively, in the form of, for example, file names. The character image information generation module 215 also displays the movement information acquired in step S520 in the movement information column 1230, in the form of, for example, file names. This allows the creator to confirm what information is being used to generate the character image. Furthermore, as necessary, the module can accept changes to the information to be set in the image generation model from the creator via the overall model information column 1210, the partial model information column 1220, the movement information column 1230, etc., and change the information to be set in the image generation model based on such changes.
[0084] For example, the character image information generation module 215 displays a preview window 1260, a preview button 1265, an image generation button 1270, a save button 1275, a cancel button 1280, and the like on the image information generation screen 1200. For example, when the character image information generation module 215 receives selection of the image generation button 1270 by the creator, the character image information generation module 215 generates a plurality of pieces of image information representing images of the character using an image generation model based on the action information and the positive and negative instruction information. The character image information generation module 215 generates a series of frame images representing a character that reflects the positive and negative instruction information and that acts in response to the action information.
[0085] This allows the character image information generation module 215 to generate, for each frame, image information representing a character image that reflects positive instruction information but is less likely to reflect negative instruction information. For example, in response to the instruction "...long-sleeved shirt with sleeves half rolled up, pendant in yellow gold, dark green flared skirt with brown belt, knee-length flared skirt..." written in the positive instruction information column 1240, a character is generated in the preview window 1260 that reflects the positive instruction information of a long-sleeved shirt with sleeves half rolled up, a yellow gold pendant, a dark green flared skirt with a brown belt, and a knee-length flared skirt.
[0086] Furthermore, due to the instruction "multiple people, ...short skirt" written in the negative instruction information field 1250, multiple characters are not generated in the preview window 1260, and no characters wearing short skirts are generated. In other words, characters that are unlikely to reflect negative instruction information are generated. Here, the image generation model generates, as multiple image information, multiple image information representing images of characters having postures corresponding to the posture of the object at a certain point in time in the action information. These multiple image information correspond, for example, to still images of the animation of the character corresponding to the action information. However, the size, aspect ratio, resolution, etc. of the images generated by the image generation model in the character image information generation module 215 are not limited to the size, shape (e.g., aspect ratio 16:9), resolution, etc. of the target animation cut or frame. The character image information generation module 215 can generate images in, for example, a size, shape, and resolution. Furthermore, the character image information generation module 215 can set the resolution of the images generated by the image generation model to, for example, standard resolution, any resolution other than standard resolution, or both.
[0087] For example, when the character image information generation module 215 receives a creator's selection of the preview button 1265, it displays an image generated by the image generation model in the preview window 1260. For example, when the character image information generation module 215 receives a creator's selection of the save button 1275, it saves the image information of the generated character in the auxiliary storage device 202 or the like. For example, when the character image information generation module 215 receives a creator's selection of the cancel button 1280, it deletes the image information of the character generated by the image generation model.
[0088] Furthermore, in the image information generation process, if image information representing the multiple character images generated in step S840 has not been obtained that satisfies the conditions for each frame (No in S850), text information that adjusts the detailed facial expressions and movements for each frame is set for the image generation model (S860), and image information representing the character image that reflects the text information that adjusts the detailed facial expressions and movements for each frame is generated again based on the movement information and text information (S840).On the other hand, if image information representing the multiple character images generated in step S840 has been obtained that satisfies the conditions for each frame (Yes in S850), the generation of image information representing the character image is terminated.
[0089] If the creator determines that image information for the character for each frame that satisfies the conditions has not been obtained, the creator can change the instruction information input in the positive instruction information field 1240 or the negative instruction information field 1250 of the image information generation screen 1200 to instruction information that adjusts the detailed facial expression and movement of the character for each frame. For example, when the text information acquisition module 214 accepts input of instruction information that adjusts the detailed facial expression and movement of the character for each frame into the positive instruction information field 1240 and the negative instruction information field 1250, and the character image information generation module 215 accepts the creator's selection of the image generation button 1270, the character image information generation module 215 sets the movement information and the instruction information that adjusts the detailed facial expression and movement for each frame (updated positive and negative instruction information) in the image generation model.
[0090] The character image information generation module 215 then generates image information for the character for each frame using the image generation model for which the new instruction information has been set. The character image information generation module 215 repeatedly executes steps S840 to S860 until it determines that image information for each frame that satisfies the conditions has been obtained. The character image information generation module 215 stores the image information for each frame of the generated character in, for example, the auxiliary storage device 202.
[0091] Note that the setting of text information for adjusting the detailed facial expressions and actions for each frame in step S860 can also be performed by other methods. Figures 13 to 15 show examples of other image information generation screens 1300, 1400, and 1500. The character image information generation module 215 displays, for example, frame image display windows 1310, 1410, and 1510 with seek bars and instruction information editing windows 1320, 1420, and 1520 on the image information generation screens 1300, 1400, and 1500. The character image information generation module 215 displays, in the frame image display windows 1310, 1410, and 1510, an image represented by image information for one frame of the generated plurality of image information, along with a seek bar indicating the time point at which the image appears among the actions represented by the plurality of image information. Character image information generation module 215 is configured to be able to, for example, sequentially display a plurality of images corresponding to a predetermined scene or cut based on an instruction from the user in frame image display windows 1310, 1410, and 1510. Character image information generation module 215 may be configured to display image information generation screens 1300, 1400, and 1500 when, for example, an operation to scroll image information generation screen 1200 downward is received.
[0092] For example, when the character image information generation module 215 receives a creator's movement of a slider on a seek bar, it displays an image corresponding to the time indicated by the slider in the frame image display window 1310, 1410, or 1510, and moves the instruction information editing window 1320, 1420, or 1520 in accordance with the slider. The character image information generation module 215 also displays text information (instruction information) set for generating image information for the frame at the time indicated by the slider in the instruction information editing window 1320, 1420, or 1520. If a large amount of text information has been set, the character image information generation module 215 displays a scroll bar in the instruction information editing window 1320, 1420, or 1520, allowing the creator to view all of the text information.
[0093] Here, the character image information generation module 215 accepts edits (e.g., changes, additions, deletions, etc.) of the text information in the instruction information editing windows 1320, 1420, and 1520 by the creator, and sets the edited instruction information in the image generation model. For example, at the time shown in FIG. 13 , the text information for adjusting the character's facial expression and movement displayed in the instruction information editing window 1320 includes "open eyes" and "close mouth." Meanwhile, at the time shown in FIG. 14 , the character image information generation module 215 accepts, for example, an edit to change the text information in the instruction information editing window 1320 to "close eyes" and "close mouth," and the selection of the image generation button 1330. Also, at the time shown in FIG. 15 , the character image information generation module 215 accepts, for example, an edit to change the text information in the instruction information editing window 1320 to "open eyes" and "open mouth," and the selection of the image generation button 1330.
[0094] As a result, character image information generation module 215 regenerates image information for each time point based on the action information at that time point using the image generation model in which the edited instruction information has been set, and displays the image represented by the generated image information in frame image display windows 1410, 1510. When the creator determines that image information for each frame that satisfies the conditions has been obtained, he or she selects save button 1440, 1540, and upon receiving the selection of save button 1440, 1540, character image information generation module 215 stores the edited instruction information and the generated image information in, for example, auxiliary storage device 202.
[0095] Furthermore, if the character image information generation module 215 determines that image information for each frame that satisfies the conditions has not been obtained, the user selects the cancel button 1450, 1550 as necessary and resets the instruction information until image information for each frame that satisfies the conditions is obtained. This configuration allows the creator to easily set text information that adjusts the detailed facial expressions and movements of the character for each frame. While the examples in FIGS. 13 to 15 show an example in which a slider bar is used to specify a time and input instruction information for image generation for each frame, instruction information may be set individually for each frame without using a slider bar. While the examples in FIGS. 13 to 15 show a background behind the character, the background may or may not be displayed at this point.
[0096] Regarding the generation of a character's facial expression, the character image information generation module 215 may set the image generation model to generate a specific facial expression, such as a smiling face, crying face, or angry face (a guide may be set), or may issue a prompt (input) to generate a specific facial expression. As another example, a character image with a facial expression corresponding to the hand-drawn input can be generated by accepting hand-drawn input, such as tears, a smiling mouth, or the open / closed state of the eyes, for a face image displayed as a line drawing, and setting the input facial expression information as a guide in the image generation model. Acquiring this hand-drawn input information allows for more accurate and detailed facial expression specification.
[0097] The character image information generation module 215 may be configured to generate another image based on an image (Image-to-Image). In this case, the character image information generation module 215 may be configured to accept a selection of a portion of the original image and convert only the selected portion into another image in accordance with arbitrary text information (instruction information). The character image information generation module 215 may also be configured to accept a selection of a portion of the original image and convert the portion other than the selected portion into another image in accordance with arbitrary text information (instruction information). This configuration makes it possible to more easily generate character images that meet the creator's wishes.
[0098] 12 to 14, the character image information generation module 215 generates character images for all frame images representing a predetermined scene that share common overall and partial characteristics, such as style and / or shape. For all frame images representing a predetermined scene, the character image information generation module 215 generates images that share, for example, the movement of hair and clothing that matches the character's actions, texture, coloring (including shading), contour lines, and peripheral lines, and generates images that have continuity among these.
[0099] The background image information generation step (S550) will be described. The background image information generation module 216 generates background image information representing an image related to the background. The background image information generation module 216 generates background image information representing an image of the background using, for example, a second image generation model. Fig. 9 shows an example of a background image information generation flow. Fig. 16 shows an example of a background image information generation screen 1600.
[0100] In the background image information generation process, the background image information generation module 216 executes the following steps: The background image information generation module 216 acquires three-dimensional information about the background (S910), cuts out an image with a field of view that matches the storyboard from the three-dimensional information (S920), inputs the cut-out image and text information into a second image generation model (S930), generates multiple pieces of background image information (S940), and extracts background image information that satisfies a condition from the multiple pieces of generated background image information (S950).
[0101] The background image information generation module 216 generates background image information based on, for example, three-dimensional information about the background. The three-dimensional information about the background is information for displaying a 3D model representing a background scene, and a scene when any region of this 3D model is viewed from any viewpoint can be extracted as a consistent background. The background image information generation module 216 extracts a background image from the 3D model in accordance with predetermined storyboard information. The storyboard information is information representing a storyboard consisting of a table illustrating each scene, which is prepared before the animation is created.
[0102] The background image information generation module 216 displays, for example, a three-dimensional information file selection field 1610 and a storyboard information file selection field 1620 on the background image information generation screen 1600, and as already explained, accepts designation of a three-dimensional information file and a storyboard information file related to the background to be acquired via these file selection fields 1610, 1620. Based on the accepted three-dimensional information file and storyboard information file, the background image information generation module 216 displays a 3D model related to the background in a 3D model display field 1615 and a storyboard for one cut in a storyboard cut display field 1625.
[0103] Next, when the background image information generation module 216 receives the creator's selection of the generate button 1650, it compares the displays in the 3D model display field 1615 and the storyboard cut display field 1625 while changing the angle, magnification, position, etc. of the display axis of the 3D model, and cuts out the image displayed in the 3D model display field 1615 at the angle, magnification, and position of the 3D model where the display in the 3D model display field 1615 best matches the display in the storyboard cut display field 1625.
[0104] The background image information generation module 216 inputs the extracted image and text information providing instructions for appropriately coloring the image into a second image generation model to generate multiple pieces of background image information. While not limited to this, the second image generation model may be the same image generation model as the image generation model (hereinafter sometimes referred to as the "first image generation model") used to generate the character image information, and may, for example, be configured with the same overall information as the first image generation model. This allows background image information to be generated that shares a common artistic style with the character. The background image information generation module 216 displays the generated multiple background images in the background candidate display field 1630 based on the generated multiple pieces of background image information.
[0105] Background image information generation module 216 accepts a creator's selection of an image suitable as a background image from among the multiple background images displayed in background candidate display field 1630, displays the selected background image in selected image display field 1640, and extracts image information representing the selected background image. For example, when background image information generation module 216 accepts the creator's selection of save button 1660, it stores the extracted image information in auxiliary storage device 202 or the like as background image information.
[0106] The background image information generation module 216 can read the conditions for determining background images and the AI model, and input multiple background images generated by the AI model. This allows the module to display only background images that meet the conditions, or highlight background images that are likely to meet the conditions and distinguish them from other background images. The size of the image generated by the second image generation model is not limited to the shape or size (e.g., 16:9) of the desired animation cut or frame, and can be larger. The resolution of the image generated by the second image generation model can be, for example, either a standard resolution or an arbitrary resolution other than the standard resolution, or both.
[0107] Next, another example of the background image information generation process (S550) will be described. FIG. 18 shows another example of a background image information generation flow. In the background image information generation process, the background image information generation module 216 executes the following steps. The background image information generation module 216 acquires point cloud data to acquire three-dimensional information (S1810), cuts out an image with a field of view that matches the storyboard from the three-dimensional information (S1820), acquires a two-dimensional image corresponding to the cut-out field of view (S1830), inputs the cut-out image, the acquired two-dimensional image, and text information into a second image generation model (S1840), generates multiple pieces of background image information (S1850), and extracts background image information that satisfies a condition from the generated multiple pieces of background image information (S1860).
[0108] The background image information generation module 216 uses a camera to acquire point cloud data of buildings and structures that are desired to be used as backgrounds for animation, such as inside a classroom, in front of an escalator, or inside a restaurant. The background image information generation module 216 generates three-dimensional information based on the acquired point cloud data (S1810). The three-dimensional information is information for displaying a 3D model generated from the captured point cloud data, and a scene when any region of this 3D model is viewed from any viewpoint can be extracted as a consistent background. The background image information generation module 216 extracts an image from the 3D model with a field of view that matches the storyboard information (S1820).
[0109] The background image information generation module 216 acquires a two-dimensional image corresponding to the cut-out image (S1830). For example, the background image information generation module 216 takes a photo of an actual building or structure with a field of view corresponding to the field of view of the image cut out from the 3D model, and acquires the photo.
[0110] The background image information generation module 216 inputs an image cut out from a 3D model, a 2D image such as a photograph acquired corresponding to the image, and text information that serves as instructions for generating a background image based on these images into a second image generation model (S1840), and generates multiple pieces of background image information (S1850). Note that the image cut out from the 3D model, the 2D image acquired corresponding to the image, and the text information that serves as instructions for generating a background image may all be input, or only one or two of these may be input. The background image information generation module 216 displays the generated multiple background images in the background candidate display field 1630 based on the generated multiple pieces of background image information.
[0111] In this way, the background image information generation module 216 generates a 3D model from point cloud data based on actual buildings and structures, and then scales, reduces, rotates, etc. this 3D model to create a field of view that matches the storyboard, from which a background image with the desired field of view can be extracted. Because this extracted background image was originally generated from point cloud data, it accurately represents the shape, but the image quality may be poor. Therefore, by taking a photograph of the actual building or structure that matches this field of view and inputting it into the second image generation model, a background image with the desired field of view based on the texture of the actual photograph can be obtained.
[0112] Background image information generation module 216 accepts a creator's selection of an image suitable as a background image from among the multiple background images displayed in background candidate display field 1630, displays the selected background image in selected image display field 1640, and extracts image information representing the selected background image. For example, when background image information generation module 216 accepts the creator's selection of save button 1660, it stores the extracted image information in auxiliary storage device 202 or the like as background image information.
[0113] The background image information generation module 216 can read the conditions for determining background images and the AI model, input multiple background images generated by the AI model, and extract and display only background images that meet the conditions, or extract and highlight background images that are likely to meet the conditions, distinguishing them from other background images (S1860). The size of the image generated by the second image generation model is not limited to the shape and size (e.g., 16:9) of the target animation cut or frame, and can be generated larger than this. The resolution of the image generated by the second image generation model can also be generated at either or both standard resolution and any other resolution.
[0114] The frame image information generation step (S560) and the video information generation step (S570) will now be described. The compositing module 217 generates a plurality of frame image information. The compositing module 217 generates a plurality of frame image information representing an image obtained by compositing a background and a character based on image information representing the character image created in step S540 and the background image information generated in step S550. The compositing module 217, for example, overlays a character image on a background image. For example, the compositing module 217 overlays the background image and the character image while relatively changing their angle, magnification (size), and position as necessary, and compares them with the display in the storyboard cut display field 1625.
[0115] The compositing module 217 then creates frame image information representing an image in which the background and character are combined at a predetermined aspect ratio (e.g., 16:9) at an angle, magnification, and position of the character image such that the display of the character superimposed on the background of the selected image display field 1640 matches the display in the storyboard cut display field 1625. The compositing module 217 then generates video information based on the multiple frame image information created in step S560. This allows an animation to be created.
[0116] The compositing module 217 may be configured to, for example, create frame image information combining a background image and a character image in a size larger than a predetermined frame image, and then extract from the large frame image a plurality of frame image information pieces of a predetermined size with positions and angles of view corresponding to the moving image of the camera viewpoint. This allows for efficient creation of moving images. Furthermore, when the camera viewpoint relative to the background image moves or the character moves in relation to a storyboard cut, the position and movement path of the character in the 3D model, as well as camerawork (the movement path of the virtual camera viewpoint), may be specified in advance during the storyboard creation stage, and the background image and character image from the virtual camera viewpoint may be geometrically composited.
[0117] 17 is an example of a storyboard information generation screen 1700. In the frame image information generation process, the compositing module 217 creates frame image information that best matches the display in the storyboard cut display field 1625, for example, based on a display in which a character is superimposed on the background of the selected image display field 1640. Here, for example, the compositing module 217 can create a storyboard using frame images by fitting a frame image 1720 corresponding to such frame image information into the corresponding cut in the storyboard table 1710. This makes it easy to grasp the progress of animation creation and the content of the creation.
[0118] The animation creation system 1 configured as described above allows for the use of an image generation model to create animations of new stories, reducing the amount of manpower and effort required by creators. Furthermore, by setting overall information and partial information for the image generation model, the feature quantities of the images generated by the image generation model can be fixedly controlled. For example, by setting art style information corresponding to information representing a two-dimensional animated image as overall information, image information representing a given two-dimensional animated image in a single work can be generated. Furthermore, by setting partial information for the image generation model that allows each character to be identified, individual characters can be drawn stably and distinctly while maintaining a common art style for all characters. When using hand-drawn illustrations, the creator's individuality (habits) may be apparent in the illustrations. However, the animation creation system 1 allows for the creation of animations while suppressing unintended image variations (habits).
[0119] <Modification> In the above embodiment, the animation creation method includes a step (S510) of setting an image generation model, and for example, the model setting module 212 sets overall information and partial information for the image generation model. However, if a set image generation model in which this information has already been set is available, it can be understood that the image generation model has been set by obtaining such a set image generation model.
[0120] In the above embodiment, when setting the overall information and partial information, the model setting module 212 applied a prepared overall information file and partial information file to the image generation model. However, the model setting module 212 may set the overall information by training the image generation model using images of characters with a predetermined shape and / or style. Alternatively, the model setting module 212 may set the partial information by introducing a trainable model or setting information into the image generation model and training the model to further learn the characteristics of a predetermined part of the character.
[0121] In the above embodiment, the model setting module 212 sets overall features and partial features for the image generation model in order to generate image information of a character having predetermined overall features and partial features. However, the model setting module 212 is not limited to generating image information of a character, and can set overall features and partial features for the image generation model when generating image information of other images such as backgrounds, for example. This allows the image generation model to have image information with a consistent style throughout the creation of a single animation.
[0122] In the above embodiment, the character image information generation module 215 and the background image information generation module 216 each generate character image information and background image information based on separate images, etc., and the compositing module 217 combines these to generate frame image information. However, the generation of character image information by the character image information generation module 215 and the generation of background image information by the background image information generation module 216 may be performed based on the same image information or video information. For example, as movement information, source video of an object moving against a three-dimensional background set for animation is acquired. Then, the character image information generation module 215 generates character image information using a first image generation model based on a portion of the source video corresponding to a target scene (cut). Furthermore, the background image information generation module 216 generates background image information using a second image generation model based on a portion of the source video corresponding to the same target scene (cut).
[0123] With this configuration, the composition module 217 can generate frame image information with a composition corresponding to the original source video by combining the generated character image information and background image information. As a result, even if the character image information and background image information are generated using different image generation models, the relative three-dimensional positions, angles, and overlaps can be accurately reproduced during composition, making it possible to create accurate video information more quickly and easily.
[0124] In the above embodiment, each screen of the animation creation application is depicted as a separate screen. However, these screens and any combination of two or more of the functions shown on these screens may be configured to be displayed on the same page. Furthermore, each screen may be configured, for example, so that the animation creation module 211 and each of the modules 212 to 217 cooperate with each other to display an appropriate screen on the display of the editing terminal 101 in accordance with the progress of the animation creation method. For example, the processing functions executed by each module may be represented by a node-based UI.
[0125] Furthermore, the animation creation module 211 and each of the modules 212 to 217 may cooperate with each other, for example, by displaying tab icons representing each screen above each screen, and when this tab icon is selected by the creator, the screen corresponding to the selected tab icon may be displayed on the display of the editing terminal 101.
[0126] In the above embodiment, when the creator selects the file selection field, for example, the animation creation system 1 displays a file selection dialog and accepts the designation of an arbitrary information file. However, the method of accepting the designation of an information file by the creator is not limited to this example. For example, when an information file is dragged and dropped onto the preview window 1070, 1260, the original frame image display field 1130, the 3D model display field 1615, etc., the information file can be accepted as the information file to be designated.
[0127] Note that the configuration of each screen and the configuration and functions of the buttons provided on each screen are not limited to the examples disclosed in the above embodiment. For example, in the above embodiment, after the text information acquisition module 214 accepts text information input by the creator into the positive instruction information field 1240 and the negative instruction information field 1250 on the image information generation screen 1200, the character image information generation module 215 causes the image generation model to generate character image information when it accepts the creator's selection of the image generation button 1270. However, for example, the character image information generation module 215 may be configured to cause the image generation model to generate image information about an image that reflects the content of the accepted text information each time the text information acquisition module 214 accepts text information input by the creator into the positive instruction information field 1240 and the negative instruction information field 1250. Furthermore, the character image information generation module 215 may be configured to display the generated image in the preview window 1260 each time the image generation model generates an image.
[0128] In the above embodiment, the background image information generation module 216 generates background image information representing an image related to a background using the second image generation model. However, the background image information generation module 216 may generate background image information using, for example, computer graphics without using an image generation model.
[0129] Furthermore, if the character image information generation module 215 creates image information of a character based on posture information (movement information) at a rate lower than the target frame rate, the character image information generation module 215 may create new image information by estimating the posture of the character at a time point between time points n and (n+1) based on, for example, the image information of the character at time point n and the image information of the character at time point (n+1). Such new image information can be created, for example, to cover the missing frame rate.
[0130] <Variation 1> Fig. 19 shows an example of a second animation creation flow. The animation creation module 211 acquires basic information necessary for animation creation, such as an original draft, plot, scenario, and script (S1901). The animation creation module 211 acquires information about the reading and information about the storyboard (S1902, S1903). The animation creation module 211 performs advance camera creation based on the storyboard information (S1904), sets up lighting, and acquires information about lighting (S1905).
[0131] The background image information generation module 216 generates three-dimensional information from the acquired point cloud data by performing a simple 3D scan of the background scene for shooting in accordance with the background image information generation flow 1800 of Figure 18 (S1907). The background image information generation module 216 adjusts the acquired three-dimensional information and adjusts the background scene for shooting (S1908). The animation creation module 211 also acquires character model information (S1909). The animation creation module 211 sets three-dimensional space information by inputting the information acquired in S1904, S1905, S1908, and S1909 into the three-dimensional space management platform (S1906).
[0132] The motion information acquisition module 213, which cooperates with the animation creation module 211, acquires motion information relating to the motion of an object performing an action, for example, based on the storyboard information (S1910). When acquiring the motion information, it is possible to set up and install a physical camera using motion capture, or it is also possible to set up and install a camera after the motion capture is completed based on the storyboard information to determine from what location the footage will be shot (S1913). The animation creation module 211 can perform editing processing and output additional material information (S1914).
[0133] Thereafter, the animation creation module 211 outputs the material that satisfies the conditions (S1911). At this time, the animation creation module 211 can write out an image of the character and background together, an image of just the character, an image of just the background, an image of just the shadow, etc., separately or all at once.
[0134] The animation creation module 211 also creates an offline video duration based on the exported material and additional material (S1920). The animation creation module 211 previews and modifies the offline first draft (S1915). The animation creation module 211 acquires dubbing sound source information and adjusts the duration based on the script (S1916). The animation creation module 211 repeats adjustments until the duration satisfies the conditions, and when the conditions are met, finalizes the duration (S1918). The animation creation module 211 then determines the generation duration and allocates work (S1919). This duration-based task allocation will be described later.
[0135] The animation creation module 211 determines the angle of view and shooting conditions of the structures and buildings in the scenes to be used in the actual animation according to the determined length, and the background image information generation module 216 acquires actual photographic information of the background material based on the angle of view and shooting conditions of the structures and buildings (S1922). This corresponds to S1830 in Fig. 18, in which a two-dimensional image (e.g., a photograph) corresponding to the angle of view cut out from the three-dimensional information is captured and acquired.
[0136] The animation creation module 211 acquires generation material information (S1921). For example, the animation creation module 211 acquires, as generation material information, facial expression information that includes facial expression information, such as a smiling mouth, tears, or angry eyebrows, for generating facial expressions, in line-drawn image information, as well as guide information for generating facial expressions. The animation creation module 211 inputs the acquired generation material information and a prompt (input information) indicating the type of output to an AI generation model (character image information generation module 315), causing the AI generation model to generate an original image of the character (S1920). The animation creation module 211 also inputs a 3D model with a specific angle of view determined by the confirmed length and background material photo information acquired to match this angle of view to the AI generation model (background image information generation module 216), causing the AI generation model to generate background information.
[0137] The animation creation module 211 retouches (corrects) the generated character original images and background information (S1924). The animation creation module 211 then generates a temporary composite for video confirmation (S1923) and performs a team preview to check the generated information (S1927). If necessary, the animation creation module 211 also performs drawing correction processing on the generated images (S1928). The animation creation module 211 acquires material information for compositing (S1929) and performs compositing (filming) (S1930). The animation creation module 211 displays a preview of the filmed animated video information for confirmation and correction by the director, etc.
[0138] The flow of acquiring audio information will be described below. The animation creation module 211 acquires post-recording sound source information (S1917). The animation creation module 211 also acquires sound effect information based on the determined length (S1925). The animation creation module 211 acquires the acquired post-recording sound source information and sound effect information, and performs sound adjustment processing (S1926). The animation creation module 211 then executes MA dubbing processing based on the adjusted audio information (S1932).
[0139] Thereafter, the animation creation module 211 acquires information such as subtitles (S1933) and executes video editing processing (S1934).The animation creation module 211 outputs delivery information (S1935) and executes distribution processing (S1936).
[0140] The allocation of work by the animation creation module 211 based on the generated scale in S1919 will now be described. Data used on the three-dimensional space management platform is managed by the number of consecutive frame images, while the various acquired video information is managed by the video shooting time. Therefore, the animation creation module 211 generates association information 2000 between them. FIG. 20 shows an example of image association information 2000. Source input 2001 indicates the time from the beginning of the acquired video material file, and source output 2002 indicates its end point. Timeline input 2003 indicates the start point of the story shown in the storyboard, and timeline output 2004 indicates its end point. Clip name 2005 indicates the storage destination of the original video material.
[0141] Clip input (frame) 2006 indicates the start frame in terms of the frame number used on the three-dimensional space management platform, and clip output (frame) 2007 indicates the end frame. That is, in the example of number 0001, frames 1102 indicated by clip input 2006 on the three-dimensional space management platform to 1162 indicated by clip output 2007 correspond to the time from 0:01:31:51 indicated by source input 2001 to 0:01:36:51 indicated by source output 2002 in the video file specified by the name A001.mp4 in clip name 2005. This time also corresponds to the time from 0:00:00:00 indicated by timeline input 2003 to 0:00:05:00 indicated by timeline output 2004 of the story created based on the storyboard.
[0142] Since errors may occur when generating images using AI, a two-frame buffer is provided as a backup, and the generation input (frame) 2008 is specified as a frame two frames before the clip input (frame) 2006. On the other hand, the generation output (frame) 2009 is specified as a frame two frames after the clip output 2007. That is, in the example of number 0001, the AI model outputs images from frame 1100 (generation input 2008) to frame 1164 (generation output), which are obtained by adding two frames before and after frame 1102 indicated by the clip input (frame) 2006 and frame 1162 indicated by the clip output (frame) 2007.
[0143] The number of frames to be generated 2010 indicates the number of frames to be generated. Confirmation 2011 confirms whether the data satisfies the conditions. All material information for editing, background images, etc. are managed with the same material length (duration), and when referencing edited data, it becomes easy to determine which parts should be drawn using the image association information 2000.
[0144] <Variant 2> In one aspect of the present technology, for example, a technology is provided for adding additional expression using an image generation model to a part of a character in image information representing an image of the character created as described above, and the animation creation method can include an additional expression adding step as an optional step.
[0145] FIG. 21 is an example of an additional expression assignment flow 2100. FIG. 22 is a schematic diagram showing an overview of the assignment of an additional expression. The additional expression assignment flow 2100 includes, for example, acquiring mask information representing a mask that specifies a portion of a character to which an additional expression is to be assigned (S2110), acquiring a first request phrase (typically, text information) requesting the creation of a predetermined additional expression image for the mask (S2120), combining the mask information and the first request phrase to create first input information (S2130), providing the first input information to an image generation model (S2140), acquiring first additional expression image information representing the additional expression image that is output from the image generation model in response to the first input information (S2150), and acquiring an image of the character to which the additional expression has been assigned by combining information representing an image of the character with the first additional expression image information (S2160).
[0146] Here, a case will be described in which an additional expression of tears (in other words, an additional expression accompanied by change or movement) is added to the face of a character in FIG. 24A. The image modification module 218, for example, acquires mask information representing a mask that specifies the part of the character to which tears are added (S2110). The mask information can be prepared in advance and stored in the auxiliary storage device 202 as image information 224, for example. FIG. 24B is a schematic diagram of a mask represented by the mask information. The mask is, for example, an element that indicates the part and path of tears that will form and flow, and can be set in an area where changes and movement of tears will appear. The mask in FIG. 24B corresponds to an area that includes the part below the character's pupil where tears form and collect, and the single teardrop that runs down the cheek from the eye.
[0147] The image modification module 218 acquires, for example, a first request phrase requesting the creation of an image (additional expression image) showing tears appearing and flowing in a region designated by a mask (S2120). The first request phrase may also include, for example, text information (text information) for conditioning the generation of an image by the image generation model. The first request phrase may be created using, for example, a screen for setting instruction information such as that shown in FIG. 12, or may be prepared in advance and stored in the auxiliary storage device 202 as text information 223. As this text information, the image modification module 218 may acquire positive instruction information indicating elements to be included in the image to be generated (e.g., in this example, phrases such as "anime," "overflowing tears: 1.2 seconds," and "tears falling down: 1.2 seconds"). The image modification module 218 may also acquire, as text information, negative instruction information indicating elements not to be included in the image to be generated (e.g., in this example, "cinematic," "low quality," etc.). Because the additional expression in this second variation involves change or movement, the first request phrase can be, for example, a phrase indicating the change or movement of tears over time. Furthermore, the first request phrase can include, for example, a specification of change or movement, or an instruction to turn the additional expression image into an animation (i.e., a collection of a series of still images representing change or movement). Furthermore, the first request phrase can include, for example, a numerical value indicating the reflection strength of the change or movement specification. Furthermore, the first request phrase can include, for example, only an instruction regarding the additional expression, and not an instruction regarding the character.
[0148] The image modification module 218, for example, combines the acquired mask information with the first request phrase to create first input information (S2130) and provides (inputs) this first input information to the image generation model (S2140). As a result, the image generation model generates an additional expression image in response to the first input information. FIG. 24(C) is an example of one additional expression image generated by the image generation model. This additional expression image includes only tears and does not include the character's face. This additional expression image also shows the stage in which tears accumulated at the top of the mask for the right eye are flowing downward along the mask, and the stage in which tears are in the middle of accumulating at the top of the mask for the left eye.
[0149] The image modification module 218 acquires first additional expression image information representing an additional expression image output from the image generation model (S2150). The image modification module 218 then acquires an image of the character to which an additional expression has been added by combining information representing the character's image with the acquired first additional expression image information (S2160). For example, the image modification module 218 can acquire multiple images of the character to which an additional expression has been added as a moving image (e.g., animation) by overlaying multiple additional expression images (e.g., additional expression images involving a series of changes or movements) on one or more images of the character.
[0150] FIG. 24(D) is an example of a character image (one scene) to which an additional expression has been added by overlaying an additional expression image on the character image. Even using an image generation model, it is difficult to generate an image with complex expressions with high accuracy. With the above configuration, it is possible to create an additional expression separately from the character image and combine the additional expression image with the character image. This makes it possible to use the image generation model to easily generate images with complex expressions with high accuracy.
[0151] <Modification 3> In one aspect of the present technology, for example, a technology is provided in which, using additional expression information prepared in advance, an additional expression is added to a part of a character in image information representing an image of the character created as described above, using an image generation model. In this case, the animation creation method can include an additional expression adding step as an optional step.
[0152] FIG. 22 is another example of the additional expression assignment flow 2200. FIG. 25 is a schematic diagram showing an overview of the assignment of an additional expression. The additional expression assignment flow 2200 includes, for example, acquiring mask information representing a mask that specifies a portion of the character to which the additional expression is to be assigned (S2210), acquiring additional expression information representing the additional expression (S2220), acquiring a second request phrase that requests the mask to create an additional expression image corresponding to the additional expression (S2230), combining the mask information, the additional expression information, and the second request phrase to create second input information (S2240), providing the second input information to an image generation model (S2250), acquiring second additional expression image information representing the additional expression image that is output from the image generation model in response to the second input information (S2260), and acquiring an image of the character to which the additional expression has been assigned by combining information representing an image of the character with the second additional expression image information (S2270).
[0153] Here, we will describe a case in which an additional expression (in other words, an additional expression involving change or movement) of flames surrounding the arm of a person in FIG. 25A is added. The image modification module 218, for example, acquires mask information representing a mask that specifies a portion of the character (here, the arm) to which flames are added (S2210). The mask information can be prepared in advance and stored in the auxiliary storage device 202 as image information 224, for example. FIG. 25B is a schematic diagram of a mask represented by the mask information. The mask is, for example, an element that indicates the portion of the arm from the elbow to which flames are surrounded, and can be set within a range where the change or movement of the flames appears. The mask in FIG. 25B corresponds to a region that completely surrounds the arm and hand from the elbow, with the surrounding region including a lighter mask density portion and a blurred mask boundary. The image modification module 218 also acquires additional expression information representing, for example, the change or movement of the flames to be added (S2220). 25C is a schematic diagram of a scene of flames represented by the additional expression information. The flame mask represented by the additional expression information is, for example, information representing a flame flickering over time.
[0154] Next, the image modification module 218 acquires a second request phrase requesting the creation of an additional expression image corresponding to the additional expression flame in the region indicated by the mask (S2230). The second request phrase may be created using, for example, a screen for setting instruction information such as that shown in FIG. 12, or may be prepared in advance and stored in the auxiliary storage device 202 as text information 223 or the like. The second request phrase may also include, for example, text-format information (text information) for conditioning the generation of an image by the image generation model. The image modification module 218 may acquire, as this text information, positive instruction information (e.g., in this example, black background, flame effect, fluid, etc.) that instructs the element to be included in the generated image. The image modification module 218 may also acquire, as text information, negative instruction information (e.g., in this example, chromatic aberration, yellow theme, etc.) that instructs the element not to be included in the generated image. Furthermore, in this third modification, since the additional expression information itself is information that involves change or movement, the second request phrase can be, for example, a phrase that instructs that the additional expression information be applied to the mask area in accordance with the shape of the mask. Furthermore, the second request phrase can include, for example, an instruction to further change or move the additional expression that involves change or movement. Furthermore, the second request phrase can include, for example, only an instruction regarding the additional expression, and not an instruction regarding the character.
[0155] The image modification module 218, for example, creates second input information by combining the acquired mask information, additional expression information, and the second request phrase (S2240), and provides (inputs) this second input information to the image generation model (S2250). This causes the image generation model to generate an additional expression image in response to the second input information. Figure 25(C) is an example of one additional expression image generated by the image generation model. This additional expression image includes only flames and does not include the character's arms or hands. This additional expression image also shows a scene of flickering flames clinging to the character's arms.
[0156] The image modification module 218 acquires second additional expression image information representing an additional expression image output from the image generation model (S2260). The image modification module 218 then acquires an image of the character to which the additional expression has been added by combining information representing the character image with the acquired second additional expression image information (S2270). For example, the image modification module 218 can acquire multiple images of the character to which the additional expression has been added as a moving image (e.g., animation) by overlaying multiple additional expression images (e.g., additional expression images involving a series of changes or movements) on one or multiple images of the character.
[0157] Figure 25 (D) is an example of a character image to which an additional expression has been added by overlaying an additional expression image on the character image. Even using an image generation model, it is difficult to generate images with complex expressions with high accuracy. With the above configuration, decorations and effect images that involve changes and movements are prepared in advance as additional expression information, and the additional expression information is formed into a desired form using the image generation model, and the formed additional expression image can be combined with a part of the character image. This makes it possible to use the image generation model to easily create images with more complex and diverse expressions with high accuracy.
[0158] <Modification 4> When generating images or videos using an image generation model, for example, it is possible to maintain consistency of characters, etc. in a one-time generation, but it is difficult (virtually impossible) to maintain complete consistency of characters, etc. across multiple generations. For example, in animation production, it is difficult to maintain consistency of characters, etc. for each cut scene. In one aspect of the present technology, for example, a technology is provided for correcting color tones for character images created multiple times. The animation creation method may include a step of correcting color tones as an optional step.
[0159] Fig. 23 is an example of a color correction flow 2300. Fig. 26 is a schematic diagram showing an overview of color correction. The color correction flow 2300 includes, for example, acquiring reference image information representing a reference image of a character (S2310), acquiring character image information representing an image of the character generated based on the reference image (S2320), acquiring a silhouette-based mask that removes the background of the reference image (S2330), acquiring color difference information regarding the color difference between the reference image with the mask applied and the image of the character with the mask applied, and correcting the color tone of the character image based on the color difference information so as to approximate the color tone of the reference image (S2340).
[0160] Here, we will explain the case where the color tone of a character image (target image) generated by an image generation model (B) is adjusted (corrected) based on a reference image (A) in Figure 26 that serves as a color tone reference. The image modification module 218 first acquires, for example, reference image information representing the character's reference image (S2310). This reference image information can be, for example, image information that can be used as input or reference image for the image generation model when generating an image to be color-corrected, such as image information with a predetermined color. Such image information can be, for example, information about an image that is faithful to the color settings and serves as a reference for generation, or a 3D model rendering image. However, since the reference image information does not provide information about texture, the texture may be simplified. The reference image (A) in Figure 26 is prepared to serve as a color tone reference, and typically has a foraging effect according to the settings, but a simplified texture. The reference image information may be, for example, prepared in advance and stored in the auxiliary storage device 202 as image information 224.
[0161] The image modification module 218 acquires character image information representing an image of a character generated by an image generation model based on the reference image (S2320). The character image information may be acquired directly or indirectly from the image generation model, or may be stored in the auxiliary storage device 202 as image information 224, for example. Although the character image information in Figure 26 (B) is given texture based on the overall features of the character set in the image generation model, some color variation from the reference image is unavoidable.
[0162] Therefore, the image modification module 218 acquires a silhouette-based mask that removes the background of the acquired reference image information (S2330). This mask is a mask that can extract the image portion of the character in the reference image information, and is, for example, information that indicates a form corresponding to the green screen background. When applied to the image of a character, such a mask can function to remove most of the background and extract the image portion of the character. Note that the mask acquisition process can be performed, for example, by automatically detecting the background (e.g., the green screen area) in the reference image information.
[0163] The image modification module 218 then obtains color difference information relating to the color difference between the reference image with the mask applied and the character image with the mask applied, and corrects the color tone of the character image based on the color difference information so that it approaches the color tone of the reference image (S2340). By performing color correction on the image from which the background has been removed using the mask in this way, it is possible to prevent changes in the background color from affecting the color difference information and the LUT.
[0164] The color difference information indicates, for example, the amount of adjustment required to adjust parameters of the color and tone of an image (e.g., hue, saturation, contrast, highlights, shadows, etc.) in order to map input color values to the desired output values. The color difference information can be obtained using, for example, a lookup table (LUT) or remapped (i.e., color correction). The image modification module 218 can perform consistent color correction for multiple images (e.g., multiple frame images) by linearly transforming color values using, for example, global parameters estimated from multiple images (e.g., corresponding to multiple frame images) generated based on reference image information. While not limited to this, the linear transformation of color values can be calculated based on the CIE Lab color space system, for example. This allows for more simple processing and natural color correction.
[0165] Figure 26 (C) shows an example of a character image after color correction. The character image after color correction is realized as a color tone very similar to that of the reference image, without compromising the texture and details of the character (target image) set in the machine learning model. By providing a reference image set to a desired color tone, a sense of unity can be achieved in the color tone of characters in animation production. Furthermore, such color correction can be performed automatically without requiring human intervention, significantly reducing the time required for manual color adjustment. Furthermore, color adjustment can be performed based on numerical calculations using a color space or other color system. This allows color correction to be performed while reducing human error. As a result, the quality of images generated by the image generation model can be improved.
[0166] In the above embodiment, the present technology has been described using an example of creating 2D animation, but animations created using the present technology are not limited to this example, and can be, for example, animations for computer games, AR animations, VR animations, etc. It should be noted that the present technology is preferably applied to animations made up of still images in a style similar to hand-drawn illustrations, known as so-called "anime," in that the advantages of the present technology are clearly exhibited.
[0167] Each module disclosed in the above embodiment may be configured by combining multiple sub-modules. Furthermore, some or all of the operations or functions performed by one module may be executed or realized by another module. In the above embodiment, one or more functional elements realized by the execution of each module of the user terminal 103 may be realized by the execution of a module of the editing terminal 101.
[0168] The present technology not only provides an invention related to an animation creation method, but also, in another aspect, provides a program for causing a computer and / or a processor to execute each step (process) of the animation creation method. This program may be composed of a single program or may be composed of two or more subprograms. Furthermore, the program may be a program for causing a computer to execute any one or two or more of the steps and / or substeps of the above method.
[0169] Furthermore, the present technology not only provides an invention related to an animation creation method, but also, in another aspect, can provide an invention related to an animation creation device, which includes one or more processors and a memory for storing one or more programs configured to be executed by the one or more processors, the one or more programs including the following instructions: set an image generation model, acquire motion information related to object motion, acquire text information related to character features, input the motion information and the text information into the image generation model to generate a plurality of image information representing images of the character, generate background image information representing a background image using a second image generation model, generate a plurality of frame image information representing an image obtained by combining the image of the character with the image of the background, and create video information based on the plurality of frame image information.
[0170] Furthermore, one or more of the one or more editing terminals 101 and one or more user terminals 103 constituting the animation creation system 1 according to the present technology may be installed in different countries. Furthermore, the editing terminal 101 may be realized by one or more computers, any of which may be installed in different countries.
[0171] The present invention is not limited to the above-described embodiments, but includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0172] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0173] Furthermore, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be considered that almost all components are interconnected. The above-mentioned embodiments disclose at least the configurations described in the claims.
[0174] 1... animation creation system, 101... editing terminal, 102... management server, 103... user terminal, 104... imaging device
Claims
1. A method for creating animation, comprising: setting an image generation model; acquiring motion information relating to the motion of an object; acquiring text information relating to the characteristics of a character; inputting the motion information and the text information into the image generation model to generate a plurality of image information representing images of the character; generating background image information representing an image of a background using a second image generation model; generating a plurality of frame image information representing an image obtained by combining the image of the character with the image of the background; and creating video information based on the plurality of frame image information.
2. The animation creation method according to claim 1, wherein the image generation model generates, as the plurality of pieces of image information, a plurality of pieces of image information representing images of the character having a posture corresponding to the posture of the object at a certain point in time of the action information.
3. The animation creation method according to claim 1, wherein the image generation model is a machine learning model trained to input text information and output an image related to the text information.
4. The animation creation method according to claim 1, wherein the step of setting the image generation model includes: acquiring an image generation model; setting overall information relating to the character's appearance and / or style in the image generation model; and setting partial information relating to some of the character's features in the image generation model.
5. The animation creation method according to claim 4, wherein the overall information includes information representing a two-dimensional animated image as information relating to the style, and image information representing a two-dimensional animated image of the character corresponding to the information representing the two-dimensional animated image is generated using the image generation model based on the overall information.
6. The animation creation method according to claim 4, wherein the partial information includes information representing at least one of personality and behavior of the character, and image information relating to an image that reflects the information representing at least one of personality and behavior is generated by the image generation model based on the partial information.
7. The animation creation method according to claim 1, wherein the step of acquiring the movement information includes: photographing or acquiring a video of the object; inputting the video into a posture estimation model that estimates posture information of the object contained in the image from the image information; and acquiring the movement information regarding the movement of the object.
8. The animation creation method according to claim 7, wherein the step of generating a plurality of pieces of image information representing images of the character comprises generating a plurality of pieces of image information representing images of the character based on movement information relating to the movement of the object obtained by inputting the video information into the posture estimation model.
9. The animation creation method according to claim 1, wherein the movement information is motion capture data of the object, and image information representing a plurality of images of the character that show movements corresponding to the motion capture data is generated by the image generation model based on the movement information.
10. The animation creation method of claim 1, wherein the step of generating image information includes: acquiring an image generation model; setting positive instruction information from the text information in the image generation model; setting negative instruction information from the text information in the image generation model; and generating, based on the action information and the positive and negative instruction information, a plurality of pieces of image information representing images of the character that reflect the positive instruction information and are less likely to reflect the negative instruction information, using the image generation model.
11. A method for creating animation as described in claim 10, comprising the steps of: when image information representing multiple generated images of the character does not have image information for each frame that satisfies the conditions, setting text information for adjusting detailed facial expressions and movements for each frame in the image generation model; and generating image information representing an image of the character that reflects the text information for adjusting detailed facial expressions and movements for each frame based on the movement information and the text information.
12. The animation creation method according to claim 1, wherein the step of generating background image information includes: acquiring three-dimensional information about the background; extracting from the three-dimensional information an image with a field of view that matches the storyboard; inputting the extracted image and text information into a second image generation model to generate multiple pieces of background image information; and extracting background image information that satisfies certain conditions from the multiple pieces of generated background image information.
13. The animation creation method according to claim 1, further comprising adding an additional expression to a part of the character, wherein the step of adding the additional expression comprises: acquiring mask information representing a mask that specifies the part of the character to which the additional expression is to be added; acquiring a first request phrase requesting creation of a predetermined additional expression image for the mask; creating first input information by combining the mask information and the first request phrase; providing the first input information to an image generation model; acquiring first additional expression image information representing the additional expression image that is output from the image generation model in response to the first input information; and acquiring an image of the character to which the additional expression has been added by combining information representing an image of the character with the first additional expression image information.
14. The animation creation method according to claim 1, further comprising adding an additional expression to a part of the character, wherein the step of adding the additional expression comprises: acquiring mask information representing a mask that specifies the part of the character to which the additional expression is to be added; acquiring additional expression information representing the additional expression; acquiring a second request phrase that requests the mask to create an additional expression image corresponding to the additional expression; creating second input information by combining the mask information, the additional expression information, and the second request phrase; providing the second input information to an image generation model; acquiring second additional expression image information representing the additional expression image that is output from the image generation model in response to the second input information; and acquiring an image of the character to which the additional expression has been added by combining information representing an image of the character with the second additional expression image information.
15. A method for creating animation as described in claim 1, further comprising correcting the color tone of the generated image of the character based on the color tone of a reference image, wherein the step of correcting the color tone comprises: acquiring reference image information representing the reference image of the character; acquiring character image information representing the image of the character generated based on the reference image; acquiring a silhouette-based mask that removes the background of the reference image; acquiring color difference information regarding the color difference between the reference image with the mask applied and the image of the character with the mask applied; and correcting the color tone of the image of the character based on the color difference information so as to approach the color tone of the reference image.
16. An animation creation system comprising: a model setting means for setting an image generation model; a motion information acquisition means for acquiring motion information relating to the motion of an object; a text information acquisition means for acquiring text information relating to the characteristics of a character; a character image information generation module for inputting the motion information and the text information into the image generation model and generating a plurality of pieces of image information representing images of the character; a background image information generation module for generating background image information representing an image of a background using a second image generation model; and a synthesis means for generating a plurality of frame image information representing an image obtained by combining the image of the character with the image of the background, and creating video information based on the plurality of frame image information.
17. A program that causes a computer to execute each step of the animation creation method described in any one of claims 1 to 15.
Citation Information
Patent Citations
Dynamic cartoon generation method and device, storage medium and electronic equipment
CN117252966A
Image processing apparatus, image processing method, and program
JP2016091100A
Image creation device, image creation method, and program
JP2019204476A
Image generation device, prompt creation support device, program, and application program
JP7398723B1
Server, system, method and program providing inpainting service using generative ai model
KR102655359B1