Inference program, inference method, inference device, learning program, learning method, and learning device
The inference and learning programs address the lack of recorded finishing data in anime production by using a trained model to generate and apply coloring operations, enhancing automation and consistency in the anime production process.
Patent Information
- Application Number
- JP2025020797
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-08-21
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The anime production process lacks recorded data for the finishing process, as it was traditionally done manually and not optimized for learning, with existing technologies failing to capture the intended coloring of areas surrounded by line drawings of various shapes.
An inference program and learning program that utilize a trained model to generate finishing images by reading learning data in tensor format, capturing drawing area changes due to coloring operations, and repeatedly generating intermediate images until output conditions are met.
Enables the inference of an appropriate finishing process for input images, facilitating automation and consistency in anime production by leveraging a trained model to learn and apply coloring operations.
Smart Images

Figure 0007727867000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an inference program, an inference method, an inference device, a learning program, a learning method, and a learning device. [Background technology]
[0002] In the animation (hereafter referred to as "anime") production process, various tasks are highly specialized, and many staff members in animation production studios are involved in the production of anime. Anime production involves processes such as "key animation," which involves drawing each scene from the storyboards written by the director or producer, "inbetween animation," which involves drawing the images (inbetween frames) between the key animations, and "finishing," which involves coloring the animation. In recent years, digital equipment has been used in the finishing process, but the finishing process itself is still often done by hand by the finishing staff. For this reason, there has been a demand for automation of the finishing process.
[0003] A known technique for adding color to an image is disclosed in Patent Document 1. Patent Document 1 discloses a technique for making an image look natural by using mask data consisting of an extraction region and a mask region to add a predetermined color to at least a part of the periphery of an extraction region in image data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2023-102665 Summary of the Invention [Problem to be solved by the invention]
[0005] The anime production process is highly specialized, and the content and meaning of each task, such as "keyboard drawings," "animation," and "finishing," are clearly defined, so a typical anime production studio has a large amount of finishing data.
[0006] However, in animation production, the finishing process, that is, the process by which the finishers apply color, has never been recorded. The reasons for this include the fact that, in the past, the finishing process was often done manually by the finishers, and there was no benefit to recording the coloring process. As a result, the only data available in animation production studios was a pair of line drawings before coloring and the finished image after coloring the line drawings. This data, as it is, was not suitable for learning about the finishing process.
[0007] Even if the process of coloring a line drawing is recorded on video, the colorist frequently enlarges, reduces, moves, etc. the line drawing during the finishing process. Therefore, the video also records screen changes that are essentially unrelated to coloring, such as the trajectory of mouse operations and the enlargement or reduction of the image editing canvas, making the data unsuitable for studying the finishing process.
[0008] Furthermore, some specialized image processing software used in animation production records in a buffer the "Undo" and "Redo" actions (hereinafter also referred to as "Undo / Redo") that are included in the work of the finishing staff. However, data that records such Undo / Redo actions is not optimized for learning the finishing process. Furthermore, while the technology disclosed in Patent Document 1 is a computer graphics technology, this technology was unable to finish areas surrounded by line drawings of various shapes in animation production as intended by the finishing staff.
[0009] The present invention has been made in view of the above circumstances, and has as its object to infer an appropriate finishing process for an input image. [Means for solving the problem]
[0010] The inference program of the present invention causes a computer to execute the following steps: reading learning data in tensor format that stores, in chronological order, drawing area information of a drawing area that has changed due to coloring operations on a learning image drawn in a fixed-size drawing area; reading from a recording unit a trained model that has learned a finishing process including coloring operations; causing the trained model to repeatedly generate intermediate finishing images until the intermediate finishing images that the trained model has inferred as the finishing process for an input image satisfy output conditions; and outputting the intermediate finishing images that satisfy the output conditions as output images. The above-described inference program is one aspect of the present invention, and an inference method and an inference device that reflect one aspect of the present invention are configured in the same manner as the above-described inference program.
[0011] In addition, the learning program of the present invention reads learning data in tensor format that stores, in chronological order, drawing area information of a drawing area that has changed due to coloring operations on a learning target image drawn in a fixed-size drawing area, and causes a computer to execute the following steps: have the learned model learn a finishing process including coloring operations; and record the learned model in a recording unit. The above learning program is one aspect of the present invention, and a learning method and learning device that reflect one aspect of the present invention are configured in the same manner as the above learning program. [Effects of the Invention]
[0012] According to the present invention, an appropriate finishing process for an input image can be inferred using a trained model that has been trained on finishing processes. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is an overall configuration diagram showing an overview of an image generation system according to an embodiment of the present invention; [Figure 2] 1 is a block diagram illustrating an example of a hardware configuration of an information processing terminal according to an embodiment of the present invention. [Figure 3] 1 is a block diagram showing a hardware configuration of a learning device according to an embodiment of the present invention. [Figure 4] 1 is a block diagram illustrating an example of a functional configuration of a learning device according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of a coloring process performed on a line drawing according to an embodiment of the present invention. [Figure 6] FIG. 2 is a diagram specifically illustrating the processing of the learning device according to one embodiment of the present invention. [Figure 7] 10 is a flowchart illustrating an example of a learning process performed by the learning device according to one embodiment of the present invention. [Figure 8] FIG. 2 is a block diagram illustrating an example of the functional configuration of a diffusion model according to an embodiment of the present invention. [Figure 9] 1 is a block diagram showing the hardware configuration of an inference device according to an embodiment of the present invention. [Figure 10] 1 is a block diagram illustrating an example of the functional configuration of an inference device according to an embodiment of the present invention. [Figure 11] 10 is a flowchart illustrating an example of an inference process performed by an inference device according to one embodiment of the present invention. [Figure 12] 1A and 1B are diagrams showing examples of in-process line drawing data and finished image data according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted.
[0015] [One embodiment] In the embodiment described below, a trained image generation model (referred to as a "trained model") is used to perform the finishing process, particularly in animation production, to generate a finished image by finishing a line drawing. A diffusion model is used as the trained model in this embodiment. The image generation system in this embodiment has the function of outputting a desirable image as an output target image by repeatedly performing the finishing process on an image generated by an image generation model that has been trained to infer the next finishing state from a specific finishing state, including an uncolored state, in order to assist specific finishing processes performed by humans.
[0016] The finisher in this embodiment is assumed to be someone who colors uncolored line drawings. By referring to the design code during drawing, the finisher can maintain consistency of characters throughout a segment of animation (e.g., an entire scene, an entire episode, or an entire work) for characters whose details vary depending on the scene or image being drawn. The design code may be any information that can be used to determine the accuracy of an object depicted in an image. The design code may be, for example, a combination of text data about the object (information about the position and size of each component) and image data of the object, or may be information extracted from such information, or learning data learned using the text data and image data. An image generation system including a learning device and an inference device according to this embodiment will now be described.
[0017] <Example of overall configuration of image generation system> First, an example of the configuration of an image generation system according to an embodiment of the present invention will be described. This image generation system is configured by combining a learning device that learns finishing processes, an inference device that generates images based on inferred finishing processes, and an information processing terminal.
[0018] <Image Generation System Overview> FIG. 1 is a diagram showing the overall configuration of an image generation system 1 according to an embodiment of the present invention. The image generation system 1 includes a tablet terminal 2A, a PC (Personal Computer) 2B, a learning device 30, and an inference device 60. The tablet terminal 2A and the PC 2B can be connected to the learning device 30 and the inference device 60 via a network N such as the Internet. In the following description, the tablet terminal 2A and the PC 2B will be collectively referred to as the information processing terminal 2.
[0019] The learning device 30 is a device that learns the finishing process in animation production. The inference device 60 is a device that causes a trained model to perform the finishing process on a line drawing to generate a finished image. For this purpose, the learning device 30 and the inference device 60 manage programs and various data used in image generation. A diffusion model 100 shown in FIG. 8 (described later) is used as the trained model. Image data of the image generated by the inference device 60 is transmitted to the information processing terminal 2.
[0020] The information processing terminal 2 can store image data generated by the inference device 60 inferring the coloring process in a recording device 22 shown in FIG. 2, which will be described later. Here, the information processing terminal 2 is described as being operated by a user. The user is assumed to be a person who operates the information processing terminal 2 and instructs the generation of an image. The user may also include the above-mentioned person in charge of finishing.
[0021] The tablet terminal 2A constituting the information processing terminal 2 uses a touch panel display device in which the input device 26 and the output device 27 are integrated. The tablet terminal 2A may be a tablet terminal. In addition, the input device 26 and the output device 27 are separate devices in the PC 2B. Note that the PC 2B may be a desktop PC, with the input device 26 and the output device 27 separately connected to the desktop PC.
[0022] The information processing terminal 2 selects a program based on an operation signal input from the input device 26 in response to an operation performed by the user, and outputs a video signal corresponding to the screen of the output device 27 to the output device 27. The output device 27 displays an image based on the video signal. The operation signal input from the input device 26 is, for example, a signal corresponding to each operation button on a keyboard. The user can input instructions through the input device 26, instruct the inference device 60 to execute an inference program, or operate the inference program recorded in the recording device 22 of the user's terminal. Operations on the inference program input from the input device 26 include, for example, various command inputs such as specifying a line drawing to be finished and the details of the finishing. Another example of an operation performed from the input device 26 is a tap operation in which the screen of the output device 27 is touched with a finger or a pen.
[0023] The information processing terminal 2 performs processes such as reading image data from the recording device 22 and executing a program, drawing a screen in accordance with an operation signal input from the input device 26, and displaying a screen by the output device 27. For example, in response to an operation by a user via the input device 26, the information processing terminal 2 displays on the output device 27 a screen on which an image based on the image data read from the recording device 22 is drawn.
[0024] The information processing terminal 2 provides the inference device 60 with line drawing data for inferring the coloring process. The information processing terminal 2 also receives input from the learning device 30, which allows the learning device 30 to learn the coloring operations performed by the finisher. In this embodiment, the coloring operations included in the finishing process refer to the coloring process of the line drawing, and are operations performed by the finisher using a mouse, digital pen, or the like to finish the line drawing, correct the line drawing, and determine the color of part of the line drawing. Therefore, the coloring operation includes the correction process for the line drawing. For example, the finisher may add lines where lines are broken or modify the shape of a character's eyes; these operations are also included in the coloring operation. The inference program according to this embodiment, which runs on the inference device 60, generates a finished image in which the finishing process has been performed on the line drawing, based on instruction information input via the input device 26 (an example of an input unit).
[0025] The information processing terminal 2 can also record the inference program used by the inference device 60 in the recording device 22 (see FIG. 2 described later) and execute the inference program read from the recording device 22 or the like. In this case, the information processing terminal 2 can generate an image in which the coloring process has been inferred within the terminal itself, without uploading image data consisting of only a line drawing to the inference device 60. Furthermore, when image data of an image generated by the inference device 60 is transmitted to the information processing terminal 2, the information processing terminal 2 can also display the image using, for example, an internet browser.
[0026] <Example of hardware configuration for image generation system> Next, an example of the hardware configuration of the image generation system 1 according to an embodiment will be described. 2 is a block diagram showing an example of the hardware configuration of the information processing terminal 2. Examples of the hardware configuration of the learning device 30 and the inference device 60 will be described later.
[0027] (Example of information processing terminal configuration) The information processing terminal 2 is an example of a computer that operates as a computer capable of executing various programs. The information processing terminal 2 includes a processor 21, a recording device 22, and a network interface 24, all of which are connected to a bus 23.
[0028] The processor 21 is configured with at least one of, for example, a central processing unit (CPU), a microprocessor unit (MPU), a graphics processing unit (GPU), and a field programmable gate array (FPGA). The processor 21 reads program code of image processing software that realizes each function according to this embodiment from the recording device 22, loads it into a temporary storage unit (not shown) provided in the recording device 22, and executes the program code. The processor 21 performs, for example, arithmetic processing for image generation and processing required to draw a GUI of the image processing software on the output device 27 of the information processing terminal 2. The processor 21 also performs processing such as processing of the OS of the information processing terminal 2 and management of input and output of data performed by each unit within the information processing terminal 2. When processing information related to image generation, the processor 21 can output an image signal to the output device 27 via the input / output interface 25.
[0029] The recording device 22 is configured by, for example, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM may be an optical disk, a magneto-optical disk, a DVD (Digital Versatile Disc)-ROM, a CD-ROM, a Blu-ray (registered trademark) disk, or the like. The RAM may be an SRAM, a DRAM, or the like. Variables, parameters, and the like generated during the arithmetic processing of the processor 21 are temporarily written to the recording device 22, and these variables, parameters, and the like are read out by the processor 21 as appropriate. The processor 21 also performs processing required to draw, for example, a GUI of image processing software on the screen of the output device 27.
[0030] The recording device 22 is configured by at least one of, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), and a flash memory. The recording device 22 stores the OS of the information processing terminal 2, various parameters, programs for causing the information processing terminal 2 to function, programs for generating images, etc. As described above, the recording device 22 stores programs, data, etc. necessary for the processor 21 to operate, and is used as an example of a computer-readable non-transitory storage medium that stores programs executed by the information processing terminal 2.
[0031] For example, a network interface card (NIC) or the like is used as the network interface 24. The network interface 24 can transmit and receive various data to and from the learning device 30 and the inference device 60 via a dedicated line or the like connected to a terminal of the NIC and via the network N, and can also communicate with other information processing terminals 2.
[0032] The input / output interface 25 converts operation signals received from the input device 26 into data in a predetermined format and passes the converted data to the processor 21. The input / output interface 25 also converts screen data drawn by the processor 21 into video signals and outputs them to the output device 27.
[0033] The input device 26 is a device that accepts input instructions or various types of information from a user. An example of the input device 26 is a pointing device that can input coordinate information of a position designated by the user. This pointing device is a mouse, a touch panel device, or the like. A touch panel device is configured by combining the input device 26 and the output device 27. The input device 26 may also be a keyboard, a mouse, or the like.
[0034] The output device 27 is a device that outputs information processed by the processor 21. The output device 27 is, for example, a display device (a display device, a touch panel device, etc.). When the output device 27 is a display device, an image based on a video signal received from the input / output interface 25 is displayed on the display device.
[0035] <Example of hardware configuration for learning device> Next, an example of the hardware configuration of the learning device 30 will be described. 3 is a block diagram showing the hardware configuration of a learning device 30 according to one embodiment of the present invention. The learning device 30 is an example of a system configured to include one or more devices and for generating a trained model 44 (see FIG. 4, described later), but in the following embodiment, for convenience of explanation, it will be described as a single device. The system for generating the trained model 44 can also refer to the learning device 30. The same applies to an inference device 60, described later.
[0036] The learning device 30 includes a processor 31, an input device 32, a display device 33, a recording device 34, and a communication device 35. These components are connected by a bus 36. Note that an interface is interposed between the bus 36 and each component device as necessary. The learning device 30 includes a configuration similar to that of a general server, PC, etc.
[0037] The processor 31 controls the overall operation of the learning device 30. For example, the processor 31 is at least one of a CPU, an MPU, a GPU, and an FPGA. The processor 31 performs various processes by reading and executing programs and data stored in the recording device 34. The processor 31 may be composed of multiple processors.
[0038] Input device 32 is a user interface that accepts input from the user to study device 30, and is, for example, a touch panel, touchpad, keyboard, mouse, or button. Display device 33 is a display that displays application screens and the like to the user of study device 30 under the control of processor 31.
[0039] The recording device 34 (an example of a recording unit) includes a main memory device and an auxiliary memory device. The main memory device is, for example, a semiconductor memory such as RAM. RAM is a volatile storage medium that allows high-speed reading and writing of information, and is used as a storage area and a working area when the processor 31 processes information. The main memory device may also include ROM, which is a read-only nonvolatile storage medium. The auxiliary storage device stores various programs and data used by the processor 31 when executing each program. The auxiliary storage device may be any nonvolatile storage or nonvolatile memory that can store information, and may be removable.
[0040] The communication device 35 transmits and receives data to and from other computers such as user terminals or servers via a network, and may be, for example, a wireless LAN module. The communication device 35 may be a device or module for other wireless communication such as a Bluetooth (registered trademark) module, or may be a device or module for wired communication such as an Ethernet (registered trademark) module or a USB interface. The system configuration and data structure of this embodiment will be described in detail below.
[0041] <Example of functional configuration of learning device> FIG. 4 is a block diagram showing an example of the functional configuration of a learning device 30 according to an embodiment of the present invention. The learning device 30 has a buffer 42, tensor learning data 43, and a trained model 44 stored in a recording device 34. Note that the history data 41 is theoretical data prepared for the purpose of explanation, and the learning device 30 does not actually store the history data 41. The trained model 44 is an example of a deep learning model. The learning device 30 also has functional blocks that have the functions of a history data generation unit 51, a history data compression unit 52, a finishing process data acquisition unit 53, a tensor encoder 54, a data acquisition unit 55, and a finishing process learning unit 56. Each function of these functional blocks is realized by the processor 31 shown in FIG. 3.
[0042] The history data generator 51 is a function of the image processing software used by the finishing artist to perform coloring operations. It generates history data 41 representing the history of coloring operations. The history data 41 is theoretical data representing a bitmap history of an image corresponding to a part of the animation production process (e.g., the coloring process) and is associated with a single line drawing. The history data 41 includes difference information representing each coloring operation performed on the canvas as a coloring operation history. The drawing area information for the drawing area changed by the coloring operation represents the difference between the previous coloring operation performed by the finishing artist and the current coloring operation. The coloring operation includes at least one of filling in the learning image drawn with line drawings, correcting the line drawing, adding line drawings, and deleting line drawings. Here, the history data generator 51 generates the history of all coloring operations performed on all pixels constituting the canvas of the image processing software during the finishing process as history data 41. For example, the canvas size is assumed to be 1920 x 1080 pixels. Thus, the historical data 41 represents coloring operations on a canvas normalized to an aspect ratio for internet or television broadcast.
[0043] The finisher performs coloring operations on multiple images simultaneously, but the history itself is linked to each individual image. In this embodiment, each image representing the history of the finishing process performed by the finisher is collectively referred to as history data 41. Here, a specific example of the coloring process performed on a line drawing drawn on a canvas will be described.
[0044] Figure 5 shows an example of a coloring process applied to a line drawing: Figure 5 shows a line drawing 40 of a car and an example of a coloring process applied to the car's windows.
[0045] The line drawing 40 is an example of a learning image whose coloring process is learned by the learning device 30, and is drawn on a fixed-size canvas cv (an example of a drawing area). The line drawing 40 is the foreground, not the background, of the animation. The canvas cv is in a bitmap format with a fixed size. In addition to line drawings, images that are already partially colored may also be used as learning images.
[0046] The order of coloring processes for a line drawing is indicated, for example, as coloring processes 1 to 4. Note that the color of the car windows is the same, but to clarify the order of the coloring processes, the coloring process is represented as coloring sections pr1 to pr4 with different hatching. The coloring process is performed by the finisher moving a digital pen multiple times on specific areas of the line drawing. The coloring operations included in the coloring process can include various aspects realized by image processing software. Any operation that can link color information to bitmap coordinates is acceptable. For example, the finisher can specify a portion of a closed area and color the entire area. To prevent uncolored areas from being left uncolored, the finisher can specify an area and only uncolored areas within the specified area are colored. Another operation involves painting the same area twice with the same color, which functions as an eraser, erasing the color of the area where the operation was performed. Another coloring process involves the finisher using an eraser tool to erase areas that have already been painted or repainting them with a different color. These operations are collectively referred to as the coloring process. The coloring process is included in the coloring operations of this embodiment.
[0047] The image processing software according to this embodiment has information about all pixels that make up the canvas cv. Therefore, the learning device 30 can uniquely identify the part of the canvas cv that has undergone the coloring process, regardless of the operations of the finishing technician.
[0048] <Specific example of processing by the learning device> Next, the compression of the history data 41 will be described with reference to FIG. 6 is a diagram specifically illustrating the processing of the learning device 30. Image processing software is provided to the user by the learning device 30. The learning device 30 enables the coloring process indicated by the history data 41 to be output as a tensor for a diffusion model.
[0049] The history data 41 is theoretical data that represents the application history of the coloring process within the image processing software as a change history of all pixels on a fixed-size canvas cv. If a finishing technician completes coloring of one line drawing image in about 10 minutes, the coloring process will be performed approximately 1,500 times. In this case, approximately 1,500 recorded images are referred to as the history data 41. When the history data 41 is played back, the process of coloring the line drawing on the fixed-size canvas cv can be displayed as a moving image on the information processing terminal 2.
[0050] As described above, in the finishing process, a coloring process is performed on one line drawing. For example, in the line drawing of a car shown in FIG. 5, four coloring processes, designated as coloring processes 1 to 4, are performed on the window portion. Each time a coloring process is performed, the difference between the current coloring process and the previous coloring process is stored as difference information in the history data 41. In other words, the difference information is information that indicates the difference between the previous coloring operation and the current coloring operation, which occurs for each coloring process (operation by the finishing technician).
[0051] The screen size of the canvas cv is, for example, 1920 x 1080 pixels, but it may be reduced from 1920 x 1080 pixels (for example, 1754 x 1060 pixels) to represent the area where the actual coloring process will be performed. Each pixel that makes up the canvas cv is a 24-bit sRGB pixel. The history data 41 shown in Figure 6 is a collection of uncompressed bitmaps for the entire canvas cv, so it contains approximately 7.8 GB of data.
[0052] The history data compression unit 52 converts only the difference information from the history data 41 into an array and stores it in the buffer 42 as finishing process data 42a. Because the data volume of the finishing process data 42a is smaller than the data volume of the history data 41, the process performed by the history data compression unit 52 is referred to as "compression." Therefore, the history data compression unit 52 compresses the coloring operation history based on the history data 41 into finishing process data 42a, which represents the differences resulting from changes made during the coloring operation. For convenience, FIG. 6 shows the history data compression process being performed on the history data 41, which includes 1,500 images. However, in reality, the difference information is stored in the buffer 42 as the finishing process data 42a each time a coloring operation is performed. In other words, the color information and position coordinate array of pixels in the history data 41 whose drawing area has changed during the coloring operation is represented as the finishing process data 42a with the difference information compressed. The changed drawing area represents changes made during one or more coloring operations.
[0053] In the example shown in FIG. 6 , the history data compression unit 52 accumulates the processing history in the buffer 42 as an array with 1500 elements. The processing history refers to the coloring process in which a finisher paints certain pixels with a color specified by RGB. Even if the process appears to a human as an eraser or a solid color, the history of painting a group of pixels with a specific color is stored in the buffer 42 without distinguishing between these operations. By storing the processing history in the buffer 42 in this way, the steps taken by the finisher to finish an image can be learned as Next Token Prediction learning in the Transformer. In this embodiment, since the learning involves predicting the next frame, Next Token Prediction learning is considered Next Frame Prediction learning.
[0054] As described above, the history data compression unit 52 saves only information about coloring change locations, which are differences from the previous frame (e.g., the previous coloring process), as an array in the buffer 42. As a result, the history data 41, which has a data volume of approximately 7.8 GB in uncompressed bitmap image data, can be compressed to approximately 205 MB. In the example shown in FIG. 5, coloring process 1 is the difference from the original line drawing. Similarly, coloring process 2 is the difference from coloring process 1. Similarly, each coloring process is the difference from the previous coloring process. In other words, for each coloring process, four arrays corresponding to the colored colors are created in the buffer 42.
[0055] The buffer 42 records all coloring operations on the canvas CV as an array consisting of 4-byte color information, a 4-byte x-coordinate array, and a 4-byte y-coordinate array (int4 rgb, int4[] x int4[] y). Each array is expressed as "RGB, X[], Y[]." Each array represents a single coloring operation, and one or more arrays are recorded in the buffer 42 for one or more coloring operations. These arrays are referred to as finishing process data 42a. Each pixel on the canvas CV, on which the line drawing is drawn, can be identified by its x- and y-coordinates, with the upper left corner as the origin. There is no limit to the size of the buffer 42 shown in Figure 6. For example, a single coloring operation often changes less than 1% of the coordinates of the bitmap image data representing the canvas CV. Therefore, on average, the data volume per coloring operation is approximately 145 KB, and the data volume per image is approximately 205 MB, multiplied by 1500. The 1500 times value used in this calculation corresponds to 1500 sheets of history data 41 shown in FIG. 6.
[0056] In the finishing process of animation production, one person usually finishes the work in a short time of about 10 to 20 minutes, and the image size is fixed and only coloring is performed. It is precisely because of this finishing process that the present embodiment can record the history of all pixels in the buffer 42.
[0057] The finishing process data acquisition unit 53 acquires, from the buffer 42, finishing process data 42a including the finishing process that the trained model 44 is to learn. For example, a person in charge of finishing may specify the finishing process that the trained model 44 is to learn, or the finishing process data acquisition unit 53 may automatically read the finishing process data 42a stored in the buffer 42. The finishing process data 42a acquired by the finishing process data acquisition unit 53 is output to the tensor encoder 54.
[0058] The tensor encoder 54 encodes the finishing process data 42a into a tensor format to generate tensor training data 43. The tensor encoder 54 is a module that converts the finishing process data 42a, which has been compressed based on history data 41 that records the color and the coordinate array of the changed pixels during each coloring process performed by the finishing technician, into tensor data that can be learned by the Transformer.
[0059] Furthermore, in the finishing process, the person in charge of finishing may manually correct the movement of the character's eyes and mouth during the coloring stage, and this correction process also contributes greatly to the beauty of the final animation. Therefore, the tensor encoder 54 has the function of efficiently tensorizing the entire coloring process, focusing on the fact that the coloring process is fast and requires little trial and error.
[0060] The tensor training data 43 output by the tensor encoder 54 is an example of training data in tensor format, storing difference information (drawing area information of the drawing area changed by the coloring operation) of the canvas cv (see FIG. 5) changed by the coloring operation performed by the finisher in chronological order. Therefore, the tensor training data 43 has a tensor structure corresponding to the trained model 44 in the subsequent stage. For example, the tensor training data 43 is an array obtained by converting a fixed-size bitmap image and is positioned as a three-dimensional tensor. In other words, the normalized pixel values of the bitmap image shown as history data 41 on the left side of FIG. 6 become a tensor.
[0061] Conventional image processing software provides a buffer for undoing and redoing processes, resulting in a buffer structure that is dependent on the specific software. In contrast, in this embodiment, since only the processing results are important, the color history of all pixels is stored as pixel color changes in buffer 42. By storing pixel color changes in buffer 42, tensor encoder 54 can directly convert buffer 42 into a tensor data structure required by trained model 44, an example of a deep learning model.
[0062] Returning to the explanation of Figure 4. The finishing process learning unit 56 receives the tensor learning data 43 acquired by the data acquisition unit 55 as input, and outputs a trained model 44 that has been trained to learn the finishing process, including coloring operations. The finishing process includes not only the coloring process described above, but also correction of line drawings, etc. The finishing process learning unit 56 is a core module of deep learning, and can use any model that can generate video data from image data.
[0063] Specifically, the trained model 44 uses a Transformer network to learn the relationships between image sequences, and uses the image sequences as input to learn the movement of objects within the sequences with high accuracy. In this embodiment, the finishing process learning unit 56 causes the trained model 44 to learn the coloring process based on tensor learning data 43 generated from video images of the coloring process performed by a finisher. Therefore, in order to reproduce the behavior of a highly skilled finisher, a time-series learning model is adopted for the trained model 44. In this embodiment, the finishing process learning unit 56 uses as input to the trained model 44 tensor learning data 43, which is a collection of individual images recorded during the coloring process, i.e., tensor learning data 43 composed of 3D tensors, which are arrays of bitmaps, which are 2D arrays. The finishing process learning unit 56 causes the trained model 44 to learn the coloring process.
[0064] As described above, the trained model 44 is a Transformer-based time-series learning model that uses an image generation architecture based on a latent diffusion model. While this embodiment does not rely on a specific diffusion model, it is preferable to use a simple diffusion model when targeting 2D animation. Furthermore, any model capable of generating video data (i.e., a sequence of related images) from image data can be used as the trained model 44.
[0065] <Example of learning process> FIG. 7 is a flowchart showing an example of the learning process performed by the learning device 30. First, the history data generation unit 51 generates history data 41 for each coloring process based on the coloring operations performed by the finisher (S1). Note that the processing by the history data generation unit 51 may be performed separately from the learning processing. In this case, the learning processing shown in FIG. 7 may start from step S2.
[0066] The history data compression unit 52 compresses the difference between the previous operation and the current operation into a coloring process sequence based on the history data 41 to create finishing process data 42a, which is stored in the buffer 42 configured in the recording device 34 (S2). The processing of steps S1 and S2 is performed for each coloring operation by the finishing technician, so there is no need for the history data 41 to be recorded in the recording device 34.
[0067] Next, the finishing process data acquisition unit 53 acquires the finishing process data 42a from the buffer 42 (S3). The finishing process data 42a acquired by the finishing process data acquisition unit 53 is output to the tensor encoder .
[0068] Next, the tensor encoder 54 encodes the finishing process data 42a into a tensor format, and stores the tensor learning data 43 in the recording device 34 (S4).
[0069] Next, the data acquisition unit 55 acquires the tensor learning data 43 corresponding to the learning target of the finishing process from the recording device 34 (S5).
[0070] Next, the finishing process learning unit 56 learns the finishing process based on the tensor learning data 43, and generates a trained model 44 that has learned the finishing process (S6). The learning process of the finishing process will be described in detail later.
[0071] Then, the finishing process learning unit 56 records the trained model 44 in the recording device 34 (S7), and ends this process. The trained model 44 recorded in the recording device 34 can be read by the inference device 60, which will be described later.
[0072] <Learning process for finishing process> Here, the learning process of the finishing process performed in step S6 will be described with reference to FIG. Fig. 8 is a block diagram showing an example of the functional configuration of the diffusion model 100. Here, a manner in which the trained model 44 shown in Fig. 4 learns a finishing step will be described. Note that Fig. 8 also includes a process in which the trained model 44 infers a finishing step for a line drawing, but this process will be described after the description of Figs. 9 to 11.
[0073] In recent years, the use of significantly advanced image generation technology for animation production has been considered. For example, by using AI in image generation technology, it is possible to generate a large number of images in a short period of time by inputting certain generation conditions into the AI. As an example of an image generation model, in particular, a diffusion model uses random numbers to generate noise, which is the starting point for image generation, and it is known that the resulting image changes significantly each time the random number is changed. This mechanism utilizing random numbers is inherently beneficial because it adds diversity to the images output by the image generation model and enables the generation of a wide variety of variations. For this reason, the diffusion model 100 described below is also adopted in this embodiment.
[0074] The diffusion model 100 includes a VAE encoder 101, a latent diffusion model (LDM) 102, a noise removal model 104, and a VAE decoder 105. The diffusion model 100 is fine-tuned to learn a finishing process (particularly a coloring process) for inputted line drawing tensor training data 43. In an inference process in an inference device 60 (described later), the diffusion model 100 generates an inferred image in which the finishing process has been applied to the line drawing based on inputted instruction information (e.g., information from a text instruction unit 106).
[0075] The diffusion model 100 includes a diffusion process in which input original data is converted into random noise, and a dediffusion process in which the original data is reconstructed from the noise. The diffusion process is a process in which the tensor training data 43 in tensor space, encoded by the VAE encoder 101, is encoded into latent space, and then the latent diffusion model 102 adds random information (e.g., noise) to the encoded tensor training data 43. For this reason, the finishing process training unit 56 uses the diffusion model 100 to gradually add noise to the original data (e.g., the tensor training data 43) to generate a noise image 103. That is, in the process of training the finishing process, the finishing process training unit 56 generates the noise image 103 by repeating a process of predicting the image of the next frame based on the image of the previous frame.
[0076] The dediffusion process is a process that uses the diffusion process to remove random information (e.g., noise) from the noisy image 103, and decodes the data from which the random information has been removed into pixel space to generate an image. In other words, the trained model 44 is a model that has learned how to remove noise from the noisy image 103, which has become a random number state, and reconstruct the original data. The trained model 44 can generate in-process finishing line drawing data 82 that infers the finishing process for the pre-finishing line drawing data 81. This trained model 44 includes a noise removal model 104, a VAE decoder 105, and a finishing process inference unit 72.
[0077] The tensor learning data 43 acquired by the finishing process learning unit 56 is input to the VAE encoder 101. As described above, the tensor learning data 43 is generated based on the history data 41, in which one coloring process performed by a finisher is represented by one image, by compressing difference information as finishing process data 42a and then encoding it into a tensor format. The VAE encoder 101 extracts features of the line drawing and the coloring process from the tensor learning data 43.
[0078] Furthermore, the finishing process learning unit 56 can output the details of the coloring process explained in text to the text instruction unit 106. For example, a detailed description of the pose, camera angle, clothing, etc. of the character depicted in the line drawing is expected. When the details of the tensor training data 43 are automatically or semi-automatically described as text by the finishing process learning unit 56, the text instruction unit 106 encodes the text and inputs it to the noise removal model 104. Therefore, the encoded text information input from the text instruction unit 106 to the noise removal model 104 is expected to become appropriate parameters for the diffusion model 100.
[0079] The operation of each part of the diffusion model 100 will be described in detail below. First, the latent diffusion model 102 has the function of a general diffusion model with the VAE encoder 101 as the input layer, and generates a noise image 103.
[0080] The noise removal model 104 removes noise from the noisy image 103. For example, a U-net is used as the noise removal model 104. The U-net is a type of FCN (fully convolution network), and is a network for estimating image segmentation (object position).
[0081] The VAE decoder 105 decodes the noise-removed image to generate an inferred image. The image decoded and generated by the VAE decoder 105 is an image similar in atmosphere to the image shown in the tensor learning data 43 previously used in the learning process by the finishing process learning unit 56. Note that the noise removal model 104 incorporates the finishing process inference unit 72 layer by layer, but to avoid complication in explanation, the noise removal model 104 and the finishing process inference unit 72 are described separately.
[0082] The finishing process inference unit 72 is a module for generating a video image that is consistent along the time axis as an additional layer to the diffusion model 100. The finishing process inference unit 72 can infer the finishing process using a trained model 44 that has learned the coloring process by a finisher as motion. Therefore, in the training process in the training device 30, the noise reduction model 104 removes noise from the noisy image 103, and the VAE decoder 105 learns the process of decoding the data into the same data as the original data input to the VAE encoder 101.
[0083] When the trained model 44 is used in the inference device 60, the finishing process inference unit 72 inputs a noise image 103, which is obtained by diffusing pre-finishing line drawing data 81 (an example of an input image) using a latent diffusion model 102, into the trained model 44 during the inference process, and outputs a line drawing for which the coloring process for the pre-finishing line drawing data 81 is inferred. This line drawing is an image that is expected to be colored for the pre-finishing line drawing data 81. For this reason, the line drawing for which the finishing process inference unit 72 infers the coloring process from the noise image 103 is an image that represents the intermediate progress of the finishing process, and is referred to as in-process finishing line drawing data 82.
[0084] (Example of hardware configuration of inference device) Next, an example of the configuration of the inference device 60 will be described. 9 is a block diagram showing the hardware configuration of an inference device 60 according to one embodiment of the present invention. The inference device 60 includes a processor 61, an input device 62, a display device 63, a recording device 64, and a communication device 65. These components are connected by a bus 66. Note that an interface is interposed between the bus 66 and each component device as necessary. The inference device 60 includes a configuration similar to that of a general server, PC, etc.
[0085] The processor 61 controls the overall operation of the inference device 60. For example, the processor 61 is at least one of a CPU, an MPU, a GPU, and an FPGA. The processor 61 performs various processes by reading and executing programs and data stored in the recording device 64. The processor 61 may be composed of multiple processors.
[0086] The input device 62 is a user interface that accepts input from a user to the inference device 60, and is, for example, a touch panel, a touch pad, a keyboard, a mouse, or a button. The display device 63 is a display that displays application screens and the like to the user of the inference device 60 under the control of the processor 61.
[0087] The recording device 64 includes a main memory device and an auxiliary memory device. The main memory device is, for example, a semiconductor memory such as RAM. RAM is a volatile storage medium that allows high-speed reading and writing of information, and is used as a storage area and a working area when the processor 61 processes information. The main memory device may also include ROM, which is a read-only nonvolatile storage medium. The auxiliary storage device stores various programs and data used by the processor 61 when executing each program. The auxiliary storage device may be any nonvolatile storage or nonvolatile memory that can store information, and may be removable.
[0088] The communication device 65 transmits and receives data to and from other computers such as user terminals or servers via a network, and may be, for example, a wireless LAN module. The communication device 65 may be a device or module for other wireless communication such as a Bluetooth (registered trademark) module, or may be a device or module for wired communication such as an Ethernet (registered trademark) module or a USB interface.
[0089] (Example of functional configuration of inference device) 10 is a block diagram showing an example of the functional configuration of an inference device 60 according to one embodiment of the present invention. The inference device 60 includes a reading unit 71, a finishing process inference unit 72, and an inference control unit 73. Each function of these functional blocks is realized by the processor 61 shown in FIG.
[0090] The reading unit 71 accesses the recording device 34 of the learning device 30 and reads the trained model 44 that has undergone series training of the finishing process from the recording device 34. The reading unit 71 also reads the pre-finishing line drawing data 81 uploaded from the information processing terminal 2 from the recording device 64.
[0091] The finishing process inference unit 72 repeatedly causes the trained model 44 to generate in-process finishing line art data 82 (an example of an in-process finishing image), which the trained model 44 has inferred as a finishing process for pre-finishing line art data 81 (an example of an input image), until the inference control unit 73 determines an output condition. For example, the finishing process inference unit 72 receives as input the in-process finishing line art data 82 generated by the trained model 44 inferring a coloring process at frame t, and causes the trained model 44 to infer a coloring process at frame t+1 and again generate in-process finishing line art data 82. The inference control unit 73 performs the process of adding frame t+1 to frame t. In the following description, the process of the finishing process inference unit 72 generating in-process finishing line art data 82 for the trained model 44 may be described as the inference process of the finishing process by the finishing process inference unit 72.
[0092] The output condition is, for example, at least one of the following: the ratio of the colored portion of the pre-finish line drawing data 81 to the entire pre-finish line drawing data 81 has reached a predetermined ratio (e.g., 100%); and an instruction to stop inference has been input. Note that in the coloring operation, even if the line drawing is not colored, a special color (e.g., white) designated as "not colored" is applied to the corresponding portion, so that there are no uncolored areas in the finished line drawing. Therefore, the above-described ratio of the colored portion of the pre-finish line drawing data 81 to the entire pre-finish line drawing data 81 having reached a predetermined ratio (e.g., 100%) also includes the fact that there are no uncolored areas in the line drawing. Note that other conditions may be set as output conditions. The instruction to stop inference may be input from the input device 62 shown in FIG. 9 or from the input device 26 of the information processing terminal 2 shown in FIG. 2.
[0093] The inference control unit 73 outputs the in-process finishing line art data 82 that satisfies the output conditions as finished image data 83 (an example of an output image). To this end, the inference control unit 73 controls the generation of the next frame of the image using the image inferred by the finishing process inference unit 72 as input. The inference control unit 73 according to this embodiment is a module that performs autoregressive inference on the automatic coloring process as a prediction process for the next frame. The inference control unit 73 recursively uses the frame generated by the finishing process inference unit 72 as an input image, causing the finishing process inference unit 72 to generate in-process finishing line art data 82 in which the coloring process of the image has been inferred almost automatically.
[0094] When inference control unit 73 receives in-process line drawing data 82, it determines whether in-process line drawing data 82 satisfies the output conditions. If inference control unit 73 determines that in-process line drawing data 82 satisfies the output conditions, it outputs in-process line drawing data 82 as finished image data 83.
[0095] The finished image data 83 is recorded in the recording device 64 of the inference device 60 and is made available for download by the information processing terminal 2 as appropriate. If the inference control unit 73 cannot find an image that satisfies the output conditions, it again causes the finishing process inference unit 72 to infer the finishing process and output the in-process finishing line drawing data 82.
[0096] (Example of inference processing) 11 is a flowchart showing an example of inference processing performed by the inference device 60. Here, the inference processing will be explained with reference to the trained model 44 and the functional blocks of the inference control unit 73 shown in FIG.
[0097] First, the reading unit 71 reads the trained model 44 from the recording device 34 of the learning device 30 (S11). Note that the trained model 44 may be copied in advance from the recording device 34 of the learning device 30 to the recording device 64 of the inference device 60, and the reading unit 71 may read the trained model 44 from the recording device 64 in the processing of step S11.
[0098] Next, the reading unit 71 reads the pre-finishing line drawing data 81 from the recording device 64 (S12). The pre-finishing line drawing data 81 is data that has been uploaded from the information processing terminal 2 and stored in the recording device 64 for image inference processing.
[0099] Next, the text instruction unit 106 acquires instruction information (e.g., a prompt character string) from the finishing process learning unit 56 (S13). This instruction information indicates the image that the noise removal model 104 will generate by removing noise. Note that the text instruction unit 106 may acquire instruction information input by the user. The text instruction unit 106 encodes the acquired instruction information into a vector format and outputs it to the noise removal model 104.
[0100] The text indicator 106 may be, for example, a text encoder such as CLIP (Contrastive Language-Image Pre-training). When a finisher uses the inference device 60, trial and error in operation is a bottleneck, so there is a need to eliminate this trial and error. Therefore, by using a mechanism called CLIP, it is expected that the user will be able to specify prompts to control parameters within the diffusion model. The trained model 44 is a model that has undergone sequential learning of the finishing process and can infer the finishing process based on text indicator information in text format. As described above, the text indicator 106 encodes the input text indicator information into the form of an embedded vector and outputs the encoded text indicator information to the noise reduction model 104. The diffusion model 100 removes noise from the noise image 103 based on the encoded text indicator information.
[0101] The text instruction information in text format input from the text instruction unit 106 is, for example, information obtained by encoding a prompt string. The prompt string is, for example, information such as "1 girl, portrait, black hair, smile" and is used as a positive prompt and a negative prompt. The positive prompt is a string of elements that the diffusion model 100 is desired to generate, and the diffusion model 100 generates an image according to the elements written in the positive prompt. On the other hand, the negative prompt is a string of elements that the diffusion model 100 is desired to exclude from the image generated by the diffusion model 100, and the elements of the image generated by the diffusion model 100 are deleted according to the elements written in the positive prompt. In this way, the diffusion model 100 generates an image according to the text instruction information.
[0102] Next, the VAE encoder 101 performs encoding processing on the input pre-finishing line drawing data 81 (S14). By this encoding processing, important features of the pre-finishing line drawing data 81 are extracted.
[0103] Next, the latent diffusion model 102 performs a diffusion process (S15) to add random information (e.g., noise) to important features of the encoded pre-finish line drawing data 81, thereby generating a noise image 103 of the pre-finish line drawing data 81.
[0104] Next, the noise removal model 104 removes noise from the noise image 103 based on the instruction information in vector format input from the text instruction unit 106 (S16).
[0105] Next, the VAE decoder 105 decodes the noise-removed image, and the finishing process inference unit 72 performs finishing inference processing on the noise-removed image (S17). As described above, the VAE decoder 105 is configured such that the finishing process inference unit 72, shown later in the figure, imports the data layer by layer, so the decoding processing and finishing inference processing are performed collectively. The image that has undergone the decoding processing and finishing inference processing is output to the inference control unit 73 as in-process finishing line drawing data 82.
[0106] The inference control unit 73 determines whether the in-process finishing line drawing data 82 satisfies the output conditions (S18). If the inference control unit 73 determines that the in-process finishing line drawing data 82 does not satisfy the output conditions (NO in S18), it instructs the finishing process inference unit 72 to perform finishing inference processing for the next frame, which is frame t+1, instead of frame t (S19). Thereafter, the processing of step S18, in which it is determined whether the in-process finishing line drawing data 82 for the next frame inferred by the finishing process inference unit 72 satisfies the output conditions, is repeated until the output conditions are satisfied.
[0107] On the other hand, if the inference control unit 73 determines that the in-process line drawing data 82 satisfies the output conditions (YES in S18), it outputs the frame of the in-process line drawing data 82 that satisfies the output conditions as finished image data 83 (S20), and terminates this processing.
[0108] The finished image data 83 is recorded in the recording device 64. The finished image data 83 is also output to the information processing terminal 2, so that the person in charge of finishing can check the finished image displayed on the output device 27 based on the finished image data 83.
[0109] Here, specific examples of the in-process line drawing data 82 and the finished image data 83 will be described. 12 is a diagram showing an example of in-process line drawing data 82 and finished image data 83. In FIG. 12, an example of a coloring process inferred for a line drawing of a dragon is shown.
[0110] An example of pre-finishing line drawing data 81 is shown in the top row of FIG. 12. The pre-finishing line drawing data 81 is data consisting of only line drawings. A coloring process is performed on the pre-finishing line drawing data 81. For example, suppose that the coloring process performed by the person in charge of finishing colors the tip of the beak, the entire beak, ears, mouth, wings, and eyes in this order for a line drawing of a dragon. In this case, as a result of the trained model 44 learning the coloring process, if the pre-finishing line drawing data 81 is a line drawing of a dragon, an image is automatically generated in which the tip of the beak, the entire beak, ears, mouth, wings, and eyes are colored in this order.
[0111] For example, the second row and subsequent rows in Figure 12 show images of the in-process line drawing data 82 from the first frame to the Nth frame. In the first frame, it is shown that part of the line drawing (the tip of the beak) has been colored. In the second frame, it is shown that part of the line drawing (the entire beak) has been colored. In this way, as the number of frames increases, the part of the line drawing that is colored increases. Note that in Figure 12, for ease of explanation, the tip of the beak and the entire beak are expressed, but in reality, a line drawing is generated in which a further portion of the tip of the beak is colored by about several tens of pixels.
[0112] The change in coloring is almost imperceptible in the (N-2)th frame, the (N-1)th frame, and the Nth frame. For example, if the inference control unit 73 determines that the inferred image at the (N-2)th frame satisfies the output conditions, it stops generating images from the next frame onwards and outputs the (N-2)th frame image as finished image data 83.
[0113] The value N specified for the number of frames may be any integer. For example, N may be set to 5, and the coloring process for the line drawing may end at the 5th frame, or N may be set to 150, and the coloring process for the line drawing may end at the 150th frame. The value of N may change when the inference control unit 73 determines that the output conditions are met.
[0114] Furthermore, the person in charge of finishing can check the quality of the finished image data 83 using the information processing terminal 2. If the finished image data 83 is not of the quality expected by the person in charge of finishing, the person in charge of finishing may give an instruction to add a coloring process to the finished image data 83, and obtain the finished image data 83.
[0115] In the image generation system 1 according to the embodiment described above, learning data extraction is built into the image processing software, enabling the extraction of learning data without any changes to the existing animation production process. This allows for natural integration into the existing animation production process. Furthermore, the process by which the trained model 44 reproduces the coloring process reflects the order of the finishing steps performed by the finishing technician. Therefore, the quality of the finished image data 83 is expected to be similar to that of the line drawing colored by the finishing technician.
[0116] The applicant of the present application can obtain training data for training the trained model 44 and construct the image generation system 1 according to this embodiment based on the following assumptions (1) to (3) specific to the animation industry. (1) Finishing (coloring) is a process of applying a pre-specified color, and there is no need for the person in charge of finishing to verify the success or failure of the finishing operation, so there is very little trial and error involved in the coloring process. Therefore, the task can be completed in a short time of about 10 minutes per image, and since the tensor does not contain trial and error noise, it is easy to train a trained model 44 to perform the finishing process. (2) Each animation studio has its own manual, and the coloring process is highly systematized, so there is little individual variation in the finishing process. (3) The finishing process is part of the process of creating a video work, and since the canvas size is completely fixed, the coordinate information within the canvas can be directly mapped to the elements of the tensor.
[0117] Furthermore, this embodiment discovers the affinity of finishing processes (including coloring processes) that have been overlooked in animation production studios as learning data, enabling unprecedented self-supervised learning. Furthermore, this embodiment can directly incorporate the work of the finishing process performed by the finishing staff in the animation production studio as tensor learning data 43. Therefore, self-supervised learning becomes possible by using tensor learning data 43 created solely by the finishing staff in the animation production studio. Furthermore, the tensor learning data 43 can be continuously supplied to a trained model 44 that is specialized for the production process specific to the animation production studio, allowing the trained model 44 to continue learning the finishing process.
[0118] While no practical self-supervised learning method has been established in the past, the applicant has an animation production studio within its group that employs many finishing staff and has a systematic image coloring process, which has enabled it to establish a practical self-supervised learning method in this embodiment.
[0119] As described above, the learning device 30 can achieve self-supervised learning by tensorizing the image coloring process within the image processing software. In a specific implementation of this embodiment, a buffer 42 for the coloring history of all pixels is introduced into the image processing software itself, allowing the coloring process to be output as video and applied to the learning data of the diffusion model 100 for video. The buffer size required by the buffer 42 according to this embodiment is estimated. In this calculation, assuming a 1920 x 1080 pixel canvas and 40% of the entire canvas is colored an average of three times, 2.5 million histories ≒ 340,000 bytes ≒ 30 MB are required. For example, if a finisher paints a certain area with one color, erases it, and paints it with a different color, this means that the area has been colored three times. As described above, the buffer size is small, and the buffer 42 can be configured using either the JavaScript memory used in web browsers or an IndexedDB. This allows the finishing process data 42a to be easily stored in the buffer 42. This method of tensorizing the coloring process as a change in coloring on a fixed-size canvas is not known anywhere other than this embodiment.
[0120] The concept of using autoregressive inference to automatically color line drawings as a next-frame prediction process according to this embodiment may be applicable to various other animation production processes. Furthermore, the finishing processes in various genres of animation produced by animation production studios can be used as learning data to train the trained model 44. As a result, even in the production of animation in genres that animation production studios have not previously handled, it is possible to obtain images in which the finishing process has been automatically performed with sufficiently high quality, thereby significantly shortening the finishing process.
[0121] Alternatively, the history data generation unit 51 may store drawing area information for the entire drawing area (information for all pixels in the drawing area) in chronological order in the history data 41. The history data compression unit 52 may then read the drawing area information in chronological order from the history data 41 in which the drawing area information for the entire drawing area is stored, and store finishing process data 42a, which is the drawing area information compressed based on the difference in the drawing area information in chronological order, in the buffer 42. Even in this form, it is possible to extract learning data without making any changes to the existing animation production process, and it is possible to generate a trained model 44 through training of the finishing process by the finishing process learning unit 56.
[0122] Furthermore, the present invention is not limited to the above-described embodiment, and it goes without saying that various other applications and modifications are possible without departing from the gist of the present invention as set forth in the claims. For example, the above-described embodiment has described the system configuration in detail and specifically to clearly explain the present invention, and is not necessarily limited to a system including all of the described configurations. Furthermore, it is also possible to add, delete, or replace part of the configuration of this embodiment with other configurations. In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0123] 1...image generation system, 2...information processing terminal, 26...input device, 27...output device, 30...learning device, 41...history data, 42...buffer, 42a...finishing process data, 43...tensor learning data, 44...trained model, 51...history data generation unit, 52...history data compression unit, 53...finishing process data acquisition unit, 54...tensor encoder, 55...data acquisition unit, 56...finishing process learning unit, 60...inference device, 71...reading unit, 72...finishing process inference unit, 73...inference control unit, 81...pre-finishing line drawing data, 82...finishing in-process line drawing data, 83...finishing image data, 100...diffusion model
Claims
1. a step of reading learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning image drawn in a drawing area of a fixed size is stored in chronological order, and reading from a recording unit a trained model that has learned a finishing process including the coloring operation; A step of repeatedly generating an intermediate finishing image, inferred by the trained model from the finishing process for the input image, until the intermediate finishing image satisfies an output condition; a step of outputting the in-process finishing image that satisfies the output conditions as an output image; An inference program to be executed by a computer.
2. The drawing area information of the drawing area changed by the coloring operation is difference information representing a difference between the current coloring operation and the previous coloring operation, and the history of the coloring operation is compressed into finishing process data representing the difference changed during the coloring operation based on history data representing the history of the coloring operation including the difference information; The finishing process data is encoded in a tensor format to generate the training data. The inference program according to claim 1 .
3. The training data in tensor space is encoded into a latent space, and then a diffusion process in which a latent diffusion model adds random information to the encoded training data, and a de-diffusion process in which the random information is removed based on the diffusion process and the data from which the random information has been removed is decoded into a pixel space. The trained model is used to generate the intermediately finished image by inferring the finishing process for the input image. The inference program according to claim 2 .
4. The output condition is at least one of: a ratio of a colored portion of the input image to the entire input image reaches a predetermined ratio; and an instruction to stop inference is input. The inference program according to claim 3.
5. the learning image includes the drawing area in a bitmap format configured with a fixed size, Among the history data, the color information and position coordinate array of the pixels in the portion where the drawing area has changed during the coloring operation are compressed as finishing process data. The inference program according to claim 4.
6. The changed portion of the drawing area is a portion of the drawing area that has been changed by one or more of the coloring operations. The inference program according to claim 5 .
7. The trained model is a model that has been trained on the finishing process in a sequential manner, and infers the finishing process based on text instruction information in a text format. The inference program according to claim 4.
8. The coloring operation includes at least one of filling in the learning image drawn with a line drawing, modifying the line drawing, adding the line drawing, and deleting the line drawing. The inference program according to claim 4.
9. A step of reading learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning image drawn in a drawing area of a fixed size is stored in chronological order, and reading out from a recording unit a trained model that has learned a finishing process including the coloring operation; A step of repeatedly generating an intermediate finishing image, inferred by the trained model from the finishing process for the input image, until the intermediate finishing image satisfies an output condition; and outputting the in-process image that satisfies the output conditions as an output image. Reasoning method.
10. a reading unit that reads learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning target image drawn in a drawing area of a fixed size is stored in chronological order, and reads from a recording unit a trained model that has learned a finishing process including the coloring operation; a finishing process inference unit that causes the trained model to repeatedly generate an intermediate finishing image inferred from the finishing process for the input image until the intermediate finishing image satisfies an output condition; an inference control unit that outputs the in-process finishing image that satisfies the output conditions as an output image; Reasoning device.
11. a step of reading learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning target image drawn in a drawing area of a fixed size is stored in chronological order, and having a trained model learn a finishing process including the coloring operation; A procedure for recording the trained model in a recording unit. A learning program for a computer to run.
12. the drawing area information of the drawing area changed by the coloring operation is difference information indicating a difference between the current coloring operation and the previous coloring operation, and among history data indicating the history of the coloring operation including the difference information, color information and an array of position coordinates of the part of the drawing area changed during the coloring operation are compressed as finishing process data and recorded in a buffer; obtaining the finishing process data from the buffer; encoding the finishing process data into the training data in tensor form. The learning program according to claim 11.
13. A step of reading learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning target image drawn in a drawing area of a fixed size is stored in chronological order, and having a trained model learn a finishing process including the coloring operation; and recording the trained model in a recording unit. How to learn.
14. a finishing process learning unit that reads learning data in a tensor format in which drawing area information of a drawing area that has been changed by a coloring operation on a learning target image drawn in a drawing area of a fixed size is stored in chronological order, and causes a trained model to learn a finishing process including the coloring operation; a recording unit that records the trained model; Learning device.
Citation Information
Patent Citations
Color information estimation model generating device, moving image colorization device, and programs for the same
JP2019117559A
Image coloring device, image coloring method, and program
JP2024029442A
Information processing system, information processor, information processing method
JP2025010938A
User interface system, user interface method, and image editing device
WO2022085775A1
Image processing program, recording medium, and image processing method
JP2023102665A