Deformation Character Generation Program and Deformation Character Generation System
The system addresses the challenge of controlling the head-to-body ratio in deformed character images by using a machine learning model to generate deformed character images based on user-input ratios and original character images, resulting in efficient and accurate image generation.
Patent Information
- Application Number
- JP2025014514
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-31
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2045-01-31
AI Technical Summary
It is difficult to control the head-to-body ratio of deformed character images when generating a deformed character from a normal character, requiring significant labor and ingenuity in selecting which parts to keep and which to delete.
A computer-based system that acquires the head-to-body ratio setting from user input, an original character image, and character information, and uses these inputs to a machine learning model to generate a deformed character image, learning from character and deformed character images as data.
The system allows for the efficient generation of deformed character images with a controlled head-to-body ratio, reducing labor and improving accuracy in reflecting the original character's features.
Smart Images

Figure 0007698856000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deformed character generation program and a deformed character generation system.
Background Art
[0002] In contents in various fields such as games, animations, education, and news reporting, so-called deformed characters such as two-headed characters are used. These give a cute impression to viewers and have effects such as making it easier to feel familiar with the content.
[0003] Although two-headed characters may be created as original characters, it is also possible to deform a character originally drawn in an eight-headed body or the like into a two-headed character.
[0004] However, when creating such deformed characters, it is necessary to devise ways to leave characteristic parts such as expressions and clothing while deleting unnecessary parts, which requires a lot of labor.
[0005] For example, Patent Document 1 discloses an image conversion device that can convert a character image into a deformed character with a two-headed or three-headed body.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Non-Patent Documents
[0007]
Non-Patent Document 1
[0008] In recent years, by using generative AI and other tools, the labor of illustrators can be reduced, for example, when an original character illustration is input, other illustrations with different expressions, clothing, etc. can be generated.
[0009] Non-Patent Document 1 discloses a technique for outputting a character image of preference by inputting a character image and a pose.
[0010] However, on the other hand, when, for example, an 8-head character is input and a 2-head character is used as the output of the generative AI, there are several problems. One is that when deforming, it is necessary to select what to keep (the emphasized part) and what not to keep (the non-emphasized part).
[0011] The physical characteristics and clothing of characters are diverse. For example, there are characters with characteristics in each of hairstyles, eye colors, clothing designs, items held, and ornaments. When deforming, what parts to keep and what parts to delete vary depending on the character. If the hairstyle changes when deforming a character with a characteristic hairstyle, it can be said that the deformation is not appropriate.
[0012] However, when trying to obtain the desired deformed character as the output of the image generation AI, there is a problem that a great deal of labor is required, such as relying on the ingenuity of the prompt.
[0013] Another is that since the impression received by the viewer (the person who views the deformed character) varies greatly depending on the head-to-body ratio of the deformed character, it is necessary to correctly control the head-to-body ratio of the deformed character. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0014] The problem to be solved is that it is difficult to control the head-to-body ratio of the deformed character image when generating a deformed character from a normal character. MEANS FOR SOLVING THE PROBLEM
[0015] The main features of the present invention are to acquire the setting of the head-to-body ratio by the user's input, acquire the original character image (for example, an 8-head-to-body character image), and acquire a deformed character image obtained by deforming the original character image to the head-to-body ratio from these pieces of information.
[0016] Non-Patent Document 1 does not describe converting to a deformed character or controlling the head-to-body ratio.
[0017] The present invention has been made in view of the above problems, and employs, for example, the following means. That is, a computer is caused to function as an original character image acquisition means for acquiring an original character image that is an image of a character, a character prompt acquisition means for acquiring character information for explaining the original character image, a head-to-body ratio acquisition means for acquiring the setting of the head-to-body ratio by the user's input, and a deformed character acquisition means for acquiring a deformed character image obtained by deforming the original character image to the head-to-body ratio, using the original character image, the character information, and the head-to-body ratio as inputs to a machine learning model. The machine learning model learns at least character images and deformed character images as learning data, and provides a deformed character generation program that takes a character image as input and outputs a deformed character image that deforms the character image.
Advantages of the Invention
[0018] The deformed character generation program of the present invention has the advantage of being able to output a deformed character image that reflects the height-to-width ratio based on the user's input. At this time, the user also has the advantage of being able to input the height-to-width ratio by a simple operation using the GUI.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0020] Embodiments of the present invention will be described based on the drawings. In the following embodiments, the same or corresponding parts may be denoted by the same reference numerals, and the description may be omitted as appropriate. Also, the drawings used below are for explaining this embodiment, and may be different from the actual device configuration, user interface (UI), data configuration, etc.
[0021] (Overview of the Embodiment) The overview of this embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an overview of the processing by the deformed character generation program P1 of this embodiment.
[0022] The deformed character generation program P1 is a program that takes an image of a tall character (the 5 - 6 head - to - body character in the upper left in FIG. 1) as input and outputs an image of a short character (the 2 - 3 head - to - body character in the lower right in FIG. 1).
[0023] The deformed character generation program P1 includes: (1) an original character image analysis means (upper right in FIG. 1) that acquires the original character image and various information (character prompt, character model, etc.); (2) a reference image analysis means (lower left in FIG. 1) that acquires the reference image and the control model (first embodiment); and (3) a deformed character acquisition means (lower right in FIG. 1) that outputs the deformed character based on the information obtained in (1) and (2).
[0024] Also, a system that performs the processing by the deformed character generation program P1 is hereinafter referred to as the "deformed character generation system 1".
[0025] In addition to the control based on the reference image corresponding to the head-to-body ratio set by the user, there is also a method of converting the head-to-body ratio set by the user into character information (prompt) for control in the control of the head-to-body ratio. Hereinafter, in the first embodiment, the former will be described, and in the second embodiment, the latter will be described.
[0026] (Details of the embodiment) Hereinafter, the deformable character generation system 1 according to the present embodiment will be described in detail. The deformable character generation system 1 includes a computer (server 10) equipped with a deformable character generation program P1, and provides an online system that deforms and outputs a character image related to an input to the user. That is, in the deformable character generation system 1, the information processing by the deformable character generation program P1 is specifically realized using hardware resources. Hereinafter, the 1. user interface, 2. program processing, and 3. hardware configuration that constitute the deformable character generation system 1 will be described in order.
[0027] (Definition of terms) Here, some terms will be defined. A "character" has a visible appearance, regardless of whether it is a living thing, non-living thing, real or imaginary. Here, in particular, it means something that can be divided into a head and the rest of the body of the whole body. For example, it is a character of a creative anime (living thing, imaginary). Here, a non-living character means, for example, a robot character. Furthermore, a real character is, for example, something like a celebrity drawn in an anime style. Note that, for example, even if the place of the neck is hidden, such as when a person is completely covered with a sheet from the head, and the head is not clear, as long as the approximate position can be inferred. Also, for example, a part with a face (eyes and mouth) may be regarded as the head. A "character image" is an image of a character, regardless of whether it is an illustration image or a photo image. A "pose" means the posture taken by a character. "Deformation" refers to intentionally changing the form of an object. For example, it refers to changing the head-to-body ratio of a character to make it from a tall-headed character to a short-headed character. "Tall-headed character" and "short-headed character" generally refer to a character with a high head-to-body ratio (e.g., 6 to 9 heads) and a character with a low head-to-body ratio (e.g., 2 to 5 heads), respectively. However, the high and low here are relative rather than absolute. When referring to deformation, the original character image (hereinafter sometimes referred to as the "input image") has a high head-to-body ratio, and the image obtained by deforming the original character image (hereinafter sometimes referred to as the "output image") has a low head-to-body ratio. For example, changing a 4-headed character to a 2-headed character is also called deformation. Generally, short-headed characters are also called chibi characters. "Feature Decoupling" is a method of independently handling different feature quantities (e.g., feature quantities related to poses, clothes, expressions, etc.) in a machine learning model, especially an image generation model. Feature quantities are also referred to as variables, etc. "User" refers to a person who uses the deformation character generation system 1. The user operates the terminal 20, regardless of whether it is an individual or a corporation. In the deformation character generation system 1, the user inputs an image of a tall-headed character and obtains the converted short-headed character as the output. "Obtain" includes not only the processor obtaining data, programs, etc. from the input unit or an external terminal, but also the processor creating, updating, etc. data, programs, etc. Please refer to the character model, etc. in the following embodiments.
[0028] The terms used in the drawings, etc. are explained below. However, the following is for helping the understanding of the invention and does not limit the content of this specification. "Stable Diffusion" is a machine learning model that converts text and the like into images and is one of the so-called image generation AIs. Stable Diffusion is characterized by using a diffusion model (especially, a latent diffusion model). Further, Stable Diffusion has a text encoder for inputting text for conditioning. The text encoder converts text into a vector and is, for example, a Transformer (a machine learning model) that has been trained by a method such as CLIP (learning combinations of a large number of images and text and calculating the similarity between images and text). The processing in Stable Diffusion includes processing for obtaining latent variables from image information by a VAE (described in the next section) and processing for obtaining image information from latent variables by a VAE. It is characteristic that the amount of calculation is suppressed by using a VAE. The specific embodiment of the diffusion model will be described in the following embodiments. "VAE (Variational Auto-Encoder)" refers to a variational auto-encoder. A VAE is a generative model having an encoder-decoder structure. However, the term VAE may also refer to a neural network having such a structure, such a network structure, a machine learning model, or a technology including these. A VAE learns to represent input data by an average and a standard deviation (or variance, etc.) (average vector, variance vector, etc.). For example, an encoder converts input data into a latent variable (a single point having a statistical distribution), and a decoder generates new data by restoring the latent variable (a single point randomly sampled from a statistical distribution). Further, a device (reparametrization trick) for error backpropagation is made during learning. "Unet" is a neural network based on CNN (Convolutional Neural Network). For example, it includes a module containing a convolutional layer and a module containing an attention layer. Generally, it is used for image segmentation, but here it is used for the reverse diffusion process of a diffusion model, such as noise removal in the latent space. Unet, for example, takes noise as input and outputs the noise to be removed. By removing the noise to be removed from the input noise, an image with the noise removed is obtained. By repeating this, the noise is gradually removed from the image. "ControlNet" is an extension of Stable Diffusion that enables the generation of (high-quality) images based on the poses and compositions specified by the user. ControlNet is an additional neural network that uses the structural features (poses, line drawings, edges, etc.) of an image as a condition (Condition) and generates an image reflecting it. While retaining the neural network blocks of Unet, additional conditions given by the user are added by the neural network on ControlNet and reflected in Unet. That is, a condition (such as composition data) etc. is input, and the blocks on Unet are the output destinations. Details are omitted, but techniques such as zero convolution and skip connections are incorporated. "LoRA (Low-Rank Adaptation)" is a technique related to the fine-tuning of a (pre-)trained (machine learning) model. It should be noted that it is different from LoRa (registered trademark), which is a low-power long-distance wireless communication technology.
[0029] In the following, when it is described as "○○" processing, it means that the computer's processor executes the processing based on the "○○" program stored in the program storage section. In this paragraph, the same word is entered in the "○○" part. That is, the "XX" program is a program that causes a computer to function as "XX" means by executing the "XX" process. At this time, the control unit including the processor also means that it functions as an "XX" unit (or an "XX" device). In this case, the "XX" unit means executing the "XX" process based on the "XX" program. Also, when indicating the processing procedure (of the "XX" process), it is described as "XX" steps.
[0030] For example, the deformed character generation program P1 is a program that causes a computer to function as deformed character generation means by executing the deformed character generation process. At this time, the control unit 12 of the computer including the processor 122 functions as a deformed character generation unit (or a deformed character generation device).
[0031] In the deformed character generation system 1, each terminal (computer) such as the terminal 20 includes a processor. However, when simply referring to the processor, it refers to the processor that performs processing by the deformed character generation program P1, which is the processor 122 of the server 10 in this embodiment.
[0032] In the following, for simplicity, "the processor 122 of the server 10 receives a request from the terminal and returns data for display on the browser of the terminal" may be described as "the processor 122 displays (causes to display) on the browser of the terminal" or "the processor 122 displays (causes to display)". Similarly, "the processor 122 of the server 10 causes the data storage unit 14b of the storage unit 14 to store data" may be described as "the processor stores (causes to store) (data)".
[0033] (First Embodiment) 1. User Interface (UI) First, an interface that the deformation character generation system 1 of the present embodiment causes the terminal 20 to display will be described with reference to the drawings. The interface described hereinafter is a simplified version of what the processor 122 causes to be displayed on the browser of the terminal 20.
[0034] Also, only icons and the like related to functions necessary for the description are displayed, and other known icons and the like are omitted. For example, a back button for returning to the page displayed immediately before is omitted.
[0035] FIG. 2 is a diagram showing a basic setting screen. The image upload column UI-11 on the left side of the basic setting screen is a UI for the user to upload the original character image which is the input image, and is also a column for displaying the uploaded original character image. The user uploads the original character image by drag and drop. Note that the method of upload is not limited to this. A folder in which the image is stored may be selected so that the image file can be selected (not shown).
[0036] The deformed character setting column UI-12 is a UI for the user to set the deformed character image to be output. The deformed character setting column UI-12 displays a body ratio setting button UI-121, a reference deformed character image UI-122, and a deformed character generation start button UI-123. Details will be described later.
[0037] FIG. 3 is a diagram showing an original character image confirmation screen. When the user uploads the original character image to the image upload column UI-11, the processor 122 acquires and creates a character prompt (character information) and a character model (machine learning model) described later from the image, and displays an original character image confirmation screen (window) UI-13. In the figure, the term "training" is used for the creation of the character model.
[0038] In this embodiment, the processor 122 displays a confirmation screen for whether to perform training (learning) after uploading, and the user starts learning by parameter update or the like by pressing the OK button icon or the like displayed there. However, this is omitted here.
[0039] As shown in FIG. 3, the processor 122 displays image information UI-131 (file name, user ID, creator, creation date and time, etc.), the original character image UI-132 uploaded by the user, the OK button icon UI-133, the retraining icon UI-134, and the delete button icon UI-135 on the original character image confirmation screen (window) UI-13.
[0040] When the user presses the OK button icon UI-133, the processor 122 displays the original character image in the image upload field UI-11 of the basic settings screen (not shown). When the user selects (presses) the original character image on the basic settings screen, the processor 122 displays the original character image confirmation screen (window) UI-13 again.
[0041] When the user presses the retraining icon UI-134 on the original character image confirmation screen (window) UI-13, the processor 122 displays a screen for selecting the original character image (not shown). When the user selects the original character image there and presses the button to execute retraining, the processor 122 creates a character prompt and a character model (again) from the original character image (input image) selected by the user. At this time, the parameters are adjusted arbitrarily or automatically by the user. In addition, the training method can be adjusted by the user (not shown).
[0042] When the user presses the delete button icon UI-135, the processor 122 deletes the uploaded original character image. The user can upload a new original character image.
[0043] FIG. 4 is an enlarged view showing the deformed character setting field (UI-12).
[0044] In this embodiment, a large number of reference deformed character images (images of the referenced deformed characters) are stored in advance in the storage unit 14 (data storage unit 14b), and the user can freely select the orientation, pose, etc. of the deformed characters.
[0045] Also, the user can upload their favorite deformed character. For example, in this embodiment, by dragging and dropping an image into the deformed character setting field UI-12, a favorite character can be used as a reference image.
[0046] The body proportion setting button UI-121 is a button icon for setting the body proportion of the output image. For example, when the 1:4 button is selected, the processor 122 causes the deformed character setting field UI-12 to output an image with a head height: full body height ratio of 1:4 (see FIG. 2. Also refer to the shaded part of the reference deformed character image UI-122 in FIGS. 2 and 4).
[0047] As shown in FIG. 4, the All button of the body proportion setting button UI-121 displays all the images with the shown (1:2.5 to 1:5 in FIG. 4) ratios.
[0048] When the user selects and presses the reference deformed character image UI-122, the processor 122 displays a check (dot) to indicate that the image has been selected. The user can select a plurality of reference deformed character images, and in that case, the processor 122 generates a plurality of deformed characters according to the selection. That is, the processor 122 can output a plurality of deformed character images at once.
[0049] The reference deformed character image UI-122 displays the pose and the standard of the head-to-body ratio of the deformed character to be output. This image is automatically selected by the processor 122 from among a number of stocks.
[0050] On the other hand, the processor 122 displays a pose display window (not shown) for selecting a pose (such as a standing pose or a sitting pose) according to the user's selection, and the user can also select a preferred pose from the group of poses listed in the pose display window.
[0051] That is, the user can easily input the head-to-body ratio by pressing the head-to-body ratio setting button UI-121 or by uploading a reference image. In other words, the user can input the head-to-body ratio by a simple operation using the GUI.
[0052] When the user presses the deformed character generation start button UI-123, the processor 122 starts processing related to deformed character generation. The processing will be described later.
[0053] Figure 5 is a diagram showing a generated image display screen. When the processor 122 acquires an image (output image) of the deformed character, it displays the output image UI-141 in the output image display column UI-14 of the generated image display screen. In addition to this, the processor also displays an output image processing related button UI-142, an output image information UI-143, and an output image saving related button UI-144.
[0054] The output image processing related button UI-142 is a button icon for further processing the output image or for regenerating the image.
[0055] The magnification button shown in Fig. 5 is a button icon for changing the displayed magnification ratio. The color button is a button icon for performing color processing such as the skin color or hair color of a character. The part separation button is a button icon for separating and displaying the parts (e.g., hairstyle, clothing, etc.) of a deformed character from other parts. The user can perform processing such as separately and independently on that part. The regeneration button is a button icon for regenerating a deformed character image. Different parameters are set in the machine learning model (random latent variables are sampled in the diffusion model), and deformed character images (usually different deformed character images) corresponding to the parameters can be generated. That is, the user has the advantage of being able to easily obtain different images with just one button.
[0056] The output image information UI - 143 is a field for displaying information about the image. The style is the file name of the reference image. The variation intensity is a parameter for adding variations. The greater the variation intensity, the greater the change from the original character image, and the output image with a greater change is output. In this embodiment, the variation intensity can be set by the user (the UI is not shown).
[0057] The output image saving - related button UI - 144 is an operation button icon related to saving the output image. DL is for downloading the output image to the terminal (terminal 20), deletion is for deleting the output image, and the save button is for saving the output image on the server (server 10) (such as on the server).
[0058] With the above configuration, the user can generate a deformed character image with a simple operation using the GUI (Graphical User Interface). In particular, since the user can select the head-to-body ratio while viewing the reference deformed character image UI-122 as a reference example, the output image is easy to predict, and the desired deformed character image can be obtained more reliably. Also, at this time, there is a remarkable effect that the process can be executed by the GUI without performing inputs such as finely adjusting the prompt.
[0059] 2. Program Processing <2-1. Deformed Character Generation Process> The program processing performed in the deformed character generation system 1 of the present embodiment will be described.
[0060] In the present embodiment, the processor 122 performs a deformed character generation process based on the deformed character generation program P1. The deformed character generation program P1 includes at least (1) a source character image analysis program P11, (2) a reference image analysis program P12, (3) a deformed character acquisition program P13, and (4) an optimization program P14. The processor 122 executes a source character image analysis process, a reference image analysis process, a deformed character acquisition process, and an optimization process based on each of these programs. In addition to the above, the deformed character generation process includes an online UI providing process for providing the above-described UI to the user online. Since the UI to be provided has been described, it is omitted here.
[0061] (1) The source character image analysis process is a process of acquiring and analyzing a source character image, including a source character image acquisition process of acquiring a source character image that is a character image, and a character prompt acquisition process of acquiring character information for explaining the source character image. (2) The reference image analysis process is a process of acquiring and analyzing a reference image of a deformed character, and includes a head-to-body ratio acquisition process of acquiring a setting of the head-to-body ratio by user input. (3) The deformed character acquisition process includes: A deformed character acquisition process of acquiring a deformed character image in which the original character image is deformed according to the body ratio by using the original character image, the character information, and the body ratio as inputs to a machine learning model. Here, the machine learning model learns at least a character image and a deformed character image as learning data, and is characterized in that, with a character image as an input, it outputs a deformed character image that deforms the character image.
[0062] More specifically, in this embodiment, the machine learning model includes a diffusion model and a control model for controlling the diffusion model. The diffusion model is a machine learning model for outputting a character image having the characteristics of the character information by using, as an input, the character information (prompt) for explaining the character image, etc. The control model is a machine learning model for controlling the diffusion model so as to output a deformed character image that is a deformed character image.
[0063] Furthermore, the diffusion model learns at least a character image, noise added to information (data) (latent variable) based on the character image, and information obtained by adding the noise as learning data. Here, the information (data) based on the character image is, for example, a plurality of images acquired from the character image that have been latent variable-transformed by VAE.
[0064] In the inference stage, noise-containing information is input, and noise to be removed from the noise-containing information is output (usually this process is repeated). As a result, information from which noise has been removed from the noise-containing information is output. As a result, with a character image and character information for explaining the character image as inputs, a deformed character image having the characteristics of the character information can be output.
[0065] Here, the height-to-width ratio based on the user's input is related to either the control model or the diffusion model to control the height-to-width ratio of the deformed character image.
[0066] That is, the height-to-width ratio may constitute the control model, or it may become the character information (prompt) describing the character image and be input to the diffusion model. Here, the former will be described. The latter form will be described later.
[0067] The above-described machine learning model is a machine learning model that learns at least a character image, character information describing the character image, a deformed character image obtained by deforming the character image, and the height-to-width ratio of the character image and the deformed character image as learning data, It can be said that it is a learned machine learning model for generating a deformed character image that takes the original character image, which is an image of a character, the character information describing the original character image, and the user's setting of the height-to-width ratio as inputs and outputs a deformed character image obtained by deforming the original character image according to the height-to-width ratio.
[0068] In the present embodiment, the processes (1) to (3) above can be further described as follows. (1) The original character image analysis process is a process of acquiring and analyzing the original character image, The original character image acquisition process of acquiring the original character image, which is an image of a character, The plurality of feature quantity acquisition process of acquiring a plurality of feature quantities by acquiring the image information of a plurality of regions including at least the face portion from one said original character image, The tag management process of attaching tags to and managing the information associated with the plurality of regions (images), A feature decoupling process that separates the plurality of feature quantities into pose feature quantities and other feature quantities (Feature Decoupling). A character prompt acquisition process for acquiring character information that explains the feature quantities of concepts including feature quantities other than pose feature quantities among the plurality of feature quantities, and A character model acquisition process for acquiring a machine learning model (diffusion model) that is learning (learning) data including at least concept feature quantities as learning data. It includes these processes.
[0069] (2) The reference image analysis process is a process of acquiring and analyzing a reference image of a deformed character, A body ratio acquisition process for acquiring a setting of the body ratio by the user's input (in this embodiment, referred to as "reference image acquisition process for acquiring a reference image of a deformed character"), and A control model acquisition process for acquiring a control model including at least line drawing information indicating a part of the shape of the deformed character from the reference image. It includes these processes. That is, in this embodiment, the processor 122 acquires the body ratio in the form of a reference image and acquires a control model from the reference image.
[0070] (3) The deformed character acquisition process is A process of inputting the character information (character prompt) to the machine learning model (diffusion model) controlled by the control model and obtaining, as an output, a deformed character image obtained by deforming the original character image.
[0071] Roughly speaking, it is a process of acquiring and analyzing an image (original character image, reference image) (the above (1) and (2)), and a process of obtaining a deformed image by image generation (the above (3)). In other words, the process of obtaining a deformed image by image generation (the above (3)) is a process of actually generating a deformed character from data, models, etc. obtained from (the above (1) and (2)).
[0072] FIG. 6 is a flowchart showing the deformed character generation process. Upon receiving an instruction to start the process of generating a deformed character (e.g., pressing the deformed character generation start button UI-123 in the deformed character setting field UI-12 shown in FIG. 4) input by the user, the processor 122 starts the deformed character generation process.
[0073] First, (1) the original character image analysis process will be described. The processor 122 acquires the original character image to be deformed, which is input by the user (step 1 - original character image acquisition process).
[0074] The processor 122 extracts and acquires a plurality of feature amounts from the acquired image (step 2 - plurality of feature amount acquisition process). In the present embodiment, the processor 122 recognizes a plurality of regions such as regions related to the head, upper body, or whole body from one original character image, acquires the image information thereof, and acquires a plurality of feature amounts.
[0075] FIG. 7 is a diagram showing an image for recognizing a plurality of regions. In FIG. 7, a plurality of rectangular regions with different scales (the portions surrounded by the quadrilateral lines - bounding boxes) indicate the plurality of regions.
[0076] In the present embodiment, the processor 122 acquires (crops) an image including the face at different scales. In particular, in the present embodiment, the face recognition (Face Detection) process for recognizing the face is performed. As the face recognition method, a known method is appropriately used. Note that the process of acquiring (cutting out) an image is generally called "crop" or "cropping".
[0077] In the case of the example shown in FIG. 7, the plurality of regions, in order from the smallest in area of the rectangular regions in FIG. 7, include the face, head, upper body, from the top of the head to the waist, above the knees (including the hat, the same applies hereinafter), above the calves, above the ankles, above the heels, and above the toes.
[0078] Obtaining (cropping) images of a plurality of regions with different scales clarifies the site (position) for which feature amounts are to be obtained and enables obtaining a plurality of feature amounts according to the site (position). For example, it is possible to obtain data regarding the face and other sites. Also, there is an advantage of increasing the amount of data from a small number (one) of images.
[0079] As a method of recognizing a plurality of regions from a single original character image and obtaining an image, known object detection techniques can be used.
[0080] Among object detection techniques, those that identify the class of an object after specifying the approximate position of the object are called two-stage models, and examples include R-CNN (Region Based Convolutional Neural Networks) and FPN (Feature Pyramid Networks). Also, those that simultaneously perform the specification of the position of an object and the identification of the class are called one-stage models, and examples include YOLO (You Only Look Once) and SSD (Single Shot multibox Detector). In addition, a cascade classifier (Open CV) or the like may be used.
[0081] The "plurality of regions" indicates a recognition unit (cropping unit) in the original character image. The processor 122 obtains one cropped image for one region. The "plurality of feature amounts" are a plurality of feature amounts obtained from the original character image. For example, they are latent variables obtained by using the cropped image as the input of the VAE.
[0082] The processor 122 attaches tags to each region (image) and / or feature amount for management (tag management process). In the present embodiment, the image obtained by cropping and the tag correspond one-to-one. In the present embodiment, the tags relate to the head, clothing, whole body, face, upper body, hair, decoration, and pose (see also FIG. 11 regarding this area).
[0083] In the present embodiment, the processor 122 manages by storing the image (file) and the text (file) including the tag in a set in the storage unit. Also, the processor 122 uses a part of this data (text) for a character prompt described later.
[0084] Examples of this text are given for each tag below. Each tag includes character information explaining the region (image) and / or feature amount. Incidentally, the character information explaining this region (image) and / or feature amount may be referred to as a prompt element or tag information for convenience. · Head: hat with feathers, pirate hat, navy hat · Cloth: miniskirt, uniform, knee boots, jacket · Full body: full body, standing, lying, sitting · Face: smile, closed mouth, bearded, blushing · Upper body: upper body · Hair: long hair, short hair, single braid, twin ponytails · Decoration: necklace, earrings, bandage, jewelry · Pose: hand on hip, leg lift, knee up
[0085] In this example, tags and text are words or phrases. However, this is not limited to this, and it may be a sentence or the like.
[0086] Generally, the technology that receives input of image data and outputs text that summarizes (explains) its content is called Image Captioning. CLIP, Flamingo, etc. are known as (multimodal) models for executing Image Captioning.
[0087] Such technology is called Image to Prompt and is a technology for generating text from an image, which is different from OCR (Optical Character Recognition) that acquires character information in the image.
[0088] Here, the method of obtaining text from an image is described, but a method of converting feature amounts based on an image into text may also be used.
[0089] Also, the processor 122 acquires feature amounts corresponding to the tags. For example, feature amounts may be acquired for each tag (that is, for each cropped image), or a plurality of acquired feature amounts may be classified according to the tags.
[0090] Note that a plurality of feature amounts may be obtained from one region (one image obtained by cropping), or one feature amount may be obtained from a plurality of regions. For example, the processor 122 may acquire feature amounts related to the head and feature amounts related to the hair from an image of the head.
[0091] Also, as shown in FIG. 7, since the image of the portion above the fingertip includes the entire body of the character, it includes the feature amount of the pose. However, the feature amount of the pose is not something that cannot be obtained without an image of the entire body. For example, an image including the upper body may include the feature amount of the pose that the character has their hand on their hip.
[0092] Also, just because there is an image of the face region, for example, it does not mean that the processor 122 cannot obtain the feature amount related to the face part from the image of the upper body region. That is, there is no limitation on what information the processor 122 extracts from a certain region (image).
[0093] In the present embodiment, it is essential that the processor 122 obtains the image of the face region of the original character image and obtains the feature amount of the face. For example, the feature amount of the face determines the feature amount of the deformed character so that the face part is particularly prominently displayed in the deformed character image.
[0094] Returning to FIG. 6, the processor 122 separates the feature amount of the pose and the feature amount of the concept including the feature amount other than the feature amount of the pose among the plurality of feature amounts obtained in step 2 (step 3 · feature amount separation process). This separation is referred to as feature decoupling.
[0095] As a method for realizing feature decoupling, designing the latent space of the generative model and training it to control different feature amounts in different dimensions can be mentioned. In particular, in the present embodiment, the feature amount of the pose is separated from other feature amounts. Due to feature decoupling, each feature amount is processed without depending on other feature amounts, so the flexibility and accuracy of the model are improved.
[0096] In this embodiment, features and elements other than poses are expressed as concepts. For example, feature amounts other than the feature amounts of poses are used as the feature amounts of concepts. In the above example, among the plurality of feature amounts, the feature amounts related to the head, clothing, whole body, decoration, face, upper body, and hair become the feature amounts of the concept. That is, in this embodiment, a plurality of feature amounts = feature amounts of poses + feature amounts of concepts.
[0097] The processor 122 acquires character information that describes the original character image from the original character image (step 4. Character prompt acquisition process). For the sake of convenience, the character information (prompt) that describes this original character image is referred to as a "character prompt".
[0098] In particular, the processor 122 acquires character information that describes elements (images and / or feature amounts) other than poses. In this embodiment, the processor 122 acquires character information that describes parts of the image other than poses, that is, the head, clothing, whole body, face, upper body, hair, or decoration, etc. (concepts).
[0099] An example of a character prompt is as follows. 「1girl,blue eyes,purple hair,long hair,police uniform,transparent background, open clothes,black skirt,mini skirt,high-heeled shoes, simple background,…(omitted)」
[0100] By using the acquired character prompt as an input for the deformed character acquisition process described later, there is an advantage that the detailed features of the character can be controlled and reflected in the deformed character.
[0101] On the other hand, the processor 122 acquires (step 5, character model acquisition process) a machine learning model (diffusion model) that learns (or is learning) at least a part of the feature amounts other than the pose feature amount among the plurality of feature amounts described above (the feature amounts of the concept including the feature amounts other than the pose feature amount) as learning data. For convenience, this machine learning model (diffusion model) is referred to as a "character model".
[0102] The training structure of this character model includes a diffusion model and a network (ControlNet) that controls the diffusion model. The character model is obtained by freezing the weights of ControlNet and updating (learning) only the parameters of the Diffusion model. There may be an expression of "(the processor 122) creates a character model each time".
[0103] Here, the general content of the "diffusion model" will be described. The diffusion model is a machine learning model that is the core of the technology related to Stable Diffusion also used in this embodiment. The diffusion model is a model capable of removing noise from a noisy image or the like to generate a noise-free image or the like. In the learning stage, the diffusion model includes a diffusion process of adding noise to the original information (pixels, latent space, etc.) and an inverse diffusion process of removing noise from the noisy information.
[0104] The diffusion model uses, as learning data, the original information such as an image, the noise to be added, the noisy information obtained by adding noise to the original information (for example, a noisy image), the noise to be removed from the noisy information, and the original information obtained by removing noise from the noisy information.
[0105] Then, the Unet included in the diffusion model learns (in the learning stage) the noise to be removed from the noisy information with the above noisy information or the like as the input. The diffusion model (Unet included) takes noisy information (such as a noisy image) as input and outputs the noise to be removed from the noisy information. As a result of repeating this process, information from which noise has been removed (such as a noise-free image) can be output from the noisy information (such as a noisy image) (inference stage).
[0106] In this embodiment, the parameters of the machine learning model (diffusion model) are updated based on the feature amount of the concept obtained by performing feature separation.
[0107] Note that what is obtained by converting the pose into a latent variable by VAE is used for maintaining the layout (the "Freezed" part in ControlNet · Figure 1).
[0108] For example, in this embodiment, the feature amount of the concept obtained by performing feature separation serves as the original information. That is, the processor 122 creates and acquires a machine learning model (diffusion model) through learning that goes through a diffusion process of adding noise to the feature amount of the concept (latent variable) and an inverse diffusion process of removing the noise.
[0109] If a deformed character image can be obtained, it may be a mode (mode 1) in which a plurality of images are not acquired from the original character image, or a method (mode 2) in which decoupling is not performed between the feature amount of the pose and other feature amounts after acquiring a plurality of images. However, according to the method of this embodiment, there is a remarkable effect that a deformed character image that accurately reflects the features of the original character image can be obtained.
[0110] Briefly, the processor 122 acquires a plurality of feature amounts from a single original character image and separates (decouples) them into the feature amount of the pose and other feature amounts (concepts). Also, character information (character prompt) for explaining this concept is acquired. Furthermore, a machine learning model (diffusion model (character model)) that learns based on the feature amount of the concept is created.
[0111] The character model of this embodiment is trained using only one image of a person (character). The trained model can be used to generate the characteristics of this person (character) with different poses and styles (for example, SD (super deformed) style, etc.).
[0112] The specific training method is summarized below. The processor uses images containing faces that are cropped at different scales. The processor extracts prompts (tag information) from these images of different scales. The training structure of the machine learning model includes a Diffusion model and ControlNet. The weights of ControlNet are frozen, and the processor updates only the parameters of the Diffusion model. The Diffusion model after parameter update is the character model of this embodiment.
[0113] Note that it may be incorporated or improved and used based on an existing machine learning model (diffusion model). For example, an already learned machine learning model (diffusion model) may be used, such as by adding noise to a character image (including a deformed character image).
[0114] Next, (2) the reference image analysis process will be described. First, the processor 122 obtains the setting of the body ratio according to the user's input (body ratio acquisition process). In this embodiment, the processor 122 obtains the setting of the body ratio by obtaining a reference image of a deformed character (step 6, reference image acquisition process). Another example will be described in the second embodiment described later.
[0115] In this embodiment, the processor 122 first deletes the concept information of the reference image as preprocessing, and acquires scribble information (scribble information called Scribble) indicating a part of the shape of the deformed character, and pose information (Pose).
[0116] As shown in the lower left of FIG. 1, the processor 122 acquires scribbles of at least the head, the tip of the hand, and the tip of the foot of the reference image as scribble information. Thereby, the pose (orientation, etc.) and the head-to-body ratio of the deformed character to be output can be controlled (described later).
[0117] In addition, the processor 122 acquires the joint position information of the reference image as pose information.
[0118] In the lower left of FIG. 1, the upper part of the image with black and white reversed is an image diagram of scribble information (Scribble), and the lower part is pose information (Pose).
[0119] In this embodiment, the processor 122 creates a control (machine learning) model from the acquired scribble information (Scribble) and pose information (Pose). This control (machine learning) model is referred to as a "control model" for convenience.
[0120] Note that it is denoted as a control model in order to distinguish it from the above-described machine learning model (character model). Therefore, for example, the character model may be referred to as a first machine learning model, the control model as a second machine learning model, and so on.
[0121] This control model is an additional neural network that controls the diffusion model. In particular, it has an excellent effect that the character of the input image can be controlled to an intended pose or the like. The control model of this embodiment controls the pose and / or the outer shape (affecting the body proportion) of the deformed character image.
[0122] The above-described processing can be executed by plugins or extensions of Stable Diffusion. For example, "ControlNet Extension" that can be added to the Stable Diffusion Web UI · Automatic1111 version of the Web UI can be preferably used.
[0123] In this embodiment, ControlNet Scribble (one of the models of ControlNet) is used to acquire line drawing information. Also, Open Pose (one of the models of ControlNet) is used to acquire pose information.
[0124] Note that Open Pose is a pose estimation method. By applying Open Pose to an image of a character, the pose of that character can be estimated. Open Pose introduces a process that takes into account the positional relationship between skeletons called Parts Affinity Fields.
[0125] Various "preprocessors" of ControlNet analyze and transform the input image (e.g., pose detection, edge detection, line drawing, etc.) and use it as a condition for the model. That is, as a function of ControlNet, there may be a mechanism to convert an image into a line drawing (edge detection / scribble detection), and the user can utilize it to generate "Scribble (line drawing)".
[0126] To summarize, in this embodiment, the processor 122 acquires the head-to-body ratio in the form of a reference image, and acquires a control model from the reference image. The control model is acquired by acquiring at least line drawing information indicating a part of the shape of the deformed character from the reference image (step 7 · control model acquisition process).
[0127] More specifically, a control (machine learning) model (which is an additional neural network for controlling the above-described diffusion model) is created and obtained by using, as learning data, line drawing information and pose information indicating (a part of the shape of) the deformed character in the reference image (more precisely, those obtained by converting these into latent variables by a VAE). That is, this line drawing information and pose information serve as inputs (conditions) to the control model.
[0128] In this embodiment, this control model is created each time. However, it may also be possible to obtain a control model that has been created in advance for each reference image and stored in the storage unit 14 of the server 10. In this case, there is an advantage that the processing speed is increased. However, when the user uploads a new image as a reference image, the processor 122 creates a control model for that new image.
[0129] Returning to FIG. 6, finally, (3) the deformed character acquisition process will be described. As described above, the processor 122 uses the original character image, the character information, and the height-to-width ratio as inputs to the machine learning model, and obtains a deformed character image in which the original character image is deformed to the height-to-width ratio (step 8, deformed character acquisition process).
[0130] More specifically, by inputting character information (= character prompt) to a diffusion model (= character model) controlled by the above-described control model (= control model), a deformed character image in which the original character image is deformed is generated and obtained as an output.
[0131] There is a remarkable effect that a deformed character image with significantly higher accuracy can be obtained compared to, for example, a novice user or the like inputting conditions by a prompt.
[0132] Then, the processor 122 causes the deformed character image being acquired to be displayed on the display unit 284a of the terminal 20. When a plurality of head-to-body ratios are selected by the user, output deformed character images corresponding to the respective head-to-body ratios. Also, as shown in the UI section, the processor 122 can accept an input by the user and regenerate a deformed character image or the like.
[0133] Note that as long as the effects of the deformed character generation process in the present embodiment can be exhibited, the order of the processes is not limited to that in FIG. 6. For example, in FIG. 6, steps 4, 5, and steps 6 to 7 are drawn sequentially, but these may be processed in parallel. Or steps 4, steps 2·3·5, and steps 6 to 7 may be processed in parallel.
[0134] <2-2. Optimization Process> In the optimization process, the processor 122 obtains and evaluates the differences between a plurality of output images obtained by changing weights in at least some of the layers of the neural network (Unet) included in the above-described diffusion model (character model), and reduces or deletes the weights of the layers that have little influence on the output images.
[0135] The processor 122 performs the optimization process based on the optimization program P14. That is, the optimization program P14 causes the computer to function as an optimization means by execution of the optimization process by the processor 122.
[0136] FIG. 8 is a flowchart showing the optimization process. The processor 122 starts the optimization process by receiving an instruction from the user to regenerate a deformed image or automatically based on a condition set in advance (see FIG. 5, generated image display screen). In the present embodiment, the user can regenerate an image in which at least a part of the image is changed.
[0137] First, the processor 122 obtains the difference between the deformed character images when weights are set to 1 and 0 respectively in each of the plurality of layers in the neural network (of the diffusion model) (step 21). While retaining the weights of the layers with a large influence, the processor 122 reduces or deletes (hereinafter referred to as "reduction, etc.") the weights of the layers with a large influence (step 22).
[0138] FIG. 9 is a diagram for explaining the reduction, etc. of weights. In FIG. 9, the part surrounded by the upper right rounded rectangle indicates that it is the same layer of the neural network. Also, the two characters surrounded by the same rounded rectangle are the deformed character images when the weights are set to 0 and 1 respectively in the said layer. Diff in FIG. 9 indicates the difference between the deformed character image when the weight is 0 and the deformed character image when the weight is 1. The blank part is the difference.
[0139] If the difference between the respective images obtained by setting the weights to 1 and 0 in a certain layer is not very large, it means that that layer does not have a great influence on the overall result. That is, by evaluating this difference in each layer, the degree of influence of each layer on the whole can be grasped. And the processor 122 deletes, etc. the weights of the layers with little influence in response to this result.
[0140] For example, in the rightmost layer (the layer represented by the four rounded rectangles) of FIG. 9, the difference between the respective images generated by the difference in weights (weights 1 and 0) is large (the difference is large). From this, it can be seen that the change in the image generated by the change in the weights of this layer is large, that is, the weights of this layer have a great influence on image generation. On the other hand, since the difference in the second layer from the left is small, it is considered that the influence of this layer on the generated image is small.
[0141] This setting optimizes the parameter layers in the fine-tuning method (LoRA) of the pre-trained model and finally produces high-quality results.
[0142] Return to the flowchart of FIG. 8. Arrange the layers with a large impact of this, that is, the layers related to the key features, in the output layer of the neural network (of Unet), and update the weights of the neural network layers by backpropagation (fine-tuning) (step 23). Through fine-tuning, a deformed character image that appropriately reproduces the features of the original character image can be obtained.
[0143] An example of the optimal image obtained by the optimization process is shown. FIG. 10 is a diagram showing deformed character images before and after the optimization process. As shown in FIG. 10, the image after optimization reproduces the original character image more faithfully than the image before optimization (especially the part surrounded by the square around the waist).
[0144] Here, the specific method used in this embodiment will be described. For example, LoRA Block Weight can be used as a method for performing the above processing. With LoRa Block Weight, the application degree of LoRA can be adjusted.
[0145] Note that LoRA (Low-Rank Adaptation) is a method for fine-tuning a pre-trained model. Among them, those that perform tuning on a concept, those that reflect the concept of the image to be generated, or those that apply a specific concept to a character are called Concept LoRA. By using Concept LoRA, the above adjustments can be made.
[0146] To summarize, through the original character image analysis process, a plurality of feature amounts are obtained from a plurality of regions of one original character image, so that the features of the original character image can be more reflected in the deformed character. Here, a character prompt and a character model are created (acquired). Also, through the reference image analysis process, the processor 122 obtains the height-to-width ratio in the form of a reference image, and through processes such as extracting a diagram and a pose from the reference image, obtains a control model for controlling the character model. Then, in the deformed character acquisition process, by inputting a character prompt to the character model controlled by the control model, a deformed character faithful to the original character image can be generated. Its high accuracy is as shown in the drawings. And through the optimization process, a higher-quality deformed image can be obtained.
[0147] Also, the above-described diffusion model (character model) is a learned model that obtains a deformed character image obtained by deforming the original character image from the original character image that is an image of a character, obtains a plurality of feature amounts from the original character image, uses, as learning data, data including at least a feature amount of a concept including feature amounts other than the pose feature amount among the plurality of feature amounts, noise added to the feature amount of the concept, and noisy information obtained by adding noise, and learns the noise to be removed from the noisy information, characterized in that, with the noisy information as an input, it outputs the noise to be removed from the noisy information, Furthermore, it is a learned model characterized in that the pose in the output deformed character image is controlled by a control model (= control model) that learns, as at least one learning data, line drawing information indicating a part (shape) of the deformed character in the reference image.
[0148] FIG. 11 is a diagram showing the above-described deformed character generation process (in a form different from FIG. 1). In addition to obtaining a deformed character image from a character image (original character image), a process of optimizing the deformed character image is shown.
[0149] 3. Hardware Configuration FIG. 12 is a network configuration diagram showing an overview of the deformed character generation system 1. As shown in FIG. 12, the deformed character generation system 1 in the present embodiment includes a system server 10 (server 10) and a (user) terminal 20. Further, these devices are connected via a network N. The network N is, for example, the Internet or the like. The server 10 includes a deformed character generation program P1, and software (application software) for operating the deformed character generation system 1 according to the present embodiment is installed. Various processes are executed by the functions of the software. Hereinafter, each hardware will be described.
[0150] <Server 10> The server 10 is a computer for executing the deformed character generation program P1. Further, the server 10 displays a website W that provides a user interface for the user.
[0151] In FIG. 12, only one server 10 is illustrated, but the number is not limited to one, and it may be realized by a plurality of servers. For example, separating a web server and a machine learning server may be mentioned. A plurality of servers may also be used from the viewpoints of load distribution and availability. Also, the process related to image generation may be entrusted to another server or may be performed in cooperation with another server. In addition, the server 10 may use a computer of a cloud service provider, or the user may prepare a computer.
[0152] Figure 13 is a hardware configuration diagram of the server 10. As shown in Figure 13, the server 10 includes a control unit 12, a storage unit 14, and a communication control unit 16. The control unit 12 further includes a processor 122, a ROM 124, a RAM 126, and a timer unit 128. The basic functions of each will be described together later.
[0153] The control unit 12 including the processor 122 also functions as a deformed character generation unit in the server 10 (not shown). The deformed character generation unit executes a deformed character generation program P1 to perform deformed character generation processing. In this embodiment, the processor 122 is a CPU (Central Processing Unit).
[0154] Also, one program may include another program. For example, in this embodiment, the deformed character generation program P1 includes an original character image analysis program P11, a reference image analysis program P12, and the like.
[0155] As shown in Figure 13, the storage unit 14 includes a program storage unit 14a and a data storage unit 14b, and stores programs and data necessary for various processes. For example, in the program storage unit 14a, in addition to the deformed character generation program P1 according to this embodiment, there are stored a control program for controlling devices connected to the server 10, such as a communication control program for controlling the communication control unit 16.
[0156] As shown in Figure 13, the communication control unit 16 is a device for performing communication between the server 10 and an external terminal, for example, the terminal 20 described later. As shown in Figure 12, the communication control unit 16 connects the server 10 to the network N.
[0157] In addition to the above, the server 10 may include an input unit (e.g., a keyboard) for inputting commands and data, an output unit (e.g., a voice output device) for outputting information in some form, etc. (not shown). Also, it may include devices additionally required for the purposes of this embodiment, or devices for improving convenience for the purposes of this embodiment.
[0158] <Terminal 20> The terminal 20 is an information processing device for a user to use the deformation character generation system 1. The user uses the deformation character generation system 1 by accessing the server 10 using the terminal 20. The terminal 20 includes a control unit 22, a storage unit 24, a communication control unit 26, and an input / output unit 28. The control unit also includes a processor, a ROM, a RAM, and a timing unit. The input / output unit 28 includes an input unit 282 and an output unit 284, and the output unit 284 includes a display unit 284a. Descriptions of those overlapping with the described content and those related to the basic functions to be described later are omitted.
[0159] In this embodiment, the terminal 20 is a desktop PC. However, the terminal 20 is not limited to this, and may be a portable terminal such as a smartphone or a tablet.
[0160] In this embodiment, the terminal 20 does not need to install a specific application program, and various processes can be executed by accessing the website W, and the deformation character generation system 1 can be used.
[0161] (Description related to the basic functions of the computer) Hereinafter, the control unit (processor, ROM, RAM, timing unit), the storage unit, the communication control unit, the input unit, and the output unit will be described. Note that, in any of the terminals of this embodiment, the connection mode (network topology) between the functional units is not particularly limited. For example, it may be a bus type, or a star type, a mesh type, etc.
[0162] The processor performs information processing and controls various devices according to programs stored in a ROM, a storage unit, etc. In this embodiment, the processor is a CPU (Central Processing Unit).
[0163] However, the processor is not limited to a CPU. Various processors such as a CPU, a DSP (Digital Signal Unit), a GPU (Graphics Processing Unit), a GPGPU (General Purpose computing on GPU), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit) may be used alone or in combination. For example, a processor integrating a CPU and a GPU is called an APU (Accelerated Processing Unit), and such a processor may be used.
[0164] The ROM is a read-only memory in which various programs and data for the processor to perform various controls and operations are stored in advance.
[0165] The RAM is a random access memory used as a working memory for the processor. Various areas for performing various processes of this embodiment can be secured in this RAM.
[0166] The timing unit performs timing processing related to acquisition of time information, etc. When the computer includes a communication control unit, time information may be acquired from the outside by NTP (Network Time Protocol).
[0167] The storage unit is a device for storing information such as programs and data. The storage unit is also referred to as a storage. The storage unit may be of an internal type or an external type.
[0168] The storage unit includes a storage medium capable of reading and writing data and a drive for reading and writing to the storage medium. The storage medium includes, for example, built-in and external types, and examples thereof include HD (hard disk), CD-ROM, flash memory, and the like. Examples of the drive include HDD (hard disk drive), SSD (solid state drive), and the like.
[0169] The storage unit includes a program storage unit and a data storage unit as functional units. The program storage unit stores control programs for controlling various devices, such as a communication control program for controlling communication.
[0170] The communication control unit is a device for performing communication between terminals and the like. The communication control unit connects a terminal including the communication control unit to the network N.
[0171] The communication method of the communication control unit is a known method, and a wired or wireless method is applied according to the device. For example, if the terminal is a desktop PC, both wired and wireless cases are considered. If the terminal is a smartphone, a wireless communication method is considered.
[0172] In the case of wired, for example, a communication method defined by IEEE802.3 (for example, bus-type or star-type wired LAN) can be preferably used. However, in addition, a communication method defined by IEEE802.5 (for example, ring-type wired LAN) or the like can also be used.
[0173] In the case of wireless, for example, a communication method defined by IEEE802.11 (for example, Wi-Fi) can be preferably used. However, in addition, IEEE802.15 (for example, Bluetooth (registered trademark), BLE (Bluetooth (registered trademark) low energy), etc.), IEEE802.16 (for example, WiMAX), ZigBee (registered trademark), 920MHz band wireless (Wi-SUN, etc.), or a communication method defined by optical communication such as infrared communication can also be used.
[0174] The input unit and the output unit are devices responsible for input to and output from the terminal, respectively. In some cases, the input unit and the output unit together may be referred to as the input / output unit. The input unit is a device that receives input from the user. Examples of such input units include a keyboard, a mouse as a pointing device, a trackpad, a tablet, or a touch panel.
[0175] When the terminal is a tablet or a smartphone, etc., and the input unit is a touch panel, the input unit is arranged on the surface of a display unit, such as a touch screen, that displays an image, etc. In this case, the input unit identifies the touch position of the user corresponding to various operation icons displayed on the display unit and receives the input from the user.
[0176] The output unit is, for example, a device for outputting an image, voice, document, etc. Examples of the output unit include a display device (display unit) such as a touch screen or a display (liquid crystal display or organic EL display), a voice output device such as a speaker, and a document output device such as a printer.
[0177] With the above configuration, the user can obtain the desired deformed character readily online.
[0178] (Second Embodiment) In the above-described embodiment, the processor 122 obtained the setting of the body proportion by obtaining a reference image of the deformed character (step 6 - reference image acquisition process). On the other hand, in the present embodiment, after the processor 122 receives an input to the deformed character setting field (UI - 12), that is, obtains the setting of the body proportion by the user's input (body proportion acquisition process), By adding the body-to-head ratio to the prompt (character information / character prompt) for the machine learning model (diffusion model) (referred to as "prompt addition process"), the body-to-head ratio of the output deformed character image is controlled.
[0179] That is, instead of controlling the body-to-head ratio in the form of a control model that controls the diffusion model, the body-to-head ratio is controlled as an input to the diffusion model (text input by the text encoder).
[0180] In the first embodiment, there was a reference image analysis process, and the reference image analysis process included a body-to-head ratio acquisition process for acquiring the setting of the body-to-head ratio by the user's input. However, in this embodiment, as another means of processing by the reference image analysis process (other than the body-to-head ratio acquisition process), there are the above-described body-to-head ratio acquisition process and prompt addition process.
[0181] Also, in this embodiment, when a reference image is uploaded by the user, the processor 122 recognizes the head and the whole body of the reference image, and calculates the body-to-head ratio by measuring their respective sizes (lengths) (referred to as "image body-to-head ratio analysis process").
[0182] Note that the programs for the processor 122 to execute the body-to-head ratio acquisition process, prompt addition process, and image body-to-head ratio analysis process are referred to as the body-to-head ratio acquisition program, prompt addition program, and image body-to-head ratio analysis program, respectively.
[0183] By the processing of this embodiment, a deformed character image that precisely reproduces the characteristics of the original character image can be obtained. Note that advantages such as more precisely reproducing the characteristics of the original character image than the aspect of the first embodiment were confirmed in the aspect of this embodiment.
[0184] (Modification example) The present invention is not limited to the above-described embodiments, and includes those obtained by making various changes to the above-described embodiments without departing from the gist of the present invention.
[0185] For example, in the above-described embodiment, the machine learning model (diffusion model) was created each time by updating the parameters. However, the learning and creation of the machine learning model are not limited to this. For example, it may be not only limited to updating the parameters, but also may be such that the model structure is changed or created from the model structure. However, updating the parameters is more preferable because the process is faster.
[0186] Also, as described above, it may be possible to obtain an already created machine learning model (diffusion model). Further, the created character model may be integrated (incorporated) into an (existing / other) Stable Diffusion diffusion model. However, by obtaining and analyzing the original character image and updating the parameters of the diffusion model to obtain a new diffusion model, the reproducibility of the original character image in the deformed character image is significantly improved.
[0187] In the above-described embodiment, the machine learning model (character model) included a diffusion model and a control model, but it is not limited to this. First, it may be composed of only the diffusion model. Also, as long as the object of the present invention is achieved, it may be composed of a machine learning model other than the diffusion model. However, the aspect including the diffusion model and the control model can generate a deformed character image that accurately reproduces the feature amount of the original character image.
[0188] In order to realize feature decoupling, the following approach can be used. In the above-described embodiment, the following (1) is used. (1) Separation of the latent space Design the latent space of the generative model and train it so that different dimensions control different feature amounts. For example, each dimension is in charge of "hairstyle", "face (expression)", or "pose", etc. (2) Architecture design Design the layers and modules of generative networks (such as GANs and Diffusion Models) so that specific layers process specific features. For example, "multi-scale feature extraction" and "style application at the layer level" for style conversion, etc. (3) Regularization of learning To maintain the independence between features, introduce a penalty to suppress the interaction between features during the learning process. (4) Conditioning between modules By adjusting the generation process based on specific conditions (text, attributes, sketches), emphasize specific features.
[0189] Aspects of the present invention including this embodiment, in other words, have the following features. The following corresponds to the scope of the claims at the time of filing of this application. However, it may be different from the description in the scope of the claims after the amendment of the scope of the claims after filing. (1) In the first aspect, cause a computer to function as original character image acquisition means for acquiring an original character image which is an image of a character, character prompt acquisition means for acquiring character information for explaining the original character image, body ratio acquisition means for acquiring a setting of the body ratio by a user's input, and deformed character acquisition means for acquiring a deformed character image in which the original character image is deformed to the body ratio, with the original character image, the character information, and the body ratio as inputs to a machine learning model, wherein the machine learning model learns at least a character image and a deformed character image as learning data, and outputs a deformed character image in which the character image is deformed with the character image as an input, and provide a deformed character generation program characterized by this. (2) In the second aspect, further provided is an optimization means for obtaining and evaluating the difference between a plurality of output images obtained by changing weights in at least a part of the layers of the neural network included in the machine learning model, and reducing or deleting the weights of the layers that have little influence on the output images. The deformation character generation program according to the first aspect is characterized by this. In this case, a deformed character image with a higher degree of reproduction of the original character image can be obtained. (3) In the third aspect, provided is a deformed character generation system including an original character image acquisition unit for acquiring an original character image that is an image of a character, a character prompt acquisition unit for acquiring character information for explaining the original character image, a body ratio acquisition unit for acquiring a body ratio setting by a user's input, and a deformed character acquisition unit for using the original character image, the character information, and the body ratio as inputs to the machine learning model and acquiring a deformed character image in which the original character image is deformed according to the body ratio. The machine learning model learns at least a character image and a deformed character image as learning data, and outputs a deformed character image in which the character image is deformed by using the character image as an input. (4) In the fourth aspect, a method for generating a deformed character includes: an original character image acquisition step of a processor acquiring an original character image which is an image of a character; a character prompt acquisition step of the processor acquiring character information explaining the original character image; a body ratio acquisition step of the processor acquiring a setting of a body ratio by a user's input; and a deformed character acquisition step of the processor acquiring a deformed character image in which the original character image is deformed to the body ratio, with the original character image, the character information, and the body ratio being inputs to a machine learning model. The machine learning model learns at least a character image and a deformed character image as learning data, and outputs a deformed character image in which the character image is deformed with the character image as an input. A method for generating a deformed character is provided, characterized by the above. (5) In the fifth aspect, a trained machine learning model for generating a deformed character is provided, which is a machine learning model that learns at least a character image, character information explaining the character image, a deformed character image in which the character image is deformed, and a body ratio of the character image and the deformed character image as learning data, and outputs a deformed character image in which the original character image is deformed to the body ratio with the original character image which is an image of a character, the character information explaining the original character image, and a setting of a body ratio by a user as inputs.
Industrial Applicability
[0190] Since cute characters are in demand throughout the industry (especially the entertainment industry) regardless of online or offline, a program capable of generating such characters enables the provision of characters that meet the needs.
Explanation of Signs
[0191] 1 Deformed Character Generation System 10 Server 12 Control Unit 122 Processor 124 ROM 126 RAM 128 Timing unit 14 Memory unit 14a Program storage unit 14b Data storage unit 16 Communication control unit 18 Input / output unit 20 Terminal 22 Control unit 24 Memory unit 26 Communication control unit 28 Input / output unit 282 Input unit 284 Output unit 284a Display unit UI-11 Upload field UI-12 Deformed character setting field UI-121 Body ratio setting button UI-122 Reference deformed character image UI-123 Deformed character generation start button UI-13 Original character image confirmation screen (window) UI-131 Image information UI-132 Original character image UI-133 OK button icon UI-14 Output image display field UI-141 Output image UI-142 Output image processing related buttons UI-143 Output image information UI-144 Output image saving related buttons P1 Deformed character generation program P11 Original character image analysis program P12 Reference image analysis program P13 Deformed character acquisition program P14 Optimization program
Claims
1. Computer, an original character image acquiring means for acquiring an original character image which is an image of a character; a character prompt acquisition means for acquiring character information explaining the original character image; a head-to-body ratio acquisition means for displaying a user interface for setting the head-to-body ratio of a deformed character image to be output, accepting an input from a user, and acquiring the head-to-body ratio setting input by the user; a deformed character acquisition means for acquiring a deformed character image obtained by deforming the original character image to the head-to-body ratio using the original character image, the character information, and the head-to-body ratio as inputs to a machine learning model; Function as a The machine learning model learns at least a character image and a deformed character image as learning data, A deformed character generating program which receives a character image as input and outputs a deformed character image obtained by deforming the character image, and which includes a diffusion model.
2. The deformed character generation program of claim 1, further comprising an optimization means for obtaining and evaluating differences between multiple output images obtained by changing weights in at least some layers of a neural network provided in the machine learning model, and reducing or deleting weights in layers that have little effect on the output image.
3. an original character image acquisition unit that acquires an original character image which is an image of a character; a character prompt acquisition unit that acquires character information explaining the original character image; a head-to-body ratio acquisition unit that displays a user interface for setting the head-to-body ratio of a deformed character image to be output, receives user input, and acquires the head-to-body ratio setting input by the user; a deformed character acquisition unit that acquires a deformed character image in which the original character image is deformed to have the head-to-body ratio by using the original character image, the character information, and the head-to-body ratio as inputs to a machine learning model, The machine learning model learns at least a character image and a deformed character image as learning data, A deformed character generating system which receives a character image as input and outputs a deformed character image obtained by deforming the character image, and which is characterized by comprising a diffusion model.
4. an original character image acquisition step in which the processor acquires an original character image which is an image of the character; a character prompt acquisition step in which a processor acquires character information explaining the original character image; a head-to-body ratio acquisition step of displaying on a display unit a user interface for setting the head-to-body ratio of a deformed character image output by the processor, receiving an input from a user, and acquiring a head-to-body ratio setting input by the user; and a deformed character acquisition step in which a processor acquires a deformed character image in which the original character image is deformed to the head-to-body ratio by using the original character image, the character information, and the head-to-body ratio as inputs to a machine learning model; Equipped with The machine learning model learns at least a character image and a deformed character image as learning data, A deformed character generating method for inputting a character image and outputting a deformed character image obtained by deforming the character image, the method comprising the steps of: inputting a character image; and outputting the deformed character image, the deformed character generating method comprising the steps of: inputting a character image;
Citation Information
Patent Citations
Arithmetic unit, arithmetic method, and learning method
JP2021047711A
Neural network optimizing method, neural network optimizing device and program
JP2021105950A
Image conversion device, image conversion method, and image conversion program
JP2023157334A
Information processing device, information processing method, and information processing program
JP2023169048A
Image acquisition device, image acquisition method, program, and recording medium
WO2024161839A1