Deformation Character Generation Program and Deformation Character Generation System

By decoupling pose and other feature amounts in the system, the challenges of preserving essential features in deformed character generation are addressed, resulting in more accurate and efficient output.

JP7698855B1Active Publication Date: 2025-06-26CTW株式会社
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025014513
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2025-06-26
Estimated Expiration
2045-01-31

AI Technical Summary

Technical Problem

Existing methods for generating deformed characters, such as converting an 8-head character to a 2-head character, require significant labor to accurately preserve essential features like expressions and clothing while removing unnecessary parts.

Method used

The system acquires multiple feature amounts from an input image and decouples the pose feature amount from other feature amounts, allowing for independent processing and improved flexibility and accuracy in generating deformed characters.

Benefits of technology

This approach enables the generation of deformed characters that accurately retain the main features of the original character image, reducing the need for manual adjustments and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698855000001_ABST
    Figure 0007698855000001_ABST
Patent Text Reader

Abstract

Generate a deformed character from a normal character while leaving the necessary features. 【Solution means】The deformed character generation program P1 generates a deformed character from the original character image through (1) original character image analysis processing, (2) reference image analysis processing, and (3) deformed character acquisition processing. In order to leave the necessary features when generating the deformed character, a plurality of feature amounts are obtained from the input image (for example, an 8-head body character image), and among them, the most important feature is to separate (decouple) the feature amount of the pose from the feature amounts other than the feature amount of the pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a deformed character generation program and a deformed character generation system.

Background Art

[0002] In contents in various fields such as games, animations, education, and news reporting, so-called deformed characters such as two-headed characters are used. These have effects such as giving a cute impression to viewers and making it easier to feel familiar with the content.

[0003] Although two-headed characters may be created as original characters, there are also cases where characters originally drawn as eight-headed or the like are deformed into two-headed characters.

[0004] However, when creating such deformed characters, it is necessary to devise ways to leave characteristic parts such as expressions and clothing while deleting unnecessary parts, which requires a lot of labor.

[0005] For example, Patent Document 1 discloses an image conversion device that can convert a character image into a deformed character with two heads or three heads.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Non-Patent Documents

[0007]

Non-Patent Document 1

[0008] In recent years, by using generative AI and other tools, the labor of illustrators and the like can be reduced, for example, when an original character illustration or the like is input, another illustration with different expressions, clothing, etc. can be generated.

[0009] Non-Patent Document 1 discloses a technique for outputting a character image of preference by inputting a character image and a pose.

[0010] However, on the other hand, when, for example, an 8-head character is input and a 2-head character is used as the output of the generative AI, there are some problems. For example, when deforming, it is necessary to select what to leave (the emphasized part) and what not to leave (the non-emphasized part).

[0011] The physical characteristics and clothing of characters are diverse. For example, there are characters with unique hairstyles, eye colors, clothing designs, items carried, and ornaments. When deforming, what parts to leave and what parts to delete vary depending on the character. If the hairstyle changes when deforming a character with a characteristic hairstyle, it can be said that the deformation is inappropriate.

[0012] However, when trying to obtain the desired deformed character as the output of the image generation AI, there are problems such as a great deal of labor being required, such as relying on the ingenuity of the prompt.

Summary of the Invention

Problems to be Solved by the Invention

[0013] The problem to be solved is that it is difficult to deform while leaving the necessary features when generating a deformed character from a normal character.

Means for Solving the Problems

[0014] In order to leave the necessary features when generating a deformed character, the present invention acquires a plurality of feature amounts from an input image (for example, an 8-head character image), and among them, separates (decouples) the feature amount of the pose (and other feature amounts) (feature amount separation) as the most main feature. By separating the feature amounts, each feature amount is processed without depending on other feature amounts, so the flexibility and accuracy of the model are significantly improved.

[0015] Non-Patent Document 1 does not describe converting to a deformed character or separating the feature amount of the pose and other feature amounts from the feature amount (feature) obtained from the input character image (Feature Decoupling, described later).

[0016] The present invention has been made in view of the above problems, and for example, the following means are adopted. That is, the computer is made to function as (1) original character image analysis means, (2) reference image analysis means, and (3) deformed character acquisition means. The above (1) original character image analysis means Original character image acquisition means for acquiring an original character image that is an image of a character, Plural feature amount acquisition means for acquiring a plurality of feature amounts from the original character image, Character prompt acquisition means for acquiring character information that describes the original character image from the original character image, and Character model acquisition means for acquiring a machine learning model that learns, as learning data, at least a part of the feature amounts other than the pose feature amount among the plurality of feature amounts, The reference image analysis means in (2) above, Reference image acquisition means for acquiring a reference image of the deformed character, and, Control model acquisition means for acquiring a control model that is learning, as learning data, data based on line drawing information indicating at least a part of the deformed character, The deformed character acquisition means in (3) above is characterized by acquiring a deformed character image obtained by deforming the original character image from the machine learning model controlled by the control model and the character information, and provides a deformed character generation program.

Advantages of the Invention

[0017] The deformed character generation program of the present invention has the advantage that it can output a deformed character while leaving the main feature amounts of the original character image because it separates (decouples) the pose feature amount from the feature amounts other than the pose feature amount.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Mode for Carrying Out the Invention

[0019] Embodiments of the present invention will be described with reference to the drawings. In the following embodiments, the same or corresponding parts may be denoted by the same reference numerals and the description may be omitted as appropriate. In addition, the drawings used below are for explaining the present embodiment, and may be different from the actual device configuration, user interface (UI), data configuration, etc.

[0020] (Overview of the Embodiment) The overview of the present embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an overview of the processing by the deformed character generation program P1 of the present embodiment.

[0021] The deformed character generation program P1 is a program that takes an image of a tall character (the 5 - 6 head - to - body character in the upper left in FIG. 1) as input and outputs an image of a short character (the 2 - 3 head - to - body character in the lower right in FIG. 1).

[0022] The deformation character generation program P1 includes: (1) an original character image analysis means (upper right in FIG. 1) for acquiring an original character image and various information (such as a character prompt and a character model); (2) a reference image analysis means (lower left in FIG. 1) for acquiring a reference image and obtaining a control model; and (3) a deformed character acquisition means (lower right in FIG. 1) for outputting a deformed character based on the information obtained in (1) and (2).

[0023] Also, a system that performs the processing by the deformation character generation program P1 is hereinafter referred to as the "deformation character generation system 1".

[0024] (Details of the Embodiment) Hereinafter, the deformation character generation system 1 according to the present embodiment will be described in detail. The deformation character generation system 1 includes a computer (server 10) equipped with the deformation character generation program P1, and provides an online system that deforms and outputs a character image related to an input for a user. That is, in the deformation character generation system 1, the information processing by the deformation character generation program P1 is specifically realized using hardware resources. Hereinafter, the 1. user interface, 2. program processing, and 3. hardware configuration that constitute the deformation character generation system 1 will be described in order.

[0025] (Definition of Terms) Here, some terms will be defined. A "character" has a visible appearance, regardless of whether it is a living or non-living thing, real or imaginary. Here, it particularly means something that can be divided into a head and the rest of the body. For example, a character in a creative anime (a living and imaginary thing). Here, a non-living character means, for example, a robot character. Furthermore, a real character means, for example, a celebrity drawn in an anime style. Note that even if the location of the neck is hidden, such as when a person is completely covered with a sheet from the head, and the head is not clear, as long as the approximate position can be inferred, it is acceptable. Also, for example, a part with a face (eyes and mouth) can be regarded as the head. A "character image" is an image of a character, regardless of whether it is an illustration image or a photo image. A "pose" means the posture taken by a character. "Deformation" means intentionally changing the form of an object. For example, it refers to changing the head-to-body ratio of a character from a tall-headed character to a short-headed character. "Tall-headed character" and "short-headed character" generally refer to a character with a high head-to-body ratio (e.g., 6 to 9 heads) and a character with a low head-to-body ratio (e.g., 2 to 5 heads), respectively. However, the high and low here are relative, not absolute. When referring to deformation, the original character image (hereinafter sometimes referred to as the "input image") has a high head-to-body ratio, and the image obtained by deforming the original character image (hereinafter sometimes referred to as the "output image") has a low head-to-body ratio. For example, changing a 4-headed character to a 2-headed character is also called deformation. Generally, a short-headed character is also called a chibi character. "Feature Decoupling" is a method in a machine learning model, particularly an image generation model, of independently handling different feature quantities (e.g., feature quantities related to pose (posture), clothing, expression, etc.). Feature quantities are also referred to as variables, etc. "User" refers to a person who uses the deformed character generation system 1. The user operates the terminal 20, regardless of whether they are an individual or a corporation. In the deformed character generation system 1, the user inputs an image of a tall character and obtains the transformed short character as the output. "Obtain" includes not only the processor obtaining data, programs, etc. from the input unit or an external terminal, but also the processor creating, updating, etc. data, programs, etc. Please refer to the character models in the following embodiments.

[0026] The terms used in the drawings and the like are explained below. However, the following is for helping the understanding of the invention and does not limit the content of this specification. "Stable Diffusion" is a machine learning model that converts text, etc. into images and is one of the so-called image generation AIs. Stable Diffusion is characterized by using a diffusion model (especially, a latent diffusion model). Also, Stable Diffusion has a text encoder for inputting text for conditioning. The text encoder converts text into a vector and is, for example, a Transformer (a machine learning model) trained by a method such as CLIP (learning combinations of a large number of images and text and calculating the similarity between images and text). The processing in Stable Diffusion includes the processing of obtaining latent variables from image information by VAE (described in the next section) and the processing of obtaining image information from latent variables by VAE. It is characterized by suppressing the amount of calculation by using VAE. The specific embodiments of the diffusion model will be described in the following embodiments. "VAE (Variational Auto-Encoder)" refers to the variational auto-encoder. VAE is a generative model with an encoder-decoder structure. However, the term VAE may also refer to a neural network with such a structure, such a network structure, a machine learning model, or a technology including these. VAE learns to represent input data in terms of the mean and standard deviation (or variance, etc.) (such as a mean vector, variance vector, etc.). For example, the encoder converts the input data into a latent variable (which becomes a single point in a statistical distribution), and the decoder generates new data by restoring the latent variable (a single point randomly sampled from a statistical distribution). Also, a device (Reparametrization trick) for error backpropagation is made during learning. "Unet" is a neural network based on CNN (Convultional Neural Network). For example, it includes a module containing a convolutional layer and a module containing an attention layer. Generally, it is used for image segmentation, but here it is used for the reverse diffusion process of a diffusion model, for example, noise removal in the latent space. Unet, for example, takes noise as input and outputs the noise to be removed. By removing the noise to be removed from the input noise, an image with the noise removed is obtained. By repeating this, the noise is gradually removed from the image. 「ControlNet」 is an extension of Stable Diffusion, enabling the generation of (high-quality) images based on the poses and compositions specified by the user. ControlNet is an additional neural network that takes the structural features (poses, line drawings, edges, etc.) of an image as conditions and generates an image reflecting them. While retaining the neural network blocks of Unet, it adds the additional conditions given by the user in the neural network on ControlNet and reflects them in Unet. That is, conditions (such as composition data) etc. are input, and the blocks on Unet are the output destinations. Although the details are omitted, technologies such as zero convolution and skip connections are incorporated. 「LoRA (Low-Rank Adaptation)」 is a technology related to the fine-tuning of (pre-)trained (machine learning) models. It should be noted that it is different from LoRa (registered trademark), which is a low-power long-distance wireless communication technology.

[0027] In the following, when described as "○○" processing, it means that the computer's processor executes the processing based on the "○○" program stored in the program storage section. In this paragraph, the same word is entered in the "○○" part. That is, the "○○" program is a program that makes the computer function as "○○" means by executing the "○○" processing. Also, at this time, it means that the control section equipped with the processor also functions as the "○○" section (or "○○" device). In this case, the "○○" section means executing the "○○" processing based on the "○○" program. Also, when indicating the processing procedure (of the "○○" processing), it is described as the "○○" step.

[0028] For example, the deformed character generation program P1 is a program that causes a computer to function as a deformed character generation means by executing the deformed character generation process. At this time, the control unit 12 of the computer including the processor 122 functions as a deformed character generation unit (or a deformed character generation device).

[0029] In the deformed character generation system 1, each terminal (computer) such as the terminal 20 includes a processor. When simply referring to the processor, it refers to the processor that performs processing by the deformed character generation program P1, which is the processor 122 of the server 10 in this embodiment.

[0030] In the following, for simplicity, "the processor 122 of the server 10 receives a request from the terminal and returns data for display on the browser of the terminal" may be described as "the processor 122 displays (causes to display) on the browser of the terminal" or "the processor 122 displays (causes to display)". Similarly, "the processor 122 of the server 10 causes the data storage unit 14b of the storage unit 14 to store data" may be described as "the processor stores (causes to store) (data)".

[0031] (First Embodiment)

[0032] 1. User Interface (UI) First, the interface that the deformed character generation system 1 of this embodiment displays on the terminal 20 will be described with reference to the drawings. The interface described hereinafter is a simplified version of what the processor 122 displays on the browser of the terminal 20.

[0033] Also, only icons and the like related to functions necessary for the explanation are displayed, and other known icons and the like are omitted. For example, the back button for returning to the previously displayed page is omitted.

[0034] Figure 2 is a diagram showing the basic setting screen. The image upload column UI-11 on the left side of the basic setting screen is a UI for the user to upload the original character image which is the input image, and is also a column for displaying the uploaded original character image. The user uploads the original character image by drag and drop. Note that the method of upload is not limited to this. It may also be possible to select the folder where the image is stored so that the image file can be selected (not shown).

[0035] The deformed character setting column UI-12 is a UI for the user to set the deformed character image to be output. The deformed character setting column UI-12 displays an equal ratio setting button UI-121, a reference deformed character image UI-122, and a deformed character generation start button UI-123. Details will be described later.

[0036] Figure 3 is a diagram showing the original character image confirmation screen. When the user uploads the original character image to the image upload column UI-11, the processor 122 acquires and creates a character prompt (character information) and a character model (machine learning model) described later from the image, and displays the original character image confirmation screen (window) UI-13. Note that in the figure, the term "training" is used for the creation of the character model.

[0037] In this embodiment, the processor 122 displays a confirmation screen regarding whether to perform training (learning) after upload, and the user starts learning by parameter update or the like by pressing the OK button icon or the like displayed there, but this is omitted here.

[0038] As shown in FIG. 3, the processor 122 displays image information UI-131 (such as file name, user ID, creator, and creation date and time), the original character image UI-132 uploaded by the user, an OK button icon UI-133, a re-training icon UI-134, and a delete button icon UI-135 on the original character image confirmation screen (window) UI-13.

[0039] When the user presses the OK button icon UI-133, the processor 122 displays the original character image in the image upload field UI-11 of the basic settings screen (not shown). If the user selects (presses) the original character image on the basic settings screen, the processor 122 displays the original character image confirmation screen (window) UI-13 again.

[0040] When the user presses the re-training icon UI-134 on the original character image confirmation screen (window) UI-13, the processor 122 displays a screen for selecting an original character image (not shown). When the user selects the original character image there and presses the button to execute re-training, the processor 122 creates a character prompt and a character model (again) from the original character image (input image) according to the user's selection. At this time, the parameters are adjusted arbitrarily or automatically by the user. In addition, the training method can be adjusted by the user (not shown).

[0041] When the user presses the delete button icon UI-135, the processor 122 deletes the uploaded original character image. The user can upload a new original character image.

[0042] FIG. 4 is a diagram (enlarged view) showing the deformed character setting column (UI-12).

[0043] In this embodiment, a large number of reference deformed character images (images of the deformed characters to be referenced) are stored in advance in the storage unit 14 (data storage unit 14b), and the user can freely select the orientation, pose, etc. of the deformed character.

[0044] Also, the user can upload their favorite deformed character. For example, in this embodiment, by dragging and dropping an image into the deformed character setting field UI-12, a favorite character can be used as a reference image.

[0045] The equal body ratio setting button UI-121 is a button icon for setting the head-to-body ratio of the output image. For example, when the 1:4 button is selected, the processor 122 causes an image with a head height: full body height of 1:4 to be output to the deformed character setting field UI-12 (see Figure 2. Also, see the shaded part of the reference deformed character image UI-122 in Figures 2 and 4).

[0046] As shown in Figure 4, the All button of the equal body ratio setting button UI-121 displays all the images with the shown (1:2.5 to 1:5 in Figure 4) ratios.

[0047] When the user selects and presses the reference deformed character image UI-122, the processor 122 displays a check (dot) to indicate that the image has been selected. The user can select multiple reference deformed character images, and in that case, the processor 122 generates multiple deformed characters according to the selection. That is, the processor 122 can output multiple deformed character images at once.

[0048] The reference deformed character image UI-122 displays a guide for the pose and head-to-body ratio of the deformed character to be output. This image is automatically selected by the processor 122 from a large number of stocks. On the other hand, the processor 122 displays a pose display window (not shown) for selecting a pose (such as a standing pose or a sitting pose) according to the user's selection, and the user can also select a preferred pose from the group of poses listed in the pose display window.

[0049] That is, the user can easily input the head-to-body ratio by pressing the equal body ratio setting button UI-121 or by uploading a reference image. In other words, the user can input the head-to-body ratio by a simple operation using the GUI.

[0050] When the user presses the deformed character generation start button UI-123, the processor 122 starts the process related to the generation of the deformed character. The process will be described later.

[0051] FIG. 5 is a diagram showing a generated image display screen. When the processor 122 acquires the image (output image) of the deformed character, the output image UI-141 is displayed in the output image display column UI-14 of the generated image display screen. In addition to this, the processor also displays an output image processing related button UI-142, an output image information UI-143, and an output image saving related button UI-144.

[0052] The output image processing related button UI-142 is a button icon for further processing the output image or regenerating the image.

[0053] The magnification button shown in FIG. 5 is a button icon for changing the displayed magnification. The color button is a button icon for performing color processing such as the skin color or hair color of the character. The part separation button is a button icon for separating and displaying the parts (such as hairstyle, clothing, etc.) of the deformed character from other parts. The user can perform separate processing on that part independently. The regenerate button is a button icon for regenerating the deformed character image. Different parameters are set in the machine learning model (random latent variables are sampled in the diffusion model), and deformed character images (usually different deformed character images) corresponding to the parameters can be generated. That is, the user has the advantage of being able to easily obtain different images with just one button.

[0054] The output image information UI-143 is a column for displaying information about the image. The style is the file name of the reference image. The variation intensity is a parameter for adding variations. The larger the variation intensity, the greater the change from the original character image, and the image with a greater change is output as the output image. In this embodiment, the variation intensity can be set by the user (the UI is not shown).

[0055] The output image save-related button UI-144 is an operation button icon related to saving the output image. DL is for downloading the output image to the terminal (terminal 20), delete is for deleting the output image, and the save button is for saving the output image on the server (server 10) (such as on the server).

[0056] With the above configuration, the user can generate a deformed character image with a simple operation using the GUI (Graphical User Interface). In particular, since the user can select the head-to-body ratio while viewing the reference deformed character image UI-122 as a reference example, the output image is easy to predict, and the desired deformed character image can be obtained more reliably. Also, at this time, there is a remarkable effect that the processing can be executed on the GUI without performing inputs such as finely adjusting the prompt.

[0057] 2. Program Processing <2-1. Deformed Character Generation Process> The program processing performed in the deformed character generation system 1 of the present embodiment will be described.

[0058] In the present embodiment, the processor 122 performs a deformed character generation process based on the deformed character generation program P1. The deformed character generation program P1 includes at least (1) the original character image analysis program P11, (2) the reference image analysis program P12, (3) the deformed character acquisition program P13, and (4) the optimization program P14. Based on these programs, the processor 122 executes the original character image analysis process, the reference image analysis process, the deformed character acquisition process, and the optimization process, respectively. In addition to the above, the deformed character generation process includes an online UI providing process for providing the UI described above to the user online. Since the UI to be provided has been described, it will be omitted here.

[0059] (1) The original character image analysis process is a process of acquiring and analyzing the original character image, The original character image acquisition process of acquiring the original character image, which is the image of the character (to be processed), The multiple feature quantity acquisition process of acquiring the image information of multiple regions from one said original character image to obtain multiple feature quantities, The tag management process of attaching tags to and managing each of the regions (of the image) and / or feature quantities, The feature quantity separation process of separating the multiple feature quantities into pose feature quantities and other feature quantities (Feature Decoupling), The character prompt acquisition process of acquiring the character information describing the original character image from the original character image (in particular, acquiring the character information describing the feature quantities of the concept including the feature quantities other than the pose feature quantities among the multiple feature quantities), and, A character model acquisition process that acquires a machine learning model (diffusion model) that learns (or is learning) at least a part of the feature amounts other than the pose feature amount among the plurality of feature amounts (feature amounts of a concept including feature amounts other than the pose feature amount) as learning data, is included.

[0060] Here, the above character information is referred to as a "character prompt", and the above machine learning model (diffusion model) is referred to as a "character model".

[0061] (2) The reference image analysis process is a process of acquiring and analyzing a reference image of a deformed character, a reference image acquisition process for acquiring a reference image of a deformed character, and a control model acquisition process for acquiring a control (machine learning) model that has learned at least one piece of line drawing information indicating a part (shape) of the deformed character in the reference image as learning data, are included.

[0062] Here, the above control (machine learning) model is referred to as a "control model". Note that it is denoted as a control model in order to distinguish it from the above-mentioned machine learning model (character model). For example, the character model may be referred to as a first machine learning model, and the control model may be referred to as a second machine learning model.

[0063] (3) The deformed character acquisition process is a process of acquiring a deformed character image obtained by deforming the original character image from the machine learning model (diffusion model) controlled by the control model and the character information.

[0064] More specifically, it is a process of inputting the character information (= character prompt) to the machine learning model (= character model) controlled by the control model (= control model), and acquiring, as an output, a deformed character image obtained by deforming the original character image.

[0065] Roughly divided, it is the process of acquiring and analyzing an image (original character image, reference image) etc. (the above (1) and (2)), and the process of obtaining a deformed character image by image generation (the above (3)). In other words, the process of obtaining a deformed image by image generation (the above (3)) is a process of actually generating a deformed character from data, a model, etc. obtained from (the above (1) and (2)).

[0066] FIG. 6 is a flowchart showing the deformed character generation process. Upon receiving an instruction to start the process of generating a deformed character (for example, pressing the deformed character generation start button UI-123 in the deformed character setting field UI-12 in FIG. 4) input by the user, the processor 122 starts the deformed character generation process.

[0067] First, the (1) original character image analysis process will be described. The processor 122 acquires the original character image to be deformed input by the user (step 1, original character image acquisition process).

[0068] The processor 122 extracts and acquires a plurality of feature amounts from the acquired image (step 2, plurality of feature amount acquisition process). In the present embodiment, the processor 122 recognizes a plurality of regions such as regions related to the head, upper body, or whole body from one original character image, acquires the image information thereof, and acquires a plurality of feature amounts.

[0069] FIG. 7 is a diagram showing an image of recognizing a plurality of regions. In FIG. 7, a plurality of rectangular regions with different scales (parts surrounded by a quadrangular line, bounding box) indicate a plurality of regions.

[0070] In the present embodiment, the processor 122 acquires (crops) an image including a face at different scales. In particular, in this embodiment, a face detection process for recognizing a face is performed. As the face detection method, a known method is appropriately used. Note that the process of acquiring an image (by cropping) in this way is generally called "crop" or "cropping".

[0071] In the case of the example shown in FIG. 7, the plurality of regions include, in order from the smallest area of the rectangular region in FIG. 7, the face, the head, the upper body, from the top of the head to the waist, above the knees (including the hat, the same applies hereinafter), above the calves, above the ankles, above the heels, and above the toes (the whole body).

[0072] Acquiring (cropping) images of a plurality of regions with different scales makes it possible to clarify the site (position) where the feature amount should be acquired and obtain a plurality of feature amounts according to the site (position). For example, it is possible to acquire data regarding the face and other parts. In addition, there is an advantage of increasing the data amount from a small number (one) of images.

[0073] As a method of recognizing a plurality of regions from a single original character image and acquiring an image, a known object detection technique can be used.

[0074] Among object detection techniques, those that identify the class of an object after specifying the position of a general object are called two-stage models, and examples include R-CNN (Region Based Convolutional Neural Networks) and FPN (Feature Pyramid Networks). In addition, those that simultaneously specify the position of an object and identify the class are called one-stage models, and examples include YOLO (You Only Look Once) and SSD (Single Shot multibox Detector). In addition, a cascade classifier (OpenCV) or the like may be used.

[0075] "A plurality of regions" indicates a recognition unit (cropping unit) in the original character image. The processor 122 obtains one cropped image for one region. "A plurality of feature amounts" are a plurality of feature amounts obtained from the original character image. In this embodiment, it is a latent variable obtained by using the cropped image as the input of the VAE.

[0076] The processor 122 attaches tags to and manages each region (image) and / or feature amount (tag management process). In this embodiment, the images obtained by cropping and the tags correspond one-to-one. In this embodiment, the tags relate to the head, clothing, whole body, face, upper body, hair, decoration, and pose (see also FIG. 11 in this regard).

[0077] In this embodiment, the processor 122 manages by storing the image (file) and the text (file) including the tags in a set in the storage unit. Also, the processor 122 uses a part of this data (text) for a character prompt described later.

[0078] Examples of this text are given below for each tag. Each tag includes character information that describes the region (image) and / or feature amount. Incidentally, the character information that describes this region (image) and / or feature amount may be referred to as a prompt element or tag information for convenience. · Head: Hat with feathers, Pirate hat, Navy hat · Cloth: Miniskirt, Uniform, Knee boots, Jacket · Full body: Full body, Standing, Lying, Sitting · Face: smile, closed mouth, beard, blush · Upper body: upper body · Hair: long hair, short hair, single braid, twin ponytails · Decoration: necklace, earrings, bandage, jewelry · Pose: hand on hip, leg lift, knee up

[0079] In this example, the tags and text are words or phrases. However, this is not limited to this, and it may also be a sentence or the like.

[0080] Generally speaking, the technology that accepts input of image data and outputs text that summarizes (describes) its content is called Image Captioning. CLIP, Flamingo, etc. are known as (multimodal) models for performing Image Captioning.

[0081] Such technology is called Image to Prompt and is a technology for generating text from an image, which is different from OCR (Optical Character Recognition) that acquires character information in the image.

[0082] Here, the method of obtaining text from an image is described, but a method of converting feature amounts based on an image into text may also be used.

[0083] Also, the processor 122 acquires feature amounts corresponding to the tags. For example, feature amounts may be obtained for each tag (i.e., for each cropped image), or a plurality of obtained feature amounts may be classified according to the tags.

[0084] Note that a plurality of feature amounts may be obtained from one region (one image obtained by cropping), or one feature amount may be obtained from a plurality of regions. For example, the processor 122 may obtain a feature amount related to the head and a feature amount related to the hair from an image of the head.

[0085] Also, as shown in FIG. 7, since the image of the portion above the fingertip includes the entire body of the character, it includes a feature amount of the pose. However, it is not the case that the feature amount of the pose cannot be obtained without an image of the entire body. For example, an image including the upper body may include a feature amount of the pose that the character is putting a hand on the hip.

[0086] Also, just because there is an image of the face region, for example, it is not the case that the processor 122 cannot obtain a feature amount related to the face portion from an image of the upper body region. That is, there is no limitation on what information the processor 122 extracts from a certain region (image).

[0087] In the present embodiment, it is essential that the processor 122 obtains an image of the face region of the original character image and obtains a feature amount of the face. For example, the feature amount of the face determines the feature amount of the deformed character so that the face portion is particularly prominently displayed in the deformed character image.

[0088] Returning to FIG. 6, the processor 122 separates, among the plurality of feature amounts obtained in step 2, the feature amount of the pose and the feature amount of the concept including the feature amounts other than the feature amount of the pose (step 3 · feature amount separation process). This separation is referred to as feature decoupling.

[0089] As a method for realizing feature separation, it is possible to design the latent space of the generative model and train it to control different features in different dimensions. In particular, in this embodiment, the pose feature is separated from other features. Due to feature separation, each feature is processed without depending on other features, thereby improving the flexibility and accuracy of the model.

[0090] In this embodiment, features and elements other than the pose are expressed as concepts. For example, features other than the pose feature are used as the feature of the concept. In the above example, among the multiple features, the features related to the head, clothing, whole body, decoration, face, upper body, and hair become the features of the concept. That is, in this embodiment, multiple features = pose feature + concept feature.

[0091] The processor 122 acquires character information that describes the original character image from the original character image (step 4, character prompt acquisition process). For the sake of convenience, the character information (prompt) that describes this original character image is referred to as a "character prompt".

[0092] In particular, the processor 122 acquires character information that describes elements (images and / or features) other than the pose. In this embodiment, the processor 122 acquires character information that describes the parts of the image other than the pose, that is, the head, clothing, whole body, face, upper body, hair, or decoration, etc. (concept).

[0093] An example of a character prompt is as follows. 「1 girl, blue eyes, purple hair, long hair, police uniform, transparent background, open clothes, black skirt, mini skirt, high - heeled shoes, simple background, …(omitted)」

[0094] By obtaining the character prompt and using it as the input for the deformation character acquisition process described later, there is an advantage that the detailed characteristics of the character can be controlled and reflected in the deformation character.

[0095] On the other hand, the processor 122 acquires (step 5 - character model acquisition process) a machine learning model (diffusion model) that learns (is learning) at least a part of the feature amounts other than the pose feature amount among the plurality of feature amounts described above (the feature amounts of the concept including the feature amounts other than the pose feature amount) as learning data. This machine learning model (diffusion model) is referred to as a "character model" for convenience.

[0096] The training structure of this character model includes a diffusion model (Diffusion model) and a network (ControlNet) that controls the diffusion model. The character model is obtained by freezing the weights of ControlNet and updating (learning) only the parameters of the Diffusion model. There may be an expression of "(the processor 122) creates a character model each time".

[0097] Here, the general content of the "diffusion model (Diffusion Model)" will be described. The diffusion model is a machine learning model that is the core of the technology related to Stable Diffusion also used in this embodiment. The diffusion model is a model that can remove noise from a noisy image or the like and generate a noise-free image or the like. In the learning stage, the diffusion model includes a diffusion process of adding noise to the original information (pixels, latent space, etc.) and a reverse diffusion process of removing noise from the noisy information.

[0098] The diffusion model uses, as learning data, the original information such as an image, the noise to be added, the noisy information obtained by adding noise to the original information (e.g., a noisy image), the noise to be removed from the noisy information, and the original information obtained by removing noise from the noisy information.

[0099] And the diffusion model (the Unet included therein) learns the noise to be removed from the noisy information, etc. by using the above-mentioned noisy information as an input (learning stage). And the diffusion model (the Unet included therein) outputs the noise to be removed from the noisy information by using the noisy information (such as a noisy image) as an input. As a result of repeating this process, it is possible to output information from which noise has been removed (such as a noise-free image) from the noisy information (such as a noisy image) (inference stage).

[0100] In this embodiment, the parameters of the machine learning model (diffusion model) are updated by the feature amount of the concept obtained by performing feature separation.

[0101] Note that what is obtained by converting the pose into a latent variable by VAE is used for maintaining the layout, etc. (the "Freezed" part in ControlNet·Figure 1).

[0102] For example, in this embodiment, the feature amount of the concept obtained by performing feature separation becomes the above-mentioned original information. That is, the processor 122 creates and acquires a machine learning model (diffusion model) through learning that goes through a diffusion process of adding noise to the feature amount (latent variable) of the concept and a reverse diffusion process of removing noise.

[0103] If a deformed character image can be obtained, it may be a mode (mode 1) in which a plurality of images are not acquired from the original character image, or a method (mode 2) in which decoupling is not performed between the feature amount of the pose and other feature amounts after acquiring a plurality of images. However, according to the method of this embodiment, there is a remarkable effect that a deformed character image that accurately reflects the features of the original character image can be obtained.

[0104] Briefly, the processor 122 acquires a plurality of feature amounts from a single original character image and separates (decouples) them into the feature amount of the pose and other feature amounts (concepts). Also, character prompt information (character prompt) explaining this concept is acquired. Furthermore, a machine learning model (diffusion model (character model)) that learns based on the feature amounts of the concepts is created.

[0105] The character model of this embodiment is trained using only one person (character) image. The model after completing this training can be used to generate the features of this person (character) with different poses and styles (for example, SD (super deformed) style, etc.).

[0106] The specific training method is summarized below. The processor uses images that include a face and are cropped at different scales. The processor extracts prompts (tag information) from these images at different scales. The training structure of the machine learning model includes a Diffusion model and ControlNet. The weights of ControlNet are frozen, and the processor updates only the parameters of the Diffusion model. The Diffusion model after parameter update is the character model of this embodiment.

[0107] Note that it may be used by incorporating or improving based on an existing machine learning model (diffusion model). For example, a machine learning model (diffusion model) that has already been learned, such as by adding noise to a character image (including a deformed character image), may be used.

[0108] Next, the (2) reference image analysis process will be described. The processor 122 acquires a reference image of the deformed character (step 6: reference image acquisition process).

[0109] In this embodiment, the processor 122 first deletes the concept information of the reference image as preprocessing, and acquires scribble information (scribble information called Scribble) indicating a part of the shape of the deformed character, and pose information (Pose).

[0110] As shown in the lower left of FIG. 1, the processor 122 acquires scribbles of at least the head, fingertips, and toes of the reference image as scribble information. Thereby, the pose (orientation, etc.) and the head-to-body ratio of the deformed character to be output can be controlled (described later).

[0111] In addition, the processor 122 acquires joint position information of the reference image as pose information.

[0112] In the lower left of FIG. 1, the upper part of the image with black and white reversed is an image diagram of scribble information (Scribble), and the lower part is pose information (Pose).

[0113] In this embodiment, the processor 122 creates a control (machine learning) model from the acquired scribble information (Scribble) and pose information (Pose). This control (machine learning) model is hereinafter referred to as a "control model" for convenience.

[0114] This control model is an additional neural network that controls the diffusion model. In particular, it has an excellent effect of being able to control the character of the input image to an intended pose or the like. The control model of this embodiment controls the pose and / or the outer shape (affecting the height-to-width ratio) of the deformed character image.

[0115] The above-described processing can be executed by plugins or extensions of Stable Diffusion. For example, "ControlNet Extension" that can be added to the Stable Diffusion Web UI·Automatic1111 version of the Web UI can be preferably used.

[0116] In this embodiment, (ControlNet) ControlNet Scribble (one of the models of ControlNet) is used to acquire line drawing information. Also, (ControlNet) Open Pose (one of the models of ControlNet) is used to acquire pose information.

[0117] Note that Open Pose is a pose estimation method. By applying Open Pose to an image of a character, the pose of that character is estimated. Open Pose incorporates processing that takes into account the positional relationship between skeletons called Parts Affinity Fields.

[0118] Various "preprocessors" of ControlNet analyze and transform the input image (e.g., pose detection, edge detection, line drawing, etc.) and use it as a condition for the model. That is, as a function of ControlNet, there may be a mechanism for converting an image into a line drawing (edge detection / scribble detection), and the user can utilize it to generate "Scribble (line drawing)".

[0119] Briefly, the processor 122 creates and acquires a control (machine learning) model that learns, as learning data, data based on line drawing information indicating at least a part of the shape of the deformed character in the reference image (step 7 · control model acquisition process).

[0120] More specifically, data based on line drawing information and pose information indicating a part of the shape of a deformed character (in the reference image) (specifically, line drawing information and pose information transformed into feature quantities (latent variables) by a VAE) are learned as learning data, and a control (machine learning) model (= control model), which is an additional neural network for controlling the above-described diffusion model, is created and obtained. That is, this line drawing information and pose information become the input (condition) of the control model.

[0121] In this embodiment, although this control model is created each time, it may be possible to obtain a control model that is created in advance for each reference image and stored in the storage unit 14 of the server 10. In this case, there is an advantage that the processing becomes faster. However, when the user uploads a new image as a reference image, the processor 122 creates a control model for the new image.

[0122] Returning to FIG. 6, finally, (3) the deformed character acquisition process will be described. The processor 122 obtains a deformed character image obtained by deforming the original character image from the machine learning model (diffusion model) controlled by the control model and the character information (step 8, deformed character acquisition process).

[0123] More specifically, by inputting character information (= character prompt) to the diffusion model (= character model) controlled by the above-described control model (= control model), a deformed character image obtained by deforming the original character image is generated and obtained as an output.

[0124] There is a remarkable effect that a deformed character image with significantly higher accuracy can be obtained compared to, for example, a novice user entering conditions based on a prompt.

[0125] Then, the processor 122 causes the obtained deformed character image to be displayed on the display unit 284a of the terminal 20. When a plurality of head-to-body ratios are selected by the user, deformed character images corresponding to the respective head-to-body ratios are output. Also, as shown in the UI section, the processor 122 can accept input from the user and regenerate a deformed character image, etc.

[0126] Note that as long as the effects of the deformed character generation process in this embodiment can be exhibited, the order of the processes is not limited to that in FIG. 6. For example, in FIG. 6, steps 4, 5, and steps 6 to 7 are drawn sequentially, but these may be processed in parallel. Alternatively, steps 4, steps 2·3·5, and steps 6 to 7 may be processed in parallel.

[0127] <2-2. Optimization Process> In the optimization process, the processor 122 obtains and evaluates the differences between a plurality of output images obtained by changing the weights in at least some of the layers of the neural network (Unet) included in the above-described diffusion model (character model), and reduces or deletes the weights of the layers that have little influence on the output images.

[0128] The processor 122 performs the optimization process based on the optimization program P14. That is, the optimization program P14 causes the computer to function as an optimization means by the execution of the optimization process by the processor 122.

[0129] FIG. 8 is a flowchart showing the optimization process. Upon receiving an instruction from the user to regenerate the deformed image or automatically based on pre-set conditions, the processor 122 starts the optimization process (see Fig. 5 - Generated Image Display Screen). In this embodiment, the user can regenerate an image in which at least a part of the image is changed.

[0130] First, the processor 122 obtains the difference of the deformed character image when the weights are set to 1 and 0 respectively in each of a plurality of layers in the neural network (of the diffusion model) (step 21). While retaining the weights of the layers with a large influence, the processor 122 reduces or deletes (hereinafter referred to as "reduction, etc.") the weights of the layers with a large influence (step 22).

[0131] Fig. 9 is a diagram for explaining the reduction, etc. of weights. In Fig. 9, the part surrounded by the upper right rounded rectangle indicates that they are the same layer of the neural network. Also, the two characters surrounded by the same rounded rectangle are the deformed character images when the weights are set to 0 and 1 respectively in that layer. Diff in Fig. 9 indicates the difference between the deformed character image when the weight is 0 and the deformed character image when the weight is 1. The blank part is the difference.

[0132] If the difference between the respective images obtained by setting the weights to 1 and 0 in a certain layer is not very large, it means that that layer does not have a large impact on the overall result. That is, by evaluating this difference in each layer, the degree of influence of each layer on the whole can be grasped. And the processor 122 deletes, etc. the weights of the layers with little influence in response to this result.

[0133] For example, in the rightmost layer (the layer represented by the four rounded rectangles in Figure 9), the difference between the respective images generated due to the difference in weights (weights 1 and 0) is large (the difference is large). From this, it can be seen that the change in the image generated by the change in the weights of this layer is large, that is, the weights of this layer have a large impact on image generation. On the other hand, since the difference in the second layer from the left is small, it is considered that the influence of this layer on the generated image is small.

[0134] This setting optimizes the parameter layer in the fine-tuning method (LoRA) of the pre-trained model and finally has the effect of generating high-quality results.

[0135] Return to the flowchart of Figure 8. Place the layer with a large influence, that is, the layer related to the key features, in the output layer of the neural network (of Unet), and update the weights of the neural network layer by backpropagation (fine-tuning) (step 23). Through fine-tuning, a deformed character image that appropriately reproduces the features of the original character image can be obtained.

[0136] An example of the optimal image obtained by the optimization process is shown. Figure 10 is a diagram showing deformed character images before and after the optimization process. As shown in Figure 10, the image after optimization more faithfully reproduces the original character image than the image before optimization (especially the part surrounded by the rectangle around the waist).

[0137] Here, the specific method of this embodiment will be described. As a method for performing the above processing, for example, LoRA Block Weight can be used. With LoRa Block Weight, the application degree of LoRA can be adjusted.

[0138] Note that LoRA (Low-Rank Adaptation) is a method for performing fine-tuning on a pre-trained model. Among them, those that tune the concept, those that reflect the concept of the image to be generated, or those that apply a specific concept to the character are called Concept LoRA. By using Concept LoRA, the above adjustments can be made.

[0139] To summarize the above, through the original character image analysis process, a plurality of feature amounts are obtained from a plurality of regions of one original character image, and the features of the original character image are reflected as much as possible in the deformed character. Here, a character prompt and a character model are created (acquired). Also, through the reference image analysis process, after processes such as extracting a diagram or a pose from the reference image, a control model for controlling the character model is obtained. Then, in the deformed character acquisition process, by inputting a character prompt to the character model controlled by the control model, a deformed character faithful to the original character image can be generated. Its high accuracy is as shown in the drawing. And through the optimization process, a higher-quality deformed image can be obtained.

[0140] Also, the above-described diffusion model (character model) is a learned model that obtains a deformed character image obtained by deforming the original character image, which is an image of a character, obtains a plurality of feature amounts from the original character image, and uses, as learning data, data including at least the feature amounts of the concept including the feature amounts other than the feature amounts of the pose among the plurality of feature amounts, the noise added to the feature amounts of the concept, and the noisy information obtained by adding the noise, and learns the noise to be removed from the noisy information, characterized by taking the noisy information as an input and outputting the noise to be removed from the noisy information, Furthermore, a learned model is characterized in that a pose in a deformed character image to be output is controlled by a control model (= control model) that has learned, as at least one piece of learning data, line drawing information indicating a part (in terms of shape) of a deformed character in a reference image.

[0141] FIG. 11 is a diagram showing the above-described deformed character generation process (in a form different from FIG. 1). In addition to obtaining a deformed character image from a character image (original character image), a process of optimizing the deformed character image is shown.

[0142] 3. Hardware Configuration FIG. 12 is a network configuration diagram showing an overview of the deformed character generation system 1. As shown in FIG. 12, the deformed character generation system 1 in the present embodiment includes a system server 10 (server 10) and a (user) terminal 20. Further, these devices are connected via a network N. The network N is, for example, the Internet or the like. The server 10 includes a deformed character generation program P1, and software (application software) for operating the deformed character generation system 1 according to the present embodiment is installed. Various processes are executed by the functions of this software. Hereinafter, each piece of hardware will be described.

[0143] <Server 10> The server 10 is a computer for executing the deformed character generation program P1. Further, the server 10 displays a website W that provides a user interface for the user.

[0144] Although only one server 10 is shown in FIG. 12, the number is not limited to one, and it may be realized by a plurality of servers. For example, separating a web server and a machine learning server. Multiple servers may also be considered for use from the perspectives of load distribution and availability. In addition, the processing related to image generation may be entrusted to another server or may be performed in cooperation with another server. In addition, the server 10 may use the computer of a cloud service provider, or the user may prepare a computer.

[0145] FIG. 13 is a hardware configuration diagram of the server 10. As shown in FIG. 13, the server 10 includes a control unit 12, a storage unit 14, and a communication control unit 16. The control unit 12 further includes a processor 122, a ROM 124, a RAM 126, and a timer unit 128. The basic functions of each will be described together later.

[0146] The control unit 12 including the processor 122 also functions as a deformed character generation unit in the server 10 (not shown). The deformed character generation unit executes a deformed character generation program P1 to perform deformed character generation processing. In the present embodiment, the processor 122 is a CPU (Central Processing Unit).

[0147] In addition, one program may include another program. For example, in the present embodiment, the deformed character generation program P1 includes an original character image analysis program P11, a reference image analysis program P12, and the like.

[0148] As shown in FIG. 13, the storage unit 14 includes a program storage unit 14a and a data storage unit 14b, and stores programs and data necessary for various processes. For example, in the program storage unit 14a, in addition to the deformed character generation program P1 according to the present embodiment, there are stored a control program for controlling the devices connected to the server 10, such as a communication control program for controlling the communication control unit 16.

[0149] As shown in FIG. 13, the communication control unit 16 is a device that communicates between the server 10 and an external terminal, for example, the terminal 20 described later. As shown in FIG. 12, the communication control unit 16 connects the server 10 to the network N.

[0150] In addition to the above, the server 10 may include an input unit (e.g., a keyboard) for inputting commands and data, an output unit (e.g., a voice output device) for outputting information in some form, etc. (not shown). Further, it may include devices additionally required for the purposes of this embodiment, or devices for improving convenience for the purposes of this embodiment.

[0151] <Terminal 20> The terminal 20 is an information processing device for a user to use the deformable character generation system 1. The user uses the deformable character generation system 1 by accessing the server 10 using the terminal 20. The terminal 20 includes a control unit 22, a storage unit 24, a communication control unit 26, and an input / output unit 28. The control unit also includes a processor, a ROM, a RAM, and a timing unit. The input / output unit 28 includes an input unit 282 and an output unit 284, and the output unit 284 includes a display unit 284a. Descriptions that overlap with the described content and those related to the basic functions described later will be omitted.

[0152] In this embodiment, the terminal 20 is a desktop PC. However, the terminal 20 is not limited to this, and may be a portable terminal such as a smartphone or a tablet.

[0153] In this embodiment, the terminal 20 does not need to install a specific application program, and various processes can be executed by accessing the website W, and the deformable character generation system 1 can be used.

[0154] (Description related to the basic functions of the computer) Hereinafter, the control unit (processor, ROM, RAM, timer unit), the storage unit, the communication control unit, the input unit, and the output unit will be described. Note that, in any of the terminals of this embodiment, the connection mode (network topology) between the functional units is not particularly limited. For example, it may be a bus type, or a star type, a mesh type, or the like.

[0155] The processor performs information processing and control of various devices according to programs stored in the ROM, the storage unit, and the like. In this embodiment, the processor is a CPU (Central Processing Unit).

[0156] However, the processor is not limited to the CPU. Various processors such as a CPU, a DSP (Digital Signal Unit), a GPU (Graphics Processing Unit), a GPGPU (General Purpose computing on GPU), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit) may be used alone or in combination. For example, a processor integrating a CPU and a GPU is called an APU (Accelerated Processing Unit), and such a processor may be used.

[0157] The ROM is a read-only memory in which various programs and data for the processor to perform various controls and operations are stored in advance.

[0158] The RAM is a random access memory used as a working memory for the processor. Various areas for performing various processes of this embodiment can be secured in this RAM.

[0159] The timer unit performs timing processing related to acquisition of time information and the like. When the computer includes a communication control unit, time information may be acquired from the outside by NTP (Network Time Protocol).

[0160] The memory unit is a device for storing information such as programs and data. The memory unit is also referred to as storage. The memory unit can be either built-in or external.

[0161] The memory unit includes a storage medium capable of reading and writing data and a drive for reading and writing to the storage medium. Examples of storage media include built-in and external types, such as hard disks (HD), CD-ROMs, and flash memories. Examples of drives include hard disk drives (HDD) and solid state drives (SSD).

[0162] The memory unit includes a program storage unit and a data storage unit as functional units. The program storage unit stores control programs for controlling various devices, such as a communication control program for controlling communication.

[0163] The communication control unit is a device for performing communication between terminals and the like. The communication control unit connects a terminal including the communication control unit to the network N.

[0164] The communication method of the communication control unit is a known method, and a wired or wireless method is applied according to the device. For example, if the terminal is a desktop PC, both wired and wireless cases are considered. If the terminal is a smartphone, a wireless communication method is considered.

[0165] If it is wired, for example, a communication method defined by IEEE802.3 (such as a bus-type or star-type wired LAN) can be preferably used. However, in addition, a communication method defined by IEEE802.5 (such as a ring-type wired LAN) can also be used.

[0166] If it is wireless, for example, a communication method defined by IEEE802.11 (e.g., Wi-Fi) can be preferably used. However, in addition to that, IEEE802.15 (e.g., Bluetooth (registered trademark), BLE (Bluetooth (registered trademark) Low Energy), etc.), IEEE802.16 (e.g., WiMAX), ZigBee (registered trademark), 920MHz band wireless (Wi-SUN, etc.), or a communication method defined by optical communication such as infrared communication may also be used.

[0167] The input unit and the output unit are devices that respectively handle input to and output from the terminal. The input unit and the output unit may be collectively referred to as the input / output unit. The input unit is a device that receives input from the user. Examples of such an input unit include a keyboard, a mouse as a pointing device, a trackpad, a tablet, or a touch panel.

[0168] When the terminal is a tablet, a smartphone, etc., and the input unit is a touch panel, the input unit is arranged on the surface of a display unit such as a touch screen that displays an image, etc. In this case, the input unit identifies the touch position of the user corresponding to various operation icons displayed on the display unit and receives the input from the user.

[0169] The output unit is, for example, a device for outputting an image, voice, document, etc. Examples of the output unit include a display device (display unit) such as a touch screen or a display (liquid crystal display or organic EL display), a voice output device such as a speaker, and a document output device such as a printer.

[0170] With the above configuration, the user can obtain a desired deformed character online easily.

[0171] (Second Embodiment) The second embodiment is an embodiment that generates only the face part of the deformed character, not the whole body.

[0172] FIG. 14 is a diagram showing an outline of the second embodiment. In this embodiment, the deformed character generation program P1 takes as input the original character image in the upper left of FIG. 14 and the reference image (face part) in the lower left, and outputs the face part of the deformed character in the lower right. In this embodiment, the "face" may include the entire head, for example, decorations such as a hat.

[0173] At this time, the UI, processing (deformed character generation processing and optimization processing), and hardware configuration are basically the same as those in the first embodiment.

[0174] However, in the processing and data of this embodiment, it is only necessary to be able to extract the feature amount of the head (face) from the original character image. For example, in the above-described plurality of feature amount acquisition processes, it is only necessary to be able to acquire an image of a region including the face. The character prompt is character information explaining the feature amount of the face part, and the character model is a model for obtaining a deformed character image (deformed character face image) of the face part.

[0175] Also, in the control model acquisition process, in the above-described embodiment, the processor 122 acquired pose information (Pose) and line drawing information (line drawing information called Scribble) indicating a part of the shape of the deformed character, but in this embodiment, only the line drawing information indicating a part of the shape of the deformed character is sufficient.

[0176] As shown in FIG. 14, in this embodiment, the processor 122 acquires line drawing information of the eyes, mouth, jaw part, and neck part as line drawing information indicating a part of the shape of the deformed character.

[0177] However, feature amounts may be acquired from information other than the line drawing information, and the feature amounts may be used for the control model (input). Examples of such feature quantities include feature quantities based on information such as the distance between the right and left eyes in the original character image, or information indicating the relationship between both eyes and the nasolabial fold (such as the crosshair in a drawing, so-called "hit"), that is, feature quantities obtained from the positional relationship of facial parts (such as eyes, nose, or mouth). By adding additional information, a deformed character image that better captures facial features can be obtained.

[0178] To summarize, in the first embodiment, the character image may be a character's face image, and the deformed character image may be a deformed character face image.

[0179] In this way, the deformed character generation program P1 can generate not only a full-body image of the deformed character but also deform the head (face) part, which is a part of the whole body. There is a high need for generating a face image of a deformed character, such as when a character's face icon is used as a profile image on SNS. The deformed character generation program P1 can also meet such needs.

[0180] FIG. 15 is a diagram showing an application example of deformation by the deformed character generation program P1 of the present embodiment (the first embodiment and the second embodiment). From a single original character image, a plurality of deformed images of the face or the whole body can be generated.

[0181] (Modification example) The present invention is not limited to the above-described embodiments, and includes those obtained by making various changes to the above-described embodiments without departing from the spirit of the present invention.

[0182] For example, in the above-described embodiment, the machine learning model (diffusion model) was created each time by updating the parameters, but the learning and creation of the machine learning model are not limited to this. For example, not only the update of the parameters, but also those that change the model structure or those that are created from the model structure may be used. However, updating the parameters is more preferable because the process is faster.

[0183] Also, as described above, you may obtain an already created machine learning model (diffusion model). Further, the created character model may be integrated (incorporated) into an (existing / other) Stable Diffusion diffusion model. However, by obtaining and analyzing the original character image and updating the parameters of the diffusion model to obtain a new diffusion model, the reproducibility of the original character image in the deformed character image is significantly improved.

[0184] In the above-described embodiment, the machine learning model (character model) included a diffusion model and a control model, but it is not limited to this. First, it may be composed of only a diffusion model. Also, as long as the object of the present invention is achieved, it may be composed of a machine learning model other than the diffusion model. However, the aspect including the diffusion model and the control model can generate a deformed character image that accurately reproduces the feature amount of the original character image.

[0185] In order to realize feature decoupling, the following approach may be used. In the above-described embodiment, the following (1) is used. (1) Separation of the latent space Design the latent space of the generative model and train it so that different dimensions control different feature amounts. For example, each dimension is in charge of "hairstyle", "face (expression)", or "pose", etc. (2) Architecture design Design the layers and modules of generative networks (such as GANs and diffusion models) so that specific layers process specific feature quantities. For example, "multi-scale feature extraction" and "style application at the layer level" for style conversion, etc. (3) Regularization of learning To maintain the independence between feature quantities, introduce a penalty to suppress the interaction between feature quantities during the learning process. (4) Conditioning between modules By adjusting the generation process based on specific conditions (text, attributes, sketches), emphasize specific feature quantities.

[0186] Aspects of the present invention including this embodiment, in other words, have the following feature quantities. The following corresponds to the scope of claims at the time of filing this application. However, due to amendments to the scope of claims after filing, it may be different from the description of the scope of claims after such amendments. (1) In the first aspect, a computer is caused to function as (1) original character image analysis means, (2) reference image analysis means, and (3) deformed character acquisition means. The (1) original character image analysis means includes original character image acquisition means for acquiring an original character image that is an image of a character, multiple feature quantity acquisition means for acquiring a plurality of feature quantities from the original character image, character prompt acquisition means for acquiring character information that describes the original character image from the original character image, and character model acquisition means for acquiring a machine learning model that has learned at least a part of the feature quantities other than the pose feature quantity among the plurality of feature quantities as learning data. The (2) reference image analysis means includes reference image acquisition means for acquiring a reference image of a deformed character, and control model acquisition means for acquiring a control model that has learned data based on at least line drawing information indicating at least a part of the deformed character as learning data. The (3) deformed character acquisition means is characterized by acquiring a deformed character image obtained by deforming the original character image from the machine learning model controlled by the control model and the character information, and provides a deformed character generation program. (2) In the second aspect, a deformed character generation program according to the first aspect is provided, characterized in that the control model that has learned data based on at least line drawing information indicating at least a part of the deformed character as learning data is a control model that has learned data based on at least line drawing information indicating at least a part of the deformed character and pose information as learning data. In this case, since the amount of information regarding the deformed character increases, it becomes easier to obtain a deformed character image with a pose or the like closer to the input reference image. (3) In a third aspect, further provided is an optimization means for obtaining and evaluating differences between a plurality of output images obtained by changing weights in at least a part of the layers of the neural network included in the machine learning model, and reducing or deleting weights of layers that have little influence on the output images, and the deformation character generation program according to the first aspect is provided. In this case, a deformed character image with a higher degree of reproduction of the original character image can be obtained. (4) In a fourth aspect, provided is the deformation character generation program according to the first aspect, characterized in that the image of the character is a face image of the character, and the deformed character image is a deformed character face image. In this case, a deformed character image for the (upward) image of the face can be obtained. (5) A deformation character generation system is provided, comprising: (1) an original character image analysis unit, (2) a reference image analysis unit, and (3) a deformed character acquisition unit. The (1) original character image analysis unit includes an original character image acquisition unit for acquiring an original character image that is an image of a character, a plurality of feature quantity acquisition units for acquiring a plurality of feature quantities from the original character image, a character prompt acquisition unit for acquiring character information explaining the original character image from the original character image, and a character model acquisition unit for acquiring a machine learning model that has learned at least a part of the feature quantities other than the pose feature quantity among the plurality of feature quantities as learning data. The (2) reference image analysis unit includes a reference image acquisition unit for acquiring a reference image of a deformed character, and a control model acquisition unit for acquiring a control model that has learned data based on line drawing information indicating at least a part of the deformed character as learning data. The (3) deformed character acquisition unit is characterized in that it acquires a deformed character image obtained by deforming the original character image from the machine learning model controlled by the control model and the character information. (6) In the sixth aspect, there is provided a deformed character generation method comprising: (1) an original character image analysis step, (2) a reference image analysis step, and (3) a deformed character acquisition step. The (1) original character image analysis step includes an original character image acquisition step in which a processor acquires an original character image that is an image of a character, a plurality of feature quantity acquisition steps in which the processor acquires a plurality of feature quantities from the original character image, a character prompt acquisition step in which the processor acquires character information describing the original character image from the original character image, and a character model acquisition step in which the processor acquires a machine learning model that has learned at least a part of the feature quantities other than the pose feature quantity among the plurality of feature quantities as learning data. The (2) reference image analysis step includes a reference image acquisition step in which a processor acquires a reference image of a deformed character, and a control model acquisition step in which the processor acquires a control model that has learned data based on line drawing information indicating at least a part of the deformed character as learning data. The (3) deformed character acquisition step is characterized in that the processor acquires a deformed character image obtained by deforming the original character image from the machine learning model controlled by the control model and the character information. (7) In the seventh aspect, there is provided a learned model that acquires a deformed character image obtained by deforming an original character image, which is an image of a character. The learned model acquires a plurality of feature amounts from the original character image, and uses, as learning data, data including at least a feature amount of a concept including feature amounts other than the feature amount of the pose among the plurality of feature amounts, noise added to the feature amount of the concept, and noise-containing information obtained by adding the noise. The learned model learns the noise to be removed from the noise-containing information, and takes the noise-containing information as an input and outputs the noise to be removed from the noise-containing information. Further, the pose in the deformed character image output by a control model that has learned, as at least one learning data, line drawing information indicating a part of the deformed character in a reference image is controlled. A learned model is provided.

Industrial Applicability

[0187] Cute characters are in demand throughout the industry (especially the entertainment industry) regardless of online or offline. Therefore, a program capable of generating such characters enables the provision of characters that meet the needs.

Explanation of Signs

[0188] 1 Deformed Character Generation System 10 Server 12 Control Unit 122 Processor 124 ROM 126 RAM 128 Timing Unit 14 Storage Unit 14a Program Storage Unit 14b Data Storage Unit 16 Communication Control Unit 18 Input / Output Unit 20 Terminal 22 Control Unit 24 Storage Unit 26 Communication Control Unit 28 Input / Output Unit 282 Input Unit 284 Output unit 284a Display unit UI-11 Upload field UI-12 Deformed character setting field UI-121 Life-size ratio setting button UI-122 Referenced deformed character image UI-123 Deformed character generation start button UI-13 Original character image confirmation screen (window) UI-131 Image information UI-132 Original character image UI-133 OK button icon UI-14 Output image display field UI-141 Output image UI-142 Output image processing related buttons UI-143 Output image information UI-144 Output image saving related buttons P1 Deformed character generation program P11 Original character image analysis program P12 Referenced image analysis program P13 Deformed character acquisition program P14 Optimization program

Claims

1. A computer is caused to function as (1) an original character image analyzing means, (2) a reference image analyzing means, and (3) a deformed character acquiring means; The (1) original character image analysis means an original character image acquiring means for acquiring an original character image which is an image of a character; a multiple feature amount acquiring means for acquiring multiple feature amounts from the original character image; a feature amount separating means for separating the plurality of feature amounts into pose feature amounts and other feature amounts; A character prompt acquisition means for acquiring character information explaining the original character image from the original character image; a character model acquisition means for acquiring a diffusion model in which at least a part of the feature amounts other than the feature amount of the pose is learned as learning data, among the plurality of feature amounts; The (2) reference image analysis means A reference image acquisition means for acquiring a reference image of a deformed character; a control model acquisition means for acquiring a control model that is learned as learning data based on line drawing information showing at least a part of the deformed character; The (3) deformed character acquisition means is characterized in that it acquires a deformed character image that deforms the original character image from the diffusion model controlled by the control model and the character information.

2. A control model that learns data based on line drawing information showing at least a part of the deformed character as learning data, A control model that learns data based on line drawing information and pose information showing at least a part of a deformed character as learning data.

2. The deformed character generating program according to claim 1.

3. The deformed character generation program according to claim 1, further comprising an optimization means for obtaining and evaluating differences between multiple output images obtained by changing weights in at least some layers of the neural network provided in the diffusion model, and reducing or deleting weights in layers that have little effect on the output image.

4. the image of the character is a face image of the character, 2. The deformed character generating program according to claim 1, wherein the deformed character image is a deformed character face image.

5. (1) an original character image analysis unit; (2) a reference image analysis unit; and (3) a deformed character acquisition unit; The (1) original character image analysis unit an original character image acquisition unit that acquires an original character image which is an image of a character; a multiple feature amount acquisition unit for acquiring multiple feature amounts from the original character image; a feature amount separation unit that separates the plurality of feature amounts into pose feature amounts and other feature amounts; a character prompt acquisition unit that acquires, from the original character image, character information explaining the original character image; a character model acquisition unit that acquires a diffusion model in which at least a part of the feature amounts other than the feature amount of the pose is learned as learning data, among the plurality of feature amounts; The (2) reference image analysis unit a reference image acquisition unit for acquiring a reference image of a deformed character; a control model acquisition unit that acquires a control model that is trained using data based on line drawing information showing at least a part of the deformed character as training data; The (3) deformed character acquisition unit is characterized in that it acquires a deformed character image that deforms the original character image from the diffusion model controlled by the control model and the character information.

6. A method for generating a deformed character, comprising: (1) an original character image analysis step; (2) a reference image analysis step; and (3) a deformed character acquisition step, The (1) original character image analysis step includes: an original character image acquisition step in which the processor acquires an original character image which is an image of the character; a multiple feature amount acquiring step in which a processor acquires multiple feature amounts from the original character image; a feature separation step in which a processor separates the plurality of feature amounts into pose feature amounts and other feature amounts; a character prompt acquisition step in which a processor acquires, from the original character image, text information describing the original character image; and a character model acquisition step of acquiring a diffusion model trained using at least a part of the feature amounts other than the feature amount of the pose as training data, The reference image analysis step (2) includes: A reference image acquisition step in which a processor acquires a reference image of a deformed character; a control model acquisition step of acquiring a control model which is learned using data based on line drawing information showing at least a part of the deformed character as learning data by a processor; The (3) deformed character acquisition step is characterized in that a processor acquires a deformed character image that deforms the original character image from the diffusion model controlled by the control model and the character information.

Citation Information

Patent Citations

  • Arithmetic unit, arithmetic method, and learning method

    JP2021047711A

  • Neural network optimizing method, neural network optimizing device and program

    JP2021105950A

  • Image conversion device, image conversion method, and image conversion program

    JP2023157334A

  • Information processing device, information processing method, and information processing program

    JP2023169048A

  • Image acquisition device, image acquisition method, program, and recording medium

    WO2024161839A1