Program, and image processing device

By pre-processing input images to combine contour and detail images before inputting them into a LoRA-adjusted generative model, unintended image generation is minimized, ensuring accurate representation of object features and styles.

JP2025144787APending Publication Date: 2025-10-03BROTHER KOGYO KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024044633
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Machine learning models, such as generative adversarial networks, may generate unintended images, particularly when transitioning an input image to a different style, leading to outputs that do not accurately represent the intended content and features.

Method used

A method involving pre-processing input images to combine contour and detail images, generating a composite image that is then input into a trained machine learning model, specifically using LoRA-adjusted generative models like Stable Diffusion, to ensure accurate representation of object contours and features in the output.

Benefits of technology

Reduces the likelihood of generating images that do not accurately depict the intended object contours and features by using a composite image pre-processing technique, enhancing the model's ability to produce consistent style transitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025144787000001_ABST
    Figure 2025144787000001_ABST
Patent Text Reader

Abstract

To reduce possibility that an unintended image is acquired from a machine learning model.SOLUTION: A third image is generated by synthesizing a first image and a second image. A new image is acquired by inputting information including the third image to a trained machine learning model. Here, the following embodiments may be adopted. That is, an object image showing an object is acquired. A contour image showing the contour of the object, and a fine portion image showing a finer feature of the object than the contour image are acquired. A composite image is generated by synthesizing the plurality of images including the contour image and the fine portion image. The new image is acquired by inputting the information including the composite image to the trained machine learning model.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification relates to techniques for using images to generate new images. [Background technology]

[0002] Various machine learning models, such as a diffusion model, a generative adversarial network, and an autoencoder, can be used to acquire new images. The machine learning model can use an image input to the machine learning model to generate various images based on the input image. For example, the machine learning model can generate an image that represents the same content as the content of the input image in a specific style that is different from the style of the input image. Here, parameters of the trained machine learning model can be adjusted to be suitable for generating images of a specific style. For example, a technique called Low-Rank Adaptation of Large Language Models (LoRA) can be used as a parameter adjustment technique. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, "LORA: LOW-RANK ADAPTATION OF LARGE LANGUAGE MODELS", arXiv:2106.09685, October 16, 2021, http: / / arxiv.org / abs / 2106.09685 Summary of the Invention [Problem to be solved by the invention]

[0004] A machine learning model may output unintended results, for example, a machine learning model may generate unintended images.

[0005] This specification discloses techniques for reducing the likelihood of unintended images being obtained from machine learning models. [Means for solving the problem]

[0006] The techniques disclosed in this specification can be implemented in the following application examples.

[0007] [Application Example 1] A program that causes a computer to realize the following functions: a first acquisition function that acquires a target image representing an object; a second acquisition function that acquires a contour image that represents the contour of the object and a detail image that represents finer features of the object than the contour image; a synthesis function that generates a composite image by synthesizing multiple images including the contour image and the detail image; and a third acquisition function that acquires a new image by inputting information including the composite image into a trained machine learning model.

[0008] According to this configuration, a composite image generated by combining multiple images including a contour image and a detail image is input into the machine learning model, thereby reducing the possibility of obtaining an image that does not represent the contour and features of the object.

[0009] [Application Example 2] A program that causes a computer to perform the following functions: generate a third image by synthesizing a first image and a second image; and obtain a new image by inputting information including the third image into a trained machine learning model.

[0010] According to this configuration, a third image generated by synthesizing the first image and the second image is input into the machine learning model, thereby reducing the possibility of obtaining an image that does not reflect the features represented by the first image and the features represented by the second image.

[0011] The technology disclosed in this specification can be realized in various forms, such as an image processing method and an image processing device, a computer program for realizing the functions of the method or device, a recording medium (e.g., a non-temporary recording medium) on which the computer program is recorded, and the like. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is an explanatory diagram illustrating an image processing device according to an embodiment; [Figure 2] FIG. 9 is a block diagram illustrating an example of a generative model 900. [Figure 3] FIG. 10 is a diagram illustrating an example of an image generated using a first adjustment parameter 990a. [Figure 4] 10 is a flowchart illustrating an example of image processing. [Figure 5] FIG. 10 is a diagram illustrating an example of an image to be processed by image processing. [Figure 6] 10 is a flowchart illustrating an example of preprocessing. [Figure 7] FIG. 10 is a diagram illustrating an example of an image to be processed in preprocessing. [Figure 8] FIG. 10 is a diagram illustrating an example of an image obtained by image processing using a sample image. [Figure 9] 10 is a flowchart illustrating another embodiment of pre-processing. [Figure 10] FIG. 10 is a diagram illustrating an example of an image to be processed in preprocessing. [Figure 11] FIG. 10 is a diagram illustrating an example of an image to be processed by image processing. [Figure 12] 10 is a flowchart illustrating another embodiment of pre-processing. [Figure 13] FIG. 10 is a diagram illustrating an example of an image to be processed in preprocessing. [Figure 14] FIG. 10 is a diagram illustrating an example of an image to be processed by image processing. [Figure 15] 10 is a flowchart illustrating another embodiment of pre-processing. [Figure 16]FIG. 10 is a diagram illustrating an example of an image to be processed in preprocessing. [Figure 17] FIG. 10 is a diagram illustrating an example of an image to be processed by image processing. DETAILED DESCRIPTION OF THE INVENTION

[0013] A. First Example: A1.Device configuration: 1 is an explanatory diagram showing an image processing apparatus according to an embodiment. The image processing apparatus 200 is, for example, a personal computer. The image processing apparatus 200 uses an image to obtain a new image.

[0014] The image processing device 200 includes a processor 210, a storage device 215, a display unit 240, an operation unit 250, a graphics processing unit 260 (referred to as GPU 260), and a communication interface 270. These elements are connected to each other via a bus. The storage device 215 includes a volatile storage device 220 and a non-volatile storage device 230.

[0015] The processor 210 is a device configured to perform data processing, and is, for example, a central processing unit (CPU) or a system on a chip (SoC). The volatile storage device 220 is, for example, a dynamic random access memory (DRAM), and the non-volatile storage device 230 is, for example, a flash memory. The non-volatile storage device 230 stores data for a program 231, a segmentation model 800, and a generative model 900. The segmentation model 800 and the generative model 900 are each program modules that form trained machine learning models. Details of the data stored in the non-volatile storage device 230 will be described later.

[0016] The display unit 240 is a device configured to display images, such as a liquid crystal display or an organic EL display. The operation unit 250 is a device configured to receive operations by a user, such as a button, a lever, or a touch panel overlaid on the display unit 240. The display unit 240 and the operation unit 250 may form a so-called touch screen. The user can input various requests and instructions to the image processing device 200 by operating the operation unit 250. The display unit 240 may display operation elements (e.g., buttons, sliders, etc.), and the displayed elements may be operated through operation of the operation unit 250.

[0017] The GPU 260 is a computing device configured to perform various numerical calculations such as image processing and machine learning. The GPU 260 performs various calculations in accordance with instructions from the processor 210. A driver program (not shown) for controlling the GPU 260 may be provided by the manufacturer of the GPU 260.

[0018] The communication interface 270 is an interface for communicating with other devices (for example, it includes one or more of a USB interface, a wired LAN interface, an IEEE802.11 wireless interface, and an industrial camera interface (for example, CameraLink, CoaXPress, etc.)).

[0019] A2. Generative Model 900: FIG. 2 is a block diagram illustrating an example of a generative model 900. The generative model 900 may be any of various models that use input image data to generate output image data based on the input image. In this embodiment, the generative model 900 is a model in which parameters of a machine learning model called Stable Diffusion are adjusted using a technique called LoRA (Low-Rank Adaptation). The generative model 900 includes a diffusion model 960, which is a Stable Diffusion model, and adjustment parameters 990a and 990b.

[0020] Stable Diffusion is a model that synthesizes high-resolution images using a latent diffusion model. The technology for synthesizing high-resolution images using a latent diffusion model is disclosed in, for example, the following paper: Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjoern Ommer, "High-Resolution Image Synthesis with Latent Diffusion Models", arXiv:2112.10752, April 13, 2022, http: / / arxiv.org / abs / 2112.10752

[0021] Data for the pre-trained Stable Diffusion model is made publicly available on the Internet by Stability AI. In this embodiment, the data for the published pre-trained model is used as the data for the diffusion model 960. The diffusion model 960 includes a text encoder 910, an image encoder 920, a latent variable model 930, and an image decoder 940. The text encoder 910 is configured to convert text Ptx (also called a prompt) into a vector tv (such a vector tv is also called a text embedding). The image encoder 920 is configured to convert an input image IMi into a latent variable lvi. The latent variable model 930 is configured to perform a noise addition process that adds noise to the latent variable lvi and a process that outputs a processed latent variable lvo by performing a de-diffusion process that removes noise from the noisy latent variable. The latent variable model 930 includes a neural network called a U-Net (not shown) for performing the de-diffusion process. The image decoder 940 is configured to generate an output image IMo using the latent variable lvo. The latent variable model 930 uses the vector tv obtained from the text encoder 910 for conditioning in the de-diffusion process. By executing such a de-diffusion process, the latent variable model 930 can generate a latent variable lvo that is associated with an image conditioned by the text Ptx. The text encoder 910 is pre-trained so that the vector tv obtained from the text Ptx and the image represented by the text Ptx can be associated. As the text encoder 910, an encoder pre-trained by a technology called CLIP is used. CLIP is a technology released by OpenAI.

[0022] The diffusion model 960 can generate a new output image IMo represented by the text Ptx by performing a de-diffusion process using randomly generated noise using the vector tv from the text encoder 910. This technique is also called txt2img. The diffusion model 960 can also generate an output image IMo by modifying the input image IMi according to the text Ptx using the input image IMi and the text Ptx. This technique is also called img2img.

[0023] The adjustment parameters 990a and 990b are configured to fine-tune parameters (e.g., weights, biases, etc.) of the latent variable model 930. For example, parameters used in the dediffusion process of the latent variable model 930 are fine-tuned. The fine-tuned parameters may include, for example, parameters of some of the layers included in the U-Net. In this embodiment, the parameters of the latent variable model 930 are fine-tuned so that the generative model 900 generates an output image IMo in which the style of the input image IMi is converted into a specific style (e.g., line art, anime art, etc.). In other words, the fine-tuned generative model 900 can generate an output image IMo that represents the same content (e.g., a person) as the input image IMi but in a style different from the style of the input image IMi. Such a generative model 900 is also referred to as a style transfer model.

[0024] The parameter adjustment method may be various methods. In this embodiment, a technique called LoRA is used. LoRA is a technique for adjusting some parameters of a pre-trained model. LoRA is disclosed in, for example, the above-mentioned paper on LoRA.

[0025] LoRA trains a set of parameter differences without changing the parameters of the pre-trained model. A fine-tuned model is formed by combining the pre-trained model and the difference set trained by LoRA. A pre-trained model can be commonly used to prepare multiple difference sets for multiple tasks. The task of the generative model 900 can be easily switched by replacing the difference set with another difference set.

[0026] The adjustment parameters 990a and 990b each represent a trained difference set. In this embodiment, the first adjustment parameters 990a are a difference set for generating an output image IMo of a line drawing. A line drawing is an image that contains lines (including contours and boundaries) representing objects and is free of shading and color. The first adjustment parameters 990a are trained using multiple line drawing images IMTa so that the generative model 900 generates line drawing images according to the LoRA technique. When the first adjustment parameters 990a are used, the generative model 900 can, for example, generate an output image IMo1a of a line drawing representing a person from an input image IMi1 representing a photograph of the same person.

[0027] The second adjustment parameters 990b are a difference set for generating an output image IMo of animated art. Animated art includes lines (including contours and borders) that represent objects and simplified color gradations. For example, a region, such as a person's skin or clothing, is colored with a small number of colors (e.g., one, two, or three). Animated art is also called a solid-color image. The second adjustment parameters 990b are used to train the generative model 900 to generate an image of animated art according to LoRA technology using multiple images IMtb of animated art. When the second adjustment parameters 990b are used, the generative model 900 can generate an output image IMo1b of animated art representing a person from an input image IMi1 that represents a photograph of the same person.

[0028] When the adjustment parameters 990a and 990b are used, the text Ptx may include various texts that represent the image to be generated. For example, the text Ptx may include text that represents the style of the image to be generated (e.g., line art, anime art, etc.). Furthermore, when training the adjustment parameters 990a and 990b, text Ptx that includes specific text associated with the adjustment parameters may be used. When generating the output image IMo using the adjustment parameters, the text Ptx may include the specific text that was used when training the adjustment parameters. Note that input of the text Ptx may be omitted.

[0029] FIG. 3 is a diagram illustrating an example of an image generated using the first adjustment parameters 990a. Image IM10 represents an example of an input image input to the generative model 900, and image IM10o represents an example of an output image generated by the generative model 900. In this embodiment, the data of each of images IM10 and IM10o is bitmap data representing color values ​​using three color components: red (R), green (G), and blue (B). Images IM10 and IM10o are rectangular images having two sides parallel to a first direction Dx and two sides parallel to a second direction Dy perpendicular to the first direction Dx. Images IM10 and IM10o are represented by the color values ​​of each of a plurality of pixels arranged in a matrix along the first direction Dx and the second direction Dy (the color values ​​represent the gradation values ​​of each of red (R), green (G), and blue (B) (e.g., values ​​greater than or equal to zero and less than or equal to 255)). The number of pixels in the first direction Dx and the number of pixels in the second direction Dy accepted by the generative model 900 are each predetermined and are common to the input image IM10 and the output image IM10o. Hereinafter, the above data format of the image data accepted by the generative model 900 will be referred to as the processed data format.

[0030] In the example of FIG. 3, the input image IM10 represents a color photograph showing an object OB and a background BG. The object OB is a person. The region representing the object OB includes a facial skin region P1 representing the skin of the face, a body skin region P2 representing the skin of the body (excluding the face), a hair region P3, and a clothing region P4. Each of the face, body, hair, and clothing is a type of object. In this way, an object may include multiple parts (i.e., multiple objects).

[0031] The processor 210 uses the input image IM10 to execute the operations of the generative model 900 to generate an output image IM10o of a line drawing (note that the processor 210 may cause the GPU 260 to execute some or all of the operations of the generative model 900). The generative model 900 may output unintended results. For example, the boundary between the hair region P3 and the background BG may be blurred in the input image IM10. In this case, the generative model 900 may not be able to generate an image that includes the boundary line between the background BG and the hair region P3. Furthermore, in the input image IM10, the background BG and the object OB may be represented in various colors. In this case, the generative model 900 may generate an image that represents the shading and color of each region.

[0032] Output image IM10o in Figure 3 shows an example of an unintended result. Output image IM10o shows object OBz, which is similar to object OB in input image IM10. Output image IM10o shows regions P1z, P2z, P3z, P4z, and BGz, which correspond to regions P1, P2, P3, P4, and BG in input image IM10, respectively. In output image IM10o, a portion of the boundary between hair region P3z and background BGz is missing. In output image IM10o, facial skin region P1z, body skin region P2z, hair region P3z, clothing region P4z, and background BGz show grayscale color gradations.

[0033] In this embodiment, in order to reduce the possibility of obtaining unintended images from the generative model 900, the processor 210 performs pre-processing of the images to be input to the generative model 900.

[0034] A3.Image Processing: FIG. 4 is a flowchart illustrating an example of image processing. The processor 210 of the image processing device 200 (FIG. 1) executes image processing according to the program 231 in response to an image processing start instruction input to the image processing device 200. The start instruction may be input by any method. In this embodiment, the user inputs the start instruction by operating the operation unit 250. The start instruction may include data information specifying input image data to be used in the image processing. The data information may specify image data stored in various storage devices. The storage device may be selected from, for example, the storage device 215 (e.g., the non-volatile storage device 230), a storage device (not shown) connected to the communication interface 270 (e.g., a USB flash drive), or a storage device of a server capable of communicating with the image processing device 200. The user may input the start instruction and input image data to the image processing device 200 via a terminal device (not shown) (e.g., a smartphone) capable of communicating with the image processing device 200.

[0035] In S110, the processor 210 acquires data of the target image in accordance with the start instruction. The processor 210 stores the acquired data of the target image in the storage device 215 (in this embodiment, the non-volatile storage device 230). If the data format of the input image data is different from the processing data format accepted by the generative model 900, the processor 210 acquires the target image data by converting the format of the input image data into the processing data format. For example, if the data format of the input image data is different from a bitmap format (e.g., a data format described in a page description language), the processor 210 acquires the target image data in the processing data format by rasterizing the input image data. If the data format of the input image data is a bitmap format (e.g., a JPEG format), the processor 210 acquires the target image data by converting the resolution of the input image data (i.e., the number of pixels in the first direction Dx and the number of pixels in the second direction Dy) to the resolution of the processing data format. If the resolution of the input image data is the same as the resolution of the processing data format, the processor 210 may use the input image data as the target image data as is.

[0036] 5 is a diagram showing an example of an image to be processed by image processing. Image IM10 in the figure shows an example of a target image. Hereinafter, the target image is assumed to be the same as input image IM10 in FIG. 3 (image IM10 will be referred to as target image IM10).

[0037] In S120 (FIG. 4), processor 210 executes preprocessing. FIG. 6 is a flowchart showing an example of preprocessing. In the figure, symbols beginning with "S" indicate steps. Symbols beginning with "IM" or "pM" indicate images, which will be described later. The symbol of an image attached to a box corresponding to each step indicates the image acquired or generated by that step. For example, the symbol IM10 is attached to the box corresponding to S210. This indicates that image IM10 is acquired in S210. The same applies to the flowcharts of preprocessing in other embodiments, which will be described later.

[0038] In this embodiment, the processor 210 executes a detailed image acquisition process PA and a contour image acquisition process PB. The detailed image acquisition process PA includes S210 and S222-S228. The contour image acquisition process PB includes S210 and S232-S238. S210 is common to the detailed image acquisition process PA and the contour image acquisition process PB (S210 may be executed once for these processes PA and PB). The processor 210 performs these processes PA and PB in parallel or in a parallel manner. Alternatively, the processor 210 may execute these processes PA and PB sequentially.

[0039] First, the detailed image acquisition process PA will be described. In S210, the processor 210 reads data of the target image from the nonvolatile storage device 230.

[0040] In S222, the processor 210 generates grayscale image data by performing grayscale processing on the target image IM10. This processing converts RGB color values ​​into luminance values ​​using a predetermined relational expression (e.g., a color conversion expression from the RGB color space to the YCbCr color space). FIG. 7 is a diagram showing an example of an image processed in preprocessing. The image pM22 on the left of the first row is an example of a grayscale image generated in S222. The grayscale image pM22 represents the object OB, similar to the target image IM10 (FIG. 5).

[0041] In S224 (FIG. 6), the processor 210 generates noise-reduced grayscale image data by performing a blurring process on the grayscale image generated in S222. The blurring process may be any of various processes that smooth color values. In this embodiment, the blurring process is a smoothing process that uses a Gaussian filter. Alternatively, various smoothing filters such as a mean filter or a median filter may be used. Although not shown, the blurring process makes fine edges (e.g., noise) that are not characteristic of the object OB in the grayscale image pM22 (FIG. 7) less noticeable.

[0042] In S226 (FIG. 6), the processor 210 generates edge image data representing the fine features of the object OB by performing edge detection processing on the grayscale image processed in S224. In this embodiment, the processor 210 performs so-called Canny edge detection. The image pM26 on the left of the second row in FIG. 7 represents an example of the edge image generated in S226. In the edge image pM26, edge pixels representing edges are represented by large pixel values ​​(e.g., the maximum value of 255), and non-edge pixels not representing edges are represented by small pixel values ​​(e.g., the minimum value of zero). In this way, in this embodiment, the edge image pM26 is generated, resembling a negative image of a photograph. The edge pixels in the edge image pM26 can represent the fine features of the object OB (details will be described later). Note that the edge detection processing may be any of various processes for detecting edge pixels in an image. For example, the processor 210 may calculate the edge strength of each pixel using a filter that calculates edge strength, such as a Laplacian filter or a Sobel filter, and detect pixels whose edge strength is greater than a threshold as edge pixels. Alternatively, a machine learning model trained to detect edges may be used (for example, a model called informative-drawings).

[0043] In S228 (FIG. 6), the processor 210 generates edge image data, such as a positive image of a photograph, by inverting pixel values ​​(here, brightness values) of the edge image generated in S226. The pixel values ​​Va before inversion are converted to inverted pixel values ​​Vb according to a predetermined relational expression. The relational expression may be, for example, Vb = maximum value (here, 255) - Va. The left image pM28 in the third row of FIG. 7 represents an example of the edge image generated in S228. In the edge image pM28, edge pixels representing edges are represented by small pixel values ​​(e.g., the minimum value, zero), and non-edge pixels not representing edges are represented by large pixel values ​​(e.g., the maximum value, 255). In the example of FIG. 7, the edge image pM28 represents small features of the object OB, such as the eyebrows, eyes Pe, nose Pn, mouth Pm, multiple hairs, and the collar of the clothes. Note that if the boundaries between multiple regions in the target image IM10 (FIG. 5) are blurred, the boundaries may not be detected by the edge detection process (S226). For example, a portion of the outline of the hair region P3 (e.g., the portion of the outline of the hair region P3 to the left of the facial skin region P1) is missing from the edge image pM28. The edge image pM28 is an example of a detail image that represents the fine features of the object represented by the target image (hereinafter, the edge image pM28 will also be referred to as the detail image pM28).

[0044] This completes the detailed image acquisition process PA (FIG. 6). Next, the outline image acquisition process PB will be described. As described above, S210 is common to processes PA and PB. When S210 is executed for the detailed image acquisition process PA, S210 for the outline image acquisition process PB may be omitted.

[0045] In S232, the processor 210 performs a segmentation process on the target image IM10. The segmentation process divides the image into multiple regions, each representing a portion of one or more objects represented by the image. In this embodiment, the processor 210 performs the segmentation process by using a segmentation model 800 ( FIG. 1 ). The segmentation model 800 may be any of a variety of models that perform segmentation processes. In this embodiment, a pre-trained model called "Multi-class selfie segmentation mode" included in a library called "MediaPipe" provided by Google is used as the segmentation model 800. This model captures an image of a person, identifies the background, hair, body (skin), face (skin), clothing, and other (accessories) regions, and outputs an image segmentation map representing each identified region. The processor 210 generates segmentation map data by executing the operations of the segmentation model 800 using the input image IM10 (note that the processor 210 may cause the GPU 260 to execute some or all of the operations of the segmentation model 800). The image pM32 on the right of the first row in FIG. 7 represents an example of a segmentation map. The segmentation map pM32 represents a facial skin region P1, a body skin region P2, a hair region P3, a clothing region P4, and a background BG in different colors.

[0046] In S234 (FIG. 6), the processor 210 performs a region contour extraction process on the segmentation map generated in S232. The region contour extraction process may be any of various processes that extract the contours of each of the multiple regions represented by the segmentation map. In this embodiment, the processor 210 extracts the contours using boundary tracking. The algorithm for this process is disclosed in, for example, the following paper: Satoshi Suzuki and Keiichi Abe, "Topological Structural Analysis of Digitized Binary Images by Border Following", Computer Vision, Graphics, and Image Processing, Volume 30, Issue 1, April 1985, Pages 32-46

[0047] The image pM34 on the right side of the second row in Figure 7 represents an example of an image generated by S234 (referred to as the contour segmentation map pM34). As shown, the contour segmentation map pM34 represents an image in which contours C1, C2, C3, and C4 of regions P1, P2, P3, and P4, respectively, are added to the segmentation map pM32. In this embodiment, the processor 210 generates the contour segmentation map pM34, which represents the extracted contours C1, C2, C3, and C4 as lines. The contours C1, C2, C3, and C4 are represented in a specific color (black in this embodiment) that is different from the color of any of the regions P1, P2, P3, P4, and BG.

[0048] In S236 (FIG. 6), the processor 210 generates grayscale image data by performing grayscale processing on the contour segmentation map generated in S234. The grayscale processing is performed in the same manner as in S222. In this embodiment, the processing of S236 generates a grayscale image that represents each region with a brightness value brighter than the contour. Image pM36 on the right of the third row in FIG. 7 represents an example of the grayscale image generated in S236. Grayscale image pM36 represents contours C1, C2, C3, and C4 represented in black, and regions P1, P2, P3, P4, and BG represented in lighter colors.

[0049] In S238 (FIG. 6), processor 210 performs an adjustment process on pixel values ​​(here, brightness values) of the grayscale image generated in S236. The adjustment process may be any of various processes that generate an image in which contours are displayed in black and parts other than the contours are displayed in white. In this embodiment, processor 210 sets pixel values ​​equal to or greater than the contour threshold to white (here, 255). As a result, the color of pixels in parts other than the contours is set to white. The color of the contours has already been set to black in S234. The contour threshold is experimentally determined in advance to be a value greater than zero (black) and less than the brightness values ​​of each region obtained in S232-S236.

[0050] Image pM38 on the right side of the fourth row in Figure 7 is an example of a grayscale image generated in S238. Grayscale image pM38 represents the contours C1, C2, C3, and C4 of regions P1, P2, P3, and P4. Grayscale image pM38 is an example of a contour image that represents the contours of the object represented by the target image (hereinafter, grayscale image pM38 will also be referred to as contour image pM38).

[0051] This completes the contour image acquisition process PB (FIG. 6). In S240, the processor 210 generates a composite image by combining the detail image pM28 and the contour image pM38. The processor 210 generates a composite image that represents an image obtained by superimposing multiple images to be combined. The composite image generated in S240 represents an image obtained by superimposing the detail image pM28 and the contour image pM38. In S245, the processor 210 sets the pixel values ​​of pixels in the composite image that have a brightness equal to or lower than the combination threshold to black. Image IM20 in FIG. 7 represents an example of a composite image generated by S240-S245. The composite image IM20 represents both the contours C1, C2, C3, and C4 represented by the contour image pM38 and the fine features of the object OB represented by the detail image pM28 (e.g., the eyes Pe, nose Pn, and mouth Pm).

[0052] In S240, the method for calculating the pixel value (here, luminance value) of each pixel of the composite image may be various. In this embodiment, the processor 210 calculates the pixel value of the composite image by alpha blending the pixel value of the detail image pM28 and the pixel value of the contour image pM38. The alpha value (i.e., the respective weights of the detail image pM28 and the contour image pM38) may be various values. For example, the alpha value may be determined so that the weights are equal between the detail image pM28 and the contour image pM38. Alternatively, the alpha value may be determined so that the weights are unequal between the detail image pM28 and the contour image pM38. The alpha value may be determined experimentally in advance so that the composite image can represent both contours and fine features. After calculating the luminance value of each pixel of the composite image, the processor 210 converts the luminance value of each pixel into a pixel value (here, R, G, B gradation value) suitable for the generative model 900 (FIG. 2). This conversion of pixel values ​​may be the inverse of the conversions in S222 and S236. Here, pixel values ​​may be converted assuming that the color of each pixel is achromatic.

[0053] In the composite image generated by S240, the color of a pixel representing an edge or contour may be set to a color brighter than black by the alpha blending operation. In S245, processor 210 sets such a bright color to black. The composite threshold is set in advance to a value greater than the possible brightness value of a pixel representing an edge or contour in the image generated by S240.

[0054] After S245, the processor 210 ends the process of FIG. 6, i.e., S120 of FIG. 4. In S130, the processor 210 obtains new image data by inputting the data of the composite image generated by S120 into the generative model 900 (FIG. 2). In this embodiment, the processor 210 uses the first adjustment parameter 990a (FIG. 2). This causes the generative model 900 to generate an output image that represents the same content as the composite image content in a line drawing style. Note that the processor 210 uses the following as the text Ptx: Text (e.g., text representing a style such as "line drawing") suitable for the adjustment parameters to be used (here, the first adjustment parameters 990a) is input to the generative model 900. The text Ptx may be set by the user. Alternatively, the processor 210 may use text that is pre-associated with the adjustment parameters to be used as the text Ptx. Note that input of the text Ptx may be omitted. The processor 210 may also cause the GPU 260 to execute some or all of the calculations of the generative model 900.

[0055] Image IM30 in Figure 5 represents an example of an output image generated by S130. As described above, output image IM30 is generated using composite image IM20 of detail image pM28 and contour image pM38. Output image IM30 represents object OBa, which is similar to object OB in composite image IM20.

[0056] The composite image IM20 represents the contours C1, C2, C3, and C4 of the object OB. Therefore, the generative model 900 can generate an output image IM30 representing contours C1a, C2a, C3a, and C4a that correspond to the contours C1, C2, C3, and C4 of the composite image IM20, respectively.

[0057] The synthesized image IM20 represents the eyes Pe, nose Pn, and mouth Pm, which are examples of fine features of the object OB. Therefore, the generative model 900 can generate an output image IM30 representing the eyes Pea, nose Pna, and mouth Pma, which correspond to the eyes Pe, nose Pn, and mouth Pm of the synthesized image IM20, respectively.

[0058] The pixel values ​​of pixels in the composite image IM20 that represent portions other than the contours C1, C2, C3, and C4 and the finer features of the object OB are set to white. That is, regions P1, P2, P3, P4, and BG are not colored. Therefore, the generative model 900 can generate an output image IM30 that represents uncolored regions P1a, P2a, P3a, P4a, and BGa corresponding to regions P1, P2, P3, P4, and BG, respectively, in the composite image IM20.

[0059] Such an output image IM30 differs from the output image IM10o of the reference example in FIG. 3, and is expressed in the style of a line drawing associated with the first adjustment parameter 990a.

[0060] After S130 (FIG. 4), in S140, the processor 210 stores the output image data in the storage device 215 (e.g., the non-volatile storage device 230). Then, the processor 210 ends the processing in FIG. 4. The output image data can be used for various processes (e.g., printing, display, etc.).

[0061] FIG. 8 is a diagram showing an example of an image obtained by image processing using a sample image. The image IM10s on the left side of the first row in the diagram is an example of a target image. The target image IM10s represents a photograph of a person. In the drawing, the target image IM10s is represented by a plurality of black dots obtained by halftoning, but in reality, the target image IM10s is a color image.

[0062] Image IM10so, located to the right of target image IM10s, is an output image obtained by inputting target image IM10s into generative model 900 (FIG. 2). In the drawing, output image IM10so is represented by multiple black dots obtained by halftoning, but in reality, output image IM10so is a grayscale image. Output image IM10so represents the outlines of each of the hair region, body skin region, facial skin region, and clothing region using lines. However, in output image IM10so, these regions and the background are represented by grayscale color gradations. The style of this output image IM10so differs from the intended line drawing style.

[0063] The left image pM28s in the second row in the figure is a detail image obtained by the detail image acquisition process PA (FIG. 6) for the target image IM10s. The detail image pM28s shows the detailed features of a person, including the eyes, nose, and mouth. However, the detail image pM28s does not show the contours of the hair region, the body skin region, the facial skin region, or the clothing region.

[0064] The image pM38s on the right side of the second row in the figure is a contour image obtained by the contour image acquisition process PB (FIG. 6) for the target image IM10s. Image pM38s shows the contours of the hair region, body skin region, facial skin region, and clothing region with lines. However, image pM38s does not show fine features such as the eyes, nose, and mouth.

[0065] Image IM20s in the third row in the figure is a composite image obtained by combining detail image pM28s and image pM38s (FIG. 6: S240-S245). Composite image IM20s shows the detailed features of a person, including the eyes, nose, and mouth, and also shows the outlines of the hair, body skin, facial skin, and clothing regions.

[0066] Image IM30s in the fourth row in the figure is an output image generated by inputting composite image IM20s into generative model 900 (FIG. 2). Output image IM30s uses lines to represent the fine features of a person, including the eyes, nose, and mouth. Output image IM30s also uses lines to represent the contours of the hair region, body skin region, facial skin region, and clothing region. In this way, by using composite image IM20s, generative model 900 can generate output image IM30s that represents the same content as the target image IM10s in a line drawing style.

[0067] As described above, in this embodiment, the processor 210 executes the following processes in accordance with the program 231. In S110 (FIG. 4), the processor 210 acquires a target image IM10 representing an object OB (FIG. 5). In S120, the processor 210 acquires a contour image pM38 and a detail image pM28, and generates a composite image IM20 of the contour image pM38 and the detail image pM28. The composite image IM20 represents an image in which the contour image pM38 and the detail image pM28 are superimposed. In this embodiment, in S120, the processor 210 executes the processes of FIG. 6. The processes of FIG. 6 include a detail image acquisition process PA and a contour image acquisition process PB. In the contour image acquisition process PB, the processor 210 acquires a contour image pM38 representing the contours C1, C2, C3, and C4 of the object OB. In the detail image acquisition process PA, the processor 210 acquires a detail image pM28 representing the features of the object OB (e.g., the eyes Pe, nose Pn, and mouth Pm). The contour image pM38 does not represent such features. Thus, the detail image pM28 represents finer features of the object OB than the contour image pM38. In S240-S245 (FIG. 6), the processor 210 generates a composite image IM20 by combining the contour image pM38 and the detail image pM28. In S130 (FIG. 4), the processor 210 inputs information including the composite image IM20 into the trained generative model 900, thereby obtaining an output image IM30, which is an example of a new image (in this embodiment, the information input to the generative model 900 includes the text Ptx). In this way, the processor 210 can reduce the possibility of obtaining an image that does not represent the contour and features of the object.

[0068] In this embodiment, the detail image acquisition process PA (FIG. 6) also includes processes (S210, S222-S228) for generating an edge image pM28 (FIG. 5) representing the edges of the object OB as a detail image. The edges of the object OB may represent fine features of the object OB, such as the eyes Pe, nose Pn, and mouth Pm. By generating the composite image IM20 using the detail image pM28 representing such edges, the processor 210 can reduce the possibility of acquiring an image that does not represent the features of the object.

[0069] In this embodiment, the contour image acquisition process PB (FIG. 6) includes processes (S210, S232-S238) for generating, as a contour image, an image pM38 that represents the contours C1, C2, C3, and C4 (FIG. 5) of the object OB using lines. By generating a composite image IM20 using the contour image pM38 that represents the contours using lines in this manner, the processor 210 can reduce the possibility of acquiring an image that does not represent the contours of the object.

[0070] In this embodiment, the object OB (FIG. 5) includes a face including N features (N is an integer equal to or greater than 1) selected from the eyes, nose, and mouth. In the example of FIG. 5, the face of the object OB in the target image IM10 includes two eyes, one nose, and one mouth, so N=4. In the detail image acquisition process PA (FIG. 6), the processor 210 generates an image pM28 (FIG. 5) representing one or more of the N features as a detail image. In the contour image acquisition process PB, the processor 210 generates an image pM38 as a contour image. The image pM38 does not represent the N features, but represents the contour C1 of the facial skin region P1 (i.e., the facial contour). Therefore, the processor 210 can acquire an output image representing one or more features selected from the eyes, nose, and mouth and the facial contour, such as the output image IM30 of FIG. 5.

[0071] In this embodiment, a combination of a diffusion model 960 and a first adjustment parameter 990a is used as the generative model 900 (FIG. 2). Such a generative model 900 is an example of a model that generates a line drawing that represents an input image with lines. By using such a generative model 900, the processor 210 can convert the style of the input image into a line drawing.

[0072] B. Second Example: FIG. 9 is a flowchart showing another embodiment of the preprocessing. The processing of FIG. 9 may be executed in S120 (FIG. 4) instead of the above preprocessing (FIG. 6). There are two major differences from the embodiment of FIG. 6. The first difference is that S224, S226, and S228 are omitted from the detailed image acquisition processing PAb. The grayscale image pM22 generated in S222 is used as is as the detailed image. The second difference is that S245 is omitted. The other parts of the preprocessing are the same as the corresponding parts of the processing of FIG. 6. For example, the contour image acquisition processing PB is the same as the contour image acquisition processing PB of FIG. 6.

[0073] FIG. 10 is a diagram showing an example of an image to be processed in preprocessing. As in the example of FIG. 7, the target image IM10 is assumed to be used. The grayscale image pM22 generated in S222 (FIG. 9) is the same as the grayscale image pM22 in FIG. 7. The images pM32, pM34, pM36, and pM38 generated by the contour image acquisition process PB are the same as the images pM32, pM34, pM36, and pM38 in FIG. 7, respectively.

[0074] At S240b (FIG. 9), processor 210 generates a composite image by combining detail image pM22 and contour image pM38. The combining method is similar to that of S240 in FIG. 6. Image IM20b in FIG. 10 shows an example of a composite image. Composite image IM20b is similar to an image obtained by superimposing the contours of contour image pM38 on grayscale image pM22. Composite image IM20b shows both the contours C1, C2, C3, and C4 represented by contour image pM38 and the fine features of object OB represented by grayscale image pM22 (e.g., eyes Pe, nose Pn, and mouth Pm).

[0075] After S240b (FIG. 9), processor 210 ends the process of FIG. 9, i.e., S120 of FIG. 4. The process of S130 is similar to the process of S130 in the embodiment of FIG. 5, except that composite image IM20b (FIG. 10) is used instead of composite image IM20 (FIG. 5).

[0076] FIG. 11 is a diagram showing an example of an image processed by image processing. Image IM30b shows an example of an output image generated by S130 (FIG. 4). Output image IM30b shows object OBb, which is the same as object OB in composite image IM20b. Output image IM30b shows contours C1b, C2b, C3b, and C4b, which correspond to contours C1, C2, C3, and C4, respectively, in composite image IM20b. Output image IM30b shows eyes Peb, nose Pnb, and mouth Pmb, which correspond to eyes Pe, nose Pn, and mouth Pm, respectively, in composite image IM20b.

[0077] Furthermore, in this embodiment, output image IM30b represents regions P1b, P2b, P3b, P4b, and BGb corresponding to regions P1, P2, P3, P4, and BG of composite image IM20b, respectively. Similar to grayscale image pM22 (FIG. 10), composite image IM20b represents grayscale color gradations within each of regions P1, P2, P3, P4, and BG. Therefore, regions P1b, P2b, P3b, P4b, and BGb of output image IM30b may represent grayscale color gradations, unlike regions P1a, P2a, P3a, P4a, and BGb of output image IM30 in FIG. 5. Note that regions P1, P2, P3, P4, and BG of composite image IM20b are not colored with chromatic colors, unlike regions P1, P2, P3, P4, and BG of target image IM10. Therefore, processor 210 can generate output image IM30b that shows regions P1b, P2b, P3b, P4b, and BGb that are not colored in dark colors but are colored in light colors, unlike output image IM10o of the reference example of Fig. 3. The style of such output image IM30b is closer to the intended line drawing style than the style of output image IM10o of the reference example of Fig. 3.

[0078] As described above, in this embodiment, the detailed image acquisition process PAb (FIG. 9) includes processes (S210, S222) for generating a grayscale image pM22 (FIG. 10) representing the target image IM10 in grayscale as a detailed image. The grayscale image pM22 may represent detailed features of the object OB, such as the eyes Pe, nose Pn, and mouth Pm. By generating a composite image IM20b using such detailed image pM22, the processor 210 can reduce the possibility of acquiring an image that does not represent the object's features.

[0079] Furthermore, in this embodiment, in the detail image acquisition process PAb, the processor 210 generates, as a detail image, an image pM22 (FIG. 10) representing one or more of the N parts selected from the eyes, nose, and mouth, similarly to the detail image acquisition process PA in Fig. 6. Therefore, in this embodiment in which the detail image acquisition process PAb and the contour image acquisition process PB are executed, the processor 210 can acquire an output image representing one or more parts selected from the eyes, nose, and mouth and the contour of the face, like the output image IM30b in Fig. 11, similarly to the embodiment in which the detail image acquisition process PA and the contour image acquisition process PB are executed.

[0080] C. Third Example: FIG. 12 is a flowchart showing another embodiment of the preprocessing. The processing of FIG. 12 may be executed in S120 (FIG. 4) instead of the above preprocessing (FIGS. 6 and 9). There are two major differences from the embodiment of FIG. 6. The first difference is that S236 and S238 are replaced with S236c and S238c, respectively. The contour image acquisition processing PBc includes S210, S232, S234, S236c, and S238c. The second difference is that S245 is omitted. The other parts of the preprocessing are the same as the corresponding parts of FIG. 6. For example, the detailed image acquisition processing PA is the same as the detailed image acquisition processing PA in FIG. 6.

[0081] FIG. 13 is a diagram showing an example of an image processed in preprocessing. FIG. 14 is a diagram showing an example of an image processed in image processing. As with the example of FIG. 7, target image IM10 (FIG. 14) is used. Images pM22, pM26, and pM28 (FIG. 13) generated by detailed image acquisition process PA (FIG. 12) are the same as images pM22, pM26, and pM28, respectively, in FIG. 7. Images pM32 and pM34 generated in S232 and S234 of contour image acquisition process PBc are the same as images pM32 and pM34, respectively, in FIG. 7.

[0082] In S236c (FIG. 12), processor 210 calculates a representative color for each region acquired in S232. The representative color of the region of interest may be any of various colors that represent the colors of the pixels included in the region of interest in target image IM10 (FIG. 14). In this embodiment, processor 210 calculates the color represented by the average value of each of RGB as the representative color. Instead of the average value, various values ​​that represent the magnitude of the gradation value, such as the mode or median, may be used.

[0083] In S238c (FIG. 12), processor 210 fills the interior of the contour of each region of contour segmentation map pM34 generated in S234 with the representative color calculated in S236c. Image pM38c in FIG. 13 shows an example of a filled image generated by S238c. In filled image pM38c, facial skin region P1 is filled with the representative color of facial skin region P1 in target image IM10 (FIG. 14). Other regions P2, P3, P4, and BG in filled image pM38c are also filled with the representative colors of corresponding regions P2, P3, P4, and BG in target image IM10, respectively. Note that the contours extracted in S234 are maintained. In this way, filled image pM38c represents contours C1, C2, C3, and C4 of each region P1, P2, P3, and P4. The filled image pM38c is used as an outline image (the filled image pM38c is also called an outline image pM38c).

[0084] In S240c (FIG. 12), processor 210 generates a composite image by combining detail image pM28s and outline image pM38c. The combining method is similar to that of S240 in FIG. 6. Image IM20c in FIG. 13 represents an example of a composite image. Composite image IM20c is similar to an image obtained by superimposing the edges represented by detail image pM28 on filled-in image pM38c. Composite image IM20c represents both the outlines C1, C2, C3, and C4 represented by outline image pM38c and the fine features of object OB represented by detail image pM28s (e.g., eyes Pe, nose Pn, and mouth Pm). Furthermore, in composite image IM20c, each region P1, P2, P3, P4, and BG is filled with a color representative of the corresponding region in target image IM10.

[0085] After S240c (FIG. 12), processor 210 ends the processing of FIG. 12, i.e., S120 of FIG. 4. The processing of S130 is similar to the processing of S130 in the embodiment of FIG. 5, except that composite image IM20c (FIG. 13) is used instead of composite image IM20 (FIG. 5), and second adjustment parameter 990b is used instead of first adjustment parameter 990a.

[0086] Image IM30c in Figure 14 represents an example of an output image generated by S130. Output image IM30c represents object OBc, which is similar to object OB in composite image IM20c. Output image IM30c represents contours C1c, C2c, C3c, and C4c, which correspond to contours C1, C2, C3, and C4, respectively, in composite image IM20c. Output image IM30c represents eyes Pec, nose Pnc, and mouth Pmc, which correspond to eyes Pe, nose Pn, and mouth Pm, respectively, in composite image IM20c.

[0087] In this embodiment, the output image IM30c represents regions P1c, P2c, P3c, P4c, and BGc, which correspond to regions P1, P2, P3, P4, and BG, respectively, of the composite image IM20c. In the composite image IM20c, the regions P1, P2, P3, P4, and BG are filled with the representative colors of the corresponding regions P1, P2, P3, P4, and BG, respectively, of the target image IM10. In S130 (FIG. 4), the processor 210 uses the second adjustment parameter 990b for anime art. Therefore, the processor 210 can generate the output image IM30c representing the regions P1c, P2c, P3c, P4c, and BGc, which are filled with the same colors as the regions P1, P2, P3, P4, and BG, respectively, of the composite image IM20c. Output image IM30c represents target image IM10 using lines and fewer colors than target image IM10, in the intended anime art style.

[0088] As described above, in this embodiment, a combination of the diffusion model 960 and the second adjustment parameter 990b is used as the generative model 900 (FIG. 2). Such a generative model 900 is an example of a model that generates an image that represents an input image using lines and fewer colors than the number of colors in the input image. By using such a generative model 900, the processor 210 can convert the style of the input image into a style (e.g., anime art) that represents an input image using lines and fewer colors than the number of colors in the input image.

[0089] In this embodiment, as shown in FIG. 14, the target image IM10 represents an object OB and a background BG. The background BG is an example of an external portion that is adjacent to the object OB. The object OB includes multiple regions P1, P2, P3, and P4, each corresponding to a different portion. The contour image acquisition process PBc (FIG. 12) includes steps S236c and S238c. In steps S236c and S238c, the processor 210 fills each region P1, P2, P3, P4, and BG with its respective representative color. That is, the image pM38c (FIG. 13) generated by the contour image acquisition process PBc represents the object OB and the background BG in different colors. Furthermore, the image pM38c represents each of the multiple regions P1, P2, P3, and P4 of the object OB as a single-color region showing the representative color representing the color of the corresponding region. In the contour image acquisition process PBc, the processor 210 generates such an image pM38c as a contour image. The processor 210 generates a composite image IM20c using such contour image pM38c, and generates an output image IM30c by inputting the composite image IM20c into the generative model 900. Thus, the processor 210 can generate an output image IM30c that represents the same content as the target image IM10 using lines and fewer colors than the number of colors in the target image IM10.

[0090] In this embodiment, the contour image acquisition process PBc (FIG. 12) includes processes (S210, S232, S234, S236c, S238c) for generating, as a contour image, an image pM38c that represents the contours C1, C2, C3, and C4 (FIG. 13) of the object OB using lines. By generating a composite image IM20c using the contour image pM38c that represents the contours using lines in this manner, the processor 210 can reduce the possibility of acquiring an image that does not represent the contours of the object.

[0091] Furthermore, in this embodiment, in the contour image acquisition process PBc, the processor 210 generates an image pM38c (FIG. 13) as a contour image, similar to the contour image acquisition process PB of FIG. 6. Unlike the detail image pM28, the image pM38c represents the contour C1 of the facial skin region P1 (i.e., the facial contour) rather than representing the N features selected from the eyes, nose, and mouth. Therefore, in this embodiment in which the detail image acquisition process PA and the contour image acquisition process PBc are executed, similar to the embodiment of FIG. 6 in which the detail image acquisition process PA and the contour image acquisition process PB are executed, the processor 210 can acquire an output image representing one or more features selected from the eyes, nose, and mouth and the facial contour, such as the output image IM30c of FIG. 14.

[0092] D. Fourth Example: FIG. 15 is a flowchart showing another embodiment of the preprocessing. The processing of FIG. 15 may be executed in S120 (FIG. 4) instead of the above preprocessing (FIGS. 6, 9, and 12). There are two major differences from the embodiment of FIG. 6. The first difference is that S222-S228 are omitted from the detailed image acquisition processing PAd. The detailed image acquisition processing PAd includes S210. The target image acquired in S210 is used as is as the detailed image. The second difference is that S245 is omitted. The other parts of the preprocessing are the same as the corresponding parts of the processing of FIG. 6. For example, the contour image acquisition processing PB is the same as the contour image acquisition processing PB of FIG. 6.

[0093] Fig. 16 is a diagram showing an example of an image processed in preprocessing. Fig. 17 is a diagram showing an example of an image processed in image processing. As with the example of Fig. 7, target image IM10 (Fig. 17) is used. Images pM32, pM34, pM36, and pM38 (Fig. 16) generated by contour image acquisition process PB are the same as images pM32, pM34, pM36, and pM38 in Fig. 7, respectively.

[0094] In S240d (FIG. 15), processor 210 generates a composite image by combining target image IM10, which is a detail image, with contour image pM38. The combining method is similar to the combining method of S240 in FIG. 6. Image IM20d in FIG. 16 represents an example of a composite image. Composite image IM20d is similar to the image obtained by superimposing the contours of contour image pM38 on target image IM10 (FIG. 17). Composite image IM20d represents both the contours C1, C2, C3, and C4 represented by contour image pM38 and the fine features of object OB represented by target image IM10 (e.g., eyes Pe, nose Pn, and mouth Pm).

[0095] After S240d (FIG. 15), processor 210 ends the process of FIG. 15, i.e., S120 of FIG. 4. The process of S130 is the same as the process of S130 in the embodiment of FIG. 5, except that composite image IM20d (FIG. 16) is used instead of composite image IM20 (FIG. 5). For example, in this embodiment, first adjustment parameter 990a is used.

[0096] Image IM30d in Figure 17 represents an example of an output image generated by S130 (Figure 4). Output image IM30d represents object OBd, which is similar to object OB in composite image IM20d. Output image IM30d represents contours C1d, C2d, C3d, and C4d, which correspond to contours C1, C2, C3, and C4, respectively, in composite image IM20d. Output image IM30d represents eyes Ped, nose Pnd, and mouth Pmd, which correspond to eyes Pe, nose Pn, and mouth Pm, respectively, in composite image IM20d.

[0097] In this embodiment, output image IM30d represents regions P1d, P2d, P3d, P4d, and BGd, respectively corresponding to regions P1, P2, P3, P4, and BG of composite image IM20d. Similar to target image IM10, composite image IM20d represents chromatic color gradations within each of regions P1, P2, P3, P4, and BG. Therefore, regions P1d, P2d, P3d, P4d, and BGd of output image IM30d may represent grayscale color gradations, similar to regions P1z, P2z, P3z, P4z, and BGz of output image IM10o of the reference example of FIG. 3. However, unlike output image IM10o of the reference example of FIG. 3, output image IM30d properly represents the contours of each of regions P1d, P2d, P3d, and P4d. The style of such output image IM30d is closer to the intended line drawing style than the style of output image IM10o in FIG.

[0098] As described above, in this embodiment, the processor 210 acquires the target image IM10 (FIG. 17) as a detail image in S210 (FIG. 15). The target image IM10 may represent detailed features of the object OB, such as the eyes Pe, nose Pn, and mouth Pm. The processor 210 uses the target image IM10 as a detail image to generate the composite image IM20d. This allows the processor 210 to reduce the possibility of acquiring an image that does not represent the features of the object.

[0099] Furthermore, in this embodiment, in the detail image acquisition process PAd, the processor 210 generates, as a detail image, an image IM10 (FIG. 17) representing one or more features of the N features selected from the eyes, nose, and mouth, similarly to the detail image acquisition process PA in Fig. 6. Therefore, in this embodiment in which the detail image acquisition process PAd and the contour image acquisition process PB are executed, the processor 210 can acquire an output image representing one or more features selected from the eyes, nose, and mouth and the contour of the face, like the output image IM30d in Fig. 17, similarly to the embodiment in Fig. 6 in which the detail image acquisition process PA and the contour image acquisition process PB are executed.

[0100] E. Variations: (1) The process of acquiring a contour image may be various processes instead of the contour image acquisition processes PB and PBc (FIGS. 6, 9, 12, and 15). For example, in S234, processor 210 may refer to multiple regions determined by the segmentation process and extract, as a portion representing a contour, a portion where multiple pixels belonging to different regions intersect. Furthermore, instead of an image representing the contour of an object using lines, the contour image may be an image representing the contour in various ways. For example, the contour image may represent two regions that intersect the contour in different colors. Such a contour image may be generated, for example, by a process obtained by omitting the region contour extraction process (S234) from contour image acquisition process PBc of FIG. 12. In S238c, processor 210 fills each region entirely with a corresponding representative color. Although not shown, the generated contour image is similar to an image obtained by omitting the contour lines from a filled-in image (for example, filled-in image pM38c of FIG. 13). In such a contour image, color changes represent contours, and by treating these color changes as boundaries, the generative model 900 can generate images in the style of line art or anime art.

[0101] (2) The process for acquiring a contour image and a detail image may be various other processes instead of the processes of Fig. 6 (processes PA, PB), Fig. 9 (processes PAb, PB), Fig. 12 (processes PA, PBc), and Fig. 15 (processes PAd, PB). A process arbitrarily selected from the detail image acquisition processes PA, PAb, and PAd may be combined with a process arbitrarily selected from the contour image acquisition processes PB and PBc. For example, the contour image acquisition process PBc may be combined with the detail image acquisition process PAb or the detail image acquisition process PAd.

[0102] (3) The images used to generate the composite image may include other images in addition to the outline image and the detail image (eg, target image IM10).

[0103] (4) The process of generating a new image using a composite image may be various processes instead of the process of FIG. 4. For example, the instruction to start image processing may include style information specifying either "line art" or "anime art." Processor 210 may proceed with image processing in accordance with the style information. For example, "line art" may be associated with any of the preprocessing processes shown in FIGS. 6, 9, and 15 and first adjustment parameters 990a. "Animation art" may be associated with the preprocessing process shown in FIG. 12 and second adjustment parameters 990b. Processor 210 may then execute the preprocessing associated with the style information and generate an output image from the composite image using the adjustment parameters associated with the style information. Note that either the first adjustment parameter 990a or the second adjustment parameter 990b may be omitted.

[0104] (5) The generative model that generates a new image using a synthetic image is not limited to the generative model 900 of FIG. 2 , but may be various machine learning models. For example, the generative model may be a model that performs a task called neural style transfer (such a model is also called a style conversion model). Neural style transfer uses a machine learning model to convert the style of an image into another style. The converted style may be selected from, for example, the following two types of styles: (Type 1 style) A style that displays the input image with lines (e.g., line drawing) (Type 2 style) A style that represents an input image using lines and fewer colors than the input image (for example, anime art).

[0105] As the first style, various other styles may be used instead of line drawings (for example, ink painting style, etc.). The first adjustment parameters 990a may be trained using a plurality of images in the first style.

[0106] Instead of anime art, various other styles may be used as the second style. For example, a style called flat color art may be used. Flat color art is a style in which each of multiple regions included in an image is represented by a single color region. Here, the outline of each region may be represented by a line. Alternatively, the outline may be omitted (in this case, the outline is represented by a portion in the image where the color changes). In either case, the second adjustment parameter 990b may be trained using multiple images in the second style. When the outline is omitted, an image that represents the outline by color changes without the outline may be used as the outline image. Such an outline image may be generated, for example, by a process obtained by omitting the region outline extraction process (S234) from the outline image acquisition process PBc of FIG. 12.

[0107] Furthermore, the style after conversion is not limited to the above two types of styles, and may be any other style (for example, watercolor style, oil painting style, etc.).

[0108] As the style transfer model, for example, a style transfer model that uses normalization called adaptive instance normalization (AdaIN) may be adopted. Furthermore, as the architecture of the style transfer model, for example, the architecture of a technology called "Fast Patch-based Style Transfer of Arbitrary Style" or the architecture of a technology called "Avatar-Net: Multi-scale Zero-shot Style Transfer by Feature Decoration" may be adopted.

[0109] In either case, the generative model may include one or both of a model that generates a line drawing that represents an input image using lines, and a model that generates an image that represents an input image using lines and fewer colors than the number of colors in the input image.

[0110] (6) The segmentation model 800 (FIG. 1) may be various models that divide an image into multiple regions, such as Mask R-CNN or YOLO, instead of the above-described models. The segmentation model 800 may be a model that performs region division called "instance segmentation" or "semantic segmentation." Furthermore, the processor 210 may divide an image into multiple regions using other methods, such as template matching, without using a machine learning model.

[0111] (7) The object represented by the target image is not limited to a person, but may be any object (for example, a pet such as a dog or cat, a vehicle such as a car or airplane, or a landscape such as the sea or mountains).

[0112] (8) The processor 210 may perform various image processes, not limited to image processing that converts the style of an image. For example, the processor 210 may perform the following process. That is, the processor 210 generates a third image by synthesizing a first image and a second image. Here, the processor 210 generates the third image representing an image obtained by superimposing the first image and the second image. Then, the processor 210 acquires a new image by inputting information including the third image into a trained machine learning model. Here, the machine learning model is not limited to a style transfer model, and may be various models that generate a new image based on an input image. With this configuration, the processor 210 can acquire a new image based on features represented by the first image and features represented by the second image.

[0113] (9) In the above embodiment and the above modification, the processor 210 may cause the GPU 260 to execute various calculations. Note that the GPU 260 may be omitted.

[0114] (10) The image processing device 200 in Fig. 1 may be a device of a type different from a personal computer (for example, a digital camera, a scanner, or a smartphone). Furthermore, multiple devices (for example, computers) that can communicate with each other via a network may share some of the image processing functions of the image processing device and collectively provide the image processing functions (a system including these devices corresponds to an image processing device).

[0115] In each of the above embodiments, a part of the configuration realized by hardware may be replaced by software, and conversely, a part or all of the configuration realized by software may be replaced by hardware. For example, the processing by the generative model 900 (FIG. 2) in FIG. 1 may be executed by a dedicated hardware circuit such as an Application Specific Integrated Circuit (ASIC).

[0116] Furthermore, when some or all of the functions of the present disclosure are realized by a computer program, the program can be provided in a form stored on a computer-readable recording medium (e.g., a non-transitory recording medium). The program can be used in a state stored on the same or a different recording medium (computer-readable recording medium) from when it was provided. The "computer-readable recording medium" is not limited to portable recording media such as memory cards and CD-ROMs, but can also include internal storage devices within a computer, such as various ROMs, and external storage devices connected to a computer, such as a hard disk drive.

[0117] The above-described examples and modifications can be combined as appropriate. The above-described examples and modifications are provided to facilitate understanding of the present disclosure and are not intended to limit the present invention. The present invention may be modified or improved without departing from the spirit thereof, and the present invention includes equivalents thereof. [Explanation of symbols]

[0118] 200...image processing device, 210...processor, 215...storage device, 220...volatile storage device, 230...non-volatile storage device, 231...program, 240...display unit, 250...operation unit, 260...graphics processing unit (GPU), 270...communication interface, 800...segmentation model, 900...generative model, 910...text encoder, 920...image encoder, 930...latent variable model, 940...image decoder, 960...diffusion model, 990a...first adjustment parameter, 990b...second adjustment parameter

Claims

1. A program, a first capture function for capturing a target image representing an object; a second acquisition function for acquiring a contour image representing a contour of the object and a detail image representing finer features of the object than the contour image; a synthesis function for generating a synthesis image by synthesizing a plurality of images including the outline image and the detail image; a third acquisition function that acquires a new image by inputting information including the synthetic image into a trained machine learning model; A program that makes the computer realize the above.

2. 2. The program according to claim 1, the second acquisition function includes a function of generating an edge image representing an edge of the object as the detail image; program.

3. 2. The program according to claim 1, the second acquisition function includes a function of acquiring the target image as the detailed image; program.

4. 2. The program according to claim 1, the second acquisition function includes a function of generating a grayscale image representing the target image in grayscale as the detail image; program.

5. 3. The program according to claim 1 or 2, the second acquisition function includes a function of generating an image representing the contour of the object using lines as the contour image; program.

6. 3. The program according to claim 1 or 2, the target image represents the object and an outer portion that is a portion in contact with the object; the object includes a plurality of distinct parts; the second acquisition function includes a function of generating, as the contour image, an image in which the object and the outer portion are represented in different colors, and in which each of the plurality of portions of the object is represented by a single-color area showing a representative color that represents the color of the portion; program.

7. 3. The program according to claim 1 or 2, the object includes a face including N features (N is an integer equal to or greater than 1) selected from eyes, a nose, and a mouth; The second acquisition function is generating an image representing one or more of the N parts as the detail image; a function of generating an image representing the contour of the face without representing the N facial features as the contour image; Including, the program.

8. 3. The program according to claim 1 or 2, The trained machine learning model is a model that generates a line drawing that represents an input image with lines, or a model that generates an image that represents the input image with lines and a number of colors that is fewer than the number of colors in the input image. program.

9. A program, a function of generating a third image by combining the first image and the second image; acquiring a new image by inputting information including the third image into a trained machine learning model; A program that makes the computer realize the above.

10. An image processing device, a first acquisition unit that acquires a target image representing an object; a second acquisition unit that acquires a contour image that represents a contour of the object and a detail image that represents finer features of the object than the contour image; a synthesis unit for generating a synthesis image by synthesizing a plurality of images including the outline image and the detail image; a third acquisition unit that acquires a new image by inputting information including the synthetic image into a trained machine learning model; An image processing device comprising:

11. An image processing device, a synthesis unit that synthesizes the first image and the second image to generate a third image; an acquisition unit that acquires a new image by inputting information including the third image into a trained machine learning model; An image processing device comprising: