Generation support device, generation support program, and generation support method
The generation support device and method address the limitations of existing image generation technologies by allowing users to input style and element images, enabling versatile and user-friendly image creation using generative models.
Patent Information
- Application Number
- JP2025155495
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-16
AI Technical Summary
Existing image generation technologies, such as those described in Patent Document 1, are limited to generating character images in arbitrary poses and lack versatility for various image generation tasks.
A generation support device and method that allows users to input style and element images, along with their positions, to a generative model for generating desired images, utilizing a generative model like GANs or CNNs, and supports image generation through a user interface that includes chat formats and drag-and-drop functionality.
Enables users to easily generate desired images by providing a user-friendly interface for inputting style and element information, enhancing the versatility and ease of image creation.
Smart Images

Figure 2025183388000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a generation support device, a generation support program, and a generation support method. [Background technology]
[0002] In recent years, image generation has been performed by various methods. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7169027 Summary of the Invention [Problem to be solved by the invention]
[0004] For example, Patent Document 1 proposes a technology for generating character images using machine learning. It is being done.
[0005] However, the technology in Patent Document 1 generates an image of a character in an arbitrary pose. This is the only thing that can be done and it cannot be applied to various image generation.
[0006] The present invention has been made in view of the above background, and is directed to a method for enabling a user to easily obtain a desired image. The purpose is to generate [Means for solving the problem]
[0007] In order to solve the above problem, the generation support device according to the present disclosure is configured to receive at least In addition, style information relating to the style of the resultant image and element images constituting a part of the resultant image are stored. a generated information acquisition unit that acquires generated information including the image; and a text generated based on the generated information. inputting store information into a generative model, and generating the and a resultant image generating unit that generates a resultant image.
[0008] Other problems and solutions disclosed in this application are described in the embodiments and drawings. will be more clearly revealed. [Effects of the Invention]
[0009] According to the present invention, a user can easily generate a desired image. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of an evaluation system according to an embodiment of the present invention. [Figure 2] 2 is a diagram illustrating an example of a hardware configuration of a server device 1 according to the embodiment. FIG. [Figure 3] 2 is a diagram illustrating an example of a functional configuration of a server device 1 according to the embodiment. FIG. [Figure 4] 10 is a diagram showing an example of basic information stored in a generation information storage unit 131. FIG. [Figure 5] 10 is an example of a screen on which the generation information acquisition unit 111 acquires position information of a partial image and a material image. [Figure 6] FIG. 10 is a diagram illustrating an example of processing performed by the server device 1 according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] <Summary of the Invention> [Item 1] A generation support device that supports generation of a result image, The user provides at least style information regarding the style of the result image and a generation information acquisition unit that acquires generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model a resultant image generating unit that generates the resultant image based on output information output from the A generation support device comprising: [Item 2] The generation information acquisition unit further acquires position information of the element image in the resultant image. thing, Item 1. A generation assistance device according to item 1. [Item 3] the style information includes web address information; The generation information acquisition unit acquires information about a website included in the website specified by the web address. the style information determined based on the block information is used as the style information; 3. The generation assistance device according to item 1 or 2, [Item 4] The information acquisition unit receives upload or selection of the element image, and acquiring the position information by arranging the element images in a frame; Item 3. The generation support device according to item 2, characterized in that: [Item 5] The information acquisition unit acquires the generation information through input by the user in a chat format. To do, 3. The generation assistance device according to item 1 or 2, [Item 6] When acquiring the input in the chat format, the information acquisition unit may include, as the generation information: providing the user with an indication of required information; Item 6. A generation assistance device according to item 5, characterized in that: [Item 7] A generation support program for supporting generation of a resultant image, The processor The user provides at least style information regarding the style of the result image and a generation information acquisition step of acquiring generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model an image generating step of generating the resultant image based on output information output from the A generation support program that executes the above. [Item 8] A generation assistance method for assisting in the generation of a resultant image, comprising: The processor: The user provides at least style information regarding the style of the result image and a generation information acquisition step of acquiring generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model an image generating step of generating the resultant image based on output information output from the A generation support method for executing the above.
[0012] FIG. 1 is a diagram showing an example of the overall configuration of an evaluation system according to one embodiment of the present invention. The morphology generation support system includes a server device 1. The server device 1 The terminal 3 is communicably connected via a communication network 2. The communication network 2 includes: For example, the Internet is a network that uses public telephone lines, mobile phone lines, wireless communication channels, and Ethernet. It is constructed using Net (registered trademark) and other technologies.
[0013] ==Server device 1== The server device 1 is a general-purpose computer such as a workstation or a personal computer. It can be implemented as a computer or logically through cloud computing. In this embodiment, for the sake of convenience, one device is illustrated, but the number of devices is not limited to this. The number of the devices is not limited, and multiple devices may be used.
[0014] ==User terminal 3== The user terminal 3 is a computer operated by a user who generates an image. These include smartphones, tablet computers, and personal computers. For example, the server device 1 is accessed by an application or a web browser executed on the user terminal 3. can be accessed.
[0015] FIG. 2 is a diagram illustrating an example of the hardware configuration of the server device 1. The illustrated configuration is This is merely an example, and the server device 1 may have other configurations. memory 102, storage device 103, communication interface 104, input device 105, output device The storage device 103 stores various data and programs. Hard disk drives, solid state drives, flash memory, etc. The interface 104 is an interface for connecting to the communication network 2, for example. For example, an adapter for connecting to an Ethernet (registered trademark), modem for wireless communication, wireless communication device for wireless communication, USB (Universal Serial Bus) for serial communication Examples include a Serial Serial Bus (Serial Serial Bus) connector and an RS232C connector. The device 105 is a device for inputting data, such as a keyboard, a mouse, a touch panel, a button, The output device 106 outputs data, for example, a display. The functional units of the server device 1, which will be described later, are processors. 101 reads a program stored in a storage device 103 into a memory 102 and executes it. Each storage unit of the server device 1 is provided by a memory 102 and a storage device 103. It is implemented as part of the storage area provided.
[0016] 3 shows the functional configuration of the server device 1. As shown in FIG. 3, the server device 1 has the following functions: Each storage unit, including a generated information storage unit 131 and a result image information storage unit 132, and a generated information acquisition unit The image processing unit 110 includes a processing unit 111 and a result image generating unit 112.
[0017] The following describes the generated information storage unit 131 and the result image information storage unit 132. .
[0018] The generation information storage unit 131 stores the resultant image (the image generated by the server device 1) as shown in FIG. The information used to generate the image (hereinafter referred to as generation information) is stored. and style information (including text information, web addresses, web information, etc.) The generation information may include information on the element image that is the basis for forming a part of the resulting image. The element image may include, for example, the subject of the resulting image (e.g., a person or object, as described below). When the generation information acquisition unit 111 generates text to be input to the generation model, the A partial image is an image of a subject (the subject of which is described as the content of the If it is an image of an object, for example, an image of a product, an image containing a product, or an image of the appearance of a product container or outer box, etc. The element images may also include source images that are not the subject of the resulting image. The generation information may include, for example, information on the position of a partial image and a material image in the resultant image. may be, but is not limited to these.
[0019] The style is a style in the design of the resultant image generated by the generation assistance device. For example, the requirements for the resulting image (such as the objects, people, and scenery contained in the image) The concept (target, story, etc.), color, texture (images that can be felt with visual sense) texture of the image surface, layout (element placement and relative positioning, etc.), font Aesthetic attributes and characteristics such as shape (sharp angles, rounded shapes, straight lines, curves, etc.) represents, but is not limited to,
[0020] The material images include images of hands, faces, plants, everyday items, tables, and geometric shapes such as circles, triangles, and squares. Free-form figures that are not bound by academic or geometric rules, and the resulting image The material image is an image that is the basis for a part of the background of the image to be generated. It may also include, but is not limited to, images (templates).
[0021] The resultant image information storage unit 132 stores the resultant image generated by the resultant image generation unit 112 .
[0022] The following describes each processing unit of the generation information acquisition unit 111 and the result image generation unit 112. do.
[0023] The generated information acquisition unit 111 receives, for example, the generated information from the user terminal 3 via the communication network 2. style information relating to the style of the result image necessary for generating the result image; A raw image including element images constituting an image and position information of the element images in the resultant image. The generation information acquisition unit 111 stores the acquired generation information in the generation information storage unit 13. The communication in the transmission and reception may be either wired or wireless. Any communication protocol may be used as long as communication can be performed.
[0024] The generation information acquisition unit 111 may acquire the generation information in the form of text information. 111 may acquire a sentence indicating the result image to be generated by a user's input operation, or The generation information acquisition unit 111 may acquire the above words. A sentence or word representing the file is presented to the user, and the sentence or word selected by the user is It may also be acquired as generation information.
[0025] The generation information acquisition unit 111 acquires information on a web address (URL, etc.) as generation information. The generated information acquisition unit 111 may acquire the generated information included in the website specified by the web address. The web information contained in the web page can be acquired, the style can be determined, and the generated information can be used. The text information, image information, video information, code information (website) contained in the website The code that makes up a website, such as HTML, CSS, and JavaScript. t, etc., but is not limited to these. 111 determines the style of the target, concept, etc. based on the text information, for example. In this case, the generated information acquisition unit 111 may perform a morphological analysis on the text information, for example. Based on the information about the words and their number, the style of the target, concept, etc. can be determined. However, the method is not limited to these. Even if you determine the style of color, texture, font, shape, etc. from information or code information, In this case, the generated information acquisition unit 111 analyzes the image information or video information and The style may be determined from the color, texture, font, shape, etc. included in the code information. The information on the color, texture, shape, and fonts used on the web background image is stored in the The style may be determined based on the above, but the method is not limited to these.
[0026] The generation information acquisition unit 111 acquires element images that are the basis for forming part of the resulting image. The generation information acquisition unit 111 may accept uploading of element images. The acquisition unit 111 acquires a material image (for example, 201 in FIG. 5) as shown in FIG. 5. The image data is stored in the server device 1 and presented to the user terminal 3, and the user selects the material image. The user may select a material image and acquire the selected material image as the generation information.
[0027] The generation information acquisition unit 111 acquires, for example, position information of element images in the resultant image. The position information indicates the coordinates of the element image in the result image, for example, a specific corner of the result image. These are X and Y coordinates with a predetermined position as the origin, such as the center of the element image. In this case, the generation information acquisition unit 111 may generate a result image as shown in FIG. The frame may be horizontal, square, vertical, etc. However, the generated information acquisition unit 111 presents the generated information (including but not limited to the above) to the user terminal 3. In the frame, information on the user's operation on the user terminal 3 is acquired, and the material image in the resulting image is In this case, the generation information acquisition unit 111 acquires the position information of the image, for example, the position information of the image acquired by the user. Accepts drag and drop operations, and allows placement of partial images (203) or material images (204). The generation information acquisition unit 111 acquires the position information of the element image. , and may also accept enlargement, reduction, rotation, inversion, and transformation of element images. The information acquisition unit 111 acquires, as position information, the anteroposterior relationship of a plurality of element images (for example, layer information In addition, the position information may be the positional relationship between element images (part information). The information may be information such as "there is a material image below the component image."
[0028] The generation information acquisition unit 111 may acquire the generation information in a chat format. The information acquisition unit 111 performs morphological analysis or the like to convert text information acquired in a chat format into word binary data. In this case, the generated information acquisition unit 111 Regarding the generated information obtained from users, Please tell me about some websites that might be useful to you." In this case, the generation information acquisition unit 111 may provide support to make it easier for the user to recognize the The image is generated from a list of information required for generating the image that has not been obtained from the user. Information that is not available or that could not be determined from the generated information obtained will be provided to the user. The method is not limited to providing guidance.
[0029] The generation information acquisition unit 111 acquires a process to be input to the image generation model based on the acquired generation information. prompts or prerequisite conditions (collectively referred to as prompt information in this specification) The prerequisites are, for example, the image size, frame shape, file size, and resolution. The generated information acquisition unit 111 may include, but is not limited to, information such as the degree of The prompt includes at least text that indicates the style. For example, prompt information is generated using a feature extraction module and a language model. The generation information acquisition unit 111 generates one or more pieces of prompt information, presents them to the user, and The user may select or edit the item.
[0030] The generation information acquisition unit 111 receives the image generated by the result image generation unit 112 when generating the prompt. The structure of the generated prompt may be changed depending on the type of generative model used for generation. The information acquisition unit 111 may generate a sentence-type prompt, or a list of words. It may also generate prompts such as bracketing important words and The order of words should appear at the beginning of the prompt, and multiple important words should be included. By indicating the importance of each word, we generate prompts that allow the generative model to recognize important words. It may be done.
[0031] The prompt generated by the generation information acquisition unit 111 includes at least text representing the style. The generation information acquisition unit 111 may generate a plurality of prompts, or may generate a plurality of prompts. For example, the generated information acquisition unit 11 may have a certain degree of randomness in the text included in the generated information acquisition unit 11. 1 is based on the semantic distance and similarity between the style and the text. Specifically, the generation information acquisition unit 111 For example, style information related to "sea" is included in the generated information acquired through user input. If the word "ocean" contains several words, it will prompt other words that are close in meaning to "ocean" or highly similar to it. The result image generating unit 112, which will be described later, generates a prompt by including the image in the prompt. By using prompts to generate a result image, the image that the user wants to generate is closer to the image. Conversely, the generation information acquisition unit 111 can generate a result image based on the user's input, etc. When the generated information obtained from the Generate prompts by including other words that are distantly related or less similar to the prompt. The result image generating unit 112, which will be described later, generates a result image using these prompts. This allows the user to choose the style of the resulting image, for example, if they do not yet have an image of the ocean. It is also possible to make it easier to consider the direction of the prompt. The randomness may have other effects as well.
[0032] The generation information acquisition unit 111 performs the following on the first resultant image generated by the resultant image generation unit 112: The additional generation information acquired by the generation information acquisition unit 111 is as follows: Used to modify or add the prompt used when generating the first result image, The generating unit 112 uses this when generating the second resultant image.
[0033] The result image generating unit 112 generates, for example, style information, element images, and position information. The result image generating unit 112 generates a result image based on at least one of the above. Generation information based on at least one of tile information, element image, and position information. The prompt information generated by the acquisition unit 111 is input to the generation model, and the image output by the generation model is The result image generation unit 112 also acquires the image output by the generative model as the result image. Alternatively, the output image may be processed to generate a result image. 112 presents the generated result image to the user. The user can download the presented image. You can download it.
[0034] The generative model used by the resultant image generating unit 112 to generate the resultant image is implemented in the server device 1. The communication network 2 may be implemented in another server accessible via the communication network 2. For this reason, the generative model may be implemented in the server device 1. If so, the result image generation unit 112 inputs the prompt information into the generation model, and If the driver is installed on another server, the result image generator 112 may generate prompt information. , and transmits it to the generative model via a communication network 2. Entering prompt information into a Generative Model, including sending prompt information to a Generative Model It is expressed as "to exert effort."
[0035] The generative model may be, for example, a specific input vector or a random nodal Any model can be used that receives the image size and generates an image from that information. The generator may be configured to convert input information into appropriate The generator converts the image into a set of features or patterns, which are then converted into an image. Convolutional Neural Network k, CNN, Transformer, or other Built using a cloud-learning architecture, but other architectures are available In addition, the generative model may include, for example, a discriminator. The classifier distinguishes whether an image is real or a fake image generated by the generator. The classifier is constructed using a network such as CNN, but is not limited to these. The generative model may comprise, for example, a genetic adversarial network (GAN). The network trains the generator to generate more realistic images, while the classifier simultaneously trains the It learns to improve its ability to distinguish between images of objects and fake images.
[0036] The result image generating unit 112 may generate two or more result images. 112 presents the generated result image to the user.
[0037] When the generated multiple result images are presented to the user terminal 3, the result image generation unit 112 In the user terminal 3, the user selects an image that is close to the desired result image from multiple images, or The selection operation of the image that is out of the range is accepted, and the result image selected by the selection operation is used as the basis. In this case, the result image generating unit 112 may further generate a result image. Based on the characteristics of result image A, which is selected as the result image close to the one you want to generate, A resultant image B similar to the image A may be generated. is the process input to the generative model when generating the result image A selected by the selection operation. Alter the information in the prompt or reproduce an image similar to or a variation of the selected result image A We generate prompts that show the However, the present invention is not limited to this method.
[0038] The resultant image generating unit 112 performs generation information acquisition on the generated resultant image (first resultant image). When the acquisition unit 111 acquires additional information from the user, a second result image may be generated. In this case, the resultant image generating unit 112 uses the generative model when generating the first resultant image. The prompt information entered in the prompt information acquisition unit 111 and the prompt information generated by the generation information acquisition unit 111 based on the additional information are input. prompt information is input to the generative model, and the first All you have to do is generate the result image of step 2.
[0039] FIG. 6 is a diagram illustrating an example of processing performed by the generation support device of this embodiment.
[0040] The server device 1 acquires the generated information from the user (1001). The server device 1 generates a prompt based on the generated information (1002). The server device 1 inputs the generated image data into the generative model (1003). ) is acquired (1004). The server device 1 presents the output information to the user (1005). .
[0041] Other examples are described below.
[0042] The server device 1 performs preprocessing on the partial image acquired by the generation information acquisition unit 111, for example. The server device 1 may, for example, determine the main subject of the partial image and remove the background other than the main subject portion. The generation information acquisition unit 111 may also, for example, emphasize a main subject from a partial image. .
[0043] As a pre-processing step, the server device 1 performs, for example, a process of determining whether a subject and a color are included in a partial image. The camera angle is detected by determining the positional relationship of the camera, and the generated information acquisition unit 1 11 may generate a prompt, which may be, for example, Examples of the display angle include, but are not limited to, a display angle specifying the angle at which the subject is displayed.
[0044] The generation information acquisition unit 111 performs the generation information acquisition for the partial image after the above-described preprocessing. The location information in the image may be obtained and a prompt may be generated.
[0045] The server device 1 may suggest styles to the user based on marketing information. The marketing information is information about the product that is the subject of the resulting image, etc., that has been acquired in advance, Information on the subject matter, such as information on the world, results of marketing surveys, etc., and information obtained from users. It may also include information such as the product's past sales record and sales of similar products. The device 1 may, for example, use information such as sales records of similar products for the product that is the subject of the image to be generated. Based on the information, the standard is determined from the sales websites and advertising images of similar products with high sales volume. In this case, the server device 1 recommends a similar product with a high sales volume to the user. The web address information of the website and the generated information acquired by the generated information acquisition unit 111 as a style are stored. The user can choose the text information to include in the prompt (e.g., "luxury" or "natural"). This can be presented or included in the prompts used to generate the image.
[0046] The server device 1 stores the resultant image generated by the user using the server device 1 in the past, or the image of interest. Recommends styles to users based on the prompt information used to generate the result image. For example, the server device 1 may analyze result images generated by the user in the past, or The style is determined based on the text information included in the prompt or the The style is presented to the user terminal 3, and the user is prompted to select whether or not to use the style in generating the result image. Specifically, the server device 1 may acquire, for example, the real-time data that the user has used in the past. If it is determined that only realistic images are being generated, the message "Do you want to generate realistic images? Yes "No," etc., is presented to the user terminal 3 via chat or the like, and the user's selection operation is acquired. , a prompt may be generated depending on the answer selected by the user.
[0047] The server device 1 may generate information related to product sales in addition to images. The information generated by 1 may be, for example, banner advertisement images for promoting products or campaigns, commercials, etc. Effective catchphrases and catchphrases that succinctly express the features of the product or brand, Product description, which is text information that describes the detailed explanation and features of the product, The design, layout, and product categories used on the top page of an e-commerce site that sells Category page design, which is the design and display method of the page, specific campaign Design landing pages to highlight your products, social media and advertising platforms This includes, but is not limited to, images and taglines for the platform. When the server device 1 generates the above-mentioned information, the generated information acquisition unit 111 generates the information based on the acquired information. If it is an image or design, the result image generation unit 112 generates the prompt and converts it into an image generation model. Enter a prompt, and if it is text information, use a text generation model (e.g., ChatGP Just input the prompt into a large-scale language model such as T and generate it.
[0048] When acquiring an element image that constitutes a part of a result image, the server device 1 Images contained in the website specified by the web address may be acquired as element images. The server device 1 acquires all images included in the website as element images and generates The image may be stored in the information storage unit 121, or may be generated from the images included in the website. Accepts a user's selection operation for the image to be acquired as information, and displays the selected image It may be stored as an element image.
[0049] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. The technical scope of the disclosure is not limited to such examples. Any person who wishes to make such a modification may make various modifications within the scope of the technical idea described in the claims. It is apparent that various modifications and variations are possible, which are within the scope of the present disclosure. It is understood to be within the scope of science.
[0050] The devices described herein may be implemented as stand-alone devices, or may be implemented in part or in whole as The part is realized by a plurality of devices (e.g., cloud servers) connected via a communication network 2. For example, the processor 101 and the storage device 103 of the server device 1 may may be realized by different servers connected to each other via a communication network 2.
[0051] The series of processes performed by the apparatus described in this specification may be implemented by software, hardware, and This implementation may be implemented using either software or a combination of software and hardware. A computer program for realizing each function of the server device 1 according to the embodiment is created, and It is possible to implement it in C, etc. Also, such a computer program can be stored In addition, a computer-readable recording medium can also be provided. Examples include magnetic disks, optical disks, magneto-optical disks, flash memory, etc. The computer program can be transmitted, for example, via a communication network 2 without using a recording medium. It may be distributed as follows.
[0052] Additionally, the processes described herein do not necessarily have to be performed in the order described. Some processing steps may be performed in parallel, and additional processing steps may be performed in parallel. A step may be adopted, and some processing steps may be omitted.
[0053] Furthermore, the effects described in this specification are merely illustrative or exemplary and are not limiting. In other words, the technology according to the present disclosure has the above-mentioned effects in addition to or instead of the above-mentioned effects. In addition, other effects that will be apparent to those skilled in the art from the description of this specification may be achieved. [Explanation of symbols]
[0054] 1. Server device 2. Communication Network 3. User terminal 101 CPU 102 memory 103 Storage device 104 Communication Interface 105 Input Device 106 Output Device 111 Generation information acquisition unit 112 Result image generation unit 131 Generation information storage unit 132 Image information storage unit
Claims
1. A generation support device that supports generation of a result image, The user provides at least style information regarding the style of the result image and a generation information acquisition unit that acquires generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model a resultant image generating unit that generates the resultant image based on output information output from the A generation support device comprising:
2. The generation information acquisition unit further acquires position information of the element image in the resultant image. thing, The generation support device according to claim 1 .
3. the style information includes web address information; The generation information acquisition unit acquires information about a website included in the website specified by the web address. the style information determined based on the block information is used as the style information; 3. The generation support device according to claim 1 or 2,
4. The information acquisition unit receives upload or selection of the element image, and acquiring the position information by arranging the element images in a frame; The generation support device according to claim 2 .
5. The information acquisition unit acquires the generation information through input by the user in a chat format. To do, 3. The generation support device according to claim 1 or 2,
6. When acquiring the input in the chat format, the information acquisition unit may include, as the generation information: providing the user with an indication of required information; The generation support device according to claim 5 .
7. A generation support program for supporting generation of a resultant image, The processor The user provides at least style information regarding the style of the result image and a generation information acquisition step of acquiring generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model an image generating step of generating the resultant image based on output information output from the A generation support program that executes the above.
8. A generation assistance method for assisting in the generation of a resultant image, comprising: The processor: The user provides at least style information regarding the style of the result image and a generation information acquisition step of acquiring generation information including an element image that constitutes a part of the image; Text information generated based on the generation information is input to a generation model, and the generation model an image generating step of generating the resultant image based on output information output from the A generation support method for executing the above.
Citation Information
Patent Citations
Character image generation device and learning model generation device
JP7169027B1