Image generation device, image generation method, and recording medium
The image generation device addresses the challenge of conveying image details by using existing images and machine learning to generate descriptive text, enabling efficient creation of new images that reflect user intent and enhance product design efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2024-11-29
- Publication Date
- 2026-06-04
AI Technical Summary
Existing image generation technologies require users to input detailed text descriptions, which can be cumbersome and may not accurately convey the desired image, especially when users lack a clear mental image of what they want to create.
An image generation device that accepts existing images as input, uses machine learning models to generate descriptive text about the images, extracts common features, and generates new images based on these features using generative AI models.
Efficiently generates new images that reflect the characteristics of the input images, allowing users to create desired images without needing to write detailed text descriptions, and facilitates rapid development of product designs that incorporate freshness and innovation while maintaining brand consistency.
Smart Images

Figure JP2024042305_04062026_PF_FP_ABST
Abstract
Description
Image generation device, image generation method, and recording medium
[0001] This disclosure relates to an image generation device, an image generation method, and a recording medium.
[0002] Technologies related to this disclosure are disclosed in Patent Document 1. Patent Document 1 discloses a technique for receiving an input of text for generating a design image from a user, identifying corresponding design conditions (conditions that restrict the range of the design), and inputting the identified conditions into a style generator to automatically generate a design image. The text input by the user is, for example, "one-piece dress, black, dots", etc. The style generator is composed of, for example, cGAN (conditional GAN).
[0003] Japanese Patent Application Laid-Open No. 2024-86945
[0004] In the case of a technique for automatically generating an image, by inputting appropriate information regarding the image to be generated into an image generation model, it becomes easier to obtain the desired image. In the technique described in Patent Document 1, the user inputs information regarding the image to be generated in text, but it is not easy to fully express the details of the image to be generated in text.
[0005] An example of the purpose of this disclosure is to provide a technique for efficiently automatically generating a desired image.
[0006] According to one aspect of this disclosure, there is provided an image generation device having: acquisition means for acquiring at least one existing image; text generation means for generating text regarding the characteristics of the existing image based on the existing image and a machine learning model; and image generation means for generating a new image based on the text and an image generation model.
[0007] Also, according to one aspect of this disclosure, there is provided an image generation method in which one or more computers acquire at least one existing image, generate text regarding the characteristics of the existing image based on the existing image and a machine learning model, and generate a new image based on the text and an image generation model.
[0008] Furthermore, according to one aspect of this disclosure, a recording medium is provided which records a program that causes a computer to function as: an acquisition means for acquiring at least one existing image; a text generation means for generating text relating to the features of the existing image based on the existing image and a machine learning model; and an image generation means for generating a new image based on the text and an image generation model.
[0009] According to one example of this disclosure, a technology for efficiently and automatically generating desired images is realized.
[0010] Figure 1 is a diagram showing an example of a functional block diagram of an image generation device. Figure 2 is a flowchart showing an example of the processing flow of an image generation device. Figure 3 is a diagram showing an example of the hardware configuration of an image generation device. Figure 4 is a diagram showing an example of the processing performed by an image generation device. Figure 5 is a flowchart showing another example of the processing flow of an image generation device. Figure 6 is a flowchart showing yet another example of the processing flow of an image generation device. Figure 7 is a diagram showing an example of a UI (user interface) screen output by an image generation device. Figure 8 is a diagram showing yet another example of a UI screen output by an image generation device.
[0011] The embodiments of this disclosure will be described below with reference to the drawings. In this disclosure, the drawings are associated with one or more embodiments. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted where appropriate.
[0012] <<First Embodiment>> Figure 1 is a functional block diagram showing an overview of the image generation device 10. Figure 2 is a flowchart showing an example of the processing flow performed by the image generation device 10.
[0013] As shown in Figure 1, the image generation device 10 includes an acquisition unit 11, a text generation unit 12, and an image generation unit 13. These functional units execute the processes shown in the flowchart of Figure 2.
[0014] In S10, the acquisition unit 11 acquires at least one existing image. In S11, the text generation unit 12 generates text about the features of the existing image based on the existing image acquired in S10 and a machine learning model. In S12, the image generation unit 13 generates a new image based on the text generated in S11 and the image generation model.
[0015] In this way, when the image generation device 10 acquires at least one existing image, it generates a new image accordingly.
[0016] In the case of image generation technology, inputting appropriate information about the image to be generated into the image generation model makes it easier to obtain the desired image. In other words, if appropriate information is not input into the image generation model, there is a possibility that the desired image will not be obtained.
[0017] In the technology disclosed in Patent Document 1, the user inputs information about the image they want to generate in "text" format. However, some users may find it troublesome to generate text that adequately describes the details of the image they want to create. In such cases, the amount of information entered tends to be small. Also, at the stage of image generation, users may not have a clear image in mind of the image they want to generate. For example, they may have a theme such as "I want to generate an image that will appeal to young people," but the specific content of the image may not be clearly defined. In such cases, it is not easy to adequately express the details of the image they want to generate in text.
[0018] To address these challenges, the image generation device 10 has the feature of accepting at least one "image" (an existing image) as input information about the image to be generated. With such an image generation device 10, the user does not need to generate text that fully expresses the details of the image to be generated. The user only needs to input an image similar to the image to be generated, or an image that has a similar atmosphere to the image to be generated, into the image generation device 10. Since an image contains a lot of information, various information about the image can be input with the input of just one image. Also, a user who does not have a clear image in mind of the image to be generated can extract an image that meets the conditions of the image to be generated from existing images and input it into the image generation device 10. For example, if the theme is "I want to generate an image that will appeal to young people," the user can collect images that are popular among young people and input them into the image generation device 10.
[0019] With such an image generation device 10, the user can efficiently generate the desired image using the image generation model.
[0020] Furthermore, when the image generation device 10 acquires at least one existing image, it uses a machine learning model to generate text describing the features of that existing image. Then, the image generation device 10 generates a new image based on that text and the image generation model. With this distinctive processing method, the image generation device 10 can efficiently generate a new image that fully reflects the content of the input at least one existing image. In addition, the user can identify areas for improvement by reviewing the intermediate product, "text describing the features of the existing image."
[0021] As described above, the image generation device 10 realizes a technology for efficiently automatically generating desired images.
[0022] <<Second Embodiment>> <Overview> The image generation device 10 of the second embodiment is a concrete implementation of the configuration of the image generation device 10 of the first embodiment. It will be described in detail below.
[0023] <Hardware Configuration> First, an example of the hardware configuration of the image generation device 10 will be described. Each functional part of the image generation device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various modifications to the implementation method and the device. The software includes programs that are stored in the device from the time of shipment, as well as programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.
[0024] Figure 3 is a block diagram illustrating the hardware configuration of the image generation device 10. As shown in Figure 3, the image generation device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. The peripheral circuitry 4A includes various modules. The image generation device 10 does not necessarily have peripheral circuitry 4A. The image generation device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0025] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to send and receive data to and from each other. The processor 1A is a processing unit such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes interfaces for connecting to a communication network such as the Internet. Input devices include, for example, a keyboard, mouse, microphone, physical buttons, touch panel, etc. Output devices include, for example, a display, projection device, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on their calculation results.
[0026] <Functional Configuration> Next, the functional configuration of the image generation device 10 will be described in detail. Figure 1 is an example of a functional block diagram of the image generation device 10. As shown in the figure, the image generation device 10 has an acquisition unit 11, a text generation unit 12, and an image generation unit 13.
[0027] The acquisition unit 11 acquires at least one (one or more) existing images.
[0028] An "existing image" is an image that already exists at the stage of generating a new image. An existing image may be a still image or a moving image. Furthermore, an existing image may be a photograph taken of a real object with a camera, an image of a hand-drawn illustration converted into digital data, an image generated by image generation software, or any other.
[0029] The existing image may be an image containing a predetermined object. For example, the existing image may be an image containing one predetermined object. The object may be an existing product such as beer, confectionery, clothing, bags, automobiles, books, or calendars.
[0030] In addition, the existing image may be an image that includes a predetermined character. For example, the existing image may be an image that includes one predetermined existing character.
[0031] In addition, existing images may include paintings such as background images or abstract paintings. Furthermore, existing images may include poster images, etc.
[0032] The acquisition unit 11 can acquire existing images input by the user. The user can collect desired images from an existing set of images (for example, a set of images owned by the user, or a set of images publicly available on the internet, etc.) and input the collected images as existing images into the image generation device 10.
[0033] In addition, the user may generate a new existing image and input the generated existing image into the image generation device 10. For example, the user may generate a new existing image by taking a picture of a predetermined object with a camera. Alternatively, the user may create a new hand-drawn illustration and generate a new existing image by converting the hand-drawn illustration into electronic data by means of scanning or taking a picture with a camera. Alternatively, the user may generate a new existing image by drawing a simple illustration with image generation software. The existing image newly generated by hand-drawing or image generation software may be a rough sketch or draft image that shows an outline or concept, such as a "rough drawing".
[0034] The user can input existing images into the image generation device 10, such as images similar to the image to be generated, images with a similar atmosphere to the image to be generated, or draft images showing the outline or concept of the image to be generated, using the means described above. In addition, the user may input images that match the theme into the image generation device 10 as existing images. For example, if the theme is "I want to generate an image that will appeal to young people," the user may collect images that include objects popular with young people, images that include characters popular with young people, or other images popular with young people (paintings, posters, etc.) and input them into the image generation device 10 as existing images.
[0035] The acquisition unit 11 may also acquire existing images using a method different from that used for user input of existing images. The details of this will be explained in the following embodiment.
[0036] The text generation unit 12 generates text about the features of the existing image based on the existing image acquired by the acquisition unit 11 and a machine learning model.
[0037] The text generation unit 12 can perform the following processes: - A first process to generate at least one (one or more) descriptive texts that describe the features of at least one (one or more) existing images. - A second process to extract commonalities among the at least one (one or more) descriptive texts. - A third process to generate requirements for a new image based on the extracted commonalities.
[0038] The following explains each process.
[0039] "A first process that generates at least one descriptive text describing the features of at least one existing image." The features of the existing image are the visual features of the existing image. The features of the existing image may also be, for example, the design features of the existing image (pattern, shape, color, logo, text, texture, composition, etc.). In addition, the features of the existing image may also be the design features of objects or characters contained in the existing image. In addition, if the existing image is a moving image, the features of the existing image may also be the features of the movement shown by the existing image. In addition, the features of the existing image may also be the features of the movement of objects or characters shown by the existing image.
[0040] One example of a machine learning model used in this process is a generative AI (Artificial Intelligence) that combines an image recognition model and a natural language processing model. The image recognition model extracts features from the input image. The natural language processing model generates text (hereinafter sometimes referred to as "explanatory text") that describes the image features extracted by the image recognition model in natural language. In other words, the natural language processing model generates explanatory text that expresses the image features extracted by the image recognition model in language that is easy for humans to understand. An example of an image recognition model is, for example, EXAONE Atelier, but it is not limited to this. An example of a natural language processing model is, for example, mistral-7b-instruct, but it is not limited to this.
[0041] Generative AI can be implemented, for example, by a neural network. A neural network contains multiple artificial neurons, each with synapses connecting them. Each synapse has a weight. When such a neural network receives input, it performs calculations using the weights associated with each synapse and produces an output corresponding to the input. The model representing the connection relationships between neurons and synapses is stored in memory, for example, in software form. Alternatively, the model may be implemented as a dedicated circuit. Similarly, the weights of each synapse are also stored in memory in software form. Alternatively, a circuit representing the weights may be implemented in a dedicated circuit. Note that when constructing generative AI using multiple models, it is not necessarily required that all models be stored in the same memory. There are many different models that use such neural networks. Generative AI can be implemented by adopting and substituting a wide variety of models, such as Transformers, Convolutional Neural Networks (CNNs), and Recurrent Neural Networks (RNNs).
[0042] The text generation unit 12 inputs the existing image acquired by the acquisition unit 11 and a predetermined prompt to the generation AI, causing the generation AI to generate the explanatory text described above. The text generation unit 12 generates the explanatory text described above for each existing image.
[0043] Prompts may be pre-prepared or the system may accept user input. Examples of prompt content include, but are not limited to, "Please describe the design of the existing image in detail" or "Please describe the design of the product included in the existing image in detail." If the existing image is a moving image, examples of prompts include, but are not limited to, "Please describe the movement shown in the existing image in detail" or "Please describe the movement of the object shown in the existing image in detail." Other examples of prompts are described below.
[0044] FIG. 4 shows a conceptual diagram of the first process. As shown in the figure, the text generation unit 12 generates "explanation text describing the features of the existing image P" as shown in FIG. 4(B) from the "existing image P" as shown in FIG. 4(A). In the illustrated example, the text generation unit 12 generates explanation text describing the design of the product included in the existing image P. In the illustrated example, an explanation text with the content "A simple-designed beer can with a large yellow crescent moon displayed in the center of a white background..." is generated.
[0045] The text generation unit 12 may generate one explanation text corresponding to one existing image, or may generate a plurality of explanation texts corresponding to one existing image. In the prompt, the number of explanation texts to be generated corresponding to one existing image may be specified.
[0046] "Second process of extracting common points of at least one explanation text" An example of the machine learning model used in this process is generative AI composed of a large language model. An example of the large language model is, for example, mistral-7b-instruct, etc., but is not limited thereto. The text generation unit 12 inputs at least one explanation text generated in the first process and a predetermined prompt into the generative AI, and causes the generative AI to generate a common point text indicating the common points of at least one explanation text.
[0047] Examples of the content of the prompt include, but are not limited to, "Please explain the common points of the input explanation text", "Please summarize the features common to the input explanation texts", etc. Other examples of the prompt will be described below.
[0048] In the second process, the text generation unit 12 generates a common point text such as "The common point is that it has a simple design. Also,...".
[0049] "A third process that generates requirements for a new image based on extracted commonalities." An example of a machine learning model used in this process is a generative AI composed of a large-scale language model. An example of a large-scale language model is, for example, mistral-7b-instruct, but is not limited to this. The text generation unit 12 inputs the commonalities text generated in the second process and a predetermined prompt to the generative AI, causing the generative AI to generate requirements text indicating the requirements for a new image.
[0050] The generating AI may include some or all of the commonalities indicated in the commonalities text as requirements for the new image. The generating AI may determine the requirements for the new image by considering the frequency of occurrence of each commonality indicated in the commonalities text in the explanatory text, and the importance of each commonality identified by any means.
[0051] The content of the prompts may include, for example, "Generate prompts for the image generation AI based on the input commonalities," or "Generate requirements for new images to include in the prompts for the image generation AI based on the input commonalities," but are not limited to these. Other examples of prompts are explained below.
[0052] As a variation, the text generation unit 12 may perform the second and third processes together. In this case, the text generation unit 12 inputs at least one explanatory text generated in the first process and a predetermined prompt to the generation AI, causing the generation AI to generate requirements text indicating the requirements for the new image. Examples of prompt content include, but are not limited to, "Generate a prompt for the image generation AI based on the commonalities of the input explanatory texts." Other examples of prompts are described below.
[0053] In the third process, the text generation unit 12 generates requirement text, such as "A design displaying cherry blossom branches on a solid-color background. Several pink petals are fluttering around..." The text generation unit 12 may generate one requirement text corresponding to one commonality text, or it may generate multiple requirement texts corresponding to one commonality text. The number of requirement texts to be generated corresponding to one commonality text may be specified in the prompt.
[0054] Returning to Figure 1, the image generation unit 13 generates a new image based on the text generated by the text generation unit 12 and the image generation model.
[0055] An image generation model is a generative AI that generates an image from input text. An image generation model may also accept image input in addition to text. Examples of image generation models include, but are not limited to, Vintedois-Diffusion.
[0056] The image generation unit 13 inputs the requirements text generated in the third process and a predetermined prompt to the generation AI, causing the generation AI to generate a new image. The generation AI can, for example, generate a new image that reflects the characteristics indicated by the requirements text. In addition, the generation AI can generate a new image that reflects the characteristics indicated by the requirements text while adding new creative elements (design elements, etc.).
[0057] The prompt may include, for example, "Generate an image that meets the entered requirements," but is not limited to these examples. Other examples of prompts are described below.
[0058] As another example, the image generation unit 13 may input the requirements text, base image, and predetermined prompt generated in the third process to the generation AI, causing the generation AI to generate a new image. The generation AI can, for example, generate a new image by retaining a part of the design of the base image while changing the other parts to reflect the features indicated by the requirements text. Alternatively, the generation AI can generate a new image that reflects the features indicated by the requirements text while adding new creative elements (design elements, etc.).
[0059] The prompts may include, for example, "Modify part of the product design in the base image to meet the entered requirements," but are not limited to these examples. Other examples of prompts are described below. The base image is the image from which the new image is derived. The generation AI can generate the new image, for example, by modifying the design of the base image.
[0060] The image generation unit 13 may generate one new image corresponding to one requirement text, or it may generate multiple new images corresponding to one requirement text. The number of new images to be generated corresponding to one requirement text may be specified in the prompt.
[0061] The image generation unit 13 can output the newly generated image. The image generation unit 13 can output the new image via an output device such as a display or projection device. The image generation unit 13 may also transmit the new image to an external device. The image generation unit 13 may also save the new image to a predetermined storage device.
[0062] Next, an example of the processing flow of the image generation device 10 will be explained using the flowchart in Figure 5. Note that the purpose here is to explain the processing flow. Details of each process have been described above, so explanations will be omitted here as appropriate.
[0063] First, the image generation device 10 acquires at least one (one or more) existing images (S20). For example, the image generation device 10 acquires at least one existing image input by the user.
[0064] Next, the image generation device 10 generates at least one descriptive text describing the features of each of the at least one existing image acquired in S20 (S21). For example, the image generation device 10 inputs the existing image and a predetermined prompt into a generation AI that combines an image recognition model and a natural language processing model, causing it to generate descriptive text describing the features of the existing image.
[0065] Next, the image generation device 10 extracts commonalities from at least one explanatory text generated in S21 (S22). For example, the image generation device 10 inputs at least one explanatory text generated in S21 and a predetermined prompt into a generation AI composed of a large-scale language model, causing it to generate commonalities text indicating the commonalities.
[0066] Next, the image generation device 10 generates requirements for a new image based on the common points extracted in S22 (S23). For example, the image generation device 10 inputs the common point text generated in S22 and a predetermined prompt into a generation AI composed of a large-scale language model, causing it to generate requirements text indicating the requirements for a new image.
[0067] Next, the image generation device 10 uses the requirements generated in S23 as input text to the image generation model and generates a new image (S24). For example, the image generation device 10 inputs the requirements text generated in S23 and a predetermined prompt to the image generation model and generates a new image that satisfies the requirements.
[0068] <Examples> Here, an example of the image generation apparatus 10 of the second embodiment will be described. Note that the examples described here are merely examples and are not limited to these examples.
[0069] "Example 2-1" In Example 2-1, a user whose job is to generate product designs utilizes the image generation device 10. The user can use the image generation device 10 when considering a new design for a given product. The user has the image generation device 10 generate a new product design. Here, the product is beer, and an example is described in which a new design (label design) for the beer container (can, bottle, etc.) is generated, but the usage is not limited to this.
[0070] First, the user collects existing images to input into the image generation device 10. For example, the user determines the target group. Here, let's assume that "young women" are determined to be the target group. Next, the user extracts at least one beer that is popular with women from existing beers and beers that have been sold in the past. The user can extract beers that are popular with women based on their own knowledge, the results of surveys targeting young women, statistical results of purchase history data, etc.
[0071] The user then collects images of at least one extracted beer by any means and inputs them into the image generation device 10 as existing images. For example, as shown in Figure 4(A), the user can input an image containing the extracted beer as an existing image P into the image generation device 10. In response to this user input, the image generation device 10 performs the following processing.
[0072] The acquisition unit 11 acquires at least one existing image that includes beer (product).
[0073] The text generation unit 12 generates multiple descriptive texts that describe the beer designs (existing designs) contained in at least one existing image. The text generation unit 12 inputs the existing images and predetermined prompts into a generation AI that combines an image recognition model and a natural language processing model, causing it to generate descriptive texts that describe the beer designs (existing designs) contained in each existing image.
[0074] Next, the text generation unit 12 extracts commonalities from the at least one generated descriptive text and generates commonalities text that indicates these commonalities. The commonalities text indicates the commonalities of the beer designs contained in each of the at least one existing input image. The text generation unit 12 inputs the at least one generated descriptive text and predetermined prompts into a generation AI composed of a large-scale language model, causing it to generate the commonalities text.
[0075] Next, the text generation unit 12 generates requirements text indicating the requirements for a new beer design based on the common points text. The text generation unit 12 inputs the generated common points text and predetermined prompts into a generation AI composed of a large-scale language model, causing it to generate requirements text indicating the requirements for a new image. The generation AI can include some or all of the common points indicated in the common points text as requirements for the new image.
[0076] The image generation unit 13 then uses the requirements text as input text for the image generation model to generate a new image including the newly designed beer. The image generation unit 13 inputs the generated requirements text and predetermined prompts into the image generation model, causing the image generation model to generate a new image that satisfies the requirements.
[0077] The image generation unit 13 can output a new image containing the newly generated beer design. The user can then view the outputted new image.
[0078] "Example 2-2" There is a service that sells products with personalized designs for gifts and the like. In Example 2-2, a customer using this service uses the image generation device 10. The user has the image generation device 10 generate a new, personalized design. Here, the product is beer, and an example is given of generating a new design (label design) for the beer container (can, bottle, etc.), but the usage scenario is not limited to this.
[0079] First, the user collects existing images to input into the image generation device 10. For example, the user identifies a beer with a design that suits the attributes and preferences of the person to whom they are giving the gift, and inputs an image containing the identified beer as an existing image into the image generation device 10. Subsequently, the image generation device 10 generates a new image using the same process as in Example 2-1 and outputs it to the user.
[0080] In Example 2-2, the image generation device 10 may function as a server in a server-client system. The image generation device 10 may receive existing images from a client terminal operated by a user and send newly generated images to the client terminal. Examples of client terminals include, but are not limited to, personal computers, smartphones, tablet terminals, and mobile phones.
[0081] <Effects and Effects> The image generation device 10 of the second embodiment can achieve the same effects and effects as the image generation device 10 of the first embodiment.
[0082] Furthermore, the image generation device 10 can accept input of an existing image that includes a predetermined object (product, etc.). In this case, the image generation device 10 can generate a new image that includes a newly designed object based on the design of the object included in the existing image.
[0083] Furthermore, the image generation device 10 can accept input of an existing image containing a predetermined character. In this case, the image generation device 10 can generate a new image containing a newly designed character based on the design of the character included in the existing image.
[0084] Furthermore, the image generation device 10 can accept input of existing images, such as background images or abstract paintings. In this case, the image generation device 10 can generate a new image, which is a painting with a new design, based on the design of the painting in the existing image.
[0085] Furthermore, the image generation device 10 can accept an existing image, which is a poster, as input. In this case, the image generation device 10 can generate a new image, which is a poster with a new design, based on the poster design in the existing image.
[0086] Furthermore, the image generation device 10 can accept input of an existing image that is a moving image. In this case, the image generation device 10 can generate a new moving image (new image) based on the motion characteristics of the existing image.
[0087] Furthermore, the image generation device 10 can acquire multiple existing images, generate multiple descriptive texts describing the features of each, extract common points among the multiple descriptive texts, and generate requirements for a new image based on those common points. With such an image generation device 10, it is possible to generate a new image that has features common to multiple existing images. The user can efficiently generate a new image with those features by simply collecting multiple existing images that have the features they want to include in the new image and inputting them into the image generation device 10.
[0088] Furthermore, the image generation device 10 can generate new images by modifying the design of a base image based on the requirements of the new image. With such an image generation device 10, for example, it is possible to generate new designs that incorporate new perspectives and creativity while utilizing the design characteristics of existing products. This makes it possible to develop product designs that incorporate freshness and innovation while maintaining brand consistency. As a specific example, it is possible to automatically generate product designs for seasonal limited editions or derivative products aimed at young people based on the product design of a long-established beer brand.
[0089] Furthermore, the image generation device 10 enables the product development process to be accelerated and made more efficient. The image generation device 10 allows for the rapid generation of numerous initial product design proposals, a process previously done manually by designers. This allows for earlier evaluation and selection processes by the marketing department and management, significantly shortening the time to commercialization.
[0090] Furthermore, the image generation device 10 can quickly respond to changes in market trends and consumer preferences. By changing existing images input to the image generation device 10 or by periodically or irregularly updating the machine learning model in accordance with changes in market trends and consumer preferences, it becomes possible to constantly generate product designs that reflect the latest trends and consumer preferences. Through these effects, the image generation device 10 greatly contributes to improving the creativity and efficiency of product design and strengthening market competitiveness.
[0091] <<Third Embodiment>> The image generation device 10 of the third embodiment acquires existing images using a method different from that used for inputting existing images by the user. This will be explained in detail below.
[0092] The acquisition unit 11 receives a target group specification from the user. The target group is defined based on a person's attributes. Examples of a person's attributes include, but are not limited to, age, gender, nationality, occupation, hobbies, family structure, etc. For example, the acquisition unit 11 accepts a target group specification such as "men in their 20s".
[0093] The acquisition unit 11 analyzes publicly available data to identify products popular with the target audience specified by the user. Publicly available data includes, for example, data published on the internet or data that can be obtained and used from designated organizations. Examples of publicly available data include, but are not limited to, large datasets that include at least one of the following: social media posts, online product reviews, purchase history data, and sales data.
[0094] The acquisition unit 11 acquires the dataset described above and extracts information on the target group specified by the user from the acquired dataset. Then, the acquisition unit 11 performs analysis processing based on the information on the target group specified by the user to obtain analysis results (information on purchasing trends, preferences, etc.) for the target group specified by the user. The analysis processing is implemented using, for example, text mining techniques or statistical methods, but is not limited to these.
[0095] For example, the acquisition unit 11 uses text mining techniques or statistical methods to extract products and features that are popular with a specific age group or gender. For example, the acquisition unit 11 may identify popular products as those that are most frequently cited with positive expressions in SNS (Social Networking Service) posts by the target group specified by the user (or a predetermined number of products from those with a high number of such citations). Alternatively, the acquisition unit 11 may analyze online product reviews published by the target group specified by the user and identify popular products as those that have received the most positive posts (or a predetermined number of products from those with a high number of such posts). Alternatively, the acquisition unit 11 may analyze purchase history data and identify popular products as those that have been purchased the most by the target group specified by the user (or a predetermined number of products from those with a high number of such purchases). Through such analysis, the acquisition unit 11 can identify specific products, such as X company's craft beer or Y company's low-calorie sparkling alcoholic beverage, as popular products among the target group specified by the user.
[0096] In addition, the acquisition unit 11 may identify products that are popular with the target group specified by the user, based on the sales data of each of the multiple products. This sales data shows the sales of each product for each target group. This sales data may be publicly available on the internet, purchased from a designated organization, or provided by a designated organization. In this case, the acquisition unit 11 can refer to this sales data to identify products that are popular with the target group specified by the user (such as the number one ranked product, or products whose ranking is above a threshold).
[0097] In addition, the acquisition unit 11 may combine the method based on SNS posts and the method based on sales data to identify products popular with the target group specified by the user. There are various ways to combine them. An example is explained below, but it is not limited to this.
[0098] For example, the acquisition unit 11 assigns SNS points according to predetermined rules to "products popular with the target group specified by the user" identified by a method based on SNS posts. The acquisition unit 11 also assigns sales points according to predetermined rules to "products popular with the target group specified by the user" identified by a method based on sales data. Weight values are set in advance for the method based on SNS posts and the method based on sales data. The acquisition unit 11 can calculate the sum of the product of the weight of SNS points and the weight of the method based on SNS posts, and the product of the weight of sales points and the weight of the method based on sales data for each product. The acquisition unit 11 can then identify the product with the highest sum (or a predetermined number of products with the highest sums, etc.) as a popular product.
[0099] After identifying popular products as described above, the acquisition unit 11 acquires images of the identified products as existing images. For example, the acquisition unit 11 may collect images of the identified products from images publicly available on the internet. For example, the acquisition unit 11 may acquire product images from online shops or product images from official websites.
[0100] Next, using the flowchart in Figure 6, we will explain an example of the processing flow of the image generation device 10, or more specifically, an example of the processing flow for acquiring existing images. An example of the processing flow after acquiring existing images is as described in the first and second embodiments. Note that the purpose here is to explain the processing flow. Details of each process have been described above, so explanations will be omitted here as appropriate.
[0101] First, the image generation device 10 receives a target group specification from the user (S30). Next, the image generation device 10 collects information about the target group specified in S30 (S31). For example, the image generation device 10 collects information about the target group (such as messages posted by the target group and sales data of the target group) from the publicly available data mentioned above.
[0102] Next, the image generation device 10 determines the target product based on the information collected in S31 (S32). For example, the image generation device 10 identifies products that are popular with the target group. Then, the image generation device 10 acquires an image of the target product determined in S32 as an existing image (S33).
[0103] Other configurations of the image generation device 10 can be the same as those in the first and second embodiments.
[0104] The image generation device 10 of the third embodiment can achieve the same effects as the image generation devices 10 of the first and second embodiments. Furthermore, the image generation device 10 can automatically collect existing images by utilizing publicly available data. Users can avoid the trouble of collecting existing images themselves. The image generation device 10 of the third embodiment is suitable, for example, when generating new images in a situation where the target layer is defined, but the specific image content is not clearly defined.
[0105] The image generation device 10 can generate new images (product images) suitable for a specific target group. Furthermore, it enables the development of product designs (labels, etc.) that reflect market trends, and is expected to generate product designs that appeal to a wider range of customers. For example, for products aimed at younger generations, it is expected that product designs incorporating more vibrant colors and modern design elements will be generated.
[0106] Furthermore, the image generation device 10 can determine existing product images based on SNS posts. With such an image generation device 10, it is possible to generate new product designs while taking into full consideration user trends and topics of interest.
[0107] Furthermore, the image generation device 10 can determine the product image to be used as an existing image based on the product's sales data. With such an image generation device 10, it is possible to generate new product designs while giving full consideration to highly reliable sales performance.
[0108] Furthermore, the image generation device 10 can determine the product image to be used as an existing image by considering both SNS posts and product sales data. With such an image generation device 10, it is possible to generate new product designs while fully considering both user trends, topicality, and reliable sales performance.
[0109] <<Fourth Embodiment>> The image generation device 10 of the fourth embodiment acquires existing images using a method different from that used for inputting existing images by the user. This will be explained in detail below.
[0110] The acquisition unit 11 receives the user's specification of the target individual. For example, the acquisition unit 11 receives the user identification information of the target individual.
[0111] Next, the acquisition unit 11 acquires information related to the target individual. The information related to the target individual includes, for example, at least one of the following: past purchase history (information indicating purchased products), past browsing history (information indicating products viewed on online shopping sites, etc.), information indicating the user's attributes, the user's health information, and information indicating the user's preferences. Such information related to the target individual may be publicly available on the internet, purchased from a designated organization, or provided by a designated organization.
[0112] Next, the acquisition unit 11 identifies products related to the target individual based on information related to that individual. For example, the acquisition unit 11 can identify at least one of the following as products related to the individual: • Products the individual has purchased in the past • Products the individual has purchased more than a predetermined number of times in the past • Products the individual has viewed in the past • Products the individual has viewed more than a predetermined number of times in the past • Products that match the individual's attributes or preferences
[0113] Searching for products that match an individual's attributes or preferences can be achieved, for example, by using recommendation features available on online shopping sites.
[0114] After identifying products related to the target individual, the acquisition unit 11 acquires images of the identified products as existing images. For example, the acquisition unit 11 may collect images of the identified products from images publicly available on the internet. For example, the acquisition unit 11 may acquire product images from online shops or product images from official websites.
[0115] Next, using the flowchart in Figure 6, we will explain an example of the processing flow of the image generation device 10, or more specifically, an example of the processing flow for acquiring existing images. An example of the processing flow after acquiring existing images is as described in the first and second embodiments. Note that the purpose here is to explain the processing flow. Details of each process have been described above, so explanations will be omitted here as appropriate.
[0116] First, the image generation device 10 receives the designation of an individual to be the user (S30). Next, the image generation device 10 collects information related to the target individual designated in S30 (S31). For example, as information related to the target individual, the image generation device 10 collects at least one of the following: past purchase history (information indicating purchased products), browsing history (information indicating viewed products), information indicating the user's attributes, the user's health information, and information indicating the user's preferences.
[0117] Next, the image generation device 10 determines the target product based on the information collected in S31 (S32). For example, the image generation device 10 can determine the "products related to an individual" as the target product. Then, the image generation device 10 acquires an image of the target product determined in S32 as an existing image (S33).
[0118] Other configurations of the image generation device 10 can be the same as those of the first to third embodiments.
[0119] The image generation device 10 of the fourth embodiment can achieve the same effects as the image generation devices 10 of the first to third embodiments. Furthermore, the image generation device 10 can automatically collect existing images by utilizing information related to the target individual. Users can avoid the trouble of collecting existing images themselves. The image generation device 10 of the third embodiment is suitable, for example, when a target individual is determined, but the specific content of the image is not clearly defined, in order to generate a new image.
[0120] The image generation device 10 can generate new images (product images) that are suitable for a specific individual. In other words, the image generation device 10 can generate new product designs that are suitable for a specific individual and reflect that individual's preferences and needs.
[0121] There are services that sell personalized products with customized designs, such as gifts. The image generation device 10 is used, for example, in such services, when a customer generates a personalized product design.
[0122] Here, we will explain an example of a process that generates new images based on a user's health information. For example, in a personalized supplement business of a health food manufacturer, the image generation device 10 can generate package designs for supplements prescribed to a specific individual based on that individual's health checkup data (blood pressure, cholesterol levels, intestinal bacteria, etc.) and daily activity data (steps, sleep time, etc.).
[0123] Specifically, first, the image generation device 10 structures health checkup data and activity data as numerical data and comparison results with reference values. Then, the image generation device 10 classifies each data item into categories such as "caution required," "appropriate range," and "improvement trend," and then analyzes the correlation between data items to extract characteristics of the health state. At this time, the image generation device 10 may also acquire separately analyzed health state characteristic data.
[0124] Alternatively, the image generation device 10 may use prompts such as the following to have the generating AI analyze the health data.
[0125] ○Example prompt: "Analyze the entered health checkup data and quantify the importance of the following items: blood pressure (high / low / normal), exercise level, and nutritional balance. For each item, quantify the deviation from the normal value on a scale of 0-1, and rearrange the items in order of need for improvement. Also, based on the combination of these health conditions, extract the most important health management points that require attention."
[0126] Next, the image generation device 10 can use prompts like the following to cause the generating AI to convert the extracted health condition characteristics into visual elements such as color, shape, and layout.
[0127] ○Example prompt: "Based on the analyzed health status, generate package design requirements that include the following elements: 1. Main color (considering the psychological effect according to the health status), 2. Sub-color (a color that harmonizes with the main color and can highlight important information), 3. Text placement (font size and placement considering age and eyesight), 4. Graphic elements (visual symbols that intuitively convey the health status). In particular, if blood pressure is elevated, use blue tones to promote psychological calmness, and if there is a lack of exercise, incorporate dynamic elements to encourage activity."
[0128] In this way, for example, for users with slightly elevated blood pressure, a design based on a calming blue color scheme can be generated to provide psychological reassurance, while for users who lack exercise, a design incorporating dynamic graphic elements that evoke an active image can be generated.
[0129] Here, when the image generation device 10 associates health conditions with design elements as described above, it can refer to external databases such as a color psychology database, a character design database, and medical design guidelines. For example, it can obtain scientific knowledge about the psychological and physiological effects that specific colors have on the human body from the color psychology database, and obtain recommended values for optimal font size and contrast ratio according to age and symptoms from the medical design guidelines.
[0130] As another example, the image generation device 10 can also generate character designs based on gut microbiota data. For instance, to generate a mascot character that reflects the state of an individual's gut microbiota for a probiotic product, the image generation device 10 can perform the following processing.
[0131] First, let's assume that for each major group of intestinal bacteria (e.g., lactic acid bacteria, bifidobacteria, bacteroides, etc.), basic character elements (shape, expression, decoration, etc.) are registered in a character design database. Specifically, we can define the following correspondences: • Lactic acid bacteria group: rounded, gentle shape, smiling expression, milky white body color • Bifidobacteria group: slender shape, lively expression, light blue body color • Bacteroides group: sturdy shape, serious expression, brown body color
[0132] Next, the system analyzes the relative abundance and activity levels of each group from the individual's gut microbiota data, and combines character elements based on the results. For example, if lactic acid bacteria are dominant, it generates an integrated character design that primarily uses the characteristics of the lactic acid bacteria group while partially incorporating the characteristics of other groups. Furthermore, if there is high bacterial diversity, it is possible to generate a colonial-type design that combines multiple characters. Specifically, the image generation device 10 can utilize the following prompts when generating such character designs.
[0133] ○Example Prompt: "Based on the state of the gut microbiota shown in the following <Data>, please generate a friendly mascot character. The character's features should follow the following <Design Rules>." <Data> ・Lactobacillus group: Abundance ratio 20% (reference value 15%), Activity level: High ・Bifidobacterium group: Abundance ratio 8% (reference value 10%), Activity level: Medium ・Bacteroides group: Abundance ratio 25% (reference value 25%), Activity level: High ・Overall gut environment score: 85 points (good) <Design Rules> 1. The higher the abundance ratio of a major bacterial group, the more strongly the characteristics of that group (shape, color, etc.) should be reflected. 2. Adjust the character's facial expressions and the liveliness of its movements according to the activity level. 3. Adjust the overall impression according to the overall score. 4. Generate as an animable 2D character.
[0134] Furthermore, the image generation device 10 can use any data as long as it corresponds to text. Here, we will explain using the example of using data related to taste and smell as information indicating individual preferences.
[0135] For example, data related to an individual's sense of smell and preferences is typically structured as a group of odor information containing multiple odor data points, each of which is obtained by quantifying the detection signals from various sensor elements. Specifically, odor information is generated by quantifying the detection signals from multiple sensor elements, each exhibiting a unique adsorption reaction to different odor substances, using a computing device. Odor information is generated by, for example, calculating the difference between the maximum value and the minimum value immediately following the maximum value of the detection signal from the raw data detected by the sensor elements, and then performing a logarithmic operation on that difference. The multiple odor information points obtained in this way are structured as an odor information group for a single odor and stored in a database along with supplementary information such as explanatory text.
[0136] The image generation device 10 can utilize the data structure of such scent information to enable a more precise analysis of individual preferences. For example, the image generation device 10 can accept multiple beverages preferred by an individual, extract common features from the scent information contained in those beverages, and automatically select design elements (color, shape, composition, etc.) that reflect those features.
[0137] More specifically, the image generation device 10 can use prompts like the following to allow the generating AI to analyze the features.
[0138] ○Example prompt: "For the scent information contained in the multiple beverages entered, analyze the descriptive text stored in the information database (e.g., 'A refreshing and sweet citrus scent. Characterized by lemon in the top notes, orange in the middle notes, and a subtle floral scent in the base notes') and the scent information group for each beverage, and extract common characteristic elements and intensities."
[0139] Furthermore, the image generation device 10 can use prompts like the following to allow the generating AI to convert the analyzed features into design elements.
[0140] ○Example prompt: "Based on the extracted features and the content of the descriptive text, convert them into design elements such as color, shape, composition, and texture. When doing so, please also consider the general correspondence between the expressions included in the descriptive text (e.g., refreshing, sweet, floral) and the design elements."
[0141] For example, in the case of an individual who prefers citrus beverages and fruit teas, the database might reference descriptive texts and scent information sets for each, and common characteristics such as "citrusy freshness (intensity 0.8)", "fruity sweetness (intensity 0.6)", and "floral aftertaste (intensity 0.3)" might be extracted. Then, based on the correspondence with the expressions contained in the descriptive text, design elements such as a bright color scheme accented with orange (#FFA500), a curved and dynamic organic form, a refreshing layout based on an upward-sloping diagonal flow, and a glossy, smooth texture might be selected.
[0142] By using corresponding descriptive text even for data such as smell and taste, it becomes possible to reflect individual preferences in the design based on a more accurate and semantic interpretation.
[0143] The image generation device 10 makes it possible to generate product designs that directly reflect the user's personal preferences and desires. This effectively addresses the demand for personalized product designs, such as gifts and commemorative items. Furthermore, by generating new product designs based on existing images selected for each individual, it becomes possible to create more satisfying product designs that capture the user's latent preferences and aesthetic sense. In addition, since personalized product designs can be generated without the involvement of professional designers, costs and production time can be reduced. Due to these effects, the image generation device 10 contributes to strengthening competitiveness and improving customer satisfaction in the personalized custom product market.
[0144] <<Fifth Embodiment>> The image generation device 10 of the fifth embodiment has the function of presenting at least one newly generated image to the user, receiving user input to select at least one from among them, and generating a new new image based on the selected at least one newly generated image. This will be described in detail below.
[0145] The image generation unit 13 presents the user with at least one newly generated image. For example, the image generation unit 13 outputs a UI (User Interface) screen containing the at least one newly generated image via an output device such as a display. Alternatively, the image generation unit 13 may transmit the UI screen to a client terminal. The image generation unit 13 may also display the at least one newly generated image in an order of priority determined by a predetermined rule.
[0146] The image generation unit 13 receives user input via the UI screen, allowing the user to select at least one of the at least one newly generated image. The image generation unit 13 can receive this user input via UI components displayed on the UI screen.
[0147] The image generation unit 13 inputs at least one new image selected by user input to the acquisition unit 11 as an existing image. The acquisition unit 11 acquires at least one new image selected by user input as an existing image.
[0148] The text generation unit 12 generates new text (explanatory text, commonalities text, requirements text, etc.) relating to the features of the existing image, based on the existing image newly acquired by the acquisition unit 11 and a machine learning model. Then, the image generation unit 13 generates a new image based on the newly generated text and the image generation model. These processes of the text generation unit 12 and the image generation unit 13 can be the same as in the first to fourth embodiments.
[0149] The image generation device 10 may repeat the above-described process multiple times.
[0150] Other configurations of the image generation device 10 can be the same as those of the first to fourth embodiments.
[0151] The image generation device 10 of the fifth embodiment can achieve the same effects and advantages as the image generation device 10 of the first to fourth embodiments. Furthermore, the image generation device 10 can generate new images by using the newly generated images as new existing images. By repeating this process, it is expected that the design of the newly generated images will converge to the user's wishes.
[0152] <<Sixth Embodiment>> The image generation device 10 of the sixth embodiment can cause the generation AI to perform various processes using characteristic prompts. Various restrictions can be imposed by the characteristic prompts. These will be explained in detail below.
[0153] "Example of a prompt used in the process of acquiring existing images" As described in the third embodiment, the acquisition unit 11 may, after receiving a target layer specification from the user, cause the large-scale language model to consider product images to be collected as existing images.
[0154] The acquisition unit 11 inputs information indicating the target group specified by the user and predetermined prompts into a large-scale language model, which can then be made to consider product images that should be collected as existing images.
[0155] Prompts can be pre-defined or they can accept user input. If the target audience is "women in their 20s" and the purpose of generating new images is "to generate label designs for plastic bottles," an example of a prompt might be the following:
[0156] ○Example prompt: "Generate a search query for products to collect in order to analyze <Objective>. <Objective> Label designs for PET bottles that sell well to women in their 20s."
[0157] When such a prompt is input into a large-scale language model, one can expect to obtain a response like the following:
[0158] ○Example answer: "Please search for the hashtag "#○○" on social media."
[0159] As a variation, the prompt example above, "Generate a search query for products to be collected in order to analyze <Purpose>," may be changed to "Search for images of products to be collected in order to analyze <Purpose>." In this case, the generative AI, which combines a large-scale language model and image search technology, generates a search method as shown in the answer example above, and then executes that search method to search for images.
[0160] "Examples of prompts used in a first process that generates at least one descriptive text describing the features of at least one existing image" Examples of prompts used in the first process include, as described in the second embodiment, "Please describe in detail the design of the existing image" and "Please describe in detail the design of the product included in the existing image." In addition, if the existing image is a moving image, examples include "Please describe in detail the movement shown in the existing image" and "Please describe in detail the movement of the object shown in the existing image."
[0161] Furthermore, by carefully crafting the prompts, it is possible to have the AI generate explanatory text that better meets the user's needs.
[0162] For example, a prompt can be modified to include more specific points of focus or instructions. Specifically, the prompt may include at least one of the following pieces of information: • Information identifying the object from which information is to be extracted (name, product type, etc.) • Information indicating elements of greater importance (design elements such as color or symbols) • Whether or not text should be read • Information specifying areas of the product in an existing image that are of greater importance (other than text information such as nutrients or expiration date, such as the top of the box)
[0163] As another example, the prompt may include instructions specifying the information to be included in the generated descriptive text. For example, the prompt may include instructions to guess the design concept and include the result of that guess in the descriptive text. The design concept may include, but is not limited to, appeal points. Specifically, the prompt may include instructions to include at least one of the following guesses in the descriptive text: • What is the main appeal of the existing image design? • What is the main message you want to convey with the existing image design? • Why do you think the existing image design was chosen?
[0164] An example of a prompt might be the following:
[0165] ○Example prompt: "Please describe the design of the existing image in detail, focusing on the key points indicated by <Key Points>. <Key Points> - No need to read colors or text - What is the most appealing aspect of the existing image's design?"
[0166] "Examples of prompts used in the second process of extracting commonalities from at least one descriptive text" Examples of prompts used in the second process include, as described in the second embodiment, "Please describe the commonalities of the input descriptive texts" and "Please summarize the common features of the input descriptive texts."
[0167] Furthermore, by carefully crafting the prompts, it is possible to have the AI generate commonalities text that is more relevant to the user's needs.
[0168] For example, a prompt could include a sample combining multiple descriptive texts and commonalities text generated from those descriptive texts. This sample is a "good example" of how commonalities text is generated according to the user's wishes.
[0169] In addition, you may add more specific points or instructions to the prompt. Specifically, the prompt may include at least one of the following pieces of information: • Information identifying the object from which you want to extract information (name, product type, etc.) • Information indicating elements of greater importance (design elements such as color or symbols) • Whether or not to read text • Information specifying areas of the product in an existing image that are of greater importance (other than text information such as nutrients and expiration dates, such as the top of the box)
[0170] Additionally, the prompt may include the reason why you want to focus on the above points. For example, the prompt may include instructions such as, "In recent years, the importance of color in product design has been increasing. Therefore, please focus on color and extract commonalities."
[0171] An example of a prompt might be the following:
[0172] ○Prompt Example: "Refer to the example of extracting common points shown in the <Sample> below, and paying attention to the <Points of Focus> below, explain the common points of explanation texts 1 through 5 shown in the <Explanatory Text>. <Points of Focus> - No need to read colors or characters - What is the most appealing point of the existing image design? <Explanatory Text> - Explanation text 1: A yellow crescent moon in the center of a white background... - Explanation text 2: ... - Explanation text 3: ... - Explanation text 4: ... - Explanation text 5: ... <Sample> - Explanation text 1': ... - Explanation text 2': ... - Explanation text 3': ... - Explanation text 4': ... - Explanation text 5': ... - Common point: The common point of these is..."
[0173] "Examples of prompts used in the third process for generating requirements for a new image based on extracted commonalities" Examples of prompts used in the third process include, as explained in the second embodiment, "Please generate requirements for a new image to be included in the prompts of the image generation AI, referring to the input commonalities."
[0174] Furthermore, by carefully crafting the prompts, it is possible to have the AI generate requirements text that better reflects the user's desires.
[0175] For example, a prompt can include information specifying the items to be included in the requirements. For instance, a prompt might specify that at least one of the following items should be included in the requirements: • Color • Logo • Concept • Number of concepts to generate • Instructions to provide a catchphrase and justification • Instructions to describe the flavor, image, and target audience the generation prompt evokes • Characteristics of the product for which you want to generate a new design (e.g., alcohol content, nutrients, selling points, catchphrase) • Concept restrictions (e.g., Japanese style, futuristic style)
[0176] In another example, the prompt may include information about the target person. For example, the prompt may instruct the system to generate requirements for a new image directed at a target having the following characteristics: For example, the user can input the following information into the image generation device 10: - The target person's language comprehension ability (information indicating the languages they understand, etc.) - The target person's age - The target person's health status - News that the target person is interested in
[0177] An example of a prompt might be the following:
[0178] ○Example Prompt: "Create five image generation prompts to create a new package design and product concept based on the following <common points>. In this case, please also include a catchphrase and the rationale for its creation. Please also explain the flavor, image, and customer base that this image generation prompt evokes. The prompts for generating images should also mention color, logo, and concept. The beer description includes the following information: ・About the product ・Appearance: An image of a can with the logo in the center. The product is clearly positioned on a white background. The image shows a round logo on the side of a beer can. ・Description: We aimed for a balance of malty flavor, a crisp aftertaste, and rich foam. Please experience our perfectly polished draft beer. ・Ingredients: Malt (imported or domestic (less than 5%)), hops, rice, corn, starch ・Country of origin of ingredients: Examples of malt production regions: Germany, France, Denmark, Canada, Australia, Japan, etc. ・Alcohol content (%): 5 ・Pure alcohol content (per 100ml) (g): 4 • Purine content (per 100ml) (mg): Approximately 7.5 • Energy (per 100ml) (kcal): 40 • Type: Beer • Volume: 135ml can, 250ml can, 350ml can, 500ml can, 334ml bottle, 500ml bottle, 633ml bottle • Sales area: Nationwide • Sales method: Year-round sales • Catchphrase: A beer that pursues the deliciousness of fresh beer.
[0179] "Examples of prompts used in the process of generating a new image using an image generation model" Examples of prompts used in this process include, as explained in the second embodiment, "Please generate an image that satisfies the input requirements."
[0180] Furthermore, by carefully crafting the prompts, it is possible to have the AI generate new images that better meet the user's needs.
[0181] For example, the prompt may include information indicating the orientation of the new image to be generated (image-related instructions). Examples of such information include, but are not limited to, "Generate an image that looks delicious."
[0182] "Example of other prompts 1" The image generation device 10 may accept feedback input from the user regarding the newly generated image. The image generation device 10 may then generate the newly generated image and a prompt that includes instructions to modify the newly generated image based on the user's feedback. The user's feedback may be, for example, "Please make it look a little more appetizing." but is not limited to this. The image generation device 10 can input the prompt into the image generation model (image generation AI) and have it generate a new image.
[0183] "Example of other prompts 2" The image generation device 10 can include various information based on user input in any of the prompts described above.
[0184] For example, when the image generation device 10 receives user input specifying the attributes (age group, gender, etc.) of the target group for a newly generated image, it generates relevant information about that target group. The image generation device 10 then receives user input to select at least one of the generated relevant information. For example, when the image generation device 10 receives the specification of a target group of men aged 25 to 40, it generates relevant information about that target group using means such as online search or the use of a large-scale language model. For example, the image generation device 10 generates relevant information such as "characteristics (high environmental awareness, high health awareness, active lifestyle), purchasing trends (organic products, sustainable products)." The image generation device 10 then presents this relevant information to the user and receives user input to select at least one of them.
[0185] The image generation device 10 can include prompts that contain instructions to generate descriptive text, commonalities text, requirements text, or a new image suitable for a target layer having features indicated by the relevant information selected by user input.
[0186] Other configurations of the image generation device 10 can be the same as those of the first to fifth embodiments.
[0187] The image generation device 10 of the sixth embodiment can achieve the same effects as the image generation device 10 of the first to fifth embodiments. Furthermore, by utilizing the characteristic prompts described above, the image generation device 10 can cause the AI to generate information that better matches the user's preferences. As a result, the user can efficiently generate the desired image.
[0188] <Examples> Here, an example of the image generation apparatus 10 of the sixth embodiment will be described. Note that the examples described here are merely examples and are not limited to these examples.
[0189] "Example 6-1" The theme of this example is to generate label designs for seasonal beers. In this example, seasonal designs are generated based on seasonal words.
[0190] The user inputs a theme such as "seasonal beer" or "limited-time beer" into the image generation device 10. The acquisition unit 11 performs an image search using keywords such as "seasonal," "limited-time," "seasonal beer," and "limited-time beer." For example, the acquisition unit 11 may perform a web search or search within a storage device specified by the user. The acquisition unit 11 acquires the images included in the search results as existing images. The text generation unit 12 generates explanatory text, commonalities text, requirements text, etc., from these existing images.
[0191] As an alternative, the acquisition unit 11 may perform an image search using the product name "beer" as a keyword. For example, the acquisition unit 11 may perform a web search or search within a storage device specified by the user. The acquisition unit 11 acquires the images included in the search results as existing images. The text generation unit 12 generates descriptive text from these existing images. In this case, the text generation unit 12 can extract descriptive texts containing words such as "seasonal limited" or "limited time," or similar words, from among the multiple descriptive texts generated. The text generation unit 12 can then generate commonalities text from the extracted descriptive texts and generate requirements text from the generated commonalities text. The text generation unit 12 can identify similar words for each word based on a pre-prepared similar word dictionary. The similar word dictionary is a dictionary that registers similar words for each of multiple words.
[0192] Furthermore, the image generation device 10 may classify the text by seasonal words based on the product text and seasonal words related to the release date, and output the seasonal words in order from most frequently occurring to least frequently occurring. In addition, the image generation device 10 may regenerate common features of designs that are highly related to the seasonal words based on the selected seasonal words, and generate specific prompts for image generation.
[0193] The image generation device 10 may also use product release date information and a seasonal word database to classify the descriptive text generated from existing images based on seasonality and generate a list of seasonal words ranked according to their frequency of appearance.
[0194] The image generation device 10 can take a seasonal word selected by the user as input, extract visual features (color, motif, composition, etc.) that are highly related to that seasonal word, and automatically generate an image generation prompt incorporating those features.
[0195] For example, when a product image is input, the image generation device 10 extracts seasonal words such as "cherry blossoms," "spring breeze," and "young leaves" from existing images such as "Sakura Beer" and "Spring Limited Beer," and identifies the most important seasonal words based on their frequency of appearance (e.g., "cherry blossoms" 8 times, "spring breeze" 5 times).
[0196] The image generation device 10 can then use correlation data between seasonality and design elements, pre-trained by a machine learning model, to extract specific colors (such as light pink #FFB7C5, light green #89C997, etc.), motifs (such as petals, drooping branches, etc.), and compositions (such as diagonal flow) from a selected seasonal word (e.g., "cherry blossoms"), and generate an image generation prompt incorporating these elements.
[0197] Here is an example of a UI screen output by the image generation device 10.
[0198] In the UI screen shown in Figure 7, the theme specified by the user is displayed in the "New Products Under Planning" section.
[0199] Furthermore, in the UI screen shown in Figure 7, the "Expected Target" column displays information about the specified target. Some or all of the information about the specified target is information specified by the user. Note that some of the information about the specified target may be information collected by the acquisition unit 11 through web searches or the like.
[0200] In the UI screen shown in Figure 7, the "Generated Design Policy" column displays the requirements text (requirements for the new image) generated by the text generation unit 12.
[0201] The user can press the "Regenerate Policy" button to have the requirements text regenerated. The image generation device 10 will regenerate the requirements text in response to this instruction. Specifically, the image generation device 10 will regenerate the same explanatory text for each existing image as before, regenerate the common points text from those explanatory texts, and then newly generate the requirements text from those common points texts. The content of the prompts when generating each text may be the same as before or may be changed. In addition, during this regeneration process, the image generation device 10 may change the prompts used in previous processes based on predetermined rules and use the changed prompts in the current process.
[0202] Furthermore, users can input corrections by pressing the "Send Feedback" button. For example, when the "Send Feedback" button is pressed, a separate window opens displaying a UI component that accepts input of corrections. The user inputs the corrections through this UI component. The image generation device 10 corrects the requirements text in accordance with these instructions. For example, the image generation device 10 may input a prompt into a large-scale language model that includes the current requirements text, the corrections entered by the user, and instructions to correct the current requirements text according to the corrections, thereby correcting the current requirements text.
[0203] Additionally, the user can generate a new image by pressing the "Generate Image" button. The image generation unit 13 generates a new image based on the requirements text displayed in the "Generated Design Policy" field in response to this instruction.
[0204] The UI screen in Figure 8 displays a list of new images generated by the image generation unit 13. Pressing the "Details" button corresponding to each image displays the requirements text used to generate that image, the design concept, target audience information, and a detailed explanation of the colors and design elements used. Pressing the "Star Mark" button corresponding to each image marks that image as a favorite and saves it for use as a reference image for future image generation. The user can press the "Regenerate Design" button to regenerate a new image. The image generation unit 13 can generate a new image based on the same requirements text used to generate the displayed new image in response to this instruction. In this regeneration process, the image generation device 10 may change the prompt used in the previous process according to predetermined rules and use the changed prompt in the current process.
[0205] "Example 6-2" The theme of this example is to generate product designs for elderly people who are interested in health foods. The target audience is, for example, people aged 60 and over who are interested in health foods, energy supplements, natural products, etc.
[0206] For example, a user inputs multiple product images related to health foods (protein bars, vitamin supplements, gluten-free cookies, etc.) as existing images into the image generation device 10. From these existing images, it is expected that descriptive text exhibiting the following characteristics will be generated: • Natural colors (green, beige) • Fonts symbolizing health and energy • Simple and easy-to-read design • Soft fonts and natural textures
[0207] The following requirements text is expected to be generated.
[0208] "We will create product labels for health supplements aimed at seniors. We will use natural colors such as green and beige, and a simple, easy-to-read layout. Soft, rounded fonts will be used to convey a friendly and gentle feel. The design will include natural textures or wood-grain backgrounds to emphasize an organic and healthy lifestyle. Strong yet subtle typographic elements will highlight the product's energy-boosting effect, appealing to those interested in maintaining vitality in later life."
[0209] "Example 6-3" The theme of this example is to generate product designs for young people who are interested in environmental issues. The target audience is, for example, teenagers and people in their twenties who are interested in environmental issues, sustainability, ethical consumption, etc.
[0210] For example, a user inputs several product images related to environmental issues (organic cotton T-shirts, recycled material bags, bamboo toothbrushes, etc.) as existing images into the image generation device 10. From these existing images, it is expected that descriptive text will be generated that exhibits the following characteristics: • Emphasis on environmentally friendly materials • Simple and easy-to-understand design • Natural colors (green, blue, beige) • Symbols such as recycling marks
[0211] The following requirements text is expected to be generated.
[0212] "We will create labels for environmentally friendly products aimed at young people interested in environmental issues. The labels will use natural colors such as green and blue, reminiscent of the Earth, and effectively incorporate symbols such as recycling marks. The font will be simple and easy to read, and will include a message that emphasizes environmental awareness. The design will clearly indicate the use of environmentally conscious materials such as organic cotton and recycled materials, creating a design that resonates with young people."
[0213] "Example 6-4" The theme of this example is to generate product designs for women with a preference for luxury goods. The target audience is women in their 30s and 40s who are interested in luxury goods, beauty products, fashion, etc.
[0214] For example, a user inputs multiple product images of luxury items (luxury perfumes, designer bags, organic cosmetics, etc.) as existing images into the image generation device 10. From such existing images, it is expected that descriptive text exhibiting the following characteristics will be generated: • Luxurious colors (gold, black, pastel colors) • Elegant and sophisticated fonts • Simple layout • High-quality material feel
[0215] The following requirements text is expected to be generated.
[0216] "We create high-end product labels for discerning women. We use sophisticated colors such as gold and black as the base, along with elegant and refined fonts. A simple layout effectively places the brand logo, conveying a sense of high-quality materials. Delicate illustrations and embellishments are added to emphasize femininity and beauty."
[0217] <<Modifications>> Below, modifications applicable to the first to sixth embodiments are described. Note that the same effects and advantages as those of the first to sixth embodiments are achieved in the following modifications as well.
[0218] <First Modification> The text generation unit 12 may generate the requirements text after generating the explanatory text, without generating the common points text. For example, the text generation unit 12 may perform this process when one existing image is acquired and one explanatory text is generated.
[0219] In this example, the text generation unit 12 generates one descriptive text from one existing image. The text generation unit 12 then inputs this descriptive text and a predetermined prompt into a large-scale language model (generating AI) to generate the required text. The prompt may be prepared in advance or it may be provided as an instruction from the user. Examples of prompt content include, but are not limited to, "Please generate a prompt for the image generation AI, referring to the input descriptive text."
[0220] <Second Modification> The image generation device 10 may present the generated descriptive texts to the user after generating descriptive texts for at least one existing image, but before generating common point texts from those descriptive texts. Presentation to the user is achieved by outputting information via an output device or by transmitting information to a client terminal. The image generation device 10 may then generate common point texts from the descriptive texts in response to receiving an instruction from the user to generate common point texts. The image generation device 10 may also accept user input to modify the presented descriptive texts. The image generation device 10 may then process the existing images again with the generation AI in response to this user input, regenerate the descriptive texts, and present them to the user.
[0221] Furthermore, the image generation device 10 may present the generated common points text to the user after generating the common points text, but before generating the requirements text from the common points text. The image generation device 10 may then generate the requirements text from the common points text in response to receiving an instruction from the user to generate the requirements text. The image generation device 10 may also accept user input to modify the presented common points text. In response to this user input, the image generation device 10 may process the existing image description text again with the generation AI to generate new common points text and present it to the user.
[0222] Furthermore, the image generation device 10 may present the generated requirement text to the user after generating the requirement text, but before generating a new image from that requirement text. Then, the image generation device 10 may generate a new image from the requirement text in response to receiving an instruction from the user to generate a new image. The image generation device 10 may also accept user input to modify the presented requirement text. Then, in response to this user input, the image generation device 10 may process the common points text again with the generation AI to generate new requirement text and present it to the user.
[0223] Although this disclosure has been described above with reference to embodiments, this disclosure is not limited to the embodiments described above. Various modifications to the structure and details of this disclosure are possible, which can be understood by those skilled in the art within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0224] Furthermore, the flowchart used in the above explanation shows multiple steps (processes) in sequence. However, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content.
[0225] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An image generation method in which one or more computers acquire at least one existing image, generate text relating to the features of the existing image based on the existing image and a machine learning model, and generate a new image based on the text and an image generation model. 2. The image generation method according to 1, wherein one or more computers output at least one of the new images generated based on the text and an image generation model, and accept user input to select at least one from the at least one of the new images. 3. The image generation method according to 2, wherein one or more computers newly acquire the at least one new image selected by the user input as the existing image. 4. The image generation method according to 3, wherein one or more computers newly generate text relating to the features of the newly acquired existing image based on the newly acquired existing image and the machine learning model, and generate a new image based on the newly generated text and the image generation model. 5. 1. An image generation method according to any one of 1 to 4, wherein one or more computers acquire multiple existing images, generate multiple descriptive texts describing the features of each of the multiple existing images, extract commonalities among the multiple descriptive texts, generate requirements for the new image based on the commonalities, and use the requirements as input text to the image generation model. 6. An image generation method according to any one of 1 to 4, wherein one or more computers acquire one existing image, generate one descriptive text describing the features of one existing image, generate requirements for the new image based on the descriptive text, and use the requirements as input text to the image generation model. 7. An image generation method according to 5 or 6, wherein one or more computers generate the requirements for the new image based on at least one of the target person's language comprehension ability, age, health condition, and news of interest.8. The image generation method according to any one of 1 to 7, wherein one or more computers acquire an image specified by user input as an existing image. 9. The image generation method according to any one of 1 to 7, wherein one or more computers receive user input specifying a target group, acquire information on the specified target group, identify products popular with the specified target group based on the acquired information on the target group, and acquire images of the identified products as existing images. 10. The image generation method according to 9, wherein one or more computers acquire information on the specified target group from publicly available data. 11. The image generation method according to 9 or 10, wherein the information on the specified target group indicates at least one of the following: social media posts by the target group, online product reviews published by the target group, purchase history data of the target group, and sales data of the target group. 12. The image generation method according to any one of 1 to 7, wherein one or more computers receive user input specifying a target individual, acquire information related to the individual designated as a target, identify products related to the individual designated as a target based on the acquired information, and acquire images of the identified products as existing images. 13. The image generation method according to 12, wherein one or more computers obtain information related to an individual designated as a target from publicly available data. 14. The image generation method according to 12 or 13, wherein the information related to the individual designated as a target includes at least one of the following: the individual's past purchase history, the individual's past browsing history, information indicating the individual's attributes, the individual's health information, and information indicating the individual's preferences.15. An image generation method according to any one of 12 to 14, wherein one or more computers identify at least one of the following as products related to the individual designated as a target: products that the individual has purchased in the past, products that the individual has purchased more than a predetermined number of times in the past, products that the individual has viewed in the past, products that the individual has viewed more than a predetermined number of times in the past, and products that suit the individual's attributes or preferences. 16. An image generation method according to any one of 1 to 15, wherein one or more computers acquire at least one existing image containing a product, generate at least one descriptive text describing the existing design of the product contained in each of the at least one existing image, extract commonalities among the at least one descriptive text, generate requirements for a new design of the product based on the commonalities, and use the requirements as input text to the image generation model to generate a new image containing the product with a new design. 17. An image generation method according to any one of 1 to 16, wherein the machine learning model is a combination of an image recognition model and a natural language processing model. 18. An image generation method according to any one of 1 to 17, wherein one or more computers acquire an image containing a predetermined object, an image containing a predetermined character, an image of a painting, or an image of a poster as the existing image. 19. An image generation apparatus having: acquisition means for acquiring at least one existing image; text generation means for generating text relating to the features of the existing image based on the existing image and a machine learning model; and image generation means for generating a new image based on the text and an image generation model. 20. A recording medium recording a program that causes a computer to function as: acquisition means for acquiring at least one existing image; text generation means for generating text relating to the features of the existing image based on the existing image and a machine learning model; and image generation means for generating a new image based on the text and an image generation model.
[0226] Some or all of the appendices 2 to 18, which are dependent on the image generation method described in appendice 1 above, may also be dependent on the image generation device in appendice 19 and the recording medium in appendice 20 in the same dependent relationship as between appendice 1 and appendices 2 to 18. Furthermore, without departing from the embodiments described above, some or all of the configurations described as appendices can be realized in various hardware, software, various recording means for recording software, or systems.
[0227] 10 Image generation device 11 Acquisition unit 12 Text generation unit 13 Image generation unit 1A Processor 2A Memory 3A Input / Output I / F 4A Peripheral circuitry 5A Bus
Claims
1. An image generation method comprising: one or more computers acquiring at least one existing image; generating text relating to the features of the existing image based on the existing image and a machine learning model; and generating a new image based on the text and an image generation model.
2. The image generation method according to claim 1, wherein one or more computers output at least one new image generated based on the text and image generation model, and accept user input to select at least one from the at least one new image.
3. The image generation method according to claim 2, wherein one or more computers newly acquire at least one of the new images selected by the user input as the existing image.
4. The image generation method according to claim 3, wherein one or more computers generate new text relating to the features of the newly acquired existing image based on the newly acquired existing image and the machine learning model, and generate the new image based on the newly generated text and the image generation model.
5. The image generation method according to any one of claims 1 to 4, wherein one or more computers acquire a plurality of existing images, generate a plurality of descriptive texts describing the characteristics of each of the plurality of existing images, extract commonalities among the plurality of descriptive texts, generate requirements for the new image based on the commonalities, and use the requirements as input text for the image generation model.
6. The image generation method according to any one of claims 1 to 4, wherein one or more computers acquire one existing image, generate one descriptive text describing the features of one existing image, generate requirements for the new image based on the descriptive text, and use the requirements as input text for the image generation model.
7. The image generation method according to claim 5 or 6, wherein one or more computers generate the requirements for the new image based on at least one of the target person's language comprehension ability, age, health condition, and news of interest.
8. The image generation method according to any one of claims 1 to 7, wherein one or more computers acquire an image specified by user input as an existing image.
9. The image generation method according to any one of claims 1 to 7, wherein one or more computers receive user input specifying a target group, acquire information on the specified target group, identify products popular with the specified target group based on the acquired information on the target group, and acquire images of the identified products as existing images.
10. The image generation method according to claim 9, wherein one or more computers obtain information of the specified target layer from publicly available data.
11. The image generation method according to claim 9 or 10, wherein the information of the designated target group is at least one of the following: social media posts by the target group, product reviews published online by the target group, purchase history data of the target group, and sales data of the target group.
12. The image generation method according to any one of claims 1 to 7, wherein one or more computers receive user input specifying a target individual, acquire information related to the individual specified as the target, identify products related to the individual specified as the target based on the acquired information, and acquire images of the identified products as existing images.
13. The image generation method according to claim 12, wherein one or more computers obtain information related to the individual designated as the target from publicly available data.
14. The image generation method according to claim 12 or 13, wherein the information relating to the individual designated as a target includes at least one of the following: the individual's past purchase history, the individual's past browsing history, information indicating the individual's attributes, the individual's health information, and information indicating the individual's preferences.
15. The image generation method according to any one of claims 12 to 14, wherein one or more computers identify at least one of the following as products related to the individual designated as a target: products that the individual has purchased in the past, products that the individual has purchased more than a predetermined number of times in the past, products that the individual has viewed in the past, products that the individual has viewed more than a predetermined number of times in the past, and products that suit the individual's attributes or preferences.
16. The image generation method according to any one of claims 1 to 15, wherein one or more computers acquire at least one existing image including a product, generate at least one descriptive text describing the existing design of the product contained in each of the at least one existing image, extract commonalities among the at least one descriptive text, generate requirements for a new design of the product based on the commonalities, and generate a new image including the product with the new design by using the requirements as input text to the image generation model.
17. The image generation method according to any one of claims 1 to 16, wherein the machine learning model is a combination of an image recognition model and a natural language processing model.
18. The image generation method according to any one of claims 1 to 17, wherein one or more computers acquire as the existing image an image containing a predetermined object, an image containing a predetermined character, an image of a painting, or an image of a poster.
19. An image generation apparatus comprising: acquisition means for acquiring at least one existing image; text generation means for generating text relating to the features of the existing image based on the existing image and a machine learning model; and image generation means for generating a new image based on the text and an image generation model.
20. A recording medium that stores a program causing a computer to function as: an acquisition means for acquiring at least one existing image; a text generation means for generating text relating to the features of the existing image based on the existing image and a machine learning model; and an image generation means for generating a new image based on the text and an image generation model.