system

The system addresses the challenge of creating high-quality, customizable images by using text-based instructions, natural language processing, and intuitive editing tools, allowing users to efficiently create and edit images without specialized knowledge.

JP2026035287APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing image generation tools require specialized knowledge and are difficult to customize, making it challenging for users to create and edit high-quality, customizable images efficiently.

Method used

A system that receives text instructions, utilizes natural language processing to analyze and generate high-resolution images, allows editing through an intuitive graphical user interface, and applies user-specified styles and partial image regeneration.

Benefits of technology

Enables users without specialized knowledge to easily generate and edit high-quality, customizable images, providing intuitive customization and efficient image creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035287000001_ABST
    Figure 2026035287000001_ABST
Patent Text Reader

Abstract

Provide a system. means for receiving a text instruction entered by a user; natural language processing means for analyzing input text instructions; a generative artificial intelligence means for generating high resolution customizable images based on the analyzed text instructions; means for transmitting the generated image to a user device; editing means for adjusting image details on the user's device; means for applying a user-specified style to the image; A means of regenerating a specified portion of an image. A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Creating high-quality images has traditionally been time-consuming, requiring specialized skills and expensive software. Existing image generation tools are difficult to customize to meet specific user needs, requiring significant effort. Therefore, there is a need for a system that allows users to easily and intuitively generate and edit high-resolution, customizable images without specialized knowledge. This invention aims to address the challenges faced by advertising agencies, graphic designers, social media managers, publishing companies, and individual content creators and artists. [Means for solving the problem]

[0005] To solve the above problems, the present invention provides the following means: a means for receiving text instructions input by a user, a natural language processing means for analyzing the input text instructions, a generating artificial intelligence means for generating a high-resolution customizable image based on the analyzed text instructions, and a means for transmitting the generated image to the user's device. The system further includes an editing means for adjusting fine details of the image on the user's device, a means for applying a style specified by the user to the image, and a means for regenerating the specified image portion. This configuration allows users without specialized knowledge to easily generate and edit high-quality images.

[0006] A "user" is a person who operates the system by entering text instructions to create and edit images.

[0007] "Text instructions" refers to inputting the content of the image the user wants to create or edit in natural language.

[0008] The "means for receiving" refers to a function for incorporating text instructions input by a user through a terminal into the system.

[0009] "Natural language processing means" refers to a function that analyzes input text instructions and creates prompts suitable for image generation.

[0010] "Generative artificial intelligence means" refers to artificial intelligence that generates high resolution images based on prompts created by the natural language processing means.

[0011] "Generated image" refers to image data created by text instructions and generating artificial intelligence means.

[0012] The "means for transmitting" refers to a function for transmitting the generated image to the user's device.

[0013] "Editing tools" are the tools and functions that a user uses to adjust the finer details of the generated image.

[0014] "Means for applying styles" refers to the ability to apply a specific art style (e.g., oil painting style, illustration style) selected by the user to an image.

[0015] The "regenerating means" refers to a function that regenerates a specified portion of an image based on a user instruction. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] DETAILED DESCRIPTION OF THE INVENTION The present invention is an image creation and editing system that allows users to input text instructions to create high-resolution, customizable images, and then fine-tune and style them through an intuitive graphical user interface (GUI).

[0038] System Overview

[0039] This system consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, image generation, image editing, and transmission. The terminal provides a GUI and an interface for users to generate and edit images. Users input text and check and edit images through the terminal.

[0040] Program processing flow

[0041] 1. Initial text input and submission

[0042] The user enters specific text instructions into the device's GUI using an input field, for example, "A silhouette of a wolf standing on a mountain top against a sunset sky."

[0043] The terminal sends this text instruction to the server.

[0044] 2. Text analysis and prompt generation

[0045] The server analyzes the received text, specifically using a natural language processing (NLP) engine to extract important keywords ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0046] Based on the extracted keywords, a generative artificial intelligence (AI) model generates suitable prompts.

[0047] 3. Image Generation

[0048] The server sends the generated prompt to a generative artificial intelligence model, such as a deep learning image generation model (e.g., DALL-E).

[0049] An image generation model generates a high-resolution image based on the prompt and returns it to the server.

[0050] 4. Image transmission and initial display

[0051] The server transmits the generated image to the terminal.

[0052] The terminal displays the received image to the user.

[0053] Specific editing and style application

[0054] 1. Editing images

[0055] The user uses the device's GUI to adjust the details of the image, for example, to move the wolf to the center, by selecting an image part and dragging and dropping it to adjust its position.

[0056] The terminal sends these editing instructions to the server, which then regenerates the image based on them.

[0057] 2. Applying Styles

[0058] The user selects a particular style, such as "oil painting," from the style change options.

[0059] The terminal transmits the selected style information to the server.

[0060] The server processes the image to apply the selected style and sends it back to the terminal.

[0061] Partial changes and final confirmation

[0062] 1. Partial Image Regeneration

[0063] The user inputs specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper."

[0064] The terminal transmits a partial change instruction to the server.

[0065] The server performs partial image regeneration according to the instructions and transmits the updated image to the terminal.

[0066] 2. Final confirmation and saving

[0067] The user reviews the final image.

[0068] The user can save the final image to their device or to a server as desired.

[0069] This allows users to easily create and edit high-quality images without specialized knowledge. The system offers high customizability to meet diverse user needs, enabling efficient image creation. The intuitive GUI and advanced AI technology also enhance the user experience.

[0070] The processing flow will be explained below.

[0071] Step 1:

[0072] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided by the terminal GUI.

[0073] Step 2:

[0074] The terminal receives the user's text instructions and transmits the data to the server.

[0075] Step 3:

[0076] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0077] Step 4:

[0078] The server generates an optimal prompt for the artificial intelligence generating means based on the extracted keywords.

[0079] Step 5:

[0080] The server sends the generated prompt to the generative artificial intelligence model and makes a request to generate an image based on the prompt.

[0081] Step 6:

[0082] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[0083] Step 7:

[0084] The server receives the image data and transmits it to the user's terminal.

[0085] Step 8:

[0086] The device displays the received image to the user, who can then use the device's editing tools to review the image and make minor adjustments as needed.

[0087] Step 9:

[0088] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[0089] Step 10:

[0090] The terminal transmits the user's editing instructions to the server.

[0091] Step 11:

[0092] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[0093] Step 12:

[0094] The server transmits the regenerated image data to the user's terminal.

[0095] Step 13:

[0096] The user selects a style such as "oil painting" from the device's toolbar.

[0097] Step 14:

[0098] The terminal transmits the user's style selection instruction to the server.

[0099] Step 15:

[0100] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[0101] Step 16:

[0102] The server transmits the image data after applying the style to the user's terminal.

[0103] Step 17:

[0104] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[0105] Step 18:

[0106] The terminal transmits a partial change instruction to the server.

[0107] Step 19:

[0108] The server regenerates the image of the relevant part based on the received partial change instruction.

[0109] Step 20:

[0110] The server transmits the regenerated partially modified image to the user's terminal.

[0111] Step 21:

[0112] The user checks the final image and saves it to the device or to the server as needed.

[0113] Example 1

[0114] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0115] Conventional image generation and editing systems have the drawback of making it difficult for users without specialized knowledge to generate the desired high-quality, customizable images. It is also often difficult for users to intuitively edit the fine details of an image or apply a specific style. Furthermore, partial image regeneration is not easily possible, making it difficult to flexibly edit images according to user needs.

[0116] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0117] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, means for generating an appropriate prompt sentence based on the analyzed text instructions, artificial intelligence generating means for generating a high-resolution, customizable image using the generated prompt sentence, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating the specified image portion. This allows users to intuitively generate and edit high-quality images without specialized knowledge. Furthermore, it is easy to apply specific styles and regenerate partial images, enabling flexible image editing according to user requests.

[0118] The "means for receiving text instructions entered by a user" is a function for transmitting specific text instructions entered by a user into an input field of a terminal to a server.

[0119] The "natural language processing means for analyzing input text instructions" is a natural language processing engine for analyzing the text entered by the user and extracting important keywords and phrases.

[0120] The "means for generating appropriate prompt sentences" is a function that generates prompt sentences in a format that is easy for the image generation model to understand, based on extracted keywords and phrases.

[0121] The "generative artificial intelligence means for generating high-resolution customizable images" is an artificial intelligence model for generating high-resolution and customizable images based on generated prompt text.

[0122] The "means for transmitting the generated image to the user's device" is a function for transmitting the generated image data from the server to the user's terminal.

[0123] The "editing means for adjusting fine details of an image" is an interface that provides a function for a user to intuitively edit fine details of a generated image.

[0124] The "means for applying a style designated by the user to an image" is a function for applying a style selected by the user (for example, "oil painting style") to a generated image.

[0125] The "means for regenerating a specified portion of an image" is a function for partially regenerating an image when a user instructs a change to a specific portion.

[0126] The image generation and editing system of this invention consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, prompt generation, and image generation and editing. The terminal provides an intuitive graphical user interface (GUI) for users to generate and edit images. Users input text instructions through the terminal and view and edit the generated images.

[0127] Text analysis and prompt generation

[0128] When the server receives the text instructions entered by the user, it analyzes the text using a natural language processing (NLP) engine. Examples of NLP engines that can be used include spaCy and NLTK. This engine extracts important keywords and phrases (e.g., "sunset," "sky," "mountain," "peak," and "silhouette of a wolf") from the text. Based on the extracted keywords, the server generates a prompt suitable for the generative AI model. This prompt is in a format that is easy for the generative AI model to understand, and can be a specific instruction such as "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0129] Image generation

[0130] The server sends the generated prompt text to a generative AI model. The generative AI model uses an image generation model such as DALL-E, which uses deep learning. This model generates a high-resolution, customizable image based on the prompt text. The generated image is returned to the server, which then sends it to the user's device. The image data sent is encoded, for example, in Base64 format.

[0131] View and edit images

[0132] The user checks the generated image on the device and edits it as necessary. The device's GUI provides an editing tool that allows the user to intuitively adjust minute details of the image. For example, the user can select an image part and drag and drop it to move the wolf to the center. The device sends the editing instructions to the server, and the server regenerates the image and sends the updated image back to the device.

[0133] Style application and partial regeneration

[0134] A user can select a specific style (e.g., "oil painting") from the GUI of the device and apply it to an image. When the device sends the style information to the server, the server applies the style to the image using a style transfer model. The regenerated image with the style applied is sent to the device. Furthermore, the user can input specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper." The server regenerates parts of the image based on the instructions and sends the updated image to the device.

[0135] An example of a prompt is "The silhouette of a wolf standing on top of a mountain against a sunset sky." After parsing this, the prompt becomes "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0136] This allows users to intuitively create and edit high-quality images without specialized knowledge. The system combines advanced AI technology with an intuitive GUI to create easily customizable images.

[0137] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0138] Step 1:

[0139] Enter and send text

[0140] The user enters specific instructions as text into an input field on the device. For example, they might enter "A silhouette of a wolf standing on a mountain top against a sunset sky." This input text is sent to the server. The device uses an HTTP request to send the text data to the server in JSON format.

[0141] Input: User's text instructions

[0142] Output: Sends text instructions to the server

[0143] Step 2:

[0144] Text analysis and keyword extraction

[0145] The server parses the received text instructions. A natural language processing (NLP) engine (e.g., spaCy or NLTK) tokenizes the text and extracts important keywords and phrases. Extracted keywords include "sunset," "sky," "mountain," "peak," and "wolf silhouette."

[0146] Input: Received text instructions

[0147] Output: Extracted keywords and phrases

[0148] Step 3:

[0149] Generate prompt statement

[0150] The server generates a prompt sentence appropriate for the generative AI model based on the extracted keywords and phrases. For example, the generated prompt sentence might be "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0151] Input: Extracted keywords or phrases

[0152] Output: Generated prompt statement

[0153] Step 4:

[0154] Sending an image generation request

[0155] The server sends the generated prompt text to a generative AI model, which then uses deep learning to generate an image based on the text (e.g., DALL-E).

[0156] Input: Generated prompt text

[0157] Output: The generated image

[0158] Step 5:

[0159] Sending the generated image

[0160] The server receives the generated image and sends the image data to the device. The image data is encoded in Base64 format. The device receives the image data in JSON format using an HTTP response.

[0161] Input: Generated image

[0162] Output: Sending image data to the device

[0163] Step 6:

[0164] Displaying images

[0165] The terminal decodes the received image data and displays it to the user. Specifically, the image after Base64 decoding is displayed as HTML. Display it with tags etc.

[0166] Input: Received image data

[0167] Output: Displaying the image to the user

[0168] Step 7:

[0169] Editing images

[0170] The user uses the GUI on the device to adjust small details of the image. For example, to move the wolf to the center, the user selects an image part and adjusts its position by dragging and dropping. The device then sends the edit instructions to the server, including the changes and new coordinates.

[0171] Input: User editing instructions

[0172] Output: Sending edit instructions to the server

[0173] Step 8:

[0174] Regenerate the image

[0175] The server regenerates the image based on the received editing instructions, and the regenerated image data is sent from the server to the terminal and displayed again.

[0176] Input: User editing instructions

[0177] Output: Regenerated image

[0178] Step 9:

[0179] Applying Styles

[0180] The user selects a specific style (e.g., "oil painting") from the device's GUI. The device sends a style application instruction to the server, which then applies the style to the image using the style transfer model. The image after the style application is sent from the server to the device.

[0181] Input: User style selection

[0182] Output: Image with style applied

[0183] Step 10:

[0184] Regenerating Partial Changes

[0185] The user inputs instructions to change a specific part (e.g., "Make the edges of the mountain peaks look sharper"). The device sends the instructions to the server, which then regenerates the image based on the instructions. The updated image is sent to the device and presented to the user.

[0186] Input: User's partial change instructions

[0187] Output: Partially regenerated image

[0188] This allows the system to combine advanced AI technology with an intuitive GUI, providing an environment in which users can easily generate and edit high-quality images.

[0189] (Application example 1)

[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0191] Conventional advertising image generation technologies are difficult for users without specialized knowledge to operate and have limited customization options, making it difficult to efficiently create high-quality advertising images. Furthermore, the process of fine-tuning and applying styles is not intuitive, making it difficult to improve the user experience.

[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0193] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generating means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating the specified image parts, means for generating and editing an advertising image based on the text instructions, and means for editing fine details of the generated image by drag and drop, thereby enabling users to intuitively and efficiently generate and edit high-quality advertising images without specialized knowledge.

[0194] A "user" is an individual or organization that has the ability to create and edit advertising images using the image generation system.

[0195] "Text instructions" refer to specific words or keywords that a user inputs for generating or editing an advertisement image.

[0196] "Natural language processing means" is a technology for analyzing input text instructions and extracting keywords and generating prompts based on those instructions.

[0197] "Generative artificial intelligence means" refers to AI technology for generating high-resolution, customizable images based on analyzed text instructions.

[0198] "High-resolution customizable images" refers to high-quality images that can be extensively edited and styled according to user instructions and preferences.

[0199] A "server" is a computing device that receives text instructions, parses them, generates images, and transmits them to a user's device.

[0200] "User's device" refers to a device (e.g., smartphone, tablet, or PC) that a user uses to access and operate the image generation system.

[0201] "Editing tools" refers to functionality that allows users to adjust the fine details of an image on their device.

[0202] "Means for applying a style" refers to a technique for reflecting a user-specified style (e.g., classic, modern, etc.) in the generated image.

[0203] The "means for regenerating a portion of an image" is a technique for regenerating an image based on additional instructions from the user for a specific portion.

[0204] "Drag-and-drop editing" refers to a feature that allows users to intuitively move and modify specific parts of an image by dragging and dropping.

[0205] MODE FOR CARRYING OUT THE INVENTION

[0206] An embodiment of the present invention will be described below: This system uses a smartphone as a main terminal as an application that can easily generate and edit advertising images.

[0207] Overall system configuration

[0208] The system consists of the following components:

[0209] User device: Smartphone (iOS or ANDROID)

[0210] Server: Responsible for receiving text instructions, analyzing them, generating images, and sending them

[0211] Generative AI methods: AI techniques that generate high-resolution, customizable images

[0212] Program processing

[0213] Initial text entry and submission

[0214] A user uses their smartphone to enter specific advertising instructions as text, such as "Summer sale, 50% off, blue background, sun icon," and the smartphone app sends the text instructions to a cloud server.

[0215] Natural Language Processing and Prompt Generation

[0216] The server analyzes the received text instructions using natural language processing tools (e.g., SpaCy, NLTK). Specifically, it extracts important keywords from the text (e.g., "summer sale," "50% off," "blue background," "sun icon") and generates a prompt suitable for a generative AI model such as GPT-4 (registered trademark).

[0217] Image generation

[0218] Using the generated prompts, the server sends instructions to a generating artificial intelligence means (e.g., DALL-E), which generates a high-resolution advertising image based on the prompts and returns it to the server, which then sends the generated image to the user's smartphone.

[0219] Editing and styling

[0220] Editing images

[0221] The smartphone app provides users with an intuitive graphical user interface (GUI) that allows them to fine-tune the image. Users can select specific parts of the image and change their position and size by dragging and dropping. Editing instructions are then sent back to the server, which then regenerates the image based on the user's instructions.

[0222] Applying Styles

[0223] Users select a style, such as "Classic" or "Modern," from the app's style change options. This information is also sent to the server, which then applies the style to the image using generative artificial intelligence methods. The resulting regenerated image is sent to the smartphone.

[0224] Partial image regeneration and final confirmation

[0225] Partial changes

[0226] When a user requests a specific change, for example, "increase the icon size," the request is also sent to the cloud server, which then regenerates the image of the corresponding part.

[0227] Final confirmation and saving

[0228] Users can view the final ad image on their smartphone and save it or share it on social media or cloud storage.

[0229] Example prompt

[0230] "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background."

[0231] This allows even non-expert users to efficiently and intuitively create and edit high-resolution, customizable advertising images. The combination of generative artificial intelligence and natural language processing technology enables high levels of customization and an improved user experience.

[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0233] Step 1:

[0234] Entering and sending text instructions

[0235] The user opens the app on their smartphone and inputs the text instructions needed to generate the ad image. For example, they might input "Summer sale, 50% off, blue background, sun icon." The device then sends this input text to the cloud server. The input in this step is the user's text instructions, and the output is the data sent to the cloud server.

[0236] Step 2:

[0237] Text analysis and keyword extraction

[0238] The cloud server analyzes the received text instructions using a natural language processing engine (e.g., SpaCy, NLTK). Specifically, it tokenizes the text and extracts important keywords ("summer sale," "50% off," "blue background," "sun icon"). The input of this step is the text instruction sent by the user, and the output is the analyzed keywords.

[0239] Step 3:

[0240] Prompt Generation

[0241] The server generates a prompt appropriate for the generative AI model (e.g., GPT-4) based on the extracted keywords. This prompt will be a sentence such as "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background." The input of this step is the parsed keywords, and the output is the generated prompt sentence.

[0242] Step 4:

[0243] Image generation

[0244] The server sends the generated prompt text to an image generation model (e.g., DALL-E). DALL-E generates a high-resolution advertising image based on the prompt and returns it to the server. The input of this step is the prompt text, and the output is a high-resolution advertising image.

[0245] Step 5:

[0246] Image transmission and initial display

[0247] The server sends the generated high-resolution advertising image to the smartphone. The device displays the received image to the user. The input of this step is the generated advertising image, and the output is the display data for the smartphone.

[0248] Step 6:

[0249] Editing images

[0250] The user uses the smartphone's GUI to adjust the details of the image, for example, by dragging and dropping the sun icon to the center. The device then sends these editing instructions to the cloud server. The input of this step is the user's editing instructions, and the output is the data sent to the cloud server.

[0251] Step 7:

[0252] Regenerate Edits

[0253] The server regenerates the image based on the received editing instructions. The regenerated image is sent to the smartphone. The input of this step is the user's editing instructions, and the output is the regenerated image.

[0254] Step 8:

[0255] Applying Styles

[0256] The user selects a style, such as "Classic" or "Modern." The device sends the style change information to the server. The server then uses a generative artificial intelligence method to apply the specified style to the image. The input to this step is the user-specified style, and the output is the image with the style applied.

[0257] Step 9:

[0258] Partial Image Regeneration

[0259] The user instructs the server to change a specific part (e.g., "increase the icon size"). The device sends the instruction to change the part to the server. The server regenerates the specific part according to the instruction and resends the image to the smartphone. The input of this step is the instruction to change the part, and the output is the regenerated image.

[0260] Step 10:

[0261] Final confirmation and saving

[0262] The user can then check the final ad image and save it on their smartphone or share it on social media or cloud storage if desired. The input of this step is the final ad image, and the output is the saved or shared image.

[0263] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0264] The interactive AI image studio of the present invention is a system that receives text instructions entered by a user, analyzes them using natural language processing means, and generates high-resolution, customizable images using generative AI means. By combining this system with an emotion engine, it can recognize the user's emotions and generate and edit images based on those emotions.

[0265] System configuration

[0266] This system is mainly composed of three elements: a server, a terminal, and a user. In particular, the addition of an emotion engine enables more personalized image generation according to the user's emotional state.

[0267] Program processing flow

[0268] 1. Initial text input and submission

[0269] The user inputs specific instructions into the device's GUI as text, such as "The silhouette of a wolf standing on top of a mountain with a sunset sky in the background."

[0270] The terminal sends this text instruction to the server.

[0271] 2. Text analysis and prompt generation

[0272] The server analyzes the received text instructions using natural language processing means, specifically by extracting important keywords from the text.

[0273] 3. Emotion analysis

[0274] The emotion engine analyzes the user's input text, on-device interactions, voice, and facial expressions to recognize the user's emotional state.

[0275] For example, if the user is expressing a "happy" emotion, a prompt containing colors and motifs that match that emotion is generated.

[0276] 4. Image Generation

[0277] The server sends prompts to the generating artificial intelligence means to generate high resolution images based on the parsed text instructions and emotion data.

[0278] An image is generated by a generating artificial intelligence means (eg, DALL-E) and sent back to the server.

[0279] 5. Sending and displaying images

[0280] The server transmits the generated image to the terminal.

[0281] The terminal displays the received image to the user.

[0282] Specific editing and style application

[0283] 1. Editing images

[0284] The user can check the image on the device's GUI and make fine adjustments as needed, for example, by selecting an image part and dragging and dropping it to move the wolf to the center.

[0285] 2. Applying Styles

[0286] The user selects a particular style, such as "oil painting," from the style change options.

[0287] The terminal transmits the selected style information to the server, and the server performs processing to apply the selected style to the image.

[0288] Partial changes and emotional readjustments

[0289] 1. Partial Image Regeneration

[0290] The user inputs instructions for partial changes such as "make the edges of the mountain peaks look sharper."

[0291] The terminal transmits a partial change instruction to the server, and the server performs partial image regeneration based on the instruction.

[0292] 2. Emotional Recalibration

[0293] The emotion engine refers to the user's input and editing history and suggests additional adjustments and optimizations based on the user's emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[0294] 3. Final confirmation and saving

[0295] The user checks the final image and saves it to the device or to the server as needed.

[0296] This system allows users to easily create and edit high-quality images that are personalized according to their emotional state. The introduction of an emotion engine further improves the user experience, making the image creation and editing process more intuitive and effective.

[0297] The processing flow will be explained below.

[0298] Specific processing steps

[0299] Step 1:

[0300] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided in the GUI of the terminal.

[0301] Step 2:

[0302] The terminal receives the user's text instructions and transmits the data to the server.

[0303] Step 3:

[0304] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0305] Step 4:

[0306] The emotion engine analyzes the user's input text and interactions to recognize the user's emotions. For example, it recognizes the emotion "fun" from the input text and user actions.

[0307] Step 5:

[0308] Based on the keywords and emotion recognition results, the server generates a prompt with details about the image to be generated. For example, if the "happy" emotion is recognized, a prompt using bright colors is generated.

[0309] Step 6:

[0310] The server sends the generated prompt to the generating artificial intelligence means, which makes a request to generate an image based on the prompt.

[0311] Step 7:

[0312] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[0313] Step 8:

[0314] The server receives the image data and transmits it to the user's terminal.

[0315] Step 9:

[0316] The device displays the received image to the user, who can review the image and make minor adjustments as needed.

[0317] Step 10:

[0318] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[0319] Step 11:

[0320] The terminal transmits the user's editing instructions to the server.

[0321] Step 12:

[0322] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[0323] Step 13:

[0324] The server transmits the regenerated image data to the user's terminal.

[0325] Step 14:

[0326] The user selects a style such as "oil painting" from the device's toolbar.

[0327] Step 15:

[0328] The terminal transmits the user's style selection instruction to the server.

[0329] Step 16:

[0330] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[0331] Step 17:

[0332] The server transmits the image data after applying the style to the user's terminal.

[0333] Step 18:

[0334] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[0335] Step 19:

[0336] The terminal transmits a partial change instruction to the server.

[0337] Step 20:

[0338] Based on the partial change instruction received by the server, the image of the relevant part is regenerated.

[0339] Step 21:

[0340] The server transmits the regenerated partially modified image to the user's terminal.

[0341] Step 22:

[0342] The emotion engine re-analyzes the user's emotional state based on their input and editing history, and suggests additional adjustments and optimizations based on their emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[0343] Step 23:

[0344] The user checks the final image and saves it to the device or to the server as needed.

[0345] Example 2

[0346] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0347] Conventional image generation systems often cannot predict the results of a generated image, even when a user inputs clear instructions, making it difficult to obtain an image that reflects the user's intentions. Furthermore, they lack the functionality to generate images that reflect the user's emotional state, making it difficult to generate personalized images. Furthermore, the functionality for partially modifying an image or applying styles is limited, which can make it time-consuming and labor-intensive to generate an image that satisfies the user.

[0348] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a text instruction input by a user, natural language processing means for analyzing the input text instruction, artificial intelligence means for generating a high-resolution customizable image based on the analyzed text instruction and the user's emotional data, an emotion engine for analyzing the user's emotional state, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating a specified portion of the image. This makes it possible to quickly generate high-quality, personalized images that correspond to the user's emotional state and to edit them intuitively and efficiently.

[0349] The "means for receiving user-entered text instructions" is a module for electronically receiving user-entered instructions in text form.

[0350] The "natural language processing means for analyzing input text instructions" is a module that includes techniques and algorithms for analyzing input text instructions and understanding their content.

[0351] A "generative artificial intelligence means for generating high-resolution customizable images" is a module that utilizes artificial intelligence technology to generate high-quality, detailed, customizable images based on specified conditions.

[0352] The "emotion engine that analyzes the user's emotional state" is a module that includes technologies and algorithms for analyzing emotions from the user's input and actions and determining their state.

[0353] The "means for transmitting the generated image to the user's device" is a module for transferring the generated image data to the user's device via a network.

[0354] The "editing tool for adjusting fine details of an image" is a module that provides an interface for users to modify and adjust the details of the generated image.

[0355] The "means for applying a user-specified style to an image" is a module for applying a particular style selected by the user (e.g., oil painting style, watercolor style, etc.) to the generated image.

[0356] The "means for regenerating a specified portion of an image" is a module for regenerating and modifying a specific portion of an image based on a user's instructions.

[0357] MODE FOR CARRYING OUT THE INVENTION

[0358] The interactive AI image studio of the present invention is a system that receives and analyzes user-entered text instructions to generate high-resolution, customizable images. The system includes an emotion engine that recognizes the user's emotional state and generates and edits images accordingly. An embodiment of the system is described in detail below.

[0359] System Components

[0360] This system mainly consists of three elements: a server, a terminal, and a user.

[0361] 1. Server:

[0362] The server has the following functions:

[0363] Means for receiving user-entered text instructions

[0364] Analyzing the text using natural language processing tools (e.g., SpaCy)

[0365] Analysis of the user's emotional state using an emotion engine (e.g., OpenAI's GPT-3)

[0366] Generation of images by artificial intelligence means (e.g., DALL-E)

[0367] A means for transmitting the generated image data to a terminal

[0368] 2. Terminal:

[0369] The device has the following features:

[0370] GUI for user input

[0371] Displaying received images

[0372] Provides tools for editing the finer details of an image

[0373] A means for sending user-specified styles to the server and applying them

[0374] Means for sending partial image regeneration instructions to a server

[0375] 3. User:

[0376] The user performs the following operations:

[0377] Enter text instructions (e.g., "A silhouette of a wolf standing on a mountain peak against a sunset sky")

[0378] View and edit the generated image

[0379] Selecting and applying a specific style

[0380] Partial changes, final confirmation and saving

[0381] System Operation

[0382] Using this system, users can easily perform the following process:

[0383] Initial input and submission:

[0384] The user inputs text into the terminal and sends it to the server. Example input: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[0385] Natural Language Processing and Sentiment Analysis:

[0386] The server analyzes the received text instructions using natural language processing to extract important keywords. The emotion engine then analyzes the user's emotional state from the text. For example, if the emotion "fun" is detected, the server generates a prompt that matches that emotion.

[0387] Image generation:

[0388] The server generates a high-resolution image using a generative AI model (e.g., DALL-E) based on the prompt and emotion data. Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors."

[0389] View and edit images:

[0390] The server sends the generated image to the terminal, where the user can view it. The user can then edit the image in detail on the terminal, adjusting the position and color tone as necessary.

[0391] Applying styles:

[0392] The user selects a style change option (e.g., "oil painting"), and the device sends that information to the server, which applies the selected style to the image and sends it back to the device.

[0393] Modify and regenerate:

[0394] The user instructs the terminal to make specific partial changes, and the terminal sends the instruction to the server, which then regenerates the image based on the instruction and sends the updated image to the terminal.

[0395] Final review and save:

[0396] The user reviews the final image and saves it to local storage or cloud storage if desired.

[0397] Specific examples

[0398] For example, if the user enters the prompt "A bird sitting in a tree under a blue sky," the following prompt sentence is generated:

[0399] Prompt: "In the clear blue sky, a bird perched on a tree, with warm and bright colors indicating happiness."

[0400] Based on this prompt, the generative AI model generates an image that reflects the user's emotional state and sends it to the user's device via the server. The user can then review the image, edit and change the style as needed, and finally save it.

[0401] The system allows users to quickly and intuitively generate and edit personalized, high-quality images that reflect their emotional state. The introduction of an emotion engine further improves the user experience and makes the image generation and editing process more effective.

[0402] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0403] Program processing flow

[0404] Step 1: User enters and submits text

[0405] Specific actions

[0406] The user enters text instructions into the terminal's GUI.

[0407] Example: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[0408] The user clicks the submit button.

[0409] The terminal sends the entered text instructions to the server.

[0410] Input: The text instruction entered by the user

[0411] Output: HTTP request sent to the server

[0412] Step 2: Parsing text on the server

[0413] Specific actions

[0414] The server passes the received text instructions to a natural language processing (NLP) module.

[0415] The NLP module analyzes the text and extracts important keywords.

[0416] Examples: "sunset," "sky," "mountain," "peak," "wolf," "silhouette"

[0417] Input: Received text instructions

[0418] Output: Extracted keywords

[0419] Step 3: Sentiment Analysis

[0420] Specific actions

[0421] The server sends the text instructions to the emotion engine.

[0422] An emotion engine analyzes the user's emotional state from the text.

[0423] Example: Detecting the emotion "happy"

[0424] Input: Text instructions

[0425] Output: Emotion data

[0426] Step 4: Generate prompts

[0427] Specific actions

[0428] The server generates a prompt sentence based on the analyzed keywords and emotion data.

[0429] Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors"

[0430] Input: Keywords, emotion data

[0431] Output: prompt statement

[0432] Step 5: Image generation

[0433] Specific actions

[0434] The server sends the generated prompt sentence to the generating artificial intelligence means (generating AI model).

[0435] Example: DALL-E

[0436] The generative AI model generates high-resolution images and sends them back to the server.

[0437] Input: prompt statement

[0438] Output: Generated image data

[0439] Step 6: Receiving and displaying images

[0440] Specific actions

[0441] The server transmits the generated image data to the terminal.

[0442] The terminal receives the image and displays it to the user.

[0443] Input: Generated image data

[0444] Output: Image displayed on the terminal

[0445] Step 7: Edit the image

[0446] Specific actions

[0447] The user edits the image to adjust the finer details.

[0448] Example: Move the wolf to the center

[0449] The device provides editing tools, allowing users to select parts of the image and adjust their position by dragging and dropping.

[0450] Input: Generated image

[0451] Output: Edited image data

[0452] Step 8: Applying Styles

[0453] Specific actions

[0454] The user selects a specific style from the style change options.

[0455] Example: "Oil painting style"

[0456] The terminal transmits style information to the server.

[0457] The server applies the style to the image and sends it back to the device.

[0458] Input: Style information

[0459] Output: Image data with style applied

[0460] Step 9: Partial Image Regeneration

[0461] Specific actions

[0462] The user inputs instructions for partial modification.

[0463] Example: "Make the edges of the mountain peaks look sharper."

[0464] The terminal transmits a partial change instruction to the server.

[0465] The server regenerates a portion of the image based on the instruction and transmits the updated image to the terminal.

[0466] Input: Partial change instruction

[0467] Output: Partially regenerated image

[0468] Step 10: Final review and save

[0469] Specific actions

[0470] The user reviews the final image.

[0471] The user clicks the Save button.

[0472] The device stores the image data in local storage or cloud storage.

[0473] Input: Last seen image

[0474] Output: Saved image data

[0475] This process flow allows users to quickly generate high-quality, personalized images according to their emotional state, and then intuitively edit and save them.

[0476] (Application example 2)

[0477] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0478] In conventional advertising production, it was difficult to customize according to the user's emotional state, making it difficult to generate intuitive and effective advertising images. It was also not easy to reflect the style and editing instructions specified in the early stages of ad production and apply them to the image immediately. As a result, the effectiveness of the advertisement was not maximized, and it was not possible to achieve advertising expressions that were in tune with the emotions of the target user.

[0479] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generation means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating a specified part of the image, emotion analysis means for recognizing the user's emotional state and generating and editing an image based on the emotion, and means for enabling application in the field of advertising production. This enables the generation and immediate editing of personalized advertising images according to the user's emotions, thereby realizing the provision of optimal advertisements to target users.

[0480] "User" means a person who uses the Interactive AI Image Studio to generate and edit images.

[0481] A "text instruction" is a sentence that describes specific details about the image that the user wants to generate.

[0482] "Natural language processing means" refers to means for analyzing input text instructions and extracting important keywords and context.

[0483] "Generative artificial intelligence means" refers to artificial intelligence techniques for generating high resolution customizable images based on analyzed instructions.

[0484] The "editing means" is a function that allows the user to adjust the fine details of the generated image.

[0485] A "style" is a particular image presentation or visual effect specified by the user.

[0486] The "regeneration means" is a function for regenerating a specified part of an image.

[0487] The "emotion analysis means" is a function for recognizing the user's emotional state and generating and editing images based on that emotion.

[0488] The "advertising production field" is the field of creating images and content to effectively promote products and services to target users.

[0489] The present invention is a system that utilizes an interactive AI image studio to generate high-resolution, customizable advertising images based on user-entered text instructions. The main components of the system include a server, a terminal, and a user.

[0490] Server Configuration

[0491] The server includes the following means:

[0492] Text receiving means: A means for receiving a text instruction input by a user from a terminal.

[0493] Natural language processing means: A means for analyzing input text instructions, extracting keywords, and generating optimal prompts for the artificial intelligence generator. Specifically, a natural language processing engine such as SpaCy is used.

[0494] Generative AI means: A means for generating high-resolution, customizable images based on analyzed instructions, specifically using a generative AI model (e.g., DALL-E).

[0495] Image transmission means: A means for transmitting the generated image to the user's terminal.

[0496] Emotion analysis means: A means for recognizing a user's emotional state by analyzing the user's input text content, interactions, voice input, and facial expression data. The IBM Watson (registered trademark) Emotion Analysis API can be used.

[0497] Application in the field of advertising production: A means to make it possible to use the generated images and emotion analysis results in advertising production.

[0498] Device configuration

[0499] The terminal includes:

[0500] GUI editing means: A graphical user interface function that allows users to adjust the fine details of the generated image. This allows users to intuitively adjust the image.

[0501] Style application: A function for applying a user-specified style to an image. Users can select styles such as "oil painting" or "vintage."

[0502] Partial regeneration means: A function for regenerating a specified part of an image.

[0503] User operations

[0504] The user performs the following steps:

[0505] 1. Enter the basic concept of the advertisement, such as "people relaxing at the beach," into the GUI of a smartphone or other device.

[0506] 2. The text instructions are sent to the server, where the server's natural language processing means analyzes the text and extracts keywords.

[0507] 3. The emotion analyzer recognizes the user's emotional state and generates appropriate prompts based on that data.

[0508] 4. The generating artificial intelligence means (such as DALL-E) generates a high-resolution image based on the prompt and sends it to the terminal via the server.

[0509] 5. The user can review the received image on their device, make minor adjustments or apply styles as needed, and even request partial regeneration.

[0510] Examples of concrete examples and prompts

[0511] As a concrete example, consider a situation where a user enters the text "A park scene with a happy family together" for an advertising image, and the emotion is "joy." An example prompt for this would be:

[0512] Scene of happy family spending time in the park, joy

[0513] In this way, users can intuitively create and edit high-quality advertising images in a short amount of time. This system enables advertising expressions that are in tune with the emotions of target users, maximizing their effectiveness.

[0514] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0515] Step 1:

[0516] The user inputs the basic idea of ​​the advertisement into the terminal's GUI (text instructions).

[0517] Input: A textual indication of user input (e.g., "People relaxing at the beach").

[0518] Output: A request to send input text instructions.

[0519] Specific operation: The user enters instructions into the terminal's GUI and presses the "Send" button, which sends a request to the server.

[0520] Step 2:

[0521] The server analyzes the received text instructions using natural language processing means.

[0522] Input: The text instruction submitted by the user.

[0523] Output: Extracted keywords and analysis results.

[0524] What happens: The server's natural language processing engine (e.g., SpaCy) analyzes the text and extracts important keywords and contextual information.

[0525] Step 3:

[0526] An emotion analyzer recognizes the user's emotional state.

[0527] Input: Text input, on-device interaction data, voice and facial expression data.

[0528] Output: Perceived emotional state (e.g., "joy").

[0529] What happens: The server analyzes the data using an emotion analysis engine (e.g., IBM Watson Emotion Analysis API) to determine the user's emotional state.

[0530] Step 4:

[0531] The server generates a prompt based on the analysis result and sends it to the generating artificial intelligence means.

[0532] Input: Keywords, contextual information, emotional state.

[0533] Output: The prompt statement.

[0534] Specific operation: The server integrates the extracted keywords and emotion data to generate a prompt sentence to send to the generative AI model (e.g., DALL-E).

[0535] Step 5:

[0536] A generating artificial intelligence means generates a high resolution image.

[0537] Input: Prompt sentence (e.g., "A park scene of a happy family spending time together, joy").

[0538] Output: High resolution customizable images.

[0539] What it does: A prompt is sent to the generative AI model, which then generates an image based on it.

[0540] Step 6:

[0541] The server transmits the generated image to the terminal.

[0542] Input: The generated image.

[0543] Output: Request to send image data.

[0544] Specific operation: The server receives the generated image and sends it to the user's device.

[0545] Step 7:

[0546] The user checks the image on the terminal and edits it as necessary.

[0547] Input: Received image data.

[0548] Output: Editing instructions and adjusted image data.

[0549] Specific behavior: The user views the image in the device's GUI and uses drag-and-drop and tools to make fine adjustments and apply styles.

[0550] Step 8:

[0551] The user requests a partial regeneration.

[0552] Input: Partial regeneration instructions.

[0553] Output: The regenerated subimage.

[0554] Specific operation: The user selects a specific image part, inputs a regeneration instruction, and sends it to the server. The server generates a prompt again, requests the generative AI model to retrieve the regenerated image part, and sends it to the user.

[0555] Step 9:

[0556] The user reviews and saves the final image.

[0557] Input: The final image data after adjustments and regeneration.

[0558] Output: Saved image file.

[0559] Specific behavior: The user confirms the final image and saves it to the device or server.

[0560] The above are the specific processing steps of the system that realizes the application example.

[0561] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0562] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0563] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0564] [Second embodiment]

[0565] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0566] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0567] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0568] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0569] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0570] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0571] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0572] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0573] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0574] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0575] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0576] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0577] DETAILED DESCRIPTION OF THE INVENTION The present invention is an image creation and editing system that allows users to input text instructions to create high-resolution, customizable images, and then fine-tune and style them through an intuitive graphical user interface (GUI).

[0578] System Overview

[0579] This system consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, image generation, image editing, and transmission. The terminal provides a GUI and an interface for users to generate and edit images. Users input text and check and edit images through the terminal.

[0580] Program processing flow

[0581] 1. Initial text input and submission

[0582] The user enters specific text instructions into the device's GUI using an input field, for example, "A silhouette of a wolf standing on a mountain top against a sunset sky."

[0583] The terminal sends this text instruction to the server.

[0584] 2. Text analysis and prompt generation

[0585] The server analyzes the received text, specifically using a natural language processing (NLP) engine to extract important keywords ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0586] Based on the extracted keywords, a generative artificial intelligence (AI) model generates suitable prompts.

[0587] 3. Image Generation

[0588] The server sends the generated prompt to a generative artificial intelligence model, such as a deep learning image generation model (e.g., DALL-E).

[0589] An image generation model generates a high-resolution image based on the prompt and returns it to the server.

[0590] 4. Image transmission and initial display

[0591] The server transmits the generated image to the terminal.

[0592] The terminal displays the received image to the user.

[0593] Specific editing and style application

[0594] 1. Editing images

[0595] The user uses the device's GUI to adjust the details of the image, for example, to move the wolf to the center, by selecting an image part and dragging and dropping it to adjust its position.

[0596] The terminal sends these editing instructions to the server, which then regenerates the image based on them.

[0597] 2. Applying Styles

[0598] The user selects a particular style, such as "oil painting," from the style change options.

[0599] The terminal transmits the selected style information to the server.

[0600] The server processes the image to apply the selected style and sends it back to the terminal.

[0601] Partial changes and final confirmation

[0602] 1. Partial Image Regeneration

[0603] The user inputs specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper."

[0604] The terminal transmits a partial change instruction to the server.

[0605] The server performs partial image regeneration according to the instructions and transmits the updated image to the terminal.

[0606] 2. Final confirmation and saving

[0607] The user reviews the final image.

[0608] The user can save the final image to their device or to a server as desired.

[0609] This allows users to easily create and edit high-quality images without specialized knowledge. The system offers high customizability to meet diverse user needs, enabling efficient image creation. The intuitive GUI and advanced AI technology also enhance the user experience.

[0610] The processing flow will be explained below.

[0611] Step 1:

[0612] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided by the terminal GUI.

[0613] Step 2:

[0614] The terminal receives the user's text instructions and transmits the data to the server.

[0615] Step 3:

[0616] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0617] Step 4:

[0618] The server generates an optimal prompt for the artificial intelligence generating means based on the extracted keywords.

[0619] Step 5:

[0620] The server sends the generated prompt to the generative artificial intelligence model and makes a request to generate an image based on the prompt.

[0621] Step 6:

[0622] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[0623] Step 7:

[0624] The server receives the image data and transmits it to the user's terminal.

[0625] Step 8:

[0626] The device displays the received image to the user, who can then use the device's editing tools to review the image and make minor adjustments as needed.

[0627] Step 9:

[0628] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[0629] Step 10:

[0630] The terminal transmits the user's editing instructions to the server.

[0631] Step 11:

[0632] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[0633] Step 12:

[0634] The server transmits the regenerated image data to the user's terminal.

[0635] Step 13:

[0636] The user selects a style such as "oil painting" from the device's toolbar.

[0637] Step 14:

[0638] The terminal transmits the user's style selection instruction to the server.

[0639] Step 15:

[0640] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[0641] Step 16:

[0642] The server transmits the image data after applying the style to the user's terminal.

[0643] Step 17:

[0644] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[0645] Step 18:

[0646] The terminal transmits a partial change instruction to the server.

[0647] Step 19:

[0648] The server regenerates the image of the relevant part based on the received partial change instruction.

[0649] Step 20:

[0650] The server transmits the regenerated partially modified image to the user's terminal.

[0651] Step 21:

[0652] The user checks the final image and saves it to the device or to the server as needed.

[0653] Example 1

[0654] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0655] Conventional image generation and editing systems have the drawback of making it difficult for users without specialized knowledge to generate the desired high-quality, customizable images. It is also often difficult for users to intuitively edit the fine details of an image or apply a specific style. Furthermore, partial image regeneration is not easily possible, making it difficult to flexibly edit images according to user needs.

[0656] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0657] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, means for generating an appropriate prompt sentence based on the analyzed text instructions, artificial intelligence generating means for generating a high-resolution, customizable image using the generated prompt sentence, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating the specified image portion. This allows users to intuitively generate and edit high-quality images without specialized knowledge. Furthermore, it is easy to apply specific styles and regenerate partial images, enabling flexible image editing according to user requests.

[0658] The "means for receiving text instructions entered by a user" is a function for transmitting specific text instructions entered by a user into an input field of a terminal to a server.

[0659] The "natural language processing means for analyzing input text instructions" is a natural language processing engine for analyzing the text entered by the user and extracting important keywords and phrases.

[0660] The "means for generating appropriate prompt sentences" is a function that generates prompt sentences in a format that is easy for the image generation model to understand, based on extracted keywords and phrases.

[0661] The "generative artificial intelligence means for generating high-resolution customizable images" is an artificial intelligence model for generating high-resolution and customizable images based on generated prompt text.

[0662] The "means for transmitting the generated image to the user's device" is a function for transmitting the generated image data from the server to the user's terminal.

[0663] The "editing means for adjusting fine details of an image" is an interface that provides a function for a user to intuitively edit fine details of a generated image.

[0664] The "means for applying a style designated by the user to an image" is a function for applying a style selected by the user (for example, "oil painting style") to a generated image.

[0665] The "means for regenerating a specified portion of an image" is a function for partially regenerating an image when a user instructs a change to a specific portion.

[0666] The image generation and editing system of this invention consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, prompt generation, and image generation and editing. The terminal provides an intuitive graphical user interface (GUI) for users to generate and edit images. Users input text instructions through the terminal and view and edit the generated images.

[0667] Text analysis and prompt generation

[0668] When the server receives the text instructions entered by the user, it analyzes the text using a natural language processing (NLP) engine. Examples of NLP engines that can be used include spaCy and NLTK. This engine extracts important keywords and phrases (e.g., "sunset," "sky," "mountain," "peak," and "silhouette of a wolf") from the text. Based on the extracted keywords, the server generates a prompt suitable for the generative AI model. This prompt is in a format that is easy for the generative AI model to understand, and can be a specific instruction such as "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0669] Image generation

[0670] The server sends the generated prompt text to a generative AI model. The generative AI model uses an image generation model such as DALL-E, which uses deep learning. This model generates a high-resolution, customizable image based on the prompt text. The generated image is returned to the server, which then sends it to the user's device. The image data sent is encoded, for example, in Base64 format.

[0671] View and edit images

[0672] The user checks the generated image on the device and edits it as necessary. The device's GUI provides an editing tool that allows the user to intuitively adjust minute details of the image. For example, the user can select an image part and drag and drop it to move the wolf to the center. The device sends the editing instructions to the server, and the server regenerates the image and sends the updated image back to the device.

[0673] Style application and partial regeneration

[0674] A user can select a specific style (e.g., "oil painting") from the GUI of the device and apply it to an image. When the device sends the style information to the server, the server applies the style to the image using a style transfer model. The regenerated image with the style applied is sent to the device. Furthermore, the user can input specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper." The server regenerates parts of the image based on the instructions and sends the updated image to the device.

[0675] An example of a prompt is "The silhouette of a wolf standing on top of a mountain against a sunset sky." After parsing this, the prompt becomes "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0676] This allows users to intuitively create and edit high-quality images without specialized knowledge. The system combines advanced AI technology with an intuitive GUI to create easily customizable images.

[0677] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0678] Step 1:

[0679] Enter and send text

[0680] The user enters specific instructions as text into an input field on the device. For example, they might enter "A silhouette of a wolf standing on a mountain top against a sunset sky." This input text is sent to the server. The device uses an HTTP request to send the text data to the server in JSON format.

[0681] Input: User's text instructions

[0682] Output: Sends text instructions to the server

[0683] Step 2:

[0684] Text analysis and keyword extraction

[0685] The server parses the received text instructions. A natural language processing (NLP) engine (e.g., spaCy or NLTK) tokenizes the text and extracts important keywords and phrases. Extracted keywords include "sunset," "sky," "mountain," "peak," and "wolf silhouette."

[0686] Input: Received text instructions

[0687] Output: Extracted keywords and phrases

[0688] Step 3:

[0689] Generate prompt statement

[0690] The server generates a prompt sentence appropriate for the generative AI model based on the extracted keywords and phrases. For example, the generated prompt sentence might be "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[0691] Input: Extracted keywords or phrases

[0692] Output: Generated prompt statement

[0693] Step 4:

[0694] Sending an image generation request

[0695] The server sends the generated prompt text to a generative AI model, which then uses deep learning to generate an image based on the text (e.g., DALL-E).

[0696] Input: Generated prompt text

[0697] Output: The generated image

[0698] Step 5:

[0699] Sending the generated image

[0700] The server receives the generated image and sends the image data to the device. The image data is encoded in Base64 format. The device receives the image data in JSON format using an HTTP response.

[0701] Input: Generated image

[0702] Output: Sending image data to the device

[0703] Step 6:

[0704] Displaying images

[0705] The terminal decodes the received image data and displays it to the user. Specifically, the image after Base64 decoding is displayed as HTML. Display it with tags etc.

[0706] Input: Received image data

[0707] Output: Displaying the image to the user

[0708] Step 7:

[0709] Editing images

[0710] The user uses the GUI on the device to adjust small details of the image. For example, to move the wolf to the center, the user selects an image part and adjusts its position by dragging and dropping. The device then sends the edit instructions to the server, including the changes and new coordinates.

[0711] Input: User editing instructions

[0712] Output: Sending edit instructions to the server

[0713] Step 8:

[0714] Regenerate the image

[0715] The server regenerates the image based on the received editing instructions, and the regenerated image data is sent from the server to the terminal and displayed again.

[0716] Input: User editing instructions

[0717] Output: Regenerated image

[0718] Step 9:

[0719] Applying Styles

[0720] The user selects a specific style (e.g., "oil painting") from the device's GUI. The device sends a style application instruction to the server, which then applies the style to the image using the style transfer model. The image after the style application is sent from the server to the device.

[0721] Input: User style selection

[0722] Output: Image with style applied

[0723] Step 10:

[0724] Regenerating Partial Changes

[0725] The user inputs instructions to change a specific part (e.g., "Make the edges of the mountain peaks look sharper"). The device sends the instructions to the server, which then regenerates the image based on the instructions. The updated image is sent to the device and presented to the user.

[0726] Input: User's partial change instructions

[0727] Output: Partially regenerated image

[0728] This allows the system to combine advanced AI technology with an intuitive GUI, providing an environment in which users can easily generate and edit high-quality images.

[0729] (Application example 1)

[0730] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] Conventional advertising image generation technologies are difficult for users without specialized knowledge to operate and have limited customization options, making it difficult to efficiently create high-quality advertising images. Furthermore, the process of fine-tuning and applying styles is not intuitive, making it difficult to improve the user experience.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0733] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generating means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating the specified image parts, means for generating and editing an advertising image based on the text instructions, and means for editing fine details of the generated image by drag and drop, thereby enabling users to intuitively and efficiently generate and edit high-quality advertising images without specialized knowledge.

[0734] A "user" is an individual or organization that has the ability to create and edit advertising images using the image generation system.

[0735] "Text instructions" refer to specific words or keywords that a user inputs for generating or editing an advertisement image.

[0736] "Natural language processing means" is a technology for analyzing input text instructions and extracting keywords and generating prompts based on those instructions.

[0737] "Generative artificial intelligence means" refers to AI technology for generating high-resolution, customizable images based on analyzed text instructions.

[0738] "High-resolution customizable images" refers to high-quality images that can be extensively edited and styled according to user instructions and preferences.

[0739] A "server" is a computing device that receives text instructions, parses them, generates images, and transmits them to a user's device.

[0740] "User's device" refers to a device (e.g., smartphone, tablet, or PC) that a user uses to access and operate the image generation system.

[0741] "Editing tools" refers to functionality that allows users to adjust the fine details of an image on their device.

[0742] "Means for applying a style" refers to a technique for reflecting a user-specified style (e.g., classic, modern, etc.) in the generated image.

[0743] The "means for regenerating a portion of an image" is a technique for regenerating an image based on additional instructions from the user for a specific portion.

[0744] "Drag-and-drop editing" refers to a feature that allows users to intuitively move and modify specific parts of an image by dragging and dropping.

[0745] MODE FOR CARRYING OUT THE INVENTION

[0746] An embodiment of the present invention will be described below: This system uses a smartphone as a main terminal as an application that can easily generate and edit advertising images.

[0747] Overall system configuration

[0748] The system consists of the following components:

[0749] User device: Smartphone (iOS or Android)

[0750] Server: Responsible for receiving text instructions, analyzing them, generating images, and sending them

[0751] Generative AI methods: AI techniques that generate high-resolution, customizable images

[0752] Program processing

[0753] Initial text entry and submission

[0754] A user uses their smartphone to enter specific advertising instructions as text, such as "Summer sale, 50% off, blue background, sun icon," and the smartphone app sends the text instructions to a cloud server.

[0755] Natural Language Processing and Prompt Generation

[0756] The server analyzes the received text instructions using natural language processing tools (e.g., SpaCy, NLTK). Specifically, it extracts important keywords from the text (e.g., "summer sale," "50% off," "blue background," "sun icon") and generates a prompt suitable for a generative AI model such as GPT-4 based on these keywords.

[0757] Image generation

[0758] Using the generated prompts, the server sends instructions to a generating artificial intelligence means (e.g., DALL-E), which generates a high-resolution advertising image based on the prompts and returns it to the server, which then sends the generated image to the user's smartphone.

[0759] Editing and styling

[0760] Editing images

[0761] The smartphone app provides users with an intuitive graphical user interface (GUI) that allows them to fine-tune the image. Users can select specific parts of the image and change their position and size by dragging and dropping. Editing instructions are then sent back to the server, which then regenerates the image based on the user's instructions.

[0762] Applying Styles

[0763] Users select a style, such as "Classic" or "Modern," from the app's style change options. This information is also sent to the server, which then applies the style to the image using generative artificial intelligence methods. The resulting regenerated image is sent to the smartphone.

[0764] Partial image regeneration and final confirmation

[0765] Partial changes

[0766] When a user requests a specific change, for example, "increase the icon size," the request is also sent to the cloud server, which then regenerates the image of the corresponding part.

[0767] Final confirmation and saving

[0768] Users can view the final ad image on their smartphone and save it or share it on social media or cloud storage.

[0769] Example prompt

[0770] "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background."

[0771] This allows even non-expert users to efficiently and intuitively create and edit high-resolution, customizable advertising images. The combination of generative artificial intelligence and natural language processing technology enables high levels of customization and an improved user experience.

[0772] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0773] Step 1:

[0774] Entering and sending text instructions

[0775] The user opens the app on their smartphone and inputs the text instructions needed to generate the ad image. For example, they might input "Summer sale, 50% off, blue background, sun icon." The device then sends this input text to the cloud server. The input in this step is the user's text instructions, and the output is the data sent to the cloud server.

[0776] Step 2:

[0777] Text analysis and keyword extraction

[0778] The cloud server analyzes the received text instructions using a natural language processing engine (e.g., SpaCy, NLTK). Specifically, it tokenizes the text and extracts important keywords ("summer sale," "50% off," "blue background," "sun icon"). The input of this step is the text instruction sent by the user, and the output is the analyzed keywords.

[0779] Step 3:

[0780] Prompt Generation

[0781] The server generates a prompt appropriate for the generative AI model (e.g., GPT-4) based on the extracted keywords. This prompt will be a sentence such as "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background." The input of this step is the parsed keywords, and the output is the generated prompt sentence.

[0782] Step 4:

[0783] Image generation

[0784] The server sends the generated prompt text to an image generation model (e.g., DALL-E). DALL-E generates a high-resolution advertising image based on the prompt and returns it to the server. The input of this step is the prompt text, and the output is a high-resolution advertising image.

[0785] Step 5:

[0786] Image transmission and initial display

[0787] The server sends the generated high-resolution advertising image to the smartphone. The device displays the received image to the user. The input of this step is the generated advertising image, and the output is the display data for the smartphone.

[0788] Step 6:

[0789] Editing images

[0790] The user uses the smartphone's GUI to adjust the details of the image, for example, by dragging and dropping the sun icon to the center. The device then sends these editing instructions to the cloud server. The input of this step is the user's editing instructions, and the output is the data sent to the cloud server.

[0791] Step 7:

[0792] Regenerate Edits

[0793] The server regenerates the image based on the received editing instructions. The regenerated image is sent to the smartphone. The input of this step is the user's editing instructions, and the output is the regenerated image.

[0794] Step 8:

[0795] Applying Styles

[0796] The user selects a style, such as "Classic" or "Modern." The device sends the style change information to the server. The server then uses a generative artificial intelligence method to apply the specified style to the image. The input to this step is the user-specified style, and the output is the image with the style applied.

[0797] Step 9:

[0798] Partial Image Regeneration

[0799] The user instructs the server to change a specific part (e.g., "increase the icon size"). The device sends the instruction to change the part to the server. The server regenerates the specific part according to the instruction and resends the image to the smartphone. The input of this step is the instruction to change the part, and the output is the regenerated image.

[0800] Step 10:

[0801] Final confirmation and saving

[0802] The user can then check the final ad image and save it on their smartphone or share it on social media or cloud storage if desired. The input of this step is the final ad image, and the output is the saved or shared image.

[0803] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0804] The interactive AI image studio of the present invention is a system that receives text instructions entered by a user, analyzes them using natural language processing means, and generates high-resolution, customizable images using generative AI means. By combining this system with an emotion engine, it can recognize the user's emotions and generate and edit images based on those emotions.

[0805] System configuration

[0806] This system is mainly composed of three elements: a server, a terminal, and a user. In particular, the addition of an emotion engine enables more personalized image generation according to the user's emotional state.

[0807] Program processing flow

[0808] 1. Initial text input and submission

[0809] The user inputs specific instructions into the device's GUI as text, such as "The silhouette of a wolf standing on top of a mountain with a sunset sky in the background."

[0810] The terminal sends this text instruction to the server.

[0811] 2. Text analysis and prompt generation

[0812] The server analyzes the received text instructions using natural language processing means, specifically by extracting important keywords from the text.

[0813] 3. Emotion analysis

[0814] The emotion engine analyzes the user's input text, on-device interactions, voice, and facial expressions to recognize the user's emotional state.

[0815] For example, if the user is expressing a "happy" emotion, a prompt containing colors and motifs that match that emotion is generated.

[0816] 4. Image Generation

[0817] The server sends prompts to the generating artificial intelligence means to generate high resolution images based on the parsed text instructions and emotion data.

[0818] An image is generated by a generating artificial intelligence means (eg, DALL-E) and sent back to the server.

[0819] 5. Sending and displaying images

[0820] The server transmits the generated image to the terminal.

[0821] The terminal displays the received image to the user.

[0822] Specific editing and style application

[0823] 1. Editing images

[0824] The user can check the image on the device's GUI and make fine adjustments as needed, for example, by selecting an image part and dragging and dropping it to move the wolf to the center.

[0825] 2. Applying Styles

[0826] The user selects a particular style, such as "oil painting," from the style change options.

[0827] The terminal transmits the selected style information to the server, and the server performs processing to apply the selected style to the image.

[0828] Partial changes and emotional readjustments

[0829] 1. Partial Image Regeneration

[0830] The user inputs instructions for partial changes such as "make the edges of the mountain peaks look sharper."

[0831] The terminal transmits a partial change instruction to the server, and the server performs partial image regeneration based on the instruction.

[0832] 2. Emotional Recalibration

[0833] The emotion engine refers to the user's input and editing history and suggests additional adjustments and optimizations based on the user's emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[0834] 3. Final confirmation and saving

[0835] The user checks the final image and saves it to the device or to the server as needed.

[0836] This system allows users to easily create and edit high-quality images that are personalized according to their emotional state. The introduction of an emotion engine further improves the user experience, making the image creation and editing process more intuitive and effective.

[0837] The processing flow will be explained below.

[0838] Specific processing steps

[0839] Step 1:

[0840] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided in the GUI of the terminal.

[0841] Step 2:

[0842] The terminal receives the user's text instructions and transmits the data to the server.

[0843] Step 3:

[0844] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[0845] Step 4:

[0846] The emotion engine analyzes the user's input text and interactions to recognize the user's emotions. For example, it recognizes the emotion "fun" from the input text and user actions.

[0847] Step 5:

[0848] Based on the keywords and emotion recognition results, the server generates a prompt with details about the image to be generated. For example, if the "happy" emotion is recognized, a prompt using bright colors is generated.

[0849] Step 6:

[0850] The server sends the generated prompt to the generating artificial intelligence means, which makes a request to generate an image based on the prompt.

[0851] Step 7:

[0852] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[0853] Step 8:

[0854] The server receives the image data and transmits it to the user's terminal.

[0855] Step 9:

[0856] The device displays the received image to the user, who can review the image and make minor adjustments as needed.

[0857] Step 10:

[0858] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[0859] Step 11:

[0860] The terminal transmits the user's editing instructions to the server.

[0861] Step 12:

[0862] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[0863] Step 13:

[0864] The server transmits the regenerated image data to the user's terminal.

[0865] Step 14:

[0866] The user selects a style such as "oil painting" from the device's toolbar.

[0867] Step 15:

[0868] The terminal transmits the user's style selection instruction to the server.

[0869] Step 16:

[0870] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[0871] Step 17:

[0872] The server transmits the image data after applying the style to the user's terminal.

[0873] Step 18:

[0874] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[0875] Step 19:

[0876] The terminal transmits a partial change instruction to the server.

[0877] Step 20:

[0878] Based on the partial change instruction received by the server, the image of the relevant part is regenerated.

[0879] Step 21:

[0880] The server transmits the regenerated partially modified image to the user's terminal.

[0881] Step 22:

[0882] The emotion engine re-analyzes the user's emotional state based on their input and editing history, and suggests additional adjustments and optimizations based on their emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[0883] Step 23:

[0884] The user checks the final image and saves it to the device or to the server as needed.

[0885] Example 2

[0886] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0887] Conventional image generation systems often cannot predict the results of a generated image, even when a user inputs clear instructions, making it difficult to obtain an image that reflects the user's intentions. Furthermore, they lack the functionality to generate images that reflect the user's emotional state, making it difficult to generate personalized images. Furthermore, the functionality for partially modifying an image or applying styles is limited, which can make it time-consuming and labor-intensive to generate an image that satisfies the user.

[0888] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a text instruction input by a user, natural language processing means for analyzing the input text instruction, artificial intelligence means for generating a high-resolution customizable image based on the analyzed text instruction and the user's emotional data, an emotion engine for analyzing the user's emotional state, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating a specified portion of the image. This makes it possible to quickly generate high-quality, personalized images that correspond to the user's emotional state and to edit them intuitively and efficiently.

[0889] The "means for receiving user-entered text instructions" is a module for electronically receiving user-entered instructions in text form.

[0890] The "natural language processing means for analyzing input text instructions" is a module that includes techniques and algorithms for analyzing input text instructions and understanding their content.

[0891] A "generative artificial intelligence means for generating high-resolution customizable images" is a module that utilizes artificial intelligence technology to generate high-quality, detailed, customizable images based on specified conditions.

[0892] The "emotion engine that analyzes the user's emotional state" is a module that includes technologies and algorithms for analyzing emotions from the user's input and actions and determining their state.

[0893] The "means for transmitting the generated image to the user's device" is a module for transferring the generated image data to the user's device via a network.

[0894] The "editing tool for adjusting fine details of an image" is a module that provides an interface for users to modify and adjust the details of the generated image.

[0895] The "means for applying a user-specified style to an image" is a module for applying a particular style selected by the user (e.g., oil painting style, watercolor style, etc.) to the generated image.

[0896] The "means for regenerating a specified portion of an image" is a module for regenerating and modifying a specific portion of an image based on a user's instructions.

[0897] MODE FOR CARRYING OUT THE INVENTION

[0898] The interactive AI image studio of the present invention is a system that receives and analyzes user-entered text instructions to generate high-resolution, customizable images. The system includes an emotion engine that recognizes the user's emotional state and generates and edits images accordingly. An embodiment of the system is described in detail below.

[0899] System Components

[0900] This system mainly consists of three elements: a server, a terminal, and a user.

[0901] 1. Server:

[0902] The server has the following functions:

[0903] Means for receiving user-entered text instructions

[0904] Analyzing the text using natural language processing tools (e.g., SpaCy)

[0905] Analysis of the user's emotional state using an emotion engine (e.g., OpenAI's GPT-3)

[0906] Generation of images by artificial intelligence means (e.g., DALL-E)

[0907] A means for transmitting the generated image data to a terminal

[0908] 2. Terminal:

[0909] The device has the following features:

[0910] GUI for user input

[0911] Displaying received images

[0912] Provides tools for editing the finer details of an image

[0913] A means for sending user-specified styles to the server and applying them

[0914] Means for sending partial image regeneration instructions to a server

[0915] 3. User:

[0916] The user performs the following operations:

[0917] Enter text instructions (e.g., "A silhouette of a wolf standing on a mountain peak against a sunset sky")

[0918] View and edit the generated image

[0919] Selecting and applying a specific style

[0920] Partial changes, final confirmation and saving

[0921] System Operation

[0922] Using this system, users can easily perform the following process:

[0923] Initial input and submission:

[0924] The user inputs text into the terminal and sends it to the server. Example input: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[0925] Natural Language Processing and Sentiment Analysis:

[0926] The server analyzes the received text instructions using natural language processing to extract important keywords. The emotion engine then analyzes the user's emotional state from the text. For example, if the emotion "fun" is detected, the server generates a prompt that matches that emotion.

[0927] Image generation:

[0928] The server generates a high-resolution image using a generative AI model (e.g., DALL-E) based on the prompt and emotion data. Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors."

[0929] View and edit images:

[0930] The server sends the generated image to the terminal, where the user can view it. The user can then edit the image in detail on the terminal, adjusting the position and color tone as necessary.

[0931] Applying styles:

[0932] The user selects a style change option (e.g., "oil painting"), and the device sends that information to the server, which applies the selected style to the image and sends it back to the device.

[0933] Modify and regenerate:

[0934] The user instructs the terminal to make specific partial changes, and the terminal sends the instruction to the server, which then regenerates the image based on the instruction and sends the updated image to the terminal.

[0935] Final review and save:

[0936] The user reviews the final image and saves it to local storage or cloud storage if desired.

[0937] Specific examples

[0938] For example, if the user enters the prompt "A bird sitting in a tree under a blue sky," the following prompt sentence is generated:

[0939] Prompt: "In the clear blue sky, a bird perched on a tree, with warm and bright colors indicating happiness."

[0940] Based on this prompt, the generative AI model generates an image that reflects the user's emotional state and sends it to the user's device via the server. The user can then review the image, edit and change the style as needed, and finally save it.

[0941] The system allows users to quickly and intuitively generate and edit personalized, high-quality images that reflect their emotional state. The introduction of an emotion engine further improves the user experience and makes the image generation and editing process more effective.

[0942] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0943] Program processing flow

[0944] Step 1: User enters and submits text

[0945] Specific actions

[0946] The user enters text instructions into the terminal's GUI.

[0947] Example: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[0948] The user clicks the submit button.

[0949] The terminal sends the entered text instructions to the server.

[0950] Input: The text instruction entered by the user

[0951] Output: HTTP request sent to the server

[0952] Step 2: Parsing text on the server

[0953] Specific actions

[0954] The server passes the received text instructions to a natural language processing (NLP) module.

[0955] The NLP module analyzes the text and extracts important keywords.

[0956] Examples: "sunset," "sky," "mountain," "peak," "wolf," "silhouette"

[0957] Input: Received text instructions

[0958] Output: Extracted keywords

[0959] Step 3: Sentiment Analysis

[0960] Specific actions

[0961] The server sends the text instructions to the emotion engine.

[0962] An emotion engine analyzes the user's emotional state from the text.

[0963] Example: Detecting the emotion "happy"

[0964] Input: Text instructions

[0965] Output: Emotion data

[0966] Step 4: Generate prompts

[0967] Specific actions

[0968] The server generates a prompt sentence based on the analyzed keywords and emotion data.

[0969] Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors"

[0970] Input: Keywords, emotion data

[0971] Output: prompt statement

[0972] Step 5: Image generation

[0973] Specific actions

[0974] The server sends the generated prompt sentence to the generating artificial intelligence means (generating AI model).

[0975] Example: DALL-E

[0976] The generative AI model generates high-resolution images and sends them back to the server.

[0977] Input: prompt statement

[0978] Output: Generated image data

[0979] Step 6: Receiving and displaying images

[0980] Specific actions

[0981] The server transmits the generated image data to the terminal.

[0982] The terminal receives the image and displays it to the user.

[0983] Input: Generated image data

[0984] Output: Image displayed on the terminal

[0985] Step 7: Edit the image

[0986] Specific actions

[0987] The user edits the image to adjust the finer details.

[0988] Example: Move the wolf to the center

[0989] The device provides editing tools, allowing users to select parts of the image and adjust their position by dragging and dropping.

[0990] Input: Generated image

[0991] Output: Edited image data

[0992] Step 8: Applying Styles

[0993] Specific actions

[0994] The user selects a specific style from the style change options.

[0995] Example: "Oil painting style"

[0996] The terminal transmits style information to the server.

[0997] The server applies the style to the image and sends it back to the device.

[0998] Input: Style information

[0999] Output: Image data with style applied

[1000] Step 9: Partial Image Regeneration

[1001] Specific actions

[1002] The user inputs instructions for partial modification.

[1003] Example: "Make the edges of the mountain peaks look sharper."

[1004] The terminal transmits a partial change instruction to the server.

[1005] The server regenerates a portion of the image based on the instruction and transmits the updated image to the terminal.

[1006] Input: Partial change instruction

[1007] Output: Partially regenerated image

[1008] Step 10: Final review and save

[1009] Specific actions

[1010] The user reviews the final image.

[1011] The user clicks the Save button.

[1012] The device stores the image data in local storage or cloud storage.

[1013] Input: Last seen image

[1014] Output: Saved image data

[1015] This process flow allows users to quickly generate high-quality, personalized images according to their emotional state, and then intuitively edit and save them.

[1016] (Application example 2)

[1017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1018] In conventional advertising production, it was difficult to customize according to the user's emotional state, making it difficult to generate intuitive and effective advertising images. It was also not easy to reflect the style and editing instructions specified in the early stages of ad production and apply them to the image immediately. As a result, the effectiveness of the advertisement was not maximized, and it was not possible to achieve advertising expressions that were in tune with the emotions of the target user.

[1019] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generation means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating a specified part of the image, emotion analysis means for recognizing the user's emotional state and generating and editing an image based on the emotion, and means for enabling application in the field of advertising production. This enables the generation and immediate editing of personalized advertising images according to the user's emotions, thereby realizing the provision of optimal advertisements to target users.

[1020] "User" means a person who uses the Interactive AI Image Studio to generate and edit images.

[1021] A "text instruction" is a sentence that describes specific details about the image that the user wants to generate.

[1022] "Natural language processing means" refers to means for analyzing input text instructions and extracting important keywords and context.

[1023] "Generative artificial intelligence means" refers to artificial intelligence techniques for generating high resolution customizable images based on analyzed instructions.

[1024] The "editing means" is a function that allows the user to adjust the fine details of the generated image.

[1025] A "style" is a particular image presentation or visual effect specified by the user.

[1026] The "regeneration means" is a function for regenerating a specified part of an image.

[1027] The "emotion analysis means" is a function for recognizing the user's emotional state and generating and editing images based on that emotion.

[1028] The "advertising production field" is the field of creating images and content to effectively promote products and services to target users.

[1029] The present invention is a system that utilizes an interactive AI image studio to generate high-resolution, customizable advertising images based on user-entered text instructions. The main components of the system include a server, a terminal, and a user.

[1030] Server Configuration

[1031] The server includes the following means:

[1032] Text receiving means: A means for receiving a text instruction input by a user from a terminal.

[1033] Natural language processing means: A means for analyzing input text instructions, extracting keywords, and generating optimal prompts for the artificial intelligence generator. Specifically, a natural language processing engine such as SpaCy is used.

[1034] Generative AI means: A means for generating high-resolution, customizable images based on analyzed instructions, specifically using a generative AI model (e.g., DALL-E).

[1035] Image transmission means: A means for transmitting the generated image to the user's terminal.

[1036] Emotion analysis means: A means of recognizing a user's emotional state by analyzing the user's input text, interactions, voice input, and facial expression data. The IBM Watson Emotion Analysis API can be used.

[1037] Application in the field of advertising production: A means to make it possible to use the generated images and emotion analysis results in advertising production.

[1038] Device configuration

[1039] The terminal includes:

[1040] GUI editing means: A graphical user interface function that allows users to adjust the fine details of the generated image. This allows users to intuitively adjust the image.

[1041] Style application: A function for applying a user-specified style to an image. Users can select styles such as "oil painting" or "vintage."

[1042] Partial regeneration means: A function for regenerating a specified part of an image.

[1043] User operations

[1044] The user performs the following steps:

[1045] 1. Enter the basic concept of the advertisement, such as "people relaxing at the beach," into the GUI of a smartphone or other device.

[1046] 2. The text instructions are sent to the server, where the server's natural language processing means analyzes the text and extracts keywords.

[1047] 3. The emotion analyzer recognizes the user's emotional state and generates appropriate prompts based on that data.

[1048] 4. The generating artificial intelligence means (such as DALL-E) generates a high-resolution image based on the prompt and sends it to the terminal via the server.

[1049] 5. The user can review the received image on their device, make minor adjustments or apply styles as needed, and even request partial regeneration.

[1050] Examples of concrete examples and prompts

[1051] As a concrete example, consider a situation where a user enters the text "A park scene with a happy family together" for an advertising image, and the emotion is "joy." An example prompt for this would be:

[1052] Scene of happy family spending time in the park, joy

[1053] In this way, users can intuitively create and edit high-quality advertising images in a short amount of time. This system enables advertising expressions that are in tune with the emotions of target users, maximizing their effectiveness.

[1054] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1055] Step 1:

[1056] The user inputs the basic idea of ​​the advertisement into the terminal's GUI (text instructions).

[1057] Input: A textual indication of user input (e.g., "People relaxing at the beach").

[1058] Output: A request to send input text instructions.

[1059] Specific operation: The user enters instructions into the terminal's GUI and presses the "Send" button, which sends a request to the server.

[1060] Step 2:

[1061] The server analyzes the received text instructions using natural language processing means.

[1062] Input: The text instruction submitted by the user.

[1063] Output: Extracted keywords and analysis results.

[1064] What happens: The server's natural language processing engine (e.g., SpaCy) analyzes the text and extracts important keywords and contextual information.

[1065] Step 3:

[1066] An emotion analyzer recognizes the user's emotional state.

[1067] Input: Text input, on-device interaction data, voice and facial expression data.

[1068] Output: Perceived emotional state (e.g., "joy").

[1069] What happens: The server analyzes the data using an emotion analysis engine (e.g., IBM Watson Emotion Analysis API) to determine the user's emotional state.

[1070] Step 4:

[1071] The server generates a prompt based on the analysis result and sends it to the generating artificial intelligence means.

[1072] Input: Keywords, contextual information, emotional state.

[1073] Output: The prompt statement.

[1074] Specific operation: The server integrates the extracted keywords and emotion data to generate a prompt sentence to send to the generative AI model (e.g., DALL-E).

[1075] Step 5:

[1076] A generating artificial intelligence means generates a high resolution image.

[1077] Input: Prompt sentence (e.g., "A park scene of a happy family spending time together, joy").

[1078] Output: High resolution customizable images.

[1079] What it does: A prompt is sent to the generative AI model, which then generates an image based on it.

[1080] Step 6:

[1081] The server transmits the generated image to the terminal.

[1082] Input: The generated image.

[1083] Output: Request to send image data.

[1084] Specific operation: The server receives the generated image and sends it to the user's device.

[1085] Step 7:

[1086] The user checks the image on the terminal and edits it as necessary.

[1087] Input: Received image data.

[1088] Output: Editing instructions and adjusted image data.

[1089] Specific behavior: The user views the image in the device's GUI and uses drag-and-drop and tools to make fine adjustments and apply styles.

[1090] Step 8:

[1091] The user requests a partial regeneration.

[1092] Input: Partial regeneration instructions.

[1093] Output: The regenerated subimage.

[1094] Specific operation: The user selects a specific image part, inputs a regeneration instruction, and sends it to the server. The server generates a prompt again, requests the generative AI model to retrieve the regenerated image part, and sends it to the user.

[1095] Step 9:

[1096] The user reviews and saves the final image.

[1097] Input: The final image data after adjustments and regeneration.

[1098] Output: Saved image file.

[1099] Specific behavior: The user confirms the final image and saves it to the device or server.

[1100] The above are the specific processing steps of the system that realizes the application example.

[1101] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1102] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1103] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1104] [Third embodiment]

[1105] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1106] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1107] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1108] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1109] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1110] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1111] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1112] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1113] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1114] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1115] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1116] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1117] DETAILED DESCRIPTION OF THE INVENTION The present invention is an image creation and editing system that allows users to input text instructions to create high-resolution, customizable images, and then fine-tune and style them through an intuitive graphical user interface (GUI).

[1118] System Overview

[1119] This system consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, image generation, image editing, and transmission. The terminal provides a GUI and an interface for users to generate and edit images. Users input text and check and edit images through the terminal.

[1120] Program processing flow

[1121] 1. Initial text input and submission

[1122] The user enters specific text instructions into the device's GUI using an input field, for example, "A silhouette of a wolf standing on a mountain top against a sunset sky."

[1123] The terminal sends this text instruction to the server.

[1124] 2. Text analysis and prompt generation

[1125] The server analyzes the received text, specifically using a natural language processing (NLP) engine to extract important keywords ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1126] Based on the extracted keywords, a generative artificial intelligence (AI) model generates suitable prompts.

[1127] 3. Image Generation

[1128] The server sends the generated prompt to a generative artificial intelligence model, such as a deep learning image generation model (e.g., DALL-E).

[1129] An image generation model generates a high-resolution image based on the prompt and returns it to the server.

[1130] 4. Image transmission and initial display

[1131] The server transmits the generated image to the terminal.

[1132] The terminal displays the received image to the user.

[1133] Specific editing and style application

[1134] 1. Editing images

[1135] The user uses the device's GUI to adjust the details of the image, for example, to move the wolf to the center, by selecting an image part and dragging and dropping it to adjust its position.

[1136] The terminal sends these editing instructions to the server, which then regenerates the image based on them.

[1137] 2. Applying Styles

[1138] The user selects a particular style, such as "oil painting," from the style change options.

[1139] The terminal transmits the selected style information to the server.

[1140] The server processes the image to apply the selected style and sends it back to the terminal.

[1141] Partial changes and final confirmation

[1142] 1. Partial Image Regeneration

[1143] The user inputs specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper."

[1144] The terminal transmits a partial change instruction to the server.

[1145] The server performs partial image regeneration according to the instructions and transmits the updated image to the terminal.

[1146] 2. Final confirmation and saving

[1147] The user reviews the final image.

[1148] The user can save the final image to their device or to a server as desired.

[1149] This allows users to easily create and edit high-quality images without specialized knowledge. The system offers high customizability to meet diverse user needs, enabling efficient image creation. The intuitive GUI and advanced AI technology also enhance the user experience.

[1150] The processing flow will be explained below.

[1151] Step 1:

[1152] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided by the terminal GUI.

[1153] Step 2:

[1154] The terminal receives the user's text instructions and transmits the data to the server.

[1155] Step 3:

[1156] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1157] Step 4:

[1158] The server generates an optimal prompt for the artificial intelligence generating means based on the extracted keywords.

[1159] Step 5:

[1160] The server sends the generated prompt to the generative artificial intelligence model and makes a request to generate an image based on the prompt.

[1161] Step 6:

[1162] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[1163] Step 7:

[1164] The server receives the image data and transmits it to the user's terminal.

[1165] Step 8:

[1166] The device displays the received image to the user, who can then use the device's editing tools to review the image and make minor adjustments as needed.

[1167] Step 9:

[1168] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[1169] Step 10:

[1170] The terminal transmits the user's editing instructions to the server.

[1171] Step 11:

[1172] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[1173] Step 12:

[1174] The server transmits the regenerated image data to the user's terminal.

[1175] Step 13:

[1176] The user selects a style such as "oil painting" from the device's toolbar.

[1177] Step 14:

[1178] The terminal transmits the user's style selection instruction to the server.

[1179] Step 15:

[1180] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[1181] Step 16:

[1182] The server transmits the image data after applying the style to the user's terminal.

[1183] Step 17:

[1184] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[1185] Step 18:

[1186] The terminal transmits a partial change instruction to the server.

[1187] Step 19:

[1188] The server regenerates the image of the relevant part based on the received partial change instruction.

[1189] Step 20:

[1190] The server transmits the regenerated partially modified image to the user's terminal.

[1191] Step 21:

[1192] The user checks the final image and saves it to the device or to the server as needed.

[1193] Example 1

[1194] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1195] Conventional image generation and editing systems have the drawback of making it difficult for users without specialized knowledge to generate the desired high-quality, customizable images. It is also often difficult for users to intuitively edit the fine details of an image or apply a specific style. Furthermore, partial image regeneration is not easily possible, making it difficult to flexibly edit images according to user needs.

[1196] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1197] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, means for generating an appropriate prompt sentence based on the analyzed text instructions, artificial intelligence generating means for generating a high-resolution, customizable image using the generated prompt sentence, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating the specified image portion. This allows users to intuitively generate and edit high-quality images without specialized knowledge. Furthermore, it is easy to apply specific styles and regenerate partial images, enabling flexible image editing according to user requests.

[1198] The "means for receiving text instructions entered by a user" is a function for transmitting specific text instructions entered by a user into an input field of a terminal to a server.

[1199] The "natural language processing means for analyzing input text instructions" is a natural language processing engine for analyzing the text entered by the user and extracting important keywords and phrases.

[1200] The "means for generating appropriate prompt sentences" is a function that generates prompt sentences in a format that is easy for the image generation model to understand, based on extracted keywords and phrases.

[1201] The "generative artificial intelligence means for generating high-resolution customizable images" is an artificial intelligence model for generating high-resolution and customizable images based on generated prompt text.

[1202] The "means for transmitting the generated image to the user's device" is a function for transmitting the generated image data from the server to the user's terminal.

[1203] The "editing means for adjusting fine details of an image" is an interface that provides a function for a user to intuitively edit fine details of a generated image.

[1204] The "means for applying a style designated by the user to an image" is a function for applying a style selected by the user (for example, "oil painting style") to a generated image.

[1205] The "means for regenerating a specified portion of an image" is a function for partially regenerating an image when a user instructs a change to a specific portion.

[1206] The image generation and editing system of this invention consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, prompt generation, and image generation and editing. The terminal provides an intuitive graphical user interface (GUI) for users to generate and edit images. Users input text instructions through the terminal and view and edit the generated images.

[1207] Text analysis and prompt generation

[1208] When the server receives the text instructions entered by the user, it analyzes the text using a natural language processing (NLP) engine. Examples of NLP engines that can be used include spaCy and NLTK. This engine extracts important keywords and phrases (e.g., "sunset," "sky," "mountain," "peak," and "silhouette of a wolf") from the text. Based on the extracted keywords, the server generates a prompt suitable for the generative AI model. This prompt is in a format that is easy for the generative AI model to understand, and can be a specific instruction such as "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1209] Image generation

[1210] The server sends the generated prompt text to a generative AI model. The generative AI model uses an image generation model such as DALL-E, which uses deep learning. This model generates a high-resolution, customizable image based on the prompt text. The generated image is returned to the server, which then sends it to the user's device. The image data sent is encoded, for example, in Base64 format.

[1211] View and edit images

[1212] The user checks the generated image on the device and edits it as necessary. The device's GUI provides an editing tool that allows the user to intuitively adjust minute details of the image. For example, the user can select an image part and drag and drop it to move the wolf to the center. The device sends the editing instructions to the server, and the server regenerates the image and sends the updated image back to the device.

[1213] Style application and partial regeneration

[1214] A user can select a specific style (e.g., "oil painting") from the GUI of the device and apply it to an image. When the device sends the style information to the server, the server applies the style to the image using a style transfer model. The regenerated image with the style applied is sent to the device. Furthermore, the user can input specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper." The server regenerates parts of the image based on the instructions and sends the updated image to the device.

[1215] An example of a prompt is "The silhouette of a wolf standing on top of a mountain against a sunset sky." After parsing this, the prompt becomes "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1216] This allows users to intuitively create and edit high-quality images without specialized knowledge. The system combines advanced AI technology with an intuitive GUI to create easily customizable images.

[1217] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1218] Step 1:

[1219] Enter and send text

[1220] The user enters specific instructions as text into an input field on the device. For example, they might enter "A silhouette of a wolf standing on a mountain top against a sunset sky." This input text is sent to the server. The device uses an HTTP request to send the text data to the server in JSON format.

[1221] Input: User's text instructions

[1222] Output: Sends text instructions to the server

[1223] Step 2:

[1224] Text analysis and keyword extraction

[1225] The server parses the received text instructions. A natural language processing (NLP) engine (e.g., spaCy or NLTK) tokenizes the text and extracts important keywords and phrases. Extracted keywords include "sunset," "sky," "mountain," "peak," and "wolf silhouette."

[1226] Input: Received text instructions

[1227] Output: Extracted keywords and phrases

[1228] Step 3:

[1229] Generate prompt statement

[1230] The server generates a prompt sentence appropriate for the generative AI model based on the extracted keywords and phrases. For example, the generated prompt sentence might be "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1231] Input: Extracted keywords or phrases

[1232] Output: Generated prompt statement

[1233] Step 4:

[1234] Sending an image generation request

[1235] The server sends the generated prompt text to a generative AI model, which then uses deep learning to generate an image based on the text (e.g., DALL-E).

[1236] Input: Generated prompt text

[1237] Output: The generated image

[1238] Step 5:

[1239] Sending the generated image

[1240] The server receives the generated image and sends the image data to the device. The image data is encoded in Base64 format. The device receives the image data in JSON format using an HTTP response.

[1241] Input: Generated image

[1242] Output: Sending image data to the device

[1243] Step 6:

[1244] Displaying images

[1245] The terminal decodes the received image data and displays it to the user. Specifically, the image after Base64 decoding is displayed as HTML. Display it with tags etc.

[1246] Input: Received image data

[1247] Output: Displaying the image to the user

[1248] Step 7:

[1249] Editing images

[1250] The user uses the GUI on the device to adjust small details of the image. For example, to move the wolf to the center, the user selects an image part and adjusts its position by dragging and dropping. The device then sends the edit instructions to the server, including the changes and new coordinates.

[1251] Input: User editing instructions

[1252] Output: Sending edit instructions to the server

[1253] Step 8:

[1254] Regenerate the image

[1255] The server regenerates the image based on the received editing instructions, and the regenerated image data is sent from the server to the terminal and displayed again.

[1256] Input: User editing instructions

[1257] Output: Regenerated image

[1258] Step 9:

[1259] Applying Styles

[1260] The user selects a specific style (e.g., "oil painting") from the device's GUI. The device sends a style application instruction to the server, which then applies the style to the image using the style transfer model. The image after the style application is sent from the server to the device.

[1261] Input: User style selection

[1262] Output: Image with style applied

[1263] Step 10:

[1264] Regenerating Partial Changes

[1265] The user inputs instructions to change a specific part (e.g., "Make the edges of the mountain peaks look sharper"). The device sends the instructions to the server, which then regenerates the image based on the instructions. The updated image is sent to the device and presented to the user.

[1266] Input: User's partial change instructions

[1267] Output: Partially regenerated image

[1268] This allows the system to combine advanced AI technology with an intuitive GUI, providing an environment in which users can easily generate and edit high-quality images.

[1269] (Application example 1)

[1270] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1271] Conventional advertising image generation technologies are difficult for users without specialized knowledge to operate and have limited customization options, making it difficult to efficiently create high-quality advertising images. Furthermore, the process of fine-tuning and applying styles is not intuitive, making it difficult to improve the user experience.

[1272] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1273] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generating means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating the specified image parts, means for generating and editing an advertising image based on the text instructions, and means for editing fine details of the generated image by drag and drop, thereby enabling users to intuitively and efficiently generate and edit high-quality advertising images without specialized knowledge.

[1274] A "user" is an individual or organization that has the ability to create and edit advertising images using the image generation system.

[1275] "Text instructions" refer to specific words or keywords that a user inputs for generating or editing an advertisement image.

[1276] "Natural language processing means" is a technology for analyzing input text instructions and extracting keywords and generating prompts based on those instructions.

[1277] "Generative artificial intelligence means" refers to AI technology for generating high-resolution, customizable images based on analyzed text instructions.

[1278] "High-resolution customizable images" refers to high-quality images that can be extensively edited and styled according to user instructions and preferences.

[1279] A "server" is a computing device that receives text instructions, parses them, generates images, and transmits them to a user's device.

[1280] "User's device" refers to a device (e.g., smartphone, tablet, or PC) that a user uses to access and operate the image generation system.

[1281] "Editing tools" refers to functionality that allows users to adjust the fine details of an image on their device.

[1282] "Means for applying a style" refers to a technique for reflecting a user-specified style (e.g., classic, modern, etc.) in the generated image.

[1283] The "means for regenerating a portion of an image" is a technique for regenerating an image based on additional instructions from the user for a specific portion.

[1284] "Drag-and-drop editing" refers to a feature that allows users to intuitively move and modify specific parts of an image by dragging and dropping.

[1285] MODE FOR CARRYING OUT THE INVENTION

[1286] An embodiment of the present invention will be described below: This system uses a smartphone as a main terminal as an application that can easily generate and edit advertising images.

[1287] Overall system configuration

[1288] The system consists of the following components:

[1289] User device: Smartphone (iOS or Android)

[1290] Server: Responsible for receiving text instructions, analyzing them, generating images, and sending them

[1291] Generative AI methods: AI techniques that generate high-resolution, customizable images

[1292] Program processing

[1293] Initial text entry and submission

[1294] A user uses their smartphone to enter specific advertising instructions as text, such as "Summer sale, 50% off, blue background, sun icon," and the smartphone app sends the text instructions to a cloud server.

[1295] Natural Language Processing and Prompt Generation

[1296] The server analyzes the received text instructions using natural language processing tools (e.g., SpaCy, NLTK). Specifically, it extracts important keywords from the text (e.g., "summer sale," "50% off," "blue background," "sun icon") and generates a prompt suitable for a generative AI model such as GPT-4 based on these keywords.

[1297] Image generation

[1298] Using the generated prompts, the server sends instructions to a generating artificial intelligence means (e.g., DALL-E), which generates a high-resolution advertising image based on the prompts and returns it to the server, which then sends the generated image to the user's smartphone.

[1299] Editing and styling

[1300] Editing images

[1301] The smartphone app provides users with an intuitive graphical user interface (GUI) that allows them to fine-tune the image. Users can select specific parts of the image and change their position and size by dragging and dropping. Editing instructions are then sent back to the server, which then regenerates the image based on the user's instructions.

[1302] Applying Styles

[1303] Users select a style, such as "Classic" or "Modern," from the app's style change options. This information is also sent to the server, which then applies the style to the image using generative artificial intelligence methods. The resulting regenerated image is sent to the smartphone.

[1304] Partial image regeneration and final confirmation

[1305] Partial changes

[1306] When a user requests a specific change, for example, "increase the icon size," the request is also sent to the cloud server, which then regenerates the image of the corresponding part.

[1307] Final confirmation and saving

[1308] Users can view the final ad image on their smartphone and save it or share it on social media or cloud storage.

[1309] Example prompt

[1310] "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background."

[1311] This allows even non-expert users to efficiently and intuitively create and edit high-resolution, customizable advertising images. The combination of generative artificial intelligence and natural language processing technology enables high levels of customization and an improved user experience.

[1312] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1313] Step 1:

[1314] Entering and sending text instructions

[1315] The user opens the app on their smartphone and inputs the text instructions needed to generate the ad image. For example, they might input "Summer sale, 50% off, blue background, sun icon." The device then sends this input text to the cloud server. The input in this step is the user's text instructions, and the output is the data sent to the cloud server.

[1316] Step 2:

[1317] Text analysis and keyword extraction

[1318] The cloud server analyzes the received text instructions using a natural language processing engine (e.g., SpaCy, NLTK). Specifically, it tokenizes the text and extracts important keywords ("summer sale," "50% off," "blue background," "sun icon"). The input of this step is the text instruction sent by the user, and the output is the analyzed keywords.

[1319] Step 3:

[1320] Prompt Generation

[1321] The server generates a prompt appropriate for the generative AI model (e.g., GPT-4) based on the extracted keywords. This prompt will be a sentence such as "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background." The input of this step is the parsed keywords, and the output is the generated prompt sentence.

[1322] Step 4:

[1323] Image generation

[1324] The server sends the generated prompt text to an image generation model (e.g., DALL-E). DALL-E generates a high-resolution advertising image based on the prompt and returns it to the server. The input of this step is the prompt text, and the output is a high-resolution advertising image.

[1325] Step 5:

[1326] Image transmission and initial display

[1327] The server sends the generated high-resolution advertising image to the smartphone. The device displays the received image to the user. The input of this step is the generated advertising image, and the output is the display data for the smartphone.

[1328] Step 6:

[1329] Editing images

[1330] The user uses the smartphone's GUI to adjust the details of the image, for example, by dragging and dropping the sun icon to the center. The device then sends these editing instructions to the cloud server. The input of this step is the user's editing instructions, and the output is the data sent to the cloud server.

[1331] Step 7:

[1332] Regenerate Edits

[1333] The server regenerates the image based on the received editing instructions. The regenerated image is sent to the smartphone. The input of this step is the user's editing instructions, and the output is the regenerated image.

[1334] Step 8:

[1335] Applying Styles

[1336] The user selects a style, such as "Classic" or "Modern." The device sends the style change information to the server. The server then uses a generative artificial intelligence method to apply the specified style to the image. The input to this step is the user-specified style, and the output is the image with the style applied.

[1337] Step 9:

[1338] Partial Image Regeneration

[1339] The user instructs the server to change a specific part (e.g., "increase the icon size"). The device sends the instruction to change the part to the server. The server regenerates the specific part according to the instruction and resends the image to the smartphone. The input of this step is the instruction to change the part, and the output is the regenerated image.

[1340] Step 10:

[1341] Final confirmation and saving

[1342] The user can then check the final ad image and save it on their smartphone or share it on social media or cloud storage if desired. The input of this step is the final ad image, and the output is the saved or shared image.

[1343] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1344] The interactive AI image studio of the present invention is a system that receives text instructions entered by a user, analyzes them using natural language processing means, and generates high-resolution, customizable images using generative AI means. By combining this system with an emotion engine, it can recognize the user's emotions and generate and edit images based on those emotions.

[1345] System configuration

[1346] This system is mainly composed of three elements: a server, a terminal, and a user. In particular, the addition of an emotion engine enables more personalized image generation according to the user's emotional state.

[1347] Program processing flow

[1348] 1. Initial text input and submission

[1349] The user inputs specific instructions into the device's GUI as text, such as "The silhouette of a wolf standing on top of a mountain with a sunset sky in the background."

[1350] The terminal sends this text instruction to the server.

[1351] 2. Text analysis and prompt generation

[1352] The server analyzes the received text instructions using natural language processing means, specifically by extracting important keywords from the text.

[1353] 3. Emotion analysis

[1354] The emotion engine analyzes the user's input text, on-device interactions, voice, and facial expressions to recognize the user's emotional state.

[1355] For example, if the user is expressing a "happy" emotion, a prompt containing colors and motifs that match that emotion is generated.

[1356] 4. Image Generation

[1357] The server sends prompts to the generating artificial intelligence means to generate high resolution images based on the parsed text instructions and emotion data.

[1358] An image is generated by a generating artificial intelligence means (eg, DALL-E) and sent back to the server.

[1359] 5. Sending and displaying images

[1360] The server transmits the generated image to the terminal.

[1361] The terminal displays the received image to the user.

[1362] Specific editing and style application

[1363] 1. Editing images

[1364] The user can check the image on the device's GUI and make fine adjustments as needed, for example, by selecting an image part and dragging and dropping it to move the wolf to the center.

[1365] 2. Applying Styles

[1366] The user selects a particular style, such as "oil painting," from the style change options.

[1367] The terminal transmits the selected style information to the server, and the server performs processing to apply the selected style to the image.

[1368] Partial changes and emotional readjustments

[1369] 1. Partial Image Regeneration

[1370] The user inputs instructions for partial changes such as "make the edges of the mountain peaks look sharper."

[1371] The terminal transmits a partial change instruction to the server, and the server performs partial image regeneration based on the instruction.

[1372] 2. Emotional Recalibration

[1373] The emotion engine refers to the user's input and editing history and suggests additional adjustments and optimizations based on the user's emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[1374] 3. Final confirmation and saving

[1375] The user checks the final image and saves it to the device or to the server as needed.

[1376] This system allows users to easily create and edit high-quality images that are personalized according to their emotional state. The introduction of an emotion engine further improves the user experience, making the image creation and editing process more intuitive and effective.

[1377] The processing flow will be explained below.

[1378] Specific processing steps

[1379] Step 1:

[1380] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided in the GUI of the terminal.

[1381] Step 2:

[1382] The terminal receives the user's text instructions and transmits the data to the server.

[1383] Step 3:

[1384] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1385] Step 4:

[1386] The emotion engine analyzes the user's input text and interactions to recognize the user's emotions. For example, it recognizes the emotion "fun" from the input text and user actions.

[1387] Step 5:

[1388] Based on the keywords and emotion recognition results, the server generates a prompt with details about the image to be generated. For example, if the "happy" emotion is recognized, a prompt using bright colors is generated.

[1389] Step 6:

[1390] The server sends the generated prompt to the generating artificial intelligence means, which makes a request to generate an image based on the prompt.

[1391] Step 7:

[1392] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[1393] Step 8:

[1394] The server receives the image data and transmits it to the user's terminal.

[1395] Step 9:

[1396] The device displays the received image to the user, who can review the image and make minor adjustments as needed.

[1397] Step 10:

[1398] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[1399] Step 11:

[1400] The terminal transmits the user's editing instructions to the server.

[1401] Step 12:

[1402] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[1403] Step 13:

[1404] The server transmits the regenerated image data to the user's terminal.

[1405] Step 14:

[1406] The user selects a style such as "oil painting" from the device's toolbar.

[1407] Step 15:

[1408] The terminal transmits the user's style selection instruction to the server.

[1409] Step 16:

[1410] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[1411] Step 17:

[1412] The server transmits the image data after applying the style to the user's terminal.

[1413] Step 18:

[1414] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[1415] Step 19:

[1416] The terminal transmits a partial change instruction to the server.

[1417] Step 20:

[1418] Based on the partial change instruction received by the server, the image of the relevant part is regenerated.

[1419] Step 21:

[1420] The server transmits the regenerated partially modified image to the user's terminal.

[1421] Step 22:

[1422] The emotion engine re-analyzes the user's emotional state based on their input and editing history, and suggests additional adjustments and optimizations based on their emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[1423] Step 23:

[1424] The user checks the final image and saves it to the device or to the server as needed.

[1425] Example 2

[1426] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1427] Conventional image generation systems often cannot predict the results of a generated image, even when a user inputs clear instructions, making it difficult to obtain an image that reflects the user's intentions. Furthermore, they lack the functionality to generate images that reflect the user's emotional state, making it difficult to generate personalized images. Furthermore, the functionality for partially modifying an image or applying styles is limited, which can make it time-consuming and labor-intensive to generate an image that satisfies the user.

[1428] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a text instruction input by a user, natural language processing means for analyzing the input text instruction, artificial intelligence means for generating a high-resolution customizable image based on the analyzed text instruction and the user's emotional data, an emotion engine for analyzing the user's emotional state, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating a specified portion of the image. This makes it possible to quickly generate high-quality, personalized images that correspond to the user's emotional state and to edit them intuitively and efficiently.

[1429] The "means for receiving user-entered text instructions" is a module for electronically receiving user-entered instructions in text form.

[1430] The "natural language processing means for analyzing input text instructions" is a module that includes techniques and algorithms for analyzing input text instructions and understanding their content.

[1431] A "generative artificial intelligence means for generating high-resolution customizable images" is a module that utilizes artificial intelligence technology to generate high-quality, detailed, customizable images based on specified conditions.

[1432] The "emotion engine that analyzes the user's emotional state" is a module that includes technologies and algorithms for analyzing emotions from the user's input and actions and determining their state.

[1433] The "means for transmitting the generated image to the user's device" is a module for transferring the generated image data to the user's device via a network.

[1434] The "editing tool for adjusting fine details of an image" is a module that provides an interface for users to modify and adjust the details of the generated image.

[1435] The "means for applying a user-specified style to an image" is a module for applying a particular style selected by the user (e.g., oil painting style, watercolor style, etc.) to the generated image.

[1436] The "means for regenerating a specified portion of an image" is a module for regenerating and modifying a specific portion of an image based on a user's instructions.

[1437] MODE FOR CARRYING OUT THE INVENTION

[1438] The interactive AI image studio of the present invention is a system that receives and analyzes user-entered text instructions to generate high-resolution, customizable images. The system includes an emotion engine that recognizes the user's emotional state and generates and edits images accordingly. An embodiment of the system is described in detail below.

[1439] System Components

[1440] This system mainly consists of three elements: a server, a terminal, and a user.

[1441] 1. Server:

[1442] The server has the following functions:

[1443] Means for receiving user-entered text instructions

[1444] Analyzing the text using natural language processing tools (e.g., SpaCy)

[1445] Analysis of the user's emotional state using an emotion engine (e.g., OpenAI's GPT-3)

[1446] Generation of images by artificial intelligence means (e.g., DALL-E)

[1447] A means for transmitting the generated image data to a terminal

[1448] 2. Terminal:

[1449] The device has the following features:

[1450] GUI for user input

[1451] Displaying received images

[1452] Provides tools for editing the finer details of an image

[1453] A means for sending user-specified styles to the server and applying them

[1454] Means for sending partial image regeneration instructions to a server

[1455] 3. User:

[1456] The user performs the following operations:

[1457] Enter text instructions (e.g., "A silhouette of a wolf standing on a mountain peak against a sunset sky")

[1458] View and edit the generated image

[1459] Selecting and applying a specific style

[1460] Partial changes, final confirmation and saving

[1461] System Operation

[1462] Using this system, users can easily perform the following process:

[1463] Initial input and submission:

[1464] The user inputs text into the terminal and sends it to the server. Example input: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[1465] Natural Language Processing and Sentiment Analysis:

[1466] The server analyzes the received text instructions using natural language processing to extract important keywords. The emotion engine then analyzes the user's emotional state from the text. For example, if the emotion "fun" is detected, the server generates a prompt that matches that emotion.

[1467] Image generation:

[1468] The server generates a high-resolution image using a generative AI model (e.g., DALL-E) based on the prompt and emotion data. Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors."

[1469] View and edit images:

[1470] The server sends the generated image to the terminal, where the user can view it. The user can then edit the image in detail on the terminal, adjusting the position and color tone as necessary.

[1471] Applying styles:

[1472] The user selects a style change option (e.g., "oil painting"), and the device sends that information to the server, which applies the selected style to the image and sends it back to the device.

[1473] Modify and regenerate:

[1474] The user instructs the terminal to make specific partial changes, and the terminal sends the instruction to the server, which then regenerates the image based on the instruction and sends the updated image to the terminal.

[1475] Final review and save:

[1476] The user reviews the final image and saves it to local storage or cloud storage if desired.

[1477] Specific examples

[1478] For example, if the user enters the prompt "A bird sitting in a tree under a blue sky," the following prompt sentence is generated:

[1479] Prompt: "In the clear blue sky, a bird perched on a tree, with warm and bright colors indicating happiness."

[1480] Based on this prompt, the generative AI model generates an image that reflects the user's emotional state and sends it to the user's device via the server. The user can then review the image, edit and change the style as needed, and finally save it.

[1481] The system allows users to quickly and intuitively generate and edit personalized, high-quality images that reflect their emotional state. The introduction of an emotion engine further improves the user experience and makes the image generation and editing process more effective.

[1482] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1483] Program processing flow

[1484] Step 1: User enters and submits text

[1485] Specific actions

[1486] The user enters text instructions into the terminal's GUI.

[1487] Example: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[1488] The user clicks the submit button.

[1489] The terminal sends the entered text instructions to the server.

[1490] Input: The text instruction entered by the user

[1491] Output: HTTP request sent to the server

[1492] Step 2: Parsing text on the server

[1493] Specific actions

[1494] The server passes the received text instructions to a natural language processing (NLP) module.

[1495] The NLP module analyzes the text and extracts important keywords.

[1496] Examples: "sunset," "sky," "mountain," "peak," "wolf," "silhouette"

[1497] Input: Received text instructions

[1498] Output: Extracted keywords

[1499] Step 3: Sentiment Analysis

[1500] Specific actions

[1501] The server sends the text instructions to the emotion engine.

[1502] An emotion engine analyzes the user's emotional state from the text.

[1503] Example: Detecting the emotion "happy"

[1504] Input: Text instructions

[1505] Output: Emotion data

[1506] Step 4: Generate prompts

[1507] Specific actions

[1508] The server generates a prompt sentence based on the analyzed keywords and emotion data.

[1509] Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors"

[1510] Input: Keywords, emotion data

[1511] Output: prompt statement

[1512] Step 5: Image generation

[1513] Specific actions

[1514] The server sends the generated prompt sentence to the generating artificial intelligence means (generating AI model).

[1515] Example: DALL-E

[1516] The generative AI model generates high-resolution images and sends them back to the server.

[1517] Input: prompt statement

[1518] Output: Generated image data

[1519] Step 6: Receiving and displaying images

[1520] Specific actions

[1521] The server transmits the generated image data to the terminal.

[1522] The terminal receives the image and displays it to the user.

[1523] Input: Generated image data

[1524] Output: Image displayed on the terminal

[1525] Step 7: Edit the image

[1526] Specific actions

[1527] The user edits the image to adjust the finer details.

[1528] Example: Move the wolf to the center

[1529] The device provides editing tools, allowing users to select parts of the image and adjust their position by dragging and dropping.

[1530] Input: Generated image

[1531] Output: Edited image data

[1532] Step 8: Applying Styles

[1533] Specific actions

[1534] The user selects a specific style from the style change options.

[1535] Example: "Oil painting style"

[1536] The terminal transmits style information to the server.

[1537] The server applies the style to the image and sends it back to the device.

[1538] Input: Style information

[1539] Output: Image data with style applied

[1540] Step 9: Partial Image Regeneration

[1541] Specific actions

[1542] The user inputs instructions for partial modification.

[1543] Example: "Make the edges of the mountain peaks look sharper."

[1544] The terminal transmits a partial change instruction to the server.

[1545] The server regenerates a portion of the image based on the instruction and transmits the updated image to the terminal.

[1546] Input: Partial change instruction

[1547] Output: Partially regenerated image

[1548] Step 10: Final review and save

[1549] Specific actions

[1550] The user reviews the final image.

[1551] The user clicks the Save button.

[1552] The device stores the image data in local storage or cloud storage.

[1553] Input: Last seen image

[1554] Output: Saved image data

[1555] This process flow allows users to quickly generate high-quality, personalized images according to their emotional state, and then intuitively edit and save them.

[1556] (Application example 2)

[1557] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1558] In conventional advertising production, it was difficult to customize according to the user's emotional state, making it difficult to generate intuitive and effective advertising images. It was also not easy to reflect the style and editing instructions specified in the early stages of ad production and apply them to the image immediately. As a result, the effectiveness of the advertisement was not maximized, and it was not possible to achieve advertising expressions that were in tune with the emotions of the target user.

[1559] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generation means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating a specified part of the image, emotion analysis means for recognizing the user's emotional state and generating and editing an image based on the emotion, and means for enabling application in the field of advertising production. This enables the generation and immediate editing of personalized advertising images according to the user's emotions, thereby realizing the provision of optimal advertisements to target users.

[1560] "User" means a person who uses the Interactive AI Image Studio to generate and edit images.

[1561] A "text instruction" is a sentence that describes specific details about the image that the user wants to generate.

[1562] "Natural language processing means" refers to means for analyzing input text instructions and extracting important keywords and context.

[1563] "Generative artificial intelligence means" refers to artificial intelligence techniques for generating high resolution customizable images based on analyzed instructions.

[1564] The "editing means" is a function that allows the user to adjust the fine details of the generated image.

[1565] A "style" is a particular image presentation or visual effect specified by the user.

[1566] The "regeneration means" is a function for regenerating a specified part of an image.

[1567] The "emotion analysis means" is a function for recognizing the user's emotional state and generating and editing images based on that emotion.

[1568] The "advertising production field" is the field of creating images and content to effectively promote products and services to target users.

[1569] The present invention is a system that utilizes an interactive AI image studio to generate high-resolution, customizable advertising images based on user-entered text instructions. The main components of the system include a server, a terminal, and a user.

[1570] Server Configuration

[1571] The server includes the following means:

[1572] Text receiving means: A means for receiving a text instruction input by a user from a terminal.

[1573] Natural language processing means: A means for analyzing input text instructions, extracting keywords, and generating optimal prompts for the artificial intelligence generator. Specifically, a natural language processing engine such as SpaCy is used.

[1574] Generative AI means: A means for generating high-resolution, customizable images based on analyzed instructions, specifically using a generative AI model (e.g., DALL-E).

[1575] Image transmission means: A means for transmitting the generated image to the user's terminal.

[1576] Emotion analysis means: A means of recognizing a user's emotional state by analyzing the user's input text, interactions, voice input, and facial expression data. The IBM Watson Emotion Analysis API can be used.

[1577] Application in the field of advertising production: A means to make it possible to use the generated images and emotion analysis results in advertising production.

[1578] Device configuration

[1579] The terminal includes:

[1580] GUI editing means: A graphical user interface function that allows users to adjust the fine details of the generated image. This allows users to intuitively adjust the image.

[1581] Style application: A function for applying a user-specified style to an image. Users can select styles such as "oil painting" or "vintage."

[1582] Partial regeneration means: A function for regenerating a specified part of an image.

[1583] User operations

[1584] The user performs the following steps:

[1585] 1. Enter the basic concept of the advertisement, such as "people relaxing at the beach," into the GUI of a smartphone or other device.

[1586] 2. The text instructions are sent to the server, where the server's natural language processing means analyzes the text and extracts keywords.

[1587] 3. The emotion analyzer recognizes the user's emotional state and generates appropriate prompts based on that data.

[1588] 4. The generating artificial intelligence means (such as DALL-E) generates a high-resolution image based on the prompt and sends it to the terminal via the server.

[1589] 5. The user can review the received image on their device, make minor adjustments or apply styles as needed, and even request partial regeneration.

[1590] Examples of concrete examples and prompts

[1591] As a concrete example, consider a situation where a user enters the text "A park scene with a happy family together" for an advertising image, and the emotion is "joy." An example prompt for this would be:

[1592] Scene of happy family spending time in the park, joy

[1593] In this way, users can intuitively create and edit high-quality advertising images in a short amount of time. This system enables advertising expressions that are in tune with the emotions of target users, maximizing their effectiveness.

[1594] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1595] Step 1:

[1596] The user inputs the basic idea of ​​the advertisement into the terminal's GUI (text instructions).

[1597] Input: A textual indication of user input (e.g., "People relaxing at the beach").

[1598] Output: A request to send input text instructions.

[1599] Specific operation: The user enters instructions into the terminal's GUI and presses the "Send" button, which sends a request to the server.

[1600] Step 2:

[1601] The server analyzes the received text instructions using natural language processing means.

[1602] Input: The text instruction submitted by the user.

[1603] Output: Extracted keywords and analysis results.

[1604] What happens: The server's natural language processing engine (e.g., SpaCy) analyzes the text and extracts important keywords and contextual information.

[1605] Step 3:

[1606] An emotion analyzer recognizes the user's emotional state.

[1607] Input: Text input, on-device interaction data, voice and facial expression data.

[1608] Output: Perceived emotional state (e.g., "joy").

[1609] What happens: The server analyzes the data using an emotion analysis engine (e.g., IBM Watson Emotion Analysis API) to determine the user's emotional state.

[1610] Step 4:

[1611] The server generates a prompt based on the analysis result and sends it to the generating artificial intelligence means.

[1612] Input: Keywords, contextual information, emotional state.

[1613] Output: The prompt statement.

[1614] Specific operation: The server integrates the extracted keywords and emotion data to generate a prompt sentence to send to the generative AI model (e.g., DALL-E).

[1615] Step 5:

[1616] A generating artificial intelligence means generates a high resolution image.

[1617] Input: Prompt sentence (e.g., "A park scene of a happy family spending time together, joy").

[1618] Output: High resolution customizable images.

[1619] What it does: A prompt is sent to the generative AI model, which then generates an image based on it.

[1620] Step 6:

[1621] The server transmits the generated image to the terminal.

[1622] Input: The generated image.

[1623] Output: Request to send image data.

[1624] Specific operation: The server receives the generated image and sends it to the user's device.

[1625] Step 7:

[1626] The user checks the image on the terminal and edits it as necessary.

[1627] Input: Received image data.

[1628] Output: Editing instructions and adjusted image data.

[1629] Specific behavior: The user views the image in the device's GUI and uses drag-and-drop and tools to make fine adjustments and apply styles.

[1630] Step 8:

[1631] The user requests a partial regeneration.

[1632] Input: Partial regeneration instructions.

[1633] Output: The regenerated subimage.

[1634] Specific operation: The user selects a specific image part, inputs a regeneration instruction, and sends it to the server. The server generates a prompt again, requests the generative AI model to retrieve the regenerated image part, and sends it to the user.

[1635] Step 9:

[1636] The user reviews and saves the final image.

[1637] Input: The final image data after adjustments and regeneration.

[1638] Output: Saved image file.

[1639] Specific behavior: The user confirms the final image and saves it to the device or server.

[1640] The above are the specific processing steps of the system that realizes the application example.

[1641] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1642] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1643] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1644] [Fourth embodiment]

[1645] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1646] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1647] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1648] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1649] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1651] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1652] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1653] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1654] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1655] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1656] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1657] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1658] DETAILED DESCRIPTION OF THE INVENTION The present invention is an image creation and editing system that allows users to input text instructions to create high-resolution, customizable images, and then fine-tune and style them through an intuitive graphical user interface (GUI).

[1659] System Overview

[1660] This system consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, image generation, image editing, and transmission. The terminal provides a GUI and an interface for users to generate and edit images. Users input text and check and edit images through the terminal.

[1661] Program processing flow

[1662] 1. Initial text input and submission

[1663] The user enters specific text instructions into the device's GUI using an input field, for example, "A silhouette of a wolf standing on a mountain top against a sunset sky."

[1664] The terminal sends this text instruction to the server.

[1665] 2. Text analysis and prompt generation

[1666] The server analyzes the received text, specifically using a natural language processing (NLP) engine to extract important keywords ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1667] Based on the extracted keywords, a generative artificial intelligence (AI) model generates suitable prompts.

[1668] 3. Image Generation

[1669] The server sends the generated prompt to a generative artificial intelligence model, such as a deep learning image generation model (e.g., DALL-E).

[1670] An image generation model generates a high-resolution image based on the prompt and returns it to the server.

[1671] 4. Image transmission and initial display

[1672] The server transmits the generated image to the terminal.

[1673] The terminal displays the received image to the user.

[1674] Specific editing and style application

[1675] 1. Editing images

[1676] The user uses the device's GUI to adjust the details of the image, for example, to move the wolf to the center, by selecting an image part and dragging and dropping it to adjust its position.

[1677] The terminal sends these editing instructions to the server, which then regenerates the image based on them.

[1678] 2. Applying Styles

[1679] The user selects a particular style, such as "oil painting," from the style change options.

[1680] The terminal transmits the selected style information to the server.

[1681] The server processes the image to apply the selected style and sends it back to the terminal.

[1682] Partial changes and final confirmation

[1683] 1. Partial Image Regeneration

[1684] The user inputs specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper."

[1685] The terminal transmits a partial change instruction to the server.

[1686] The server performs partial image regeneration according to the instructions and transmits the updated image to the terminal.

[1687] 2. Final confirmation and saving

[1688] The user reviews the final image.

[1689] The user can save the final image to their device or to a server as desired.

[1690] This allows users to easily create and edit high-quality images without specialized knowledge. The system offers high customizability to meet diverse user needs, enabling efficient image creation. The intuitive GUI and advanced AI technology also enhance the user experience.

[1691] The processing flow will be explained below.

[1692] Step 1:

[1693] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided by the terminal GUI.

[1694] Step 2:

[1695] The terminal receives the user's text instructions and transmits the data to the server.

[1696] Step 3:

[1697] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1698] Step 4:

[1699] The server generates an optimal prompt for the artificial intelligence generating means based on the extracted keywords.

[1700] Step 5:

[1701] The server sends the generated prompt to the generative artificial intelligence model and makes a request to generate an image based on the prompt.

[1702] Step 6:

[1703] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[1704] Step 7:

[1705] The server receives the image data and transmits it to the user's terminal.

[1706] Step 8:

[1707] The device displays the received image to the user, who can then use the device's editing tools to review the image and make minor adjustments as needed.

[1708] Step 9:

[1709] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[1710] Step 10:

[1711] The terminal transmits the user's editing instructions to the server.

[1712] Step 11:

[1713] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[1714] Step 12:

[1715] The server transmits the regenerated image data to the user's terminal.

[1716] Step 13:

[1717] The user selects a style such as "oil painting" from the device's toolbar.

[1718] Step 14:

[1719] The terminal transmits the user's style selection instruction to the server.

[1720] Step 15:

[1721] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[1722] Step 16:

[1723] The server transmits the image data after applying the style to the user's terminal.

[1724] Step 17:

[1725] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[1726] Step 18:

[1727] The terminal transmits a partial change instruction to the server.

[1728] Step 19:

[1729] The server regenerates the image of the relevant part based on the received partial change instruction.

[1730] Step 20:

[1731] The server transmits the regenerated partially modified image to the user's terminal.

[1732] Step 21:

[1733] The user checks the final image and saves it to the device or to the server as needed.

[1734] Example 1

[1735] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1736] Conventional image generation and editing systems have the drawback of making it difficult for users without specialized knowledge to generate the desired high-quality, customizable images. It is also often difficult for users to intuitively edit the fine details of an image or apply a specific style. Furthermore, partial image regeneration is not easily possible, making it difficult to flexibly edit images according to user needs.

[1737] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1738] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, means for generating an appropriate prompt sentence based on the analyzed text instructions, artificial intelligence generating means for generating a high-resolution, customizable image using the generated prompt sentence, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating the specified image portion. This allows users to intuitively generate and edit high-quality images without specialized knowledge. Furthermore, it is easy to apply specific styles and regenerate partial images, enabling flexible image editing according to user requests.

[1739] The "means for receiving text instructions entered by a user" is a function for transmitting specific text instructions entered by a user into an input field of a terminal to a server.

[1740] The "natural language processing means for analyzing input text instructions" is a natural language processing engine for analyzing the text entered by the user and extracting important keywords and phrases.

[1741] The "means for generating appropriate prompt sentences" is a function that generates prompt sentences in a format that is easy for the image generation model to understand, based on extracted keywords and phrases.

[1742] The "generative artificial intelligence means for generating high-resolution customizable images" is an artificial intelligence model for generating high-resolution and customizable images based on generated prompt text.

[1743] The "means for transmitting the generated image to the user's device" is a function for transmitting the generated image data from the server to the user's terminal.

[1744] The "editing means for adjusting fine details of an image" is an interface that provides a function for a user to intuitively edit fine details of a generated image.

[1745] The "means for applying a style designated by the user to an image" is a function for applying a style selected by the user (for example, "oil painting style") to a generated image.

[1746] The "means for regenerating a specified portion of an image" is a function for partially regenerating an image when a user instructs a change to a specific portion.

[1747] The image generation and editing system of this invention consists of three main components: a server, a terminal, and a user. The server is responsible for text analysis, prompt generation, and image generation and editing. The terminal provides an intuitive graphical user interface (GUI) for users to generate and edit images. Users input text instructions through the terminal and view and edit the generated images.

[1748] Text analysis and prompt generation

[1749] When the server receives the text instructions entered by the user, it analyzes the text using a natural language processing (NLP) engine. Examples of NLP engines that can be used include spaCy and NLTK. This engine extracts important keywords and phrases (e.g., "sunset," "sky," "mountain," "peak," and "silhouette of a wolf") from the text. Based on the extracted keywords, the server generates a prompt suitable for the generative AI model. This prompt is in a format that is easy for the generative AI model to understand, and can be a specific instruction such as "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1750] Image generation

[1751] The server sends the generated prompt text to a generative AI model. The generative AI model uses an image generation model such as DALL-E, which uses deep learning. This model generates a high-resolution, customizable image based on the prompt text. The generated image is returned to the server, which then sends it to the user's device. The image data sent is encoded, for example, in Base64 format.

[1752] View and edit images

[1753] The user checks the generated image on the device and edits it as necessary. The device's GUI provides an editing tool that allows the user to intuitively adjust minute details of the image. For example, the user can select an image part and drag and drop it to move the wolf to the center. The device sends the editing instructions to the server, and the server regenerates the image and sends the updated image back to the device.

[1754] Style application and partial regeneration

[1755] A user can select a specific style (e.g., "oil painting") from the GUI of the device and apply it to an image. When the device sends the style information to the server, the server applies the style to the image using a style transfer model. The regenerated image with the style applied is sent to the device. Furthermore, the user can input specific instructions for partial changes, such as "make the edges of the mountain peaks look sharper." The server regenerates parts of the image based on the instructions and sends the updated image to the device.

[1756] An example of a prompt is "The silhouette of a wolf standing on top of a mountain against a sunset sky." After parsing this, the prompt becomes "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1757] This allows users to intuitively create and edit high-quality images without specialized knowledge. The system combines advanced AI technology with an intuitive GUI to create easily customizable images.

[1758] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1759] Step 1:

[1760] Enter and send text

[1761] The user enters specific instructions as text into an input field on the device. For example, they might enter "A silhouette of a wolf standing on a mountain top against a sunset sky." This input text is sent to the server. The device uses an HTTP request to send the text data to the server in JSON format.

[1762] Input: User's text instructions

[1763] Output: Sends text instructions to the server

[1764] Step 2:

[1765] Text analysis and keyword extraction

[1766] The server parses the received text instructions. A natural language processing (NLP) engine (e.g., spaCy or NLTK) tokenizes the text and extracts important keywords and phrases. Extracted keywords include "sunset," "sky," "mountain," "peak," and "wolf silhouette."

[1767] Input: Received text instructions

[1768] Output: Extracted keywords and phrases

[1769] Step 3:

[1770] Generate prompt statement

[1771] The server generates a prompt sentence appropriate for the generative AI model based on the extracted keywords and phrases. For example, the generated prompt sentence might be "Generate an image with a silhouette of a wolf standing on the mountain top against a sunset sky."

[1772] Input: Extracted keywords or phrases

[1773] Output: Generated prompt statement

[1774] Step 4:

[1775] Sending an image generation request

[1776] The server sends the generated prompt text to a generative AI model, which then uses deep learning to generate an image based on the text (e.g., DALL-E).

[1777] Input: Generated prompt text

[1778] Output: The generated image

[1779] Step 5:

[1780] Sending the generated image

[1781] The server receives the generated image and sends the image data to the device. The image data is encoded in Base64 format. The device receives the image data in JSON format using an HTTP response.

[1782] Input: Generated image

[1783] Output: Sending image data to the device

[1784] Step 6:

[1785] Displaying images

[1786] The terminal decodes the received image data and displays it to the user. Specifically, the image after Base64 decoding is displayed as HTML. Display it with tags etc.

[1787] Input: Received image data

[1788] Output: Displaying the image to the user

[1789] Step 7:

[1790] Editing images

[1791] The user uses the GUI on the device to adjust small details of the image. For example, to move the wolf to the center, the user selects an image part and adjusts its position by dragging and dropping. The device then sends the edit instructions to the server, including the changes and new coordinates.

[1792] Input: User editing instructions

[1793] Output: Sending edit instructions to the server

[1794] Step 8:

[1795] Regenerate the image

[1796] The server regenerates the image based on the received editing instructions, and the regenerated image data is sent from the server to the terminal and displayed again.

[1797] Input: User editing instructions

[1798] Output: Regenerated image

[1799] Step 9:

[1800] Applying Styles

[1801] The user selects a specific style (e.g., "oil painting") from the device's GUI. The device sends a style application instruction to the server, which then applies the style to the image using the style transfer model. The image after the style application is sent from the server to the device.

[1802] Input: User style selection

[1803] Output: Image with style applied

[1804] Step 10:

[1805] Regenerating Partial Changes

[1806] The user inputs instructions to change a specific part (e.g., "Make the edges of the mountain peaks look sharper"). The device sends the instructions to the server, which then regenerates the image based on the instructions. The updated image is sent to the device and presented to the user.

[1807] Input: User's partial change instructions

[1808] Output: Partially regenerated image

[1809] This allows the system to combine advanced AI technology with an intuitive GUI, providing an environment in which users can easily generate and edit high-quality images.

[1810] (Application example 1)

[1811] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1812] Conventional advertising image generation technologies are difficult for users without specialized knowledge to operate and have limited customization options, making it difficult to efficiently create high-quality advertising images. Furthermore, the process of fine-tuning and applying styles is not intuitive, making it difficult to improve the user experience.

[1813] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1814] In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generating means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating the specified image parts, means for generating and editing an advertising image based on the text instructions, and means for editing fine details of the generated image by drag and drop, thereby enabling users to intuitively and efficiently generate and edit high-quality advertising images without specialized knowledge.

[1815] A "user" is an individual or organization that has the ability to create and edit advertising images using the image generation system.

[1816] "Text instructions" refer to specific words or keywords that a user inputs for generating or editing an advertisement image.

[1817] "Natural language processing means" is a technology for analyzing input text instructions and extracting keywords and generating prompts based on those instructions.

[1818] "Generative artificial intelligence means" refers to AI technology for generating high-resolution, customizable images based on analyzed text instructions.

[1819] "High-resolution customizable images" refers to high-quality images that can be extensively edited and styled according to user instructions and preferences.

[1820] A "server" is a computing device that receives text instructions, parses them, generates images, and transmits them to a user's device.

[1821] "User's device" refers to a device (e.g., smartphone, tablet, or PC) that a user uses to access and operate the image generation system.

[1822] "Editing tools" refers to functionality that allows users to adjust the fine details of an image on their device.

[1823] "Means for applying a style" refers to a technique for reflecting a user-specified style (e.g., classic, modern, etc.) in the generated image.

[1824] The "means for regenerating a portion of an image" is a technique for regenerating an image based on additional instructions from the user for a specific portion.

[1825] "Drag-and-drop editing" refers to a feature that allows users to intuitively move and modify specific parts of an image by dragging and dropping.

[1826] MODE FOR CARRYING OUT THE INVENTION

[1827] An embodiment of the present invention will be described below: This system uses a smartphone as a main terminal as an application that can easily generate and edit advertising images.

[1828] Overall system configuration

[1829] The system consists of the following components:

[1830] User device: Smartphone (iOS or Android)

[1831] Server: Responsible for receiving text instructions, analyzing them, generating images, and sending them

[1832] Generative AI methods: AI techniques that generate high-resolution, customizable images

[1833] Program processing

[1834] Initial text entry and submission

[1835] A user uses their smartphone to enter specific advertising instructions as text, such as "Summer sale, 50% off, blue background, sun icon," and the smartphone app sends the text instructions to a cloud server.

[1836] Natural Language Processing and Prompt Generation

[1837] The server analyzes the received text instructions using natural language processing tools (e.g., SpaCy, NLTK). Specifically, it extracts important keywords from the text (e.g., "summer sale," "50% off," "blue background," "sun icon") and generates a prompt suitable for a generative AI model such as GPT-4 based on these keywords.

[1838] Image generation

[1839] Using the generated prompts, the server sends instructions to a generating artificial intelligence means (e.g., DALL-E), which generates a high-resolution advertising image based on the prompts and returns it to the server, which then sends the generated image to the user's smartphone.

[1840] Editing and styling

[1841] Editing images

[1842] The smartphone app provides users with an intuitive graphical user interface (GUI) that allows them to fine-tune the image. Users can select specific parts of the image and change their position and size by dragging and dropping. Editing instructions are then sent back to the server, which then regenerates the image based on the user's instructions.

[1843] Applying Styles

[1844] Users select a style, such as "Classic" or "Modern," from the app's style change options. This information is also sent to the server, which then applies the style to the image using generative artificial intelligence methods. The resulting regenerated image is sent to the smartphone.

[1845] Partial image regeneration and final confirmation

[1846] Partial changes

[1847] When a user requests a specific change, for example, "increase the icon size," the request is also sent to the cloud server, which then regenerates the image of the corresponding part.

[1848] Final confirmation and saving

[1849] Users can view the final ad image on their smartphone and save it or share it on social media or cloud storage.

[1850] Example prompt

[1851] "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background."

[1852] This allows even non-expert users to efficiently and intuitively create and edit high-resolution, customizable advertising images. The combination of generative artificial intelligence and natural language processing technology enables high levels of customization and an improved user experience.

[1853] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1854] Step 1:

[1855] Entering and sending text instructions

[1856] The user opens the app on their smartphone and inputs the text instructions needed to generate the ad image. For example, they might input "Summer sale, 50% off, blue background, sun icon." The device then sends this input text to the cloud server. The input in this step is the user's text instructions, and the output is the data sent to the cloud server.

[1857] Step 2:

[1858] Text analysis and keyword extraction

[1859] The cloud server analyzes the received text instructions using a natural language processing engine (e.g., SpaCy, NLTK). Specifically, it tokenizes the text and extracts important keywords ("summer sale," "50% off," "blue background," "sun icon"). The input of this step is the text instruction sent by the user, and the output is the analyzed keywords.

[1860] Step 3:

[1861] Prompt Generation

[1862] The server generates a prompt appropriate for the generative AI model (e.g., GPT-4) based on the extracted keywords. This prompt will be a sentence such as "Create an advertisement image for a summer sale with a 50% off text, and a sun icon on a blue background." The input of this step is the parsed keywords, and the output is the generated prompt sentence.

[1863] Step 4:

[1864] Image generation

[1865] The server sends the generated prompt text to an image generation model (e.g., DALL-E). DALL-E generates a high-resolution advertising image based on the prompt and returns it to the server. The input of this step is the prompt text, and the output is a high-resolution advertising image.

[1866] Step 5:

[1867] Image transmission and initial display

[1868] The server sends the generated high-resolution advertising image to the smartphone. The device displays the received image to the user. The input of this step is the generated advertising image, and the output is the display data for the smartphone.

[1869] Step 6:

[1870] Editing images

[1871] The user uses the smartphone's GUI to adjust the details of the image, for example, by dragging and dropping the sun icon to the center. The device then sends these editing instructions to the cloud server. The input of this step is the user's editing instructions, and the output is the data sent to the cloud server.

[1872] Step 7:

[1873] Regenerate Edits

[1874] The server regenerates the image based on the received editing instructions. The regenerated image is sent to the smartphone. The input of this step is the user's editing instructions, and the output is the regenerated image.

[1875] Step 8:

[1876] Applying Styles

[1877] The user selects a style, such as "Classic" or "Modern." The device sends the style change information to the server. The server then uses a generative artificial intelligence method to apply the specified style to the image. The input to this step is the user-specified style, and the output is the image with the style applied.

[1878] Step 9:

[1879] Partial Image Regeneration

[1880] The user instructs the server to change a specific part (e.g., "increase the icon size"). The device sends the instruction to change the part to the server. The server regenerates the specific part according to the instruction and resends the image to the smartphone. The input of this step is the instruction to change the part, and the output is the regenerated image.

[1881] Step 10:

[1882] Final confirmation and saving

[1883] The user can then check the final ad image and save it on their smartphone or share it on social media or cloud storage if desired. The input of this step is the final ad image, and the output is the saved or shared image.

[1884] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1885] The interactive AI image studio of the present invention is a system that receives text instructions entered by a user, analyzes them using natural language processing means, and generates high-resolution, customizable images using generative AI means. By combining this system with an emotion engine, it can recognize the user's emotions and generate and edit images based on those emotions.

[1886] System configuration

[1887] This system is mainly composed of three elements: a server, a terminal, and a user. In particular, the addition of an emotion engine enables more personalized image generation according to the user's emotional state.

[1888] Program processing flow

[1889] 1. Initial text input and submission

[1890] The user inputs specific instructions into the device's GUI as text, such as "The silhouette of a wolf standing on top of a mountain with a sunset sky in the background."

[1891] The terminal sends this text instruction to the server.

[1892] 2. Text analysis and prompt generation

[1893] The server analyzes the received text instructions using natural language processing means, specifically by extracting important keywords from the text.

[1894] 3. Emotion analysis

[1895] The emotion engine analyzes the user's input text, on-device interactions, voice, and facial expressions to recognize the user's emotional state.

[1896] For example, if the user is expressing a "happy" emotion, a prompt containing colors and motifs that match that emotion is generated.

[1897] 4. Image Generation

[1898] The server sends prompts to the generating artificial intelligence means to generate high resolution images based on the parsed text instructions and emotion data.

[1899] An image is generated by a generating artificial intelligence means (eg, DALL-E) and sent back to the server.

[1900] 5. Sending and displaying images

[1901] The server transmits the generated image to the terminal.

[1902] The terminal displays the received image to the user.

[1903] Specific editing and style application

[1904] 1. Editing images

[1905] The user can check the image on the device's GUI and make fine adjustments as needed, for example, by selecting an image part and dragging and dropping it to move the wolf to the center.

[1906] 2. Applying Styles

[1907] The user selects a particular style, such as "oil painting," from the style change options.

[1908] The terminal transmits the selected style information to the server, and the server performs processing to apply the selected style to the image.

[1909] Partial changes and emotional readjustments

[1910] 1. Partial Image Regeneration

[1911] The user inputs instructions for partial changes such as "make the edges of the mountain peaks look sharper."

[1912] The terminal transmits a partial change instruction to the server, and the server performs partial image regeneration based on the instruction.

[1913] 2. Emotional Recalibration

[1914] The emotion engine refers to the user's input and editing history and suggests additional adjustments and optimizations based on the user's emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[1915] 3. Final confirmation and saving

[1916] The user checks the final image and saves it to the device or to the server as needed.

[1917] This system allows users to easily create and edit high-quality images that are personalized according to their emotional state. The introduction of an emotion engine further improves the user experience, making the image creation and editing process more intuitive and effective.

[1918] The processing flow will be explained below.

[1919] Specific processing steps

[1920] Step 1:

[1921] The user inputs the text instruction "A silhouette of a wolf standing on top of a mountain against a sunset sky" into an input field provided in the GUI of the terminal.

[1922] Step 2:

[1923] The terminal receives the user's text instructions and transmits the data to the server.

[1924] Step 3:

[1925] The server receives the text instruction from the device and analyzes it using natural language processing means. Specifically, it extracts important keywords from the text ("sunset," "sky," "mountain," "peak," "wolf silhouette").

[1926] Step 4:

[1927] The emotion engine analyzes the user's input text and interactions to recognize the user's emotions. For example, it recognizes the emotion "fun" from the input text and user actions.

[1928] Step 5:

[1929] Based on the keywords and emotion recognition results, the server generates a prompt with details about the image to be generated. For example, if the "happy" emotion is recognized, a prompt using bright colors is generated.

[1930] Step 6:

[1931] The server sends the generated prompt to the generating artificial intelligence means, which makes a request to generate an image based on the prompt.

[1932] Step 7:

[1933] The image generation model receives the prompt, generates a high-resolution image according to the instructions, and sends the generated image data back to the server.

[1934] Step 8:

[1935] The server receives the image data and transmits it to the user's terminal.

[1936] Step 9:

[1937] The device displays the received image to the user, who can review the image and make minor adjustments as needed.

[1938] Step 10:

[1939] For example, if the user wants to move the wolf to the center, he or she can select a part of the image and adjust its position by dragging and dropping.

[1940] Step 11:

[1941] The terminal transmits the user's editing instructions to the server.

[1942] Step 12:

[1943] The server regenerates the image based on the editing instructions received, regenerating only the necessary parts.

[1944] Step 13:

[1945] The server transmits the regenerated image data to the user's terminal.

[1946] Step 14:

[1947] The user selects a style such as "oil painting" from the device's toolbar.

[1948] Step 15:

[1949] The terminal transmits the user's style selection instruction to the server.

[1950] Step 16:

[1951] The server processes the image to apply the selected style, for example, by applying an oil painting filter.

[1952] Step 17:

[1953] The server transmits the image data after applying the style to the user's terminal.

[1954] Step 18:

[1955] The user inputs instructions for modifying a portion of an image, for example, "make the edges of the mountain peaks look sharper."

[1956] Step 19:

[1957] The terminal transmits a partial change instruction to the server.

[1958] Step 20:

[1959] Based on the partial change instruction received by the server, the image of the relevant part is regenerated.

[1960] Step 21:

[1961] The server transmits the regenerated partially modified image to the user's terminal.

[1962] Step 22:

[1963] The emotion engine re-analyzes the user's emotional state based on their input and editing history, and suggests additional adjustments and optimizations based on their emotions. For example, if the user is feeling stressed, it will suggest readjusting the color tone or composition to make it more relaxing.

[1964] Step 23:

[1965] The user checks the final image and saves it to the device or to the server as needed.

[1966] Example 2

[1967] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1968] Conventional image generation systems often cannot predict the results of a generated image, even when a user inputs clear instructions, making it difficult to obtain an image that reflects the user's intentions. Furthermore, they lack the functionality to generate images that reflect the user's emotional state, making it difficult to generate personalized images. Furthermore, the functionality for partially modifying an image or applying styles is limited, which can make it time-consuming and labor-intensive to generate an image that satisfies the user.

[1969] The specification process by the specification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving a text instruction input by a user, natural language processing means for analyzing the input text instruction, artificial intelligence means for generating a high-resolution customizable image based on the analyzed text instruction and the user's emotional data, an emotion engine for analyzing the user's emotional state, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, and means for regenerating a specified portion of the image. This makes it possible to quickly generate high-quality, personalized images that correspond to the user's emotional state and to edit them intuitively and efficiently.

[1970] The "means for receiving user-entered text instructions" is a module for electronically receiving user-entered instructions in text form.

[1971] The "natural language processing means for analyzing input text instructions" is a module that includes techniques and algorithms for analyzing input text instructions and understanding their content.

[1972] A "generative artificial intelligence means for generating high-resolution customizable images" is a module that utilizes artificial intelligence technology to generate high-quality, detailed, customizable images based on specified conditions.

[1973] The "emotion engine that analyzes the user's emotional state" is a module that includes technologies and algorithms for analyzing emotions from the user's input and actions and determining their state.

[1974] The "means for transmitting the generated image to the user's device" is a module for transferring the generated image data to the user's device via a network.

[1975] The "editing tool for adjusting fine details of an image" is a module that provides an interface for users to modify and adjust the details of the generated image.

[1976] The "means for applying a user-specified style to an image" is a module for applying a particular style selected by the user (e.g., oil painting style, watercolor style, etc.) to the generated image.

[1977] The "means for regenerating a specified portion of an image" is a module for regenerating and modifying a specific portion of an image based on a user's instructions.

[1978] MODE FOR CARRYING OUT THE INVENTION

[1979] The interactive AI image studio of the present invention is a system that receives and analyzes user-entered text instructions to generate high-resolution, customizable images. The system includes an emotion engine that recognizes the user's emotional state and generates and edits images accordingly. An embodiment of the system is described in detail below.

[1980] System Components

[1981] This system mainly consists of three elements: a server, a terminal, and a user.

[1982] 1. Server:

[1983] The server has the following functions:

[1984] Means for receiving user-entered text instructions

[1985] Analyzing the text using natural language processing tools (e.g., SpaCy)

[1986] Analysis of the user's emotional state using an emotion engine (e.g., OpenAI's GPT-3)

[1987] Generation of images by artificial intelligence means (e.g., DALL-E)

[1988] A means for transmitting the generated image data to a terminal

[1989] 2. Terminal:

[1990] The device has the following features:

[1991] GUI for user input

[1992] Displaying received images

[1993] Provides tools for editing the finer details of an image

[1994] A means for sending user-specified styles to the server and applying them

[1995] Means for sending partial image regeneration instructions to a server

[1996] 3. User:

[1997] The user performs the following operations:

[1998] Enter text instructions (e.g., "A silhouette of a wolf standing on a mountain peak against a sunset sky")

[1999] View and edit the generated image

[2000] Selecting and applying a specific style

[2001] Partial changes, final confirmation and saving

[2002] System Operation

[2003] Using this system, users can easily perform the following process:

[2004] Initial input and submission:

[2005] The user inputs text into the terminal and sends it to the server. Example input: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[2006] Natural Language Processing and Sentiment Analysis:

[2007] The server analyzes the received text instructions using natural language processing to extract important keywords. The emotion engine then analyzes the user's emotional state from the text. For example, if the emotion "fun" is detected, the server generates a prompt that matches that emotion.

[2008] Image generation:

[2009] The server generates a high-resolution image using a generative AI model (e.g., DALL-E) based on the prompt and emotion data. Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors."

[2010] View and edit images:

[2011] The server sends the generated image to the terminal, where the user can view it. The user can then edit the image in detail on the terminal, adjusting the position and color tone as necessary.

[2012] Applying styles:

[2013] The user selects a style change option (e.g., "oil painting"), and the device sends that information to the server, which applies the selected style to the image and sends it back to the device.

[2014] Modify and regenerate:

[2015] The user instructs the terminal to make specific partial changes, and the terminal sends the instruction to the server, which then regenerates the image based on the instruction and sends the updated image to the terminal.

[2016] Final review and save:

[2017] The user reviews the final image and saves it to local storage or cloud storage if desired.

[2018] Specific examples

[2019] For example, if the user enters the prompt "A bird sitting in a tree under a blue sky," the following prompt sentence is generated:

[2020] Prompt: "In the clear blue sky, a bird perched on a tree, with warm and bright colors indicating happiness."

[2021] Based on this prompt, the generative AI model generates an image that reflects the user's emotional state and sends it to the user's device via the server. The user can then review the image, edit and change the style as needed, and finally save it.

[2022] The system allows users to quickly and intuitively generate and edit personalized, high-quality images that reflect their emotional state. The introduction of an emotion engine further improves the user experience and makes the image generation and editing process more effective.

[2023] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2024] Program processing flow

[2025] Step 1: User enters and submits text

[2026] Specific actions

[2027] The user enters text instructions into the terminal's GUI.

[2028] Example: "A silhouette of a wolf standing on top of a mountain against a sunset sky."

[2029] The user clicks the submit button.

[2030] The terminal sends the entered text instructions to the server.

[2031] Input: The text instruction entered by the user

[2032] Output: HTTP request sent to the server

[2033] Step 2: Parsing text on the server

[2034] Specific actions

[2035] The server passes the received text instructions to a natural language processing (NLP) module.

[2036] The NLP module analyzes the text and extracts important keywords.

[2037] Examples: "sunset," "sky," "mountain," "peak," "wolf," "silhouette"

[2038] Input: Received text instructions

[2039] Output: Extracted keywords

[2040] Step 3: Sentiment Analysis

[2041] Specific actions

[2042] The server sends the text instructions to the emotion engine.

[2043] An emotion engine analyzes the user's emotional state from the text.

[2044] Example: Detecting the emotion "happy"

[2045] Input: Text instructions

[2046] Output: Emotion data

[2047] Step 4: Generate prompts

[2048] Specific actions

[2049] The server generates a prompt sentence based on the analyzed keywords and emotion data.

[2050] Example: "In the sunset sky, a silhouette of a wolf standing on the mountain peak, reflecting a joyful emotion with bright colors"

[2051] Input: Keywords, emotion data

[2052] Output: prompt statement

[2053] Step 5: Image generation

[2054] Specific actions

[2055] The server sends the generated prompt sentence to the generating artificial intelligence means (generating AI model).

[2056] Example: DALL-E

[2057] The generative AI model generates high-resolution images and sends them back to the server.

[2058] Input: prompt statement

[2059] Output: Generated image data

[2060] Step 6: Receiving and displaying images

[2061] Specific actions

[2062] The server transmits the generated image data to the terminal.

[2063] The terminal receives the image and displays it to the user.

[2064] Input: Generated image data

[2065] Output: Image displayed on the terminal

[2066] Step 7: Edit the image

[2067] Specific actions

[2068] The user edits the image to adjust the finer details.

[2069] Example: Move the wolf to the center

[2070] The device provides editing tools, allowing users to select parts of the image and adjust their position by dragging and dropping.

[2071] Input: Generated image

[2072] Output: Edited image data

[2073] Step 8: Applying Styles

[2074] Specific actions

[2075] The user selects a specific style from the style change options.

[2076] Example: "Oil painting style"

[2077] The terminal transmits style information to the server.

[2078] The server applies the style to the image and sends it back to the device.

[2079] Input: Style information

[2080] Output: Image data with style applied

[2081] Step 9: Partial Image Regeneration

[2082] Specific actions

[2083] The user inputs instructions for partial modification.

[2084] Example: "Make the edges of the mountain peaks look sharper."

[2085] The terminal transmits a partial change instruction to the server.

[2086] The server regenerates a portion of the image based on the instruction and transmits the updated image to the terminal.

[2087] Input: Partial change instruction

[2088] Output: Partially regenerated image

[2089] Step 10: Final review and save

[2090] Specific actions

[2091] The user reviews the final image.

[2092] The user clicks the Save button.

[2093] The device stores the image data in local storage or cloud storage.

[2094] Input: Last seen image

[2095] Output: Saved image data

[2096] This process flow allows users to quickly generate high-quality, personalized images according to their emotional state, and then intuitively edit and save them.

[2097] (Application example 2)

[2098] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2099] In conventional advertising production, it was difficult to customize according to the user's emotional state, making it difficult to generate intuitive and effective advertising images. It was also not easy to reflect the style and editing instructions specified in the early stages of ad production and apply them to the image immediately. As a result, the effectiveness of the advertisement was not maximized, and it was not possible to achieve advertising expressions that were in tune with the emotions of the target user.

[2100] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text instructions input by a user, natural language processing means for analyzing the input text instructions, artificial intelligence generation means for generating a high-resolution customizable image based on the analyzed text instructions, means for transmitting the generated image to the user's device, editing means for adjusting fine details of the image on the user's device, means for applying a style specified by the user to the image, means for regenerating a specified part of the image, emotion analysis means for recognizing the user's emotional state and generating and editing an image based on the emotion, and means for enabling application in the field of advertising production. This enables the generation and immediate editing of personalized advertising images according to the user's emotions, thereby realizing the provision of optimal advertisements to target users.

[2101] "User" means a person who uses the Interactive AI Image Studio to generate and edit images.

[2102] A "text instruction" is a sentence that describes specific details about the image that the user wants to generate.

[2103] "Natural language processing means" refers to means for analyzing input text instructions and extracting important keywords and context.

[2104] "Generative artificial intelligence means" refers to artificial intelligence techniques for generating high resolution customizable images based on analyzed instructions.

[2105] The "editing means" is a function that allows the user to adjust the fine details of the generated image.

[2106] A "style" is a particular image presentation or visual effect specified by the user.

[2107] The "regeneration means" is a function for regenerating a specified part of an image.

[2108] The "emotion analysis means" is a function for recognizing the user's emotional state and generating and editing images based on that emotion.

[2109] The "advertising production field" is the field of creating images and content to effectively promote products and services to target users.

[2110] The present invention is a system that utilizes an interactive AI image studio to generate high-resolution, customizable advertising images based on user-entered text instructions. The main components of the system include a server, a terminal, and a user.

[2111] Server Configuration

[2112] The server includes the following means:

[2113] Text receiving means: A means for receiving a text instruction input by a user from a terminal.

[2114] Natural language processing means: A means for analyzing input text instructions, extracting keywords, and generating optimal prompts for the artificial intelligence generator. Specifically, a natural language processing engine such as SpaCy is used.

[2115] Generative AI means: A means for generating high-resolution, customizable images based on analyzed instructions, specifically using a generative AI model (e.g., DALL-E).

[2116] Image transmission means: A means for transmitting the generated image to the user's terminal.

[2117] Emotion analysis means: A means of recognizing a user's emotional state by analyzing the user's input text, interactions, voice input, and facial expression data. The IBM Watson Emotion Analysis API can be used.

[2118] Application in the field of advertising production: A means to make it possible to use the generated images and emotion analysis results in advertising production.

[2119] Device configuration

[2120] The terminal includes:

[2121] GUI editing means: A graphical user interface function that allows users to adjust the fine details of the generated image. This allows users to intuitively adjust the image.

[2122] Style application: A function for applying a user-specified style to an image. Users can select styles such as "oil painting" or "vintage."

[2123] Partial regeneration means: A function for regenerating a specified part of an image.

[2124] User operations

[2125] The user performs the following steps:

[2126] 1. Enter the basic concept of the advertisement, such as "people relaxing at the beach," into the GUI of a smartphone or other device.

[2127] 2. The text instructions are sent to the server, where the server's natural language processing means analyzes the text and extracts keywords.

[2128] 3. The emotion analyzer recognizes the user's emotional state and generates appropriate prompts based on that data.

[2129] 4. The generating artificial intelligence means (such as DALL-E) generates a high-resolution image based on the prompt and sends it to the terminal via the server.

[2130] 5. The user can review the received image on their device, make minor adjustments or apply styles as needed, and even request partial regeneration.

[2131] Examples of concrete examples and prompts

[2132] As a concrete example, consider a situation where a user enters the text "A park scene with a happy family together" for an advertising image, and the emotion is "joy." An example prompt for this would be:

[2133] Scene of happy family spending time in the park, joy

[2134] In this way, users can intuitively create and edit high-quality advertising images in a short amount of time. This system enables advertising expressions that are in tune with the emotions of target users, maximizing their effectiveness.

[2135] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2136] Step 1:

[2137] The user inputs the basic idea of ​​the advertisement into the terminal's GUI (text instructions).

[2138] Input: A textual indication of user input (e.g., "People relaxing at the beach").

[2139] Output: A request to send input text instructions.

[2140] Specific operation: The user enters instructions into the terminal's GUI and presses the "Send" button, which sends a request to the server.

[2141] Step 2:

[2142] The server analyzes the received text instructions using natural language processing means.

[2143] Input: The text instruction submitted by the user.

[2144] Output: Extracted keywords and analysis results.

[2145] What happens: The server's natural language processing engine (e.g., SpaCy) analyzes the text and extracts important keywords and contextual information.

[2146] Step 3:

[2147] An emotion analyzer recognizes the user's emotional state.

[2148] Input: Text input, on-device interaction data, voice and facial expression data.

[2149] Output: Perceived emotional state (e.g., "joy").

[2150] What happens: The server analyzes the data using an emotion analysis engine (e.g., IBM Watson Emotion Analysis API) to determine the user's emotional state.

[2151] Step 4:

[2152] The server generates a prompt based on the analysis result and sends it to the generating artificial intelligence means.

[2153] Input: Keywords, contextual information, emotional state.

[2154] Output: The prompt statement.

[2155] Specific operation: The server integrates the extracted keywords and emotion data to generate a prompt sentence to send to the generative AI model (e.g., DALL-E).

[2156] Step 5:

[2157] A generating artificial intelligence means generates a high resolution image.

[2158] Input: Prompt sentence (e.g., "A park scene of a happy family spending time together, joy").

[2159] Output: High resolution customizable images.

[2160] What it does: A prompt is sent to the generative AI model, which then generates an image based on it.

[2161] Step 6:

[2162] The server transmits the generated image to the terminal.

[2163] Input: The generated image.

[2164] Output: Request to send image data.

[2165] Specific operation: The server receives the generated image and sends it to the user's device.

[2166] Step 7:

[2167] The user checks the image on the terminal and edits it as necessary.

[2168] Input: Received image data.

[2169] Output: Editing instructions and adjusted image data.

[2170] Specific behavior: The user views the image in the device's GUI and uses drag-and-drop and tools to make fine adjustments and apply styles.

[2171] Step 8:

[2172] The user requests a partial regeneration.

[2173] Input: Partial regeneration instructions.

[2174] Output: The regenerated subimage.

[2175] Specific operation: The user selects a specific image part, inputs a regeneration instruction, and sends it to the server. The server generates a prompt again, requests the generative AI model to retrieve the regenerated image part, and sends it to the user.

[2176] Step 9:

[2177] The user reviews and saves the final image.

[2178] Input: The final image data after adjustments and regeneration.

[2179] Output: Saved image file.

[2180] Specific behavior: The user confirms the final image and saves it to the device or server.

[2181] The above are the specific processing steps of the system that realizes the application example.

[2182] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2183] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2184] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2185] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2186] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2187] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2188] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2189] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2190] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2191] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2192] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2193] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2194] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2195] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2196] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2197] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2198] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2199] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2200] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2201] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2202] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2203] The following is further disclosed regarding the above embodiment.

[2204] (Claim 1)

[2205] means for receiving a text instruction entered by a user;

[2206] natural language processing means for analyzing input text instructions;

[2207] a generative artificial intelligence means for generating high resolution customizable images based on the analyzed text instructions;

[2208] means for transmitting the generated image to a user device;

[2209] editing means for adjusting image details on the user's device;

[2210] means for applying a user-specified style to the image;

[2211] A means of regenerating a specified portion of an image.

[2212] A system including:

[2213] (Claim 2)

[2214] 2. The system of claim 1, wherein the natural language processing means for analyzing the text instructions is configured to extract keywords to generate optimal prompts for the generating artificial intelligence means.

[2215] (Claim 3)

[2216] 2. The system of claim 1, wherein the editing means provides a graphical user interface for modifying portions of the image based on user instructions.

[2217] "Example 1"

[2218] (Claim 1)

[2219] means for receiving a text instruction entered by a user;

[2220] natural language processing means for analyzing input text instructions;

[2221] means for generating an appropriate prompt sentence based on the parsed text instructions;

[2222] a generating artificial intelligence means for generating high-resolution customizable images using the generated prompt text;

[2223] means for transmitting the generated image to a user device;

[2224] editing means for adjusting image details on the user's device;

[2225] means for applying a user-specified style to the image;

[2226] A means of regenerating a specified portion of an image.

[2227] A system including:

[2228] (Claim 2)

[2229] 2. The system according to claim 1, wherein the natural language processing means is configured to extract keywords and generate an optimal prompt sentence for the artificial intelligence generating means.

[2230] (Claim 3)

[2231] 2. The system of claim 1, wherein the editing means provides a graphical user interface for modifying portions of the image based on user instructions.

[2232] "Application Example 1"

[2233] (Claim 1)

[2234] means for receiving a text instruction entered by a user;

[2235] natural language processing means for analyzing input text instructions;

[2236] a generative artificial intelligence means for generating high resolution customizable images based on the analyzed text instructions;

[2237] means for transmitting the generated image to a user device;

[2238] editing means for adjusting image details on the user's device;

[2239] means for applying a user-specified style to the image;

[2240] means for regenerating a specified portion of the image;

[2241] means for generating and editing advertising images based on text instructions;

[2242] Drag-and-drop means to fine-tune the generated images

[2243] A system including:

[2244] (Claim 2)

[2245] 2. The system of claim 1, wherein the natural language processing means for analyzing the text instructions is configured to extract keywords to generate optimal prompts for the generating artificial intelligence means.

[2246] (Claim 3)

[2247] 2. The system of claim 1, wherein the editing means provides a graphical user interface for modifying portions of the image based on user instructions.

[2248] "Example 2: Combining Emotion Engines"

[2249] (Claim 1)

[2250] means for receiving a text instruction entered by a user;

[2251] natural language processing means for analyzing input text instructions;

[2252] a generative artificial intelligence means for generating high-resolution customizable images based on the analyzed text instructions and the user's emotional data;

[2253] an emotion engine that analyzes the user's emotional state;

[2254] means for transmitting the generated image to a user device;

[2255] editing means for adjusting image details on the user's device;

[2256] means for applying a user-specified style to the image;

[2257] A means of regenerating a specified portion of an image.

[2258] A system including:

[2259] (Claim 2)

[2260] 2. The system of claim 1, wherein the natural language processing means for analyzing the text instructions is configured to extract keywords to generate optimal prompts for the generating artificial intelligence means.

[2261] (Claim 3)

[2262] 2. The system of claim 1, wherein the editing means provides a graphical user interface for modifying portions of the image based on user instructions.

[2263] "Application example 2 when combining emotion engines"

[2264] (Claim 1)

[2265] means for receiving a text instruction entered by a user;

[2266] natural language processing means for analyzing input text instructions;

[2267] a generative artificial intelligence means for generating high resolution customizable images based on the analyzed text instructions;

[2268] means for transmitting the generated image to a user device;

[2269] editing means for adjusting image details on the user's device;

[2270] means for applying a user-specified style to the image;

[2271] means for regenerating a specified portion of the image;

[2272] emotion analysis means for recognizing the user's emotional state and generating and editing images based on the emotions;

[2273] A means to enable application in the field of advertising production

[2274] A system including:

[2275] (Claim 2)

[2276] 2. The system of claim 1, wherein the natural language processing means for analyzing the text instructions is configured to extract keywords to generate optimal prompts for the generating artificial intelligence means.

[2277] (Claim 3)

[2278] 2. The system of claim 1, wherein the editing means provides a graphical user interface for making partial changes to the image based on user instructions, and further provides additional editing suggestions based on results of the sentiment analysis means. [Explanation of symbols]

[2279] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving a text instruction entered by a user; natural language processing means for analyzing input text instructions; a generative artificial intelligence means for generating high resolution customizable images based on the analyzed text instructions; means for transmitting the generated image to a user device; editing means for adjusting image details on the user's device; means for applying a user-specified style to the image; A means of regenerating a specified portion of an image. A system including:

2. 2. The system of claim 1, wherein the natural language processing means for analyzing the text instructions is configured to extract keywords to generate optimal prompts for the generating artificial intelligence means.

3. 2. The system according to claim 1, wherein the editing means provides a graphical user interface for partially modifying the image based on a user instruction.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A