System

The system addresses the inefficiencies of conventional illustration generation by using AI to analyze handwritten drafts and user-specified tastes, allowing for quick and high-quality illustrations with consistent styles.

JP2026017387APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118169
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Conventional illustration generation systems face difficulties in generating ideal illustrations efficiently, requiring numerous trials and advanced design skills, and struggle to produce high-quality illustrations with consistent styles.

Method used

A system that receives a handwritten draft image and text information specifying taste, analyzes the image, and uses AI to generate illustrations based on the analysis results, allowing resizing and normalization for optimal processing, enabling quick and high-quality illustration creation with consistent styles.

Benefits of technology

Enables users to easily and efficiently obtain high-quality illustrations in desired styles by simplifying the process and improving illustration quality through AI-generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017387000001_ABST
    Figure 2026017387000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system includes a means for receiving a draft image drawn by handwriting, a means for receiving character information designating a taste together with the draft image, a means for generating an illustration on the basis of an analysis result of the draft image and the designated taste, and a means for transmitting the generated illustration to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional illustration generation systems face the problem of making it difficult to generate ideal illustrations using only prompts, resulting in a large number of trials. Another issue is the difficulty of generating multiple patterns with the same style. This consumes a lot of time and effort, making it difficult for users to efficiently obtain the illustrations they desire. Furthermore, there is also the problem that it is difficult to obtain high-quality illustrations if the user does not have advanced design skills. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for receiving a handwritten draft image, a means for receiving text information specifying a taste along with the draft image, a means for generating an illustration based on the analysis results of the draft image and the specified taste, and a means for transmitting the generated illustration to a user terminal. This system allows users to quickly obtain an illustration in a desired style simply by using the draft and specifying the taste. Furthermore, by providing a means for resizing and normalizing the draft image and a means for generating an illustration based on a taste selected from multiple taste options, the quality of the generated illustration can be improved and multiple patterns with the same taste can be easily generated.

[0006] A "hand-drawn draft image" is an image that shows the initial composition or motif of an illustration that a user has manually drawn on paper or a digital device.

[0007] "Text information specifying taste" is text data that describes the style, color tone, and other characteristics of the illustration desired by the user.

[0008] "Draft image analysis results" refers to information about shapes and objects obtained by using algorithms or AI models to analyze hand-drawn draft images.

[0009] "Generating an illustration" means creating an illustration as final image data based on the analysis results of the draft image and the specified taste.

[0010] A "user terminal" is a device that a user operates, and includes a personal computer, a smartphone, and the like.

[0011] "Resizing and normalizing" refers to processing to change the size of the draft image and standardize the data format.

[0012] "Multiple taste options" refers to the various illustration styles and expression methods that the system provides. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0035] Program processing

[0036] User operations

[0037] Users create a rough draft of the illustration's composition and motif on their device (PC or smartphone).

[0038] Upload the completed draft image to the system's web application or dedicated app by clicking the Upload button and selecting the appropriate image file from the file dialog.

[0039] Next, enter the style of illustration you want in the text input field, such as "soft watercolor colors" or "bright manga colors."

[0040] Once you have completed the draft image and taste specification, press the send button to send the data to the server.

[0041] Server Processing

[0042] The server receives the draft image and character information specifying the taste sent by the user.

[0043] The received draft image is analyzed and resized and normalized, optimizing the image resolution and format for AI processing.

[0044] It also maps the text information of the taste specification to predefined taste options, which determine the style and color tone of the generated illustration.

[0045] Processing AI models

[0046] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input.

[0047] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the specified taste. For example, if the draft is a "composition of a boy reading a book" and the taste is specified as "bright, cartoon-style colors," the AI ​​model will generate a cartoon-style illustration of a boy reading a book.

[0048] Server provides results

[0049] The server receives the illustration data returned by the AI ​​model and reformats it into a format that can be viewed by the user.

[0050] The generated illustration data is sent to the user's terminal, and the results are provided to the user.

[0051] User Verification

[0052] The user checks the illustration sent back from the server on the terminal.

[0053] You can save the generated illustrations as needed, and make further adjustments or requests.

[0054] Specific examples

[0055] For example, if a user uploads a "picture of a cat sleeping on a sofa" as a draft and specifies the taste as "warm, hand-drawn," the process will proceed as follows.

[0056] 1. The user simply draws the shape of a cat and a sofa and uploads the image file to the system.

[0057] 2. The user enters the taste as "Warm hand-drawn" in the text input field and clicks the submit button.

[0058] 3. The server receives the draft image and taste information, resizes and normalizes the image, and maps the taste information.

[0059] 4. The pre-processed data is sent to an AI model to generate a hand-drawn illustration.

[0060] 5. The server receives the generated illustration data and sends it to the user.

[0061] 6. The user checks the generated illustration on their device and requests saving or addition as needed.

[0062] In this way, the system allows users to easily generate high-quality illustrations.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] Users use their devices (PC or smartphone) to create a rough draft of the composition and motif of the illustration, which represents a specific scene, such as a cat sleeping on a sofa.

[0066] Step 2:

[0067] The user uploads the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0068] Step 3:

[0069] After uploading is complete, users can enter the desired style of the illustration in a text input field that appears, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0070] Step 4:

[0071] After the user inputs the taste specification, the user clicks the send button to send the draft image and taste information to the server.

[0072] Step 5:

[0073] The server receives the draft image and taste specification character information sent by the user, and stores the received data in a format that can be processed immediately.

[0074] Step 6:

[0075] The server analyzes the received draft images and resizes and normalizes them, converting them into a format that is optimally processed by the AI ​​model.

[0076] Step 7:

[0077] The server analyzes the textual information of the taste specification and maps it to predefined taste options, which are then converted into a format that is easy for the AI ​​model to understand.

[0078] Step 8:

[0079] The server passes the preprocessed draft image data and taste information to the AI ​​model as input.

[0080] Step 9:

[0081] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the user's preferences, such as a scene of a cat sleeping on a sofa in a warm, hand-drawn style.

[0082] Step 10:

[0083] The AI ​​model returns the generated illustration data to the server.

[0084] Step 11:

[0085] The server then verifies the illustration data and reformats it into a user-viewable format, such as JPEG or PNG.

[0086] Step 12:

[0087] The server sends the optimized illustration file to the user's terminal.

[0088] Step 13:

[0089] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0090] Through the above processing steps, the user can easily and efficiently obtain an illustration with the style they desire.

[0091] Example 1

[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0093] While the demand for digital art has increased in recent years, generating high-quality digital illustrations from handwritten sketches requires a high level of specialized knowledge and skill. It is particularly difficult to efficiently generate works that reflect a specific style. To solve this problem, a system that can be intuitively operated by users and automatically generate illustrations in a variety of styles is needed.

[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0095] In this invention, the server includes a means for receiving a handwritten draft image, a means for receiving the draft image along with text information specifying a taste, a means for passing data to a generative AI model that generates an illustration based on the analysis results of the draft image and the specified taste, and a means for transmitting the illustration returned from the generative AI model to a user terminal. This makes it possible to automatically generate high-quality illustrations with the taste specified by the user based on the handwritten draft image.

[0096] A "hand-drawn draft image" refers to draft data of an illustration or composition that a user has hand-drawn using calligraphy implements or a digital device.

[0097] "Text information specifying taste" refers to information entered in text form by the user that specifies the style and color tone of the illustration desired.

[0098] A "generative AI model" is an artificial intelligence model that generates illustrations based on draft images and taste specifications. For example, a model based on deep learning technology falls into this category.

[0099] "Resizing" refers to the process of changing the resolution or size of an original image.

[0100] "Normalization" is a process that standardizes the format of image data, adjusts color tone, standardizes resolution, etc., to prepare the data in a format that is easy for AI models to use.

[0101] A "user terminal" is a digital device such as a computer, smartphone, or tablet that is operated by a user.

[0102] "Illustration" refers to visual artwork created by a generative AI model based on a specified taste or style.

[0103] A "server" is a computer system that receives data from users, processes it, works with an AI model to generate illustrations, and sends the generated results to the user's device.

[0104] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0105] First, the user uses a device such as a PC or smartphone to create a handwritten draft using digital painting software (general name: digital painting software). The created draft image is then uploaded to the system's web application or dedicated app. To do this, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from the file dialog.

[0106] Next, the user enters the desired style of the illustration in the text input field, for example, "soft watercolor-style colors" or "bright manga-style colors," and once the draft image and style specification are complete, the data is sent to the server by pressing the send button.

[0107] The server receives the draft image and text information specifying the taste sent by the user. The received draft image is resized and normalized to optimize the image resolution and format for AI processing. This preprocessing is performed using the Python Pillow library (general name: image processing library). The text information specifying the taste is also analyzed using a natural language processing library (general name: natural language processing library) and mapped to predefined taste options.

[0108] The server passes the preprocessed draft image data and the mapped taste information as input to a generative AI model. These models utilize deep learning technology (commonly known as generative AI models), such as Stable Diffusion and DALL-E. These AI models analyze shapes and objects from the draft image and generate the final illustration based on the specified taste.

[0109] The illustrations returned by the generative AI model are received by the server and reformatted into a format that can be viewed by the user. This reformatting includes image format conversion and compression. The reformatted illustrations are temporarily stored in cloud storage (commonly known as cloud storage services).

[0110] The user can view the illustration returned from the server on their device, save the generated illustration, or request further adjustments or additions. For example, they can input a prompt such as, "Please generate a warm, hand-drawn illustration based on this draft image." In this way, the system allows users to intuitively operate the system and easily generate high-quality illustrations.

[0111] As described above, this system makes it possible to automatically generate high-quality illustrations in a style specified by the user based on a handwritten draft image.

[0112] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0113] Step 1:

[0114] A user uses a digital device (PC or smartphone) to create a handwritten draft using paint software. The created draft image is then uploaded to the system's web application or a dedicated app. Specifically, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from a file dialog.

[0115] Input: A user-created draft image file.

[0116] Output: The uploaded draft image data.

[0117] Step 2:

[0118] The user enters the desired style of the illustration in a text input field on the web page and presses the send button to send the draft image and style specification to the server. Specifically, the user enters styles such as "soft watercolor-style colors" or "bright manga-style colors."

[0119] Input: The text information of the taste specified by the user.

[0120] Output: Draft image data and text information with specified tastes sent to the server.

[0121] Step 3:

[0122] The server receives the draft image and text information specified by the user. The received draft image is resized and normalized using the Python Pillow library. The image resolution and format are optimized for the AI ​​model.

[0123] Input: Draft image data sent by the user.

[0124] Output: Resized and normalized draft image data.

[0125] Step 4:

[0126] The server analyzes the textual information of the taste specification using a natural language processing library (e.g., NLTK or spaCy) and maps it to predefined taste options.

[0127] Input: Text information of taste specification sent by user.

[0128] Output: Mapped taste information.

[0129] Step 5:

[0130] The server passes the resized and normalized draft image data and the mapped taste information as input to a generative AI model (e.g., Stable Diffusion, DALL-E).

[0131] Input: Preprocessed draft image data and mapped taste information.

[0132] Output: Illustration data generated based on the specified taste.

[0133] Step 6:

[0134] The server receives the illustration data returned by the generative AI model and reformats it into a format that can be viewed by users. It uses the Pillow library to convert and compress the image format.

[0135] Input: Illustration data returned from the generative AI model.

[0136] Output: Reformatted illustration data.

[0137] Step 7:

[0138] The server temporarily stores the reformatted illustration data in cloud storage and sends it to the user's device. The user can then check the generated illustration on the web page and save it if necessary. Specifically, the user looks at the preview of the illustration and clicks the "Save" button.

[0139] Input: Reformatted illustration data.

[0140] Output: The illustration displayed on the user's device and, if necessary, saved as illustration data.

[0141] This is the specific flow of the program processing for this system, which allows users to automatically generate high-quality illustrations in a specified style based on a handwritten draft image.

[0142] (Application example 1)

[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0144] Conventional illustration generation systems have the ability to generate illustrations by specifying tastes based on handwritten drafts, but lack a means to display the generated illustrations in conjunction with real space. As a result, users cannot check the generated illustrations by overlaying them on real space in real time. The present invention aims to provide a system that overlays illustrations generated from handwritten drafts on the user's visual device, allowing users to enjoy art more intuitively.

[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0146] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for sending the generated illustration to a user terminal, means for overlaying the generated illustration in real space, and means for displaying the illustration on the user's visual device based on taste options specified by the user. This allows the user to check the generated illustration by overlaying it in real space in real time.

[0147] A "hand-drawn draft image" is an image of the composition or motif of an illustration that a user has hand-drawn on paper or a digital device.

[0148] "Text information specifying taste" is information that specifies the style and color tone of the illustration desired by the user in text format.

[0149] "Draft image analysis results" refers to information about shapes and objects obtained by the AI ​​model analyzing handwritten draft images.

[0150] The "specified taste" refers to the style and color tone of the illustration determined based on the character information specifying the taste by the user.

[0151] The "means of generating an illustration" is a function in which the AI ​​model generates the final illustration based on the analysis results of the draft image and the specified taste.

[0152] A "user terminal" is a digital device that receives and displays the generated illustrations, and includes smartphones, PCs, tablets, etc.

[0153] "Means for overlaying and displaying a generated illustration in real space" refers to a function that uses smart glasses or a head-mounted display to display a generated illustration superimposed on real space.

[0154] A "visual device" is a display device worn by a user on the eye, including smart glasses and head-mounted displays.

[0155] "Resizing and normalization" is the process of adjusting the size of the draft image and converting it into a format that is easy for the AI ​​model to use as input.

[0156] "Multiple taste options" are multiple style and color options that the user can choose from, such as "vintage style" and "modern art style."

[0157] The system for implementing this invention allows users to upload a draft image drawn by hand and generates an illustration that is overlaid on real space based on the user's specified taste. This system can overlay the illustration on real space in real time using visual devices such as smart glasses or a head-mounted display.

[0158] System configuration

[0159] The system consists of the following main components:

[0160] 1. User Device:

[0161] The user terminal includes smart glasses and a head-mounted display, and has the functions of taking handwritten images, specifying tastes, and uploading data.

[0162] The user takes a photo of their handwritten draft using the camera on their smart glasses or head-mounted display and saves it as an image file on their device.

[0163] 2. Server:

[0164] The server receives the draft image and character information specifying the taste sent from the user terminal.

[0165] The server resizes and normalizes the images, converting them into a format suitable for the AI ​​model.

[0166] The character information of the taste specification is analyzed and mapped to predefined taste options.

[0167] The server passes the preprocessed draft image and taste information as input to the AI ​​model, and generates an illustration based on the specified taste.

[0168] The generated illustration data is sent to the user's visual device to realize an overlay display.

[0169] Hardware and software used

[0170] Hardware:

[0171] Smart glasses (e.g. Google Glass)

[0172] Head-mounted displays (e.g. Microsoft HoloLens)

[0173] Digital camera devices (cameras built into smart glasses or HoloLens)

[0174] software:

[0175] On-device applications (e.g., developed with Unity)

[0176] Server-side image processing and resizing / normalization functions (e.g. OpenCV)

[0177] AI models (e.g., deep learning models using TensorFlow or PyTorch)

[0178] A description of what the program does

[0179] The server receives the user's handwritten draft image and the specified tastes. It first resizes and normalizes the image. This process uses the image processing library OpenCV. Next, the text information specified by the taste is mapped to predefined taste options and input into an AI model (a model using TensorFlow or PyTorch). The AI ​​model analyzes the shapes and objects in the draft image and generates the final illustration based on the specified tastes. The server then sends the generated illustration to the user's visual device, where the user can view the illustration overlaid on the real world in real time through smart glasses or a head-mounted display.

[0180] Specific examples

[0181] For example, if a user attending a workshop at an art supply store draws a handwritten sketch of a "flowering park scene," photographs it with smart glasses, and specifies the style as "vintage," the following prompt sentence will be entered:

[0182] Example prompt sentence:

[0183] Input image: A hand-drawn sketch of a park landscape with flowers in bloom

[0184] Style: Vintage

[0185] The AI ​​model generates a vintage-style park landscape illustration based on the user's specifications. This illustration is displayed on the user's smart glasses, allowing them to see the real world and digital art blended together in their own field of vision. This system allows users to intuitively enjoy their own original art.

[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0187] Step 1:

[0188] The user uses a handheld visual device (smart glasses or a head-mounted display) to take a picture of a handwritten draft and save it in the device. At this time, the camera of the visual device is activated and an operation to capture the draft image is performed. The input is the handwritten draft image, and the output is the captured image file.

[0189] Step 2:

[0190] The user uploads a captured draft image to the server through an application on the visual device. The application inputs text information for selecting an image file and specifying tastes through a user interface. The input is the captured draft image and the specified taste information, and the output is data transmission to the server.

[0191] Step 3:

[0192] The server receives the draft image and text information of taste specification sent by the user. The received data is prepared to be passed to the analysis process within the server. The input is the draft image and taste specification information, and the output is the readiness status for data analysis.

[0193] Step 4:

[0194] The server resizes and normalizes the received draft images. Specifically, it uses an image processing library such as OpenCV to adjust the image size and normalize the pixel values. The input is the original draft image, and the output is the resized and normalized image data.

[0195] Step 5:

[0196] The server maps the textual information of the taste specification to predefined taste options. This process uses a string matching algorithm to match the specified taste to an internal taste option list. The input is the textual information of the taste specification, and the output is the mapped taste information.

[0197] Step 6:

[0198] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input. The AI ​​model (using TensorFlow or PyTorch) analyzes the draft image and generates the final illustration based on the specified taste. The input is the preprocessed draft image and the mapped taste information, and the output is the generated illustration.

[0199] Step 7:

[0200] The server receives the generated illustration data and sends it to the user's visual device. The visual device processes the illustration to overlay it in real time. The input is the generated illustration data, and the output is the data transmission and display instructions to the visual device.

[0201] Step 8:

[0202] The user can view the illustrations overlaid on the real world in real time through a visual device, and can save or request additional illustrations as needed. The input is the illustration displayed on the visual device, and the output is the user's visual perception and operational instructions.

[0203] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0204] This invention combines a system in which a user uploads a hand-drawn draft image and an AI generates an illustration based on the user's specified taste with an emotion engine that recognizes the user's emotions, allowing the system to generate an illustration with a taste that matches the user's emotions.

[0205] Program processing

[0206] User operations

[0207] First, the user creates a rough draft of the illustration's composition and motif on a device at hand (PC or smartphone).

[0208] Next, upload the draft image to the system's web application or dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0209] Users can enter the style of the illustration they want in a text input field, but the emotion engine can also analyze the user's emotions and automatically select the appropriate style.

[0210] Manipulating the Emotion Engine

[0211] The emotion engine analyzes emotions from facial expressions, voice, input text, etc. acquired from the user's device. Examples of emotions include joy, sadness, surprise, and anger.

[0212] Based on the analysis results, the emotion engine determines the user's current emotional state and selects a taste that suits it. For example, if the user is feeling happy, a bright color tone taste will be selected.

[0213] Server Processing

[0214] The server receives the draft image sent by the user and the taste information selected by the emotion engine.

[0215] The received draft image is analyzed, resized, and normalized, converting it into a format optimized for AI processing.

[0216] The text information of the taste specification is analyzed and mapped to predefined taste options. The taste selected by the emotion engine is applied with priority.

[0217] Processing AI models

[0218] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the AI ​​model.

[0219] The AI ​​model analyzes shapes and objects from the draft image and generates the final illustration based on the specified taste. For example, if a user sketches a "cat sleeping on a sofa" and the emotion engine selects a "warm, hand-drawn" taste, the AI ​​model will generate an illustration according to that instruction.

[0220] Server provides results

[0221] The server receives the illustration data returned by the AI ​​model and reformats it into a user-viewable format, such as a common image format like JPEG or PNG.

[0222] The optimized illustration file is sent to the user's device.

[0223] User Verification

[0224] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0225] Specific examples

[0226] For example, if a user uploads a draft of a "boy reading a book" and selects the emotion joy, the emotion engine will analyze it and automatically select a bright, cartoon-style image. Based on this information, the server passes the data to an AI model, which then generates a final "bright, cartoon-style" illustration of a "boy reading a book." The generated illustration is then sent to the user's device, where they can view it.

[0227] In this way, by combining emotion engines, a system configuration is provided that can quickly generate illustrations with an appropriate style according to the user's emotions.

[0228] The processing flow will be explained below.

[0229] Step 1:

[0230] Users use their device (PC or smartphone) to create a rough draft of the illustration's composition and motif, which can then be freely drawn on paper or using digital tools.

[0231] Step 2:

[0232] Users can upload the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0233] Step 3:

[0234] When users upload a draft image, they can enter the style of their illustration they want in a text input field or choose an auto-select option, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0235] Step 4:

[0236] The device temporarily stores the draft image uploaded by the user and the taste information entered, and if the user selects the automatic selection option, it starts the emotion engine.

[0237] Step 5:

[0238] The emotion engine analyzes facial expression video and audio data acquired from the user's device, as well as input text data, to determine the user's emotions. For example, it uses facial recognition technology to recognize smiling and surprised expressions, and performs text analysis.

[0239] Step 6:

[0240] The emotion engine categorizes the user's emotional state based on the analysis results and selects the appropriate taste option. For example, if the user is expressing joy, a bright and positive taste will be selected.

[0241] Step 7:

[0242] The terminal sends the taste information selected by the emotion engine and the draft image to the server. Similarly, if the user manually inputs tastes, the terminal also sends the draft image and taste information to the server.

[0243] Step 8:

[0244] The server receives the draft image and taste specification character information sent by the user and stores the received data in a format that can be processed internally.

[0245] Step 9:

[0246] The server analyzes the received draft images and performs resizing and normalization processes, converting the draft images into a format optimal for AI processing.

[0247] Step 10:

[0248] The server analyzes the character information of the taste specification and maps it to predefined taste options. The taste information selected by the emotion engine is used preferentially.

[0249] Step 11:

[0250] The server passes the preprocessed draft image data and selected taste information to the AI ​​model as input.

[0251] Step 12:

[0252] The AI ​​model generates illustrations in a specified style based on the input draft image and taste information. For example, it generates an illustration by applying a "warm, hand-drawn" taste to a draft of a "cat sleeping on a sofa."

[0253] Step 13:

[0254] The illustration data generated by the AI ​​model is sent back to the server.

[0255] Step 14:

[0256] The server receives the generated illustration data and reformats it into a format that can be viewed by the user, for example, converting it into JPEG or PNG format.

[0257] Step 15:

[0258] The server sends the optimized illustration file to the user's terminal.

[0259] Step 16:

[0260] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0261] Through the above processing steps, users can efficiently obtain illustrations with the style they desire. Furthermore, by using the emotion engine, illustrations that match the user's emotional state can be automatically generated.

[0262] Example 2

[0263] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0264] Conventional illustration generation systems could generate illustrations based on tastes specified by a user's hand-drawn draft image, but it was difficult to automatically select an appropriate taste that took the user's emotional state into consideration, which resulted in the problem of users being unable to quickly obtain an illustration that matched their emotions.

[0265] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0266] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for analyzing emotions from the user's facial expression, voice, and input text, means for selecting a taste based on the emotion analysis results, and means for transmitting the generated illustration to the user terminal. This makes it possible to quickly generate an illustration with an appropriate taste that takes into account the user's emotional state.

[0267] A "draft image" is an image of an illustration that has been hand-drawn by a user in the initial stage.

[0268] "Taste" is setting information that indicates the style, color tone, and design direction of an illustration.

[0269] The "emotion engine" is a system that analyzes the user's facial expressions, voice, input text, etc. to determine the user's emotional state, and selects an appropriate taste based on the results.

[0270] The "server" is a computer system that processes the data received from the user, uses an AI model to generate the final illustration, and returns it to the user.

[0271] An "AI model" is an algorithm that uses machine learning techniques such as deep learning to take a rough sketch image and taste information as input and generate an illustration in a specified style.

[0272] "Resizing" is the process of changing the size of a draft image and adjusting it to the optimal resolution for processing by the AI ​​model.

[0273] "Normalization" is the process of converting image data into a uniform format to make analysis by AI models more efficient.

[0274] "Emotion analysis" is a technology that determines a user's emotional state based on data such as the user's facial expressions, voice, and input text.

[0275] A "generated illustration" is a final digital image generated by an AI model based on a draft image and a specified style.

[0276] A "user terminal" is a device used by a user, such as a PC or smartphone, on which the generated illustration is displayed.

[0277] This invention is a system in which a user uploads a hand-drawn draft image and a generative AI model generates an illustration based on the user's specified taste, and further recognizes the user's emotion and generates an illustration with a taste corresponding to that emotion. This system is composed of a server, a terminal, and an emotion engine.

[0278] Program Overview

[0279] User operations

[0280] First, the user creates a rough sketch of the illustration's composition and motif using a device (PC or smartphone). For example, they can use a drawing app on their smartphone. After creating the sketch, the user uploads the rough image to the system's web application or a dedicated app. During the upload process, the user clicks a dedicated upload button and selects the image file from a file dialog.

[0281] Manipulating the Emotion Engine

[0282] The emotion engine acquires facial expressions, voice, input text, etc. from the user's device. It uses the device's camera and microphone to collect the user's facial expressions and tone of voice, and analyzes the acquired data in real time. This allows it to determine the user's emotional state (happiness, sadness, surprise, anger, etc.) and select an appropriate illustration style accordingly.

[0283] Server Processing

[0284] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it uses an image analysis module (e.g., OpenCV) to resize and normalize the draft image. This converts the image into a format optimal for AI model processing. It also analyzes the text information of the taste specification and maps the results based on the taste selected by the emotion engine.

[0285] Processing AI models

[0286] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. For example, using a deep learning framework (TensorFlow or PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste. The AI ​​model recognizes shapes and objects from the image and draws them in a style that matches the taste.

[0287] Server provides results

[0288] The server receives the illustration data returned by the AI ​​model and converts it into a format that the user can view, such as JPEG or PNG, optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[0289] User review and feedback of results

[0290] Users can view the generated illustrations on their devices, save them, or share them on social media. If necessary, they can request regeneration or make additional adjustments.

[0291] Examples of concrete examples and prompts

[0292] For example, if a user uploads a draft of a "boy reading a book" and expresses joy as an emotion, the emotion engine analyzes it and selects a bright color tone. Based on this information, the server passes the data to an AI model, which ultimately generates an illustration of a "boy reading a book in bright colors." The generated illustration is sent to the user's device, where the user can view it.

[0293] Prompt Sentence Examples

[0294] Prompt: "If a user uploads a draft composition of dogs playing in a park and experiences happiness as an emotion, the sentiment analysis engine should select a bright, vibrant, cartoon-like design."

[0295] Prompt: "If the composition shows a person relaxing by the sea, the emotion analysis engine should detect the relaxed emotion and generate a calm, soft-colored, hand-drawn illustration."

[0296] In this way, the system uses user input and emotion-based prompts to provide instructions to the AI ​​model for generating optimal illustrations.

[0297] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0298] Step 1:

[0299] The user creates a draft image using a drawing app on their device (PC or smartphone). The created draft image is then uploaded to the system using a dedicated application. Specifically, the user clicks a dedicated upload button and selects the desired image file from a file dialog.

[0300] Input: User-created draft image file

[0301] Output: Draft image file sent to the server

[0302] Step 2:

[0303] The emotion engine captures facial expressions, voice, and input text from the user's device in real time and analyzes the data to determine the user's emotional state. For example, it uses the device's camera and microphone to collect facial expressions and voice data and sends that data to an analysis algorithm.

[0304] Input: User's facial expressions, voice, input text

[0305] Output: Parsed emotional state (e.g., happy, sad, surprised, angry)

[0306] Step 3:

[0307] Based on the emotion analysis results, the emotion engine selects a taste that matches the user's emotional state. For example, if the emotion of joy is detected, a bright color tone taste will be selected. This information is saved as text and sent to the system.

[0308] Input: Parsed emotional state

[0309] Output: Selected taste information (e.g. bright colors, cartoon style)

[0310] Step 4:

[0311] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it resizes and normalizes the draft image using an image analysis module (e.g., OpenCV), converting the image into a format suitable for processing by the AI ​​model.

[0312] Input: Draft image file, selected taste information

[0313] Output: Preprocessed draft image data, taste information for analysis

[0314] Step 5:

[0315] The server analyzes the character information of the taste specification and maps the results based on the taste selected by the emotion engine, thereby determining the setting value corresponding to the specified taste.

[0316] Input: Selected taste information

[0317] Output: Mapped taste setting value

[0318] Step 6:

[0319] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. Using a deep learning framework (e.g., TensorFlow, PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste.

[0320] Input: Preprocessed draft image data, mapped taste settings

[0321] Output: Generated illustration data

[0322] Step 7:

[0323] The server receives the illustration data returned by the AI ​​model and converts it into a format that can be viewed by the user (e.g., JPEG or PNG), optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[0324] Input: Generated illustration data

[0325] Output: The converted illustration data sent to the user's device

[0326] Step 8:

[0327] Users can check the generated illustrations on their devices, save them or share them on social media as needed, and can even request a re-generation if they are not satisfied with the quality or style of the illustration.

[0328] Input: Illustration data sent to the user's device

[0329] Output: User feedback and regeneration requests

[0330] In this way, by explaining the specific operations and data flow at each processing step, the system's functions can be understood in detail.

[0331] (Application example 2)

[0332] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0333] A problem with modern digital advertising is the lack of technology that can generate personalized visuals that respond to a user's emotions. Traditional advertisements have a static, uniform design and cannot adapt to individual user emotions. This limits the effectiveness of advertisements and makes it difficult to attract users' attention.

[0334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0335] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis result of the draft image and the specified taste, emotion analysis means for analyzing a user's emotions, means for the emotion analysis means to select a taste based on the user's emotions, and means for transmitting the generated illustration to a user terminal, thereby enabling the generation of personalized advertising visuals according to the user's emotions.

[0336] "Hand-drawn draft images" refer to early-stage illustrations or blueprints created by a user using hand-drawn techniques on paper or a digital device.

[0337] "Text information specifying taste" refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[0338] "Draft image analysis results" refers to information about the image features and structure obtained after processing and analyzing the draft image.

[0339] "Generating an illustration" refers to the digital creation of a final visual work based on input data.

[0340] "User terminal" refers to a digital device, such as a personal computer, smartphone, or tablet, that is used to receive and view system output.

[0341] "Emotion analysis means" refers to a combination of hardware and software for analyzing a user's emotions, specifically using data such as facial expressions, voice, and text input.

[0342] "Selecting a taste based on emotion" refers to a process of automatically selecting the appropriate taste option according to the result of analyzing the user's emotion.

[0343] "Resize and normalize" refers to the process of resizing and standardizing the input draft image to convert it into a format optimal for processing by the AI ​​model.

[0344] "Multiple taste options" refers to the various styles and atmospheres the system offers.

[0345] "Generating advertising visuals" refers to generating visual materials for the purpose of advertising or promotion.

[0346] This invention is applied to a system in which a user uploads a draft image drawn by hand and AI generates advertising visuals based on the user's specified tastes and emotions.

[0347] The server first receives a handwritten draft image from the user's device. A handwritten draft image refers to an early stage illustration or blueprint created by the user using paper or a digital device. Next, the server receives text information specifying the style along with the draft image. The text information specifying the style refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[0348] The user's emotions are then analyzed using an emotion analysis means. The emotion analysis means refers to a combination of hardware and software for analyzing the user's emotions, specifically using data such as facial expressions, voice, and text input. Based on the analysis results, a taste corresponding to the emotion is automatically selected. Selecting a taste based on emotion refers to the process of automatically selecting the appropriate taste option according to the result of the user's emotion analysis.

[0349] Next, the draft image is resized and normalized. Resizing and normalization refers to the process of changing the size and standardizing the input draft image to convert it into a format that is optimal for processing by the AI ​​model. This converts the draft image into a format suitable for processing and prepares it for input into the AI ​​model.

[0350] The server uses a generative AI model to generate advertising visuals based on the analysis results of the draft image and the specified taste. Here, generating advertising visuals refers to generating visual materials for advertising and promotional purposes. The generated advertising visuals are converted into common image formats (e.g., JPEG, PNG) and sent to the user's device. A user's device refers to a digital device, such as a personal computer, smartphone, or tablet, used to receive and view the system's output.

[0351] As a concrete example, consider a case where a user takes a photo of "new shoes" and inputs the emotion "excitement." The emotion engine analyzes this, and the AI ​​model generates bright, dynamic advertising visuals. The generated visuals are sent to the user's device.

[0352] An example of a prompt for the generative AI model is, "Generate a bright and dynamic advertising visual from a product image of 'new shoes' and the emotion 'excitement.'" This system enables the rapid generation of personalized advertising visuals according to the user's emotions.

[0353] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0354] Step 1:

[0355] The user creates a handwritten draft image on their device, then photographs or scans it and uploads it.

[0356] Input: A draft image drawn by the user.

[0357] How it works: Digitizes an image using the device's camera or scanner and sends it to a server through the application

[0358] Output: Digital draft images are sent to the server

[0359] Step 2:

[0360] The user can enter the desired taste information in the text input field, or the taste will be automatically selected based on sentiment analysis.

[0361] Input: User-entered taste information or user sentiment

[0362] Operation: Enter taste information into the device's text input field, or obtain the user's facial expressions and voice to perform emotion analysis.

[0363] Output: Taste information is sent to the server.

[0364] Step 3:

[0365] The server resizes and normalizes the received draft image.

[0366] Input: Draft image received by the server

[0367] How it works: Using an image processing library such as OpenCV, we resize and normalize the image to convert it into a format suitable for AI models.

[0368] Output: Resized and normalized image data

[0369] Step 4:

[0370] Select the appropriate taste based on sentiment analysis data.

[0371] Input: User sentiment analysis data and specified taste information

[0372] How it works: Uses sentiment analysis to automatically select tastes that match the user's emotions.

[0373] Output: Selected taste information

[0374] Step 5:

[0375] The server inputs the draft image and taste information into the generated AI model to generate advertising visuals.

[0376] Input: Resized and normalized draft image, selected taste information

[0377] How it works: Input data into a generative AI model to generate ad visuals

[0378] Output: Generated ad visuals

[0379] Step 6:

[0380] The server converts the generated advertising visuals into a common image format.

[0381] Input: Generated ad visual data

[0382] How it works: Converts to JPEG or PNG format using an image format conversion library.

[0383] Output: Converted advertising visual image

[0384] Step 7:

[0385] The server sends the final advertisement visual to the user's terminal.

[0386] Input: Transformed ad visual image

[0387] Behavior: Sends image data to the user's device using an HTTP request, etc.

[0388] Output: Ad visual image sent to user's device

[0389] Step 8:

[0390] The user checks the received advertisement visuals on the terminal.

[0391] Input: Ad visual image sent to user terminal

[0392] Behavior: Display advertising visuals in the device's image viewer or application.

[0393] Output: Ad visuals for the user to see

[0394] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0395] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0396] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0397] [Second embodiment]

[0398] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0399] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0400] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0401] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0402] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0404] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0405] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0406] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0407] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0408] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0409] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0410] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0411] Program processing

[0412] User operations

[0413] Users create a rough draft of the illustration's composition and motif on their device (PC or smartphone).

[0414] Upload the completed draft image to the system's web application or dedicated app by clicking the Upload button and selecting the appropriate image file from the file dialog.

[0415] Next, enter the style of illustration you want in the text input field, such as "soft watercolor colors" or "bright manga colors."

[0416] Once you have completed the draft image and taste specification, press the send button to send the data to the server.

[0417] Server Processing

[0418] The server receives the draft image and character information specifying the taste sent by the user.

[0419] The received draft image is analyzed and resized and normalized, optimizing the image resolution and format for AI processing.

[0420] It also maps the text information of the taste specification to predefined taste options, which determine the style and color tone of the generated illustration.

[0421] Processing AI models

[0422] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input.

[0423] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the specified taste. For example, if the draft is a "composition of a boy reading a book" and the taste is specified as "bright, cartoon-style colors," the AI ​​model will generate a cartoon-style illustration of a boy reading a book.

[0424] Server provides results

[0425] The server receives the illustration data returned by the AI ​​model and reformats it into a format that can be viewed by the user.

[0426] The generated illustration data is sent to the user's terminal, and the results are provided to the user.

[0427] User Verification

[0428] The user checks the illustration sent back from the server on the terminal.

[0429] You can save the generated illustrations as needed, and make further adjustments or requests.

[0430] Specific examples

[0431] For example, if a user uploads a "picture of a cat sleeping on a sofa" as a draft and specifies the taste as "warm, hand-drawn," the process will proceed as follows.

[0432] 1. The user simply draws the shape of a cat and a sofa and uploads the image file to the system.

[0433] 2. The user enters the taste as "Warm hand-drawn" in the text input field and clicks the submit button.

[0434] 3. The server receives the draft image and taste information, resizes and normalizes the image, and maps the taste information.

[0435] 4. The pre-processed data is sent to an AI model to generate a hand-drawn illustration.

[0436] 5. The server receives the generated illustration data and sends it to the user.

[0437] 6. The user checks the generated illustration on their device and requests saving or addition as needed.

[0438] In this way, the system allows users to easily generate high-quality illustrations.

[0439] The processing flow will be explained below.

[0440] Step 1:

[0441] Users use their devices (PC or smartphone) to create a rough draft of the composition and motif of the illustration, which represents a specific scene, such as a cat sleeping on a sofa.

[0442] Step 2:

[0443] The user uploads the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0444] Step 3:

[0445] After uploading is complete, users can enter the desired style of the illustration in a text input field that appears, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0446] Step 4:

[0447] After the user inputs the taste specification, the user clicks the send button to send the draft image and taste information to the server.

[0448] Step 5:

[0449] The server receives the draft image and taste specification character information sent by the user, and stores the received data in a format that can be processed immediately.

[0450] Step 6:

[0451] The server analyzes the received draft images and resizes and normalizes them, converting them into a format that is optimally processed by the AI ​​model.

[0452] Step 7:

[0453] The server analyzes the textual information of the taste specification and maps it to predefined taste options, which are then converted into a format that is easy for the AI ​​model to understand.

[0454] Step 8:

[0455] The server passes the preprocessed draft image data and taste information to the AI ​​model as input.

[0456] Step 9:

[0457] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the user's preferences, such as a scene of a cat sleeping on a sofa in a warm, hand-drawn style.

[0458] Step 10:

[0459] The AI ​​model returns the generated illustration data to the server.

[0460] Step 11:

[0461] The server then verifies the illustration data and reformats it into a user-viewable format, such as JPEG or PNG.

[0462] Step 12:

[0463] The server sends the optimized illustration file to the user's terminal.

[0464] Step 13:

[0465] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0466] Through the above processing steps, the user can easily and efficiently obtain an illustration with the style they desire.

[0467] Example 1

[0468] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0469] While the demand for digital art has increased in recent years, generating high-quality digital illustrations from handwritten sketches requires a high level of specialized knowledge and skill. It is particularly difficult to efficiently generate works that reflect a specific style. To solve this problem, a system that can automatically generate illustrations in a variety of styles and that can be intuitively operated by users is needed.

[0470] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0471] In this invention, the server includes a means for receiving a handwritten draft image, a means for receiving the draft image along with text information specifying a taste, a means for passing data to a generative AI model that generates an illustration based on the analysis results of the draft image and the specified taste, and a means for transmitting the illustration returned from the generative AI model to a user terminal. This makes it possible to automatically generate high-quality illustrations with the taste specified by the user based on the handwritten draft image.

[0472] A "hand-drawn draft image" refers to draft data of an illustration or composition that a user has hand-drawn using calligraphy implements or a digital device.

[0473] "Text information specifying taste" refers to information entered in text form by the user that specifies the style and color tone of the illustration desired.

[0474] A "generative AI model" is an artificial intelligence model that generates illustrations based on draft images and taste specifications. For example, a model based on deep learning technology falls into this category.

[0475] "Resizing" refers to the process of changing the resolution or size of an original image.

[0476] "Normalization" is a process that standardizes the format of image data, adjusts color tone, standardizes resolution, etc., to prepare the data in a format that is easy for AI models to use.

[0477] A "user terminal" is a digital device such as a computer, smartphone, or tablet that is operated by a user.

[0478] "Illustration" refers to visual artwork created by a generative AI model based on a specified taste or style.

[0479] A "server" is a computer system that receives data from users, processes it, works with an AI model to generate illustrations, and sends the generated results to the user's device.

[0480] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0481] First, the user uses a device such as a PC or smartphone to create a handwritten draft using digital painting software (general name: digital painting software). The created draft image is then uploaded to the system's web application or dedicated app. To do this, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from the file dialog.

[0482] Next, the user enters the desired style of the illustration in the text input field, for example, "soft watercolor-style colors" or "bright manga-style colors," and once the draft image and style specification are complete, the data is sent to the server by pressing the send button.

[0483] The server receives the draft image and text information specifying the taste sent by the user. The received draft image is resized and normalized to optimize the image resolution and format for AI processing. This preprocessing is performed using the Python Pillow library (general name: image processing library). The text information specifying the taste is also analyzed using a natural language processing library (general name: natural language processing library) and mapped to predefined taste options.

[0484] The server passes the preprocessed draft image data and the mapped taste information as input to a generative AI model. These models utilize deep learning technology (commonly known as generative AI models), such as Stable Diffusion and DALL-E. These AI models analyze shapes and objects from the draft image and generate the final illustration based on the specified taste.

[0485] The illustrations returned by the generative AI model are received by the server and reformatted into a format that can be viewed by the user. This reformatting includes image format conversion and compression. The reformatted illustrations are temporarily stored in cloud storage (commonly known as cloud storage services).

[0486] The user can view the illustration returned from the server on their device, save the generated illustration, or request further adjustments or additions. For example, they can input a prompt such as, "Please generate a warm, hand-drawn illustration based on this draft image." In this way, the system allows users to intuitively operate the system and easily generate high-quality illustrations.

[0487] As described above, this system makes it possible to automatically generate high-quality illustrations in a style specified by the user based on a handwritten draft image.

[0488] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0489] Step 1:

[0490] A user uses a digital device (PC or smartphone) to create a handwritten draft using paint software. The created draft image is then uploaded to the system's web application or a dedicated app. Specifically, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from a file dialog.

[0491] Input: A user-created draft image file.

[0492] Output: The uploaded draft image data.

[0493] Step 2:

[0494] The user enters the desired style of the illustration in a text input field on the web page and presses the send button to send the draft image and style specification to the server. Specifically, the user enters styles such as "soft watercolor-style colors" or "bright manga-style colors."

[0495] Input: The text information of the taste specified by the user.

[0496] Output: Draft image data and text information with specified tastes sent to the server.

[0497] Step 3:

[0498] The server receives the draft image and text information specified by the user. The received draft image is resized and normalized using the Python Pillow library. The image resolution and format are optimized for the AI ​​model.

[0499] Input: Draft image data sent by the user.

[0500] Output: Resized and normalized draft image data.

[0501] Step 4:

[0502] The server analyzes the textual information of the taste specification using a natural language processing library (e.g., NLTK or spaCy) and maps it to predefined taste options.

[0503] Input: Text information of taste specification sent by user.

[0504] Output: Mapped taste information.

[0505] Step 5:

[0506] The server passes the resized and normalized draft image data and the mapped taste information as input to a generative AI model (e.g., Stable Diffusion, DALL-E).

[0507] Input: Preprocessed draft image data and mapped taste information.

[0508] Output: Illustration data generated based on the specified taste.

[0509] Step 6:

[0510] The server receives the illustration data returned by the generative AI model and reformats it into a format that can be viewed by users. It uses the Pillow library to convert and compress the image format.

[0511] Input: Illustration data returned from the generative AI model.

[0512] Output: Reformatted illustration data.

[0513] Step 7:

[0514] The server temporarily stores the reformatted illustration data in cloud storage and sends it to the user's device. The user can then check the generated illustration on the web page and save it if desired. Specifically, the user looks at the preview of the illustration and clicks the "Save" button.

[0515] Input: Reformatted illustration data.

[0516] Output: The illustration displayed on the user's device and, if necessary, saved as illustration data.

[0517] This is the specific flow of the program processing for this system, which allows users to automatically generate high-quality illustrations in a specified style based on a handwritten draft image.

[0518] (Application example 1)

[0519] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0520] Conventional illustration generation systems have the ability to generate illustrations by specifying tastes based on handwritten drafts, but lack a means to display the generated illustrations in conjunction with real space. As a result, users cannot check the generated illustrations by overlaying them on real space in real time. The present invention aims to provide a system that overlays illustrations generated from handwritten drafts on the user's visual device, allowing users to enjoy art more intuitively.

[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0522] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for sending the generated illustration to a user terminal, means for overlaying the generated illustration in real space, and means for displaying the illustration on the user's visual device based on taste options specified by the user. This allows the user to check the generated illustration by overlaying it in real space in real time.

[0523] A "hand-drawn draft image" is an image of the composition or motif of an illustration that a user has hand-drawn on paper or a digital device.

[0524] "Text information specifying taste" is information that specifies the style and color tone of the illustration desired by the user in text format.

[0525] "Draft image analysis results" refers to information about shapes and objects obtained by the AI ​​model analyzing handwritten draft images.

[0526] The "specified taste" refers to the style and color tone of the illustration determined based on the character information specifying the taste by the user.

[0527] The "means of generating an illustration" is a function in which the AI ​​model generates the final illustration based on the analysis results of the draft image and the specified taste.

[0528] A "user terminal" is a digital device that receives and displays the generated illustrations, and includes smartphones, PCs, tablets, etc.

[0529] "Means for overlaying and displaying a generated illustration in real space" refers to a function that uses smart glasses or a head-mounted display to display a generated illustration superimposed on real space.

[0530] A "visual device" is a display device worn by a user on the eye, including smart glasses and head-mounted displays.

[0531] "Resizing and normalization" is the process of adjusting the size of the draft image and converting it into a format that is easy for the AI ​​model to use as input.

[0532] "Multiple taste options" are multiple style and color options that the user can choose from, such as "vintage style" and "modern art style."

[0533] The system for implementing this invention allows users to upload a draft image drawn by hand and generates an illustration that is overlaid on real space based on the user's specified taste. This system can overlay the illustration on real space in real time using visual devices such as smart glasses or a head-mounted display.

[0534] System configuration

[0535] The system consists of the following main components:

[0536] 1. User Device:

[0537] The user terminal includes smart glasses and a head-mounted display, and has the functions of taking handwritten images, specifying tastes, and uploading data.

[0538] The user takes a photo of their handwritten draft using the camera on their smart glasses or head-mounted display and saves it as an image file on their device.

[0539] 2. Server:

[0540] The server receives the draft image and character information specifying the taste sent from the user terminal.

[0541] The server resizes and normalizes the images, converting them into a format suitable for the AI ​​model.

[0542] The character information of the taste specification is analyzed and mapped to predefined taste options.

[0543] The server passes the preprocessed draft image and taste information as input to the AI ​​model, and generates an illustration based on the specified taste.

[0544] The generated illustration data is sent to the user's visual device to realize an overlay display.

[0545] Hardware and software used

[0546] Hardware:

[0547] Smart glasses (e.g. Google Glass)

[0548] Head-mounted displays (e.g. Microsoft HoloLens)

[0549] Digital camera devices (cameras built into smart glasses or HoloLens)

[0550] software:

[0551] On-device applications (e.g., developed with Unity)

[0552] Server-side image processing and resizing / normalization functions (e.g. OpenCV)

[0553] AI models (e.g., deep learning models using TensorFlow or PyTorch)

[0554] A description of what the program does

[0555] The server receives the user's handwritten draft image and the specified tastes. It first resizes and normalizes the image. This process uses the image processing library OpenCV. Next, the text information specified by the taste is mapped to predefined taste options and input into an AI model (a model using TensorFlow or PyTorch). The AI ​​model analyzes the shapes and objects in the draft image and generates the final illustration based on the specified tastes. The server then sends the generated illustration to the user's visual device, where the user can view the illustration overlaid on the real world in real time through smart glasses or a head-mounted display.

[0556] Specific examples

[0557] For example, if a user attending a workshop at an art supply store draws a handwritten sketch of a "flowering park scene," photographs it with smart glasses, and specifies the style as "vintage," the following prompt sentence will be entered:

[0558] Example prompt sentence:

[0559] Input image: A hand-drawn sketch of a park landscape with flowers in bloom

[0560] Style: Vintage

[0561] The AI ​​model generates a vintage-style park landscape illustration based on the user's specifications. This illustration is displayed on the user's smart glasses, allowing them to see the real world and digital art blended together in their own field of vision. This system allows users to intuitively enjoy their own original art.

[0562] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0563] Step 1:

[0564] The user uses a handheld visual device (smart glasses or a head-mounted display) to take a picture of a handwritten draft and save it in the device. At this time, the camera of the visual device is activated and an operation to capture the draft image is performed. The input is the handwritten draft image, and the output is the captured image file.

[0565] Step 2:

[0566] The user uploads a captured draft image to the server through an application on the visual device. The application inputs text information for selecting an image file and specifying tastes through a user interface. The input is the captured draft image and the specified taste information, and the output is data transmission to the server.

[0567] Step 3:

[0568] The server receives the draft image and text information of taste specification sent by the user. The received data is prepared to be passed to the analysis process within the server. The input is the draft image and taste specification information, and the output is the readiness status for data analysis.

[0569] Step 4:

[0570] The server resizes and normalizes the received draft images. Specifically, it uses an image processing library such as OpenCV to adjust the image size and normalize the pixel values. The input is the original draft image, and the output is the resized and normalized image data.

[0571] Step 5:

[0572] The server maps the textual information of the taste specification to predefined taste options. This process uses a string matching algorithm to match the specified taste to an internal taste option list. The input is the textual information of the taste specification, and the output is the mapped taste information.

[0573] Step 6:

[0574] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input. The AI ​​model (using TensorFlow or PyTorch) analyzes the draft image and generates the final illustration based on the specified taste. The input is the preprocessed draft image and the mapped taste information, and the output is the generated illustration.

[0575] Step 7:

[0576] The server receives the generated illustration data and sends it to the user's visual device, which then processes it to overlay the illustration in real time. The input is the generated illustration data, and the output is the data transmission and display instructions to the visual device.

[0577] Step 8:

[0578] The user can view the illustrations overlaid on the real world in real time through a visual device, and can save or request additional illustrations as needed. The input is the illustration displayed on the visual device, and the output is the user's visual perception and operational instructions.

[0579] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0580] This invention combines a system in which a user uploads a hand-drawn draft image and an AI generates an illustration based on the user's specified taste with an emotion engine that recognizes the user's emotions, allowing the system to generate an illustration with a taste that matches the user's emotions.

[0581] Program processing

[0582] User operations

[0583] First, the user creates a rough draft of the illustration's composition and motif on a device at hand (PC or smartphone).

[0584] Next, upload the draft image to the system's web application or dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0585] Users can enter the style of the illustration they want in a text input field, but the emotion engine can also analyze the user's emotions and automatically select the appropriate style.

[0586] Manipulating the Emotion Engine

[0587] The emotion engine analyzes emotions from facial expressions, voice, input text, etc. acquired from the user's device. Examples of emotions include joy, sadness, surprise, and anger.

[0588] Based on the analysis results, the emotion engine determines the user's current emotional state and selects a taste that suits it. For example, if the user is feeling happy, a bright color tone taste will be selected.

[0589] Server Processing

[0590] The server receives the draft image sent by the user and the taste information selected by the emotion engine.

[0591] The received draft image is analyzed, resized, and normalized, converting it into a format optimized for AI processing.

[0592] The text information of the taste specification is analyzed and mapped to predefined taste options. The taste selected by the emotion engine is applied with priority.

[0593] Processing AI models

[0594] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the AI ​​model.

[0595] The AI ​​model analyzes shapes and objects from the draft image and generates the final illustration based on the specified taste. For example, if a user sketches a "cat sleeping on a sofa" and the emotion engine selects a "warm, hand-drawn" taste, the AI ​​model will generate an illustration according to that instruction.

[0596] Server provides results

[0597] The server receives the illustration data returned by the AI ​​model and reformats it into a user-viewable format, such as a common image format like JPEG or PNG.

[0598] The optimized illustration file is sent to the user's device.

[0599] User Verification

[0600] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0601] Specific examples

[0602] For example, if a user uploads a draft of a "boy reading a book" and selects the emotion joy, the emotion engine will analyze it and automatically select a bright, cartoon-style image. Based on this information, the server passes the data to an AI model, which then generates a final "bright, cartoon-style" illustration of a "boy reading a book." The generated illustration is then sent to the user's device, where they can view it.

[0603] In this way, by combining emotion engines, a system configuration is provided that can quickly generate illustrations with an appropriate style according to the user's emotions.

[0604] The processing flow will be explained below.

[0605] Step 1:

[0606] Users use their device (PC or smartphone) to create a rough draft of the illustration's composition and motif, which can then be freely drawn on paper or using digital tools.

[0607] Step 2:

[0608] Users can upload the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0609] Step 3:

[0610] When users upload a draft image, they can enter the style of their illustration they want in a text input field or choose an auto-select option, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0611] Step 4:

[0612] The device temporarily stores the draft image uploaded by the user and the taste information entered, and if the user selects the automatic selection option, it starts the emotion engine.

[0613] Step 5:

[0614] The emotion engine analyzes facial expression video and audio data acquired from the user's device, as well as input text data, to determine the user's emotions. For example, it uses facial recognition technology to recognize smiling and surprised expressions, and performs text analysis.

[0615] Step 6:

[0616] The emotion engine categorizes the user's emotional state based on the analysis results and selects the appropriate taste option. For example, if the user is expressing joy, a bright and positive taste will be selected.

[0617] Step 7:

[0618] The terminal sends the taste information selected by the emotion engine and the draft image to the server. Similarly, if the user manually inputs tastes, the terminal also sends the draft image and taste information to the server.

[0619] Step 8:

[0620] The server receives the draft image and taste specification character information sent by the user and stores the received data in a format that can be processed internally.

[0621] Step 9:

[0622] The server analyzes the received draft images and performs resizing and normalization processes, converting the draft images into a format optimal for AI processing.

[0623] Step 10:

[0624] The server analyzes the character information of the taste specification and maps it to predefined taste options. The taste information selected by the emotion engine is used preferentially.

[0625] Step 11:

[0626] The server passes the preprocessed draft image data and selected taste information to the AI ​​model as input.

[0627] Step 12:

[0628] The AI ​​model generates illustrations in a specified style based on the input draft image and taste information. For example, it generates an illustration by applying a "warm, hand-drawn" taste to a draft of a "cat sleeping on a sofa."

[0629] Step 13:

[0630] The illustration data generated by the AI ​​model is sent back to the server.

[0631] Step 14:

[0632] The server receives the generated illustration data and reformats it into a format that can be viewed by the user, for example, converting it into JPEG or PNG format.

[0633] Step 15:

[0634] The server sends the optimized illustration file to the user's terminal.

[0635] Step 16:

[0636] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0637] Through the above processing steps, users can efficiently obtain illustrations with the style they desire. Furthermore, by using the emotion engine, illustrations that match the user's emotional state can be automatically generated.

[0638] Example 2

[0639] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0640] Conventional illustration generation systems could generate illustrations based on tastes specified by a user's hand-drawn draft image, but it was difficult to automatically select an appropriate taste that took the user's emotional state into consideration, which resulted in the problem of users being unable to quickly obtain an illustration that matched their emotions.

[0641] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0642] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for analyzing emotions from the user's facial expression, voice, and input text, means for selecting a taste based on the emotion analysis results, and means for transmitting the generated illustration to the user terminal. This makes it possible to quickly generate an illustration with an appropriate taste that takes into account the user's emotional state.

[0643] A "draft image" is an image of an illustration that has been hand-drawn by a user in the initial stage.

[0644] "Taste" is setting information that indicates the style, color tone, and design direction of an illustration.

[0645] The "emotion engine" is a system that analyzes the user's facial expressions, voice, input text, etc. to determine the user's emotional state, and selects an appropriate taste based on the results.

[0646] The "server" is a computer system that processes the data received from the user, uses an AI model to generate the final illustration, and returns it to the user.

[0647] An "AI model" is an algorithm that uses machine learning techniques such as deep learning to take a rough sketch image and taste information as input and generate an illustration in a specified style.

[0648] "Resizing" is the process of changing the size of a draft image and adjusting it to the optimal resolution for processing by the AI ​​model.

[0649] "Normalization" is the process of converting image data into a uniform format to make analysis by AI models more efficient.

[0650] "Emotion analysis" is a technology that determines a user's emotional state based on data such as the user's facial expressions, voice, and input text.

[0651] A "generated illustration" is a final digital image generated by an AI model based on a draft image and a specified style.

[0652] A "user terminal" is a device used by a user, such as a PC or smartphone, on which the generated illustration is displayed.

[0653] This invention is a system in which a user uploads a hand-drawn draft image and a generative AI model generates an illustration based on the user's specified taste, and further recognizes the user's emotion and generates an illustration with a taste corresponding to that emotion. This system is composed of a server, a terminal, and an emotion engine.

[0654] Program Overview

[0655] User operations

[0656] First, the user creates a rough sketch of the illustration's composition and motif using a device (PC or smartphone). For example, they can use a drawing app on their smartphone. After creating the sketch, the user uploads the rough image to the system's web application or a dedicated app. During the upload process, the user clicks a dedicated upload button and selects the image file from a file dialog.

[0657] Manipulating the Emotion Engine

[0658] The emotion engine acquires facial expressions, voice, input text, etc. from the user's device. It uses the device's camera and microphone to collect the user's facial expressions and tone of voice, and analyzes the acquired data in real time. This allows it to determine the user's emotional state (happiness, sadness, surprise, anger, etc.) and select an appropriate illustration style accordingly.

[0659] Server Processing

[0660] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it uses an image analysis module (e.g., OpenCV) to resize and normalize the draft image. This converts the image into a format optimal for AI model processing. It also analyzes the text information of the taste specification and maps the results based on the taste selected by the emotion engine.

[0661] Processing AI models

[0662] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. For example, using a deep learning framework (TensorFlow or PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste. The AI ​​model recognizes shapes and objects from the image and draws them in a style that matches the taste.

[0663] Server provides results

[0664] The server receives the illustration data returned by the AI ​​model and converts it into a format that the user can view, such as JPEG or PNG, optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[0665] User review and feedback of results

[0666] Users can view the generated illustrations on their devices, save them, or share them on social media. If necessary, they can request regeneration or make additional adjustments.

[0667] Examples of concrete examples and prompts

[0668] For example, if a user uploads a draft of a "boy reading a book" and expresses joy as an emotion, the emotion engine analyzes it and selects a bright color tone. Based on this information, the server passes the data to an AI model, which ultimately generates an illustration of a "boy reading a book in bright colors." The generated illustration is sent to the user's device, where the user can view it.

[0669] Prompt Sentence Examples

[0670] Prompt: "If a user uploads a draft composition of dogs playing in a park and experiences happiness as an emotion, the sentiment analysis engine should select a bright, vibrant, cartoon-like design."

[0671] Prompt: "If the composition shows a person relaxing by the sea, the emotion analysis engine should detect the relaxed emotion and generate a calm, soft-colored, hand-drawn illustration."

[0672] In this way, the system uses user input and emotion-based prompts to provide instructions to the AI ​​model for generating the best illustrations.

[0673] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0674] Step 1:

[0675] The user creates a draft image using a drawing app on their device (PC or smartphone). The created draft image is then uploaded to the system using a dedicated application. Specifically, the user clicks a dedicated upload button and selects the desired image file from a file dialog.

[0676] Input: User-created draft image file

[0677] Output: Draft image file sent to the server

[0678] Step 2:

[0679] The emotion engine captures facial expressions, voice, and input text from the user's device in real time and analyzes the data to determine the user's emotional state. For example, it uses the device's camera and microphone to collect facial expressions and voice data and sends that data to an analysis algorithm.

[0680] Input: User's facial expressions, voice, input text

[0681] Output: Parsed emotional state (e.g., happy, sad, surprised, angry)

[0682] Step 3:

[0683] Based on the emotion analysis results, the emotion engine selects a taste that matches the user's emotional state. For example, if the emotion of joy is detected, a bright color tone taste will be selected. This information is saved as text and sent to the system.

[0684] Input: Parsed emotional state

[0685] Output: Selected taste information (e.g. bright colors, cartoon style)

[0686] Step 4:

[0687] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it resizes and normalizes the draft image using an image analysis module (e.g., OpenCV), converting the image into a format suitable for processing by the AI ​​model.

[0688] Input: Draft image file, selected taste information

[0689] Output: Preprocessed draft image data, taste information for analysis

[0690] Step 5:

[0691] The server analyzes the character information of the taste specification and maps the results based on the taste selected by the emotion engine, thereby determining the setting value corresponding to the specified taste.

[0692] Input: Selected taste information

[0693] Output: Mapped taste setting value

[0694] Step 6:

[0695] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. Using a deep learning framework (e.g., TensorFlow, PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste.

[0696] Input: Preprocessed draft image data, mapped taste settings

[0697] Output: Generated illustration data

[0698] Step 7:

[0699] The server receives the illustration data returned by the AI ​​model and converts it into a format that can be viewed by the user (e.g., JPEG or PNG), optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[0700] Input: Generated illustration data

[0701] Output: The converted illustration data sent to the user's device

[0702] Step 8:

[0703] Users can check the generated illustrations on their devices, save them or share them on social media as needed, and can even request a re-generation if they are not satisfied with the quality or style of the illustration.

[0704] Input: Illustration data sent to the user's device

[0705] Output: User feedback and regeneration requests

[0706] In this way, by explaining the specific operations and data flow at each processing step, the system's functions can be understood in detail.

[0707] (Application example 2)

[0708] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0709] A problem with modern digital advertising is the lack of technology that can generate personalized visuals that respond to a user's emotions. Traditional advertisements have a static, uniform design and cannot adapt to individual user emotions. This limits the effectiveness of advertisements and makes it difficult to attract users' attention.

[0710] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0711] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis result of the draft image and the specified taste, emotion analysis means for analyzing a user's emotions, means for the emotion analysis means to select a taste based on the user's emotions, and means for transmitting the generated illustration to a user terminal, thereby enabling the generation of personalized advertising visuals according to the user's emotions.

[0712] "Hand-drawn draft images" refer to early-stage illustrations or blueprints created by users using hand-drawn techniques on paper or digital devices.

[0713] "Text information specifying taste" refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[0714] "Draft image analysis results" refers to information about the image features and structure obtained after processing and analyzing the draft image.

[0715] "Generating an illustration" refers to the digital creation of a final visual work based on input data.

[0716] "User terminal" refers to a digital device, such as a personal computer, smartphone, or tablet, that is used to receive and view system output.

[0717] "Emotion analysis means" refers to a combination of hardware and software for analyzing a user's emotions, specifically using data such as facial expressions, voice, and text input.

[0718] "Selecting a taste based on emotion" refers to a process of automatically selecting the appropriate taste option according to the result of analyzing the user's emotion.

[0719] "Resize and normalize" refers to the process of resizing and standardizing the input draft image to convert it into a format optimal for processing by the AI ​​model.

[0720] "Multiple taste options" refers to the various styles and atmospheres the system offers.

[0721] "Generating advertising visuals" refers to generating visual materials for the purpose of advertising or promotion.

[0722] This invention is applied to a system in which a user uploads a draft image drawn by hand and AI generates advertising visuals based on the user's specified tastes and emotions.

[0723] The server first receives a handwritten draft image from the user's device. A handwritten draft image refers to an early stage illustration or blueprint created by the user using paper or a digital device. Next, the server receives text information specifying the style along with the draft image. The text information specifying the style refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[0724] The user's emotions are then analyzed using an emotion analysis means. The emotion analysis means refers to a combination of hardware and software for analyzing the user's emotions, specifically using data such as facial expressions, voice, and text input. Based on the analysis results, a taste corresponding to the emotion is automatically selected. Selecting a taste based on emotion refers to the process of automatically selecting the appropriate taste option according to the result of the user's emotion analysis.

[0725] Next, the draft image is resized and normalized. Resizing and normalization refers to the process of changing the size and standardizing the input draft image to convert it into a format that is optimal for processing by the AI ​​model. This converts the draft image into a format suitable for processing and prepares it for input into the AI ​​model.

[0726] The server uses a generative AI model to generate advertising visuals based on the analysis results of the draft image and the specified taste. Here, generating advertising visuals refers to generating visual materials for advertising and promotional purposes. The generated advertising visuals are converted into common image formats (e.g., JPEG, PNG) and sent to the user's device. A user's device refers to a digital device, such as a personal computer, smartphone, or tablet, used to receive and view the system's output.

[0727] As a concrete example, consider a case where a user takes a photo of "new shoes" and inputs the emotion "excitement." The emotion engine analyzes this, and the AI ​​model generates bright, dynamic advertising visuals. The generated visuals are sent to the user's device.

[0728] An example of a prompt for the generative AI model is, "Generate a bright and dynamic advertising visual from a product image of 'new shoes' and the emotion 'excitement.'" This system enables the rapid generation of personalized advertising visuals according to the user's emotions.

[0729] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0730] Step 1:

[0731] The user creates a handwritten draft image on their device, then photographs or scans it and uploads it.

[0732] Input: A draft image drawn by the user.

[0733] How it works: Digitizes an image using the device's camera or scanner and sends it to a server through the application

[0734] Output: Digital draft images are sent to the server

[0735] Step 2:

[0736] The user can enter the desired taste information in the text input field, or the taste will be automatically selected based on sentiment analysis.

[0737] Input: User-entered taste information or user sentiment

[0738] Operation: Enter taste information into the device's text input field, or obtain the user's facial expressions and voice to perform emotion analysis.

[0739] Output: Taste information is sent to the server.

[0740] Step 3:

[0741] The server resizes and normalizes the received draft image.

[0742] Input: Draft image received by the server

[0743] How it works: Using an image processing library such as OpenCV, we resize and normalize the image to convert it into a format suitable for AI models.

[0744] Output: Resized and normalized image data

[0745] Step 4:

[0746] Select the appropriate taste based on sentiment analysis data.

[0747] Input: User sentiment analysis data and specified taste information

[0748] How it works: Uses sentiment analysis to automatically select tastes that match the user's emotions.

[0749] Output: Selected taste information

[0750] Step 5:

[0751] The server inputs the draft image and taste information into the generated AI model to generate advertising visuals.

[0752] Input: Resized and normalized draft image, selected taste information

[0753] How it works: Input data into a generative AI model to generate ad visuals

[0754] Output: Generated ad visuals

[0755] Step 6:

[0756] The server converts the generated advertising visuals into a common image format.

[0757] Input: Generated ad visual data

[0758] How it works: Converts to JPEG or PNG format using an image format conversion library.

[0759] Output: Converted advertising visual image

[0760] Step 7:

[0761] The server sends the final advertisement visual to the user's terminal.

[0762] Input: Transformed ad visual image

[0763] Behavior: Sends image data to the user's device using an HTTP request, etc.

[0764] Output: Ad visual image sent to user's device

[0765] Step 8:

[0766] The user checks the received advertisement visuals on the terminal.

[0767] Input: Ad visual image sent to user terminal

[0768] Behavior: Display advertising visuals in the device's image viewer or application.

[0769] Output: Ad visuals for the user to see

[0770] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0771] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0772] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0773] [Third embodiment]

[0774] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0775] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0776] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0777] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0778] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0779] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0780] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0781] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0782] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0783] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0784] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0785] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0786] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0787] Program processing

[0788] User operations

[0789] Users create a rough draft of the illustration's composition and motif on their device (PC or smartphone).

[0790] Upload the completed draft image to the system's web application or dedicated app by clicking the Upload button and selecting the appropriate image file from the file dialog.

[0791] Next, enter the style of illustration you want in the text input field, such as "soft watercolor colors" or "bright manga colors."

[0792] Once you have completed the draft image and taste specification, press the send button to send the data to the server.

[0793] Server Processing

[0794] The server receives the draft image and character information specifying the taste sent by the user.

[0795] The received draft image is analyzed and resized and normalized, optimizing the image resolution and format for AI processing.

[0796] It also maps the text information of the taste specification to predefined taste options, which determine the style and color tone of the generated illustration.

[0797] Processing AI models

[0798] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input.

[0799] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the specified taste. For example, if the draft is a "composition of a boy reading a book" and the taste is specified as "bright, cartoon-style colors," the AI ​​model will generate a cartoon-style illustration of a boy reading a book.

[0800] Server provides results

[0801] The server receives the illustration data returned by the AI ​​model and reformats it into a format that can be viewed by the user.

[0802] The generated illustration data is sent to the user's terminal, and the results are provided to the user.

[0803] User Verification

[0804] The user checks the illustration sent back from the server on the terminal.

[0805] You can save the generated illustrations as needed, and make further adjustments or requests.

[0806] Specific examples

[0807] For example, if a user uploads a "picture of a cat sleeping on a sofa" as a draft and specifies the taste as "warm, hand-drawn," the process will proceed as follows.

[0808] 1. The user simply draws the shape of a cat and a sofa and uploads the image file to the system.

[0809] 2. The user enters the taste as "Warm hand-drawn" in the text input field and clicks the submit button.

[0810] 3. The server receives the draft image and taste information, resizes and normalizes the image, and maps the taste information.

[0811] 4. The pre-processed data is sent to an AI model to generate a hand-drawn illustration.

[0812] 5. The server receives the generated illustration data and sends it to the user.

[0813] 6. The user checks the generated illustration on their device and requests saving or addition as needed.

[0814] In this way, the system allows users to easily generate high-quality illustrations.

[0815] The processing flow will be explained below.

[0816] Step 1:

[0817] Users use their devices (PC or smartphone) to create a rough draft of the composition and motif of the illustration, which represents a specific scene, such as a cat sleeping on a sofa.

[0818] Step 2:

[0819] The user uploads the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0820] Step 3:

[0821] After uploading is complete, users can enter the desired style of the illustration in a text input field that appears, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0822] Step 4:

[0823] After the user inputs the taste specification, the user clicks the send button to send the draft image and taste information to the server.

[0824] Step 5:

[0825] The server receives the draft image and taste specification character information sent by the user, and stores the received data in a format that can be processed immediately.

[0826] Step 6:

[0827] The server analyzes the received draft images and resizes and normalizes them, converting them into a format that is optimally processed by the AI ​​model.

[0828] Step 7:

[0829] The server analyzes the textual information of the taste specification and maps it to predefined taste options, which are then converted into a format that is easy for the AI ​​model to understand.

[0830] Step 8:

[0831] The server passes the preprocessed draft image data and taste information to the AI ​​model as input.

[0832] Step 9:

[0833] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the user's preferences, such as a cat sleeping on a sofa in a warm, hand-drawn style.

[0834] Step 10:

[0835] The AI ​​model returns the generated illustration data to the server.

[0836] Step 11:

[0837] The server then verifies the illustration data and reformats it into a user-viewable format, such as JPEG or PNG.

[0838] Step 12:

[0839] The server sends the optimized illustration file to the user's terminal.

[0840] Step 13:

[0841] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0842] Through the above processing steps, the user can easily and efficiently obtain an illustration with the style they desire.

[0843] Example 1

[0844] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0845] While the demand for digital art has increased in recent years, generating high-quality digital illustrations from handwritten sketches requires a high level of specialized knowledge and skill. It is particularly difficult to efficiently generate works that reflect a specific style. To solve this problem, a system that can be intuitively operated by users and automatically generate illustrations in a variety of styles is needed.

[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0847] In this invention, the server includes a means for receiving a handwritten draft image, a means for receiving the draft image along with text information specifying a taste, a means for passing data to a generative AI model that generates an illustration based on the analysis results of the draft image and the specified taste, and a means for transmitting the illustration returned from the generative AI model to a user terminal. This makes it possible to automatically generate high-quality illustrations with the taste specified by the user based on the handwritten draft image.

[0848] A "hand-drawn draft image" refers to draft data of an illustration or composition that a user has hand-drawn using calligraphy implements or a digital device.

[0849] "Text information specifying taste" refers to information entered in text form by the user that specifies the style and color tone of the illustration desired.

[0850] A "generative AI model" is an artificial intelligence model that generates illustrations based on draft images and taste specifications. For example, a model based on deep learning technology falls into this category.

[0851] "Resizing" refers to the process of changing the resolution or size of an original image.

[0852] "Normalization" is a process that standardizes the format of image data, adjusts color tone, standardizes resolution, etc., to prepare the data in a format that is easy for AI models to use.

[0853] A "user terminal" is a digital device such as a computer, smartphone, or tablet that is operated by a user.

[0854] "Illustration" refers to visual artwork created by a generative AI model based on a specified taste or style.

[0855] A "server" is a computer system that receives data from users, processes it, works with an AI model to generate illustrations, and sends the generated results to the user's device.

[0856] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[0857] First, the user uses a device such as a PC or smartphone to create a handwritten draft using digital painting software (general name: digital painting software). The created draft image is then uploaded to the system's web application or dedicated app. To do this, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from the file dialog.

[0858] Next, the user enters the desired style of the illustration in the text input field, for example, "soft watercolor-style colors" or "bright manga-style colors," and once the draft image and style specification are complete, the data is sent to the server by pressing the send button.

[0859] The server receives the draft image and text information specifying the taste sent by the user. The received draft image is resized and normalized to optimize the image resolution and format for AI processing. This preprocessing is performed using the Python Pillow library (general name: image processing library). The text information specifying the taste is also analyzed using a natural language processing library (general name: natural language processing library) and mapped to predefined taste options.

[0860] The server passes the preprocessed draft image data and the mapped taste information as input to a generative AI model. These models utilize deep learning technology (commonly known as generative AI models), such as Stable Diffusion and DALL-E. These AI models analyze shapes and objects from the draft image and generate the final illustration based on the specified taste.

[0861] The illustrations returned by the generative AI model are received by the server and reformatted into a format that can be viewed by the user. This reformatting includes image format conversion and compression. The reformatted illustrations are temporarily stored in cloud storage (commonly known as cloud storage services).

[0862] The user can view the illustration returned from the server on their device, save the generated illustration, or request further adjustments or additions. For example, they can input a prompt such as, "Please generate a warm, hand-drawn illustration based on this draft image." In this way, the system allows users to intuitively operate the system and easily generate high-quality illustrations.

[0863] As described above, this system makes it possible to automatically generate high-quality illustrations in a style specified by the user based on a handwritten draft image.

[0864] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0865] Step 1:

[0866] A user uses a digital device (PC or smartphone) to create a handwritten draft using paint software. The created draft image is then uploaded to the system's web application or a dedicated app. Specifically, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from a file dialog.

[0867] Input: A user-created draft image file.

[0868] Output: The uploaded draft image data.

[0869] Step 2:

[0870] The user enters the desired style of the illustration in a text input field on the web page and presses the send button to send the draft image and style specification to the server. Specifically, the user enters styles such as "soft watercolor-style colors" or "bright manga-style colors."

[0871] Input: The text information of the taste specified by the user.

[0872] Output: Draft image data and text information with specified tastes sent to the server.

[0873] Step 3:

[0874] The server receives the draft image and text information specified by the user. The received draft image is resized and normalized using the Python Pillow library. The image resolution and format are optimized for the AI ​​model.

[0875] Input: Draft image data sent by the user.

[0876] Output: Resized and normalized draft image data.

[0877] Step 4:

[0878] The server analyzes the textual information of the taste specification using a natural language processing library (e.g., NLTK or spaCy) and maps it to predefined taste options.

[0879] Input: Text information of taste specification sent by user.

[0880] Output: Mapped taste information.

[0881] Step 5:

[0882] The server passes the resized and normalized draft image data and the mapped taste information as input to a generative AI model (e.g., Stable Diffusion, DALL-E).

[0883] Input: Preprocessed draft image data and mapped taste information.

[0884] Output: Illustration data generated based on the specified taste.

[0885] Step 6:

[0886] The server receives the illustration data returned by the generative AI model and reformats it into a format that can be viewed by users. It uses the Pillow library to convert and compress the image format.

[0887] Input: Illustration data returned from the generative AI model.

[0888] Output: Reformatted illustration data.

[0889] Step 7:

[0890] The server temporarily stores the reformatted illustration data in cloud storage and sends it to the user's device. The user can then check the generated illustration on the web page and save it if desired. Specifically, the user looks at the preview of the illustration and clicks the "Save" button.

[0891] Input: Reformatted illustration data.

[0892] Output: The illustration displayed on the user's device and, if necessary, saved as illustration data.

[0893] This is the specific flow of the program processing for this system, which allows users to automatically generate high-quality illustrations in a specified style based on a handwritten draft image.

[0894] (Application example 1)

[0895] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0896] Conventional illustration generation systems have the ability to generate illustrations by specifying tastes based on handwritten drafts, but lack a means to display the generated illustrations in conjunction with real space. As a result, users cannot check the generated illustrations by overlaying them on real space in real time. The present invention aims to provide a system that overlays illustrations generated from handwritten drafts on the user's visual device, allowing users to enjoy art more intuitively.

[0897] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0898] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for sending the generated illustration to a user terminal, means for overlaying the generated illustration in real space, and means for displaying the illustration on the user's visual device based on taste options specified by the user. This allows the user to check the generated illustration by overlaying it in real space in real time.

[0899] A "hand-drawn draft image" is an image of the composition or motif of an illustration that a user has hand-drawn on paper or a digital device.

[0900] "Text information specifying taste" is information that specifies the style and color tone of the illustration desired by the user in text format.

[0901] "Draft image analysis results" refers to information about shapes and objects obtained by the AI ​​model analyzing handwritten draft images.

[0902] The "specified taste" refers to the style and color tone of the illustration determined based on the character information specifying the taste by the user.

[0903] The "means of generating an illustration" is a function in which the AI ​​model generates the final illustration based on the analysis results of the draft image and the specified taste.

[0904] A "user terminal" is a digital device that receives and displays the generated illustrations, and includes smartphones, PCs, tablets, etc.

[0905] "Means for overlaying and displaying a generated illustration in real space" refers to a function that uses smart glasses or a head-mounted display to display a generated illustration superimposed on real space.

[0906] A "visual device" is a display device worn by a user on the eye, including smart glasses and head-mounted displays.

[0907] "Resizing and normalization" is the process of adjusting the size of the draft image and converting it into a format that is easy for the AI ​​model to use as input.

[0908] "Multiple taste options" are multiple style and color options that the user can choose from, such as "vintage style" and "modern art style."

[0909] The system for implementing this invention allows users to upload a draft image drawn by hand and generates an illustration that is overlaid on real space based on the user's specified taste. This system can overlay the illustration on real space in real time using visual devices such as smart glasses or a head-mounted display.

[0910] System configuration

[0911] The system consists of the following main components:

[0912] 1. User Device:

[0913] The user terminal includes smart glasses and a head-mounted display, and has the functions of taking handwritten images, specifying tastes, and uploading data.

[0914] The user takes a photo of their handwritten draft using the camera on their smart glasses or head-mounted display and saves it as an image file on their device.

[0915] 2. Server:

[0916] The server receives the draft image and character information specifying the taste sent from the user terminal.

[0917] The server resizes and normalizes the images, converting them into a format suitable for the AI ​​model.

[0918] The character information of the taste specification is analyzed and mapped to predefined taste options.

[0919] The server passes the preprocessed draft image and taste information as input to the AI ​​model, and generates an illustration based on the specified taste.

[0920] The generated illustration data is sent to the user's visual device to realize an overlay display.

[0921] Hardware and software used

[0922] Hardware:

[0923] Smart glasses (e.g. Google Glass)

[0924] Head-mounted displays (e.g. Microsoft HoloLens)

[0925] Digital camera devices (cameras built into smart glasses or HoloLens)

[0926] software:

[0927] On-device applications (e.g., developed with Unity)

[0928] Server-side image processing and resizing / normalization functions (e.g. OpenCV)

[0929] AI models (e.g., deep learning models using TensorFlow or PyTorch)

[0930] A description of what the program does

[0931] The server receives the user's handwritten draft image and the specified tastes. It first resizes and normalizes the image. This process uses the image processing library OpenCV. Next, the text information specified by the taste is mapped to predefined taste options and input into an AI model (a model using TensorFlow or PyTorch). The AI ​​model analyzes the shapes and objects in the draft image and generates the final illustration based on the specified tastes. The server then sends the generated illustration to the user's visual device, where the user can view the illustration overlaid on the real world in real time through smart glasses or a head-mounted display.

[0932] Specific examples

[0933] For example, if a user attending a workshop at an art supply store draws a handwritten sketch of a "flowering park scene," photographs it with smart glasses, and specifies the style as "vintage," the following prompt sentence will be entered:

[0934] Example prompt sentence:

[0935] Input image: A hand-drawn sketch of a park landscape with flowers in bloom

[0936] Style: Vintage

[0937] The AI ​​model generates a vintage-style park landscape illustration based on the user's specifications. This illustration is displayed on the user's smart glasses, allowing them to see the real world and digital art blended together in their own field of vision. This system allows users to intuitively enjoy their own original art.

[0938] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0939] Step 1:

[0940] The user uses a handheld visual device (smart glasses or a head-mounted display) to take a picture of a handwritten draft and save it in the device. At this time, the camera of the visual device is activated and an operation to capture the draft image is performed. The input is the handwritten draft image, and the output is the captured image file.

[0941] Step 2:

[0942] The user uploads a captured draft image to the server through an application on the visual device. The application inputs text information for selecting an image file and specifying tastes through a user interface. The input is the captured draft image and the specified taste information, and the output is data transmission to the server.

[0943] Step 3:

[0944] The server receives the draft image and text information of taste specification sent by the user. The received data is prepared to be passed to the analysis process within the server. The input is the draft image and taste specification information, and the output is the readiness status for data analysis.

[0945] Step 4:

[0946] The server resizes and normalizes the received draft images. Specifically, it uses an image processing library such as OpenCV to adjust the image size and normalize the pixel values. The input is the original draft image, and the output is the resized and normalized image data.

[0947] Step 5:

[0948] The server maps the textual information of the taste specification to predefined taste options. This process uses a string matching algorithm to match the specified taste to an internal taste option list. The input is the textual information of the taste specification, and the output is the mapped taste information.

[0949] Step 6:

[0950] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input. The AI ​​model (using TensorFlow or PyTorch) analyzes the draft image and generates the final illustration based on the specified taste. The input is the preprocessed draft image and the mapped taste information, and the output is the generated illustration.

[0951] Step 7:

[0952] The server receives the generated illustration data and sends it to the user's visual device, which then processes it to overlay the illustration in real time. The input is the generated illustration data, and the output is the data transmission and display instructions to the visual device.

[0953] Step 8:

[0954] The user can view the illustrations overlaid on the real world in real time through a visual device, and can save or request additional illustrations as needed. The input is the illustration displayed on the visual device, and the output is the user's visual perception and operational instructions.

[0955] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0956] This invention combines a system in which a user uploads a hand-drawn draft image and an AI generates an illustration based on the user's specified taste with an emotion engine that recognizes the user's emotions, allowing the system to generate an illustration with a taste that matches the user's emotions.

[0957] Program processing

[0958] User operations

[0959] First, the user creates a rough draft of the illustration's composition and motif on a device at hand (PC or smartphone).

[0960] Next, upload the draft image to the system's web application or dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0961] Users can enter the style of the illustration they want in a text input field, but the emotion engine can also analyze the user's emotions and automatically select the appropriate style.

[0962] Manipulating the Emotion Engine

[0963] The emotion engine analyzes emotions from facial expressions, voice, input text, etc. acquired from the user's device. Examples of emotions include joy, sadness, surprise, and anger.

[0964] Based on the analysis results, the emotion engine determines the user's current emotional state and selects a taste that suits it. For example, if the user is feeling happy, a bright color tone taste will be selected.

[0965] Server Processing

[0966] The server receives the draft image sent by the user and the taste information selected by the emotion engine.

[0967] The received draft image is analyzed, resized, and normalized, converting it into a format optimized for AI processing.

[0968] The text information of the taste specification is analyzed and mapped to predefined taste options. The taste selected by the emotion engine is applied with priority.

[0969] Processing AI models

[0970] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the AI ​​model.

[0971] The AI ​​model analyzes shapes and objects from the draft image and generates the final illustration based on the specified taste. For example, if a user sketches a "cat sleeping on a sofa" and the emotion engine selects a "warm, hand-drawn" taste, the AI ​​model will generate an illustration according to that instruction.

[0972] Server provides results

[0973] The server receives the illustration data returned by the AI ​​model and reformats it into a user-viewable format, such as a common image format like JPEG or PNG.

[0974] The optimized illustration file is sent to the user's device.

[0975] User Verification

[0976] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[0977] Specific examples

[0978] For example, if a user uploads a draft of a "boy reading a book" and selects the emotion joy, the emotion engine will analyze it and automatically select a bright, cartoon-style image. Based on this information, the server passes the data to an AI model, which then generates a final "bright, cartoon-style" illustration of a "boy reading a book." The generated illustration is then sent to the user's device, where they can view it.

[0979] In this way, by combining emotion engines, a system configuration is provided that can quickly generate illustrations with an appropriate style according to the user's emotions.

[0980] The processing flow will be explained below.

[0981] Step 1:

[0982] Users use their device (PC or smartphone) to create a rough draft of the illustration's composition and motif, which can then be freely drawn on paper or using digital tools.

[0983] Step 2:

[0984] Users can upload the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[0985] Step 3:

[0986] When users upload a draft image, they can enter the style of their illustration they want in a text input field or choose an auto-select option, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[0987] Step 4:

[0988] The device temporarily stores the draft image uploaded by the user and the taste information entered, and if the user selects the automatic selection option, it starts the emotion engine.

[0989] Step 5:

[0990] The emotion engine analyzes facial expression video and audio data acquired from the user's device, as well as input text data, to determine the user's emotions. For example, it uses facial recognition technology to recognize smiling and surprised expressions, and performs text analysis.

[0991] Step 6:

[0992] The emotion engine categorizes the user's emotional state based on the analysis results and selects the appropriate taste option. For example, if the user is expressing joy, a bright and positive taste will be selected.

[0993] Step 7:

[0994] The terminal sends the taste information selected by the emotion engine and the draft image to the server. Similarly, if the user manually inputs tastes, the terminal also sends the draft image and taste information to the server.

[0995] Step 8:

[0996] The server receives the draft image and taste specification character information sent by the user and stores the received data in a format that can be processed internally.

[0997] Step 9:

[0998] The server analyzes the received draft images and performs resizing and normalization processes, converting the draft images into a format optimal for AI processing.

[0999] Step 10:

[1000] The server analyzes the character information of the taste specification and maps it to predefined taste options. The taste information selected by the emotion engine is used preferentially.

[1001] Step 11:

[1002] The server passes the preprocessed draft image data and selected taste information to the AI ​​model as input.

[1003] Step 12:

[1004] The AI ​​model generates illustrations in a specified style based on the input draft image and taste information. For example, it generates an illustration by applying a "warm, hand-drawn" taste to a draft of a "cat sleeping on a sofa."

[1005] Step 13:

[1006] The illustration data generated by the AI ​​model is sent back to the server.

[1007] Step 14:

[1008] The server receives the generated illustration data and reformats it into a format that can be viewed by the user, for example, converting it into JPEG or PNG format.

[1009] Step 15:

[1010] The server sends the optimized illustration file to the user's terminal.

[1011] Step 16:

[1012] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[1013] Through the above processing steps, users can efficiently obtain illustrations with the style they desire. Furthermore, by using the emotion engine, illustrations that match the user's emotional state can be automatically generated.

[1014] Example 2

[1015] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1016] Conventional illustration generation systems could generate illustrations based on tastes specified by a user's hand-drawn draft image, but it was difficult to automatically select an appropriate taste that took the user's emotional state into consideration, which resulted in the problem of users being unable to quickly obtain an illustration that matched their emotions.

[1017] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1018] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for analyzing emotions from the user's facial expression, voice, and input text, means for selecting a taste based on the emotion analysis results, and means for transmitting the generated illustration to the user terminal. This makes it possible to quickly generate an illustration with an appropriate taste that takes into account the user's emotional state.

[1019] A "draft image" is an image of an illustration that has been hand-drawn by a user in the initial stage.

[1020] "Taste" is setting information that indicates the style, color tone, and design direction of an illustration.

[1021] The "emotion engine" is a system that analyzes the user's facial expressions, voice, input text, etc. to determine the user's emotional state, and selects an appropriate taste based on the results.

[1022] The "server" is a computer system that processes the data received from the user, uses an AI model to generate the final illustration, and returns it to the user.

[1023] An "AI model" is an algorithm that uses machine learning techniques such as deep learning to take a rough sketch image and taste information as input and generate an illustration in a specified style.

[1024] "Resizing" is the process of changing the size of a draft image and adjusting it to the optimal resolution for processing by the AI ​​model.

[1025] "Normalization" is the process of converting image data into a uniform format to make analysis by AI models more efficient.

[1026] "Emotion analysis" is a technology that determines a user's emotional state based on data such as the user's facial expressions, voice, and input text.

[1027] A "generated illustration" is a final digital image generated by an AI model based on a draft image and a specified style.

[1028] A "user terminal" is a device used by a user, such as a PC or smartphone, on which the generated illustration is displayed.

[1029] This invention is a system in which a user uploads a hand-drawn draft image and a generative AI model generates an illustration based on the user's specified taste, and further recognizes the user's emotion and generates an illustration with a taste corresponding to that emotion. This system is composed of a server, a terminal, and an emotion engine.

[1030] Program Overview

[1031] User operations

[1032] First, the user creates a rough sketch of the illustration's composition and motif using a device (PC or smartphone). For example, they can use a drawing app on their smartphone. After creating the sketch, the user uploads the rough image to the system's web application or a dedicated app. During the upload process, the user clicks a dedicated upload button and selects the image file from a file dialog.

[1033] Manipulating the Emotion Engine

[1034] The emotion engine acquires facial expressions, voice, input text, etc. from the user's device. It uses the device's camera and microphone to collect the user's facial expressions and tone of voice, and analyzes the acquired data in real time. This allows it to determine the user's emotional state (happiness, sadness, surprise, anger, etc.) and select an appropriate illustration style accordingly.

[1035] Server Processing

[1036] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it uses an image analysis module (e.g., OpenCV) to resize and normalize the draft image. This converts the image into a format optimal for AI model processing. It also analyzes the text information of the taste specification and maps the results based on the taste selected by the emotion engine.

[1037] Processing AI models

[1038] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. For example, using a deep learning framework (TensorFlow or PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste. The AI ​​model recognizes shapes and objects from the image and draws them in a style that matches the taste.

[1039] Server provides results

[1040] The server receives the illustration data returned by the AI ​​model and converts it into a format that the user can view, such as JPEG or PNG, optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[1041] User review and feedback of results

[1042] Users can view the generated illustrations on their devices, save them, or share them on social media. If necessary, they can request regeneration or make additional adjustments.

[1043] Examples of concrete examples and prompts

[1044] For example, if a user uploads a draft of a "boy reading a book" and expresses joy as an emotion, the emotion engine analyzes it and selects a bright color tone. Based on this information, the server passes the data to an AI model, which ultimately generates an illustration of a "boy reading a book in bright colors." The generated illustration is sent to the user's device, where the user can view it.

[1045] Prompt Sentence Examples

[1046] Prompt: "If a user uploads a draft composition of dogs playing in a park and experiences happiness as an emotion, the sentiment analysis engine should select a bright, vibrant, cartoon-like design."

[1047] Prompt: "If the composition shows a person relaxing by the sea, the emotion analysis engine should detect the relaxed emotion and generate a calm, soft-colored, hand-drawn illustration."

[1048] In this way, the system uses user input and emotion-based prompts to provide instructions to the AI ​​model for generating the best illustrations.

[1049] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1050] Step 1:

[1051] The user creates a draft image using a drawing app on their device (PC or smartphone). The created draft image is then uploaded to the system using a dedicated application. Specifically, the user clicks a dedicated upload button and selects the desired image file from a file dialog.

[1052] Input: User-created draft image file

[1053] Output: Draft image file sent to the server

[1054] Step 2:

[1055] The emotion engine captures facial expressions, voice, and input text from the user's device in real time and analyzes the data to determine the user's emotional state. For example, it uses the device's camera and microphone to collect facial expressions and voice data and sends that data to an analysis algorithm.

[1056] Input: User's facial expressions, voice, input text

[1057] Output: Parsed emotional state (e.g., happy, sad, surprised, angry)

[1058] Step 3:

[1059] Based on the emotion analysis results, the emotion engine selects a taste that matches the user's emotional state. For example, if the emotion of joy is detected, a bright color tone taste will be selected. This information is saved as text and sent to the system.

[1060] Input: Parsed emotional state

[1061] Output: Selected taste information (e.g. bright colors, cartoon style)

[1062] Step 4:

[1063] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it resizes and normalizes the draft image using an image analysis module (e.g., OpenCV), converting the image into a format suitable for processing by the AI ​​model.

[1064] Input: Draft image file, selected taste information

[1065] Output: Preprocessed draft image data, taste information for analysis

[1066] Step 5:

[1067] The server analyzes the character information of the taste specification and maps the results based on the taste selected by the emotion engine, thereby determining the setting value corresponding to the specified taste.

[1068] Input: Selected taste information

[1069] Output: Mapped taste setting value

[1070] Step 6:

[1071] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. Using a deep learning framework (e.g., TensorFlow, PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste.

[1072] Input: Preprocessed draft image data, mapped taste settings

[1073] Output: Generated illustration data

[1074] Step 7:

[1075] The server receives the illustration data returned by the AI ​​model and converts it into a format that can be viewed by the user (e.g., JPEG or PNG), optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[1076] Input: Generated illustration data

[1077] Output: The converted illustration data sent to the user's device

[1078] Step 8:

[1079] Users can check the generated illustrations on their devices, save them or share them on social media as needed, and can even request a re-generation if they are not satisfied with the quality or style of the illustration.

[1080] Input: Illustration data sent to the user's device

[1081] Output: User feedback and regeneration requests

[1082] In this way, by explaining the specific operations and data flow at each processing step, the system's functions can be understood in detail.

[1083] (Application example 2)

[1084] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1085] A problem with modern digital advertising is the lack of technology that can generate personalized visuals that respond to a user's emotions. Traditional advertisements have a static, uniform design and cannot adapt to individual user emotions. This limits the effectiveness of advertisements and makes it difficult to attract users' attention.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1087] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis result of the draft image and the specified taste, emotion analysis means for analyzing a user's emotions, means for the emotion analysis means to select a taste based on the user's emotions, and means for transmitting the generated illustration to a user terminal, thereby enabling the generation of personalized advertising visuals according to the user's emotions.

[1088] "Hand-drawn draft images" refer to early-stage illustrations or blueprints created by users using hand-drawn techniques on paper or digital devices.

[1089] "Text information specifying taste" refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[1090] "Draft image analysis results" refers to information about the image features and structure obtained after processing and analyzing the draft image.

[1091] "Generating an illustration" refers to the digital creation of a final visual work based on input data.

[1092] "User terminal" refers to a digital device, such as a personal computer, smartphone, or tablet, that is used to receive and view system output.

[1093] "Emotion analysis means" refers to a combination of hardware and software for analyzing a user's emotions, specifically using data such as facial expressions, voice, and text input.

[1094] "Selecting a taste based on emotion" refers to a process of automatically selecting the appropriate taste option according to the result of analyzing the user's emotion.

[1095] "Resize and normalize" refers to the process of resizing and standardizing the input draft image to convert it into a format optimal for processing by the AI ​​model.

[1096] "Multiple taste options" refers to the various styles and atmospheres the system offers.

[1097] "Generating advertising visuals" refers to generating visual materials for the purpose of advertising or promotion.

[1098] This invention is applied to a system in which a user uploads a draft image drawn by hand and AI generates advertising visuals based on the user's specified tastes and emotions.

[1099] The server first receives a handwritten draft image from the user's device. A handwritten draft image refers to an early stage illustration or blueprint created by the user using paper or a digital device. Next, the server receives text information specifying the style along with the draft image. The text information specifying the style refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[1100] The user's emotions are then analyzed using an emotion analysis means. The emotion analysis means refers to a combination of hardware and software for analyzing the user's emotions, specifically using data such as facial expressions, voice, and text input. Based on the analysis results, a taste corresponding to the emotion is automatically selected. Selecting a taste based on emotion refers to the process of automatically selecting the appropriate taste option according to the result of the user's emotion analysis.

[1101] Next, the draft image is resized and normalized. Resizing and normalization refers to the process of changing the size and standardizing the input draft image to convert it into a format that is optimal for processing by the AI ​​model. This converts the draft image into a format suitable for processing and prepares it for input into the AI ​​model.

[1102] The server uses a generative AI model to generate advertising visuals based on the analysis results of the draft image and the specified taste. Here, generating advertising visuals refers to generating visual materials for advertising and promotional purposes. The generated advertising visuals are converted into common image formats (e.g., JPEG, PNG) and sent to the user's device. A user's device refers to a digital device, such as a personal computer, smartphone, or tablet, used to receive and view the system's output.

[1103] As a concrete example, consider a case where a user takes a photo of "new shoes" and inputs the emotion "excitement." The emotion engine analyzes this, and the AI ​​model generates bright, dynamic advertising visuals. The generated visuals are sent to the user's device.

[1104] An example of a prompt for the generative AI model is, "Generate a bright and dynamic advertising visual from a product image of 'new shoes' and the emotion 'excitement.'" This system enables the rapid generation of personalized advertising visuals according to the user's emotions.

[1105] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1106] Step 1:

[1107] The user creates a handwritten draft image on their device, then photographs or scans it and uploads it.

[1108] Input: A draft image drawn by the user.

[1109] How it works: Digitizes an image using the device's camera or scanner and sends it to a server through the application

[1110] Output: Digital draft images are sent to the server

[1111] Step 2:

[1112] The user can enter the desired taste information in the text input field, or the taste will be automatically selected based on sentiment analysis.

[1113] Input: User-entered taste information or user sentiment

[1114] Operation: Enter taste information into the device's text input field, or obtain the user's facial expressions and voice to perform emotion analysis.

[1115] Output: Taste information is sent to the server.

[1116] Step 3:

[1117] The server resizes and normalizes the received draft image.

[1118] Input: Draft image received by the server

[1119] How it works: Using an image processing library such as OpenCV, we resize and normalize the image to convert it into a format suitable for AI models.

[1120] Output: Resized and normalized image data

[1121] Step 4:

[1122] Select the appropriate taste based on sentiment analysis data.

[1123] Input: User sentiment analysis data and specified taste information

[1124] How it works: Uses sentiment analysis to automatically select tastes that match the user's emotions.

[1125] Output: Selected taste information

[1126] Step 5:

[1127] The server inputs the draft image and taste information into the generated AI model to generate advertising visuals.

[1128] Input: Resized and normalized draft image, selected taste information

[1129] How it works: Input data into a generative AI model to generate ad visuals

[1130] Output: Generated ad visuals

[1131] Step 6:

[1132] The server converts the generated advertising visuals into a common image format.

[1133] Input: Generated ad visual data

[1134] How it works: Converts to JPEG or PNG format using an image format conversion library.

[1135] Output: Converted advertising visual image

[1136] Step 7:

[1137] The server sends the final advertisement visual to the user's terminal.

[1138] Input: Transformed ad visual image

[1139] Behavior: Sends image data to the user's device using an HTTP request, etc.

[1140] Output: Ad visual image sent to user's device

[1141] Step 8:

[1142] The user checks the received advertisement visuals on the terminal.

[1143] Input: Ad visual image sent to user terminal

[1144] Behavior: Display advertising visuals in the device's image viewer or application.

[1145] Output: Ad visuals for the user to see

[1146] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1147] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1148] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1149] [Fourth embodiment]

[1150] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1151] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1152] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1153] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1154] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1155] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1156] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1157] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1158] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1159] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1160] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1161] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1162] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1163] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[1164] Program processing

[1165] User operations

[1166] Users create a rough draft of the illustration's composition and motif on their device (PC or smartphone).

[1167] Upload the completed draft image to the system's web application or dedicated app by clicking the Upload button and selecting the appropriate image file from the file dialog.

[1168] Next, enter the style of illustration you want in the text input field, such as "soft watercolor colors" or "bright manga colors."

[1169] Once you have completed the draft image and taste specification, press the send button to send the data to the server.

[1170] Server Processing

[1171] The server receives the draft image and character information specifying the taste sent by the user.

[1172] The received draft image is analyzed and resized and normalized, optimizing the image resolution and format for AI processing.

[1173] It also maps the text information of the taste specification to predefined taste options, which determine the style and color tone of the generated illustration.

[1174] Processing AI models

[1175] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input.

[1176] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the specified taste. For example, if the draft is a "composition of a boy reading a book" and the taste is specified as "bright, cartoon-style colors," the AI ​​model will generate a cartoon-style illustration of a boy reading a book.

[1177] Server provides results

[1178] The server receives the illustration data returned by the AI ​​model and reformats it into a format that can be viewed by the user.

[1179] The generated illustration data is sent to the user's terminal, and the results are provided to the user.

[1180] User Verification

[1181] The user checks the illustration sent back from the server on the terminal.

[1182] You can save the generated illustrations as needed, and make further adjustments or requests.

[1183] Specific examples

[1184] For example, if a user uploads a "picture of a cat sleeping on a sofa" as a draft and specifies the taste as "warm, hand-drawn," the process will proceed as follows.

[1185] 1. The user simply draws the shape of a cat and a sofa and uploads the image file to the system.

[1186] 2. The user enters the taste as "Warm hand-drawn" in the text input field and clicks the submit button.

[1187] 3. The server receives the draft image and taste information, resizes and normalizes the image, and maps the taste information.

[1188] 4. The pre-processed data is sent to an AI model to generate a hand-drawn illustration.

[1189] 5. The server receives the generated illustration data and sends it to the user.

[1190] 6. The user checks the generated illustration on their device and requests saving or addition as needed.

[1191] In this way, the system allows users to easily generate high-quality illustrations.

[1192] The processing flow will be explained below.

[1193] Step 1:

[1194] Users use their devices (PC or smartphone) to create a rough draft of the composition and motif of the illustration, which represents a specific scene, such as a cat sleeping on a sofa.

[1195] Step 2:

[1196] The user uploads the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[1197] Step 3:

[1198] After uploading is complete, users can enter the desired style of the illustration in a text input field that appears, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[1199] Step 4:

[1200] After the user inputs the taste specification, the user clicks the send button to send the draft image and taste information to the server.

[1201] Step 5:

[1202] The server receives the draft image and taste specification character information sent by the user, and stores the received data in a format that can be processed immediately.

[1203] Step 6:

[1204] The server analyzes the received draft images and resizes and normalizes them, converting them into a format that is optimally processed by the AI ​​model.

[1205] Step 7:

[1206] The server analyzes the textual information of the taste specification and maps it to predefined taste options, which are then converted into a format that is easy for the AI ​​model to understand.

[1207] Step 8:

[1208] The server passes the preprocessed draft image data and taste information to the AI ​​model as input.

[1209] Step 9:

[1210] The AI ​​model analyzes shapes and objects in the draft image and generates the final illustration based on the user's preferences, such as a cat sleeping on a sofa in a warm, hand-drawn style.

[1211] Step 10:

[1212] The AI ​​model returns the generated illustration data to the server.

[1213] Step 11:

[1214] The server then verifies the illustration data and reformats it into a user-viewable format, such as JPEG or PNG.

[1215] Step 12:

[1216] The server sends the optimized illustration file to the user's terminal.

[1217] Step 13:

[1218] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[1219] Through the above processing steps, the user can easily and efficiently obtain an illustration with the style they desire.

[1220] Example 1

[1221] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1222] While the demand for digital art has increased in recent years, generating high-quality digital illustrations from handwritten sketches requires a high level of specialized knowledge and skill. It is particularly difficult to efficiently generate works that reflect a specific style. To solve this problem, a system that can be intuitively operated by users and automatically generate illustrations in a variety of styles is needed.

[1223] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1224] In this invention, the server includes a means for receiving a handwritten draft image, a means for receiving the draft image along with text information specifying a taste, a means for passing data to a generative AI model that generates an illustration based on the analysis results of the draft image and the specified taste, and a means for transmitting the illustration returned from the generative AI model to a user terminal. This makes it possible to automatically generate high-quality illustrations with the taste specified by the user based on the handwritten draft image.

[1225] A "hand-drawn draft image" refers to draft data of an illustration or composition that a user has hand-drawn using calligraphy implements or a digital device.

[1226] "Text information specifying taste" refers to information entered in text form by the user that specifies the style and color tone of the illustration desired.

[1227] A "generative AI model" is an artificial intelligence model that generates illustrations based on draft images and taste specifications. For example, a model based on deep learning technology falls into this category.

[1228] "Resizing" refers to the process of changing the resolution or size of an original image.

[1229] "Normalization" is a process that standardizes the format of image data, adjusts color tone, standardizes resolution, etc., to prepare the data in a format that is easy for AI models to use.

[1230] A "user terminal" is a digital device such as a computer, smartphone, or tablet that is operated by a user.

[1231] "Illustration" refers to visual artwork created by a generative AI model based on a specified taste or style.

[1232] A "server" is a computer system that receives data from users, processes it, works with an AI model to generate illustrations, and sends the generated results to the user's device.

[1233] The system for implementing this invention allows users to upload a draft image they have drawn by hand, and AI generates an illustration based on the user's specified taste. The system begins when a user creates a draft image and uploads it to the system from a digital device (terminal).

[1234] First, the user uses a device such as a PC or smartphone to create a handwritten draft using digital painting software (general name: digital painting software). The created draft image is then uploaded to the system's web application or dedicated app. To do this, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from the file dialog.

[1235] Next, the user enters the desired style of the illustration in the text input field, for example, "soft watercolor-style colors" or "bright manga-style colors," and once the draft image and style specification are complete, the data is sent to the server by pressing the send button.

[1236] The server receives the draft image and text information specifying the taste sent by the user. The received draft image is resized and normalized to optimize the image resolution and format for AI processing. This preprocessing is performed using the Python Pillow library (general name: image processing library). The text information specifying the taste is also analyzed using a natural language processing library (general name: natural language processing library) and mapped to predefined taste options.

[1237] The server passes the preprocessed draft image data and the mapped taste information as input to a generative AI model. These models utilize deep learning technology (commonly known as generative AI models), such as Stable Diffusion and DALL-E. These AI models analyze shapes and objects from the draft image and generate the final illustration based on the specified taste.

[1238] The illustrations returned by the generative AI model are received by the server and reformatted into a format that can be viewed by the user. This reformatting includes image format conversion and compression. The reformatted illustrations are temporarily stored in cloud storage (commonly known as cloud storage services).

[1239] The user can view the illustration returned from the server on their device, save the generated illustration, or request further adjustments or additions. For example, they can input a prompt such as, "Please generate a warm, hand-drawn illustration based on this draft image." In this way, the system allows users to intuitively operate the system and easily generate high-quality illustrations.

[1240] As described above, this system makes it possible to automatically generate high-quality illustrations in a style specified by the user based on a handwritten draft image.

[1241] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1242] Step 1:

[1243] A user uses a digital device (PC or smartphone) to create a handwritten draft using paint software. The created draft image is then uploaded to the system's web application or a dedicated app. Specifically, the user opens a web browser, accesses the system's web page, clicks the upload button, and selects the appropriate image file from a file dialog.

[1244] Input: A user-created draft image file.

[1245] Output: The uploaded draft image data.

[1246] Step 2:

[1247] The user enters the desired style of the illustration in a text input field on the web page and presses the send button to send the draft image and style specification to the server. Specifically, the user enters styles such as "soft watercolor-style colors" or "bright manga-style colors."

[1248] Input: The text information of the taste specified by the user.

[1249] Output: Draft image data and text information with specified tastes sent to the server.

[1250] Step 3:

[1251] The server receives the draft image and text information specified by the user. The received draft image is resized and normalized using the Python Pillow library. The image resolution and format are optimized for the AI ​​model.

[1252] Input: Draft image data sent by the user.

[1253] Output: Resized and normalized draft image data.

[1254] Step 4:

[1255] The server analyzes the textual information of the taste specification using a natural language processing library (e.g., NLTK or spaCy) and maps it to predefined taste options.

[1256] Input: Text information of taste specification sent by user.

[1257] Output: Mapped taste information.

[1258] Step 5:

[1259] The server passes the resized and normalized draft image data and the mapped taste information as input to a generative AI model (e.g., Stable Diffusion, DALL-E).

[1260] Input: Preprocessed draft image data and mapped taste information.

[1261] Output: Illustration data generated based on the specified taste.

[1262] Step 6:

[1263] The server receives the illustration data returned by the generative AI model and reformats it into a format that can be viewed by users. It uses the Pillow library to convert and compress the image format.

[1264] Input: Illustration data returned from the generative AI model.

[1265] Output: Reformatted illustration data.

[1266] Step 7:

[1267] The server temporarily stores the reformatted illustration data in cloud storage and sends it to the user's device. The user can then check the generated illustration on the web page and save it if desired. Specifically, the user looks at the preview of the illustration and clicks the "Save" button.

[1268] Input: Reformatted illustration data.

[1269] Output: The illustration displayed on the user's device and, if necessary, saved as illustration data.

[1270] This is the specific flow of the program processing for this system, which allows users to automatically generate high-quality illustrations in a specified style based on a handwritten draft image.

[1271] (Application example 1)

[1272] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1273] Conventional illustration generation systems have the ability to generate illustrations by specifying tastes based on handwritten drafts, but lack a means to display the generated illustrations in conjunction with real space. As a result, users cannot check the generated illustrations by overlaying them on real space in real time. The present invention aims to provide a system that overlays illustrations generated from handwritten drafts on the user's visual device, allowing users to enjoy art more intuitively.

[1274] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1275] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for sending the generated illustration to a user terminal, means for overlaying the generated illustration in real space, and means for displaying the illustration on the user's visual device based on taste options specified by the user. This allows the user to check the generated illustration by overlaying it in real space in real time.

[1276] A "hand-drawn draft image" is an image of the composition or motif of an illustration that a user has hand-drawn on paper or a digital device.

[1277] "Text information specifying taste" is information that specifies the style and color tone of the illustration desired by the user in text format.

[1278] "Draft image analysis results" refers to information about shapes and objects obtained by the AI ​​model analyzing handwritten draft images.

[1279] The "specified taste" refers to the style and color tone of the illustration determined based on the character information specifying the taste by the user.

[1280] The "means of generating an illustration" is a function in which the AI ​​model generates the final illustration based on the analysis results of the draft image and the specified taste.

[1281] A "user terminal" is a digital device that receives and displays the generated illustrations, and includes smartphones, PCs, tablets, etc.

[1282] "Means for overlaying and displaying a generated illustration in real space" refers to a function that uses smart glasses or a head-mounted display to display a generated illustration superimposed on real space.

[1283] A "visual device" is a display device worn by a user on the eye, including smart glasses and head-mounted displays.

[1284] "Resizing and normalization" is the process of adjusting the size of the draft image and converting it into a format that is easy for the AI ​​model to use as input.

[1285] "Multiple taste options" are multiple style and color options that the user can choose from, such as "vintage style" and "modern art style."

[1286] The system for implementing this invention allows users to upload a draft image drawn by hand and generates an illustration that is overlaid on real space based on the user's specified taste. This system can overlay the illustration on real space in real time using visual devices such as smart glasses or a head-mounted display.

[1287] System configuration

[1288] The system consists of the following main components:

[1289] 1. User Device:

[1290] The user terminal includes smart glasses and a head-mounted display, and has the functions of taking handwritten images, specifying tastes, and uploading data.

[1291] The user takes a photo of their handwritten draft using the camera on their smart glasses or head-mounted display and saves it as an image file on their device.

[1292] 2. Server:

[1293] The server receives the draft image and character information specifying the taste sent from the user terminal.

[1294] The server resizes and normalizes the images, converting them into a format suitable for the AI ​​model.

[1295] The character information of the taste specification is analyzed and mapped to predefined taste options.

[1296] The server passes the preprocessed draft image and taste information as input to the AI ​​model, and generates an illustration based on the specified taste.

[1297] The generated illustration data is sent to the user's visual device to realize an overlay display.

[1298] Hardware and software used

[1299] Hardware:

[1300] Smart glasses (e.g. Google Glass)

[1301] Head-mounted displays (e.g. Microsoft HoloLens)

[1302] Digital camera devices (cameras built into smart glasses or HoloLens)

[1303] software:

[1304] On-device applications (e.g., developed with Unity)

[1305] Server-side image processing and resizing / normalization functions (e.g. OpenCV)

[1306] AI models (e.g., deep learning models using TensorFlow or PyTorch)

[1307] A description of what the program does

[1308] The server receives the user's handwritten draft image and the specified tastes. It first resizes and normalizes the image. This process uses the image processing library OpenCV. Next, the text information specified by the taste is mapped to predefined taste options and input into an AI model (a model using TensorFlow or PyTorch). The AI ​​model analyzes the shapes and objects in the draft image and generates the final illustration based on the specified tastes. The server then sends the generated illustration to the user's visual device, where the user can view the illustration overlaid on the real world in real time through smart glasses or a head-mounted display.

[1309] Specific examples

[1310] For example, if a user attending a workshop at an art supply store draws a handwritten sketch of a "flowering park scene," photographs it with smart glasses, and specifies the style as "vintage," the following prompt sentence will be entered:

[1311] Example prompt sentence:

[1312] Input image: A hand-drawn sketch of a park landscape with flowers in bloom

[1313] Style: Vintage

[1314] The AI ​​model generates a vintage-style park landscape illustration based on the user's specifications. This illustration is displayed on the user's smart glasses, allowing them to see the real world and digital art blended together in their own field of vision. This system allows users to intuitively enjoy their own original art.

[1315] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1316] Step 1:

[1317] The user uses a handheld visual device (smart glasses or a head-mounted display) to take a picture of a handwritten draft and save it in the device. At this time, the camera of the visual device is activated and an operation to capture the draft image is performed. The input is the handwritten draft image, and the output is the captured image file.

[1318] Step 2:

[1319] The user uploads a captured draft image to the server through an application on the visual device. The application inputs text information for selecting an image file and specifying tastes through a user interface. The input is the captured draft image and the specified taste information, and the output is data transmission to the server.

[1320] Step 3:

[1321] The server receives the draft image and text information of taste specification sent by the user. The received data is prepared to be passed to the analysis process within the server. The input is the draft image and taste specification information, and the output is the readiness status for data analysis.

[1322] Step 4:

[1323] The server resizes and normalizes the received draft images. Specifically, it uses an image processing library such as OpenCV to adjust the image size and normalize the pixel values. The input is the original draft image, and the output is the resized and normalized image data.

[1324] Step 5:

[1325] The server maps the textual information of the taste specification to predefined taste options. This process uses a string matching algorithm to match the specified taste to an internal taste option list. The input is the textual information of the taste specification, and the output is the mapped taste information.

[1326] Step 6:

[1327] The server passes the preprocessed draft image data and the mapped taste information to the AI ​​model as input. The AI ​​model (using TensorFlow or PyTorch) analyzes the draft image and generates the final illustration based on the specified taste. The input is the preprocessed draft image and the mapped taste information, and the output is the generated illustration.

[1328] Step 7:

[1329] The server receives the generated illustration data and sends it to the user's visual device, which then processes it to overlay the illustration in real time. The input is the generated illustration data, and the output is the data transmission and display instructions to the visual device.

[1330] Step 8:

[1331] The user can view the illustrations overlaid on the real world in real time through a visual device, and can save or request additional illustrations as needed. The input is the illustration displayed on the visual device, and the output is the user's visual perception and operational instructions.

[1332] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1333] This invention combines a system in which a user uploads a hand-drawn draft image and an AI generates an illustration based on the user's specified taste with an emotion engine that recognizes the user's emotions, allowing the system to generate an illustration with a taste that matches the user's emotions.

[1334] Program processing

[1335] User operations

[1336] First, the user creates a rough draft of the illustration's composition and motif on a device at hand (PC or smartphone).

[1337] Next, upload the draft image to the system's web application or dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[1338] Users can enter the style of the illustration they want in a text input field, but the emotion engine can also analyze the user's emotions and automatically select the appropriate style.

[1339] Manipulating the Emotion Engine

[1340] The emotion engine analyzes emotions from facial expressions, voice, input text, etc. acquired from the user's device. Examples of emotions include joy, sadness, surprise, and anger.

[1341] Based on the analysis results, the emotion engine determines the user's current emotional state and selects a taste that suits it. For example, if the user is feeling happy, a bright color tone taste will be selected.

[1342] Server Processing

[1343] The server receives the draft image sent by the user and the taste information selected by the emotion engine.

[1344] The received draft image is analyzed, resized, and normalized, converting it into a format optimized for AI processing.

[1345] The text information of the taste specification is analyzed and mapped to predefined taste options. The taste selected by the emotion engine is applied with priority.

[1346] Processing AI models

[1347] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the AI ​​model.

[1348] The AI ​​model analyzes shapes and objects from the draft image and generates the final illustration based on the specified taste. For example, if a user sketches a "cat sleeping on a sofa" and the emotion engine selects a "warm, hand-drawn" taste, the AI ​​model will generate an illustration according to that instruction.

[1349] Server provides results

[1350] The server receives the illustration data returned by the AI ​​model and reformats it into a user-viewable format, such as a common image format like JPEG or PNG.

[1351] The optimized illustration file is sent to the user's device.

[1352] User Verification

[1353] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[1354] Specific examples

[1355] For example, if a user uploads a draft of a "boy reading a book" and selects the emotion joy, the emotion engine will analyze it and automatically select a bright, cartoon-style image. Based on this information, the server passes the data to an AI model, which then generates a final "bright, cartoon-style" illustration of a "boy reading a book." The generated illustration is then sent to the user's device, where they can view it.

[1356] In this way, by combining emotion engines, a system configuration is provided that can quickly generate illustrations with an appropriate style according to the user's emotions.

[1357] The processing flow will be explained below.

[1358] Step 1:

[1359] Users use their device (PC or smartphone) to create a rough draft of the illustration's composition and motif, which can then be freely drawn on paper or using digital tools.

[1360] Step 2:

[1361] Users can upload the created draft image to the system's web application or a dedicated app by clicking the upload button and selecting the draft image from the file dialog.

[1362] Step 3:

[1363] When users upload a draft image, they can enter the style of their illustration they want in a text input field or choose an auto-select option, such as "warm, hand-drawn" or "futuristic, cyberpunk-inspired."

[1364] Step 4:

[1365] The device temporarily stores the draft image uploaded by the user and the taste information entered, and if the user selects the automatic selection option, it starts the emotion engine.

[1366] Step 5:

[1367] The emotion engine analyzes facial expression video and audio data acquired from the user's device, as well as input text data, to determine the user's emotions. For example, it uses facial recognition technology to recognize smiling and surprised expressions, and performs text analysis.

[1368] Step 6:

[1369] The emotion engine categorizes the user's emotional state based on the analysis results and selects the appropriate taste option. For example, if the user is expressing joy, a bright and positive taste will be selected.

[1370] Step 7:

[1371] The terminal sends the taste information selected by the emotion engine and the draft image to the server. Similarly, if the user manually inputs tastes, the terminal also sends the draft image and taste information to the server.

[1372] Step 8:

[1373] The server receives the draft image and taste specification character information sent by the user and stores the received data in a format that can be processed internally.

[1374] Step 9:

[1375] The server analyzes the received draft images and performs resizing and normalization processes, converting the draft images into a format optimal for AI processing.

[1376] Step 10:

[1377] The server analyzes the character information of the taste specification and maps it to predefined taste options. The taste information selected by the emotion engine is used preferentially.

[1378] Step 11:

[1379] The server passes the preprocessed draft image data and selected taste information to the AI ​​model as input.

[1380] Step 12:

[1381] The AI ​​model generates illustrations in a specified style based on the input draft image and taste information. For example, it generates an illustration by applying a "warm, hand-drawn" taste to a draft of a "cat sleeping on a sofa."

[1382] Step 13:

[1383] The illustration data generated by the AI ​​model is sent back to the server.

[1384] Step 14:

[1385] The server receives the generated illustration data and reformats it into a format that can be viewed by the user, for example, converting it into JPEG or PNG format.

[1386] Step 15:

[1387] The server sends the optimized illustration file to the user's terminal.

[1388] Step 16:

[1389] The user can check the generated illustration on their device and, if necessary, save it or make further adjustments or requests for additions.

[1390] Through the above processing steps, users can efficiently obtain illustrations with the style they desire. Furthermore, by using the emotion engine, illustrations that match the user's emotional state can be automatically generated.

[1391] Example 2

[1392] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1393] Conventional illustration generation systems could generate illustrations based on tastes specified by a user's hand-drawn draft image, but it was difficult to automatically select an appropriate taste that took the user's emotional state into consideration, which resulted in the problem of users being unable to quickly obtain an illustration that matched their emotions.

[1394] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1395] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis results of the draft image and the specified taste, means for analyzing emotions from the user's facial expression, voice, and input text, means for selecting a taste based on the emotion analysis results, and means for transmitting the generated illustration to the user terminal. This makes it possible to quickly generate an illustration with an appropriate taste that takes into account the user's emotional state.

[1396] A "draft image" is an image of an illustration that has been hand-drawn by a user in the initial stage.

[1397] "Taste" is setting information that indicates the style, color tone, and design direction of an illustration.

[1398] The "emotion engine" is a system that analyzes the user's facial expressions, voice, input text, etc. to determine the user's emotional state, and selects an appropriate taste based on the results.

[1399] The "server" is a computer system that processes the data received from the user, uses an AI model to generate the final illustration, and returns it to the user.

[1400] An "AI model" is an algorithm that uses machine learning techniques such as deep learning to take a rough sketch image and taste information as input and generate an illustration in a specified style.

[1401] "Resizing" is the process of changing the size of a draft image and adjusting it to the optimal resolution for processing by the AI ​​model.

[1402] "Normalization" is the process of converting image data into a uniform format to make analysis by AI models more efficient.

[1403] "Emotion analysis" is a technology that determines a user's emotional state based on data such as the user's facial expressions, voice, and input text.

[1404] A "generated illustration" is a final digital image generated by an AI model based on a draft image and a specified style.

[1405] A "user terminal" is a device used by a user, such as a PC or smartphone, on which the generated illustration is displayed.

[1406] This invention is a system in which a user uploads a hand-drawn draft image and a generative AI model generates an illustration based on the user's specified taste, and further recognizes the user's emotion and generates an illustration with a taste corresponding to that emotion. This system is composed of a server, a terminal, and an emotion engine.

[1407] Program Overview

[1408] User operations

[1409] First, the user creates a rough sketch of the illustration's composition and motif using a device (PC or smartphone). For example, they can use a drawing app on their smartphone. After creating the sketch, the user uploads the rough image to the system's web application or a dedicated app. During the upload process, the user clicks a dedicated upload button and selects the image file from a file dialog.

[1410] Manipulating the Emotion Engine

[1411] The emotion engine acquires facial expressions, voice, input text, etc. from the user's device. It uses the device's camera and microphone to collect the user's facial expressions and tone of voice, and analyzes the acquired data in real time. This allows it to determine the user's emotional state (happiness, sadness, surprise, anger, etc.) and select an appropriate illustration style accordingly.

[1412] Server Processing

[1413] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it uses an image analysis module (e.g., OpenCV) to resize and normalize the draft image. This converts the image into a format optimal for AI model processing. It also analyzes the text information of the taste specification and maps the results based on the taste selected by the emotion engine.

[1414] Processing AI models

[1415] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. For example, using a deep learning framework (TensorFlow or PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste. The AI ​​model recognizes shapes and objects from the image and draws them in a style that matches the taste.

[1416] Server provides results

[1417] The server receives the illustration data returned by the AI ​​model and converts it into a format that the user can view, such as JPEG or PNG, optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[1418] User review and feedback of results

[1419] Users can view the generated illustrations on their devices, save them, or share them on social media. If necessary, they can request regeneration or make additional adjustments.

[1420] Examples of concrete examples and prompts

[1421] For example, if a user uploads a draft of a "boy reading a book" and expresses joy as an emotion, the emotion engine analyzes it and selects a bright color tone. Based on this information, the server passes the data to an AI model, which ultimately generates an illustration of a "boy reading a book in bright colors." The generated illustration is sent to the user's device, where the user can view it.

[1422] Prompt Sentence Examples

[1423] Prompt: "If a user uploads a draft composition of dogs playing in a park and experiences happiness as an emotion, the sentiment analysis engine should select a bright, vibrant, cartoon-like design."

[1424] Prompt: "If the composition shows a person relaxing by the sea, the emotion analysis engine should detect the relaxed emotion and generate a calm, soft-colored, hand-drawn illustration."

[1425] In this way, the system uses user input and emotion-based prompts to provide instructions to the AI ​​model for generating the best illustrations.

[1426] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1427] Step 1:

[1428] The user creates a draft image using a drawing app on their device (PC or smartphone). The created draft image is then uploaded to the system using a dedicated application. Specifically, the user clicks a dedicated upload button and selects the desired image file from a file dialog.

[1429] Input: User-created draft image file

[1430] Output: Draft image file sent to the server

[1431] Step 2:

[1432] The emotion engine captures facial expressions, voice, and input text from the user's device in real time and analyzes the data to determine the user's emotional state. For example, it uses the device's camera and microphone to collect facial expressions and voice data and sends that data to an analysis algorithm.

[1433] Input: User's facial expressions, voice, input text

[1434] Output: Parsed emotional state (e.g., happy, sad, surprised, angry)

[1435] Step 3:

[1436] Based on the emotion analysis results, the emotion engine selects a taste that matches the user's emotional state. For example, if the emotion of joy is detected, a bright color tone taste will be selected. This information is saved as text and sent to the system.

[1437] Input: Parsed emotional state

[1438] Output: Selected taste information (e.g. bright colors, cartoon style)

[1439] Step 4:

[1440] The server receives the draft image sent by the user and the taste information selected by the emotion engine. After receiving it, it resizes and normalizes the draft image using an image analysis module (e.g., OpenCV), converting the image into a format suitable for processing by the AI ​​model.

[1441] Input: Draft image file, selected taste information

[1442] Output: Preprocessed draft image data, taste information for analysis

[1443] Step 5:

[1444] The server analyzes the character information of the taste specification and maps the results based on the taste selected by the emotion engine, thereby determining the setting value corresponding to the specified taste.

[1445] Input: Selected taste information

[1446] Output: Mapped taste setting value

[1447] Step 6:

[1448] The server passes the preprocessed draft image data and the taste information selected by the emotion engine as input to the generative AI model. Using a deep learning framework (e.g., TensorFlow, PyTorch), the AI ​​model analyzes the draft image and generates the final illustration based on the specified taste.

[1449] Input: Preprocessed draft image data, mapped taste settings

[1450] Output: Generated illustration data

[1451] Step 7:

[1452] The server receives the illustration data returned by the AI ​​model and converts it into a format that can be viewed by the user (e.g., JPEG or PNG), optimizing image quality and file size. The converted illustration file is then sent to the user's device.

[1453] Input: Generated illustration data

[1454] Output: The converted illustration data sent to the user's device

[1455] Step 8:

[1456] Users can check the generated illustrations on their devices, save them or share them on social media as needed, and can even request a re-generation if they are not satisfied with the quality or style of the illustration.

[1457] Input: Illustration data sent to the user's device

[1458] Output: User feedback and regeneration requests

[1459] In this way, by explaining the specific operations and data flow at each processing step, the system's functions can be understood in detail.

[1460] (Application example 2)

[1461] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1462] A problem with modern digital advertising is the lack of technology that can generate personalized visuals that respond to a user's emotions. Traditional advertisements have a static, uniform design and cannot adapt to individual user emotions. This limits the effectiveness of advertisements and makes it difficult to attract users' attention.

[1463] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1464] In this invention, the server includes means for receiving a handwritten draft image, means for receiving text information specifying a taste along with the draft image, means for generating an illustration based on the analysis result of the draft image and the specified taste, emotion analysis means for analyzing a user's emotions, means for the emotion analysis means to select a taste based on the user's emotions, and means for transmitting the generated illustration to a user terminal, thereby enabling the generation of personalized advertising visuals according to the user's emotions.

[1465] "Hand-drawn draft images" refer to early-stage illustrations or blueprints created by users using hand-drawn techniques on paper or digital devices.

[1466] "Text information specifying taste" refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[1467] "Draft image analysis results" refers to information about the image features and structure obtained after processing and analyzing the draft image.

[1468] "Generating an illustration" refers to the digital creation of a final visual work based on input data.

[1469] "User terminal" refers to a digital device, such as a personal computer, smartphone, or tablet, that is used to receive and view system output.

[1470] "Emotion analysis means" refers to a combination of hardware and software for analyzing a user's emotions, specifically using data such as facial expressions, voice, and text input.

[1471] "Selecting a taste based on emotion" refers to a process of automatically selecting the appropriate taste option according to the result of analyzing the user's emotion.

[1472] "Resize and normalize" refers to the process of resizing and standardizing the input draft image to convert it into a format optimal for processing by the AI ​​model.

[1473] "Multiple taste options" refers to the various styles and atmospheres the system offers.

[1474] "Generating advertising visuals" refers to generating visual materials for the purpose of advertising or promotion.

[1475] This invention is applied to a system in which a user uploads a draft image drawn by hand and AI generates advertising visuals based on the user's specified tastes and emotions.

[1476] The server first receives a handwritten draft image from the user's device. A handwritten draft image refers to an early stage illustration or blueprint created by the user using paper or a digital device. Next, the server receives text information specifying the style along with the draft image. The text information specifying the style refers to information that specifically expresses in text the style and atmosphere of the illustration desired by the user.

[1477] The user's emotions are then analyzed using an emotion analysis means. The emotion analysis means refers to a combination of hardware and software for analyzing the user's emotions, specifically using data such as facial expressions, voice, and text input. Based on the analysis results, a taste corresponding to the emotion is automatically selected. Selecting a taste based on emotion refers to the process of automatically selecting the appropriate taste option according to the result of the user's emotion analysis.

[1478] Next, the draft image is resized and normalized. Resizing and normalization refers to the process of changing the size and standardizing the input draft image to convert it into a format that is optimal for processing by the AI ​​model. This converts the draft image into a format suitable for processing and prepares it for input into the AI ​​model.

[1479] The server uses a generative AI model to generate advertising visuals based on the analysis results of the draft image and the specified taste. Here, generating advertising visuals refers to generating visual materials for advertising and promotional purposes. The generated advertising visuals are converted into common image formats (e.g., JPEG, PNG) and sent to the user's device. A user's device refers to a digital device, such as a personal computer, smartphone, or tablet, used to receive and view the system's output.

[1480] As a concrete example, consider a case where a user takes a photo of "new shoes" and inputs the emotion "excitement." The emotion engine analyzes this, and the AI ​​model generates bright, dynamic advertising visuals. The generated visuals are sent to the user's device.

[1481] An example of a prompt for the generative AI model is, "Generate a bright and dynamic advertising visual from a product image of 'new shoes' and the emotion 'excitement.'" This system enables the rapid generation of personalized advertising visuals according to the user's emotions.

[1482] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1483] Step 1:

[1484] The user creates a handwritten draft image on their device, then photographs or scans it and uploads it.

[1485] Input: A draft image drawn by the user.

[1486] How it works: Digitizes an image using the device's camera or scanner and sends it to a server through the application

[1487] Output: Digital draft images are sent to the server

[1488] Step 2:

[1489] The user can enter the desired taste information in the text input field, or the taste will be automatically selected based on sentiment analysis.

[1490] Input: User-entered taste information or user sentiment

[1491] Operation: Enter taste information into the device's text input field, or obtain the user's facial expressions and voice to perform emotion analysis.

[1492] Output: Taste information is sent to the server.

[1493] Step 3:

[1494] The server resizes and normalizes the received draft image.

[1495] Input: Draft image received by the server

[1496] How it works: Using an image processing library such as OpenCV, we resize and normalize the image to convert it into a format suitable for AI models.

[1497] Output: Resized and normalized image data

[1498] Step 4:

[1499] Select the appropriate taste based on sentiment analysis data.

[1500] Input: User sentiment analysis data and specified taste information

[1501] How it works: Uses sentiment analysis to automatically select tastes that match the user's emotions.

[1502] Output: Selected taste information

[1503] Step 5:

[1504] The server inputs the draft image and taste information into the generated AI model to generate advertising visuals.

[1505] Input: Resized and normalized draft image, selected taste information

[1506] How it works: Input data into a generative AI model to generate ad visuals

[1507] Output: Generated ad visuals

[1508] Step 6:

[1509] The server converts the generated advertising visuals into a common image format.

[1510] Input: Generated ad visual data

[1511] How it works: Converts to JPEG or PNG format using an image format conversion library.

[1512] Output: Converted advertising visual image

[1513] Step 7:

[1514] The server sends the final advertisement visual to the user's terminal.

[1515] Input: Transformed ad visual image

[1516] Behavior: Sends image data to the user's device using an HTTP request, etc.

[1517] Output: Ad visual image sent to user's device

[1518] Step 8:

[1519] The user checks the received advertisement visuals on the terminal.

[1520] Input: Ad visual image sent to user terminal

[1521] Behavior: Display advertising visuals in the device's image viewer or application.

[1522] Output: Ad visuals for the user to see

[1523] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1524] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1525] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1526] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1527] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1528] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1529] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1530] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1531] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1532] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1533] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1534] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1535] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1536] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1537] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1538] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1539] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1540] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1541] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1542] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1543] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1544] The following is further disclosed regarding the above embodiment.

[1545] (Claim 1)

[1546] a means for receiving a handwritten draft image;

[1547] A means for receiving text information specifying a taste along with a draft image;

[1548] [Generate an illustration based on the analysis results of the draft image and the specified taste]

[1549] A means for sending the generated illustration to a user terminal;

[1550] A system including:

[1551] (Claim 2)

[1552] The system of claim 1, wherein the system resizes and normalizes the draft image.

[1553] (Claim 3)

[1554] The system of claim 1 [generates illustrations based on a taste selected from a plurality of taste options].

[1555] "Example 1"

[1556] (Claim 1)

[1557] a means for receiving a handwritten draft image;

[1558] A means for receiving text information specifying a taste along with a draft image;

[1559] [Transferring data to a generative AI model that generates illustrations based on the analysis results of the draft image and the specified taste],

[1560] [Transmitting the illustration returned from the generative AI model to the user's device];

[1561] A system including:

[1562] (Claim 2)

[1563] The system of claim 1, wherein the system resizes and normalizes the draft image.

[1564] (Claim 3)

[1565] The system of claim 1 [including a generative AI model that generates illustrations based on a taste selected from a plurality of taste options].

[1566] "Application Example 1"

[1567] (Claim 1)

[1568] a means for receiving a handwritten draft image;

[1569] A means for receiving text information specifying a taste along with a draft image;

[1570] [Generate an illustration based on the analysis results of the draft image and the specified taste]

[1571] A means for sending the generated illustration to a user terminal;

[1572] A means to overlay the generated illustrations in real space,

[1573] a means for displaying an illustration on a user's visual device based on user-specified taste options;

[1574] A system including:

[1575] (Claim 2)

[1576] The system of claim 1, wherein the system resizes and normalizes the draft image.

[1577] (Claim 3)

[1578] The system of claim 1 [generates illustrations based on a taste selected from a plurality of taste options].

[1579] "Example 2: Combining Emotion Engines"

[1580] (Claim 1)

[1581] a means for receiving a handwritten draft image;

[1582] A means for receiving text information specifying a taste along with a draft image;

[1583] [Generate an illustration based on the analysis results of the draft image and the specified taste]

[1584] A means for sending the generated illustration to a user terminal;

[1585] A means to analyze emotions from the user's facial expressions, voice, and input text,

[1586] A means for selecting tastes based on the results of sentiment analysis;

[1587] A system including:

[1588] (Claim 2)

[1589] The system of claim 1, wherein the system resizes and normalizes the draft image.

[1590] (Claim 3)

[1591] The system of claim 1 [generates illustrations based on a taste selected from a plurality of taste options].

[1592] "Application example 2 when combining emotion engines"

[1593] (Claim 1)

[1594] a means for receiving a handwritten draft image;

[1595] A means for receiving text information specifying a taste along with a draft image;

[1596] [Generate an illustration based on the analysis results of the draft image and the specified taste]

[1597] [equipped with emotion analysis means for analyzing user emotions, the emotion analysis means selecting tastes based on the user emotions] means;

[1598] A means for sending the generated illustration to a user terminal;

[1599] A system including:

[1600] (Claim 2)

[1601] The system of claim 1, wherein the system resizes and normalizes the draft image.

[1602] (Claim 3)

[1603] The system of claim 1 [generates advertising visuals based on a taste selected from a plurality of taste options]. [Explanation of symbols]

[1604] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving a handwritten draft image; A means for receiving text information specifying a taste together with a draft image; A means for generating an illustration based on the analysis result of the draft image and a specified taste; means for transmitting the generated illustration to a user terminal; A system including:

2. The system of claim 1 , wherein the draft image is resized and normalized.

3. The system of claim 1 , wherein the system generates an illustration based on a taste selected from a plurality of taste options.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A